Methods for labeling and analyzing single-cell nucleic acids

JP2024546177A5Pending Publication Date: 2025-12-09SHENZHEN HUADA GENE INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024538042
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-24
Filing Date
2022-11-30
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing single-cell transcriptome sequencing technologies, such as 10x Chromium, face limitations in throughput and cell capture rate, particularly when sequencing large numbers of cells or rare cell types, due to Poisson distribution and microdroplet formation constraints.

Method used

A method involving a nucleic acid array with microdots on a solid support, where cells are positioned on microdots, and nucleic acid molecules are labeled with tag sequences using oligonucleotide probes, enabling high-throughput single cell transcriptome sequencing by reverse transcription and ligation of positioning tags.

Benefits of technology

Enhances cell capture rates and sequencing throughput beyond existing limits, allowing for efficient sequencing of large numbers of cells or rare cell types without loss of transcriptome information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided is a method for positioning and labeling nucleic acid molecules, and a method for constructing a nucleic acid molecule library for single-cell transcriptome sequencing, which relates to the technical field of single-cell transcriptome sequencing and biomolecule spatial information detection.Further provided is a nucleic acid molecule library constructed by using the method, and a kit for carrying out the method.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present application relates to the technical field of single-cell transcriptome sequencing and biomolecular spatial information detection. In particular, the present application relates to a method for positionally labeling nucleic acid molecules in individual cells and a method for constructing a single-cell transcriptome sequencing library. In addition, the present application relates to a kit for carrying out the method. [Background technology]

[0002] Single-cell transcriptome sequencing technology is an important tool for identifying cellular heterogeneity. The importance of single-cell transcriptome sequencing technology has promoted the rapid development of this technology in terms of throughput, easy operation, etc. The development of single-cell transcriptome sequencing technology has encouraged the international community to spend a large amount of funds to undertake the Human Cell Atlas Project to create a three-dimensional atlas of human cells. The undertaking of the Human Cell Atlas Project has presented higher requirements and challenges for the throughput of single-cell transcriptome sequencing technology. In addition to scientific research needs, single-cell transcriptome sequencing technology is also used by medical workers to discover the small number of "cancer stem cells" in cancer to find targeted drugs and therapies to overcome malignant tumors. Since malignant tumor cells are relatively rare, single-cell transcriptome sequencing technology requires high cell utilization and capture rate to avoid losing the transcriptome information of malignant tumor cells with low abundance.

[0003] Existing single-cell transcriptome sequencing technologies mainly include two categories: one is low-throughput sequencing technologies based on multi-well plates, in which single cells are assigned to individual wells of a multi-well plate, such as smart-seq and CEL-seq; the other is magnetic bead-based sequencing technologies, such as 10x Chromium, Drop-seq, Seq-well and other technologies, in which cells and labeled magnetic beads are simultaneously entrapped in microdroplets or microwells by microfluidics. Among single-cell transcriptome sequencing technologies, 10x Chromium has the highest throughput, and its throughput in a single run is 5,000-7,000 cells, and can reach up to 10,000 cells. Moreover, its cell capture rate is 30%-60%, depending on the cell type.

[0004] Taking 10x Chromium, the most widely used in the market, as an example, its technical feature is the use of a microfluidic system for cell sorting. In brief, gel beads with labeling molecules or barcode molecules (Barcode) are allowed to enter the microfluidic system at a uniform speed, and the cells and enzymes to be sorted are allowed to bind to the gel beads at a specific time interval, forming a GEM (gel beads in emulsion) in the oil phase. Ideally, each cell binds to one gel bead to form one GEM. Therefore, this method can achieve the purpose of single-cell transcriptome sequencing. However, the formation of GEM follows a Poisson distribution. That is, a single GEM may contain 0 or more cells. The sequencing data created by this GEM does not correspond to the state of a single cell, so it cannot be used later and needs to be filtered by an algorithm. The throughput of this technology is limited by the number of microdroplets formed in the oil phase, making it difficult to exceed the level of 10,000, and at the same time, due to the characteristics of the Poisson distribution, the maximum cell capture rate of this technology can reach up to 60%. Therefore, when single-cell transcriptome sequencing with a throughput of 100,000 or even more cells is required, or rare cells need to be captured and sequenced, this technology still has great shortcomings and is difficult to meet practical needs.Therefore, there is a need in this technical field to develop a new single-cell transcriptome sequencing method with higher cell capture rate. [Brief description of the drawings]

[0005] [Figure 1]FIG. 1A shows an exemplary structure of a chip used in the present application for capturing and labeling nucleic acid molecules, including a chip and oligonucleotide probes (also called chip sequences) bound to the chip. Each oligonucleotide probe contains a tag sequence Y corresponding to its location on the chip, and the area of ​​the chip to which one type of oligonucleotide probe is bound can be called a microdot. Each oligonucleotide probe can contain a single copy or multiple copies. FIG. 1B shows that after cells in a sample contact the chip, they are labeled by one or more microdots on the chip. [Diagram 2] 2 shows an exemplary scheme 1 for preparing a cDNA strand using a sample RNA (eg, mRNA) as a template and an exemplary structure of the cDNA strand. CA: consensus sequence A, CB: consensus sequence B. [Diagram 3] 3 shows an exemplary scheme for using a chip sequence to label the 5' end of a cDNA strand (i.e., ligating the 5' end of the cDNA strand to the 3' end of the chip sequence) to form a novel nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence), and an exemplary structure of the novel nucleic acid molecule containing chip sequence information: CA: consensus sequence A, CB: consensus sequence B, X1: consensus sequence X1, Y: tag sequence Y, X2: consensus sequence X2. [Figure 4] 4 shows an exemplary scheme 1 for preparing a complementary strand of a cDNA strand using a sample RNA (e.g., mRNA) as a template and an exemplary structure of the complementary strand of a cDNA strand: CA: consensus sequence A, CB: consensus sequence B, EP: extension primer. [Diagram 5]5 shows an exemplary scheme for forming a novel nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) by using a chip sequence to label the 5' end of a complementary strand of a cDNA strand (i.e., ligating the 5' end of the complementary strand of a cDNA strand to the 3' end of the chip sequence), and an exemplary structure of the novel nucleic acid molecule containing chip sequence information: CA: consensus sequence A, CB: consensus sequence B, X1: consensus sequence X1, Y: tag sequence Y, X2: consensus sequence X2. [Figure 6] 6 shows an exemplary scheme 2 for preparing a cDNA strand using a sample RNA (eg, mRNA) as a template and an exemplary structure of the cDNA strand. CA: consensus sequence A, CB: consensus sequence B. [Figure 7] 7 shows an exemplary scheme 1 for forming a novel nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) using a complementary sequence of a chip sequence to label the 3' end of a cDNA strand, and an exemplary structure of the novel nucleic acid molecule containing chip sequence information: CA: consensus sequence A, CB: consensus sequence B, X1: consensus sequence X1, Y: tag sequence Y, X2: consensus sequence X2, P1: first region, P2: second region. [Figure 8] 8 shows an exemplary scheme 2 for forming a novel nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) using a complementary sequence of a chip sequence to label the 3' end of a cDNA strand, and an exemplary structure of the novel nucleic acid molecule containing chip sequence information. CA: consensus sequence A, CB: consensus sequence B, X1: consensus sequence X1, Y: tag sequence Y, X2: consensus sequence X2. [Figure 9] 9 shows an exemplary scheme 2 for preparing a complementary strand of a cDNA strand using a sample RNA (e.g., mRNA) as a template and an exemplary structure of the complementary strand of a cDNA strand: CA: consensus sequence A, CB: consensus sequence B, EP: extension primer. [Figure 10]10 shows an exemplary scheme 1 for forming a novel nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) using a complementary sequence of a chip sequence to label the 3' end of a complementary strand of a cDNA strand, and an exemplary structure of the novel nucleic acid molecule containing chip sequence information. CA: consensus sequence A, CB: consensus sequence B, X1: consensus sequence X1, Y: tag sequence Y, X2: consensus sequence X2, P1: first region, P2: second region. [Figure 11] 11 shows an exemplary scheme 2 for forming a novel nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) using a complementary sequence of a chip sequence to label the 3' end of a complementary strand of a cDNA strand, and an exemplary structure of the novel nucleic acid molecule containing chip sequence information. CA: consensus sequence A, CB: consensus sequence B, X1: consensus sequence X1, Y: tag sequence Y, X2: consensus sequence X2. [Figure 12] FIG. 12 shows gene expression profiles of a subset of Hek293 cells obtained by the method of Example 1. [Figure 13] FIG. 13 shows an enlarged partial view of the gene expression profile of a subset of Hek293 cells obtained by the method of Example 1. [Figure 14] FIG. 14 shows the length distribution of the cDNA amplification products in Example 2. [Figure 15] FIG. 15 shows the gene expression profile of Hek293 cells obtained by the method of Example 2. [Figure 16] FIG. 16 shows the average gene number and UMI as captured from single cells by the method of Example 2. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0006] The present application provides a method for positionally labeling nucleic acid molecules in a cell sample and a method for constructing a single-cell transcriptome sequencing library based on this method. Furthermore, the present application also relates to a kit for carrying out this method.

[0007] Methods for generating a population of labeled nucleic acid molecules

[0008] In one aspect, the present application provides a method for generating a population of labeled nucleic acid molecules, comprising: (1) providing a sample comprising one or more cells and a nucleic acid array, the sample is a single cell suspension, the cells comprising (e.g., on their surface) a first binding molecule; The nucleic acid array comprises a solid support, the solid support comprising (e.g., on its surface) a first labeled molecule, and a first binding molecule capable of forming an interactive pair with the first labeled molecule; Additionally, the solid support further comprises a plurality of microdots, the size (e.g., equivalent diameter) of the microdots being less than 5 μm and the center-to-center distance between adjacent microdots being less than 10 μm, each microdot being bound to one type of oligonucleotide probe, each type of oligonucleotide probe comprising at least one copy, the oligonucleotide probe comprising or consisting of, in the 5' to 3' direction, a consensus sequence X1, a tag sequence Y and a consensus sequence X2; the oligonucleotide probes bound to different microdots have different tag sequences Y; (2) contacting one or more cells with the solid support of the nucleic acid array, whereby each cell occupies at least one microdot of the nucleic acid array (i.e., each cell contacts at least one microdot of the nucleic acid array) and allows a first binding molecule of the cell to interact with a first label molecule of the solid support; performing a pre-treatment, before or after contacting the one or more cells with the nucleic acid array, comprising reverse transcription of RNA (e.g., mRNA) of the one or more cells to generate a first population of nucleic acid molecules; and (3) associating the first population of nucleic acid molecules derived from each cell obtained in the previous step with oligonucleotide probes bound to the microdots occupied by the cells from which the first population of nucleic acid molecules was derived, thereby generating a second population of nucleic acid molecules labeled with tag sequence Y. The present invention provides a method comprising:

[0009] In some embodiments, the center-to-center distance between adjacent microdots is less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, or less than 0.01 μm, and the size (e.g., equivalent diameter) of the microdots is less than 5 μm, less than 1 μm, less than 0.3 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, less than 0.01 μm, or less than 0.001 μm.

[0010] In some embodiments, the center-to-center distance between adjacent microdots is between 0.5 μm and 1 μm, for example, between 0.5 μm and 0.9 μm, between 0.5 μm and 0.8 μm.

[0011] In some embodiments, the size (eg, equivalent diameter) of the microdots is between 0.001 μm and 0.5 μm (eg, between 0.01 μm and 0.1 μm, between 0.01 μm and 0.2 μm, between 0.2 μm and 0.5 μm, between 0.2 μm and 0.4 μm, between 0.2 μm and 0.3 μm).

[0012] In certain embodiments, the first binding molecule is capable of forming a specific or non-specific interaction pair with the first labeled molecule.

[0013] In certain embodiments, the interaction pair is selected from the group consisting of positive and negative charge interaction pairs, affinity interaction pairs (e.g., biotin / avidin, biotin / streptavidin, antigen / antibody, receptor / ligand, enzyme / cofactor), pairs of molecules capable of undergoing click chemistry reactions (e.g., alkynyl-containing compound / azide compound), N-hydroxysulfosuccinate (NHS) ester / amino-containing compound, and any combination thereof.

[0014] In certain embodiments, the first labeled molecule is polylysine and the first binding molecule is a protein capable of binding to polylysine; the first labeled molecule is an antibody and the first binding molecule is an antigen capable of binding to the antibody; the first labeled molecule is an amino-containing compound and the first binding molecule is an N-hydroxysulfosuccinate (NHS) ester; or the first labeled molecule is biotin and the first binding molecule is streptavidin.

[0015] In certain embodiments, the first binding molecule is naturally occurring in a cell.

[0016] In certain embodiments, the first binding molecule is non-naturally occurring in the cell.

[0017] In certain embodiments, the method further comprises binding the first binding molecule to one or more cells or expressing the first binding molecule in one or more cells to provide the cell sample of step (1).

[0018] In certain embodiments, the method further comprises binding a first labeled molecule of a solid support to provide a nucleic acid array of step (1).

[0019] Scheme I

[0020] In some embodiments, in step (2), the pretreatment comprises: (i) reverse transcribing RNA (e.g., mRNA) of one or more cells using primer IA to generate extension products as first nucleic acid molecules to be labeled, thereby generating a first population of nucleic acid molecules, where primer IA comprises consensus sequence A and capture sequence A, where capture sequence A is capable of annealing to the RNA (e.g., mRNA) to be captured to initiate an extension reaction, and where consensus sequence A is located upstream of capture sequence A (e.g., located at the 5' end of primer IA); or (ii) (a) reverse transcribing RNA (e.g., mRNA) of one or more cells using primer IA to generate a cDNA strand, the cDNA strand being formed by reverse transcription primed by primer IA and comprising a cDNA sequence that is complementary to the RNA (e.g., mRNA) and a 3'-end overhang, primer IA comprising consensus sequence A and capture sequence A, capture sequence A capable of annealing to the RNA (e.g., mRNA) to be captured to initiate an extension reaction, consensus sequence A being located upstream of capture sequence A (e.g., between the 5' end of primer IA and the 5' end of primer IA); (b) annealing primer IB with the cDNA strand generated in (a) to initiate an extension reaction to generate a first extension product as a first nucleic acid molecule to be labeled, thereby generating a first population of nucleic acid molecules, wherein primer IB comprises consensus sequence B, a complementary sequence of the 3'-end overhang and optionally a tag sequence B, wherein the complementary sequence of the 3'-end overhang is located at the 3' end of primer IB, and consensus sequence B is located upstream of the complementary sequence of the 3'-end overhang (e.g., at the 5' end of primer IB); or (iii) (a) reverse transcribing RNA (e.g., mRNA) of one or more cells using primer I-A' to generate a cDNA strand, the cDNA strand being formed by reverse transcription primed by primer I-A' and comprising a cDNA sequence that is complementary to the RNA (e.g., mRNA) and a 3'-end overhang, primer I-A' comprising capture sequence A, which is capable of annealing to the RNA (e.g., mRNA) to be captured to initiate an extension reaction; and (b) annealing primer IB to the cDNA strand generated in (a) to initiate an extension reaction. (c) providing an extension primer and performing an extension reaction using the first extension products as templates to generate second extension products as labeled first nucleic acid molecules, thereby generating a population of first nucleic acid molecules. Including, In step (3), a first population of nucleic acid molecules derived from each cell obtained in the previous step is associated with an oligonucleotide probe bound to a microdot occupied by the cell from which the first population of nucleic acid molecules is derived, thereby generating a second population of nucleic acid molecules labeled with a tag sequence Y, which comprises: contacting the bridging oligonucleotide I with the first nucleic acid molecule derived from each cell and obtained in step (2) and the oligonucleotide probes bound to the microdots occupied by the cells under conditions that allow annealing; annealing (e.g., in situ annealing) the bridging oligonucleotide I with the first nucleic acid molecule derived from each cell and obtained in step (2) and the oligonucleotide probes bound to the microdots occupied by the cells; and ligating the first nucleic acid molecule and the oligonucleotide probes of the array annealed with the bridging oligonucleotide I to obtain a ligation product as a second nucleic acid molecule having a positioning tag, thereby generating a second population of nucleic acid molecules. Including, Bridge oligonucleotide I comprises a first region and a second region, and optionally a third region located between the first region and the second region, the first region being upstream of the second region (e.g., 5' of the second region); the first region is capable of annealing to all or part of consensus sequence A of primer IA described in step (2)(i) or step (2)(ii), or is capable of annealing to all or part of consensus sequence B of primer IB described in step (2)(iii); The second region is capable of annealing to all or part of the consensus sequence X2.

[0021] In certain embodiments, in step (3), the first region and the second region of the bridging oligonucleotide I are immediately adjacent, and ligating the first nucleic acid molecule to the oligonucleotide probe comprises using a nucleic acid ligase to ligate the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide I to obtain a ligation product as a second nucleic acid molecule having a positioning tag; or The bridging oligonucleotide I comprises a first region, a second region and a third region located therebetween, and ligating the first nucleic acid molecule to the oligonucleotide probe comprises using a nucleic acid polymerase to carry out a polymerization reaction using the third region as a template, and using a nucleic acid ligase to ligate the nucleic acid molecule hybridized with the first region of the same bridging oligonucleotide I to the nucleic acid molecule hybridized with the third region and the second region to obtain a ligation product as a second nucleic acid molecule having a positioning tag, and preferably the nucleic acid polymerase does not have 5' to 3' exonuclease activity or strand displacement activity.

[0022] In certain embodiments, each type of oligonucleotide probe comprises one copy.

[0023] In certain embodiments, each type of oligonucleotide probe comprises multiple copies.

[0024] It is easy to understand that when each type of oligonucleotide probe contains one copy, each microdot is bound to one oligonucleotide probe, and the oligonucleotide probes of different microdots have different tag sequences Y, and when each type of oligonucleotide probe contains multiple copies, each microdot is bound to multiple oligonucleotide probes, and the oligonucleotide probes of the same microdot have the same tag sequence Y, and the oligonucleotide probes of different microdots have different tag sequences Y.

[0025] In certain embodiments, the solid support comprises a plurality of microdots, each microdot being bound to one type of oligonucleotide probe, and each type of oligonucleotide probe may comprise one or multiple copies.

[0026] In certain embodiments, the solid support comprises a plurality of (e.g., at least 10, at least 10 2 Pieces, at least 10 3Pieces, at least 10 4 Pieces, at least 10 5 Pieces, at least 10 6 Pieces, at least 10 7 Pieces, at least 10 8 In certain embodiments, the solid support comprises at least 10 4 (e.g., at least 10 4 Pieces, at least 10 5 Pieces, at least 10 6 Pieces, at least 10 7 Pieces, at least 10 8 Pieces, at least 10 9 Pieces, at least 10 10 Pieces, at least 10 11 or at least 10 12 pcs) microdots / mm 2 Includes.

[0027] An embodiment including steps (1), (2)(i) and (3) In certain embodiments, the method includes steps (1), (2)(i) and (3), and the ligation product obtained in step (3) is taken as a second nucleic acid molecule having a positioning tag at 5' to 3' thereof, the positioning tag including the consensus sequence X1, the tag sequence Y, the consensus sequence X2, optionally the complementary sequence of the third region of the bridging oligonucleotide I, and the sequence of the first nucleic acid molecule to be labeled.

[0028] In certain embodiments, in step (2)(i) of the method, capture sequence A is a random oligonucleotide sequence.

[0029] In some embodiments, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have a different capture sequence A, and capture sequence A serves as a unique molecular identifier (UMI) for the second nucleic acid molecule.

[0030] In certain embodiments, the extension product in step (2)(i) (the first nucleic acid molecule to be labeled) comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by consensus sequence A, primer IA, and that is complementary to the RNA.

[0031] In certain embodiments, in step (2)(i) of the method, capture sequence A is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0032] In certain embodiments, primer IA further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0033] In certain embodiments, capture sequence A is located at the 3' end of primer IA and consensus sequence A is located upstream of tag sequence A (eg, located at the 5' end of primer IA).

[0034] In certain embodiments, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have a different tag sequence A as the UMI.

[0035] In some embodiments, the extension product in step (2)(i) comprises from 5' to 3' the consensus sequence A, the tag sequence A and a cDNA sequence formed by reverse transcription primed by primer IA, and that is complementary to the RNA.

[0036] An embodiment including steps (1), (2)(ii) and (3) In some embodiments, the method includes steps (1), (2)(ii) and (3), and the ligation product obtained in step (3) is taken as a second nucleic acid molecule having a positioning tag comprising, from 5' to 3', the consensus sequence X1, the tag sequence Y and the consensus sequence X2, optionally a complementary sequence of the third region of the bridging oligonucleotide I, and the sequence of the first nucleic acid molecule to be labeled.

[0037] In certain embodiments, in step (2)(ii)(a) of the method, capture sequence A is a random oligonucleotide sequence.

[0038] In some embodiments, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have a different capture sequence A, and capture sequence A serves as a unique molecular identifier (UMI) for the second nucleic acid molecule.

[0039] In certain embodiments, the first extension product in step (2)(ii) (the first nucleic acid molecule to be labeled) comprises, from 5' to 3', consensus sequence A, a cDNA sequence formed by reverse transcription primed by primer IA and complementary to the RNA, and a 3'-end overhang sequence, optionally a complement of tag sequence B and a complement of consensus sequence B.

[0040] In certain embodiments, in step (2)(ii)(a), capture sequence A is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0041] In certain embodiments, primer IA further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0042] In certain embodiments, capture sequence A is located at the 3' end of primer IA and consensus sequence A is located upstream of tag sequence A (eg, located at the 5' end of primer IA).

[0043] In certain embodiments, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have a different tag sequence A as the UMI.

[0044] In certain embodiments, the first extension product in step (2)(ii) (the first nucleic acid molecule to be labeled) comprises, from 5' to 3', consensus sequence A, tag sequence A, a cDNA sequence formed by reverse transcription primed by primer IA and complementary to the RNA and a 3'-end overhang sequence, optionally a complement of tag sequence B and a complement of consensus sequence B.

[0045] In certain embodiments, in the method, primer IA comprises a 5' phosphate at the 5'-end.

[0046] An exemplary embodiment of the present application, including step (1), step (2)(ii) and step (3), is described in detail as follows:

[0047] I. An exemplary embodiment for preparing a cDNA strand using RNA (e.g., mRNA) of a sample as a template includes the following steps (as shown in FIG. 2):

[0048] (1) Using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyl transferase activity) and primer IA, RNA molecules (e.g., mRNA molecules) of a permeabilized cell sample are reverse transcribed to generate cDNA, and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. A variety of reverse transcriptases having terminal deoxynucleotidyl transferase activity can be used to perform reverse transcription. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity. Includes.

[0049] Primer IA comprises a poly(T) sequence and a consensus sequence A (denoted as CA in the figures). In certain embodiments (e.g., when the method is used to construct a 3' transcriptome library), primer IA further comprises a unique molecular identifier (UMI) sequence. Typically, the poly(T) sequence is located at the 3' end of primer IA to initiate reverse transcription. In a preferred embodiment, the UMI sequence is located upstream (e.g., 5') of the poly(T) sequence and the consensus sequence A is located upstream (e.g., 5') of the UMI sequence.

[0050] (2) Primer IB, which contains consensus sequence B (denoted as CB in the figure), can be used to anneal or hybridize with the cDNA strand, and then the nucleic acid fragment hybridized or annealed with primer IB can be extended in the presence of a nucleic acid polymerase using consensus sequence B as a template to add the complementary sequence of consensus sequence B to the 3' end of the cDNA strand, thereby generating a nucleic acid molecule bearing consensus sequence A and tag sequence A at the 5' end and the complementary sequence of consensus sequence B at the 3' end.

[0051] Primer IB may comprise a sequence that is complementary to the 3'-end overhang of the cDNA strand. For example, if the cDNA strand comprises an overhang of three cytosine nucleotides at its 3' end, primer IB may comprise GGG at its 3' end. Furthermore, the nucleotides of primer IB may be modified (e.g., primer IB may be modified to comprise one or more locked nucleic acids) to enhance the binding affinity for complementary pairing between primer IB and the 3'-end overhang of the cDNA strand.

[0052] Without being limited by any theory, the extension reaction can be carried out using various suitable nucleic acid polymerases (e.g., DNA polymerases or reverse transcriptases) as long as the hybridized or annealed nucleic acid fragment (reverse transcription product) can be extended using a partial sequence of primer IB as a template. In certain exemplary embodiments, the hybridized or annealed nucleic acid fragment (reverse transcription product) can be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0053] In certain preferred embodiments, step (2) and step (1) are carried out simultaneously.

[0054] In certain embodiments, the method optionally further comprises step (3): adding RNase H to digest the RNA strand of the RNA / cDNA hybrid to form a single-stranded cDNA.

[0055] In certain preferred embodiments, the method does not include step (3).

[0056] An exemplary structure of a cDNA strand prepared by the above exemplary embodiment includes a consensus sequence A, a UMI sequence, a sequence that is complementary to a sequence of an RNA (eg, an mRNA), and a complementary sequence of consensus sequence B.

[0057] II. An exemplary embodiment of labeling the 5' end of a cDNA strand with an oligonucleotide probe (also called a chip sequence) (i.e., ligating the 5' end of the cDNA strand to the 3' end of the chip sequence) to form a new nucleic acid molecule bearing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) comprises the following steps (as shown in FIG. 3): A bridging oligonucleotide I is provided that comprises at its 5' end a sequence (first region, P1) that is at least partially complementary to the 5' end of the cDNA sequence (e.g., at least partially complementary to consensus sequence A (CA)) and at its 3' end a sequence (second region, P2) that is at least partially complementary to the 3' end of the chip sequence (e.g., at least partially complementary to consensus sequence X2).

[0058] In certain preferred embodiments, the P1 and P2 sequences of bridging oligonucleotide I are immediately adjacent with no intervening nucleotide between them.

[0059] In certain preferred embodiments, the P1 sequence, the P2 sequence, the consensus sequence A, and the consensus sequence X2 each independently have a length of 20-100 nt (e.g., 20-70 nt). The bridging oligonucleotide I is annealed or hybridized with the oligonucleotide probe and the cDNA strand, and then the 5' end of the cDNA strand is ligated to the 3' end of the oligonucleotide probe by a DNA ligase and / or a DNA polymerase, thereby forming a novel nucleic acid molecule (i.e., a nucleic acid molecule labeled with the oligonucleotide probe) that contains the sequence information of the oligonucleotide probe. In certain preferred embodiments, the DNA polymerase does not have 5' to 3' exonuclease activity or strand displacement activity.

[0060] An exemplary structure of a novel nucleic acid molecule containing chip sequence information formed by the above exemplary embodiment includes a consensus sequence X1, a tag sequence Y, a consensus sequence X2, a consensus sequence A, a UMI sequence, and a sequence complementary to a sequence of an RNA (e.g., mRNA), and a complementary sequence of the consensus sequence B.

[0061] An embodiment including steps (1), (2)(iii) and (3) In some embodiments, the method includes steps (1), (2)(iii) and (3), and the ligation product obtained in step (3) is taken as a second nucleic acid molecule having a positioning tag comprising, from 5' to 3', the consensus sequence X1, the tag sequence Y, the consensus sequence X2, optionally the complementary sequence of the third region of the bridging oligonucleotide I and the sequence of the first nucleic acid molecule to be labeled.

[0062] In certain embodiments, in step (2)(iii)(c) of the method, the extension primer is primer IB or primer B″, where primer B″ is capable of annealing to the complementary sequence of consensus sequence B or a subsequence thereof to initiate the extension reaction.

[0063] In certain embodiments, in step (2)(iii)(c), the extension primer is primer B″.

[0064] In certain embodiments, in step (2)(iii)(a) of the method, capture sequence A of primer IA' is a random oligonucleotide sequence.

[0065] In certain embodiments, in step (2)(iii)(b), primer IB comprises consensus sequence B, a complementary sequence of the 3′-end overhang and tag sequence B.

[0066] In certain embodiments, the first extension product is formed from 5' to 3' by reverse transcription primed by primer I-A' and comprises a cDNA sequence complementary to the RNA sequence, a 3'-end overhang sequence, a complement of tag sequence B and a complement of consensus sequence B, where the complement of tag sequence B serves as a unique molecular identifier (UMI) for the second nucleic acid molecule.

[0067] In certain embodiments, in step (2)(iii)(c), the second extension product (first nucleic acid molecule to be labeled) comprises, from 5' to 3' to the first extension product, consensus sequence B or a partial sequence thereof at its 3' end, tag sequence B, a complementary sequence of the 3'-end overhang sequence, and a complementary sequence of the cDNA sequence, wherein tag sequence B serves as a unique molecular identifier (UMI) of the second nucleic acid molecule.

[0068] In certain embodiments, in step (2)(iii)(a) of the method, capture sequence A of primer IA' is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0069] In certain embodiments, primer IA' further comprises a tag sequence A, eg, a random oligonucleotide sequence and a consensus sequence A.

[0070] In certain embodiments, capture sequence A is located at the 3' end of primer IA'.

[0071] In certain embodiments, consensus sequence A is located upstream of capture sequence A (eg, at the 5' end of primer IA').

[0072] In certain embodiments, primer IB comprises consensus sequence B, a complementary sequence of the 3′-end overhang and tag sequence B.

[0073] In certain embodiments, in step (2)(iii)(b), the first extension product comprises, from 5' to 3', consensus sequence A, optionally tag sequence A, a cDNA sequence formed by reverse transcription primed by primer I-A' and that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0074] In certain embodiments, in step (2)(iii)(c), the second extension product (first nucleic acid molecule to be labeled) comprises, from 5' to 3', consensus sequence B or a partial sequence thereof at its 3' end, tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence of the first extension product, and optionally, a complement of tag sequence A and a complement of consensus sequence A.

[0075] In some embodiments, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have a different tag sequence B as a UMI.

[0076] In certain embodiments, the extension primer comprises a 5' phosphate at the 5'-end.

[0077] In certain embodiments, prior to step (2)(iii)(c), the method further comprises treating the product of step (2)(iii)(a) or step (2)(iii)(b) to remove RNA (e.g., by heat treatment).

[0078] In certain embodiments, in step (2)(iii)(b) of the method, the cDNA strand is annealed with primer IB via its 3'-end overhang, and the cDNA strand is extended using primer IB as a template to generate a first extension product in the presence of a nucleic acid polymerase (e.g., a DNA polymerase or a reverse transcriptase).

[0079] An exemplary embodiment of the present application, including step (1), step (2)(iii) and step (3), is described in detail as follows:

[0080] I. An exemplary embodiment for preparing a complementary strand of a cDNA strand using RNA (e.g., mRNA) in a sample as a template includes the following steps (as shown in FIG. 4):

[0081] (1) Using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyl transferase activity) and primer I-A', RNA molecules (e.g., mRNA molecules) of the permeabilized sample are reverse transcribed to generate cDNA, and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. A variety of reverse transcriptases having terminal deoxynucleotidyl transferase activity can be used to perform the reverse transcription. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity.

[0082] The reverse transcription primer IA' contains a poly(T) sequence and a consensus sequence A (CA). Usually, the poly(T) sequence is located at the 3' end of the primer IA to initiate reverse transcription.

[0083] (2) A primer IB is used to anneal or hybridize with the cDNA strand, and the primer IB comprises a consensus sequence B (CB) and a complementary sequence of the 3'-end overhang of the cDNA. In certain embodiments (e.g., when the method is used to construct a 5' transcriptome library), the primer IB further comprises a unique molecular identifier (UMI) sequence. The nucleic acid fragment hybridized or annealed with the primer IB can then be extended in the presence of a nucleic acid polymerase using the consensus sequence B and the UMI sequence as templates to add the complementary sequence of the consensus sequence B and the complementary sequence of the UMI sequence to the 3' end of the cDNA strand, thereby generating a nucleic acid molecule that retains the consensus sequence A at its 5' end and the complementary sequence of the consensus sequence B and the complementary sequence of the UMI molecule at its 3' end.

[0084] In the case where the cDNA strand comprises an overhang of three cytosine nucleotides at its 3' end, primer IB may comprise GGG at its 3' end. Furthermore, the nucleotides of primer IB may also be modified to enhance the binding affinity for complementary pairing between primer IB and the 3'-end overhang of the cDNA strand (e.g., primer IB may be modified to comprise one or more locked nucleic acids).

[0085] Without being limited by any theory, the extension reaction can be carried out using a variety of suitable nucleic acid polymerases (e.g., DNA polymerases or reverse transcriptases) as long as the sequence of primer IB or a subsequence thereof can be used as a template to extend the annealed or hybridized nucleic acid fragment (reverse transcription product). In certain exemplary embodiments, the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0086] In certain preferred embodiments, step (2) and step (1) are carried out simultaneously.

[0087] In certain embodiments, the method optionally further comprises step (3): digesting the RNA strand of the RNA / cDNA hybrid to form a single-stranded cDNA.

[0088] In certain preferred embodiments, the method does not include step (3).

[0089] (4) Using an extension primer, an extension reaction is carried out using the single-stranded cDNA obtained in step (3) as a template to obtain an extension product, and the extension primer is capable of annealing to the complementary sequence of consensus sequence B or a partial sequence thereof to initiate the extension reaction.

[0090] In certain embodiments, the extension primer is identical to primer IB.

[0091] An exemplary structure comprising a complement of a cDNA strand prepared by the above exemplary embodiment comprises consensus sequence B, a UMI sequence, a sequence complementary to the 3'-end overhang sequence of the cDNA, a complement of the cDNA sequence, and a complement of consensus sequence A.

[0092] II. An exemplary embodiment of using an oligonucleotide probe (also referred to as a chip sequence) to label the 5' end of a complementary strand of a cDNA strand (i.e., ligating the 5' end of the complementary strand of a cDNA strand to the 3' end of the chip sequence) to generate a new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) comprises the following steps (as shown in FIG. 5): A bridging oligonucleotide I is provided that comprises at its 5' end a sequence (first region, P1) that is at least partially complementary to consensus sequence B (CB) and at its 3' end a sequence (second region, P2) that is at least partially complementary to consensus sequence X2.

[0093] In certain preferred embodiments, the P1 and P2 sequences of bridging oligonucleotide I are immediately adjacent with no intervening nucleotide between them.

[0094] In certain preferred embodiments, the P1 sequence and the P2 sequence each independently have a length of 20 to 100 nt (eg, 20 to 70 nt).

[0095] The bridge oligonucleotide I is annealed or hybridized with the oligonucleotide probe and the complementary strand of the cDNA strand, and then the 5' end of the complementary strand of the cDNA strand is ligated to the 3' end of the chip sequence by DNA ligase and / or DNA polymerase to form a new nucleic acid molecule containing the sequence information of the oligonucleotide probe (i.e., a nucleic acid molecule labeled with the oligonucleotide probe). In certain preferred embodiments, the DNA polymerase does not have 5' to 3' exonuclease activity or strand displacement activity.

[0096] An exemplary structure of a novel nucleic acid molecule containing chip sequence information formed by the above exemplary embodiment includes a consensus sequence X1, a tag sequence Y, a consensus sequence X2, a consensus sequence B, a UMI sequence, a complementary sequence of a cDNA sequence, and a complementary sequence of a consensus sequence A.

[0097] Scheme II

[0098] In some embodiments, in step (2), the pretreatment comprises: (i) (a) reverse transcribing RNA (e.g., mRNA) of one or more cells using primer II-A to generate a cDNA strand, the cDNA strand being formed by reverse transcription primed by primer II-A and comprising a cDNA sequence that is complementary to the RNA (e.g., mRNA) and a 3'-end overhang, primer II-A comprising capture sequence A, which is capable of annealing to the RNA (e.g., mRNA) to be captured and initiating an extension reaction; (b) applying primer II-B to (a); and performing an extension reaction to generate a first extension product as a first nucleic acid molecule to be labeled, thereby generating a first population of nucleic acid molecules, wherein primer II-B comprises consensus sequence B, a complementary sequence of the 3'-end overhang and optionally a tag sequence B, wherein the complementary sequence of the 3'-end overhang is located at the 3' end of primer II-B, and consensus sequence B is located upstream of the complementary sequence of the 3'-end overhang (e.g., located at the 5' end of primer II-B); or (ii)(a) reverse transcribing RNA (e.g., mRNA) of one or more cells using primer II-A' to generate a cDNA strand, the cDNA strand being formed by reverse transcription primed by primer II-A' and comprising a cDNA sequence that is complementary to the RNA (e.g., mRNA) and a 3'-end overhang, primer II-A' comprising consensus sequence A and capture sequence A, where capture sequence A is capable of annealing to the RNA (e.g., mRNA) to be captured and initiating an extension reaction, and where consensus sequence A is located upstream of capture sequence A (e.g., located at the 5' end of primer II-A'); (b) primer II-B' is added to (a); (c) providing an extension primer and performing an extension reaction using the first extension product as a template to generate a second extension product as a labeled first nucleic acid molecule, thereby generating a first population of nucleic acid molecules; Including, Further, in step (3), a first population of nucleic acid molecules derived from each cell obtained in the previous step is associated with an oligonucleotide probe bound to a microdot occupied by the cell from which the first population of nucleic acid molecules is derived, thereby generating a second population of nucleic acid molecules labeled with a tag sequence Y, which is (i) annealing (e.g., in situ annealing) a first nucleic acid molecule derived from each cell obtained in step (2) with an oligonucleotide probe bound to a microdot occupied by the cell by applying annealing conditions to the product of step (2) and performing an extension reaction to generate an extension product as a second nucleic acid molecule having a positioning tag, thereby generating a second population of nucleic acid molecules, wherein the consensus sequence X2 of the oligonucleotide probe or a subsequence thereof is (a) capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence B of the first extension product obtained in step (2)(i), or (b) capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence A of the second extension product obtained in step (2)(ii); or (ii) contacting a bridging oligonucleotide pair with the first nucleic acid molecule derived from each cell obtained in step (2) and the oligonucleotide probe bound to the microdots occupied by the cells under conditions that allow annealing, and allowing the bridging oligonucleotide pair to anneal (e.g., in situ anneal) with the first nucleic acid molecule derived from each cell obtained in step (2) and the oligonucleotide probe bound to the microdots occupied by the cells; wherein the bridging oligonucleotide pair is comprised of a bridging oligonucleotide II-I and a bridging oligonucleotide II-II, each of which independently comprises a first region, a second region, and optionally a third region located between the first region and the second region, wherein the first region is located upstream of the second region (e.g., 5' of the second region); a first region of the bridging oligonucleotide II-I capable of annealing to a first region of the bridging oligonucleotide II-II, and a second region of the bridging oligonucleotide II-I capable of annealing to a consensus sequence X2 of the oligonucleotide probe or a subsequence thereof; the second region of bridging oligonucleotide II-II is capable of annealing (a) to a complementary sequence or a subsequence thereof of consensus sequence B of the first extension product obtained in step (2)(i); or (b) to a complementary sequence or a subsequence thereof of consensus sequence A of the second extension product obtained in step (2)(ii); In the bridging oligonucleotide pair contacting the first nucleic acid molecule and the oligonucleotide probe, the bridging oligonucleotide II-I and the bridging oligonucleotide II-II of the bridging oligonucleotide pair are each present in a single-stranded form, or the bridging oligonucleotide II-I and the bridging oligonucleotide II-II of the bridging oligonucleotide pair are annealed to each other and present in a partially double-stranded form; performing a ligation reaction: ligating the nucleic acid molecules hybridized with the first region and the second region of the same bridging oligonucleotide II-I, and / or ligating the nucleic acid molecules hybridized with the first region and the second region of the same bridging oligonucleotide II-II, and performing an extension reaction to obtain the reaction product as a second nucleic acid molecule having a positioning tag, thereby generating a second population of nucleic acid molecules, wherein the ligation reaction and the extension reaction are performed in any order. Includes.

[0099] In certain embodiments, in step (3)(ii): (1) the first region and the second region of the bridging oligonucleotide II-I are immediately adjacent, and ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I comprises ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I using a nucleic acid ligase; or the bridging oligonucleotide II-I comprises a first region, a second region and a third region therebetween, and ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I comprises performing a polymerization reaction using a nucleic acid polymerase (e.g., a nucleic acid polymerase that does not have 5' to 3' exonuclease activity or strand displacement activity) with the third region as a template, and ligating the nucleic acid molecule hybridized with the first region of the same bridging oligonucleotide II-I to the nucleic acid molecule hybridized with the third region and the second region using a nucleic acid ligase; and / or (2) the first region and the second region of the bridging oligonucleotide II-II are immediately adjacent, and ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II comprises ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II using a nucleic acid ligase; or The bridging oligonucleotide II-II comprises a first region, a second region and a third region therebetween, and ligation of the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II comprises performing a polymerization reaction using a nucleic acid polymerase (e.g., a nucleic acid polymerase that does not have 5' to 3' exonuclease activity or strand displacement activity) using the third region as a template, and ligating the nucleic acid molecule hybridized with the first region of the same bridging oligonucleotide II-II to the nucleic acid molecule hybridized with the third region and the second region using a nucleic acid ligase.

[0100] In certain embodiments, the method comprises steps (1), (2)(i) and (3), wherein in step (2)(i)(b), primer II-B comprises consensus sequence B, a complementary sequence of the 3'-end overhang and tag sequence B.

[0101] In some embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0102] In certain embodiments, in step (3), the second nucleic acid molecules derived from each copy of the oligonucleotide probe bound to the same microdot have a different tag sequence B as a UMI.

[0103] An embodiment including steps (1), (2)(i) and (3)(i) In certain embodiments, the method comprises steps (1), (2)(i) and (3)(i), wherein the consensus sequence X2 or a subsequence thereof is capable of annealing to a complementary sequence of the consensus sequence B or a subsequence thereof, and the extension product obtained in step (3)(i) is taken as a labeled nucleic acid molecule comprising a first strand comprising the sequence of the first nucleic acid molecule to be labeled and / or a second strand comprising the sequence of the oligonucleotide probe.

[0104] It is easy to understand that "a subsequence of XX (sequence)" or "a subsequence of XX (sequence)" refers to the nucleotide sequence of at least one segment of "XX (sequence)."

[0105] For example, the consensus sequence X2 can anneal with the complementary sequence of the consensus sequence B or a partial segment thereof at its entire nucleotide sequence, and the consensus sequence X2 can also anneal with the complementary sequence of the consensus sequence B or a partial segment thereof at its partial segment nucleotide sequence.

[0106] "Annealing" means that, between two nucleotide sequences that anneal to each other, each base of one nucleotide sequence can pair with the base of the other nucleotide sequence without mismatch or gap, or, between two nucleotide sequences that anneal to each other, most of the bases of one nucleotide sequence can pair with the base of the other nucleotide sequence, allowing mismatch or gap (e.g., mismatch or gap of one or several nucleotides). That is, two nucleotide sequences that can be annealed may be completely complementary or partially complementary. The description "annealing" is applied to the entire text of this application unless otherwise specified or clearly contradicted in the context.

[0107] In certain embodiments, the first strand comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0108] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, tag sequence B, a complement of the 3'-end overhang sequence, and a complement of the cDNA sequence formed by reverse transcription primed by primer II-A and that is complementary to the RNA.

[0109] An embodiment comprising steps (1), (2)(i) and (3)(i) for producing the first strand In certain embodiments, the consensus sequence X2 or a subsequence thereof is capable of annealing to a complementary sequence of the consensus sequence B or a subsequence thereof (e.g., a subsequence thereof at its 3' end), and the complementary sequence of the consensus sequence B of the first extension product in step (2)(i) has a free 3' end.

[0110] In certain embodiments, the extension product obtained in step (3)(i) is a labeled nucleic acid molecule comprising a first strand.

[0111] In certain embodiments, the first strand comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0112] In certain embodiments, in step (3)(i), the oligonucleotide probe is incapable of initiating an extension reaction (eg, the 3' end of the oligonucleotide probe is blocked).

[0113] In certain embodiments, in step (2)(i)(a) of the method, capture sequence A of primer II-A is a random oligonucleotide sequence.

[0114] In some embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0115] In certain embodiments, the first strand comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0116] In certain embodiments, in step (2)(i)(a) of the method, capture sequence A of primer II-A is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0117] In certain embodiments, primer II-A comprises a consensus sequence A, and optionally further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0118] In certain embodiments, capture sequence A is located at the 3' end of primer II-A.

[0119] In certain embodiments, consensus sequence A is located upstream of capture sequence A (eg, located at the 5' end of primer II-A).

[0120] In certain embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', consensus sequence A, optionally tag sequence A, a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0121] In certain embodiments, the first strand comprises, from 5' to 3', consensus sequence A, optionally tag sequence A, a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0122]

[0033] Embodiments including steps (1), (2)(i), and (3)(i) for producing the second strand In certain embodiments, the consensus sequence X2 or a subsequence thereof (e.g., a subsequence thereof at its 3' end) is capable of annealing to a complementary sequence of the consensus sequence B or a subsequence thereof, and the consensus sequence X2 of the oligonucleotide probe has a free 3' end.

[0123] In certain embodiments, the extension product obtained in step (3)(i) is a labeled nucleic acid molecule comprising a second strand.

[0124] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, tag sequence B, a complement of the 3'-end overhang sequence, and a complement of the cDNA sequence formed by reverse transcription primed by primer II-A and that is complementary to the RNA.

[0125] In certain embodiments, the first extension product obtained in step (2)(i) is incapable of initiating an extension reaction (e.g., the 3' end of the first extension product obtained in step (2)(i) is blocked).

[0126] In certain embodiments, in step (2)(i)(a) of the method, capture sequence A of primer II-A is a random oligonucleotide sequence.

[0127] In some embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0128] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, tag sequence B, a complement of the 3'-end overhang sequence, and a complement of the cDNA sequence formed by reverse transcription primed by primer II-A and that is complementary to the RNA.

[0129] In certain embodiments, in step (2)(i)(a) of the method, capture sequence A of primer II-A is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0130] In certain embodiments, primer II-A comprises a consensus sequence A, and optionally further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0131] In certain embodiments, capture sequence A is located at the 3' end of primer II-A.

[0132] In certain embodiments, consensus sequence A is located upstream of capture sequence A (eg, located at the 5' end of primer II-A).

[0133] In certain embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', consensus sequence A, optionally tag sequence A, a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0134] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A and that is complementary to the RNA, and optionally a complement of tag sequence A and a complement of consensus sequence A.

[0135] An exemplary embodiment of the present application including steps (1), (2)(i) and (3)(i) is detailed as follows:

[0136] I. An exemplary embodiment using a sample's RNA (e.g., mRNA) as a template to prepare a cDNA strand that contains a complementary sequence of a UMI proximal to its 3' end includes the following steps (as shown in FIG. 6):

[0137] (1) Using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyl transferase activity) and primer II-A, RNA molecules (e.g., mRNA molecules) of the permeabilized sample are reverse transcribed to generate cDNA, and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. Various reverse transcriptases having terminal deoxynucleotidyl transferase activity can be used to perform reverse transcription. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity.

[0138] In certain embodiments, primer II-A comprises a poly(T) sequence and a consensus sequence A (CA). Typically, the poly(T) sequence is located at the 3' end of primer II-A to initiate reverse transcription.

[0139] In certain embodiments, primer II-A contains a random oligonucleotide sequence that can be used to capture RNA that does not have a poly(A) tail. Typically, the random oligonucleotide sequence is located at the 3' end of primer II-A to initiate reverse transcription.

[0140] (2) Primer II-B is used to anneal or hybridize with the cDNA strand, where primer II-B contains consensus sequence B (CB), a unique molecular identifier (UMI) sequence, and the complement of the 3'-end overhang of the cDNA. The nucleic acid fragment hybridized or annealed with primer II-B can then be extended in the presence of a nucleic acid polymerase using the UMI sequence and consensus sequence B as templates, thereby generating a nucleic acid molecule bearing at its 3' end the complement of the UMI sequence and the complement of consensus sequence B.

[0141] Typically, consensus sequence B is located upstream of the UMI sequence (eg, located 5' of the UMI sequence), and a sequence complementary to the 3'-end overhang of the cDNA strand is located at the 3' end of primer II-B.

[0142] For example, if the cDNA strand contains an overhang of three cytosine nucleotides at its 3' end, primer II-B may contain GGG at its 3' end. Additionally, the nucleotides of primer II-B may be modified (e.g., primer II-B may be modified to contain one or more locked nucleic acids) to enhance the binding affinity for complementary pairing between primer II-B and the 3'-end overhang of the cDNA strand.

[0143] Without being limited by any theory, the extension reaction can be carried out using a variety of suitable nucleic acid polymerases (e.g., DNA polymerases or reverse transcriptases) as long as the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the sequence of primer II-B or a subsequence thereof as a template. In certain exemplary embodiments, the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0144] In some embodiments, this step is carried out simultaneously with step (1) (eg, in the same reaction system).

[0145] In certain embodiments, the method optionally further comprises step (3): adding RNase H to digest the RNA strand of the RNA / cDNA hybrid to form a single-stranded cDNA.

[0146] In certain embodiments, the method does not include step (3).

[0147] An exemplary structure of a cDNA strand prepared by the above exemplary embodiment includes consensus sequence A, a cDNA sequence, a 3'-end overhang sequence, a complement of a UMI sequence, and a complement of consensus sequence B.

[0148] II. An exemplary embodiment of labeling the 3' end of a cDNA strand with the complementary sequence of an oligonucleotide probe (also called a chip sequence) to form a new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) comprises the following steps (as shown in FIG. 8): In some embodiments, the consensus sequence X2 of the chip sequence or a subsequence thereof can be annealed to the complementary sequence or a subsequence thereof of the consensus sequence B of the cDNA strand obtained in the above step I. Thereby, a new nucleic acid molecule containing the chip sequence information (i.e., a nucleic acid molecule labeled with the chip sequence) can be obtained by annealing or hybridizing the cDNA strand with the chip sequence and performing an extension reaction in the presence of a polymerase.

[0149] An exemplary structure of a novel nucleic acid molecule containing chip sequence information formed by the above exemplary embodiment includes a nucleic acid strand and / or its complementary nucleic acid strand, the nucleic acid strand including, from 5' to 3', a consensus sequence A, a cDNA sequence, a 3'-end overhang sequence, a complementary sequence of a UMI sequence, a complementary sequence of a consensus sequence B, a complementary sequence of a tag sequence Y, and a complementary sequence of a consensus sequence X1.

[0150] An embodiment including steps (1), (2)(i) and (3)(ii) In certain embodiments, the method comprises steps (1), (2)(i) and (3)(ii), wherein the second region of the bridging oligonucleotide II-II is capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence B of the first extension product obtained in step (2)(i), and the reaction product obtained in step (3)(ii) is taken as a labeled nucleic acid molecule comprising a first strand comprising the sequence of the first nucleic acid molecule to be labeled and / or a second strand comprising the sequence of the oligonucleotide probe.

[0151] It is easy to understand that the second region of the bridging oligonucleotide II-II is capable of annealing to the complementary sequence of the consensus sequence B of the first extension product obtained in step (2)(i) or a partial segment thereof.

[0152] In certain embodiments, the first strand comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, optionally a complement of the third region of bridging oligonucleotide II-II, a sequence of bridging oligonucleotide II-I, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0153] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, optionally a complement of the third region of bridging oligonucleotide II-I, a sequence of bridging oligonucleotide II-II, tag sequence B, a complement of the 3'-end overhang sequence, and a complement of the cDNA sequence formed by reverse transcription primed by primer II-A and that is complementary to the RNA.

[0154] An embodiment comprising steps (1), (2)(i) and (3)(ii) for producing the first strand In certain embodiments, the second region of the bridging oligonucleotide II-II is capable of annealing to the complementary sequence of the consensus sequence B of the first extension product obtained in step (2)(i) or a subsequence thereof (e.g., a 3'-terminal subsequence thereof), and the second region of the bridging oligonucleotide II-I has a free 3'-end.

[0155] In certain embodiments, the reaction product obtained in step (3)(ii) is a labeled nucleic acid molecule comprising the first strand.

[0156] In certain embodiments, the first strand comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, optionally a complement of the third region of bridging oligonucleotide II-II, a sequence of bridging oligonucleotide II-I, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0157] In certain embodiments, the second region of the bridging oligonucleotide II-I is located at the 3' end of the bridging oligonucleotide II-I.

[0158] In certain embodiments, the first region of the bridging oligonucleotide II-I is located at the 5' end of the bridging oligonucleotide II-I.

[0159] In certain embodiments, bridging oligonucleotide II-I does not comprise the third region and / or bridging oligonucleotide II-II does not comprise the third region.

[0160] In certain embodiments, bridging oligonucleotide II-I comprises a 5' phosphate at the 5' end.

[0161] In certain embodiments, bridging oligonucleotide II-I comprises a free --OH at the 3' end.

[0162] In certain embodiments, in step (3)(ii), the bridging oligonucleotide II-II is incapable of initiating an extension reaction (e.g., the 3' end of the bridging oligonucleotide II-II is blocked) and / or the oligonucleotide probe is incapable of initiating an extension reaction (i.e., the 3' end of the oligonucleotide probe is blocked).

[0163] In some embodiments, in step (2)(i)(a) of the method, capture sequence A of primer II-A is a random oligonucleotide sequence.

[0164] In some embodiments, the first extension product in step (2)(i)(b) of the method comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0165] In certain embodiments, the first strand comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, optionally a complement of the third region of bridging oligonucleotide II-II, a sequence of bridging oligonucleotide II-I, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0166] In certain embodiments, in step (2)(i)(a), capture sequence A of primer II-A is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0167] In certain embodiments, primer II-A comprises a consensus sequence A, and optionally further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0168] In certain embodiments, capture sequence A is located at the 3' end of primer II-A.

[0169] In certain embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', consensus sequence A, optionally tag sequence A, a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0170] In certain embodiments, the first strand comprises, from 5' to 3', consensus sequence A, optionally tag sequence A, a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, a complement of consensus sequence B, optionally a complement of the third region of bridging oligonucleotide II-II, a sequence of bridging oligonucleotide II-I, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0171] It is easy to understand that in step (3)(ii), after the bridging oligonucleotides II-I and II-II are annealed with the oligonucleotide probe and the first nucleic acid molecule to be labeled at the corresponding position of the oligonucleotide probe, the ligation reaction for ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I and / or the extension reaction in step (3)(ii) for ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II can be performed in any order as long as a second nucleic acid molecule having a positioning tag is obtained.

[0172] For example, when the ligation reaction and the extension reaction are carried out in the same system, the first strand can be obtained by ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II, and extending the bridging oligonucleotide II-I by the extension reaction. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0173] For example, when the ligation reaction and the extension reaction are carried out in different systems, a ligation reaction followed by an extension reaction can be carried out and the first strand can be obtained in the following exemplary manner: (A) ligating a nucleic acid molecule hybridized with a first region and a nucleic acid molecule hybridized with a second region of the same bridging oligonucleotide II-II, and extending the bridging oligonucleotide II-I by an extension reaction to obtain a first strand, wherein the polymerase used for the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity. Or, (B) ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-I, and extending the labeled first nucleic acid molecule by an extension reaction to obtain a first strand, wherein the polymerase used for the extension reaction preferably has strand displacement activity or 5' to 3' exonuclease activity.

[0174] For example, when the ligation reaction and the extension reaction are carried out in different systems, an extension reaction is followed by a ligation reaction, and the first strand can be obtained by extending the bridging oligonucleotide II-I by an extension reaction, and then ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-II. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0175]

[0033] Embodiments including steps (1), (2)(i), and (3)(ii) for producing the second strand In certain embodiments, the second region of the bridging oligonucleotide II-II is capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence B of the first extension product obtained in step (2)(i), and the second region of the bridging oligonucleotide II-II has a free 3' end.

[0176] In certain embodiments, the reaction product obtained in step (3)(ii) is a labeled nucleic acid molecule comprising the second strand.

[0177] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, optionally the complement of the third region of bridging oligonucleotide II-I, the sequence of bridging oligonucleotide II-II, tag sequence B, the complement of the 3'-end overhang sequence, and the complement of the cDNA sequence formed by reverse transcription primed by primer II-A and which is complementary to the RNA.

[0178] In certain embodiments, the second region of the bridging oligonucleotide II-II is located at the 3' end of the bridging oligonucleotide II-II.

[0179] In certain embodiments, the first region of the bridging oligonucleotide II-II is located at the 5' end of the bridging oligonucleotide II-II.

[0180] In certain embodiments, bridging oligonucleotide II-I does not comprise the third region and / or bridging oligonucleotide II-II does not comprise the third region.

[0181] In certain embodiments, bridging oligonucleotide II-II comprises a 5' phosphate at the 5' end.

[0182] In certain embodiments, bridging oligonucleotide II-II comprises a free --OH at the 3' end.

[0183] In certain embodiments, in step (3)(ii), the bridging oligonucleotide II-I is incapable of initiating an extension reaction (e.g., the 3' end of the bridging oligonucleotide II-I is blocked), and / or the first extension product obtained in step (2)(i) is incapable of initiating an extension reaction (e.g., the 3' end of the first extension product obtained in step (2)(i) is blocked).

[0184] In certain embodiments, in step (2)(i)(a), capture sequence A of primer II-A is a random oligonucleotide sequence.

[0185] In some embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', a cDNA sequence formed by reverse transcription primed by primer II-A that is complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0186] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, optionally a complement of the third region of bridging oligonucleotide II-I, a sequence of bridging oligonucleotide II-II, tag sequence B, a complement of the 3'-end overhang sequence, and a complement of the cDNA sequence formed by reverse transcription primed by primer II-A and that is complementary to the RNA.

[0187] In certain embodiments, in step (2)(i)(a), capture sequence A of primer II-A is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0188] In certain embodiments, primer II-A comprises a consensus sequence A, and optionally further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0189] In certain embodiments, capture sequence A is located at the 3' end of primer II-A.

[0190] In certain embodiments, the first extension product in step (2)(i)(b) comprises, from 5' to 3', consensus sequence A, optionally tag sequence A, a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, a 3'-end overhang sequence, a complement of tag sequence B, and a complement of consensus sequence B.

[0191] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, and consensus sequence X2, optionally a complement of the third region of bridging oligonucleotide II-I, a sequence of bridging oligonucleotide II-II, tag sequence B, a complement of the 3'-end overhang sequence, a complement of a cDNA sequence formed by reverse transcription primed by primer II-A and complementary to the RNA, and optionally a complement of tag sequence A and a complement of consensus sequence A.

[0192] It is easy to understand that in step (3)(ii), after the bridging oligonucleotides II-I and II-II are annealed with the oligonucleotide probe and the first nucleic acid molecule to be labeled at the corresponding position of the oligonucleotide probe, the ligation reaction for ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I and / or the extension reaction in step (3)(ii) for ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II can be performed in any order as long as a second nucleic acid molecule having a positioning tag is obtained.

[0193] For example, when the ligation reaction and the extension reaction are carried out in the same system, the second strand can be obtained by ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I, and extending the bridging oligonucleotide II-II by an extension reaction. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0194] For example, when the ligation reaction and the extension reaction are carried out in different systems, a ligation reaction followed by an extension reaction can be carried out and the second strand can be obtained in the following exemplary manner: (A) ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I, and extending the bridging oligonucleotide II-II by an extension reaction to obtain a second strand, wherein the polymerase used for the extension reaction preferably has or does not have strand displacement activity or 5' to 3' exonuclease activity. (B) ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-II, and extending the oligonucleotide probe by an extension reaction to obtain a second strand, wherein the polymerase used for the extension reaction preferably has strand displacement activity or 5' to 3' exonuclease activity.

[0195] For example, when the ligation reaction and the extension reaction are carried out in different systems, an extension reaction is followed by a ligation reaction, and the second strand can be obtained by extending the bridging oligonucleotide II-II by an extension reaction, and then ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-I. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0196] The exemplary embodiment of the present application, including step (1), step (2)(i) and step (3)(ii), is detailed as follows:

[0197] I. An exemplary embodiment in which RNA (e.g., mRNA) of a sample is used as a template to prepare a cDNA strand includes the following steps (as shown in FIG. 6):

[0198] (1) Using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyl transferase activity) and primer II-A, RNA molecules (e.g., mRNA molecules) of the permeabilized sample are reverse transcribed to generate cDNA, and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. Various reverse transcriptases having terminal deoxynucleotidyl transferase activity can be used to perform reverse transcription. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity.

[0199] In certain embodiments, primer II-A comprises a poly(T) sequence and a consensus sequence A (CA). Typically, the poly(T) sequence is located at the 3' end of primer II-A to initiate reverse transcription.

[0200] In certain embodiments, primer II-A contains a random oligonucleotide sequence that can be used to capture RNA that does not have a poly(A) tail. Typically, the random oligonucleotide sequence is located at the 3' end of primer II-A to initiate reverse transcription.

[0201] (2) Primer II-B is used to anneal or hybridize with the cDNA strand, where primer II-B contains consensus sequence B (CB), a unique molecular identifier (UMI) sequence, and the complement of the 3'-end overhang of the cDNA. The nucleic acid fragment hybridized or annealed with primer II-B can then be extended in the presence of a nucleic acid polymerase using the UMI sequence and consensus sequence B as templates, thereby generating a nucleic acid molecule bearing at its 3' end the complement of the UMI sequence and the complement of consensus sequence B.

[0202] Typically, consensus sequence B is located upstream of the UMI sequence (eg, located 5' of the UMI sequence) and a sequence complementary to the 3'-end overhang of the cDNA strand is located at the 3' end of primer II-B.

[0203] For example, if the cDNA strand contains an overhang of three cytosine nucleotides at its 3' end, primer II-B may contain GGG at its 3' end. Additionally, the nucleotides of primer II-B may be modified (e.g., primer II-B may be modified to contain one or more locked nucleic acids) to enhance the binding affinity for complementary pairing between primer II-B and the 3'-end overhang of the cDNA strand.

[0204] Without being limited by any theory, the extension reaction can be carried out using a variety of suitable nucleic acid polymerases (e.g., DNA polymerases or reverse transcriptases) as long as the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the sequence of primer II-B or a subsequence thereof as a template. In certain exemplary embodiments, the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0205] In some embodiments, this step is carried out simultaneously with step (1) (eg, in the same reaction system).

[0206] In certain embodiments, the method optionally further comprises step (3): adding RNase H to digest the RNA strand of the RNA / cDNA hybrid to form a single-stranded cDNA.

[0207] In certain embodiments, the method does not include step (3).

[0208] An exemplary structure of a cDNA strand prepared by the above exemplary embodiment includes consensus sequence A, a cDNA sequence, a 3'-end overhang sequence, a complement of a UMI sequence, and a complement of consensus sequence B.

[0209] II. An exemplary embodiment of labeling the 3' end of a cDNA strand with the complementary sequence of an oligonucleotide probe (also referred to as a chip sequence) to form a new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) comprises the following steps (as shown in FIG. 7): providing a bridging oligonucleotide pair consisting of a bridging oligonucleotide II-I and a bridging oligonucleotide II-II, each of which independently comprises a first region (P1) and a second region (P2), the first region being located upstream of the second region (e.g., located 5' of the second region); a first region of the bridging oligonucleotide II-I capable of annealing to a first region of the bridging oligonucleotide II-II, and a second region of the bridging oligonucleotide II-I capable of annealing to a consensus sequence X2 of the oligonucleotide probe or a subsequence thereof; A second region of the bridging oligonucleotide II-II is capable of annealing to the complementary sequence or a subsequence thereof of the consensus sequence B of the cDNA strand obtained in step I above.

[0210] In certain embodiments, the bridging oligonucleotide II-I comprises an intermediate nucleotide sequence between the first and second regions, e.g., an intermediate nucleotide sequence of 1 nt to 5 nt or 5 nt to 10 nt, i.e., the bridging oligonucleotide II-I comprises a third region located between the first and second regions. In certain preferred embodiments, the first and second regions of the bridging oligonucleotide II-I are directly adjacent with no extra nucleotides between them, i.e., the bridging oligonucleotide II-I does not comprise a third region located between the first and second regions.

[0211] In certain embodiments, the bridging oligonucleotide II-II comprises an intermediate nucleotide sequence between the first and second regions, e.g., an intermediate nucleotide sequence of 1 nt to 5 nt or 5 nt to 10 nt, i.e., the bridging oligonucleotide II-II comprises a third region located between the first and second regions. In certain preferred embodiments, the first and second regions of the bridging oligonucleotide II-II are directly adjacent with no extra nucleotides between them, i.e., the bridging oligonucleotide II-II does not comprise a third region located between the first and second regions.

[0212] A new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled by chip sequence) can be obtained by annealing or hybridizing bridging oligonucleotide II-I and bridging oligonucleotide II-II with chip sequence and the cDNA strand obtained in step I above, and using DNA ligase to ligate the nucleic acid molecule hybridized with the first region and the second region of the same bridging oligonucleotide II-I, and / or ligate the nucleic acid molecule hybridized with the first region and the second region of the same bridging oligonucleotide II-II, and performing an extension reaction in the presence of DNA polymerase. The ligation process and the extension reaction can be performed in any order.

[0213] An exemplary structure of a novel nucleic acid molecule containing chip sequence information formed by the above exemplary embodiment includes a nucleic acid strand and / or its complementary nucleic acid strand, the nucleic acid strand including, from 5' to 3', consensus sequence A, a cDNA sequence, a 3'-end overhang sequence, a complementary sequence of the UMI sequence, a complementary sequence of the consensus sequence B, a sequence of bridging oligonucleotide II-I, a complementary sequence of tag sequence Y, and a complementary sequence of consensus sequence X1.

[0214] In certain embodiments, the method comprises steps (1), (2)(ii) and (3). In certain embodiments, in step (2)(ii)(b), the first extension product comprises, from 5' to 3', consensus sequence A, a cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, a 3'-end overhang sequence, optionally a complement of tag sequence B and a complement of consensus sequence B.

[0215] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B' or primer B" as described above, where primer B" is capable of annealing to the complementary sequence of consensus sequence B or a subsequence thereof to initiate the extension reaction.

[0216] In certain embodiments, in step (2)(ii)(c), a second extension product is formed 5' to 3' by an extension reaction primed by the extension primer and comprises a sequence complementary to the cDNA sequence and a complement of consensus sequence A.

[0217] An embodiment including steps (1), (2)(ii) and (3)(i) In certain embodiments, the method comprises steps (1), (2)(ii) and (3)(i), wherein the consensus sequence X2 or a subsequence thereof is capable of annealing to a complementary sequence of the consensus sequence A or a subsequence thereof, and the extension product obtained in step (3)(i) is a labeled nucleic acid molecule, comprising a first strand comprising the sequence of the first nucleic acid molecule to be labeled and / or a second strand comprising the sequence of the oligonucleotide probe.

[0218] It is easy to understand that the consensus sequence X2 can anneal to the complementary sequence of the consensus sequence A or a partial segment thereof and the entire nucleotide sequence thereof, and that the consensus sequence X2 can also anneal to the complementary sequence of the consensus sequence A or a partial segment thereof and the nucleotide sequence thereof.

[0219] In certain embodiments, the first strand comprises, from 5' to 3', the sequence of the first nucleic acid molecule to be labeled, the complement of tag sequence Y, and the complement of consensus sequence X1.

[0220] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, and a cDNA sequence that is complementary to the sequence of the first nucleic acid molecule to be labeled.

[0221] An embodiment comprising steps (1), (2)(ii) and (3)(i) for producing the first strand In certain embodiments, the consensus sequence X2 or a subsequence thereof is capable of annealing to a complementary sequence of the consensus sequence A or a subsequence thereof (e.g., a subsequence at the 3' end), and the extension product obtained in step (3)(i) is a labeled nucleic acid molecule and comprises a first strand comprising the sequence of the first nucleic acid molecule to be labeled.

[0222] In certain embodiments, in step (3)(i), the oligonucleotide probe is incapable of initiating an extension reaction (eg, the 3' end of the oligonucleotide probe is blocked).

[0223] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a random oligonucleotide sequence.

[0224] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, and a complement of consensus sequence A. In certain embodiments, the first strand comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, a complement of consensus sequence A, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0225] In some embodiments, in step (3), the first strand from each copy of the oligonucleotide probe bound to the same microdot has a different complementary sequence of capture sequence A as a UMI.

[0226] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0227] In certain embodiments, primer II-A' further comprises a tag sequence A, eg, a random oligonucleotide sequence.

[0228] In certain embodiments, capture sequence A is located at the 3' end of primer II-A'.

[0229] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, a complement of tag sequence A, and a complement of consensus sequence A. In certain embodiments, the first strand comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, a complement of tag sequence A, a complement of consensus sequence A, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0230] In some embodiments, in step (3), the first strand from each copy of the oligonucleotide probe bound to the same microdot has a different complementary sequence of tag sequence A as a UMI.

[0231] An embodiment comprising steps (1), (2)(ii), and (3)(i) for producing the second strand. In some embodiments, the consensus sequence X2 or a subsequence thereof (e.g., a subsequence at its 3' end) is capable of annealing to a complementary sequence of the consensus sequence A or a subsequence thereof, and the extension product obtained in step (3)(i) is a labeled nucleic acid molecule and comprises a second strand comprising the sequence of the oligonucleotide probe.

[0232] In certain embodiments, the second extension product obtained in step (2)(ii) is incapable of initiating an extension reaction (e.g., the 3' end of the second extension product obtained in step (2)(ii) is blocked).

[0233] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a random oligonucleotide sequence.

[0234] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, and a complement of consensus sequence A. In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, a cDNA sequence complementary to the sequence of the first nucleic acid molecule to be labeled, a 3'-end overhang sequence, optionally a complement of tag sequence B, and a complement of consensus sequence B.

[0235] In certain embodiments, in step (3), the second strand from each copy of the oligonucleotide probe bound to the same microdot has a different capture sequence A as a UMI.

[0236] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0237] In certain embodiments, primer II-A' further comprises a tag sequence A, eg, a random oligonucleotide sequence.

[0238] In certain embodiments, capture sequence A is located at the 3' end of primer II-A'.

[0239] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, a complement of tag sequence A, and a complement of consensus sequence A. In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, tag sequence A, a cDNA sequence complementary to the sequence of the first nucleic acid molecule to be labeled, a 3'-end overhang sequence, optionally a complement of tag sequence B, and a complement of consensus sequence B.

[0240] In certain embodiments, in step (3), the second strand from each copy of the oligonucleotide probe bound to the same microdot has a different tag sequence A as a UMI.

[0241] Exemplary embodiments of the present application including steps (1), (2)(ii) and (3)(i) are detailed below:

[0242] I. An exemplary embodiment using a sample's RNA (e.g., mRNA) as a template to prepare a complementary strand of a cDNA strand that contains a complementary sequence of a UMI proximal to its 3' end includes the following steps (as shown in FIG. 9):

[0243] (1) Using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyl transferase activity) and primer II-A', the RNA molecules (e.g., mRNA molecules) of the permeabilized sample are reverse transcribed to generate cDNA, and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. Various reverse transcriptases having terminal deoxynucleotidyl transferase activity can be used to perform the reverse transcription. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity.

[0244] In certain embodiments, primer II-A' comprises a poly(T) sequence, a UMI sequence, and a consensus sequence A (CA). Typically, the poly(T) sequence is located at the 3' end of primer II-A' to initiate reverse transcription, and the consensus sequence A is located upstream of the UMI sequence (e.g., 5' of the UMI sequence).

[0245] In certain embodiments, primer II-A' comprises a random oligonucleotide sequence and a consensus sequence A and can be used to capture RNA that does not have a poly-A tail. Typically, the random oligonucleotide sequence is located at the 3' end of primer II-A' to initiate reverse transcription.

[0246] (2) Primer II-B' is used to anneal or hybridize with the cDNA strand, where primer II-B' contains consensus sequence B(CB) and the complementary sequence of the 3'-end overhang of the cDNA. The nucleic acid fragment hybridized or annealed with primer II-B' can then be extended in the presence of a nucleic acid polymerase using consensus sequence B as a template to add the complementary sequence of consensus sequence B (c(CB)) to the 3' end of the cDNA strand, thereby generating a nucleic acid molecule bearing the complementary sequence of consensus sequence B at its 3' end.

[0247] Typically, a sequence complementary to the 3'-end overhang of the cDNA strand is located at the 3' end of primer II-B'.

[0248] For example, if the cDNA strand contains an overhang of three cytosine nucleotides at its 3' end, primer II-B' may contain GGG at its 3' end. Additionally, the nucleotides of primer II-B' may be modified (e.g., primer II-B' may be modified to contain one or more locked nucleic acids) to enhance the binding affinity for complementary pairing between primer II-B' and the 3'-end overhang of the cDNA strand.

[0249] Without being limited by any theory, the extension reaction can be carried out using a variety of suitable nucleic acid polymerases (e.g., DNA polymerases or reverse transcriptases) as long as the sequence of primer II-B' or the partial sequence thereof can be used as a template to extend the annealed or hybridized nucleic acid fragment (reverse transcription product). In certain exemplary embodiments, the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0250] In some embodiments, this step is carried out simultaneously (eg, in the same reaction system) with step (1).

[0251] In certain embodiments, the method optionally further comprises step (3): adding RNase H to digest the RNA strand of the RNA / cDNA hybrid to form a single-stranded cDNA.

[0252] In certain embodiments, the method does not include step (3).

[0253] (4) Using an extension primer, an extension reaction is carried out using the cDNA strand obtained in the previous step as a template to obtain an extension product, the extension primer being the above-mentioned primer II-B' or primer B", which is capable of annealing to the consensus sequence B or a partial sequence thereof to initiate the extension reaction.

[0254] An exemplary structure of the complementary strand of the cDNA strand prepared by the above exemplary embodiment includes consensus sequence B, a complementary sequence of the 3'-end overhang, a complementary sequence of the cDNA sequence, a complementary sequence of the UMI sequence, and a complementary sequence of consensus sequence A.

[0255] II. An exemplary embodiment of using the complementary sequence of an oligonucleotide probe (also called a chip sequence) to label the 3' end of the complementary strand of a cDNA strand to form a new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) comprises the following steps (as shown in FIG. 11): In some embodiments, the consensus sequence X2 of the chip sequence or a subsequence thereof can be annealed to the complementary sequence or a subsequence thereof of the consensus sequence A of the complementary strand of the cDNA strand obtained in the above step I. Thus, a new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled by the chip sequence) can be obtained by annealing or hybridizing the complementary strand of the cDNA strand with the chip sequence and performing an extension reaction in the presence of a polymerase.

[0256] An exemplary structure of a novel nucleic acid molecule containing chip sequence information formed by the above exemplary embodiment includes a nucleic acid strand and / or its complementary nucleic acid strand, the nucleic acid strand including, from 5' to 3', consensus sequence B, a complementary sequence of the 3'-end overhang, a complementary sequence of the cDNA sequence, a complementary sequence of the UMI sequence, a complementary sequence of consensus sequence A, a complementary sequence of tag sequence Y, and a complementary sequence of consensus sequence X1.

[0257] An embodiment including steps (1), (2)(ii) and (3)(ii) In certain embodiments, the method comprises steps (1), (2)(ii) and (3)(ii), wherein the second region of the bridging oligonucleotide II-II is capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence A of the second extension product obtained in step (2)(ii), and the reaction product obtained in step (3)(ii) is a labeled nucleic acid molecule, which comprises a first strand comprising the sequence of the first nucleic acid molecule to be labeled and / or a second strand comprising the sequence of the oligonucleotide probe.

[0258] It is easy to understand that the second region of the bridging oligonucleotide II-II is capable of annealing to the complementary sequence or a partial segment thereof of the consensus sequence A of the second extension product obtained in step (2)(ii).

[0259] In certain embodiments, the first strand comprises, from 5' to 3', the sequence of the first nucleic acid molecule to be labeled and, optionally, the complement of the third region of bridging oligonucleotide II-II, the sequence of bridging oligonucleotide II-I, the complement of tag sequence Y, and the complement of consensus sequence X1.

[0260] In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, optionally the complementary sequence of the third region of bridging oligonucleotide II-I, the sequence of bridging oligonucleotide II-II, and a cDNA sequence complementary to the sequence of the first nucleic acid molecule to be labeled.

[0261] An embodiment comprising steps (1), (2)(ii) and (3)(ii) for producing the first strand In certain embodiments, the second region of the bridging oligonucleotide II-II is capable of annealing to the complementary sequence of the consensus sequence A of the second extension product obtained in step (2)(ii) or a subsequence thereof at its 3' end, and the second region of the bridging oligonucleotide II-I has a free 3' end.

[0262] In certain embodiments, the reaction product obtained in step (3)(ii) is a labeled nucleic acid molecule that comprises the first strand.

[0263] In certain embodiments, the second region of the bridging oligonucleotide II-I is located at the 3' end of the bridging oligonucleotide II-I.

[0264] In certain embodiments, the first region of the bridging oligonucleotide II-I is located at the 5' end of the bridging oligonucleotide II-I.

[0265] In certain embodiments, bridging oligonucleotide II-I does not comprise the third region and / or bridging oligonucleotide II-II does not comprise the third region.

[0266] In certain embodiments, bridging oligonucleotide II-I comprises a 5' phosphate at the 5' end.

[0267] In certain embodiments, bridging oligonucleotide II-I comprises a free --OH at the 3' end.

[0268] In certain embodiments, in step (3)(ii), the bridging oligonucleotide II-II is incapable of initiating an extension reaction (e.g., the 3' end of the bridging oligonucleotide II-II is blocked) and / or the oligonucleotide probe is incapable of initiating an extension reaction (i.e., the 3' end of the oligonucleotide probe is blocked).

[0269] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a random oligonucleotide sequence.

[0270] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and that is complementary to the RNA, and a complement of consensus sequence A. In certain embodiments, the first strand comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of a cDNA sequence formed by reverse transcription primed by primer II-A' and which is complementary to the RNA, a complement of consensus sequence A, optionally a complement of the third region of bridging oligonucleotide II-II, the sequence of bridging oligonucleotide II-I, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0271] In some embodiments, in step (3), the first strand from each copy of the oligonucleotide probe bound to the same microdot has a different complementary sequence of capture sequence A as a UMI.

[0272] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0273] In certain embodiments, primer II-A' further comprises a tag sequence A, eg, a random oligonucleotide sequence.

[0274] In certain embodiments, capture sequence A is located at the 3' end of primer II-A'.

[0275] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and that is complementary to the RNA, a complement of tag sequence A, and a complement of consensus sequence A. In certain embodiments, the first strand comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of a cDNA sequence formed by reverse transcription primed by primer II-A' and which is complementary to the RNA, a complement of tag sequence A, a complement of consensus sequence A, optionally a complement of the third region of bridging oligonucleotide II-II, a sequence of bridging oligonucleotide II-I, a complement of tag sequence Y, and a complement of consensus sequence X1.

[0276] In some embodiments, in step (3), the first strand from each copy of the oligonucleotide probe bound to the same microdot has a different complementary sequence of tag sequence A as a UMI.

[0277] It is easy to understand that in step (3)(ii), after the bridging oligonucleotide II-I and the bridging oligonucleotide II-II are annealed with the oligonucleotide probe and the first nucleic acid molecule to be labeled at the corresponding position of the oligonucleotide probe, the ligation reaction for ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I and / or the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II, and the extension reaction in step (3)(ii) can be performed in any order as long as a second nucleic acid molecule having a positioning tag is obtained.

[0278] For example, when the ligation reaction and the extension reaction are carried out in the same system, the first strand can be obtained by ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II, and extending the bridging oligonucleotide II-I by the extension reaction. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0279] For example, when the ligation reaction and the extension reaction are carried out in different systems, a ligation reaction followed by an extension reaction can be carried out and the first strand can be obtained in the following exemplary manner: (A) ligating a nucleic acid molecule hybridized with a first region and a nucleic acid molecule hybridized with a second region of the same bridging oligonucleotide II-II, and extending the bridging oligonucleotide II-I by an extension reaction to obtain a first strand, wherein the polymerase used for the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity. Or, (B) ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-I, and extending the labeled first nucleic acid molecule by an extension reaction to obtain a first strand, wherein the polymerase used for the extension reaction preferably has strand displacement activity or 5' to 3' exonuclease activity.

[0280] For example, when the ligation reaction and the extension reaction are carried out in different systems, an extension reaction is followed by a ligation reaction, and the first strand can be obtained by extending the bridging oligonucleotide II-I by an extension reaction, and then ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-II. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0281]

[0033] Embodiments including steps (1), (2)(ii), and (3)(ii) for producing the second strand In certain embodiments, the second region of the bridging oligonucleotide II-II is capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence A of the second extension product obtained in step (2)(ii), and the second region of the bridging oligonucleotide II-II has a free 3' end.

[0282] In certain embodiments, the reaction product obtained in step (3)(ii) is a labeled nucleic acid molecule comprising the second strand.

[0283] In certain embodiments, the second region of the bridging oligonucleotide II-II is located at the 3' end of the bridging oligonucleotide II-II.

[0284] In certain embodiments, the first region of the bridging oligonucleotide II-II is located at the 5' end of the bridging oligonucleotide II-II.

[0285] In certain embodiments, bridging oligonucleotide II-I does not comprise the third region and / or bridging oligonucleotide II-II does not comprise the third region.

[0286] In certain embodiments, bridging oligonucleotide II-II comprises a 5' phosphate at the 5' end.

[0287] In certain embodiments, bridging oligonucleotide II-II comprises a free --OH at the 3' end.

[0288] In certain embodiments, in step (3)(ii), the bridging oligonucleotide II-I is incapable of initiating an extension reaction (e.g., the 3' end of the bridging oligonucleotide II-I is blocked), and / or the second extension product obtained in step (2)(ii) is incapable of initiating an extension reaction (e.g., the 3' end of the second extension product obtained in step (2)(ii) is blocked).

[0289] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a random oligonucleotide sequence.

[0290] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, and a complement of consensus sequence A. In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, optionally a complement of the third region of bridging oligonucleotide II-I, the sequence of bridging oligonucleotide II-II, a cDNA sequence complementary to the sequence of the first nucleic acid molecule to be labeled, a 3'-end overhang sequence, optionally a complement of tag sequence B, and a complement of consensus sequence B.

[0291] In certain embodiments, in step (3), the second strand from each copy of the oligonucleotide probe bound to the same microdot has a different capture sequence A as a UMI.

[0292] In certain embodiments, in step (2)(ii)(a), capture sequence A of primer II-A' is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0293] In certain embodiments, primer II-A' further comprises a tag sequence A, eg, a random oligonucleotide sequence.

[0294] In certain embodiments, capture sequence A is located at the 3' end of primer II-A'.

[0295] In certain embodiments, in step (2)(ii)(c), the extension primer is primer II-B'. In certain embodiments, in step (2)(ii)(c), the second extension product comprises, from 5' to 3', consensus sequence B, optionally tag sequence B, a complement of the 3'-end overhang sequence, a complement of the cDNA sequence formed by reverse transcription primed by primer II-A' and complementary to the RNA, a complement of tag sequence A, and a complement of consensus sequence A. In certain embodiments, the second strand comprises, from 5' to 3', consensus sequence X1, tag sequence Y, consensus sequence X2, optionally a complement of the third region of bridging oligonucleotide II-I, the sequence of bridging oligonucleotide II-II, tag sequence A, a cDNA sequence complementary to the sequence of the first nucleic acid molecule to be labeled, a 3'-end overhang sequence, optionally a complement of tag sequence B, and a complement of consensus sequence B.

[0296] In certain embodiments, in step (3), the second strand from each copy of the oligonucleotide probe bound to the same microdot has a different tag sequence A as a UMI.

[0297] It is easy to understand that in step (3)(ii), after the bridging oligonucleotide II-I and the bridging oligonucleotide II-II are annealed with the oligonucleotide probe and the first nucleic acid molecule to be labeled at the corresponding position of the oligonucleotide probe, the ligation reaction for ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I and / or the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II, and the extension reaction in step (3)(ii) can be performed in any order as long as a second nucleic acid molecule having a positioning tag is obtained.

[0298] For example, when the ligation reaction and the extension reaction are carried out in the same system, the second strand can be obtained by ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I, and extending the bridging oligonucleotide II-II by an extension reaction. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0299] For example, when the ligation reaction and the extension reaction are carried out in different systems, a ligation reaction followed by an extension reaction can be carried out and the second strand can be obtained in the following exemplary manner: (A) ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I, and extending the bridging oligonucleotide II-II by an extension reaction to obtain a second strand, wherein the polymerase used for the extension reaction preferably has or does not have strand displacement activity or 5' to 3' exonuclease activity, Or, (B) ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-II, and extending the oligonucleotide probe by an extension reaction to obtain a second strand, wherein the polymerase used for the extension reaction preferably has strand displacement activity or 5' to 3' exonuclease activity.

[0300] For example, when the ligation reaction and the extension reaction are carried out in different systems, an extension reaction is followed by a ligation reaction, and the second strand can be obtained by extending the bridging oligonucleotide II-II by an extension reaction, and then ligating the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide II-I. In this case, the polymerase used in the extension reaction preferably does not have strand displacement activity or 5' to 3' exonuclease activity.

[0301] An exemplary embodiment of the present application, including step (1), step (2)(ii) and step (3)(ii), is described in detail as follows:

[0302] I. An exemplary embodiment for preparing a complementary strand of a cDNA strand using RNA (e.g., mRNA) of a sample as a template includes the following steps (as shown in FIG. 9):

[0303] (1) Using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyl transferase activity) and primer II-A', the RNA molecules (e.g., mRNA molecules) of the permeabilized sample are reverse transcribed to generate cDNA, and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. Various reverse transcriptases having terminal deoxynucleotidyl transferase activity can be used to perform the reverse transcription. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity.

[0304] In certain embodiments, primer II-A' comprises a poly(T) sequence, a UMI sequence, and a consensus sequence A (CA). Typically, the poly(T) sequence is located at the 3' end of primer II-A' to initiate reverse transcription, and the consensus sequence A is located upstream of the UMI sequence (e.g., 5' of the UMI sequence).

[0305] In certain embodiments, primer II-A' comprises a random oligonucleotide sequence and a consensus sequence A and can be used to capture RNA that does not have a polyA tail. Typically, the random oligonucleotide sequence is located at the 3' end of primer II-A' to initiate reverse transcription.

[0306] (2) Primer II-B' is used to anneal or hybridize with the cDNA strand, where primer II-B' contains consensus sequence B(CB) and the complementary sequence of the 3'-end overhang of the cDNA. The nucleic acid fragment hybridized or annealed with primer II-B' can then be extended in the presence of a nucleic acid polymerase using consensus sequence B as a template to add the complementary sequence of consensus sequence B (c(CB)) to the 3' end of the cDNA strand, thereby generating a nucleic acid molecule that retains the complementary sequence of consensus sequence B at its 3' end.

[0307] Typically, a sequence complementary to the 3'-end overhang of the cDNA strand is located at the 3' end of primer II-B'.

[0308] For example, if the cDNA strand contains an overhang of three cytosine nucleotides at its 3' end, primer II-B' may contain GGG at its 3' end. Additionally, the nucleotides of primer II-B' may be modified (e.g., primer II-B' may be modified to contain one or more locked nucleic acids) to enhance the binding affinity for complementary pairing between primer II-B' and the 3'-end overhang of the cDNA strand.

[0309] Without being limited by any theory, the extension reaction can be carried out using a variety of suitable nucleic acid polymerases (e.g., DNA polymerases or reverse transcriptases) as long as the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the sequence of primer II-B' or a subsequence thereof as a template. In certain exemplary embodiments, the annealed or hybridized nucleic acid fragment (reverse transcription product) can be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0310] In some embodiments, this step is carried out simultaneously (eg, in the same reaction system) with step (1).

[0311] In certain embodiments, the method optionally further comprises step (3): adding RNase H to digest the RNA strand of the RNA / cDNA hybrid to form a single-stranded cDNA.

[0312] In certain embodiments, the method does not include step (3).

[0313] (4) Using an extension primer, an extension reaction is carried out using the cDNA strand obtained in the previous step as a template to obtain an extension product, the extension primer being primer II-B' or primer B" as described above, and primer B" is capable of annealing to consensus sequence B or a partial sequence thereof to initiate the extension reaction.

[0314] An exemplary structure of the complementary strand of the cDNA strand prepared by the above exemplary embodiment includes consensus sequence B, a complementary sequence of the 3'-end overhang, a complementary sequence of the cDNA sequence, a complementary sequence of the UMI sequence, and a complementary sequence of consensus sequence A.

[0315] II. An exemplary embodiment of using the complementary sequence of an oligonucleotide probe (also called a chip sequence) to label the 3' end of the complementary strand of a cDNA strand to form a new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled with a chip sequence) comprises the following steps (as shown in FIG. 10): providing a bridging oligonucleotide pair consisting of a bridging oligonucleotide II-I and a bridging oligonucleotide II-II, wherein the bridging oligonucleotide II-I and the bridging oligonucleotide II-II each independently comprise a first region (P1) and a second region (P2), the first region being located upstream of the second region (e.g., located 5' of the second region); a first region of the bridging oligonucleotide II-I capable of annealing to a first region of the bridging oligonucleotide II-II, and a second region of the bridging oligonucleotide II-I capable of annealing to a consensus sequence X2 of the oligonucleotide probe or a subsequence thereof; A step in which the second region of the bridging oligonucleotide II-II is capable of annealing to the complementary sequence or a subsequence thereof of the consensus sequence A of the complementary strand of the cDNA strand obtained in the above step I.

[0316] In certain embodiments, the bridging oligonucleotide II-I comprises an intermediate nucleotide sequence between the first and second regions, e.g., an intermediate nucleotide sequence of 1 nt to 5 nt or 5 nt to 10 nt, i.e., the bridging oligonucleotide II-I comprises a third region located between the first and second regions. In certain preferred embodiments, the first and second regions of the bridging oligonucleotide II-I are directly adjacent with no extra nucleotides between them, i.e., the bridging oligonucleotide II-I does not comprise a third region located between the first and second regions.

[0317] In certain embodiments, the bridging oligonucleotide II-II comprises an intermediate nucleotide sequence between the first and second regions, e.g., an intermediate nucleotide sequence of 1 nt to 5 nt or 5 nt to 10 nt, i.e., the bridging oligonucleotide II-II comprises a third region located between the first and second regions. In certain preferred embodiments, the first and second regions of the bridging oligonucleotide II-II are directly adjacent with no extra nucleotides between them, i.e., the bridging oligonucleotide II-II does not comprise a third region located between the first and second regions.

[0318] A new nucleic acid molecule containing chip sequence information (i.e., a nucleic acid molecule labeled by chip sequence) can be obtained by annealing or hybridizing bridging oligonucleotide II-I and bridging oligonucleotide II-II with the chip sequence and the complementary strand of the cDNA strand obtained in step I above, and using DNA ligase to ligate the nucleic acid molecule hybridized with the first region and the second region of the same bridging oligonucleotide II-I, and / or ligate the nucleic acid molecule hybridized with the first region and the second region of the same bridging oligonucleotide II-II, and performing an extension reaction in the presence of DNA polymerase. The ligation process and the extension reaction can be performed in any order.

[0319] An exemplary structure of a novel nucleic acid molecule containing chip sequence information formed by the above exemplary embodiment includes a nucleic acid strand and / or its complementary nucleic acid strand, the nucleic acid strand including, from 5' to 3', consensus sequence B, a complementary sequence of the 3'-end overhang, a complementary sequence of the cDNA sequence, a complementary sequence of the UMI sequence, a complementary sequence of the consensus sequence A, a sequence of bridging oligonucleotide II-I, a complementary sequence of tag sequence Y, and a complementary sequence of consensus sequence X1.

[0320] In certain embodiments, in step (2)(i)(b), the cDNA strand is annealed with primer II-B via its 3'-end overhang, and the cDNA strand is extended in the presence of a nucleic acid polymerase (e.g., a DNA polymerase or a reverse transcriptase) using primer II-B as a template to generate a first extension product.

[0321] In certain embodiments, in step (2)(ii)(b), the cDNA strand anneals with primer II-B via its 3'-end overhang, and the cDNA strand is extended in the presence of a nucleic acid polymerase (e.g., a DNA polymerase or a reverse transcriptase) using primer II-B' as a template to generate a first extension product.

[0322] In certain embodiments, the 3'-overhang has a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides. In certain embodiments, the 3'-overhang is a 3'-overhang of 2-5 cytosine nucleotides (e.g., a CCC overhang).

[0323] In certain embodiments of Scheme I or Scheme II, in step (2), the pretreatment is carried out intracellularly.

[0324] In certain embodiments of Scheme I or Scheme II, RNA (e.g., mRNA) of one or more cells is subjected to pretreatment before or after contacting the one or more cells with the solid support of a nucleic acid array to generate a first population of nucleic acid molecules.

[0325] In certain embodiments of Scheme I or Scheme II, prior to pretreatment, the cells are permeabilized.

[0326] In certain embodiments of Scheme I or Scheme II, in step (2), the pretreatment is carried out extracellularly.

[0327] In certain embodiments of Scheme I or Scheme II, after contacting one or more cells with the solid support of a nucleic acid array, RNA (e.g., mRNA) of the one or more cells is subjected to pre-treatment to generate a first population of nucleic acid molecules.

[0328] In certain embodiments of Scheme I or Scheme II, prior to performing the pretreatment, the method further comprises releasing intracellular RNA (e.g., mRNA), preferably, the intracellular RNA (e.g., mRNA) is released by cell permeabilization or cell lysis treatment.

[0329] In certain embodiments of Scheme I or Scheme II, the reverse transcription in step (2) is carried out by using a reverse transcriptase.

[0330] In certain embodiments of Scheme I or Scheme II, the reverse transcriptase has terminal deoxynucleotidyl transferase activity.

[0331] In certain embodiments of Scheme I or Scheme II, the reverse transcriptase is capable of synthesizing a cDNA strand using an RNA (e.g., an mRNA) as a template and adding an overhang to the 3' end of the cDNA strand.

[0332] In certain embodiments of Scheme I or Scheme II, the reverse transcriptase is capable of adding an overhang having a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides to the 3' end of the cDNA strand.

[0333] In certain embodiments of Scheme I or Scheme II, the reverse transcriptase is capable of adding an overhang of 2-5 cytosine nucleotides (eg, a CCC overhang) to the 3' end of the cDNA strand.

[0334] In certain embodiments of Scheme I or Scheme II, the reverse transcriptase is selected from the group consisting of M-MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and variants, modified products, and derivatives thereof that have the reverse transcription activity of the above reverse transcriptases.

[0335] In certain embodiments of Scheme I or Scheme II, step (2) and step (3) have one or more features selected from the following: (1) Primer IA, primer II-A, primer I-A', primer II-A', primer IB, primer II-B, primer II-B', bridging oligonucleotide I, bridging oligonucleotide II-I, and bridging oligonucleotide II-II each independently comprise or consist of natural nucleotides (e.g., deoxyribonucleotides or ribonucleotides), modified nucleotides, non-natural nucleotides, or any combination thereof. In certain embodiments, primer IA, primer II-A, primer I-A', and primer II-A' are capable of initiating an extension reaction; (2) Primer IB, primer II-B, and primer II-B' each independently contain a modified nucleotide (e.g., a locked nucleic acid). In certain embodiments, primer IB, primer II-B, and primer II-B' each independently contain one or more modified nucleotides (e.g., one or more locked nucleic acids) at their 3' ends. (3) tag sequence A and tag sequence B each independently have a length of 5 to 200 nt (e.g., 5 to 30 nt, 6 to 15 nt); (4) The consensus sequence A and the consensus sequence B each independently have a length of 10 to 200 nt (e.g., 10 to 100 nt, 20 to 100 nt, 25 to 100 nt, 5 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 50 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt), (5) primer IA, primer II-A, primer I-A', primer II-A', primer IB, primer II-B and primer II-B' each independently have a length of 4 to 200 nt (e.g., 5 to 200 nt, 15 to 230 nt, 26 to 115 nt, 10 to 130 nt, 10 to 20 nt, 20 to 50 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt); (6) The first region and the second region of the bridging oligonucleotide I, the bridging oligonucleotide II-I, and the bridging oligonucleotide II-II each independently have a length of 3 to 100 nt (e.g., 20 to 100 nt, 3 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 70 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, or 50 to 100 nt). (7) The third regions of the bridging oligonucleotide I, the bridging oligonucleotide II-I, and the bridging oligonucleotide II-II each independently have a length of 0 to 50 nt (e.g., 0 nt, 0 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt), (8) The bridging oligonucleotide I, the bridging oligonucleotide II-I, and the bridging oligonucleotide II-II each independently have a length of 6 to 200 nt (e.g., 20 to 100 nt, 20 to 70 nt, 6 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt), (9) The poly(T) sequence contains at least 5 or at least 20 (e.g., 6 to 100, 10 to 50) deoxythymidine residues; (10) The random oligonucleotide sequence has a length of 5 to 200 (e.g., 5 nt, 5 to 30 nt, 6 to 15 nt).

[0336] In certain embodiments of Scheme I or Scheme II, the method further comprises the step of (4) recovering and purifying the second population of nucleic acid molecules.

[0337] In certain embodiments of Scheme I or Scheme II, the resulting second population of nucleic acid molecules and / or their complements are used to construct a transcriptome library or for transcriptome sequencing.

[0338] In certain embodiments of Scheme I or Scheme II, the oligonucleotide probe in step (1) has one or more characteristics selected from the following: (1) The consensus sequence X1, the tag sequence Y, and the consensus sequence X2 each independently comprise or consist of natural nucleotides (e.g., deoxyribonucleotides or ribonucleotides), modified nucleotides, non-natural nucleotides (e.g., peptide nucleic acid (PNA) or locked nucleic acid), or any combination thereof; (2) The consensus sequence X1, the tag sequence Y, and the consensus sequence X2 each independently have a length of 2 to 200 nt (e.g., 10 to 200 nt, 25 to 100 nt, 10 to 30 nt, 10 to 100 nt, 5 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt).

[0339] In certain embodiments of Scheme I or Scheme II, the nucleic acid array of step (1) is provided by the following steps: (1) providing a plurality of types of carrier sequences, each type of carrier sequence comprising at least one copy (e.g., multiple copies) of a carrier sequence, the carrier sequence comprising, in a 5' to 3' direction, a complement of a consensus sequence X2, a complement of a tag sequence Y, and an immobilization sequence, each type of carrier sequence having a different complement of the tag sequence Y; (2) binding a plurality of types of carrier sequences to the surface of a solid support (e.g., a chip); (3) providing an immobilized primer and performing a primer extension reaction using the carrier sequence as a template to generate an extension product to obtain an oligonucleotide probe, wherein the immobilized primer comprises a sequence of a consensus sequence X1 and is capable of annealing to the immobilized sequence of the carrier sequence and initiating an extension reaction, and in some embodiments, the extension product comprises or consists of, in the 5' to 3' direction, the consensus sequence X1, the tag sequence Y, and the consensus sequence X2; (4) coupling the immobilized primer to the surface of a solid support, wherein steps (3) and (4) are performed in any order; (5) Optionally, the immobilized sequence of the carrier sequence further comprises a cleavage site, and the cleavage may be selected from nicking enzyme digestion, USER enzyme digestion, light-responsive excision, chemical excision or CRISPR-mediated excision, and cleavage is performed at the cleavage site contained in the immobilized sequence of the carrier sequence to digest the carrier sequence and separate the extension product in step (3) from the template (i.e., carrier sequence) on which the extension product is generated, thereby linking the oligonucleotide probe to the surface of a solid support (e.g., chip). In certain embodiments, the method further comprises separating the extension product in step (3) from the template (i.e., carrier sequence) on which the extension product is generated by high temperature denaturation.

[0340] In certain embodiments of Scheme I or Scheme II, the various carrier sequences are DNBs formed from concatemers of multiple copies of the carrier sequences.

[0341] In certain embodiments of Scheme I or Scheme II, multiple types of carrier sequences are prepared in step (1) by the steps of: (i) providing a plurality of carrier-template sequences, each of which comprises a complementary sequence of a carrier sequence; (ii) performing a nucleic acid amplification reaction using the various carrier-template sequences as templates to obtain amplification products of the various carrier-template sequences, the amplification products comprising at least one copy of the carrier sequence, and in certain embodiments, rolling circle replication is performed to obtain DNBs formed from concatemers of the carrier sequences; Offered by.

[0342] Scheme III

[0343] In certain embodiments, in step (1) of the method, the consensus sequence X2 comprises a capture sequence, the capture sequence is capable of hybridizing to all or a portion of the nucleic acid to be captured, the capture sequence comprises a poly(T) sequence or a specific sequence or a random oligonucleotide sequence that targets the target nucleic acid, and the capture sequence has a free 3' end, such that the consensus sequence X2 can act as an extension primer.

[0344] In such embodiments, step (2) comprises contacting one or more cells with the solid support of the nucleic acid array, whereby each cell individually occupies at least one microdot of the nucleic acid array (i.e., each cell contacts at least one microdot of the nucleic acid array) and allows a first binding molecule of the cell to interact with a first label molecule of the solid support, and applying annealing conditions to anneal the nucleic acid of the one or more cells with the capture sequence, whereby the location of the nucleic acid is mapped to the location of the oligonucleotide probe of the nucleic acid array; Further, step (3) includes performing a primer extension reaction using the oligonucleotide probe as a primer and the captured nucleic acid molecule as a template under conditions that allow primer extension to generate a labeled nucleic acid molecule (e.g., a nucleic acid molecule labeled with tag sequence Y), and / or performing a primer extension reaction using the captured nucleic acid molecule as a primer and the oligonucleotide probe as a template to generate an extended captured nucleic acid molecule, thereby forming a labeled nucleic acid molecule (e.g., a nucleic acid molecule labeled with the complement of tag sequence Y).

[0345] In certain embodiments, the oligonucleotide probe in step (1) further comprises a unique molecular identifier (UMI) sequence. Preferably, the UMI sequence is located upstream of the capture sequence. Preferably, the oligonucleotide probes bound to the same microdot comprise different UMI sequences.

[0346] In some embodiments, the nucleic acid array of step (1) is provided by the following steps: (1) providing a plurality of types of carrier sequences, each type of carrier sequence comprising a plurality of copies of a carrier sequence, the carrier sequence comprising, in a 5' to 3' direction, a positioning sequence and a first immobilization sequence; The positioning sequence is the complement of the tag sequence Y, The first immobilized sequence is allowed to anneal with its complementary nucleotide sequence and initiate an extension reaction; (2) binding a plurality of types of carrier sequences to the surface of a solid support (e.g., a chip); (3) providing a first primer and performing a primer extension reaction using the carrier sequence as a template to obtain a first nucleic acid molecule hybridized with the first immobilized sequence and the positioning sequence of the carrier sequence, thereby forming a duplex region with the first immobilized sequence and the positioning sequence of the carrier sequence, wherein the first nucleic acid molecule comprises, in a 5' to 3' direction, a complementary sequence of the first immobilized sequence and a complementary sequence of the positioning sequence, and the first primer comprises a region at its 3' end that is complementary to the first immobilized sequence, and the region of the first primer comprises the complementary sequence of the first immobilized sequence or a fragment thereof and has a free 3' end; (4) providing a second nucleic acid molecule, the second nucleic acid molecule comprising a consensus sequence X2 (i.e., a capture sequence) having a free 3' end such that the second nucleic acid molecule can be used as an extension primer; (5) ligating the second nucleic acid molecule to the first nucleic acid molecule (e.g., using a ligase to ligate the second nucleic acid molecule to the first nucleic acid molecule), wherein the ligation product is an oligonucleotide probe comprising, in the 5' to 3' direction, consensus sequence X1, tag sequence Y, and consensus sequence X2.

[0347] In certain embodiments, the carrier sequence is optionally digested such that the ligation product is separated from the carrier sequence in step (5), thereby linking the oligonucleotide probe to the surface of the solid support.

[0348] In certain embodiments, the first nucleic acid molecule or the second nucleic acid molecule further comprises a UMI sequence. In certain embodiments, the second nucleic acid molecule comprises a UMI sequence located 5' of the capture sequence. In certain embodiments, the plurality of carrier sequences is provided by the steps of: (i) providing a plurality of carrier-template sequences, the carrier-template sequences including complementary sequences of the carrier sequences; (ii) performing a nucleic acid amplification reaction using the various carrier-template sequences as templates to obtain amplification products of the various carrier-template sequences, the amplification products comprising multiple copies of the carrier sequences; Preferably, the amplification is selected from the group consisting of rolling circle replication (RCA), bridge PCR amplification, multiple displacement amplification (MDA) or emulsion PCR amplification, preferably rolling circle replication is performed to obtain DNBs formed by concatemers of carrier sequences, or bridge PCR amplification, emulsion PCR amplification or multiple displacement amplification is performed to obtain DNA clusters formed by clonal populations of carrier sequences.

[0349] In certain embodiments of Scheme I, Scheme II, or Scheme III, the oligonucleotide probe is attached to the solid support via a linker.

[0350] In certain embodiments of Scheme I, Scheme II, or Scheme III, the linker is a linking group capable of binding to an activating group, and the surface of the solid support is modified with an activating group.

[0351] In certain embodiments of Scheme I, Scheme II, or Scheme III, the linker comprises -SH, -DBCO, or -NHS.

[0352] In certain embodiments of Scheme I, Scheme II, or Scheme III, the linker is -DBCO and the surface of the solid support is [ka] (Azido-dPEG® 8-NHS ester).

[0353] In some embodiments of Scheme I, Scheme II, or Scheme III, the nucleic acid array of step (1) has one or more features selected from the following: (1) The oligonucleotide probes bound to the same solid support have the same consensus sequence X1 and / or the same consensus sequence X2; (2) The consensus sequence X1 of the oligonucleotide probe comprises a cleavage site, and in some embodiments, the cleavage site can be cleaved or destroyed by nicking enzyme digestion, USER enzyme digestion, light-responsive excision, chemical excision, or CRISPR-mediated excision.

[0354] In certain embodiments of Scheme I, Scheme II, or Scheme III, the solid support in step (1) has one or more characteristics selected from the following: (1) the solid support is selected from the group consisting of latex beads, dextran beads, polystyrene surfaces, polypropylene surfaces, polyacrylamide gels, gold surfaces, glass surfaces, chips, sensors, electrodes, and silicon wafers, and in some embodiments the solid support is a chip; (2) The solid support is planar, spherical or porous; (3) The solid support can be used as a sequencing platform, e.g., a sequencing chip. In some embodiments, the solid support is a sequencing chip for an Illumina, MGI, or Thermo Fisher sequencing platform; and (4) The solid support is capable of releasing the oligonucleotide probes spontaneously or upon exposure to one or more stimuli (e.g., temperature change, pH change, exposure to certain chemicals or phases, exposure to light, exposure to a reducing agent, etc.).

[0355] Methods for constructing libraries of nucleic acid molecules In another aspect, the present application also provides a method for constructing a library of nucleic acid molecules, comprising: (a) generating a population of labeled nucleic acid molecules according to a method as described above; (b) randomly fragmenting nucleic acid molecules in the population of labeled nucleic acid molecules and ligating adapters thereto; and (c) optionally amplifying and / or enriching the product of step (b) thereby obtaining a library of nucleic acid molecules.

[0356] In certain embodiments, the library of nucleic acid molecules comprises nucleic acid molecules from a plurality of single cells, wherein the nucleic acid molecules of different single cells have different tag sequences Y.

[0357] In certain embodiments, the library of nucleic acid molecules is used for sequencing, e.g., transcriptome sequencing, e.g., single cell transcriptome sequencing (e.g., 5' or 3' transcriptome sequencing).

[0358] In certain embodiments, prior to performing step (b), the method further comprises the step (pre-b): amplifying and / or enriching the population of labeled nucleic acid molecules.

[0359] In certain embodiments, in step (pre-b), the population of labeled nucleic acid molecules is subjected to a nucleic acid amplification reaction to generate an amplification product.

[0360] In certain embodiments, the amplification reaction is carried out using at least primer C and / or primer D, where primer C is capable of hybridizing or annealing to a complementary sequence of consensus sequence X1 or a subsequence thereof to initiate an extension reaction, and primer D is capable of hybridizing or annealing to a nucleic acid molecule strand comprising tag sequence Y in a population of labeled nucleic acid molecules to initiate an extension reaction.

[0361] In certain embodiments, the nucleic acid amplification reaction in step (pre-b) is carried out by using a nucleic acid polymerase (e.g., a DNA polymerase, e.g., a DNA polymerase with strand displacement activity and / or high fidelity).

[0362] In certain embodiments, in step (b) of the method, the nucleic acid molecule is randomly fragmented and the resulting fragments are ligated with adapters by using a transposase.

[0363] In some embodiments, in step (b) of the method, the nucleic acid molecule obtained in the previous step is randomly fragmented and the resulting fragments are ligated at both ends with an adapter and a second adapter, respectively, by using a transposase.

[0364] In certain embodiments, the transposase is selected from the group consisting of Tn5 transposase, MuA transposase, Sleeping Beauty transposase, Mariner transposase, Tn7 transposase, Tn10 transposase, Ty1 transposase, Tn552 transposase, and variants, modified products, and derivatives thereof that have the transposition activity of the above transposases.

[0365] In certain embodiments, the transposase is a Tn5 transposase.

[0366] In some embodiments, in step (c), the product of step (b) is amplified using at least primer C' and / or primer D', where primer C' is capable of hybridizing or annealing to a first adapter to initiate an extension reaction, and primer D' is capable of hybridizing or annealing to a second adapter to initiate an extension reaction.

[0367] In some embodiments, in step (c), the product of step (b) is amplified using at least primer C and / or primer D' as described above, where primer D' is capable of hybridizing or annealing to the first adaptor or the second adaptor to initiate the extension reaction.

[0368] How to Conduct Transcriptome Sequencing In another aspect, the present application also provides a method for transcriptome sequencing of cells in a sample, comprising: (1) constructing a library of nucleic acid molecules according to the method described above; and (2) sequencing the library of nucleic acid molecules. Also provided is a method comprising:

[0369] How to Perform Single-Cell Transcriptome Analysis In another aspect, the present application also provides a method of performing single cell transcriptome analysis, comprising: (1) performing transcriptome sequencing of a single cell in a sample according to the above method; and (2) A step of analyzing the sequencing data, comprising matching the sequencing results of the sequencing library with the tag sequence Y or its complementary sequence of the oligonucleotide probe bound to each microdot of the nucleic acid array, whereby the microdot is identified as a positive microdot if the matching is successful, and the sequencing data derived from the positive microdots having regional continuity in the nucleic acid array are identified as the transcription data of the same cell, thereby performing single-cell transcriptome analysis. Also provided is a method comprising:

[0370] kit In another aspect, the present application also provides a kit, the kit comprising: a nucleic acid array for labeling nucleic acids and optionally a first binding molecule, the nucleic acid array comprising a solid support, the solid support comprising (e.g., on its surface) a first label molecule, the first binding molecule being capable of forming an interactive pair with the first label molecule; The solid support further comprises a plurality of microdots, the size (e.g., equivalent diameter) of the microdots being less than 5 μm and the center-to-center distance between adjacent microdots being less than 10 μm, each microdot being bound to one type of oligonucleotide probe, each type of oligonucleotide probe comprising at least one copy, the oligonucleotide probe comprising or consisting of, in the 5' to 3' direction, a consensus sequence X1, a tag sequence Y and a consensus sequence X2; The oligonucleotide probes bound to the different microdots have different tag sequences Y.

[0371] In certain embodiments, the center-to-center distance between adjacent microdots is less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, or less than 0.01 μm, and the size (e.g., equivalent diameter) of the microdots is less than 5 μm, less than 1 μm, less than 0.3 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, less than 0.01 μm, or less than 0.001 μm.

[0372] Preferably, the center-to-center distance between adjacent microdots is from 0.5 μm to 1 μm, for example, from 0.5 μm to 0.9 μm, for example, from 0.5 μm to 0.8 μm.

[0373] Preferably, the size (eg, equivalent diameter) of the microdots is 0.001 μm to 0.5 μm (eg, 0.01 μm to 0.1 μm, 0.01 μm to 0.2 μm, 0.2 μm to 0.5 μm, 0.2 μm to 0.4 μm, 0.2 μm to 0.3 μm).

[0374] In certain embodiments, the solid support comprises a plurality of (e.g., at least 10, at least 10 2 , at least 10 3 , at least 104 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 In certain embodiments, the solid support comprises at least 10 4 (e.g., at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 Or at least 10 12 ) microdots / mm 2 Includes.

[0375] In certain embodiments, the first binding molecule is capable of forming a specific or non-specific interaction pair with the first labeled molecule.

[0376] In certain embodiments, the interaction pair is selected from the group consisting of positive and negative charge interaction pairs, affinity interaction pairs (e.g., biotin / avidin, biotin / streptavidin, antigen / antibody, receptor / ligand, enzyme / cofactor), pairs of molecules capable of undergoing click chemistry reactions (e.g., alkynyl-containing compound / azide compound), N-hydroxysulfosuccinate (NHS) ester / amino-containing compound, and any combination thereof.

[0377] For example, the first labeled molecule is polylysine and the first binding molecule is a protein capable of binding to polylysine; the first labeled molecule is an antibody and the first binding molecule is an antigen capable of binding to the antibody; the first labeled molecule is an amino-containing compound and the first binding molecule is an N-hydroxysulfosuccinate (NHS) ester; or the first labeled molecule is biotin and the first binding molecule is streptavidin.

[0378] In certain embodiments, the kit comprises: (i) a primer set comprising primer IA or primer I-A' and primer IB, or a primer set comprising primer IA and primer IB; Primer IA comprises a consensus sequence A and a capture sequence A, where the capture sequence A is capable of annealing to an RNA (e.g., an mRNA) to be captured to initiate an extension reaction, and preferably the consensus sequence A is located upstream of the capture sequence A (e.g., located at the 5' end of primer IA); Primer I-A' comprises a capture sequence A, which is capable of annealing to the RNA (e.g., mRNA) to be captured to initiate an extension reaction; primer IB comprises consensus sequence B, a complementary sequence of a 3'-end overhang, and optionally a tag sequence B, wherein the complementary sequence of the 3'-end overhang is located at the 3' end of primer IB, and consensus sequence B is located upstream of the complementary sequence of the 3'-end overhang (e.g., at the 5' end of primer IB), and the 3'-end overhang refers to one or more non-templated nucleotides contained at the 3' end of a cDNA strand generated by reverse transcription using the RNA captured by capture sequence A of primer I-A' as a template; (ii) further comprising a bridging oligonucleotide I, the bridging oligonucleotide I comprising a first region and a second region, and optionally a third region located between the first region and the second region, the first region being located upstream of the second region (e.g., located 5' of the second region); the first region is capable of annealing (a) to all or a portion of consensus sequence A of primer IA; or (b) to all or a portion of consensus sequence B of primer IB; The second region is capable of annealing to all or part of the consensus sequence X2.

[0379] In certain embodiments, the kit comprises a primer IA as described in (i) and a bridging oligonucleotide I as described in (ii), wherein a first region of the bridging oligonucleotide I is capable of annealing to all or a portion of the consensus sequence A of the primer IA and a second region of the bridging oligonucleotide I is capable of annealing to all or a portion of the consensus sequence X2; The capture sequence A of primer IA is a random oligonucleotide sequence, or the capture sequence A of primer IA is a poly(T) sequence or a specific sequence that targets the target nucleic acid, and primer IA further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0380] In certain embodiments, primer IA comprises a 5' phosphate at the 5' end.

[0381] In certain embodiments, the kit comprises a primer set comprising primer I-A' and primer IB as described in (i) and a bridging oligonucleotide I as described in (ii), wherein a first region of the bridging oligonucleotide I is capable of annealing to all or a portion of the consensus sequence B of primer IB and a second region of the bridging oligonucleotide I is capable of annealing to all or a portion of the consensus sequence X2; The capture sequence A of the primer I-A' is a random oligonucleotide sequence, or the capture sequence A of the primer I-A' is a poly(T) sequence or a specific sequence targeting the target nucleic acid, and the primer I-A' further comprises a tag sequence A and a consensus sequence A; Primer IB contains consensus sequence B, a complementary sequence of the 3'-end overhang and tag sequence B.

[0382] In certain embodiments, the kit further comprises a primer B″, which is capable of annealing to a complementary sequence of consensus sequence B or a subsequence thereof to initiate an extension reaction.

[0383] In certain embodiments, primer IB or primer B″ comprises a 5′ phosphate at the 5′ end.

[0384] In certain embodiments, primer IB comprises a modified nucleotide (eg, a locked nucleic acid), and preferably primer IB comprises one or more modified nucleotides (eg, one or more locked nucleic acids) at the 3' end.

[0385] In certain embodiments, the kit comprises a primer set comprising primer IA and primer IB as described in (i), and a bridging oligonucleotide I as described in (ii), wherein a first region of the bridging oligonucleotide I is capable of annealing to all or a portion of the consensus sequence A of primer IA, and a second region of the bridging oligonucleotide I is capable of annealing to all or a portion of the consensus sequence X2 of primer IA; The capture sequence A of primer IA is a random oligonucleotide sequence, or the capture sequence A of primer IA is a poly(T) sequence or a specific sequence that targets the target nucleic acid, and primer IA further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0386] In certain embodiments, primer IA comprises a 5' phosphate at the 5' end.

[0387] In certain embodiments, primer IB comprises a modified nucleotide (eg, a locked nucleic acid), and preferably primer IB comprises one or more modified nucleotides (eg, one or more locked nucleic acids) at the 3' end.

[0388] In certain embodiments, the kit comprises: (i) a primer set comprising primer II-A and primer II-B, or a primer set comprising primer II-A' and primer II-B'; Primer II-A comprises a capture sequence A, which is capable of annealing to the RNA (e.g., mRNA) to be captured to initiate an extension reaction; primer II-B comprises a consensus sequence B, a complementary sequence of a 3'-end overhang, and optionally a tag sequence B, the complementary sequence of the 3'-end overhang being located at the 3' end of primer II-B, and the consensus sequence B being located upstream of the complementary sequence of the 3'-end overhang (e.g., at the 5' end of primer II-B), the 3'-end overhang referring to one or more non-templated nucleotides contained at the 3' end of a cDNA strand generated by reverse transcription using the RNA captured by the capture sequence A of primer II-A as a template; Primer II-A' comprises consensus sequence A and capture sequence A, where capture sequence A is located at the 3' end of primer II-A' and consensus sequence A is located upstream of capture sequence A (e.g., at the 5' end of primer II-A'); Primer II-B' comprises consensus sequence B, a complementary sequence of the 3'-end overhang, and optionally, a tag sequence B, where the complementary sequence of the 3'-end overhang is located at the 3' end of primer II-B', and consensus sequence B is located upstream of the complementary sequence of the 3'-end overhang (e.g., at the 5' end of primer II-B'), and the 3'-end overhang refers to one or more non-templated nucleotides contained at the 3' end of a cDNA strand generated by reverse transcription using the RNA captured by capture sequence A of primer II-A' as a template.

[0389] In certain embodiments, the kit comprises: (i) a primer set comprising primer II-A and primer II-B as described above; and (ii) bridging oligonucleotides II-I and II-II, each of which independently comprises a first region and a second region, and optionally a third region located between the first region and the second region, wherein the first region is located upstream of the second region (e.g., 5′ of the second region); a first region of the bridging oligonucleotide II-I capable of annealing to a first region of the bridging oligonucleotide II-II, and a second region of the bridging oligonucleotide II-I capable of annealing to a consensus sequence X2 of the oligonucleotide probe or a subsequence thereof; a second region of bridging oligonucleotide II-II capable of annealing to a complementary sequence or a subsequence thereof of consensus sequence B of primer II-B; The capture sequence A of primer II-A is a random oligonucleotide sequence, or the capture sequence A of primer II-A is a poly(T) sequence or a specific sequence targeting a target nucleic acid, and primer II-A preferably further comprises a consensus sequence A and optionally a tag sequence A, e.g., a random oligonucleotide sequence; Primer II-B contains consensus sequence B, a complementary sequence of the 3'-end overhang and tag sequence B.

[0390] In certain embodiments, primer II-B comprises modified nucleotides (e.g., locked nucleic acids), and preferably primer II-B comprises one or more modified nucleotides (e.g., one or more locked nucleic acids) at the 3' end.

[0391] In certain embodiments, the kit comprises a primer set comprising primer II-A and primer II-B as described in (i), The capture sequence A of primer II-A is a random oligonucleotide sequence, or the capture sequence A of primer II-A is a poly(T) sequence or a specific sequence targeting a target nucleic acid, and primer II-A preferably further comprises a consensus sequence A and optionally a tag sequence A, e.g., a random oligonucleotide sequence; Primer II-B contains consensus sequence B, a complementary sequence of the 3'-end overhang and tag sequence B.

[0392] In certain embodiments, primer II-B comprises modified nucleotides (e.g., locked nucleic acids), and preferably primer II-B comprises one or more modified nucleotides (e.g., one or more locked nucleic acids) at the 3' end.

[0393] In certain embodiments, the kit comprises: (i) a primer set comprising primer II-A' and primer II-B' as described above; and (ii) bridging oligonucleotides II-I and II-II, each of which independently comprises a first region and a second region, and optionally a third region located between the first region and the second region, wherein the first region is located upstream of the second region (e.g., 5' of the second region); a first region of the bridging oligonucleotide II-I capable of annealing to a first region of the bridging oligonucleotide II-II, and a second region of the bridging oligonucleotide II-I capable of annealing to a consensus sequence X2 of the oligonucleotide probe or a subsequence thereof; a second region of the bridging oligonucleotide II-II capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence A of the primer II-A'; The capture sequence A of primer II-A' is a random oligonucleotide sequence, or the capture sequence A of primer II-A' is a poly(T) sequence or a specific sequence that targets the target nucleic acid, and primer II-A' further comprises a tag sequence A, e.g., a random oligonucleotide sequence.

[0394] In certain embodiments, primer II-B' comprises modified nucleotides (e.g., locked nucleic acids), and preferably primer II-B' comprises one or more modified nucleotides (e.g., one or more locked nucleic acids) at the 3' end.

[0395] In certain embodiments, the kit further comprises a primer B″, which is capable of annealing to a complementary sequence of consensus sequence B or a subsequence thereof to initiate an extension reaction. In certain embodiments, the kit comprises a primer set comprising primer II-A' and primer II-B' as described in (i), The capture sequence A of primer II-A' is a random oligonucleotide sequence, or the capture sequence A of primer II-A' is a poly(T) sequence or a specific sequence that targets the target nucleic acid, and primer II-A' further comprises a tag sequence A, e.g., a random oligonucleotide sequence; Primer II-B' contains the consensus sequence B, a complementary sequence of the 3'-end overhang and the tag sequence B.

[0396] In certain embodiments, primer II-B' comprises modified nucleotides (e.g., locked nucleic acids), and preferably primer II-B' comprises one or more modified nucleotides (e.g., one or more locked nucleic acids) at the 3' end.

[0397] In certain embodiments, the kit further comprises a primer B″, which is capable of annealing to a complementary sequence of consensus sequence B or a subsequence thereof to initiate an extension reaction.

[0398] In certain embodiments, the kit has one or more features selected from the following: (1) the oligonucleotide probes, primer IA, primer II-A, primer I-A', primer II-A', primer IB, primer II-B, primer II-B', primer B", bridging oligonucleotide I, bridging oligonucleotide II-I and bridging oligonucleotide II-II each independently comprise or consist of natural nucleotides (e.g., deoxyribonucleotides or ribonucleotides), modified nucleotides, non-natural nucleotides or any combination thereof; (2) The oligonucleotide probes each independently have a length of 15 to 300 nt (e.g., 15 to 200 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt), (3) primer IA, primer II-A, primer I-A', primer II-A', primer IB, primer II-B, primer II-B' and primer B" each independently have a length of 4 to 200 nt (e.g., 5 to 200 nt, 15 to 230 nt, 26 to 115 nt, 10 to 130 nt, 10 to 20 nt, 20 to 50 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt), (4) The bridging oligonucleotide I, the bridging oligonucleotide II-I, and the bridging oligonucleotide II-II each independently have a length of 6 to 200 nt (e.g., 20 to 100 nt, 20 to 70 nt, 6 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt), (5) The oligonucleotide probes bound to the same solid support have the same consensus sequence X1 and / or the same consensus sequence X2; (6) The consensus sequence X1 of the oligonucleotide probe comprises a cleavage site, and in some embodiments, the cleavage site can be cleaved or destroyed by nicking enzyme digestion, USER enzyme digestion, light-responsive excision, chemical excision, or CRISPR-mediated excision.

[0399] In certain embodiments, the kit further comprises a reverse transcriptase, a nucleic acid ligase, a nucleic acid polymerase, and / or a transposase.

[0400] In certain embodiments, the reverse transcriptase has terminal deoxynucleotidyl transferase activity. In certain embodiments, the reverse transcriptase is capable of synthesizing a cDNA strand using RNA (e.g., mRNA) as a template and adding a 3'-end overhang to the 3' end of the cDNA strand. In certain embodiments, the reverse transcriptase is capable of adding an overhang having a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides to the 3' end of the cDNA strand. In certain embodiments, the reverse transcriptase is capable of adding an overhang of 2-5 cytosine nucleotides (e.g., a CCC overhang) to the 3' end of the cDNA strand. In certain embodiments, the reverse transcriptase is selected from the group consisting of M-MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and variants, modified products, and derivatives thereof having the reverse transcription activity of the above reverse transcriptases.

[0401] In certain embodiments, the nucleic acid polymerase does not have 5' to 3' exonuclease activity or strand displacement activity.

[0402] In certain embodiments, the nucleic acid polymerase has 5' to 3' exonuclease activity or strand displacement activity.

[0403] In certain embodiments, the transposase is selected from the group consisting of Tn5 transposase, MuA transposase, Sleeping Beauty transposase, Marina transposase, Tn7 transposase, Tn10 transposase, Ty1 transposase, Tn552 transposase, and variants, modified products, and derivatives thereof that have the transposition activity of the above transposases.

[0404] In certain embodiments, the kit further comprises primer C, primer D, primer C', and / or primer D'. For example, the kit further comprises primer C, primer D, and primer D'. For example, the kit further comprises primer C, primer D, primer C', and primer D'.

[0405] In certain embodiments, the kit further comprises reagents for nucleic acid hybridization, reagents for nucleic acid extension, reagents for nucleic acid amplification, reagents for recovering or purifying nucleic acids, reagents for constructing a transcriptome sequencing library, reagents for sequencing (e.g., second or third generation sequencing), or any combination thereof.

[0406] Definition of Terms In this application, unless otherwise specified, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps of molecular biology, biochemistry, nucleic acid chemistry, cell culture, etc., as used herein, are all routine steps widely used in the corresponding technical fields. Meanwhile, in order to better understand this application, the definitions and explanations of relevant terms are provided below.

[0407] When the terms "eg," "for example," "such as," "comprise," "include," or variations thereof are used herein, these terms are not intended to be limiting terms and instead are intended to mean "including but not limited to" or "not limiting."

[0408] Unless otherwise indicated herein or clearly contradicted by context, the terms "a" and "an" as well as "the" and similar referents in the context of describing this application (particularly in the context of the claims which follow) should be construed to cover the singular and plural.

[0409] As used herein, cells that can be applied to the methods of the present application (e.g., cells that can be processed using the methods of the present application to generate a population of labeled nucleic acid molecules) can be any cell of interest, such as cancer cells, stem cells, neural cells, fetal cells, and immune cells involved in immune responses. A cell can include one cell or multiple cells. A cell can be a mixture of cells of the same type or a completely heterogeneous mixture of different types. Different cell types can include cells from different tissues of an individual or cells from the same tissue from different individuals or cells derived from different genera, species, strains, variants of microorganisms or any combination of any or all of the above. For example, different cell types can include normal cells and cancer cells from an individual, different cell types obtained from a human species, such as different immune cells, different bacterial species, strains and / or variants from environmental, forensic, microbiome or other samples, or any other different mixtures of cell types.

[0410] As used herein, the term "UMI" refers to a "unique molecular identifier" that can be used to characterize and / or quantify a nucleic acid molecule. Unless otherwise stated herein or clearly contradicted by context, the present application does not limit the location and amount of the UMI or its complementary sequence in a nucleic acid molecule. For example, if a cDNA strand contains a UMI or its complementary sequence, the UMI or its complementary sequence may be located 3' of the cDNA sequence of the cDNA strand, or may be located 5' of the cDNA sequence, or the UMI or its complementary sequence may be included in both the 3' and 5' of the cDNA sequence. If the complementary strand of a cDNA strand contains a UMI or its complementary sequence, the UMI or its complementary sequence may be located 3' of the complementary sequence of the cDNA sequence of the complementary strand of the cDNA strand, or may be located 5' of the complementary sequence of the cDNA sequence, or the UMI or its complementary sequence may be included in both the 3' and 5' of the complementary sequence of the cDNA sequence.

[0411] As used in this application, "DNB" (DNA nanoball) is a typical RCA (rolling circle amplification) product, which has the characteristics of an RCA product, where the RCA product is a single-stranded DNA with multiple copies of a specific sequence, which can form a "globular"-like structure due to the interactions between the bases contained in the DNA. Typically, the library molecules are circularized to form single-stranded circular DNA, which can then be amplified by several orders of magnitude using rolling circle amplification technology, thereby generating an amplification product called DNB.

[0412] As used herein, a "nucleic acid molecule population" refers to a collection or set of nucleic acid molecules, e.g., nucleic acid molecules that are derived directly or indirectly from a target nucleic acid molecule (e.g., double-stranded DNA, RNA / cDNA hybrid, single-stranded DNA, or single-stranded RNA). In some embodiments, the nucleic acid molecule population comprises a library of nucleic acid molecules, the library of nucleic acid molecules comprising sequences that are qualitatively and / or quantitatively representative of the target nucleic acid molecule sequences. In other embodiments, the population of nucleic acid molecules comprises a subset of the library of nucleic acid molecules.

[0413] As used herein, a "library of nucleic acid molecules" refers to a collection or population of labeled nucleic acid molecules (e.g., labeled double-stranded DNA, labeled RNA / cDNA hybrid, labeled single-stranded DNA, or labeled single-stranded RNA) or fragments thereof that are generated directly or indirectly from a target nucleic acid molecule, and the combination of labeled nucleic acid molecules or fragments thereof in the collection or population is shown to be qualitatively and / or quantitatively representative of the sequence of the target nucleic acid molecule sequence from which the labeled nucleic acid molecule was generated. In certain embodiments, the library of nucleic acid molecules is a sequencing library. In certain embodiments, the library of nucleic acid molecules can be used to construct a sequencing library.

[0414] As used herein, "cDNA" or "cDNA strand" refers to a "complementary DNA" synthesized by using at least a portion of an RNA molecule of interest as a template and extending through a primer that anneals to the RNA molecule of interest under the catalysis of an RNA-dependent DNA polymerase or reverse transcriptase (this process is also called "reverse transcription"). The synthesized cDNA molecule is "homologous" or "complementary" to at least a portion of the template, or "base-pairs" or "complexes" with at least a portion of the template.

[0415] As used herein, the term "upstream" is used to describe the relative positional relationship of two nucleic acid sequences (or two nucleic acid molecules) and has the meaning commonly understood by those skilled in the art. For example, the expression "a nucleic acid sequence is located upstream of another nucleic acid sequence" means that the former is located at a more forward position (i.e., closer to the 5'-end) than the latter when aligned in the 5' to 3' direction. As used herein, the term "downstream" has the opposite meaning to "upstream".

[0416] As used herein, "tag sequence Y", "tag sequence A", "tag sequence B", "consensus sequence X1", "consensus sequence X2", "consensus sequence A", "consensus sequence B", etc. refer to an oligonucleotide having a non-target nucleic acid component that provides a means for identification, recognition and / or molecular or biochemical manipulation of the nucleic acid molecule ligated thereto or a derivative of the nucleic acid molecule ligated thereto (e.g., a complementary fragment of a nucleic acid molecule, a short fragment of a nucleic acid molecule, etc.) (e.g., by providing a site for annealing with an oligonucleotide. The oligonucleotide may be, for example, a primer for DNA polymerase extension or an oligonucleotide for a capture or ligation reaction.) The oligonucleotide may consist of at least 2 (preferably about 6-100, although there is no definite limit to the length of the oligonucleotide, the exact size depends on a number of factors, which in turn depend on the ultimate function or use of the oligonucleotide) nucleotides, and may be composed of multiple oligonucleotide fragments arranged contiguously or non-contiguously. The oligonucleotide sequence may be unique to each nucleic acid molecule to which it is ligated, or may be unique to a particular type of nucleic acid molecule to which it is ligated. The oligonucleotide sequence may be reversibly or irreversibly ligated to the polynucleotide sequence to be "labeled" by any method, including ligation, hybridization, or other methods. The process of ligating an oligonucleotide sequence to a nucleic acid molecule may be referred to as "labeling," and a nucleic acid molecule to which a label has been attached or which contains a label sequence is referred to as a "labeled nucleic acid molecule" or a "tagged nucleic acid molecule."

[0417] For a variety of reasons, the nucleic acids or polynucleotides of the present application (e.g., "tag sequence Y", "tag sequence A", "tag sequence B", "consensus sequence X1", "consensus sequence X2", "consensus sequence A", "consensus sequence B", "primer IA", "primer I-A'", "primer IB", "primer II-A", "primer II-A'", "primer II-B", "primer II-B'", "primer B", "primer C", "primer D", "primer C'", "primer D'", "random primer", "bridging oligonucleotide I", "bridging oligonucleotide sequence II-I", "bridging oligonucleotide sequence II-II", etc.) may comprise one or more modified nucleobases, sugar moieties or internucleoside linkages. For example, some reasons for using nucleic acids or polynucleotides containing modified nucleobases, sugar moieties, or internucleoside linkages include, but are not limited to, (1) changing the Tm, (2) changing the susceptibility of the polynucleotide to one or more nucleases, (3) providing a moiety for linking a label, (4) providing a label or label quencher, or (5) providing a moiety such as biotin for attachment to another molecule in solution or bound to a surface. For example, in some embodiments, oligonucleotides, e.g., primers, can be synthesized to include one or more nucleic acid analogs with constrained conformations in the random portion, including, but not limited to, one or more ribonucleic acid analogs in which the ribose ring is "locked" by a methylene bridge connecting the 2'-O atom to the 4'-C atom. These modified nucleotides provide an increase in the Tm or melting temperature of each molecule of about 2 degrees Celsius to about 8 degrees Celsius. For example, in some embodiments in which an oligonucleotide primer containing ribonucleotides is used, one indication for using modified nucleotides in the method may be that oligonucleotides containing modified nucleotides may be digested by single-strand specific RNases.

[0418] As used herein, a "first binding molecule" can specifically or non-specifically interact with a "first labeled molecule." In certain embodiments, the first binding molecule interacts with the first labeled molecule in a manner selected from the group consisting of interactions between positive and negative charges, affinity interactions (e.g., interactions between biotin and avidin, biotin and streptavidin, antigens and antibodies, receptors and ligands, enzymes and cofactors), click chemistry reactions (e.g., click chemistry reactions between alkynyl-containing compounds and azide compounds), and any combination thereof.

[0419] For example, the first labeled molecule is polylysine and the first binding molecule is a protein capable of binding to the polylysine; the first labeled molecule is an antibody and the first binding molecule is an antigen capable of binding to the antibody; the first labeled molecule is biotin and the first binding molecule is streptavidin; the first binding molecule is a compound containing an alkynyl group and the labeled molecule is an azide compound; or the first binding molecule is an N-hydroxysulfosuccinate (NHS) ester and the first labeled molecule is an amino-containing compound.

[0420] For example, the first labeled molecule is an antigen and the first binding molecule is an antibody capable of binding to the antigen; the first labeled molecule is streptavidin and the first binding molecule is biotin; the first labeled molecule is an azide compound and the first labeled molecule is an alkynyl-containing compound; or the first binding molecule is an amino-containing compound and the first labeled molecule is an N-hydroxysulfosuccinate (NHS) ester.

[0421] In the methods of the present application, for example, the nucleobase of a single nucleotide at one or more positions of a polynucleotide or oligonucleotide may include guanine, adenine, uracil, thymine or cytosine, or optionally, one or more of the nucleobases may include modified bases, such as, but not limited to, xanthine, allylamino-uracil, allylamino-thymine nucleoside, hypoxanthine, 2-aminoadenine, 5-propynyluracil, 5-propynylcytosine, 4-thiouracil, 6-thioguanine, azauracil, deazauracil, thymine nucleoside, cytosine, adenine or guanine. Furthermore, they may include nucleobases derivatized with the following moieties: biotin moiety, digoxigenin moiety, fluorescent or chemiluminescent moiety, quenching moiety or some other moiety. The present application is not limited to the listed nucleobases, and the given list illustrates a wide range of examples of bases that may be used in the methods of the present application.

[0422] With respect to the nucleic acids or polynucleotides of the present application, one or more of the sugar moieties may comprise 2'-deoxyribose, or optionally, one or more of the sugar moieties may comprise some other sugar moieties, such as, but not limited to, ribose or 2'-fluoro-2'-deoxyribose or 2'-O-methyl-ribose, which are resistant to some nucleases, or 2'-amino-2'-deoxyribose or 2'-azido-2'-deoxyribose, which are labeled by reaction with a visible, fluorescent, infrared fluorescent or other detectable dye or a chemical having an electrophilic, photoreactive, alkynyl or other reactive chemical moiety.

[0423] The internucleoside linkages of the nucleic acids or polynucleotides of the present application may be phosphodiester linkages, or optionally, one or more of the internucleoside linkages may comprise modified linkages, such as, but not limited to, phosphorothioate, phosphorodithioate, phosphoroselenate, or phosphorodiselenate linkages, which are resistant to some nucleases.

[0424] As used herein, the term "terminal deoxynucleotidyl transferase activity" refers to the ability to catalyze the template-independent addition (or "tailing") of one or more deoxyribonucleoside triphosphates (dNTPs) or a single dideoxyribonucleoside triphosphate to the 3'-end of a cDNA. Examples of reverse transcriptases with terminal deoxynucleotidyl transferase activity include, but are not limited to, -MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and variants, modified products, and derivatives thereof that have reverse transcription activity and terminal deoxynucleotidyl transferase activity of reverse transcriptase. Reverse transcriptases may or may not have RNase activity (particularly RNase H activity). In a preferred embodiment, the reverse transcriptase used for reverse transcription of RNA to generate cDNA does not have RNase activity. Thus, in a preferred embodiment, the reverse transcriptase used for reverse transcription of RNA to produce cDNA has terminal deoxynucleotidyl transferase activity and no RNase activity.

[0425] As used herein, a nucleic acid polymerase having "strand displacement activity" refers to a nucleic acid polymerase that, when it encounters a downstream nucleic acid strand that is complementary to the template strand during the process of extending a new nucleic acid strand, can continue the extension reaction and replace (rather than degrade) the nucleic acid strand that is complementary to the template strand.

[0426] As used herein, a nucleic acid polymerase with "5' to 3' exonuclease activity" refers to a nucleic acid polymerase that can catalyze the hydrolysis of 3,5-phosphodiester bonds in the 5' to 3' order of a polynucleotide, thereby degrading nucleotides.

[0427] As used herein, a nucleic acid polymerase (or DNA polymerase) with "high fidelity" refers to a nucleic acid polymerase (or DNA polymerase) that is less likely to introduce an incorrect nucleotide (i.e., error rate) during amplification of a nucleic acid than a wild-type Taq enzyme (e.g., the Taq enzyme whose sequence is set forth in the UniProt Accession: P19821.1).

[0428] As used herein, the terms "annealed", "annealing", "annealing", "hybridized" or "hybridizing" and the like refer to the formation of a complex between nucleotide sequences that have sufficient complementarity to form a complex via Watson-Crick base pairing. For the purposes of this application, nucleic acid sequences that are "complementary" or "hybridize" or "anneal" to one another must be capable of forming a sufficiently stable "hybrid" or "complex" for the intended purpose. It is not necessary that both nucleic acid molecules or corresponding sequences presented therein be "complementary" or "anneal" or "hybridize" to one another by having all nucleobases in the sequence presented by a nucleic acid molecule capable of base pairing or pairing or complexing with all nucleobases in the sequence presented by another nucleic acid molecule. As used herein, the terms "complementary" or "complementarity" are used when referring to a sequence of nucleotides related by the rules of base pairing. For example, the sequence 5'-AGT-3' is complementary to the sequence 3'-TCA-5'. Complementarity can be "partial", where only a portion of the nucleic acid bases match according to the rules of base pairing. Optionally, there can be "complete" or "total" complementarity between the nucleic acids. The degree of complementarity between nucleic acid strands has a large effect on the efficiency and strength of hybridization between nucleic acid strands. The degree of complementarity is particularly important in amplification reactions and detection methods that rely on nucleic acid hybridization. The term "homology" refers to the degree of complementarity of one nucleic acid sequence to another. There can be partial homology (i.e., complementarity) or complete homology (i.e., complementarity). A partially complementary sequence is one that at least partially inhibits the hybridization of a completely complementary sequence to a target nucleic acid and is referred to using the functional term "substantially homologous". Inhibition of hybridization of the completely complementary sequence with the target sequence can be tested under low stringency conditions using hybridization assays (eg, Southern or Northern blotting, solution hybridization, etc.).A substantially homologous sequence or probe competes for or inhibits the binding (i.e., hybridization) of a completely homologous sequence to a target under low stringency conditions. This does not mean that low stringency conditions are conditions that allow non-specific binding, but rather that low stringency conditions require that two sequences bind to each other through specific (i.e., selective) interactions. Absence of non-specific binding can be tested by using a second target that lacks complementarity or has only a low degree of complementarity (e.g., less than about 30% complementarity). If specific binding is low or absent, the probe will not hybridize to the nucleic acid target. The term "substantially homologous" when used in reference to a double-stranded nucleic acid sequence, such as a cDNA or genomic clone, means that any oligonucleotide or probe can hybridize to one or both strands of a double-stranded nucleic acid sequence under low stringency conditions as described herein. As used herein, the term "annealing" or "hybridization" is used to refer to the pairing of complementary nucleic acid strands. Hybridization and hybridization force (i.e., the force of association between nucleic acid strands) are affected by a number of factors known in the art, including the degree of complementarity between nucleic acids, including the stringency of conditions, which is affected by factors such as salt concentration, Tm (melting temperature) for forming hybrids, the presence of other components (e.g., the presence or absence of polyethylene glycol or betaine), the molar concentration of hybridized strands, and the G:C content of the nucleic acid strands.

[0429] As described herein, the solid support is capable of releasing the oligonucleotide probe spontaneously or upon exposure to one or more stimuli (e.g., temperature change, pH change, exposure to certain chemicals or phases, exposure to light, exposure to a reducing agent, etc.) It is understood that the oligonucleotide probe may be released by cleavage of the bond between the oligonucleotide probe and the solid support or by degradation of the solid support itself, or both, and that the oligonucleotide probe may allow itself to be accessed or be accessible by other reagents.

[0430] The addition of multiple types of labile bonds to a solid support allows for the ability of the solid support to respond to different stimuli: each type of labile bond may be sensitive to a relevant stimulus (e.g., chemical stimuli, light, temperature, etc.), such that the release of the material attached to the solid support via each labile bond can be controlled by application of the appropriate stimulus. In addition to thermally cleavable bonds, disulfide bonds, and UV-sensitive bonds, other non-limiting examples of labile bonds that can be attached to a solid support include an ester bond (e.g., an ester bond that can be cleaved with acid, base, or hydroxylamine), an ortho-diol bond (e.g., an ortho-diol bond that can be cleaved by sodium periodate), a Diels-Alder bond (e.g., a Diels-Alder bond that can be cleaved thermally), a sulfone bond (e.g., a sulfone bond that can be cleaved by alkali), a silicyl ether bond (e.g., a silicyl ether bond that can be cleaved by acid), a glycosidic bond (e.g., a glycosidic bond that can be cleaved by an amylase), a peptide bond (e.g., a peptide bond that can be cleaved by a protease), or a phosphodiester bond (e.g., a phosphodiester bond that can be cleaved by a nuclease, such as a DNA enzyme).

[0431] In addition to or as an alternative to the cleavable bond between the solid support and the oligonucleotide described above, the solid support may be degradable, destructible or soluble, either spontaneously or upon exposure to one or more stimuli (e.g., temperature change, pH change, exposure to certain chemicals or phases, exposure to light, exposure to reducing agents, etc.). In some cases, the solid support may be soluble, such that the material components of the solid support dissolve upon exposure to certain chemicals or environmental changes (e.g., temperature change or pH change). In some cases, the solid support may degrade or dissolve under high temperature and / or alkaline conditions. In some cases, the solid support may be thermally degradable, such that the solid support degrades when exposed to an appropriate temperature change (e.g., heating). Decomposition or dissolution of the solid support bound to a substance (e.g., an oligonucleotide probe) may result in the release of the substance from the solid support.

[0432] As used herein, the terms "transposase" and "reverse transcriptase" and "nucleic acid polymerase" refer to a protein molecule or an aggregate of protein molecules involved in catalyzing certain chemical and biological reactions. In general, the methods, compositions, or kits of the present application are not limited to the use of a specific transposase, reverse transcriptase, or nucleic acid polymerase from a specific source. Rather, the methods, compositions, or kits of the present application may include any transposase, reverse transcriptase, or nucleic acid polymerase from any source that has an enzymatic activity equivalent to the specific enzyme of the specific method, composition, or kit disclosed herein. In addition, the methods of the present application also include the following embodiments: any one specific enzyme provided and used in the steps of the method is replaced by a combination of two or more enzymes, whether used separately, stepwise, or simultaneously together, such that the reaction mixture produces the same results as would be obtained using that specific enzyme when the two or more enzymes are used in combination. The methods, buffers, and reaction conditions provided herein, including those in the examples, are currently preferred for the embodiments of the methods, compositions, and kits of the present application. However, other enzyme storage buffers, reaction buffers and reaction conditions may be used for some of the enzymes of the present application and are known in the art and may be suitable for use in the present application and are included herein.

[0433] [Beneficial Effects of the Present Application] The present application provides high-resolution nucleic acid arrays (e.g., chips) and methods that allow positional labeling of nucleic acid molecules, as well as methods using nucleic acid arrays or methods for high-throughput sequencing (particularly high-throughput single-cell transcriptome sequencing). The methods of the present application have one or more beneficial technical effects selected from the following: (1) The nucleic acid array (e.g., chip) has high resolution and can measure the single-cell region (e.g., 80-100 μm 2) can contain at least 50 (e.g., at least 50, at least 100, at least 200, at least 300, at least 400, or at least 500) microdots, each microdot being bound to one type of oligonucleotide probe for labeling (e.g., an oligonucleotide probe containing tag sequence Y) containing position information, and each type of oligonucleotide probe contains at least one copy. Thus, using a nucleic acid array, it can be realized that different cells of a sample (e.g., a cell suspension) are labeled with a specific positioning sequence (e.g., tag sequence Y), and thus, by detecting the specific positioning sequence (e.g., tag sequence Y) of the labeled nucleic acid molecule, the spatial position information of the nucleic acid molecule of the nucleic acid array can be detected, and then the nucleic acid molecule from the same single cell can be identified, thereby realizing the analysis of a single cell sample.

[0434] (2) When a nucleic acid array (e.g., chip) is used to construct a high-throughput sequencing library, it is not necessary to perform single-cell sorting of cells of a sample (e.g., cell suspension), and multiple cells can be directly adsorbed onto the nucleic acid array (e.g., chip). Due to the high resolution of the nucleic acid array (e.g., chip), the size and spacing of the microdots are much smaller than the size of a single cell. Therefore, each cell (or nucleic acid molecule from a cell) is captured and labeled using one or more oligonucleotide probes for labeling (e.g., oligonucleotide probes containing tag sequence Y) that have position information and are located on the nucleic acid array (e.g., chip). In other words, the nucleic acid array (e.g., chip) can theoretically capture and label every cell of a sample, which effectively avoids the loss of rare cell information.

[0435] (3) Based on the unique design of the nucleic acid array (e.g., chip), the method of the present application can capture millions of cells on a single chip for single-cell sequencing, and the cell capture efficiency can theoretically reach 100%. That is, the cell capture throughput of the method of the present application can reach the scale of millions, and the cell capture efficiency can reach nearly 100%, which greatly exceeds existing technologies (existing technologies such as the 10x Chromium cell sorting platform are limited by the number of microdroplets formed in the oil phase, and as a result, their throughput is difficult to exceed the scale of tens of thousands, and due to the characteristics of the Poisson distribution, their cell capture rate is theoretically up to 60%).

[0436] Preferred embodiments of the present application are described in detail below with reference to the accompanying drawings and examples, but those skilled in the art will understand that the following drawings and examples are merely used to illustrate the present application and do not limit the scope of the present application. Various objects and advantageous aspects of the present application will become apparent to those skilled in the art from the following detailed description of the accompanying drawings and preferred embodiments. EXAMPLES

[0437] The present application will now be described with reference to the following examples, which are intended to illustrate, but not limit, the present application. Unless otherwise specified, the experiments and methods described in the examples were essentially carried out according to conventional methods well known in the art and described in various references. Furthermore, if no specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. If the manufacturer of the reagents or equipment used was not specified, they were all conventional products that could be purchased commercially. Those skilled in the art will understand that the examples are intended to illustrate the present application and are not intended to limit the scope of the protection sought by the present application. All publications and other references mentioned herein are incorporated by reference in their entirety.

[0438] Example 1 The sequence information involved in this example is shown in Table 1-1:

[0439] [Table 1-1]

[0440] I. Preparation of the Capture Chip 1. A sequence of DNA library molecule containing chip position information was designed, which was composed of the coding sequence (X1) of consensus sequence X1, the coding sequence (Y) of tag sequence and the coding sequence (X2) of consensus sequence X2 from 5' to 3'. An exemplary nucleotide sequence of the DNA library molecule is shown in SEQ ID NO: 1. Beijing Liuhe BGI Co., Ltd. was commissioned to synthesize the DNA library molecule.

[0441] 2. Amplification and Loading of Library Molecules (1) DNA nanoballs (DNBs) were prepared using DNBSEQ sequencing kit (purchased from MGI, catalog number 1000019840). A specific embodiment is briefly described below.

[0442] Briefly, 40 μL of the reaction system shown in Table 1-2 was prepared. The reaction system was placed in a PCR machine and reacted according to the following reaction conditions: 95°C for 3 minutes, 40°C for 3 minutes. After the reaction was completed, the reaction product was placed on ice and 40 μL of mixed enzyme I, 2 μL of mixed enzyme II (from DNBSEQ sequencing kit), 1 μL of ATP (100 mM stock solution, obtained from Thermo Fisher) and 0.1 μL of T4 ligase (obtained from NEB, catalog number: M0202S) were added. After mixing the wells, the above reaction system was placed in a PCR machine and reacted at 30°C for 20 minutes to generate DNB.

[0443] [Table 1-2]

[0444] (2) The DNBs were then loaded onto a BGISEQ 500 sequencing chip according to the method described in the BGISEQ 500 High Throughput Sequencing Reagent Set (SE50) (purchased from MGI, catalog number: 1000012551).

[0445] In the sequencing chip, MDA reagent from BGISEQ500 PE50 sequencing kit (purchased from MGI, catalog number: 1000012554) was added and incubated at 37° C. for 30 minutes, then the chip was washed with 5× SSC.

[0446] (3) The chip surface was modified with N3-PEG3500-NHS (modification reagent purchased from Sigma, catalog number: JKA5086) by incubating for 30 minutes, and then the DBCO-modified primer for chip sequence synthesis (sequence shown in SEQ ID NO: 3) was injected and incubated at room temperature overnight.

[0447] 3. Sequencing and deciphering of position sequence information was performed. DNB was sequenced according to the instruction manual of BGISEQ-500 high-throughput sequencing reagent set, and the length of SE read was set to 25bp. During the sequencing process, the above DBCO modified primer was extended to obtain a strand containing position sequence information, and this strand was deciphered to obtain the position sequence information corresponding to DNB.

[0448] 4. Continued extension of the strand obtained during step 3 during the sequencing process: Based on step 3 above, the cPAS reaction of 15 bases was continued to obtain the chip sequence (SEQ ID NO: 8, consisting of consensus sequence X1 (SEQ ID NO: 4), tag sequence Y and consensus sequence X2 (SEQ ID NO: 5)).

[0449] 5. The DNB was excised using the restriction endonuclease HaeIII, and the remaining fragments of the DNB were removed by denaturation at high temperature, so that in step 4 only the chip sequence remained on the chip.

[0450] 6. Chip cutting: The prepared chip was cut into several small pieces. The size of the small pieces was adjusted according to the experimental needs. The chip was immersed in 50 mM tris buffer, pH 8.0 at 4°C for later use.

[0451] II. Single cell fixation 1. Approximately 3,000 Hek293 cells were harvested and prepared into a cell suspension (PBS solution) according to a conventional method.

[0452] 2. The prepared chip was taken out and treated with polylysine for 30 minutes, then dried for 30 minutes.

[0453] 3. The cell suspension was dropped onto the treated chip and incubated at room temperature for 30 minutes, then the chip was fixed with frozen methanol at -20°C for 30 minutes.

[0454] 4. Cells were permeabilized by treatment with 0.1% triton-100 for 15 minutes.

[0455] III. In situ synthesis of cDNA 1. cDNA Synthesis 200 μL of the reverse transcriptase reaction system shown in Table 1-3 was prepared, and the reaction solution was added to the chip so as to completely cover it, and reacted at 42° C. for 90 to 180 minutes. cDNA was synthesized using reverse transcriptase, mRNA as a template, and a primer containing polyT (the sequence of the primer is shown in SEQ ID NO: 6 and was composed of the consensus sequence A (CA), UMI sequence (NNNNNNNNNN), and polyT sequence), and a CCC overhang was added to the 3' end of the cDNA strand. After hybridization and annealing (via complementary pairing between the GGG at the end of the TSO sequence and the CCC overhang of the cDNA strand) of the TSO sequence (SEQ ID NO: 7, composed of consensus sequence B(CB) and a GGG overhang) with the cDNA strand, reverse transcriptase was used to continuously extend the cDNA strand using consensus sequence B as a template, resulting in labeling of the 3' end of the cDNA with a c(CB) tag (the complementary sequence of consensus sequence B).

[0456] [Table 1-3]

[0457] The synthesized cDNA strand had the following sequence structure: reverse transcription primer sequence (SEQ ID NO: 6)-cDNA sequence-c(TSO) sequence (complementary sequence of SEQ ID NO: 7).

[0458] 2. Ligation of Sequencing Chip with cDNA of ChIP Sequence After cDNA was synthesized, the chip was washed twice with 5X SSC, 1 ml of the reaction system shown in Tables 1-4 was prepared and the appropriate volume was injected into the chip to ensure that the chip was filled with the following ligation reaction solutions, and the reaction was carried out at room temperature for 30 minutes.

[0459] By the above reaction, the 5' end of the cDNA sequence was ligated to the 3' end of the chip sequence of the single-cell sequencing chip (i.e., the 5' end of the cDNA sequence was labeled with the chip sequence) to obtain a new nucleic acid molecule containing positional information (i.e., tag sequence Y), which was composed of the following sequence structure: chip sequence (sequence number 8)-reverse transcription primer sequence (sequence number 6)-cDNA sequence-c(TSO) sequence (complementary sequence of sequence number 7).

[0460] After the reaction was completed, the chip was washed with 5X SSC. 200 μL of Bst polymerization reaction solution (NEB, M0275S) was prepared according to the manufacturer's instructions, injected into the chip, and reacted at 65° C. for 60 minutes to obtain a single-stranded nucleic acid molecule containing positional information.

[0461] [Table 1-4]

[0462] 3. Release of cDNA The chip was incubated with 75 μL of 80 mM KOH for 5 min at room temperature, the resulting liquid was collected, and then 10 μL of 1 M, pH 8.0 Tris-HCl was added to neutralize the cDNA recovery solution.

[0463] 4. Amplification of cDNA 200 μL of the reactions shown in Tables 1-5 were prepared for 3'-end transcriptome sequencing and library construction, respectively, and divided into two tubes for PCR:

[0464] [Table 1-5]

[0465] The above reaction system was placed in a PCR machine and the reaction program was set as follows: 95°C for 3 min, 11 cycles (98°C for 20 sec, 58°C for 20 sec, 72°C for 3 min), 72°C for 5 min, and 4°C for ∞. After the reaction was completed, XP beads (purchased from AMPure) were used for magnetic bead-based purification and recovery. The dsDNA concentration was determined using a Qubit instrument, and the length distribution of the cDNA amplification products was detected using a 2100 Bioanalyzer (purchased from Agilent).

[0466] IV. cDNA Library Construction and Sequencing 1.Tn5 tagmentation According to the cDNA concentration, 20ng of cDNA (obtained in step III) was taken, 0.5μM of Tn5 transposase and corresponding buffer solution (purchased from BGI, catalog number: 10000028493, the method of coating Tn5 transposase was according to the operation of Stereomics Library Preparation Kit-S1) were added, mixed thoroughly to prepare 20μL of reaction system, and the reaction was carried out at 55℃ for 10 minutes, then 5μL of 0.1% SDS was added, mixed thoroughly at room temperature for 5 minutes to terminate Tn5 tagmentation.

[0467] 2. PCR Amplification 100 μL of the following reaction mixture was prepared: [Table 1-6]

[0468] After mixing, it was placed in a PCR machine and the program was set as follows: 95°C for 3 min, 11 cycles (98°C for 20 sec, 58°C for 20 sec, 72°C for 3 min), 72°C for 5 min, hold at 4°C for ∞. After the reaction was completed, XP beads were used for magnetic bead-based purification and recovery. The dsDNA concentration was determined using a Qubit instrument.

[0469] 3. Sequencing 80 fmol of the amplified product was removed and used to prepare DNB. 40 μL of the following reaction mixture was prepared: [Table 1-7] The above reaction volume was placed in a PCR machine for reaction, and the reaction conditions were as follows: 95°C for 3 minutes, 40°C for 3 minutes. After the reaction was completed, the resulting reaction solution was placed on ice, and 40μL of mixed enzyme I, 2μL of mixed enzyme II, 1μL of ATP and 0.1μL of T4 ligase required for DNB preparation of DNBSEQ sequencing kit were added, mixed thoroughly, and the above reaction system was placed in a PCR machine and reacted at 30°C for 20 minutes to form DNB.

[0470] According to the method described in the PE50 kit supporting MGISEQ 2000, the DNB was loaded onto the sequencing chip of MGISEQ 2000, and sequencing was performed according to the related instruction manual. The PE50 sequencing model was selected, in which the first strand sequencing was divided into two sections of sequencing, i.e., 25bp sequencing was performed first, followed by 15 cycles of dark reaction, then 10bp UMI sequence sequencing was performed, and second strand sequencing was performed by setting 50bp for sequencing.

[0471] V. Data Analysis: 1. Data analysis was performed by logging in to the website http: / / stereomap.cngb.org / Stereo-Draftsman / report / index and following the operation guide of the website. For the read 1 sequence obtained in PE50 sequencing (from first strand sequencing), the first 25bp of sequence was aligned with the 25bp position information during the chip preparation process, and the reads that were well aligned with the chip position information were kept and mapped to the corresponding chip position. Read 2 (from second strand sequencing) corresponding to the reads mapped to the chip position was found, and read 2 was aligned with the human genome, and duplicate reads were removed based on the UMI information, thus obtaining the genes captured from each cell and the number of reads per gene.

[0472] 2. A part of the chip was taken out and the analysis results are shown in Figure 12 and Figure 13, where Figure 13 shows a part of the enlarged view of the cell in Figure 12. From the figure, it could be seen that single cells were scattered on the chip. By circling one of the cells, the average number of captured genes and the number of UMIs of the cell could be obtained.

[0473] Example 2 The sequence information pertaining to this example is shown in Tables 2-1 and 1-1:

[0474] [Table 2-1]

[0475] I. Preparation of the Capture Chip 1. A DNA library sequence containing positional information was designed and synthesized, and its nucleotide sequence is shown in SEQ ID NO: 1. Beijing Liuhe BGI Co., Ltd. was commissioned to carry out the sequence synthesis.

[0476] 2. In Situ Amplification of Libraries (1) Preparation of DNA nanoballs (DNBs): 40 μL of a reaction system as shown in Table 2-2 was prepared, and 80 fmol of a DNA library sequence containing positional sequence information was added.

[0477] [Table 2-2]

[0478] The above reaction volume was placed in a PCR machine for reaction, and the reaction conditions were as follows: 95°C for 3 minutes, 40°C for 3 minutes. After the reaction was completed, the resulting reaction solution was placed on ice, and 40μL of mixed enzyme I, 2μL of mixed enzyme II, 1μL of ATP (100mM stock solution, Thermo Fisher) and 0.1μL of T4 ligase (purchased from NEB, catalog number: M0202S) required for DNB preparation of DNBSEQ sequencing kit were added, mixed thoroughly, and the above reaction system was placed in a PCR machine and reacted at 30°C for 20 minutes to form DNB. The DNB was loaded onto the SEQ 500 sequencing chip according to the method described in the BGISEQ-500 High Throughput Sequencing Reagent Set (SE50).

[0479] (2) MDA reagent from a PE50 sequencing kit (purchased from MGI, catalog number: 1000012554) was added to the sequencing chip and incubated at 37° C. for 30 minutes, then the chip was washed with 5× SSC.

[0480] (3) The chip surface was incubated with N3-PEG3500-NHS (purchased from Sigma, catalog number: JKA5086) to modify it with N3-PEG3500-NHS, and after incubation for 30 minutes, it was injected with DBCO-modified primer (SEQ ID NO: 3) and incubated at room temperature overnight.

[0481] 3. Sequencing and deciphering of position sequence information was performed. DNB was sequenced according to the instruction manual of BGISEQ-500 high-throughput sequencing reagent set, and the length of SE read was set to 25bp. During the sequencing process, the above DBCO modified primer was extended to obtain a strand containing position sequence information, and this strand was deciphered to obtain the position sequence information corresponding to DNB.

[0482] 4. Liuhe BGI was commissioned to synthesize a probe capture sequence with UMI (SEQ ID NO: 15, with phosphorylation modification at its 5' end), and the capture sequence was ligated with the strand obtained during the sequencing process using T4 ligase according to the following reaction system: The ligation reaction system is shown in Table 2-3.

[0483] [Table 2-3]

[0484] 5. The DNB was excised using restriction endonucleases HaeIII and MboI, and the remaining fragments of the DNB were removed by denaturation at high temperature, so that only the probe containing the positional information and the capture sequence remained on the chip (the sequence was shown in SEQ ID NO: 16): CTGCTGACGTACTGAGAGGCATGGCGACCTTATCAG NNNNNNNNNNNNNNNNNNNNN TTGTCTTCCTAAGACNNNNNNNNNNNTTTTTTTTTTTTTTTTTTV, where NNNNNNNNNNNNNNNNNNNNNNN represents the location information.

[0485] 6. Chip cutting: The prepared chip was cut into several small pieces. The size of the small pieces was adjusted according to the experimental needs. The chip was immersed in 50 mM tris buffer, pH 8.0 at 4°C for later use.

[0486] II. Single cell fixation 1. Approximately 3,000 Hek293 cells were harvested and prepared into a cell suspension (PBS solution) according to a conventional method.

[0487] 2. The prepared chip was taken out and treated with polylysine for 30 minutes, then dried for 30 minutes.

[0488] 3. The cell suspension was dropped onto the treated chip and incubated at room temperature for 30 minutes, then the chip was fixed with frozen methanol at -20°C for 30 minutes.

[0489] 4. Cells were permeabilized by treatment with 0.1% triton-100 for 15 minutes.

[0490] III. In situ synthesis of cDNA 1. cDNA synthesis. The chip was washed twice with 5X SSC at room temperature. 200μL of reverse transcriptase reaction system as shown in Table 2-4 was prepared, and the reaction solution was added to the chip containing cells so as to completely cover it, and the reaction was carried out at 42℃ for 90-180 minutes, and cDNA was synthesized using the polyT of the chip probe as a primer, and the 3' end of the cDNA was labeled with a TSO tag for the synthesis of the complementary strand of cDNA. The cDNA strand was as follows: CTGCTGACGTACTGAGAGGCATGGCGACCTTATCAGNNNNNNNNNNNNNNNNNNNNNNNTTGTCTTCcTAAGACNNNNNNNNNTTTTTTTTTTTTTTTTTTTTV (cDNA)CCCGCCTCTCAGTACGTCAGCAG, RNase H treatment was performed for 30 minutes to digest RNA.

[0491] [Table 2-4]

[0492] 3. cDNA release: 75 μL of 200 mM KOH was used to incubate the chip at room temperature for 5 minutes, the resulting liquid was collected, and then 15 μL of 1 M, pH 8.0 Tris-HCl was added to neutralize the cDNA recovery solution.

[0493] 4. cDNA Amplification. 200 μL of reaction systems as shown in Tables 2-5 were prepared for the construction of 3' and 5' transcriptome libraries, respectively, and each of them was divided into two tubes for PCR:

[0494] [Table 2-5]

[0495] The above reaction system was placed in a PCR machine, and the reaction program was set as follows: 95°C for 3 min, 11 cycles (98°C for 20 sec, 58°C for 20 sec, 72°C for 3 min), 72°C for 5 min, and 4°C for ∞. After the reaction was completed, XP beads were used for magnetic bead-based purification and recovery. The dsDNA concentration was determined using Qubit kit, and the cDNA fragment distribution was detected using 2100 Bioanalyzer (purchased from Agilent). The detection results are shown in Figure 14, and the cDNA length was predicted.

[0496] IV. Library construction and sequencing of cDNA libraries 1. Cleavage and tagmentation with Tn5. According to the cDNA concentration, take 20ng of cDNA, add 0.5μM Tn5 transposase (coated with the first strand as shown in SEQ ID NO:19 and the second strand as shown in SEQ ID NO:20) and corresponding buffer (purchased from BGI, catalog number: 10000028493, the method of coating Tn5 transposase was according to the operation of Stereomics library preparation kit), mix well to prepare 20μL reaction system, carry out the reaction at 55℃ for 10 minutes, then add 5μL of 0.1% SDS, mix well at room temperature for 5 minutes to terminate the cleavage and tagmentation.

[0497] 2. PCR Amplification. 100 μL of reaction system was prepared as shown in Table 2-6: [Table 2-6]

[0498] After mixing, it was loaded into the PCR machine and the following program was set: 95°C for 3 min, 11 cycles (98°C for 20 sec, 58°C for 20 sec, 72°C for 3 min), 72°C for 5 min, hold at 4°C for ∞. After the reaction was completed, XP beads were used for magnetic bead-based purification and recovery. dsDNA concentration was determined using the Qubit kit.

[0499] 3. Sequencing. 80 fmol of the above amplification product was taken out to prepare DNB. 40 μL of the reaction system shown in Table 2-7 was prepared: [Table 2-7]

[0500] The above reaction volume was placed in a PCR machine for reaction, and the reaction conditions were as follows: 95°C for 3 minutes, 40°C for 3 minutes. After the reaction was completed, the resulting reaction solution was placed on ice, and 40μL of mixed enzyme I, 2μL of mixed enzyme II, 1μL of ATP (100mM stock solution, Thermo Fisher) and 0.1μL of T4 ligase required for DNB preparation of DNBSEQ sequencing kit were added, mixed thoroughly, and then the above reaction system was placed in a PCR machine and reacted at 30°C for 20 minutes to form DNB.

[0501] According to the method described in the PE50 kit supporting MGISEQ 2000, the DNB was loaded onto the sequencing chip of MGISEQ 2000, and sequencing was performed according to the related instruction manual. The PE50 sequencing model was selected, in which the first strand sequencing was divided into two sections of sequencing, i.e., 25bp sequencing was performed first, followed by 15 cycles of dark reaction, then 10bp UMI sequence sequencing was performed, and second strand sequencing was performed by setting 50bp for sequencing.

[0502] V. Data Analysis: 1. Data analysis was performed by logging in to the website http: / / stereomap.cngb.org / Stereo-Draftsman / report / index and following the operation guide of the website. For the read 1 sequence in PE50, the first 25bp were aligned with the 25bp position information during the chip preparation process, and the reads that were well aligned with the chip position information were kept and mapped to the corresponding chip position. Read 2, which corresponds to the reads mapped to the chip position, was aligned with the human genome and duplicate reads were removed based on the UMI information, thus obtaining the genes captured from each cell and the number of reads per gene.

[0503] 2. The analysis results are shown in Figure 15 and Figure 16, where a part of the chip was taken, where we could see that many single cells were scattered on the chip. By circling one of the cells, we could get the average number of captured genes and the number of UMIs of the cell.

[0504] Although the specific embodiments of the present application are described in detail, those skilled in the art will understand that various modifications and changes can be made in the details based on all the technologies disclosed, and all these modifications are within the scope of protection of the present application. The full scope of the present application is given by the appended claims and any equivalents thereof.

Claims

1. 1. A method for generating a population of labeled nucleic acid molecules, comprising the steps of: (1) providing a sample comprising one or more cells and a nucleic acid array; the sample is a single cell suspension, the cells comprising a first binding molecule; the nucleic acid array comprises a solid support, the solid support comprises a first labeled molecule, and the first binding molecule is capable of forming an interactive pair with the first labeled molecule; the solid support further comprises a plurality of microdots, the size of the microdots being less than 5 μm, the center-to-center distance between adjacent microdots being less than 10 μm, each microdot being bound to one type of oligonucleotide probe, each type of oligonucleotide probe comprising at least one copy, the oligonucleotide probes comprising or consisting of, in a 5' to 3' direction, a consensus sequence X1, a tag sequence Y and a consensus sequence X2; the oligonucleotide probes bound to different microdots have different tag sequences Y; (2) contacting the one or more cells with the solid support of the nucleic acid array, whereby each cell occupies at least one microdot of the nucleic acid array and allows a first binding molecule of the cell to interact with the first labeled molecule of the solid support; performing a pretreatment, including reverse transcription, on RNA from the one or more cells before or after contacting the one or more cells with the nucleic acid array to generate a first population of nucleic acid molecules; and (3) associating the first population of nucleic acid molecules derived from each cell obtained in the previous step with the oligonucleotide probes bound to the microdots occupied by the cells from which the first population of nucleic acid molecules was derived, thereby generating a second population of nucleic acid molecules labeled with tag sequence Y. A method comprising:

2. the center-to-center distance between adjacent microdots is less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm or less than 0.01 μm, and the size of the microdots is less than 5 μm, less than 1 μm, less than 0.3 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, less than 0.01 μm or less than 0.001 μm; Preferably, the center-to-center distance between adjacent microdots is between 0.5 μm and 1 μm, for example between 0.5 μm and 0.9 μm, for example between 0.5 μm and 0.8 μm; The method of claim 1, wherein the size of the microdots is preferably between 0.001 μm and 0.5 μm.

3. the first binding molecule is capable of forming a specific interaction pair or a non-specific interaction pair with the first labeled molecule; Preferably, the interaction pair is selected from the group consisting of positive and negative charge interaction pairs, affinity interaction pairs (e.g., biotin / avidin, biotin / streptavidin, antigen / antibody, receptor / ligand, enzyme / cofactor), pairs of molecules capable of undergoing click chemistry reactions (e.g., alkynyl-containing compound / azide compound), N-hydroxysulfosuccinate (NHS) ester / amino-containing compound, and any combination thereof; The method of claim 1, wherein the first labeled molecule is polylysine and the first binding molecule is a protein capable of binding to polylysine; the first labeled molecule is an antibody and the first binding molecule is an antigen capable of binding to the antibody; the first labeled molecule is an amino-containing compound and the first binding molecule is an N-hydroxysulfosuccinate (NHS) ester; or the first labeled molecule is biotin and the first binding molecule is streptavidin. (a) the first binding molecule is naturally present in the cell; or (b) the first binding molecule is non-naturally occurring in the cell, and the method further comprises the step of binding the first binding molecule to the one or more cells or expressing the first binding molecule in the one or more cells to provide the cell sample of step (1). The method of claim 1.

5. In step (2), the preprocessing is (i) reverse transcribing the RNA of the one or more cells using primer I-A to generate extension products as labeled first nucleic acid molecules, thereby generating the first population of nucleic acid molecules, wherein primer I-A comprises consensus sequence A and capture sequence A, wherein capture sequence A is capable of annealing to the RNA to be captured to initiate an extension reaction, and wherein consensus sequence A is located upstream of capture sequence A; or (ii) (a) reverse transcribing the RNA of the one or more cells using primer I-A to generate a cDNA strand, wherein the cDNA strand is formed by reverse transcription primed by primer I-A and comprises a cDNA sequence complementary to the RNA and a 3'-end overhang, wherein primer I-A comprises consensus sequence A and capture sequence A, wherein capture sequence A is capable of annealing to the RNA to be captured to initiate an extension reaction, and wherein consensus sequence A is located upstream of capture sequence A; and (b) annealing primer I-B to the cDNA strand generated in (a) to initiate an extension reaction to generate a first extension product as a labeled first nucleic acid molecule, thereby generating the first population of nucleic acid molecules, wherein primer I-B comprises consensus sequence B, a complementary sequence of the 3'-end overhang, and optionally a tag sequence B, wherein the complementary sequence of the 3'-end overhang is located at the 3' end of primer I-B, and the consensus sequence B is located upstream of the complementary sequence of the 3'-end overhang; or (iii) (a) reverse transcribing the RNA of the one or more cells using primer I-A' to generate a cDNA strand, wherein the cDNA strand is formed by reverse transcription primed by primer I-A' and comprises a cDNA sequence complementary to the RNA and a 3'-end overhang, and wherein primer I-A' comprises capture sequence A, which is capable of annealing to the RNA to be captured to initiate an extension reaction; (b) annealing primer I-B to the cDNA strand generated in (a) and performing an extension reaction to generate a cDNA strand; (c) providing an extension primer and performing an extension reaction using the first extension product as a template to generate a second extension product as a labeled first nucleic acid molecule, thereby generating the first population of nucleic acid molecules. Including, In step (3), the first population of nucleic acid molecules derived from each cell obtained in the previous step is associated with oligonucleotide probes bound to the microdots occupied by the cells from which the first population of nucleic acid molecules was derived, thereby generating a second population of nucleic acid molecules labeled with the tag sequence Y, which comprises: contacting bridging oligonucleotide I with the first nucleic acid molecules derived from each cell and obtained in step (2) and the oligonucleotide probes bound to the microdots occupied by the cells under conditions that allow annealing; annealing (e.g., in situ annealing) bridging oligonucleotide I with the first nucleic acid molecules derived from each cell and obtained in step (2) and the oligonucleotide probes bound to the microdots occupied by the cells; and ligating the first nucleic acid molecules and the oligonucleotide probes of the array annealed with bridging oligonucleotide I to obtain ligation products as second nucleic acid molecules having positioning tags, thereby generating the second population of nucleic acid molecules. Including, Bridge oligonucleotide I comprises a first region and a second region, and optionally a third region located between the first region and the second region, and the first region is located upstream of the second region; the first region is capable of annealing to all or part of the consensus sequence A of the primer IA in step (2)(i) or step (2)(ii), or is capable of annealing to all or part of the consensus sequence B of the primer I-B in step (2)(iii); 2. The method of claim 1, wherein the second region is capable of annealing to all or part of the consensus sequence X2.

6. In step (3), the first region and the second region of the bridging oligonucleotide I are immediately adjacent, and ligating the first nucleic acid molecule to the oligonucleotide probe comprises using a nucleic acid ligase to ligate the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide I to obtain a ligation product as a second nucleic acid molecule having a positioning tag; or 6. The method of claim 5, wherein the bridging oligonucleotide I comprises a first region, a second region, and a third region located therebetween, and ligating the first nucleic acid molecule to the oligonucleotide probe comprises using a nucleic acid polymerase to perform a polymerization reaction using the third region as a template, and using a nucleic acid ligase to ligate the nucleic acid molecule hybridized with the first region of the same bridging oligonucleotide I to the nucleic acid molecule hybridized with the third region and the second region to obtain a ligation product as a second nucleic acid molecule having a positioning tag, and preferably the nucleic acid polymerase does not have 5' to 3' exonuclease activity or strand displacement activity.

7. A method comprising step (1), step (2)(i) or (2)(ii), and step (3), (I) the capture sequence A is a random oligonucleotide sequence, and preferably, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have different capture sequences A, and the capture sequence A serves as a unique molecular identifier (UMI) of the second nucleic acid molecule; or (II) The capture sequence A is a poly(T) sequence or a specific sequence targeting a target nucleic acid, and the primer I-A further comprises a tag sequence A such as a random oligonucleotide sequence. Preferably, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have different tag sequences A as UMIs. The method of claim 5.

8. Comprising step (1), step (2)(iii) and step (3), (I) In step (2)(iii)(c), the extension primer is primer IB or primer B″, and primer B″ is capable of annealing to a complementary sequence of consensus sequence B or a subsequence thereof to initiate the extension reaction; (II) in step (2)(iii)(a), the capture sequence A of the primer IA′ is a random oligonucleotide sequence; and in step (2)(iii)(b), the primer IB comprises the consensus sequence B, a complementary sequence of the 3′-end overhang, and the tag sequence B; and (III) In step (2)(iii)(a), the capture sequence A of the primer IA' is a poly(T) sequence or a specific sequence targeting a target nucleic acid, and the primer I-B comprises the consensus sequence B, a complementary sequence of the 3'-end overhang, and the tag sequence B; preferably, in step (3), the ligation products derived from each copy of the oligonucleotide probe bound to the same microdot have different tag sequences B as UMIs. The method of claim 5 , having one or more features selected from the group consisting of:

9. In step (2), the pretreatment is (i) (a) reverse transcribing the RNA of the one or more cells using primer II-A to generate a cDNA strand, wherein the cDNA strand is formed by reverse transcription primed by primer II-A and comprises a cDNA sequence complementary to the RNA and a 3'-end overhang, and wherein primer II-A comprises capture sequence A, which is capable of annealing to the captured RNA and initiating the extension reaction; and (b) generating primer II-B in (a). annealing the primer II-B with the cDNA strand that has been labeled, and performing an extension reaction to generate first extension products as labeled first nucleic acid molecules, thereby generating the first population of nucleic acid molecules, wherein the primer II-B comprises a consensus sequence B, a complementary sequence of the 3'-end overhang, and optionally a tag sequence B, the complementary sequence of the 3'-end overhang being located at the 3' end of the primer II-B, and the consensus sequence B being located upstream of the complementary sequence of the 3'-end overhang; or (ii) (a) reverse transcribing the RNA of one or more cells using primer II-A' to generate a cDNA strand, wherein the cDNA strand is formed by reverse transcription primed by primer II-A' and comprises a cDNA sequence complementary to the RNA and a 3'-end overhang, wherein primer II-A' comprises consensus sequence A and capture sequence A, wherein capture sequence A is capable of annealing to the RNA to be captured and initiating an extension reaction, and wherein consensus sequence A is located upstream of capture sequence A; (b) annealing primer II-B' to the cDNA strand generated in (a); (c) providing an extension primer and performing an extension reaction using the first extension product as a template to generate a second extension product as a labeled first nucleic acid molecule, thereby generating the first population of nucleic acid molecules. Including, In step (3), the first population of nucleic acid molecules derived from each cell obtained in the previous step is associated with the oligonucleotide probes bound to the microdots occupied by the cells from which the first population of nucleic acid molecules was derived, thereby generating a second population of nucleic acid molecules labeled with the tag sequence Y, which comprises: (i) applying annealing conditions to the products of step (2) to anneal (e.g., in situ anneal) the first nucleic acid molecules from each cell obtained in step (2) to the oligonucleotide probes bound to the microdots occupied by the cells, and performing an extension reaction to generate extension products as second nucleic acid molecules having positioning tags, thereby generating the second population of nucleic acid molecules, wherein the consensus sequence X2 or a subsequence thereof of the oligonucleotide probes is (a) capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence B of the first extension products obtained in step (2)(i), or (b) capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence A of the second extension products obtained in step (2)(ii); or (ii) contacting a bridging oligonucleotide pair with the first nucleic acid molecule derived from each cell obtained in step (2) and the oligonucleotide probes bound to the microdots occupied by the cells under conditions that allow annealing, and allowing the bridging oligonucleotide pair to anneal (e.g., in situ anneal) with the first nucleic acid molecule derived from each cell obtained in step (2) and the oligonucleotide probes bound to the microdots occupied by the cells; wherein the bridging oligonucleotide pair is composed of a bridging oligonucleotide II-I and a bridging oligonucleotide II-II, and the bridging oligonucleotide II-I and the bridging oligonucleotide II-II each independently comprise a first region, a second region, and optionally a third region located between the first region and the second region, and the first region is located upstream of the second region; the first region of the bridging oligonucleotide II-I is capable of annealing to the first region of the bridging oligonucleotide II-II, and the second region of the bridging oligonucleotide II-I is capable of annealing to the consensus sequence X2 of the oligonucleotide probe or a subsequence thereof; the second region of the bridging oligonucleotide II-II is (a) capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence B of the first extension product obtained in step (2)(i), or (b) capable of annealing to a complementary sequence or a subsequence thereof of the consensus sequence A of the second extension product obtained in step (2)(ii); In the bridging oligonucleotide pair contacted with the first nucleic acid molecule and the oligonucleotide probe, the bridging oligonucleotide II-I and the bridging oligonucleotide II-II of the bridging oligonucleotide pair are each present in a single-stranded form, or the bridging oligonucleotide II-I and the bridging oligonucleotide II-II of the bridging oligonucleotide pair are annealed to each other and present in a partially double-stranded form; performing a ligation reaction: ligating nucleic acid molecules hybridized to the first region and nucleic acid molecules hybridized to the second region of the same bridging oligonucleotide II-I, and / or ligating nucleic acid molecules hybridized to the first region and nucleic acid molecules hybridized to the second region of the same bridging oligonucleotide II-II, and performing an extension reaction to obtain reaction products as second nucleic acid molecules having positioning tags, thereby generating a second population of nucleic acid molecules, wherein the ligation reaction and the extension reaction are performed in any order. Including, The method of claim 1.

10. In step (3)(ii), (1) the first region and the second region of the bridging oligonucleotide III-I are immediately adjacent, and ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide III-I comprises ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I using a nucleic acid ligase; or the bridging oligonucleotide III-I comprises a first region, a second region and a third region therebetween, and ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-I comprises carrying out a polymerization reaction using a nucleic acid polymerase with the third region as a template, and ligating the nucleic acid molecule hybridized with the first region of the same bridging oligonucleotide II-I to the nucleic acid molecule hybridized with the third region and the second region using a nucleic acid ligase; and / or (2) the first region and the second region of the bridging oligonucleotide II-II are immediately adjacent, and the ligation of the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II comprises ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II using a nucleic acid ligase; or the bridging oligonucleotide II-II comprises a first region, a second region and a third region therebetween, and ligating the nucleic acid molecule hybridized with the first region and the nucleic acid molecule hybridized with the second region of the same bridging oligonucleotide II-II comprises carrying out a polymerization reaction using a nucleic acid polymerase with the third region as a template, and using a nucleic acid ligase to ligate the nucleic acid molecule hybridized with the first region of the same bridging oligonucleotide II-II to the nucleic acid molecule hybridized with the third region and the second region; 10. The method of claim 9.

11. The method includes steps (1), (2)(i) and (3), (I) in step (2)(i)(b), the primer II-B comprises the consensus sequence B, a complementary sequence of the 3′-end overhang, and the tag sequence B; Preferably, in step (3), the second nucleic acid molecules derived from each copy of the oligonucleotide probe bound to the same microdot have different tag sequences B as UMIs. (II) in step (2)(i)(a), the capture sequence A of the primer II-A is a random oligonucleotide sequence; and (III) In step (2)(i)(a), the capture sequence A of the primer II-A is a poly(T) sequence or a specific sequence targeting a target nucleic acid, and preferably, the primer II-A further comprises a consensus sequence A and, optionally, a tag sequence A such as a random oligonucleotide sequence; 10. The method of claim 9, having one or more features selected from the group consisting of:

12. comprising step (1), step (2)(ii) and step (3), (I) In step (2)(ii)(b), the first extension product comprises, from 5' to 3', the consensus sequence A, a cDNA sequence formed by reverse transcription primed by the primer II-A' and complementary to RNA, the 3'-end overhang sequence, optionally a complementary sequence of the tag sequence B, and a complementary sequence of the consensus sequence B; (II) in step (2)(ii)(c), the extension primer is primer II-B' or primer B", and primer B" is capable of annealing to a complementary sequence of consensus sequence B or a partial sequence thereof to initiate an extension reaction; (III) in step (2)(ii)(a), the capture sequence A of the primer II-A' is a random oligonucleotide sequence, and preferably in step (3), the second nucleic acid molecule derived from each copy of the oligonucleotide probe bound to the same microdot has a different capture sequence A as a UMI; and (IV) In step (2)(ii)(a), the capture sequence A of the primer II-A' is a poly(T) sequence or a specific sequence targeting a target nucleic acid, and the primer II-A' further comprises a tag sequence A such as a random oligonucleotide sequence; and preferably, in step (3), the second nucleic acid molecules derived from each copy of the oligonucleotide probe bound to the same microdot have different tag sequences A as UMIs.

10. The method of claim 9, having one or more features selected from the group consisting of:

13. The method of claim 1, wherein in step (2), the pretreatment is performed intracellularly or extracellularly.

14. The nucleic acid array of step (1) is subjected to the following steps: (1) providing a plurality of types of carrier sequences, each type of carrier sequence comprising at least one copy of the carrier sequence, the carrier sequence comprising, in a 5' to 3' direction, a complementary sequence of the consensus sequence X2, a complementary sequence of the tag sequence Y, and an immobilization sequence, each type of carrier sequence having a different complementary sequence of the tag sequence Y; (2) binding the plurality of types of carrier sequences to the surface of the solid support (e.g., chip); (3) providing an immobilized primer and performing a primer extension reaction using the carrier sequence as a template to generate an extension product to obtain the oligonucleotide probe, wherein the immobilized primer comprises the sequence of the consensus sequence X1 and is capable of annealing to the immobilized sequence of the carrier sequence to initiate an extension reaction, and preferably the extension product comprises or consists of, in a 5' to 3' direction, the consensus sequence X1, the tag sequence Y, and the consensus sequence X2; (4) linking the immobilized primer to the surface of the solid support, wherein steps (3) and (4) are performed in any order; (5) Optionally, the immobilized sequence of the carrier sequence further comprises a cleavage site, and the cleavage may be selected from nicking enzyme digestion, USER enzyme digestion, light-responsive excision, chemical excision, or CRISPR-mediated excision, and cleaving at the cleavage site comprised in the immobilized sequence of the carrier sequence to digest the carrier sequence and separate the extension product in step (3) from the template (i.e., the carrier sequence) on which the extension product was generated, thereby linking the oligonucleotide probe to the surface of the solid support (e.g., chip). and preferably, the method further comprises separating the extension products in step (3) from the templates on which they are generated by high temperature denaturation; Preferably, each carrier sequence is a DNB formed from a concatemer of multiple copies of said carrier sequence, Preferably, the plurality of types of carrier sequences are prepared by the following steps: (i) providing a plurality of carrier-template sequences, wherein the carrier-template sequences comprise complementary sequences of carrier sequences; (ii) performing a nucleic acid amplification reaction using the various carrier-template sequences as templates to obtain amplification products of the various carrier-template sequences, wherein the amplification products comprise at least one copy of the carrier sequences, and preferably, rolling circle replication is performed to obtain DNBs formed from concatemers of the carrier sequences.

2. The method of claim 1, wherein the step (1) is provided by:

15. 1. A method for constructing a library of nucleic acid molecules, comprising: (a) generating a population of labeled nucleic acid molecules according to the method of claim 1; (b) randomly fragmenting the nucleic acid molecules in the population of labeled nucleic acid molecules and ligating adapters thereto; and (c) optionally amplifying and / or enriching the product of step (b); thereby obtaining a library of nucleic acid molecules, Preferably, the library of nucleic acid molecules comprises nucleic acid molecules from a plurality of single cells, wherein the nucleic acid molecules from different single cells have different tag sequences Y; Preferably, the method, wherein the library of nucleic acid molecules is used for sequencing, e.g., transcriptome sequencing, e.g., single-cell transcriptome sequencing.

16. 1. A method for transcriptome sequencing of cells in a sample, comprising: (1) constructing a library of nucleic acid molecules according to the method of claim 15; and (2) sequencing the library of nucleic acid molecules; A method comprising:

17. 1. A method for performing single cell transcriptome analysis, comprising: (1) performing transcriptome sequencing of a single cell in a sample according to the method of claim 16; and (2) A step of analyzing sequencing data, comprising matching the sequencing results of the sequencing library with the tag sequence Y or its complementary sequence of the oligonucleotide probe bound to each microdot of the nucleic acid array; if the matching is successful, the microdot is identified as a positive microdot; and sequencing data derived from positive microdots having regional continuity in the nucleic acid array are identified as transcription data of the same cell, thereby performing single-cell transcriptome analysis. A method comprising:

18. 1. A kit comprising a nucleic acid array for labeling nucleic acids and optionally a first binding molecule, the nucleic acid array comprises a solid support, the solid support comprises a first labeled molecule, and the first binding molecule is capable of forming an interactive pair with the first labeled molecule; the solid support further comprises a plurality of microdots, the size of the microdots being less than 5 μm, the center-to-center distance between adjacent microdots being less than 10 μm, each microdot being bound to one type of oligonucleotide probe, each type of oligonucleotide probe comprising at least one copy, the oligonucleotide probe comprising or consisting of, in a 5′ to 3′ direction, a consensus sequence X1, a tag sequence Y and a consensus sequence X2; A kit, wherein the oligonucleotide probes bound to different microdots have different tag sequences Y.

19. the center-to-center distance between adjacent microdots is less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm or less than 0.01 μm, and the size of the microdots is less than 5 μm, less than 1 μm, less than 0.3 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, less than 0.01 μm or less than 0.001 μm; Preferably, the center-to-center distance between adjacent microdots is between 0.5 μm and 1 μm, e.g., between 0.5 μm and 0.9 μm, e.g., between 0.5 μm and 0.8 μm; The kit according to claim 18, wherein the size of the microdots is preferably between 0.001 μm and 0.5 μm.

20. the first binding molecule is capable of forming a specific interaction pair or a non-specific interaction pair with the first labeled molecule; Preferably, the interaction pair is selected from the group consisting of positive and negative charge interaction pairs, affinity interaction pairs (e.g., biotin / avidin, biotin / streptavidin, antigen / antibody, receptor / ligand, enzyme / cofactor), pairs of molecules capable of undergoing click chemistry reactions (e.g., alkynyl-containing compound / azide compound), N-hydroxysulfosuccinate (NHS) ester / amino-containing compound, and any combination thereof; The kit of claim 18, wherein the first labeled molecule is polylysine and the first binding molecule is a protein capable of binding to polylysine; the first labeled molecule is an antibody and the first binding molecule is an antigen capable of binding to the antibody; the first labeled molecule is an amino-containing compound and the first binding molecule is an N-hydroxysulfosuccinate (NHS) ester; or the first labeled molecule is biotin and the first binding molecule is streptavidin.

21. (i) a primer set comprising primer IA, or primer IA′ and primer IB, or a primer set comprising primer IA and primer IB, the primer IA comprises a consensus sequence A and a capture sequence A, the capture sequence A is capable of annealing to the RNA to be captured to initiate an extension reaction, and preferably the consensus sequence A is located upstream of the capture sequence A; the primer IA' comprises a capture sequence A, which is capable of annealing to the RNA to be captured to initiate an extension reaction; a primer set comprising primer A, or primer A' and primer B, or primer set comprising primer A and primer B, wherein primer I-B comprises consensus sequence B, a complementary sequence of a 3'-end overhang, and optionally a tag sequence B, the complementary sequence of the 3'-end overhang being located at the 3' end of primer I-B, and the consensus sequence B being located upstream of the complementary sequence of the 3'-end overhang, and the 3'-end overhang refers to one or more non-template nucleotides contained at the 3' end of the cDNA strand generated by reverse transcription using the RNA captured by the capture sequence A of primer IA' as a template; and (ii) a bridge oligonucleotide I comprising a first region and a second region, and optionally a third region located between the first region and the second region, wherein the first region is located upstream of the second region; the first region is capable of (a) annealing to all or part of the consensus sequence A of the primer IA, or (b) annealing to all or part of the consensus sequence B of the primer IB; The second region is a bridging oligonucleotide I capable of annealing to all or part of the consensus sequence X2.

20. The kit of claim 18, further comprising:

22. (i) a primer set comprising primer II-A and primer II-B, or a primer set comprising primer II-A' and primer II-B'; the primer II-A comprises a capture sequence A, which is capable of annealing to the RNA to be captured to initiate an extension reaction; the primer II-B comprises a consensus sequence B, a complementary sequence of a 3'-end overhang, and optionally a tag sequence B, the complementary sequence of the 3'-end overhang being located at the 3' end of the primer II-B, and the consensus sequence B being located upstream of the complementary sequence of the 3'-end overhang, and the 3'-end overhang refers to one or more non-template nucleotides contained at the 3' end of a cDNA strand generated by reverse transcription using the RNA captured by the capture sequence A of the primer II-A as a template; the primer II-A' comprises a consensus sequence A and a capture sequence A, the capture sequence A is located at the 3' end of the primer II-A', and the consensus sequence A is located upstream of the capture sequence A; 19. The kit of claim 18, wherein the primer II-B' comprises a consensus sequence B, a complementary sequence of a 3'-end overhang, and optionally a tag sequence B, the complementary sequence of the 3'-end overhang is located at the 3'-end of the primer II-B', the consensus sequence B is located upstream of the complementary sequence of the 3'-end overhang, and the 3'-end overhang refers to one or more non-template nucleotides contained at the 3'-end of a cDNA strand generated by reverse transcription using the RNA captured by the capture sequence A of the primer II-A' as a template. (ii) further comprising a bridging oligonucleotide II-I and a bridging oligonucleotide II-II; the bridging oligonucleotide II-I and the bridging oligonucleotide II-II each independently comprise a first region and a second region, and optionally a third region located between the first region and the second region, the first region being located upstream of the second region; the first region of the bridging oligonucleotide II-I is capable of annealing to the first region of the bridging oligonucleotide II-II, and the second region of the bridging oligonucleotide II-I is capable of annealing to the consensus sequence X2 of the oligonucleotide probe or a subsequence thereof; the second region of the bridging oligonucleotide II-II is capable of annealing to (a) the consensus sequence B of the primer II-B, or (b) a complementary sequence or a subsequence thereof of the consensus sequence A of the primer II-A'; 23. The kit of claim 22.

24. Use of the method according to any one of claims 1 to 14 or the kit according to any one of claims 18 to 23 for constructing a library of nucleic acid molecules or for performing transcriptome sequencing, comprising: Preferably, the method or kit is used to construct a single-cell library of nucleic acid molecules or to perform single-cell transcriptome sequencing.