Method and kit for generating a population of labeled nucleic acid molecules

JP7918267B2Active Publication Date: 2026-09-09STOMICS TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2024538195
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2026-09-09
Estimated Expiration
2041-12-24

Smart Images

  • Figure 0007918267000010
    Figure 0007918267000010
  • Figure 0007918267000011
    Figure 0007918267000011
  • Figure 0007918267000012
    Figure 0007918267000012
Patent Text Reader

Abstract

The present application relates to transcriptome sequencing and biomolecular spatial information detection. In particular, the present application relates to a method for positioning and labeling nucleic acid molecules, a method for constructing a nucleic acid molecule library for transcriptome sequencing, and a kit for carrying out the method.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to the technical fields of transcriptome sequencing and biomolecular spatial information detection. More specifically, this application relates to a method for positionally labeling nucleic acid molecules and a method for constructing nucleic acid molecular libraries for transcriptome sequencing. Furthermore, this application also relates to nucleic acid molecular libraries constructed by the present method and kits for carrying out the present method. [Background technology]

[0002] The spatial location of cells in a tissue significantly impacts their function. To investigate this spatial heterogeneity, it is necessary to quantify and analyze the cellular genome or transcriptome using knowledge of spatial coordinates. However, collecting small tissue regions or even single cells for genome or transcriptome analysis is extremely laborious, costly, and inaccurate. Therefore, there is a need to develop methods that can achieve high-rate detection of spatial information of biomolecules (e.g., nucleic acid localization, distribution, and / or expression) at the single-cell level, or even at the intracellular level. [Brief explanation of the drawing]

[0003] [Figure 1] Figure 1 shows an exemplary structure of a chip used in this application to capture and label nucleic acid molecules, and includes the chip and an oligonucleotide probe (also called a chip sequence) bound to the chip. Each oligonucleotide probe includes a tag sequence Y corresponding to its position on the chip, and the region on the chip to which each oligonucleotide probe is bound may be called a microdot. Each oligonucleotide probe may contain a single copy or multiple copies. [Figure 2] Figure 2 shows an exemplary scheme for preparing a cDNA strand using sample RNA (e.g., mRNA) as a template, as well as an exemplary structure of the cDNA strand. CA: Consensus sequence A, CB: Consensus sequence B. [Figure 3] Figure 3 shows an exemplary scheme for generating a novel nucleic acid molecule containing information from a chip sequence (i.e., a nucleic acid molecule labeled by a chip sequence) by labeling the 5' end of a cDNA strand with a chip sequence (i.e., ligating the 5' end of the cDNA strand to the 3' end of the chip sequence), as well as an exemplary structure of the novel nucleic acid molecule containing chip sequence information. CA: Consensus sequence A, CB: Consensus sequence B, X1: Consensus sequence X1, Y: Tag sequence Y, X2: Consensus sequence X2. [Figure 4] Figure 4 shows an exemplary scheme for preparing a complementary strand of a cDNA chain using sample RNA (e.g., mRNA) as a template, and an exemplary structure of the complementary strand of the cDNA chain. CA: Consensus sequence A, CB: Consensus sequence B, EP: Extension primer. [Figure 5] Figure 5 shows an exemplary scheme for generating a novel nucleic acid molecule containing information from a chip sequence (i.e., a nucleic acid molecule labeled by a chip sequence) by labeling the 5'-end of the complementary strand of a cDNA strand with a chip sequence (i.e., ligating the 5'-end of the complementary strand of a cDNA strand to the 3'-end of the chip sequence), as well as an exemplary structure of the novel nucleic acid molecule containing chip sequence information. CA: Consensus sequence A, CB: Consensus sequence B, X1: Consensus sequence X1, Y: Tag sequence Y, X2: Consensus sequence X2. [Figure 6] Figure 6 shows the length distribution of the cDNA amplification product prepared in Example 2. [Figure 7] Figure 7 shows a mouse brain gene expression map obtained by the method of this application. [Figure 8] Figure 8 shows an enlarged mouse brain gene expression map obtained by the method of this application. [Figure 9] Figure 9 shows the number of genes captured and the UMI analysis results using the method of this application. [Modes for carrying out the invention]

[0004] This application provides a novel method for generating a population of labeled nucleic acid molecules, a method for constructing a nucleic acid molecule library based on this method, and a method for performing high-rate sequencing.

[0005] Method for generating a collection of labeled nucleic acid molecules

[0006] In one embodiment, the present application relates to a method for generating a population of labeled nucleic acid molecules, the following (1) A step of providing a biological sample and a nucleic acid array, wherein the nucleic acid array includes a solid support, the solid support is bound to a plurality of types of oligonucleotide probes, each type of oligonucleotide probe includes at least one copy, and the oligonucleotide probe includes or consists of a consensus sequence X1, a tag sequence Y and a consensus sequence X2 in the 5' to 3' direction. Each type of oligonucleotide probe has a different tag sequence Y, and the tag sequence Y has a nucleotide sequence specific to the position of that type of oligonucleotide probe on the solid support, step, (2) A step of bringing a biological sample into contact with a nucleic acid array so that the position of RNA (e.g., mRNA) in the biological sample is mapped to the position of oligonucleotide probes on the nucleic acid array, and pre-treating the RNA (e.g., mRNA) in the biological sample to generate a first population of nucleic acid molecules, wherein the pre-treatment is (i) A step of reverse transcription of RNA (e.g., mRNA) of a biological sample using primer A to generate an elongation product as a first nucleic acid molecule to be labeled, thereby generating a first population of nucleic acid molecules, wherein primer A comprises a consensus sequence A and a capture sequence A, the capture sequence A is capable of annealing with the RNA to be captured (e.g., mRNA) to initiate the elongation reaction, and the consensus sequence A is located upstream of the capture sequence A (e.g., located at the 5' end of primer A), or (ii)(a) A step of generating a cDNA strand by performing reverse transcription of RNA (e.g., mRNA) of a biological sample using primer A, wherein the cDNA strand comprises a cDNA sequence that is generated by reverse transcription primed by primer A and is complementary to RNA (e.g., mRNA), and a 3'-terminal overhang, wherein primer A comprises a consensus sequence A and a capture sequence A, wherein the capture sequence A is capable of annealing with the captured RNA (e.g., mRNA) to initiate an extension reaction, and the consensus sequence A is upstream of the capture sequence A (e.g., at the 5' end of primer A) Step (b) Annealing primer IB with the cDNA strand generated in (a) to carry out an extension reaction to produce a first extension product as a first nucleic acid molecule to be labeled, thereby generating a first population of nucleic acid molecules, wherein primer B comprises a consensus sequence B, a complementary sequence for the 3'-terminal overhang, and optionally a tag sequence B, the complementary sequence for the 3'-terminal overhang located at the 3'-end of primer B, and the consensus sequence B located upstream of the complementary sequence for the 3'-terminal overhang (for example, located at the 5'-end of primer B), or (iii) (a) using primer A', performing reverse transcription of RNA (e.g., mRNA) from a biological sample to generate a cDNA strand, wherein the cDNA strand is generated by reverse transcription primed by primer A', and comprises a cDNA sequence complementary to RNA (e.g., mRNA) and a 3'-terminal overhang, primer A' comprises a capture sequence A, and the capture sequence A is capable of annealing with the RNA (e.g., mRNA) to be captured and initiating an extension reaction, (b) annealing primer B to the cDNA strand generated in (a), and performing an extension reaction to generate a first extension product, wherein primer B comprises a consensus sequence B, a complementary sequence of the 3'-terminal overhang, and optionally a tag sequence B, the complementary sequence of the 3'-terminal overhang is located at the 3'-end of primer B, and the consensus sequence B is located upstream of the complementary sequence of the 3'-terminal overhang (e.g., located at the 5'-end of primer B), and (c) providing an extension primer, performing an extension reaction using the first extension product as a template to generate a second extension product as a first nucleic acid molecule to be labeled, thereby generating a first population of nucleic acid molecules comprising the step, (3) contacting a bridging oligonucleotide with the product of step (2) under conditions allowing annealing, annealing the bridging oligonucleotide (e.g., in situ annealing) to an oligonucleotide probe and a first nucleic acid molecule to be labeled located at a position corresponding to the oligonucleotide probe, and ligating the first nucleic acid molecule annealed to the bridging oligonucleotide and the oligonucleotide probe on the array to obtain a ligation product as a second nucleic acid molecule having a position tag, thereby generating a second population of nucleic acid molecules providing a method comprising, the bridging oligonucleotide comprises a first region, a second region, and optionally a third region located between the first region and the second region, the first region is located upstream of the second region (e.g., located 5' to the second region), The first region is capable of annealing to all or part of the consensus sequence A of primer A in step (2)(i) or step (2)(ii), or is capable of annealing to all or part of the consensus sequence B of primer B in step (2)(iii), The second region is capable of annealing to all or part of the consensus sequence X2.

[0007] In a specific embodiment, in step (3) of the method, when the first region and the second region of the bridging oligonucleotide are directly adjacent, the ligation of the first nucleic acid molecule to the oligonucleotide probe comprises using a nucleic acid ligase to ligate the nucleic acid molecule hybridized to the first region and the nucleic acid molecule hybridized to the second region of the same bridging oligonucleotide, to obtain a ligation product as the second nucleic acid molecule having a positioning tag, or When the bridging oligonucleotide comprises a first region, a second region, and a third region located between the foregoing two regions, the ligation of the first nucleic acid molecule to the oligonucleotide probe comprises carrying out a polymerization reaction using a nucleic acid polymerase with the third region as a template, and using a nucleic acid ligase to ligate the nucleic acid molecule hybridized to the first region of the same bridging oligonucleotide and the nucleic acid molecule hybridized to the third region and the second region, to obtain a ligation product as the second nucleic acid molecule having a positioning tag. In a specific embodiment, the nucleic acid polymerase does not have 5' to 3' exonuclease activity or strand displacement activity.

[0008] In a specific embodiment, each type of oligonucleotide probe comprises one copy.

[0009] In a specific embodiment, each type of oligonucleotide probe comprises a plurality of copies.

[0010] In a specific embodiment, the regions where various oligonucleotide probes bind to the solid support are called microdots. When the various oligonucleotide probes each comprise one copy, each microdot is bound to one oligonucleotide probe, and the oligonucleotide probes on different microdots have different tag sequences Y. It is easy to understand that when the various oligonucleotide probes comprise multiple copies, each microdot is bound to multiple oligonucleotide probes, the oligonucleotide probes on the same microdot have the same tag sequence Y, and the oligonucleotide probes on different microdots have different tag sequences Y.

[0011] In a specific embodiment, the solid support comprises a plurality of microdots, each microdot is bound to one type of oligonucleotide probe, and each of the various oligonucleotide probes may comprise one or more copies.

[0012] In a specific embodiment, the solid support comprises a plurality of (e.g., at least 10, at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 or more) microdots. In a specific embodiment, the solid support comprises at least 10 4 (e.g., at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 or at least 10 12 ) microdots / mm 2 .

[0013] In some embodiments, the spacing between adjacent microdots is less than 100 μm, less than 50 μm, less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, or less than 0.01 μm.

[0014] In certain embodiments, the microdots are of a size (e.g., equivalent diameter) of less than 100 μm, less than 50 μm, less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, or less than 0.01 μm.

[0015] Embodiments including step (1), step (2)(i), and step (3)

[0016] In some embodiments, the method comprises steps (1), (2)(i), and (3), wherein the ligation product obtained in step (3) is taken as a second nucleic acid molecule having a positioning tag from 5' to 3' that includes a consensus sequence X1, a tag sequence Y, a consensus sequence X2, optionally a complementary sequence of a third region of a crosslinked oligonucleotide, and the sequence of the first nucleic acid molecule to be labeled.

[0017] In a particular embodiment, in step (2)(i) of the method, the capture sequence A is a random oligonucleotide sequence.

[0018] In some embodiments, in step (3), each ligation product derived from a copy of the same type of oligonucleotide probe has a different capture sequence A, which acts as a unique molecular identifier (UMI) for the second nucleic acid molecule.

[0019] In a particular embodiment, the extension product (the first nucleic acid molecule to be labeled) in step (2)(i) comprises, from 5' to 3', a consensus sequence A, a cDNA sequence generated by reverse transcription primed with primer A and complementary to RNA.

[0020] In a particular embodiment, in step (2)(i) of the method, the capture sequence A is a poly(T) sequence or a specific sequence that targets a target nucleic acid.

[0021] In a particular embodiment, primer A further comprises a tag sequence A, for example, a random oligonucleotide sequence, where tag sequence A acts as a unique molecular identifier (UMI) for a second nucleic acid molecule.

[0022] In a particular embodiment, the capture sequence A is located at the 3' end of primer A, and the consensus sequence A is located upstream of the tag sequence A (for example, at the 5' end of primer A).

[0023] In a particular embodiment, in step (3), the ligation products derived from each copy of the same type of oligonucleotide probe have different tag sequences A as UMIs.

[0024] In some embodiments, the extension product described in step (2)(i) comprises, from 5' to 3', a consensus sequence A, a tag sequence A, and a cDNA sequence that is generated by reverse transcription primed with primer A and is complementary to the RNA.

[0025] Embodiments including step (1), step (2)(ii), and step (3)

[0026] In some embodiments, the method comprises steps (1), (2)(ii), and (3), wherein the ligation product obtained in step (3) is taken as a second nucleic acid molecule having a positioning tag from 5' to 3' that includes a consensus sequence X1, a tag sequence Y, a consensus sequence X2, optionally a complementary sequence of a third region of a crosslinked oligonucleotide, and the sequence of the first nucleic acid molecule to be labeled.

[0027] In a particular embodiment, in step (2)(ii)(a) of the method, the capture sequence A is a random oligonucleotide sequence.

[0028] In some embodiments, in step (3), each ligation product derived from a copy of the same type of oligonucleotide probe has a different capture sequence A, which acts as a unique molecular identifier (UMI) for the second nucleic acid molecule.

[0029] In a particular embodiment, the first extension product (the first nucleic acid molecule to be labeled) in step (2)(ii) comprises, from 5' to 3', a consensus sequence A, a cDNA sequence generated by reverse transcription primed with primer A and complementary to RNA, a 3'-terminal overhang sequence, optionally a complementary sequence to tag sequence B, and a complementary sequence to consensus sequence B.

[0030] In a particular embodiment, in step (2)(ii)(a), the capture sequence A is a poly(T) sequence or a specific sequence that targets a target nucleic acid.

[0031] In a particular embodiment, primer A further comprises a tag sequence A, for example, a random oligonucleotide sequence, where tag sequence A acts as a unique molecular identifier (UMI) for a second nucleic acid molecule.

[0032] In a particular embodiment, the capture sequence A is located at the 3' end of primer A, and the consensus sequence A is located upstream of the tag sequence A (for example, at the 5' end of primer A).

[0033] In a particular embodiment, in step (3), the ligation products derived from each copy of the same type of oligonucleotide probe have different tag sequences A as UMIs.

[0034] In a particular embodiment, the first extension product (the first nucleic acid molecule to be labeled) in step (2)(ii) comprises, from 5' to 3', a consensus sequence A, a tag sequence A, a cDNA sequence generated by reverse transcription primed with primer A and complementary to RNA, a 3'-terminal overhang sequence, optionally a complementary sequence to tag sequence B, and a complementary sequence to consensus sequence B.

[0035] In a particular embodiment, in the method, primer A contains a 5'-phosphate at its 5'-terminus.

[0036] In certain embodiments, prior to step (3), the method further includes a step of processing the product of step (2)(i) or step (2)(ii) to remove RNA (e.g., a heat treatment step).

[0037] Exemplary embodiments of this application, including steps (1), (2)(ii), and (3), are described in detail below:

[0038] 1. An exemplary embodiment of preparing a cDNA strand using sample RNA (e.g., mRNA) as a template includes the following steps (as shown in Figure 2):

[0039] (1) Using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyltransferase activity) and primer A, RNA molecules (e.g., mRNA molecules) from a permeabilized tissue sample are reverse transcribed to produce cDNA, and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. Various reverse transcriptases having terminal deoxynucleotidyltransferase activity may be used to carry out the reverse transcription reaction. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity. Includes.

[0040] Primer A comprises a poly(T) sequence and a consensus sequence A (labeled CA in the figure). In certain embodiments (for example, when the method is used to construct a 3' transcriptome library), primer A further comprises a unique molecular identifier (UMI). Typically, the poly(T) sequence is located at the 3' end of primer A to initiate reverse transcription. In preferred embodiments, the UMI sequence is located upstream of the poly(T) sequence (e.g., 5'), and the consensus sequence A is located upstream of the UMI sequence (e.g., 5').

[0041] (2) A template switch oligo (TSO, which can be used as primer B and contains consensus sequence B (indicated as CB in the figure)) is used to anneal or hybridize with a cDNA strand. Subsequently, the nucleic acid fragment hybridized with or annealed with TSO is extended in the presence of nucleic acid polymerase using consensus sequence B as a template, thereby adding the complementary sequence of consensus sequence B to the 3' end of the cDNA strand. This produces a nucleic acid molecule that holds consensus sequence A and tag sequence A at the 5' end and the complementary sequence of consensus sequence B at the 3' end.

[0042] The TSO sequence may contain a sequence complementary to the 3'-terminal overhang of the cDNA strand. For example, if the cDNA strand contains an overhang of three cytosine nucleotides at its 3'-terminus, the template switch oligo may contain GGG at its 3'-terminus. Furthermore, the nucleotides in the TSO sequence may also be modified to enhance the binding affinity for complementary pair formation between the TSO sequence and the 3'-terminal overhang of the cDNA strand (for example, the TSO sequence may be modified to contain locked nucleic acids).

[0043] While not limited by any particular theory, the extension reaction can be carried out using a variety of suitable nucleic acid polymerases (e.g., DNA polymerase or reverse transcriptase) as long as the hybridized or annealed nucleic acid fragment (reverse transcript) can be extended using the TSO sequence or a subsequence thereof as a template. In certain exemplary embodiments, the hybridized or annealed nucleic acid fragment (reverse transcript) may be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0044] In a particular preferred embodiment, steps (2) and (1) are performed simultaneously.

[0045] In certain embodiments, the method optionally further includes step (3): adding RNase H to digest the RNA strand of the RNA / cDNA hybrid to form a single cDNA strand.

[0046] In a particular preferred embodiment, the method does not include step (3).

[0047] The exemplary structure of the cDNA strand prepared by the above exemplary embodiment includes a consensus sequence A, a UMI sequence, a sequence complementary to the RNA (e.g., mRNA) sequence, and a complementary sequence to consensus sequence B.

[0048] 2. An exemplary embodiment of forming a novel nucleic acid molecule containing chip sequence information (i.e., a nucleic acid labeled using a chip sequence) by labeling the 5' end of a cDNA strand with an oligonucleotide probe (also called a chip sequence) (i.e., ligating the 5' end of the cDNA strand to the 3' end of the chip sequence) includes the following steps (shown in Figure 3): The step of providing a crosslinked oligonucleotide having a 5'-terminus containing a sequence (first region, P1) that is at least partially complementary to the 5'-terminus of the cDNA sequence (e.g., at least partially complementary to consensus sequence A(CA)) and a 3'-terminus containing a sequence (second region, P2) that is at least partially complementary to the 3'-terminus of the chip sequence (e.g., at least partially complementary to consensus sequence X2).

[0049] In a particular preferred embodiment, the P1 and P2 sequences of the crosslinked oligonucleotide are directly adjacent, with no intermediate nucleotides between them.

[0050] In a particular preferred embodiment, the P1 sequence, P2 sequence, consensus sequence A, and consensus sequence X2 each independently have a length of 20 to 100 nt (e.g., 20 to 70 nt). The crosslinked oligonucleotide is annealed or hybridized with the oligonucleotide probe and the cDNA strand, and then the 5'-end of the cDNA strand is ligated to the 3'-end of the oligonucleotide probe by a DNA ligase and / or DNA polymerase, thereby generating a novel nucleic acid molecule containing the sequence information of the oligonucleotide probe (i.e., a nucleic acid molecule labeled with the oligonucleotide probe). In a particular preferred embodiment, the DNA polymerase does not have 5'-to-3' exonuclease activity or strand displacement activity.

[0051] An exemplary structure of a novel nucleic acid molecule containing chip sequence information, as formed by the exemplary embodiment described above, includes a consensus sequence X1, a tag sequence Y, a consensus sequence X2, a consensus sequence A, a UMI sequence, a complementary sequence of an RNA (e.g., mRNA) sequence, and a complementary sequence of a consensus sequence B.

[0052] Embodiments including step (1), step (2)(iii), and step (3)

[0053] In some embodiments, the method comprises steps (1), (2)(iii), and (3), wherein the ligation product obtained in step (3) is taken as a second nucleic acid molecule having a positioning tag from 5' to 3' that includes a consensus sequence X1, a tag sequence Y, a consensus sequence X2, optionally a complementary sequence of a third region of a crosslinked oligonucleotide, and the sequence of the first nucleic acid molecule to be labeled.

[0054] In some embodiments, in step (2)(iii)(c) of the method, the extension primer is primer B or primer B', and primer B can anneal to all or part of the complementary sequence of consensus sequence B to initiate the extension reaction.

[0055] In a particular embodiment, in step (2)(iii)(c), the extension primer is primer B'.

[0056] In a particular embodiment, in step (2)(iii)(a) of the method, the capture sequence A of primer A' is a random oligonucleotide sequence.

[0057] In a particular embodiment, in step (2)(iii)(b), primer B includes consensus sequence B, a complementary sequence for the 3' end overhang, and tag sequence B.

[0058] In some embodiments, the first extension product comprises a cDNA sequence, a 3'-terminal overhang sequence, a complementary sequence of tag sequence B, and a complementary sequence of consensus sequence B, which are generated by reverse transcription primed with primer A' from 5' to 3' and are complementary to the RNA sequence, with the complementary sequence of tag sequence B acting as a unique molecular identifier (UMI) for the second nucleic acid molecule.

[0059] In a particular embodiment, in step (2)(iii)(c), the second extension product (the labeled first nucleic acid molecule) comprises, from 5' to 3', a consensus sequence B or its 3'-terminal subsequence, a tag sequence B, a complementary sequence of the 3'-terminal overhang sequence, and a complementary sequence of the cDNA sequence, wherein the tag sequence B serves as a unique molecular identifier (UMI) of the second nucleic acid molecule.

[0060] In a particular embodiment, in step (2)(iii)(a) of the method, the capture sequence A of primer A' is a poly(T) sequence or a specific sequence that targets the target nucleic acid.

[0061] In a particular embodiment, primer A' further comprises a tag sequence A, for example, a random oligonucleotide sequence, and a consensus sequence A.

[0062] In a particular embodiment, capture sequence A is located at the 3'-end of primer A'.

[0063] In a particular embodiment, consensus sequence A is located upstream of capture sequence A (for example, at the 5'-end of primer A').

[0064] In a particular embodiment, primer B includes a consensus sequence B, a complementary sequence for the 3'-terminal overhang, and a tag sequence B.

[0065] In a particular embodiment, in step (2)(iii)(b), the first extension product comprises, from 5' to 3', a consensus sequence A, optionally a tag sequence A, a cDNA sequence generated by reverse transcription primed with primer A' and complementary to RNA, a 3'-terminal overhang sequence, a complementary sequence to tag sequence B, and a complementary sequence to consensus sequence B.

[0066] In a particular embodiment, in step (2)(iii)(c), the second extension product (the labeled first nucleic acid molecule) comprises, from 5' to 3', a consensus sequence B or its 3'-terminal subsequence, a tag sequence B, a complementary sequence to the 3'-terminal overhang sequence, a complementary sequence to the cDNA sequence in the first extension product, and optionally, a complementary sequence to the tag sequence A and a complementary sequence to the consensus sequence A.

[0067] In a particular embodiment, in step (3), the ligation products derived from each copy of the same type of oligonucleotide probe have different tag sequences B as UMIs.

[0068] In certain embodiments, the extension primer contains a 5'-phosphate at its 5'-terminus.

[0069] In certain embodiments, prior to step (2)(iii)(c), the method further includes a step of processing the product of step (2)(iii)(a) or step (2)(iii)(b) to remove RNA (e.g., a heat treatment step).

[0070] In a particular embodiment, in step (2)(iii)(b) of the method, a cDNA strand is annealed to primer B via its 3'-terminal overhang, and the cDNA strand is extended using primer B as a template in the presence of a nucleic acid polymerase (e.g., DNA polymerase or reverse transcriptase) to produce a first extension product.

[0071] Exemplary embodiments of this application, including steps (1), (2)(iii), and (3), are described in detail below:

[0072] 1. An exemplary embodiment of preparing a complementary strand of a cDNA strand using RNA (e.g., mRNA) in a sample as a template includes the following steps (as shown in Figure 4):

[0073] (1) To perform reverse transcription of RNA molecules (e.g., mRNA molecules) from a permeabilized sample, cDNA is generated using a reverse transcriptase (e.g., a reverse transcriptase having terminal deoxynucleotidyltransferase activity) and primer A', and an overhang (e.g., an overhang containing three cytosine nucleotides) is added to the 3' end of the cDNA. Various reverse transcriptases having terminal deoxynucleotidyltransferase activity can be used to carry out the reverse transcription reaction. In certain preferred embodiments, the reverse transcriptase used does not have RNase H activity.

[0074] Reverse transcription primer A' contains a poly(T) sequence and a consensus sequence A(CA). Typically, the poly(T) sequence is located at the 3' end of primer A to initiate reverse transcription.

[0075] (2) Primer B is annealed or hybridized with the cDNA strand, and primer B contains consensus sequence B (CB) and a complementary sequence to the 3'-terminal overhang of the cDNA. In certain embodiments (for example, when the method is used to construct a 5' transcriptome library), primer B further contains a unique molecular identifier (UMI). The nucleic acid fragment hybridized or annealed with primer B is then extended in the presence of nucleic acid polymerase using consensus sequence B and the UMI sequence as templates, so that the complementary sequences of consensus sequence B and the UMI sequence are added to the 3'-terminus of the cDNA strand, thereby producing a nucleic acid molecule that holds consensus sequence A at the 5'-terminus and the complementary sequences of consensus sequence B and the UMI molecule at the 3'-terminus.

[0076] If the cDNA strand contains an overhang of three cytosine nucleotides at its 3' end, primer B may contain GGG at its 3' end. Furthermore, the nucleotides of primer B may also be modified to enhance the binding affinity for complementary pair formation between primer B and the 3'-terminal overhang of the cDNA strand (for example, primer B may be modified to contain locked nucleic acids).

[0077] While not limited by theory, the extension reaction can be carried out using a variety of suitable nucleic acid polymerases (e.g., DNA polymerase or reverse transcriptase) as long as they can extend the annealed or hybridized nucleic acid fragment (reverse transcript) using the sequence or subsequence of primer B as a template. In certain exemplary embodiments, the annealed or hybridized nucleic acid fragment (reverse transcript) may be extended using the same reverse transcriptase used in the reverse transcription step described above.

[0078] In a particular preferred embodiment, steps (2) and (1) are performed simultaneously.

[0079] In certain embodiments, the method optionally further includes step (3): adding RNase H to digest the RNA strand of the RNA / cDNA hybrid to form a single cDNA strand.

[0080] In a particular preferred embodiment, the method does not include step (3).

[0081] (4) Using the extension primer, the single-stranded cDNA obtained in (3) is used as a template to carry out the extension reaction and obtain the extension product. The extension primer can anneal to all or part of the complementary sequence of consensus sequence B and initiate the extension reaction.

[0082] In certain embodiments, the extension primer is identical to the template switch oligo.

[0083] An exemplary structure, including a complementary strand of the cDNA strand prepared by the above exemplary embodiment, includes consensus sequence B, UMI sequence, complementary sequence of the cDNA 3'-terminal overhang sequence, complementary sequence of the cDNA sequence, and complementary sequence of consensus sequence A.

[0084] 2. An exemplary embodiment of generating a novel nucleic acid molecule containing information from a chip sequence (i.e., a nucleic acid molecule labeled with a chip sequence) by using an oligonucleotide probe (also called a chip sequence) to label the 5'-end of the complementary strand of a cDNA strand (i.e., the 5'-end of the complementary strand of the cDNA strand is ligated with the 3'-end of the chip sequence) includes the following steps (as shown in Figure 5): The step of providing a crosslinked oligonucleotide comprising a sequence (first region, P1) at its 5' end that is at least partially complementary to consensus sequence B(CB), and a sequence (second region, P2) at its 3' end that is at least partially complementary to consensus sequence X2.

[0085] In a particular preferred embodiment, the P1 and P2 sequences of the crosslinked oligonucleotide are directly adjacent, with no intermediate nucleotides between them.

[0086] In a particular preferred embodiment, the P1 sequence and the P2 sequence each independently have a length of 20 to 100 nt (e.g., 20 to 70 nt).

[0087] The crosslinked oligonucleotide is annealed or hybridized with the oligonucleotide probe and the complementary strand of the cDNA chain, and then the 5'-end of the complementary strand of the cDNA chain is ligated to the 3'-end of the tip sequence by DNA ligase and / or DNA polymerase to form a novel nucleic acid molecule containing the sequence information of the oligonucleotide probe (i.e., a nucleic acid molecule labeled with the oligonucleotide probe). In certain preferred embodiments, the DNA polymerase does not have 5'-to-3' exonuclease activity or strand displacement activity.

[0088] An exemplary structure of a novel nucleic acid molecule containing chip sequence information formed by the above exemplary embodiment includes a consensus sequence X1, a tag sequence Y, a consensus sequence X2, a consensus sequence B, a UMI sequence, a complementary sequence of the cDNA sequence, and a complementary sequence of the consensus sequence A.

[0089] In certain embodiments, the 3'-terminal overhang has a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides. In certain embodiments, the 3'-terminal overhang is a 3'-terminal overhang of 2 to 5 cytosine nucleotides (e.g., a CCC overhang).

[0090] In some embodiments, in step (2) of the method, the biological sample is permeabilized before pretreatment.

[0091] In certain embodiments, the biological sample is a tissue sample.

[0092] In a particular embodiment, the tissue sample is a tissue section.

[0093] In certain embodiments, tissue sections are prepared from fixed tissue, such as formalin-fixed paraffin-embedded (FFPE) tissue or rapidly frozen tissue.

[0094] In a particular embodiment, when a biological sample is brought into contact with a nucleic acid array, each cell in the biological sample occupies one or more microdots of the nucleic acid array individually (i.e., each cell is in contact with one or more microdots of the nucleic acid array individually).

[0095] In some embodiments, reverse transcription is performed in step (2) by using reverse transcriptase.

[0096] In certain embodiments, the reverse transcriptase has terminal deoxynucleotidyltransferase activity.

[0097] In certain embodiments, the reverse transcriptase can synthesize a cDNA strand using RNA (e.g., mRNA) as a template and add an overhang to the 3' end of the cDNA strand.

[0098] In certain embodiments, the reverse transcriptase can add an overhang of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides in length to the 3' end of the cDNA strand.

[0099] In certain embodiments, reverse transcriptase can add an overhang of 2-5 cytosine nucleotides (e.g., a CCC overhang) to the 3' end of a cDNA strand.

[0100] In certain embodiments, the reverse transcriptase is selected from the group consisting of M-MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and variants, modified products, and derivatives thereof having reverse transcription activity.

[0101] In a particular embodiment, steps (2) and (3) of the method have one or more features selected from the following: (1) Primer A, Primer A', Primer B and crosslinking oligonucleotide each independently contain or consist of a native nucleotide (e.g., deoxyribonucleotide or ribonucleotide), a modified nucleotide, a non-native nucleotide, or any combination thereof. In certain embodiments, Primer A and Primer A' are capable of initiating the extension reaction. (2) Primer B contains a modified nucleotide (e.g., locked nucleic acid). In certain embodiments, primer B contains one or more modified nucleotides (e.g., locked nucleic acid) at its 3' end. (3) Tag sequence A and tag sequence B each have a length of 5 to 200 nt (e.g., 5 to 30 nt, 6 to 15 nt), (4) Consensus sequence A and consensus sequence B each independently have lengths of 10 to 200 nt (for example, 10 to 100 nt, 20 to 100 nt, 25 to 100 nt, 5 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt). (5) Each primer A, primer A', and primer B independently have a length of 10 to 200 nt (for example, 10 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt). (6) The first and second regions of the cross-linked oligonucleotide each independently have a length of 3 to 100 nt (e.g., 3 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt). (7) The third region of the cross-linked oligonucleotide has a length of 0 to 100 nt (e.g., 0 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt). (8) The cross-linked oligonucleotides have a length of 6 to 200 nt (for example, 20 to 70 nt, 6 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt), (9) The poly(T) sequence contains at least 10 or at least 20 (e.g., at least 30) deoxythymidine residues, (10) The random oligonucleotide sequence has a length of 5 to 200 nt (e.g., 5 to 30 nt, 6 to 15 nt).

[0102] In a particular embodiment, the method further includes the step of (4) recovering and purifying the second population of nucleic acid molecules.

[0103] In a particular embodiment, the method uses the obtained second population of nucleic acid molecules and / or their complements to construct a transcriptome library or for transcriptome sequencing.

[0104] In a particular embodiment, the oligonucleotide probe in step (1) has one or more features selected from the following: (1) The consensus sequence X1, tag sequence Y, and consensus sequence X2 each independently comprise a native nucleotide (e.g., deoxyribonucleotide or ribonucleotide), a modified nucleotide, a non-native nucleotide (e.g., peptide nucleic acid (PNA) or locked nucleic acid) or any combination thereof, and in a particular embodiment, the consensus sequence X2 has a free hydroxyl group (-OH) at its 3'-terminus. (2) The consensus sequence X1, the tag sequence Y, and the consensus sequence X2 each have a length of 2 to 100 nt (for example, 10 to 200 nt, 10 to 100 nt, 20 to 100 nt, 25 to 100 nt, 50 to 100 nt, 5 to 30 nt, 6 to 15 nt, 5 to 10 nt, 10 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt). (3) Each oligonucleotide probe independently has a length of 15-200 nt (e.g., 15-20 nt, 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-150 nt, 150-200 nt).

[0105] In a particular embodiment, the oligonucleotide probe is bonded to a solid support via a linker.

[0106] In a particular embodiment, the linker is a linking group capable of bonding with an activating group, and the surface of the solid support is modified with the activating group.

[0107] In certain embodiments, the linker includes -SH, -DBCO, or -NHS.

[0108] In a particular embodiment, the linker is -DBCO, and the surface of the solid-phase support is [ka] It is modified with (Azido-dPEG (registered trademark) 8-NHS ester).

[0109] In some embodiments, the nucleic acid array of step (1) has one or more features selected from the following:

[0110] In a particular embodiment, oligonucleotide probes bound to the same solid support have the same consensus sequence X1 and / or the same consensus sequence X2.

[0111] In certain embodiments, the consensus sequence X1 of the oligonucleotide probe includes a cleavage site, and in some embodiments, the cleavage site can be cleaved or destroyed by a method selected from Nicking enzyme digestion, USER enzyme digestion, photoresponsive excision, chemical excision, or CRISPR-mediated excision.

[0112] In some specific embodiments, the nucleic acid array of step (1) is provided by the following steps: (1) A step of providing multiple types of carrier sequences, wherein each type of carrier sequence comprises at least one copy (e.g., multiple copies) of a carrier sequence, and each carrier sequence comprises, in the 5' to 3' direction, a complementary sequence of the consensus sequence X2, a complementary sequence of the tag sequence Y, and a fixed sequence, and the complementary sequences of the tag sequence Y of each type of carrier sequence are different from each other. (2) The step of bonding multiple types of carrier arrays to the surface of a solid support (e.g., a chip), (3) A step of providing an immobilized primer, carrying out a primer extension reaction using a carrier sequence as a template to produce an extension product and obtain an oligonucleotide probe, wherein the immobilized primer contains the sequence of consensus sequence X1 and is capable of annealing with the immobilized sequence of the carrier sequence to initiate the extension reaction, and in some embodiments, the extension product in the 5' to 3' direction contains or consists of consensus sequence X1, tag sequence Y and consensus sequence X2, (4) A step of connecting the immobilized primer to the surface of the solid support, wherein steps (3) and (4) are carried out in any order, (5) Optionally, the immobilized carrier sequence further includes cleavage sites, and the cleavage may be selected from Nicking enzyme digestion, USER enzyme digestion, photoresponsive excision, chemical excision or CRISPR-mediated excision, and the cleavage is carried out at the cleavage sites included in the immobilized carrier sequence to digest the carrier sequence and separate the extension product in step (3) from the template (i.e., the carrier sequence) from which the extension product is generated, thereby ligating the oligonucleotide probe to the surface of a solid support (e.g., a tip).

[0113] In a particular embodiment, the various carrier sequences are DNBs formed from concatemers of multiple copies of the carrier sequence.

[0114] In a particular embodiment, multiple types of carrier arrays are used in step (1) as follows: (i) A step of providing a plurality of types of carrier-template sequences, wherein the carrier-template sequence includes a complementary sequence of the carrier sequence, (ii) A step of carrying out a nucleic acid amplification reaction using various carrier-template sequences as templates to obtain amplification products of various carrier-template sequences, wherein the amplification products include at least one copy (e.g., multiple copies) of the carrier sequence, and in a particular embodiment, rolling circle replication is performed to obtain a DNB formed from concatemers of the carrier sequence. Provided by [company name].

[0115] In some embodiments, the solid support in step (1) has one or more features selected from the following: (1) The solid phase support is selected from the group consisting of latex beads, dextran beads, polystyrene surfaces, polypropylene surfaces, polyacrylamide gel, gold surfaces, glass surfaces, chips, sensors, electrodes, and silicon wafers, and in some embodiments the solid phase support is a chip. (2) The solid phase support is planar, spherical, or porous. (3) The solid-phase support can be used as a sequencing platform, for example, as a sequencing chip. In some embodiments, the solid-phase support is a sequencing chip for Illumina, MGI or Thermo Fisher sequencing platforms, and (4) The solid support is capable of spontaneously releasing oligonucleotide probes or upon exposure to one or more stimuli (e.g., temperature changes, pH changes, exposure to specific chemicals or phases, exposure to light, exposure to reducing agents, etc.).

[0116] Method for constructing a nucleic acid molecule library

[0117] In another aspect, the present application also relates to a method for constructing a nucleic acid molecule library, (a) A step of generating a group of labeled nucleic acid molecules according to the method described above, and (b) The steps of randomly fragmenting nucleic acid molecules in a population of labeled nucleic acid molecules and attaching adapters, and (c) Optionally, a step of amplifying and / or concentrating the product of step (b). This invention provides a method for obtaining a nucleic acid molecule library, which includes [a specific component / method].

[0118] In certain embodiments, a nucleic acid molecule library is used for sequencing, for example, transcriptome sequencing, for example, single-cell transcriptome sequencing.

[0119] In a particular embodiment, before performing step (b), the method further includes the step (pre-b): amplifying and / or enriching a population of labeled nucleic acid molecules.

[0120] In a particular embodiment, in step (pre-b), a group of labeled nucleic acid molecules is subjected to a nucleic acid amplification reaction to generate an amplified product.

[0121] In a particular embodiment, the nucleic acid amplification reaction is carried out using at least primer C and / or primer D, wherein primer C can hybridize or anneal to a complementary sequence of consensus sequence X1 or a complementary sequence of the 3'-terminal sequence of consensus sequence X1 to initiate the extension reaction, and primer D can hybridize or anneal to nucleic acid molecules of a population of labeled nucleic acid molecules to initiate the extension reaction.

[0122] In a particular embodiment, the nucleic acid amplification reaction in step (preb) is carried out using a nucleic acid polymerase (e.g., DNA polymerase, e.g., a DNA polymerase having strand displacement activity and / or high fidelity).

[0123] In certain embodiments, in step (b), the nucleic acid molecule is randomly fragmented into fragments, and adapters are attached to the fragments using a transposase. In some embodiments, in step (b) of the method, the nucleic acid molecule obtained in the previous step is randomly fragmented into fragments, and a first adapter and a second adapter are attached to both ends of the fragments, respectively, using a transposase.

[0124] In certain embodiments, the transposase is selected from the group consisting of Tn5 transposase, MuA transposase, Sleeping Beauty transposase, Mariner transposase, Tn7 transposase, Tn10 transposase, Ty1 transposase, Tn552 transposase, and their variants, modified products, and derivatives having transposase activity.

[0125] In a particular embodiment, the transposase is Tn5 transposase.

[0126] In some embodiments, in step (c), the product of step (b) is amplified using at least primer C' and / or primer D', wherein primer C' can hybridize or anneal with the first adapter to initiate the extension reaction, and primer D' can hybridize or anneal with the second adapter to initiate the extension reaction.

[0127] In some embodiments, in step (c), the product of step (b) can be amplified using at least primer C and / or primer D', and primer D' can hybridize or anneal with the first or second adapter to initiate the extension reaction.

[0128] Sequencing method

[0129] In another aspect, the present application also provides a method for sequencing nucleic acid samples, the method being: (1) The step of constructing a nucleic acid molecule library according to the method described above, and (2) Step of sequencing the nucleic acid molecule library Includes.

[0130] kit

[0131] In another embodiment, this application also provides a kit, the kit is, (i) A nucleic acid array for labeling nucleic acids, comprising a solid support, the solid support being bound to a plurality of types of oligonucleotide probes, each type of oligonucleotide probe comprising at least one copy, and each oligonucleotide probe comprising or comprising a consensus sequence X1, a tag sequence Y, and a consensus sequence X2 in the 5' to 3' direction, Various oligonucleotide probes have different tag sequences Y, and tag sequence Y has a nucleotide sequence specific to the position on the solid support of that type of oligonucleotide probe, for use in nucleic acid arrays for labeling nucleic acids. (ii) A primer set comprising primer A, or primer A' and primer B, or a primer set comprising primer A and primer B, Primer A comprises a consensus sequence A and a capture sequence A, the capture sequence A being capable of annealing with the captured RNA (e.g., mRNA) to initiate an elongation reaction, and in some embodiments, the consensus sequence A is located upstream of the capture sequence A (e.g., located at the 5' end of primer A), Primer A' contains capture sequence A, which can anneal to the captured RNA (e.g., mRNA) to initiate the elongation reaction. Primer B comprises a consensus sequence B, a complementary sequence for the 3'-terminal overhang, and optionally, a tag sequence B, wherein in certain embodiments, the complementary sequence for the 3'-terminal overhang is located at the 3'-end of primer B, and in certain embodiments, the consensus sequence B is located upstream of the complementary sequence for the 3'-terminal overhang (e.g., at the 5'-end of primer B), where the 3'-terminal overhang refers to one or more non-template nucleotides contained in the 3'-end of a cDNA strand produced by reverse transcription using RNA captured by the capture sequence A of primer A' as a template, or a primer set comprising primer A' and primer B, or a primer set comprising primer A and primer B and (iii) A crosslinked oligonucleotide comprising a first region and a second region, and optionally a third region located between the first region and the second region, wherein the first region is located upstream of the second region (for example, located at 5' of the second region), The first region is capable of (a) annealing with all or part of the consensus sequence A of primer A, or (b) annealing with all or part of the consensus sequence B of primer B. The second region is a cross-linked oligonucleotide that can anneal to all or part of the consensus sequence X2. Includes.

[0132] In a particular embodiment, each oligonucleotide probe contains one copy.

[0133] In a particular embodiment, various oligonucleotide probes include multiple copies.

[0134] In a particular embodiment, the region to which various oligonucleotide probes are bound to a solid support is called a microdot. If the various oligonucleotide probes contain one copy, each microdot is bound to one oligonucleotide probe, and oligonucleotide probes in different microdots have different tag sequences Y. If the various oligonucleotide probes contain multiple copies, each microdot is bound to multiple oligonucleotide probes, and oligonucleotide probes in the same microdot have the same tag sequence Y, while oligonucleotide probes in different microdots have different tag sequences Y.

[0135] In a particular embodiment, the solid-phase support comprises a plurality of microdots, each microdot bound to one type of oligonucleotide probe, and each type of oligonucleotide probe may comprise one or more copies.

[0136] In a particular embodiment, the solid phase support is a plurality of (e.g., at least 10, at least 10) 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 The solid-phase support comprises (or more than) microdots, and in certain embodiments, the solid-phase support comprises at least 10 4 (For example, at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 or at least 10 12 ) microdots / mm 2 Includes.

[0137] In some embodiments, the spacing between adjacent microdots is less than 100 μm, less than 50 μm, less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, or less than 0.01 μm.

[0138] In certain embodiments, the microdots have a size (e.g., equivalent diameter) of less than 100 μm, less than 50 μm, less than 10 μm, less than 5 μm, less than 1 μm, less than 0.5 μm, less than 0.1 μm, less than 0.05 μm, or less than 0.01 μm.

[0139] In a particular embodiment, the kit comprises (i) a nucleic acid array for labeling the nucleic acid described above, (ii) a primer A, and (iii) a crosslinked oligonucleotide, wherein a first region of the crosslinked oligonucleotide is annealable with all or part of consensus sequence A of primer A, and a second region of the crosslinked oligonucleotide is annealable with all or part of consensus sequence X2.

[0140] In a particular embodiment, the capture sequence A of primer A is a random oligonucleotide sequence.

[0141] In certain embodiments, the capture sequence A of primer A is a poly(T) sequence or a specific sequence that targets a target nucleic acid; in certain embodiments, primer A further comprises a tag sequence A, for example, a random oligonucleotide sequence; and in certain embodiments, the capture sequence A is located at the 3' end of primer A, and the consensus sequence A is located upstream of the tag sequence A (for example, at the 5' end of primer A).

[0142] In a particular embodiment, primer A contains a 5'-phosphate at its 5'-terminus.

[0143] In a particular embodiment, the kit comprises (i) a nucleic acid array for labeling the nucleic acid described above, a primer set comprising primer A' and primer B described above, and (iii) a crosslinked oligonucleotide described above, wherein a first region of the crosslinked oligonucleotide is annealable with all or part of consensus sequence B of primer B, and a second region of the crosslinked oligonucleotide is annealable with all or part of consensus sequence X2.

[0144] In a particular embodiment, the capture sequence of primer A' is a random oligonucleotide sequence.

[0145] In certain embodiments, the capture sequence A of primer A' is a poly(T) sequence or a sequence that targets a target nucleic acid; in certain embodiments, primer A' further comprises a tag sequence A and a consensus sequence A; in certain embodiments, the capture sequence A is located at the 3' end of primer A'; and in certain embodiments, the consensus sequence A is located upstream of the capture sequence A (for example, at the 5' end of primer A').

[0146] In a particular embodiment, primer B includes a consensus sequence B, a complementary sequence for the 3'-terminal overhang, and a tag sequence B.

[0147] In certain embodiments, the kit further comprises primer B', which can anneal to all or part of the complementary sequence of consensus sequence B to initiate an extension reaction.

[0148] In certain embodiments, primer B or primer B' contains 5'-phosphate at its 5'-terminus.

[0149] In certain embodiments, primer B includes a modified nucleotide (e.g., locked nucleic acid), and in certain embodiments, primer B includes one or more modified nucleotides (e.g., one or more locked nucleic acids) at its 3' end.

[0150] In a particular embodiment, the kit comprises (i) a nucleic acid array for labeling the nucleic acid described above, a primer set comprising primer A and primer B described above, and (iii) a crosslinked oligonucleotide described above, wherein a first region of the crosslinked oligonucleotide is annealable with all or part of consensus sequence A of primer A, and a second region of the crosslinked oligonucleotide is annealable with all or part of consensus sequence X2.

[0151] In a particular embodiment, the capture sequence A of primer A is a random oligonucleotide sequence.

[0152] In certain embodiments, the capture sequence A of primer A is a poly(T) sequence or a sequence that targets a target nucleic acid; in certain embodiments, primer A further comprises a tag sequence A, for example, a random oligonucleotide sequence; and in certain embodiments, the capture sequence A is located at the 3' end of primer A, and the consensus sequence A is located upstream of the tag sequence A (for example, at the 5' end of primer A).

[0153] In a particular embodiment, primer A contains a 5'-phosphate at its 5'-terminus.

[0154] In certain embodiments, primer B includes a modified nucleotide (e.g., locked nucleic acid). In certain embodiments, primer B includes one or more modified nucleotides (e.g., locked nucleic acid) at its 3' end.

[0155] In a particular embodiment, the kit has one or more features selected from the following: (1) The oligonucleotide probe, primer A, primer A', primer B, primer B', and crosslinked oligonucleotide each independently contain or consist of a native nucleotide (e.g., deoxyribonucleotide or ribonucleotide), a modified nucleotide, a non-native nucleotide, or any combination thereof. In certain embodiments, primer A, primer A', and primer B' are capable of initiating an extension reaction. In certain embodiments, the consensus sequence X2 has a free hydroxyl group (-OH) at its 3'-terminus. (2) Each oligonucleotide probe independently has a length of 15-300 nt (e.g., 15-200 nt, 15-20 nt, 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-150 nt, 150-200 nt), (3) Each primer A, primer A', primer B, and primer B' independently has a length of 10 to 200 nt (for example, 10 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt). (4) The cross-linked oligonucleotides have a length of 6 to 200 nt (for example, 20 to 70 nt, 6 to 15 nt, 15 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 100 nt, 100 to 150 nt, 150 to 200 nt). (5) The oligonucleotide probes bound to the same solid support have the same consensus sequence X1 and / or the same consensus sequence X2. (6) The consensus sequence X1 of the oligonucleotide probe includes a cleavage site, which in some embodiments may be cleaved or destroyed by a method selected from Nicking enzyme digestion, USER enzyme digestion, photoresponsive excision, chemical excision, or CRISPR-mediated excision.

[0156] In certain embodiments, the kit further comprises a reverse transcriptase, a nucleic acid ligase, a nucleic acid polymerase, and / or a transposase.

[0157] In certain embodiments, the reverse transcriptase has terminal deoxynucleotidyltransferase activity; in certain embodiments, the reverse transcriptase is capable of synthesizing a cDNA strand using RNA (e.g., mRNA) as a template and adding a 3'-terminal overhang to the 3'-end of the cDNA strand; in certain embodiments, the reverse transcriptase is capable of adding an overhang to the 3'-end of the cDNA strand having a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides; in certain embodiments, the reverse transcriptase is capable of adding an overhang of 2-5 cytosine nucleotides (e.g., a CCC overhang) to the 3'-end of the cDNA strand; and in certain embodiments, the reverse transcriptase is selected from the group consisting of M-MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and their variants, modified products and derivatives having the reverse transcription activity of the above reverse transcriptases.

[0158] In certain embodiments, the nucleic acid polymerase lacks 5'-to-3' exonuclease activity or strand displacement activity.

[0159] In certain embodiments, the transposase is selected from the group consisting of Tn5 transposase, MuA transposase, Sleeping Beauty transposase, Marina transposase, Tn7 transposase, Tn10 transposase, Ty1 transposase, Tn552 transposase, and their variants, modified products, and derivatives having transposase activity.

[0160] In a particular embodiment, the kit further comprises primer C, primer D, primer C' and / or primer D'. For example, the kit further comprises primer C, primer D and primer D'. For example, the kit further comprises primer C, primer D, primer C' and primer D'.

[0161] In a particular embodiment, the kit further includes reagents for nucleic acid hybridization, reagents for nucleic acid extension, reagents for nucleic acid amplification, reagents for recovering or purifying nucleic acids, reagents for constructing transcriptome sequencing libraries, reagents for sequencing (e.g., second or third-generation sequencing), or any combination thereof.

[0162] use

[0163] In another embodiment, the application also provides a method for generating such labeled nucleic acid molecule populations, or the use of such kits for constructing nucleic acid molecule libraries or for performing transcriptome sequencing.

[0164] Definition of Terms

[0165] In this application, unless otherwise specified, scientific and technical terms used herein have meanings that are generally understood by those skilled in the art. Furthermore, the operational steps, such as those in molecular biology, biochemistry, nucleic acid chemistry, and cell culture, as used herein, are all commonplace steps widely used in the corresponding art. On the other hand, for the purpose of further understanding this application, definitions and explanations of relevant terms are provided below.

[0166] Where the terms “e.g.”, “for example,” “such as,” “comprise,” “include,” or variations thereof are used herein, these terms are not considered restrictive terms, but rather are construed to mean “not limited to” or “not restrictive.”

[0167] Unless otherwise indicated herein or unless explicitly stated otherwise by context, the terms “a” and “an” and “the” and similar referents in the context describing this application (in particular in the context of the following claims) should be interpreted as referring to singular and plural forms.

[0168] As used herein, "DNB" (DNA nanoball) is a typical RCA (rolling circle amplification) product and possesses the characteristics of an RCA product. Here, an RCA product is a single-stranded DNA sequence having multiple copies, which can form a "globular" structure through interactions between bases contained in the DNA. Typically, library molecules are circularized to form a single-stranded circular DNA, which can then be amplified by several orders of magnitude using rolling circle amplification techniques, thereby producing an amplified product called a DNB.

[0169] As used herein, “nucleic acid molecule population” refers to a group or collection of nucleic acid molecules that are directly or indirectly derived from a target nucleic acid molecule, such as double-stranded DNA, RNA / cDNA hybrid, single-stranded DNA, or single-stranded RNA. In some embodiments, the nucleic acid molecule population includes a nucleic acid molecule library, which includes sequences that qualitatively and / or quantitatively represent the target nucleic acid molecule sequence. In other embodiments, the nucleic acid molecule population includes a subset of the nucleic acid molecule library.

[0170] As used herein, “nucleic acid molecule library” means a collection or group of labeled nucleic acid molecules (e.g., labeled double-stranded DNA, labeled RNA / cDNA hybrid, labeled single-stranded DNA, or labeled single-stranded RNA) or fragments thereof generated directly or indirectly from a target nucleic acid molecule, where the collection or group of labeled nucleic acid molecules or combinations thereof are indicated to qualitatively and / or quantitatively represent the sequence of the target nucleic acid molecule sequence from which the labeled nucleic acid molecule was generated. In certain embodiments, the nucleic acid molecule library is a sequencing library. In certain embodiments, the nucleic acid molecule library may be used to construct a sequencing library.

[0171] As used herein, “cDNA” or “cDNA strand” means “complementary DNA” synthesized by using at least a portion of the RNA molecule of interest as a template and extending it via a primer that anneals with the RNA molecule of interest under the catalysis of RNA-dependent DNA polymerase or reverse transcriptase (this process is also called “reverse transcription”). The synthesized cDNA molecule is “homologous” or “complementary” to at least a portion of the template, or “base-pairs” or “complexes” with at least a portion of the template.

[0172] As used herein, the term “upstream” is used to describe the relative positional relationship between two nucleic acid sequences (or two nucleic acid molecules) and has a meaning that is generally understood by those skilled in the art. For example, the expression “one nucleic acid sequence is located upstream of another nucleic acid sequence” means that, when aligned in the 5'-to-3' direction, the former is located further forward (i.e., closer to the 5' end) than the latter. As used herein, the term “downstream” has the opposite meaning of “upstream.”

[0173] As used herein, "Tag Sequence Y," "Tag Sequence A," "Tag Sequence B," "Consensus Sequence X1," "Consensus Sequence X2," "Consensus Sequence A," "Consensus Sequence B," etc., refer to oligonucleotides having non-target nucleic acid components that provide means for identification, recognition, and / or molecular or biochemical manipulation of nucleic acid molecules ligated thereto or derivative products of nucleic acid molecules ligated thereto (e.g., complementary fragments of nucleic acid molecules, short fragments of nucleic acid molecules, etc.) (e.g., by providing a site for annealing with an oligonucleotide, where the oligonucleotide is, for example, a primer for DNA polymerase elongation or an oligonucleotide for a capture or ligation reaction). An oligonucleotide may consist of at least two nucleotides (preferably about 6 to 100, but there is no clear limit to the length of the oligonucleotide, and the exact size depends on a number of factors, while these factors depend on the final function or use of the oligonucleotide), and may consist of a plurality of oligonucleotide segments arranged contiguously or discontinuously. An oligonucleotide sequence may be specific to each nucleic acid molecule it ligates, or it may be specific to a particular type of nucleic acid molecule it ligates. Oligonucleotide sequences can be reversibly or irreversibly ligated into polynucleotide sequences that are “labeled” by any method, including ligation, hybridization, or other methods. The process of ligating an oligonucleotide sequence with a nucleic acid molecule is sometimes called “labeling,” and the nucleic acid molecule that undergoes labeling, or that contains a labeled sequence, is called a “labeled nucleic acid molecule” or “tagged nucleic acid molecule.”

[0174] For various reasons, the nucleic acids or polynucleotides of this application (e.g., "Tag Sequence Y", "Tag Sequence A", "Tag Sequence B", "Consensus Sequence X1", "Consensus Sequence X2", "Consensus Sequence A", "Consensus Sequence B", "Primer A", "Primer A'", "Primer B", "Primer C", "Primer D", "Primer D'", "Cross-linked oligonucleotide", etc.) may contain one or more modified nucleic acid bases, sugar moieties, or internucleoside linkages. Some, but not limited to, reasons for using nucleic acids or polynucleotides containing modified nucleic acid bases, sugar moieties, or internucleoside linkages include, (1) changes in Tm, (2) changes in the sensitivity of the polynucleotide to one or more nucleases, (3) providing a portion for linking labels, (4) providing a label or label quencher, or (5) providing a portion such as biotin for attachment to another molecule in solution or bound to the surface. For example, in some embodiments, oligonucleotides, e.g., primers, can be synthesized to contain, in a random portion, one or more nucleic acid analogs having a constrained conformation, which may include, but are not limited to, one or more ribonucleic acid analogs in which the ribose ring is "locked" by a methylene bridge linking the 2'-O atom to the 4'-C atom. These modified nucleotides result in an increase of about 2 to about 8 degrees Celsius in the Tm or melting temperature of each molecule. For example, in some embodiments in which oligonucleotide primers containing ribonucleotides are used, one indicator of using modified nucleotides in a method may be that the oligonucleotides containing the modified nucleotides can be digested by single-strand specific RNases.

[0175] In the methods of this application, for example, the nucleic acid bases of a single nucleotide at one or more positions of a polynucleotide or oligonucleotide may include guanine, adenine, uracil, thymine, or cytosine, or optionally, one or more of the nucleic acid bases may include modified bases, for example, but not limited to, xanthine, allylaminouracil, allylaminothymine nucleoside, hypoxanthine, 2-aminoadenine, 5-propynyluracil, 5-propynylcytosine, 4-thiouracil, 6-thioguanine, azauracil, deazauracil, thymine nucleoside, cytosine, adenine, or guanine. Furthermore, they may include nucleic acid bases derivatized with the following parts: biotin moiety, digoxigenin moiety, fluorescent or chemiluminescent moiety, quenching moiety, or several other parts. This application is not limited to the enumerated nucleic acid bases, and the given list exemplifies a broad range of bases that may be used in the methods of this application.

[0176] With respect to nucleic acids or polynucleotides of this application, one or more of the sugar moieties may comprise 2'-deoxyribose, or optionally, one or more of the sugar moieties may comprise several other sugar moieties, for example, ribose, 2'-fluoro-2'-deoxyribose, or 2'-O-methyl-ribose, which are resistant to several nucleases, or 2'-amino-2'-deoxyribose or 2'-azido-2'-deoxyribose, which are labeled by a visible, fluorescent, infrared fluorescent, or other detectable dye or by a chemical having an electrophilic, photoreactive, alkynyl, or other reactive chemical moiety.

[0177] The nucleoside linkages of the nucleic acids or polynucleotides of this application may be phosphodiester links, or optionally, one or more of the nucleoside linkages may include modified linkages, such as phosphorothioates, phosphorodithioates, phosphoroselenates, or phosphorodiserenates, which are resistant to several nucleases, but are not limited to these.

[0178] As used herein, the term “terminal deoxynucleotidyltransferase activity” refers to the ability to catalyze the template-independent addition (or “tailing”) of one or more deoxyribonucleoside triphosphates (dNTPs) or a single dideoxyribonucleoside triphosphate to the 3'-end of cDNA. Examples of reverse transcriptases having terminal deoxynucleotidyltransferase activity include, but are not limited to, MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and their variants, modified products, and derivatives having reverse transcriptase activity and terminal deoxynucleotidyltransferase activity. Reverse transcriptases may or may not have RNase activity (in particular, RNase H activity). In preferred embodiments, the reverse transcriptase used for reverse transcription of RNA to produce cDNA does not have RNase activity (in particular, RNase H activity). Therefore, in preferred embodiments, the reverse transcriptase used for reverse transcription of RNA to generate cDNA has terminal deoxynucleotidyltransferase activity and does not have RNase activity (in particular, RNase H activity).

[0179] As used herein, a nucleic acid polymerase having "chain displacement activity" refers to a nucleic acid polymerase that, when encountering a downstream nucleic acid chain complementary to the template chain during the process of extending a new nucleic acid chain, can continue the extension reaction and replace the complementary nucleic acid chain (rather than decompose it).

[0180] As used herein, a nucleic acid polymerase having "5'-to-3' exonuclease activity" refers to a nucleic acid polymerase that can catalyze the hydrolysis of the 3,5-phosphate diester bond of a polynucleotide from the 5'-terminus to the 3'-terminus, thereby degrading the nucleotide.

[0181] As used herein, a nucleic acid polymerase (or DNA polymerase) having "high fidelity" refers to a nucleic acid polymerase (or DNA polymerase) that, during nucleic acid amplification, is less likely to introduce incorrect nucleotides (i.e., has a lower error rate) than a wild-type Taq enzyme (e.g., a Taq enzyme whose sequence is shown in UniProt Commission: P19821.1).

[0182] As used herein, the terms “annealed,” “annealing,” “to anneal,” “hybridized,” or “to hybridize” refer to the formation of a complex between nucleotide sequences that are sufficiently complementary to form a complex via Watson-Crick base pairing. For the purposes of this application, nucleic acid sequences that are “complementary,” “hybridize,” or “anneal” to each other must be able to form a sufficiently stable “hybrid” or “complex” for the intended purposes. It is not necessary for both nucleic acid molecules or the corresponding sequences presented therein to be “complementary,” “annealing,” or “hybridize” to each other by the fact that all nucleic acid bases in a sequence presented by one nucleic acid molecule can base pair, pair, or complex with all nucleic acid bases in a sequence presented by another nucleic acid molecule. As used herein, the terms “complementary” or “complementarity” are used when referring to sequences of nucleotides associated by the rules of base pairing. For example, the sequence 5'-AGT-3' is complementary to the sequence 3'-TCA-5'. Complementarity can be "partial," in which case only some of the nucleic acid bases match according to the rules of base pairing. Optionally, there may be "complete" or "total" complementarity between nucleic acids. The degree of complementarity between nucleic acid strands has a significant impact on the efficiency and strength of hybridization between nucleic acid strands. The degree of complementarity is particularly important in amplification and detection methods that rely on nucleic acid hybridization. The term "homology" refers to the degree of complementarity of one nucleic acid sequence to another. There can be partial homology (i.e., complementarity) or complete homology (i.e., complementarity). A partially complementary sequence is one that at least partially inhibits the hybridization of a fully complementary sequence with its target nucleic acid, and is referred to using the functional term "substantially homologous." Inhibition of hybridization of a completely complementary sequence with a target sequence can be tested under low stringency conditions using hybridization assays (e.g., Southern blotting or Northern blotting, solution hybridization, etc.).Substantially homologous sequences or probes will compete for or inhibit binding (i.e., hybridization) with fully homologous target sequences under low stringency conditions. This does not mean that low stringency conditions allow for nonspecific binding; rather, low stringency conditions require the two sequences to bind to each other by specific (i.e., selective) interactions. The absence of nonspecific binding can be tested by using a second target that lacks complementarity or has only low complementarity (e.g., less than approximately 30%). If specific binding is low or absent, the probe will not hybridize with the nucleic acid target. The term “substantially homologous” means, when used with respect to double-stranded nucleic acid sequences, e.g., cDNA or genomic clones, any oligonucleotide or probe that can hybridize with one or both strands of a double-stranded nucleic acid sequence under the low stringency conditions described herein. As used herein, the terms “annealing” or “hybridization” are used to refer to the pairing of complementary nucleic acid chains. Hybridization and hybridization force (i.e., the association force between nucleic acid chains) are influenced by a number of factors known in the art, including the degree of complementarity between nucleic acids, which includes stringency, a condition influenced by factors such as salt concentration, Tm (melting temperature) for hybridization, presence of other components (e.g., polyethylene glycol or betaine), molar concentration of the hybridized chain, and GC content of the nucleic acid chains.

[0183] As described herein, a solid support can spontaneously release oligonucleotide probes or upon exposure to one or more stimuli (e.g., temperature changes, pH changes, exposure to specific chemicals or phases, exposure to light, reducing agents, etc.). The oligonucleotide probes may be released by cleavage of the bond between the oligonucleotide probe and the solid support, or by decomposition of the solid support itself, or both, and it is understood that the oligonucleotide probes may or may be made accessible by other reagents.

[0184] By adding multiple types of unstable bonds to a solid support, the ability of the solid support to respond to different stimuli becomes possible. Each type of unstable bond may be sensitive to the relevant stimuli (e.g., chemical stimuli, light, temperature, etc.), and as a result, the release of substances attached to the solid support via each unstable bond can be controlled by applying the appropriate stimulus. In addition to thermally cleavable bonds, disulfide bonds, and UV-sensitive bonds, other, not limited to, examples of unstable bonds that can be bonded to a solid support include ester bonds (e.g., ester bonds that can be cleaved with acids, bases, or hydroxylamines), orthodiol bonds (e.g., orthodiol bonds that can be cleaved with sodium periodate), Diels-Alder bonds (e.g., Diels-Alder bonds that can be cleaved thermally), sulfone bonds (e.g., sulfone bonds that can be cleaved with alkalis), silicyl ether bonds (e.g., silicyl ether bonds that can be cleaved with acids), glycosidic bonds (e.g., glycosidic bonds that can be cleaved with amylases), peptide bonds (e.g., peptide bonds that can be cleaved with proteases), or phosphate diester bonds (e.g., phosphate diester bonds that can be cleaved with nucleases (e.g., DNA enzymes)).

[0185] In addition to the cleavable bond between the solid support and the oligonucleotide described above, or as an alternative thereto, the solid support may be degradable, destructible, or soluble spontaneously or in response to exposure to one or more stimuli (e.g., temperature changes, pH changes, exposure to specific chemicals or phases, exposure to light, exposure to reducing agents, etc.). In some cases, the solid support may be soluble, and as a result, the material components of the solid support may dissolve upon exposure to specific chemicals or environmental changes (e.g., temperature changes or pH changes). In some cases, the solid support may decompose or dissolve under high temperature and / or alkaline conditions. In some cases, the solid support may be thermally decomposable, and as a result, the solid support may decompose when exposed to appropriate temperature changes (e.g., heating). Decomposition or dissolution of a solid support bound to a substance (e.g., an oligonucleotide probe) may result in the release of the substance from the solid support.

[0186] As used herein, the terms “transposase,” “reverse transcriptase,” and “nucleic acid polymerase” refer to protein molecules or aggregates of protein molecules involved in catalyzing specific chemical and biological reactions. In general, the methods, compositions, or kits of this application are not limited to the use of specific transposases, reverse transcriptases, or nucleic acid polymerases from specific sources. Rather, the methods, compositions, or kits of this application may include any transposase, reverse transcriptase, or nucleic acid polymerase from any source having enzymatic activity equivalent to that of the specific enzyme in the specific method, composition, or kit disclosed herein. Furthermore, the methods of this application also include the following embodiments: any one specific enzyme provided and used in a step of the method is replaced by a combination of two or more enzymes, and when two or more enzymes are used in combination, whether separately and stepwise or simultaneously together, the reaction mixture yields the same results as those that would be obtained using that specific enzyme. The methods, buffers, and reaction conditions provided herein, including those in the examples, are currently preferred for embodiments of the methods, compositions, and kits of this application. However, other enzyme preservation buffers, reaction buffers, and reaction conditions may be used for some of the enzymes of this application, and may be known in the art and suitable for use in this application, and are included herein.

[0187] [Beneficial effects of this application] This application provides a novel method for generating a population of labeled nucleic acid molecules, and a method for constructing a nucleic acid molecule library based on this method and performing high-rate sequencing, thereby achieving highly accurate intracellular-level spatial positioning of a sample. The method of this application has one or more beneficial technical effects selected from the following:

[0188] (1) Probes in traditional nucleic acid arrays (e.g., chips) used for spatial transcriptome sequencing include fixed capture sequences. Typically, a specific capture sequence can capture only the corresponding specific target nucleic acid molecule. For example, if the capture sequence is poly(T), it will correspondingly capture target nucleic acid molecules containing poly(A). If the target nucleic acid molecule changes, the probe sequence containing the capture sequence must be modified to match it, i.e., the entire nucleic acid array (e.g., chip) must be modified, which is costly and inefficient in practical applications. The nucleic acid array (e.g., chip) of this application does not include a capture sequence, and the capture sequence resides in a reverse transcription primer independent of the nucleic acid array (i.e., the capture sequence and probe are independent of each other). After the capture sequence captures the target nucleic acid molecule, it ligates to the probe via a cross-linked oligonucleotide.

[0189] Therefore, this application enables the design of corresponding capture sequences for different target nucleic acid molecules without changing the probe sequence (i.e., without changing the nucleic acid array (e.g., the chip)), and achieves capture of different target nucleic acid molecules by changing the capture sequence and crosslinking oligonucleotides.

[0190] (2) All traditional spatial transcriptome methods use poly(T) as a capture sequence and cannot capture RNA that does not have a poly(A) tail. However, this application can capture target nucleic acid molecules that do not have a poly(A) tail by replacing the poly(T) of the capture sequence with a sequence of random sequences (e.g., random primer sequences, e.g., N6, N8, etc.), and the sequence of random sequences can also simultaneously function as a unique molecular identifier (UMI) sequence.

[0191] (3) Traditional nucleic acid arrays (e.g., chips) used for spatial transcriptome sequencing have immobilized capture probes. Generally, tissue permeabilization is performed first to release intracellular RNA. If permeabilization is excessive, the RNA spreads to adjacent cells, even to the periphery of the tissue sample, and is captured by the probe, making it impossible to achieve insight capture of mRNA. If permeabilization is incomplete, the mRNA capture efficiency is affected. With the method of this application, the nucleic acid array (e.g., chip) does not contain capture sequences (the nucleic acid array contains spatial information and does not contain capture sequences), and the purpose of tissue permeabilization is to allow reverse transcription primers to enter cells and hybridize with mRNA and insights without requiring strong permeabilization reagent treatment, thereby reducing the spread of the sample.

[0192] Preferred embodiments of this application are described below in detail with reference to the accompanying drawings and examples, but those skilled in the art will understand that the following drawings and examples are provided merely to illustrate this application and do not limit its scope. Various purposes and advantageous aspects of this application will become apparent to those skilled in the art from the accompanying drawings and the detailed description of the following preferred embodiments. [Examples]

[0193] This application will be described here with reference to the following examples, which are intended to be illustrative (but not limiting) to this application. Unless otherwise noted, the experiments and methods described in the examples were carried out in essentially accordance with the prior art methods described in various references. Furthermore, where specific conditions are not specified in the examples, the prior art conditions or conditions recommended by the manufacturer should be followed. Where the manufacturer of the reagents or equipment used is not specified, they were all conventional products that can be purchased commercially. Those skilled in the art will understand that the examples are illustrative and not intended to limit the scope of the application to be protected. All publications and other references mentioned herein are incorporated by reference in their entirety.

[0194] Sequence information

[0195] Information regarding some of the sequences included in this application is provided in Table 1 below. [Table 1]

[0196] Example 1: Preparation of the capture tip 1. The sequence of the DNA library molecule containing the chip's positional information was designed, consisting of a consensus sequence X1 (X1), a tag sequence (Y), and a consensus sequence X2 (X2) from 5' to 3'. The normal nucleotide sequence of the DNA library molecule is shown in Sequence ID No. 1. The DNA library molecule was commissioned to Beijing Liuhe BGI Co., Ltd. for synthesis.

[0197] 2. Amplification and loading of library molecules (1) DNA nanoballs (DNBs) were prepared using the DNBSEQ sequencing kit (purchased from MGI, catalog number 1000019840). A specific embodiment is briefly described below.

[0198] Briefly, a 40 μL reaction system as shown in Table 2 was prepared. The reaction system was placed in a PCR instrument and the reaction was carried out according to the following conditions: 3 minutes at 95°C, then 3 minutes at 40°C. After the reaction was complete, the reaction product was placed on ice and 40 μL of mixed enzyme I, 2 μL of mixed enzyme II (from the DNBSEQ sequencing kit), 1 μL of ATP (100 mM storage solution, obtained from Thermo Fisher), and 0.1 μL of T4 ligase (obtained from NEB, catalog number: M0202S) were added. After mixing the wells, the above reaction system was placed in a PCR instrument and reacted at 30°C for 20 minutes to produce DNB.

[0199] [Table 2]

[0200] (2) Subsequently, the DNB was loaded onto the SEQ500 sequencing chip (purchased from MGI) using the SEQ500 SE50 kit set (purchased from MGI, 1000012551).

[0201] In the sequencing chips, the MDA reagent from the PE50 sequencing kit (purchased from MGI, 1000012554) was added, incubated at 37°C for 30 minutes, and then the chips were washed with 5X SSC.

[0202] (3) The surface of the chip was modified with N3-PEG3500-NHS (modification reagent purchased from Sigma, catalog number: JKA5086). After incubation for 30 minutes, DBCO modification primers for chip sequence synthesis (sequence shown in SEQ ID NO: 3) were injected into it, and it was incubated overnight at room temperature.

[0203] 3. Sequencing and Decoding of Positional Sequence Information: Sequencing was performed according to the instructions for the BGISEQ500 SE50 sequencing kit, with the SE read length set to 25 bp. The fq file generated by sequencing was saved for later use.

[0204] 4. Extension of the consensus sequence X2 (SEQ ID NO: 5) of the chip sequence: Based on step 3 above, the 13-base cPAS reaction was continued to obtain the chip sequence (SEQ ID NO: 8, consisting of consensus sequence X1 (SEQ ID NO: 4), tag sequence Y, and consensus sequence X2 (SEQ ID NO: 5)).

[0205] 5. Tip cutting: The prepared tips were cut into several small pieces. The size of the pieces was adjusted according to experimental requirements. The tips were immersed in 50 mM Tris buffer at pH 8.0 at 4°C for later use.

[0206] Example 2: Synthesis of cDNA insights 1. cDNA synthesis Mouse tissue sections were prepared according to the standard method for frozen sections, and the frozen sections were mounted on the chip prepared in Example 1. After fixing with frozen methanol for 30 minutes, the tissue was permeabilized using 0.5% triton x-100. The chip was washed twice at room temperature using 5X SSC, and 200 μL of the reverse transcriptase reaction system as shown in Table 3 was prepared. The reaction solution was added to the chip so as to completely cover it, and the reaction was carried out at 42°C for 90–180 minutes. cDNA synthesis was carried out by reverse transcriptase using mRNA and poly-T containing primers (sequence shown in SEQ ID NO: 6, consisting of consensus sequence A(CA), UMI sequence, and poly-T sequence) as templates, and a CCC overhang was added to the 3' end of the cDNA strand. After hybridization and annealing of the TSO sequence (Sequence ID 7, composed of consensus sequence B(CB) and a GGG overhang) with a cDNA strand (via complementary pair formation between the GGG at the end of the TSO sequence and the CCC overhang of the cDNA strand), the cDNA strand was continuously extended by reverse transcriptase using consensus sequence B as a template. As a result, the 3' end of the cDNA was labeled with a c(CB) tag (the complementary sequence of consensus sequence B).

[0207] [Table 3]

[0208] The synthesized cDNA strand consisted of the following sequences: reverse transcription primer sequence (SEQ ID NO: 6) - cDNA sequence - c(TSO) sequence (complementary sequence of SEQ ID NO: 7).

[0209] 2. Ligation of the chip sequence of the sequencing chip with cDNA After cDNA synthesis, the chip was washed twice with 5X SSC, and a 1 mL reaction system as shown in Table 4 was prepared. An appropriate volume of this system was injected into the chip to ensure that the chip was filled with the ligation reaction solution, and the reaction was carried out at room temperature for 30 minutes.

[0210] The above reaction ligated the 5' end of the cDNA sequence to the 3' end of the single-cell sequencing chip sequence (i.e., the 5' end of the cDNA sequence was labeled with the chip sequence) to obtain a novel nucleic acid molecule containing positional information (i.e., tag sequence Y), which consisted of the following sequence structure: chip sequence (SEQ ID NO: 8) - reverse transcription primer sequence (SEQ ID NO: 6) - cDNA sequence - c(TSO) sequence (complementary sequence of SEQ ID NO: 7).

[0211] After the reaction was complete, the chip was washed with 5X SSC. 200 μL of Bst polymerization reaction solution (NEB, M0275S) was prepared according to the instructions for use, injected into the chip, and reacted at 65°C for 60 minutes to obtain a single-stranded nucleic acid molecule with positional information.

[0212] [Table 4]

[0213] 3. Release of cDNA The chips were incubated at room temperature for 5 minutes using 75 μL of 80 mM KOH. After collecting the liquid, the cDNA recovery solution was neutralized by adding 10 μL of 1 M, pH 8.0 Tris-HCl.

[0214] 4. Amplification of cDNA 200 μL of the reaction system shown in Table 5 was prepared and used for 3' transcriptome sequencing and library construction, respectively, and then divided into two tubes for PCR:

[0215] [Table 5]

[0216] The above reaction system was placed in a PCR instrument, and the reaction program was set as follows: 3 minutes at 95°C, 11 cycles (20 seconds at 98°C, 20 seconds at 58°C, 3 minutes at 72°C), 5 minutes at 72°C, and infinity at 4°C. After the reaction was complete, XP beads (purchased from AMPure) were used for magnetic bead-based purification and recovery. The dsDNA concentration was quantified using a Qubit instrument, and the length distribution of the cDNA amplification product was detected using a 2100 bioanalyzer (purchased from Agilent). The detection results are shown in Figure 6.

[0217] Example 3: cDNA library construction and sequencing 1.Tn5 fragmentation According to the cDNA concentration, 20 ng of cDNA (obtained in step 4 of Example 2) was taken out, 0.5 μM Tn5 transposase and the corresponding buffer (purchased from BGI, catalog number: 10000028493; the Tn5 transposase was coated according to the instructions for use of Stereomics Library Preparation Kit-S1) were added, and the mixture was thoroughly mixed to obtain a 20 μL reaction system. The reaction was carried out at 55°C for 10 minutes, and then 5 μL of 0.1% SDS was added and the mixture was thoroughly mixed at room temperature for 5 minutes to terminate the Tn5 rearrangement.

[0218] 2. PCR amplification The following reaction system was prepared in 100 μL: [Table 6]

[0219] After mixing, the mixture was placed in a PCR instrument and the following program was set: 3 minutes at 95°C, 11 cycles (20 seconds at 98°C, 20 seconds at 58°C, 3 minutes at 72°C), 5 minutes at 72°C, 4°C infinity. After the reaction was complete, XP beads were used for magnetic bead-based purification and recovery. dsDNA concentration was quantified using a Qubit instrument.

[0220] 3. Sequencing 80 fmol of the above fragmented amplification product was isolated to prepare DNB. A 40 μL reaction system was prepared as follows: [Table 7]

[0221] The above reaction volume was placed in a PCR instrument for the reaction, and the reaction conditions were as follows: 95°C for 3 minutes, then 40°C for 3 minutes. After the reaction was complete, the resulting reaction solution was placed on ice, and 40 μL of mixed enzyme I, 2 μL of mixed enzyme II, 1 μL of ATP, and 0.1 μL of T4 ligase, which are necessary for DNB preparation in the DNBSEQ sequencing kit, were added. After mixing, the above reaction system was placed in a PCR instrument and reacted at 30°C for 20 minutes to form DNB.

[0222] Following the procedure described in the PE50 kit supporting MGISEQ 2000, the DNB was loaded onto the MGISEQ2000 sequencing chip, and sequencing was performed according to the relevant instructions for use. The PE50 sequencing model was selected, which divided the first strand sequencing into two sections: first measuring 25 bp, then performing a 15-cycle dark reaction, and then measuring a 10 bp UMI sequence. Second strand sequencing was then performed, measuring 50 bp.

[0223] Example 4: Data Analysis Alignment was used to match the 25 bp first strand sequence obtained by cDNA sequencing with the fq of the chip's positioning sequence (sequencing result in Example 1), and 25 bp matching reads from the two sequencing sequences were extracted for subsequent analysis. The second strand of the 25 bp matching reads from the two sequencing sequences was analyzed, and the second strand sequence was aligned with the mouse genome to extract unique mapping reads.

[0224] By mapping unique mapping reads to the spatial position of the chip, a mouse brain gene expression map as shown in Figure 7 was obtained. When Figure 7 is magnified to DNB-level (i.e., nanometer-level) resolution (Figure 8), clear single-cell aggregation of the tissue sample can be found, demonstrating that the method of this application reduces RNA diffusion, enables the detection of the spatial transcriptome, and achieves nanometer-level resolution.

[0225] Statistics on the average number of captured genes and UMIs per 40,000 DNBs indicate that approximately 4,000 genes and 10,000 mRNA molecules were captured on average, demonstrating that the method of this application was able to capture a sufficient number of genes and mRNA molecules for downstream analysis.

[0226] While specific embodiments of this application are described in detail, those skilled in the art will understand that various modifications and alterations can be made in detail based on all the disclosed technology, and all such alterations will be within the scope of protection of this application. The entire scope of this application is given by the appended claims and any equivalents thereof.

Claims

1. A method for generating a population of labeled nucleic acid molecules, comprising the following steps: (1) A step of providing a biological sample and a nucleic acid array, wherein the nucleic acid array includes a solid support, the solid support is bound to a plurality of types of oligonucleotide probes, each type of oligonucleotide probe includes at least one copy, and the oligonucleotide probe includes or comprises a consensus sequence X1, a tag sequence Y, and a consensus sequence X2 in the 5' to 3' direction. Steps include: various oligonucleotide probes having different tag sequences Y, and the tag sequence Y having a nucleotide sequence specific to the position of that type of oligonucleotide probe on the solid support; (2) A step of bringing the biological sample into contact with the nucleic acid array so that the position of RNA in the biological sample is mapped to the position of the oligonucleotide probe on the nucleic acid array, and pre-treating the RNA in the biological sample to generate a first nucleic acid molecule population, wherein the pre-treatment is (i) A step of performing reverse transcription of the RNA of the biological sample using primer A to generate an extension product as a first nucleic acid molecule to be labeled, thereby generating the first nucleic acid molecule population, wherein primer A includes a consensus sequence A and a capture sequence A, the capture sequence A is capable of annealing with the RNA to be captured to initiate the extension reaction, and the consensus sequence A is located upstream of the capture sequence A. or (ii) (a) A step of generating a cDNA strand by performing a reverse transcription of the RNA of the biological sample using primer A, wherein the cDNA strand comprises a cDNA sequence that is generated by the reverse transcription primed by primer A and is complementary to the RNA, and a 3'-terminal overhang, wherein primer A comprises a consensus sequence A and a capture sequence A, wherein the capture sequence A is capable of annealing with the RNA to be captured to initiate an extension reaction, and the consensus sequence A is located upstream of the capture sequence A, and (b) A step of generating primer B in (a) A step of annealing a cDNA strand and carrying out an extension reaction to produce a first extension product as a labeled first nucleic acid molecule, thereby generating a first nucleic acid molecule population, wherein the primer B includes a consensus sequence B and a complementary sequence of the 3'-terminal overhang, or the primer B includes a consensus sequence B, a complementary sequence of the 3'-terminal overhang, and a tag sequence B, wherein the complementary sequence of the 3'-terminal overhang is located at the 3'-end of the primer B, and the consensus sequence B is located upstream of the complementary sequence of the 3'-terminal overhang. or (iii) (a) A step of performing a reverse transcription of the RNA of the biological sample using primer A' to generate a cDNA strand, wherein the cDNA strand comprises a cDNA sequence generated by the reverse transcription primed by primer A' and complementary to the RNA, and a 3'-terminal overhang, wherein primer A' includes a capture sequence A, and the capture sequence A is capable of annealing with the captured RNA to initiate an extension reaction, (b) A step of annealing primer B with the cDNA strand generated in (a) to perform an extension reaction to generate a first extension product, wherein the (c) Providing an extension primer, carrying out an extension reaction using the first extension product as a template to produce a second extension product as a first nucleic acid molecule to be labeled, thereby producing a first nucleic acid molecule population. Steps including (3) A step in which a crosslinked oligonucleotide is brought into contact with the product of step (2) under conditions that enable annealing, the crosslinked oligonucleotide is annealed with the oligonucleotide probe and the labeled first nucleic acid molecule located at a position corresponding to the oligonucleotide probe, and the crosslinked oligonucleotide is ligated with the first nucleic acid molecule and the oligonucleotide probe on the array to obtain a ligation product as a second nucleic acid molecule having a positioning tag, thereby generating a second population of nucleic acid molecules. Includes, The crosslinked oligonucleotide comprises a first region and a second region, or the crosslinked oligonucleotide comprises a first region, a second region, and a third region located between the first region and the second region, wherein the first region is located upstream of the second region. The first region is annealable to all or part of the consensus sequence A of primer A in step (2)(i) or step (2)(ii), or is annealable to all or part of the consensus sequence B of primer B in step (2)(iii), A method wherein the second region is annealable with all or part of the consensus array X2.

2. In step (3), if the first and second regions of the crosslinked oligonucleotide are directly adjacent, the ligation of the first nucleic acid molecule to the oligonucleotide probe includes using a nucleic acid ligase to ligate the nucleic acid molecule hybridized to the first region of the same crosslinked oligonucleotide with the nucleic acid molecule hybridized to the second region, thereby obtaining a ligation product as a second nucleic acid molecule having a positioning tag, or The method according to claim 1, wherein the crosslinked oligonucleotide includes a first region, a second region and a third region located between them, the ligation of the first nucleic acid molecule to the oligonucleotide probe comprises carrying out a polymerization reaction using the third region as a template with a nucleic acid polymerase, and using a nucleic acid ligase to ligate the nucleic acid molecule hybridized to the first region of the same crosslinked oligonucleotide with the nucleic acid molecule hybridized to the third region and the second region to obtain a ligation product as the second nucleic acid molecule having a positioning tag.

3. The method according to claim 1 or 2, comprising steps (1), (2)(i), and (3), wherein the ligation product obtained in step (3) is taken as the second nucleic acid molecule having a positioning tag, wherein the 5' to 3' comprises (a) the consensus sequence X1, the tag sequence Y, the consensus sequence X2, and the sequence of the first nucleic acid molecule to be labeled, or (b) the consensus sequence X1, the tag sequence Y, the consensus sequence X2, the complementary sequence of the third region of the crosslinked oligonucleotide, and the sequence of the first nucleic acid molecule to be labeled.

4. The method according to claim 3, having one or more features selected from the following: (I) In step (2)(i), the capture sequence A is a random oligonucleotide sequence. (II) In step (2)(i), the capture sequence A is a random oligonucleotide sequence, and in step (3), the ligation product derived from each copy of the same oligonucleotide probe has a different capture sequence A, and the capture sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule. (III) In step (2)(i), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid. (IV) In step (2)(i), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, and the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule. (V) In step (2)(i), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule, and the tag sequence A is a random oligonucleotide sequence. (VI) In step (2)(i), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule, the capture sequence A is located at the 3'-terminus of the primer A, and the consensus sequence A is located upstream of the tag sequence A. (VII) In step (2)(i), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule, and in step (3), the ligation products derived from each copy of the same oligonucleotide probe have different tag sequences A as UMIs.

5. The method according to claim 1 or 2, comprising steps (1), (2)(ii), and (3), wherein the ligation product obtained in step (3) is taken as the second nucleic acid molecule having a positioning tag, wherein the 5' to 3' comprises (a) the consensus sequence X1, the tag sequence Y, the consensus sequence X2, and the sequence of the first nucleic acid molecule to be labeled, or (b) the consensus sequence X1, the tag sequence Y, the consensus sequence X2, the complementary sequence of the third region of the crosslinked oligonucleotide, and the sequence of the first nucleic acid molecule to be labeled.

6. The method according to claim 5, having one or more features selected from the following: (I) In step (2)(ii)(a), the capture sequence A is a random oligonucleotide sequence. (II) In step (2)(ii)(a), the capture sequence A is a random oligonucleotide sequence, and in step (3), the ligation product derived from each copy of the same oligonucleotide probe has a different capture sequence A, and the capture sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule. (III) In step (2)(ii)(a), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid. (IV) In step (2)(ii)(a), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, and the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule. (V) In step (2)(ii)(a), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule, and the tag sequence A is a random oligonucleotide sequence. (VI) In step (2)(ii)(a), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule, the capture sequence A is located at the 3'-terminus of the primer A, and the consensus sequence A is located upstream of the tag sequence A. (VII) In step (2)(ii)(a), the capture sequence A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further comprises a tag sequence A, the tag sequence A acts as a unique molecular identifier (UMI) of the second nucleic acid molecule, and in step (3), the ligation products derived from each copy of the same oligonucleotide probe have different tag sequences A as UMIs.

7. The method according to any one of claims 3 to 6, having one or more features selected from the following: (I) The primer A contains 5'-phosphate at its 5'-terminus, (II) Prior to step (3), the method further comprises a step of treating the product of step (2)(i) or step (2)(ii) to remove RNA.

8. The method according to claim 1 or 2, comprising steps (1), (2)(iii) and (3), wherein the ligation product obtained in step (3) is taken as the second nucleic acid molecule having a positioning tag, wherein the 5' to 3' comprises (a) the consensus sequence X1, the tag sequence Y, the consensus sequence X2, and the sequence of the first nucleic acid molecule to be labeled, or (b) the consensus sequence X1, the tag sequence Y, the consensus sequence X2, the complementary sequence of the third region of the crosslinked oligonucleotide, and the sequence of the first nucleic acid molecule to be labeled.

9. The method according to claim 8, having one or more features selected from the following: (I) In step (2) (iii) (c), the extension primer is the primer B. (II) In step (2)(iii)(c), the extension primer is primer B', and primer B' is capable of annealing with all or part of the complementary sequence of the consensus sequence B and initiating the extension reaction. (III) The primer B comprises the consensus sequence B, the complementary sequence of the 3'-terminal overhang, and the tag sequence B. (IV) In step (2)(iii)(a), the capture sequence A of the primer A' is a random oligonucleotide sequence. (V) The primer B comprises a consensus sequence B, a complementary sequence of the 3'-terminal overhang, and a tag sequence B, wherein the first extension product comprises a cDNA sequence complementary to the RNA, the 3'-terminal overhang sequence, a complementary sequence of the tag sequence B, and a complementary sequence of the consensus sequence B, generated from 5' to 3' by reverse transcription primed with the primer A', the complementary sequence of the tag sequence B, and the complementary sequence of the tag sequence B acts as a unique molecular identifier (UMI) of the second nucleic acid molecule. (VI) In step (2)(iii)(a), the capture sequence A of the primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid. (VII) In step (2)(iii)(a), the capture sequence A of the primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A' further comprises a tag sequence A and a consensus sequence A. (VIII) In step (2)(iii)(a), the capture sequence A of the primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A' further comprises a tag sequence A and a consensus sequence A, wherein the tag sequence A is a random oligonucleotide. (IX) The capture sequence A is located at the 3'-terminus of the primer A', (X) In step (2)(iii)(a), the capture sequence A of the primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A' further includes a tag sequence A and a consensus sequence A, wherein the consensus sequence A is located upstream of the capture sequence A. (XI) In step (2)(iii)(a), the capture sequence A of primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid, and primer A' further comprises a tag sequence A and a consensus sequence A, and primer B comprises a consensus sequence B, a complementary sequence of the 3'-terminal overhang and a tag sequence B, and in step (2)(iii)(b), the first extension product comprises, from 5' to 3', the consensus sequence A, the tag sequence A, a cDNA sequence generated by reverse transcription primed by primer A' and complementary to the RNA, the 3'-terminal overhang sequence, a complementary sequence of the tag sequence B and a complementary sequence of the consensus sequence B, (XII) The primer B comprises a consensus sequence B, a complementary sequence for the 3'-terminal overhang, and a tag sequence B, and in step (3), the ligation products derived from each copy of the same oligonucleotide probe have different tag sequences B as UMIs. (XIII) The extension primer contains 5'-phosphate at its 5'-terminus, (XIV) Prior to step (2)(iii)(c), the method further comprises the step of treating the product of step (2)(iii)(a) or step (2)(iii)(b) to remove RNA, (XV) In step (2)(iii)(b), the cDNA strand is annealed to the primer B via its 3'-terminal overhang, and the cDNA strand is extended using the primer B as a template in the presence of nucleic acid polymerase to produce a first extension product. (XVI) The 3'-terminal overhang has a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides. (XVII) The aforementioned 3'-terminal overhangs are the 3'-terminal overhangs of cytosine nucleotides 2 to 5.

10. The method according to any one of claims 1 to 9, having one or more features selected from the following: (I) In step (3), the annealing is insight annealing, (II) In step (3), the crosslinked oligonucleotide comprises the first region, the second region and the third region located between them, and the ligation of the first nucleic acid molecule to the oligonucleotide probe comprises carrying out a polymerization reaction using the third region as a template with a nucleic acid polymerase, and ligating the nucleic acid molecule hybridized to the first region of the same crosslinked oligonucleotide with the nucleic acid molecule hybridized to the third region and the second region using a nucleic acid ligase to obtain a ligation product as the second nucleic acid molecule having a positioning tag, wherein the nucleic acid polymerase does not have 5' to 3' exonuclease activity or chain displacement activity. (III) In step (2), the biological sample is subjected to permeabilization before the pretreatment. (IV) The biological sample is a tissue sample. (V) The biological sample is a tissue sample, and the tissue sample is a tissue section. (VI) The biological sample is a tissue sample, the tissue sample is a tissue section, and the tissue section is prepared from fixed tissue. (VII) The biological sample is a tissue sample, the tissue sample is a tissue section, and the tissue section is prepared from formalin-fixed paraffin-embedded (FFPE) tissue or rapidly frozen tissue. (VIII) In step (2), the reverse transcription is carried out using reverse transcriptase. (IX) In step (2), the reverse transcription is carried out using a reverse transcriptase, wherein the reverse transcriptase has terminal deoxynucleotidyltransferase activity. (X) In step (2), the reverse transcription is carried out using a reverse transcriptase, the reverse transcriptase is capable of synthesizing a cDNA strand using RNA as a template and adding an overhang to the 3' end of the cDNA strand. (XI) In step (2), the reverse transcription is carried out using a reverse transcriptase, wherein the reverse transcriptase is capable of adding an overhang to the 3'-end of the cDNA strand having a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides. (XII) In step (2), the reverse transcription is carried out using a reverse transcriptase, wherein the reverse transcriptase is capable of adding an overhang of 2 to 5 cytosine nucleotides to the 3'-end of the cDNA strand. (XIII) In step (2), the reverse transcription is carried out using a reverse transcriptase, the reverse transcriptase being selected from the group consisting of M-MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and their variants, modified products and derivatives having reverse transcription activity.

11. The method according to any one of claims 1 to 10, wherein steps (2) and (3) have one or more features selected from the following: (1) Each of the primers A, A', B and crosslinked oligonucleotides independently contains or consists of a natural nucleotide, a modified nucleotide, a non-natural nucleotide, or any combination thereof. (2) Primers A and A' can initiate the extension reaction. (3) The primer B contains a modified nucleotide, (4) The primer B contains one or more modified nucleotides at its 3'-terminus. (5) The primer B contains locked nucleic acid nucleotides, (6) The primer B contains one or more locked nucleic acid nucleotides at its 3'-terminus. (7) The tag array B has a length of 5 to 200 nt, (8) The primer A includes a tag sequence A, and the tag sequence A has a length of 5 to 200 nt. (9) The primer A includes a tag sequence A, the tag sequence A has a length of 5 to 200 nt, and the tag sequence A is a random oligonucleotide. (10) The primer A' includes a tag sequence A, and the tag sequence A has a length of 5 to 200 nt. (11) The primer A' includes a tag sequence A, the tag sequence A has a length of 5 to 200 nt, and the tag sequence A is a random oligonucleotide. (12) The consensus sequence A and consensus sequence B each independently have a length of 10 to 200 nt. (13) Each of the primers A, A' and B has a length of 10 to 200 nt independently. (14) The first region and the second region of the crosslinked oligonucleotide each independently have a length of 3 to 100 nt. (15) The third region of the crosslinked oligonucleotide has a length of 0 to 100 nt. (16) The crosslinked oligonucleotide has a length of 6 to 200 nt, (17) The capture sequence A of the primer A is a poly(T) sequence, and the poly(T) sequence contains at least 10, at least 20, or at least 30 deoxythymidine residues. (18) The capture sequence A of the primer A' is a poly(T) sequence, and the poly(T) sequence contains at least 10, at least 20, or at least 30 deoxythymidine residues.

12. The method according to any one of claims 1 to 11, having one or more features selected from the following: (i) The method further comprises (4) the step of recovering and purifying the second nucleic acid molecule population, (II) The obtained second population of nucleic acid molecules and / or its complements are used to construct a transcriptome library or for transcriptome sequencing. (III) The consensus sequence X1, tag sequence Y, and consensus sequence X2 each independently contain a natural nucleotide, a modified nucleotide, a non-natural nucleotide, or any combination thereof. (IV) The consensus sequence X2 has a free hydroxyl group (-OH) at its 3'-terminus. (V) The consensus array X1, tag array Y, and consensus array X2 each independently have a length of 2 to 100 nt. (VI) Each of the oligonucleotide probes independently has a length of 15 to 200 nt. (VII) The oligonucleotide probe is bound to the solid support via a linker, (VIII) The oligonucleotide probe is bound to the solid support via a linker, the linker is a linking group capable of binding to an activating group, and the surface of the solid support is modified using the activating group. (IX) The oligonucleotide probe is bound to the solid support via a linker, the linker comprising -SH, -DBCO, or -NHS, (X) The oligonucleotide probe is bonded to the solid support via a linker, the linker is -DBCO, and the surface of the solid support is 【Chemistry 1】 Modified with (azido-dPEG® 8-NHS ester), (XI) The oligonucleotide probes bound to the same solid support have the same consensus sequence X1 and / or the same consensus sequence X2, (XII) The consensus sequence X1 of the oligonucleotide probe includes a cleavage site, (XIII) The consensus sequence X1 of the oligonucleotide probe includes a cleavage site, the cleavage site can be cleaved or destroyed by a method selected from Nicking enzyme digestion, USER enzyme digestion, photoresponsive excision, chemical excision or CRISPR-mediated excision.

13. The nucleic acid array in step (1) (1) A step of providing a plurality of types of carrier sequences, wherein each type of carrier sequence includes at least one copy of the carrier sequence, and the carrier sequence includes, in the 5' to 3' direction, a complementary sequence of the consensus sequence X2, a complementary sequence of the tag sequence Y, and an immobilization sequence, and the complementary sequences of the tag sequence Y of each type of carrier sequence are different from each other. (2) The step of bonding the plurality of carrier arrays to the surface of the solid support, (3) A step of providing an immobilized primer, using the carrier sequence as a template to carry out a primer extension reaction to produce an extension product and obtain the oligonucleotide probe, wherein the immobilized primer includes the sequence of the consensus sequence X1 and is capable of annealing with the immobilized sequence of the carrier sequence to initiate the extension reaction. (4) A step of connecting the immobilized primer to the surface of the solid support, wherein steps (3) and (4) are carried out in any order, The method according to any one of claims 1 to 12, provided in step (1).

14. The method according to claim 13, wherein the immobilized sequence of the carrier sequence further includes cleavage sites, the cleavage sites can be cleaved by a method selected from the group consisting of Nicking enzyme digestion, USER enzyme digestion, photoresponsive excision, chemical excision and CRISPR-mediated excision, and the cleavage is carried out at the cleavage sites included in the immobilized sequence of the carrier sequence to digest the carrier sequence, thereby separating the extension product in step (3) from the template on which the extension product is generated, thereby linking the oligonucleotide probe to the surface of the solid support.

15. The method according to claim 13 or 14, having one or more features selected from the following: (I) Various carrier sequences are DNBs formed from concatemers of multiple copies of the carrier sequence. (II) The above-mentioned multiple types of carrier sequences are as follows: (i) A step of providing a plurality of types of carrier-template sequences, wherein the carrier-template sequence includes a complementary sequence of a carrier sequence, (ii) A step of carrying out a nucleic acid amplification reaction using various carrier-template sequences as templates to obtain amplification products of various carrier-template sequences, wherein the amplification product includes at least one copy of the carrier sequence. Provided in step (1), (III) The above-mentioned multiple types of carrier sequences are as follows: (i) A step of providing a plurality of types of carrier-template sequences, wherein the carrier-template sequence includes a complementary sequence of a carrier sequence, (ii) A step of performing rolling circle replication using various carrier-template arrays as templates to obtain a DNB formed from concatemers of the carrier arrays. This is provided in step (1).

16. The method according to any one of claims 1 to 15, wherein the solid support in step (1) has one or more features selected from the following: (1) The solid-phase support is selected from the group consisting of latex beads, dextran beads, polystyrene surface, polypropylene surface, polyacrylamide gel, gold surface, glass surface, chip, sensor, electrode, and silicon wafer. (2) The solid support is a chip, (3) The solid phase support is planar, spherical or porous, (4) The solid-phase support can be used as a sequencing platform, (5) The solid-phase support is capable of releasing the oligonucleotide probe spontaneously or upon exposure to one or more stimuli.

17. A method for constructing a nucleic acid molecule library, (a) the step of generating a population of labeled nucleic acid molecules by the method described in any one of claims 1 to 16, and (b) The step of randomly fragmenting the nucleic acid molecules in the labeled nucleic acid molecule population and adding adapters to them. A method comprising obtaining a library of nucleic acid molecules.

18. The method according to claim 17, having one or more features selected from the following: (1) The method further comprises a step of amplifying and / or concentrating the product of step (b), (2) The nucleic acid molecule library is used for sequencing, (3) The nucleic acid molecule library is used for transcriptome sequencing or single-cell transcriptome sequencing.

19. A method for sequencing nucleic acid samples, (1) The step of constructing a nucleic acid molecule library by the method described in claim 17 or 18, and (2) Step of sequencing the nucleic acid molecule library A method that includes this.

20. (i) A nucleic acid array for labeling nucleic acids, comprising a solid support, wherein the solid support is bound to a plurality of types of oligonucleotide probes, each type of oligonucleotide probe comprising at least one copy, wherein the oligonucleotide probe comprises or consists of a consensus sequence X1, a tag sequence Y, and a consensus sequence X2 in the 5' to 3' direction. A nucleic acid array in which various oligonucleotide probes have different tag sequences Y, and the tag sequence Y has a nucleotide sequence specific to the position of that type of oligonucleotide probe on the solid support, (ii) A primer set comprising primer A, or primer A' and primer B, or a primer set comprising primer A and primer B, The primer A comprises a consensus sequence A and a capture sequence A, and the capture sequence A is capable of annealing with the captured RNA and initiating an extension reaction. The primer A' includes a capture sequence A, and the capture sequence A can anneal to the captured RNA and initiate an extension reaction. The primer B comprises a consensus sequence B and a complementary sequence for the 3'-terminal overhang, or the primer B comprises a consensus sequence B, a complementary sequence for the 3'-terminal overhang, and a tag sequence B, and the primer A is a set comprising primer A' and primer B, or a set comprising primer A and primer B. and (iii) A crosslinked oligonucleotide comprising (a) a first region and a second region, or (b) a first region, a second region, and a third region located between the first region and the second region, wherein the first region is located upstream of the second region. The first region is (a) annealable to all or part of the consensus sequence A of the primer A, or (b) annealable to all or part of the consensus sequence B of the primer B. The second region is a cross-linked oligonucleotide that can anneal to all or part of the consensus sequence X2. A kit that includes this.

21. The kit according to claim 20, having one or more features selected from the following: (1) The consensus sequence A is located upstream of the capture sequence A, (2) The 3'-terminal overhang refers to one or more non-template nucleotides included in the 3'-end of the cDNA strand produced by reverse transcription using the RNA captured by the capture sequence A of the primer A' as a template. (3) The complementary arrangement of the 3'-terminal overhang is located at the 3'-terminal side of primer B. (4) The consensus sequence B is located upstream of the complementary sequence in the 3'-terminal overhang.

22. (i) A nucleic acid array for labeling the nucleic acid described in (i), a primer A described in (ii), and a crosslinked oligonucleotide described in (iii), wherein the first region of the crosslinked oligonucleotide is annealable with all or part of the consensus sequence A of the primer A, and the second region of the crosslinked oligonucleotide is annealable with all or part of the consensus sequence X2. The kit according to claim 20 or 21.

23. The kit according to claim 22, having one or more features selected from the following: (1) The capture sequence A of the primer A is a random oligonucleotide sequence. (2) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid. (3) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A further comprises a tag sequence A. (4) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A further includes a tag sequence A, and the tag sequence A is a random oligonucleotide sequence. (5) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further includes a tag sequence A, the capture sequence A is located at the 3' end of the primer A, and the consensus sequence A is located upstream of the tag sequence A. (6) The primer A contains 5'-phosphate at its 5'-terminus.

24. (i) A nucleic acid array for labeling the nucleic acid described in (i), a primer set comprising primer A' and primer B described in (ii), and a crosslinked oligonucleotide described in (iii), wherein the first region of the crosslinked oligonucleotide is annealable with all or part of the consensus sequence B of primer B, and the second region of the crosslinked oligonucleotide is annealable with all or part of the consensus sequence X2. The kit according to claim 20 or 21.

25. The kit according to claim 24, having one or more features selected from the following: (1) The capture sequence A of the primer A' is a random oligonucleotide sequence. (2) The capture sequence A of the primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid. (3) The capture sequence A of the primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A' further comprises a tag sequence A and a consensus sequence A. (4) The capture sequence A is located at the 3' end of the primer A', (5) The capture sequence A of the primer A' is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A' further includes a tag sequence A and a consensus sequence A, wherein the consensus sequence A is located upstream of the capture sequence A. (6) The primer B includes the consensus sequence B, the complementary sequence of the 3'-terminal overhang, and the tag sequence B. (7) The kit further comprises primer B', wherein primer B' is capable of annealing with all or part of the complementary sequence of consensus sequence B to initiate an extension reaction. (8) The primer B contains 5'-phosphate at its 5'-terminus, (9) The kit further comprises primer B', wherein primer B' is capable of annealing with all or part of the complementary sequence of consensus sequence B to initiate an extension reaction, and primer B' contains 5'-phosphate at its 5'-terminus. (10) The primer B contains a modified nucleotide, (11) The primer B contains locked nucleic acid nucleotides, (12) The primer B contains one or more modified nucleotides at its 3' end, (13) The primer B contains one or more locked nucleic acid nucleotides at its 3'-terminus.

26. (i) A nucleic acid array for labeling the nucleic acid described in (i), a primer set comprising primer A and primer B described in (ii), and a crosslinked oligonucleotide described in (iii), wherein the first region of the crosslinked oligonucleotide is annealable with all or part of the consensus sequence A of primer A, and the second region of the crosslinked oligonucleotide is annealable with all or part of the consensus sequence X2. The kit according to claim 20 or 21.

27. ​​The kit according to claim 26, having one or more features selected from the following: (1) The capture sequence A of the primer A is a random oligonucleotide sequence. (2) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid. (3) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A further comprises a tag sequence A. (4) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid, and the primer A further includes a tag sequence A, and the tag sequence A is a random oligonucleotide sequence. (5) The capture sequence A of the primer A is a poly(T) sequence or a specific sequence of the target nucleic acid, the primer A further includes a tag sequence A, the capture sequence A is located at the 3' end of the primer A, and the consensus sequence A is located upstream of the tag sequence A. (6) The primer A contains 5'-phosphate at its 5'-terminus, (7) The primer B contains a modified nucleotide, (8) The primer B contains locked nucleic acid nucleotides, (9) The primer B contains one or more modified nucleotides at its 3'-terminus, (10) The primer B contains one or more locked nucleic acid nucleotides at its 3'-terminus.

28. A kit according to any one of claims 20 to 27, having one or more features selected from the following: (1) The oligonucleotide probe, primer A, primer A', primer B, and crosslinked oligonucleotide each independently contain or consist of a natural nucleotide, a modified nucleotide, a non-natural nucleotide, or any combination thereof. (2) The kit further comprises primer B', wherein primer B' is capable of annealing with all or part of the complementary sequence of consensus sequence B to initiate an extension reaction, wherein primer B' contains a 5'-phosphate at its 5'-terminus, and primer B' contains or consists of a natural nucleotide, a modified nucleotide, a non-natural nucleotide, or any combination thereof. (3) The primer A or primer A' is capable of initiating an extension reaction. (4) The consensus sequence X2 has a free hydroxyl group (-OH) at its 3'-terminus. (5) Each of the oligonucleotide probes independently has a length of 15 to 300 nt. (6) Each of the primers A, A', and B has a length of 10 to 200 nt independently. (7) The kit further comprises a primer B' which is capable of annealing with all or part of the complementary sequence of the consensus sequence B to initiate an extension reaction, wherein the primer B' contains a 5'-phosphate at its 5'-terminus, and the primer B' has a length of 10 to 200 nt. (8) The crosslinked oligonucleotide has a length of 6 to 200 nt, (9) Oligonucleotide probes bound to the same solid support have the same consensus sequence X1 and / or the same consensus sequence X2, (10) The consensus sequence X1 of the oligonucleotide probe includes a cleavage site. (11) The consensus sequence X1 of the oligonucleotide probe includes a cleavage site, and the cleavage site can be cleaved or destroyed by a method selected from Nicking enzyme digestion, USER enzyme digestion, photoresponsive excision, chemical excision or CRISPR-mediated excision. (12) The kit further comprises a reverse transcriptase, a nucleic acid ligase, a nucleic acid polymerase and / or a transposase, (13) The kit comprises a reverse transcriptase, wherein the reverse transcriptase has terminal deoxynucleotidyltransferase activity. (14) The kit includes a reverse transcriptase capable of synthesizing a cDNA strand using RNA as a template and adding a 3'-terminal overhang to the 3' end of the cDNA strand. (15) The kit comprises a reverse transcriptase capable of adding an overhang to the 3'-end of the cDNA strand having a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides. (16) The kit comprises a reverse transcriptase capable of adding an overhang of 2 to 5 cytosine nucleotides to the 3'-end of the cDNA strand. (17) The kit comprises a reverse transcriptase, wherein the reverse transcriptase is selected from the group consisting of M-MLV reverse transcriptase, HIV-1 reverse transcriptase, AMV reverse transcriptase, telomerase reverse transcriptase, and variants, modified products and derivatives thereof having reverse transcription activity of the above reverse transcriptases. (18) The kit comprises a nucleic acid polymerase, and the nucleic acid polymerase does not have 5' to 3' exonuclease activity or chain displacement activity. (19) The kit comprises a transposase, wherein the transposase is selected from the group consisting of Tn5 transposase, MuA transposase, Sleeping Beauty transposase, Mariner transposase, Tn7 transposase, Tn10 transposase, Ty1 transposase, Tn552 transposase, and variants, modified products and derivatives thereof having transposase activity. (20) The kit further comprises reagents for nucleic acid hybridization, reagents for nucleic acid extension, reagents for nucleic acid amplification, reagents for recovery or purification of nucleic acids, reagents for constructing a transcriptome sequencing library, reagents for sequencing, or any combination thereof.

29. Use of the method according to any one of claims 1 to 16 or the kit according to any one of claims 20 to 28 for constructing a nucleic acid molecular library or for performing transcriptome sequencing.

Citation Information

Patent Citations

  • Space transcriptome database building and sequencing method and device adopted for same

    CN105505755A

  • Methods and products for localized or spatial detection of nucleic acids in tissue samples.

    JP2014513523A

  • Methods for determining a location of a biological analyte in a biological sample

    US20200277663A1

  • Deterministic barcoding for spatial omics sequencing

    US20210095331A1

  • High resolution spatial genomic analysis of tissues and cell aggregates

    WO2018075436A1