Spatial Transcriptomics Library Preparation Materials and Methods

By employing capture oligonucleotides to hybridize and ligate gene-specific probes to mRNA transcripts, the method overcomes RNA degradation issues in FFPE samples, resulting in more complete transcriptomics libraries for accurate gene expression analysis.

JP2026501486APending Publication Date: 2026-01-16ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024576648
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2023-12-29
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing methods for generating spatial transcriptomics mRNA libraries from formalin-fixed, paraffin-embedded (FFPE) tissue samples face challenges due to RNA degradation and chemical modification during tissue processing, making poly(A) capture of mRNA more difficult compared to fresh-frozen tissue.

Method used

The method involves using capture oligonucleotides with specific sequences to hybridize and ligate gene-specific probes to mRNA transcripts, followed by ligation and capture on a substrate, improving the efficiency of mRNA transcript capture from FFPE and fresh-frozen tissue samples.

Benefits of technology

This approach enhances the generation of more complete transcriptomics libraries by effectively capturing mRNA transcripts from FFPE and fresh-frozen tissues, enabling accurate gene expression profiling and correlation with disease states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501486000004
    Figure 2026501486000004
  • Figure 2026501486000005
    Figure 2026501486000005
  • Figure 2026501486000006
    Figure 2026501486000006
Patent Text Reader

Abstract

The present disclosure generally relates to methods for improving the preparation of spatial transcriptomics RNA libraries, such as mRNA libraries, by improving the capture of RNA transcript information from tissue samples in situ. Spatial transcriptomics libraries from tissue samples are useful for determining genetic profiles, aiding in the diagnosis of individuals having or at risk of having diseases such as cancer, genetic diseases, autoimmune diseases, and other indications, and for improving treatment of subjects.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 477,726, filed December 29, 2022, and U.S. Provisional Patent Application No. 63 / 612,819, filed December 20, 2023, which are incorporated by reference in their entireties.

[0002] Incorporation by reference of sequence disclosure A sequence listing, which is part of this disclosure, is submitted herewith as a computer-readable file. The file containing the sequence listing is named "IP-2535-PC_SeqListing.xml," created on December 21, 2023, and is 20,735 bytes in size. The subject matter of the sequence listing is incorporated herein by reference in its entirety.

[0003] The present disclosure relates generally to methods for generating spatial transcriptomics mRNA libraries by improving methods for capturing mRNA transcripts from in situ samples, and to mRNA libraries produced by these methods. [Background technology]

[0004] Spatial transcriptomics enables high-resolution in situ gene expression profiling, where cellular relationships are captured within complex tissue structures. Formalin-fixed, paraffin-embedded (FFPE) tissue is an invaluable resource for cancer research because it is the most widely available material for which patient outcomes are known (recent estimates suggest over 1 billion FFPE samples worldwide). However, formalin fixation and subsequent decrosslinking are known to cause RNA degradation and chemical modification during tissue processing, making poly(A) capture of mRNA more challenging than with fresh-frozen tissue. Summary of the Invention

[0005] The present disclosure provides improved methods for generating mRNA transcript libraries from in situ samples, such as fresh-frozen or formalin-fixed, paraffin-embedded tissue samples, by improving the efficiency of capture of mRNA transcripts from tissue samples, thereby generating more complete transcriptomics libraries. The methods are useful for isolating genomic information from samples such as tumor biopsies or other tissues in patients suffering from disease and correlating the genetic information with having or being at risk of having, or developing, the disease.

[0006] In one aspect, the disclosure provides a method for preparing an mRNA transcript expression library from a tissue sample, the method comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, wherein the capture oligonucleotides comprise a first clustering sequence (e.g., P7), a spatial barcode sequence (SBC), and a first universal adaptor sequence (e.g., Rd2 adaptor); b) contacting the tissue sample with i) a plurality of 5' gene-specific probes comprising a 5' gene-specific primer and a sequence complementary to the first universal adaptor sequence, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence (e.g., Rd1 adaptor), under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample; c) contacting the 5' gene-specific probes that are hybridized to the mRNA transcripts in close proximity to each other. contacting the tissue sample of (b) with a ligation reagent such that the 5' gene-specific probe and the 3' gene-specific probe are ligated together to form one or more ligated gene-specific probe pairs; d) removing mRNA transcripts hybridized to the ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; and e) capturing the ligated gene-specific probe pair oligonucleotides of (d) on a substrate by binding a sequence complementary to a first universal adapter sequence in the 5' gene-specific probe to a first universal adapter sequence (e.g., an Rd2 adapter) of the capture oligonucleotide.

[0007] A method for determining mRNA transcript expression in a tissue sample, comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, wherein the capture oligonucleotides comprise a first clustered sequence (e.g., P7), a spatial barcode sequence (SBC), and a first universal adaptor sequence (e.g., Rd2 adaptor); b) coupling the tissue sample to i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence and a 5' gene-specific primer, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence (e.g., Rd1 adaptor), such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample. c) contacting the tissue sample of (b) with a ligation reagent such that the 5' gene-specific probe and the 3' gene-specific probe that are hybridized to the mRNA transcripts adjacent to each other are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts that hybridized to the ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; and e) capturing the ligated gene-specific probe pair oligonucleotides of (d) on a substrate by binding a sequence complementary to the first universal adapter sequence in the 5' gene-specific probe to a first universal adapter sequence (e.g., an Rd2 adapter) of the capture oligonucleotide.

[0008] In another aspect, the disclosure provides a method for preparing an mRNA transcript expression library from a tissue sample and / or a method for determining mRNA transcript expression from a tissue sample, the method comprising: a) loading the tissue sample onto a substrate comprising a plurality of capture oligonucleotides, wherein the capture oligonucleotides comprise a first clustering sequence (e.g., P7), a spatial barcode sequence (SBC), and a first universal adaptor sequence (e.g., Rd2 adaptor); b) sequencing the tissue sample with: i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence, a unique molecular index, and a 5' gene-specific primer; and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer and a second universal adaptor sequence (e.g., Rd1 adaptor), and wherein one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes are present in a single sequence in the tissue sample. c) contacting the tissue sample of (b) with a ligation reagent such that the 5' gene-specific probe and the 3' gene-specific probe that are hybridized to the mRNA transcripts adjacent to each other are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts that hybridized to the ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; and e) capturing the ligated gene-specific probe pair oligonucleotides of (d) on a substrate by binding a sequence complementary to the first universal adapter sequence in the 5' gene-specific probe to a first universal adapter sequence (e.g., an Rd2 adapter) of the capture oligonucleotide.

[0009] In various embodiments, the disclosure provides a method for preparing an mRNA transcript expression library from a tissue sample, comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, wherein the capture oligonucleotides comprise a first clustering sequence (e.g., P7), a spatial barcode sequence (SBC), and a first universal adaptor sequence (e.g., Rd2 adaptor); and b) coupling the tissue sample to: i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence and a 5' gene-specific primer, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence (e.g., Rd1 adaptor), under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample, wherein hybridization of the 5' gene-specific probes and the 3' gene-specific probes on the mRNA transcripts results in hybridized c) contacting the tissue sample of (b) with nucleotide bases and a ligation reagent such that gaps between the 5' gene-specific probe and the 3' gene-specific probe hybridized to the mRNA transcripts are filled with nucleotide bases complementary to the mRNA transcripts, and the 5' gene-specific probe and the 3' gene-specific probe are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts hybridized to the ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; and e) capturing the ligated gene-specific probe pair oligonucleotide sequences of (d) on a substrate by binding a sequence complementary to a first universal adaptor sequence in the 5' gene-specific probe to a first universal adaptor sequence (e.g., an Rd2 adaptor) of a capture oligonucleotide.

[0010] The present disclosure provides a method for determining mRNA transcript expression in a tissue sample, comprising: a) loading the tissue sample onto a substrate comprising a plurality of capture oligonucleotides, wherein the capture oligonucleotides comprise a first clustered sequence (e.g., P7), a spatial barcode sequence (SBC), and a first universal adaptor sequence (e.g., Rd2 adaptor); and b) coupling the tissue sample to: i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence and a 5' gene-specific primer; and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence (e.g., Rd1 adaptor), under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample, wherein hybridization of the one or more 5' gene-specific probes and the one or more 3' gene-specific probes on the mRNA transcripts results in nucleic acid sequences between the hybridized molecules. c) contacting the tissue sample of (b) with nucleotide bases and a ligation reagent such that nucleotide gaps between the 5' gene-specific probe and the 3' gene-specific probe hybridized to the mRNA transcript are filled with nucleotide bases complementary to the mRNA transcript, and the 5' gene-specific probe and the 3' gene-specific probe are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts hybridized to the ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; and e) capturing the ligated gene-specific probe pair oligonucleotide sequences of (d) on a substrate by binding a sequence complementary to the first universal adapter sequence in the 5' gene-specific probe to a first universal adapter sequence (e.g., an Rd2 adapter) of a capture oligonucleotide.

[0011] In various embodiments, the disclosure provides a method for preparing an mRNA transcript expression library from a tissue sample and / or a method for determining mRNA transcript expression from a tissue sample, the method comprising: a) loading the tissue sample onto a substrate comprising a plurality of capture oligonucleotides, wherein the capture oligonucleotides comprise a first clustering sequence (e.g., P7), a spatial barcode sequence (SBC), and a first universal adaptor sequence (e.g., Rd2 adaptor); and b) coupling the tissue sample to i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence, a unique molecular index, and a 5' gene-specific primer, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer and a second universal adaptor sequence (e.g., Rd1 adaptor), under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample, wherein hybridization of the 5' gene-specific probes and the 3' gene-specific probes on the mRNA transcripts occurs. c) contacting the tissue sample of (b) with nucleotide bases and a ligation reagent under conditions such that the ligation results in a nucleotide gap between the hybridized molecules; c) contacting the tissue sample of (b) with nucleotide bases and a ligation reagent such that gaps between the 5' gene-specific probe and the 3' gene-specific probe hybridized to the mRNA transcript are filled with nucleotide bases complementary to the mRNA transcript and the 5' gene-specific probe and the 3' gene-specific probe are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts hybridized to the ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; and e) capturing the ligated gene-specific probe pair oligonucleotide sequences of (d) on a substrate by binding a sequence complementary to a first universal adaptor sequence in the 5' gene-specific probe to a first universal adaptor sequence (e.g., an Rd2 adaptor) of a capture oligonucleotide.

[0012] In various embodiments, the nucleotide gap is between 1 and 50 or more nucleotides, including 50 or more nucleotides, between 1 and 50 nucleotides, between 1 and 40 nucleotides, between 1 and 30 nucleotides, between 1 and 20 nucleotides, or between 1 and 10 nucleotides.

[0013] In various embodiments, the tissue sample is a fresh tissue sample, a fresh frozen tissue sample, or a formalin-fixed paraffin-embedded (FFPE) tissue sample.

[0014] It is contemplated that the method may further include indexing and sequencing the ligated gene-specific probe pairs, including: f) performing an extension reaction and PCR on the oligonucleotides of (e) to obtain PCR templates representing one or more mRNA transcripts in the tissue sample; g) eluting the PCR templates; and h) performing indexing PCR to generate double-stranded PCR products comprising a first-strand PCR product and a second strand complementary to the first-strand PCR product. In various embodiments, the method may further include sequencing the PCR products of (h) and determining the location of the mRNA transcripts in the tissue based on the spatial barcode (SBC) sequences.

[0015] The present disclosure provides improved methods for generating RNA libraries, e.g., mRNA libraries, from tissue samples, e.g., fresh-frozen or formalin-fixed, paraffin-embedded tissue samples, by improving the efficiency of capture of mRNA transcripts from the tissue sample, thereby creating a more complete transcriptomics library.

[0016] Existing targeting ex situ spatial approaches often involve the ligation of probe pairs to RNA targets in tissue.Unless gap filling and then ligation are performed, no sequence information can be obtained from RNA, and instead the ligated probes are counted through sequencing.For example, if mutations (such as SNVs or altered splice junctions) exist in RNA, they will not be detected.

[0017] Multiple methods have been proposed to capture RNA using targeted probes that can hybridize to substrate-linked probes containing spatially barcoded sequences for RNA library preparation.

[0018] In one aspect, the disclosure provides a method for preparing a spatially barcoded RNA library from a tissue sample, the method comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide sequence complementary to an RNA in the sample and a first substrate capture oligonucleotide complementary to a first domain of a plurality of splint oligonucleotides; (b) hybridizing the RNA capture oligonucleotides of the RNA capture probes to the RNA in the tissue sample to form RNA-RNA capture probe hybrids; and (c) using a reverse transcriptase to extend the RNA capture oligonucleotides of the RNA-RNA capture probe hybrids to produce a plurality of first strand cDNA molecules, the first strand cDNA molecules comprising the first strand cDNA molecules. (d) forming a plurality of first strand cDNA molecules, each of the DNA molecules comprising an RNA capture oligonucleotide and a first substrate capture oligonucleotide; (d) capturing the first strand cDNA molecules onto a substrate, the substrate comprising a plurality of substrate capture probes, each of which comprises a spatial barcode and a second substrate capture oligonucleotide complementary to a second domain of the splint oligonucleotide, wherein capturing comprises hybridizing the splint oligonucleotide to the first substrate capture oligonucleotide of the first strand cDNA molecule and the second substrate capture oligonucleotide of the substrate capture probe; and (e) ligating the captured first strand cDNA molecules to the substrate capture probes, thereby forming spatially barcoded first strand cDNA molecules.

[0019] In various embodiments, the substrate capture probe further comprises a substrate anchor portion.

[0020] In various embodiments, the surface oligonucleotides further comprise a P7 adaptor and an RNA capture probe primer for reading the spatial barcode sequence.

[0021] A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, the RNA capture probes comprising an RNA capture oligonucleotide and a handle sequence complementary to the RNA in the sample; (b) hybridizing the RNA capture oligonucleotides of the RNA capture probes to the RNA in the tissue sample to form RNA-RNA capture probe hybrids; and (c) using a reverse transcriptase to extend the RNA capture oligonucleotides of the RNA-RNA capture probe hybrids to produce a plurality of first strand cDNA molecules, each of the first strand cDNA molecules comprising an RNA capture oligonucleotide and a handle sequence. (d) adding a 3' terminal oligonucleotide to the 3' end of each first strand cDNA molecule, wherein the 3' terminal oligonucleotide comprises a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on the substrate, each of the plurality of substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, and a first domain; (e) hybridizing the substrate capture oligonucleotide of the first strand cDNA molecule to the first domain of the substrate capture probe; and (f) extending the first domain of the hybridized substrate capture probe to form the plurality of spatially barcoded first strand cDNA molecules.

[0022] In various embodiments, the handle sequence is a PCR handle sequence, a molecular identifier, a UMI, or any combination thereof. In various embodiments, the handle sequence is a P5 adapter sequence.

[0023] In various embodiments, the 3'-terminal oligonucleotide is added by tagging. In various embodiments, the 3'-terminal oligonucleotide is added by click chemistry or oNTP-directed adaptorization. In various embodiments, the 3'OH is added by terminating the extension reaction with a click-labeled nucleotide. In various embodiments, the click-labeled nucleotide is an azide- or alkyne-labeled oligonucleotide. In various embodiments, the extension reaction adds a polyA sequence to the 3'-extended sequence.

[0024] In various embodiments, the first strand cDNA is captured at a poly-T sequence on a surface capture oligonucleotide.

[0025] 1. A method for preparing a spatially barcoded RNA library from a tissue sample, the method comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, the RNA capture probes comprising an RNA capture oligonucleotide complementary to the RNA in the sample and a handle sequence; (b) hybridizing the RNA capture oligonucleotides of the RNA capture probes to the RNA in the tissue sample to form RNA-RNA capture probe hybrids; (c) using a reverse transcriptase to extend the RNA capture oligonucleotides of the RNA-RNA capture probe hybrids to form a plurality of first strand cDNA molecules, each first strand cDNA molecule comprising an RNA capture oligonucleotide and a handle sequence; and (d) synthesizing the first strand cDNA molecules with a reverse transcriptase (RT) and a template switch oligonucleotide (TSO), wherein the RT incorporates a non-templated cytosine nucleotide at the 3' end of the first cDNA and the TSO incorporates a non-templated cytosine nucleotide at the 3' end of the first cDNA. a 3'-end of each first strand cDNA molecule via template switching, comprising contacting the first strand cDNA molecule with a reverse transcriptase (RT) and a template switch oligonucleotide (TSO) comprising a sequence capable of hybridizing to a non-template cytosine nucleotide, wherein the 3'-end oligonucleotide is added to the 3'-end of the first cDNA and the RT extends to generate a TSO complement, the 3'-end oligonucleotide comprising a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the plurality of substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, and a first domain; (e) hybridizing the substrate capture oligonucleotide of the first strand cDNA molecule to the first domain of the substrate capture probe; and (f) extending the first domain of the hybridized substrate capture probe to form a plurality of spatially barcoded first strand cDNA molecules.

[0026] In various embodiments, the substrate capture probe is released from the substrate before hybridizing the substrate capture oligonucleotide to the first domain of the substrate capture probe.

[0027] In various embodiments, the first domain is a poly-T sequence.

[0028] In another aspect, the disclosure provides a method for preparing a spatially barcoded RNA library from a tissue sample, the method comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, wherein the RNA capture probes comprise an RNA capture oligonucleotide complementary to the RNA in the sample and a handle sequence; (b) hybridizing the RNA capture oligonucleotides of the RNA capture probes to the RNA in the tissue sample to form RNA-RNA capture probe hybrids; (c) using a reverse transcriptase to extend the RNA capture oligonucleotides of the RNA-RNA capture probe hybrids to form a plurality of first strand cDNA molecules, each first strand cDNA molecule comprising an RNA capture oligonucleotide and a handle sequence; and (d) hybridizing the first strand cDNA molecules with a reverse transcriptase (RT) and a template switch oligonucleotide (TSO), wherein the RT incorporates a non-template cytosine nucleotide at the 3' end of the first cDNA and the TSO hybridizes to the non-template cytosine nucleotide. and contacting the 3' end of each first strand cDNA molecule with a reverse transcriptase (RT) and a template switch oligonucleotide (TSO) comprising a sequence capable of: (e) releasing the substrate capture probe from the substrate; (f) hybridizing the substrate capture oligonucleotide of the first strand cDNA molecule to the first domain of the substrate capture probe; (g) contacting the first strand with a second strand synthesis mix comprising a TSO primer and extending the TSO primer using the first strand as a template to generate a second strand complementary to the first strand, wherein the second strand comprises TSO, a second cDNA complementary to the first cDNA,and generating a second strand comprising second strand barcode information comprising a spatial barcode sequence complement (SBC') complementary to the spatial barcode sequence (SBC).

[0029] In various embodiments, the first domain is a poly-G sequence that hybridizes with a poly-C sequence on the TSO. In various embodiments, the handle is a P5 sequence and the second handle is a P7 sequence.

[0030] In another aspect, the disclosure provides a method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) capturing the tissue sample with a plurality of RNA capture probes that bind to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide complementary to an RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, the RNA capture oligonucleotides complementary to the RNA being blocked at their 3' ends, each of the substrate capture probes comprising, in a 5' to 3' orientation, a first domain and a first substrate anchor sequence and adjacent to one or more barcoded substrate probes on the substrate, each of the barcoded substrate probes comprising, in a 5' to 3' orientation, a second substrate anchor sequence, a spatial barcode, and a random priming sequence; (b) hybridizing the RNA capture oligonucleotides of the RNA capture probes to RNA in the tissue sample to form RNA-RNA capture probe hybrids having a 5' single-stranded RNA region; (c) hybridizing the substrate capture oligonucleotides of the RNA-RNA capture probe hybrids to a first domain of the substrate capture probe; (d) hybridizing the 5' single-stranded RNA region of the RNA-RNA capture probe hybrid to a random priming sequence of a barcoded substrate probe; and (e) using a reverse transcriptase to extend the random priming sequence hybridized to the 5' single-stranded RNA region to form a plurality of spatially barcoded first strand cDNA molecules.

[0031] In various embodiments, the nucleotide sequence complementary to the RNA in the sample is a poly-T oligonucleotide, a randomer, a semi-randomer, or a target-specific sequence. In various embodiments, the nucleotide sequence complementary to the RNA in the sample is a poly-T oligonucleotide.

[0032] In various embodiments, the method further comprises removing RNA from the sample. In various embodiments, the RNA is removed from the sample after extension to form first strand cDNA. In various embodiments, the RNA is removed by enzymatic or thermal methods.

[0033] A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide complementary to an RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a first domain, a linker, a spatial barcode, and a random priming sequence; and (b) contacting the RNA capture probes with the RNA in the tissue sample. Also provided is a method comprising: (a) hybridizing a substrate capture oligonucleotide of the RNA-RNA capture probe hybrid with a first domain of the substrate capture probe to form an RNA-RNA capture probe hybrid having a 5' single-stranded RNA region; (b) hybridizing a substrate capture oligonucleotide of the RNA-RNA capture probe hybrid with a first domain of the substrate capture probe; (c) hybridizing the 5' single-stranded RNA region of the RNA-RNA capture probe hybrid with a random priming sequence of the substrate capture probe; and (d) using a reverse transcriptase to extend the random priming sequence hybridized to the 5' single-stranded RNA region to form a plurality of spatially barcoded first strand cDNA molecules.

[0034] In various embodiments, the linker is a linker that cannot be read through by a polymerase.

[0035] In another aspect, the disclosure provides a method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that bind to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide sequence complementary to RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the substrate capture probes comprising, in a 5' to 3' orientation, a first domain and a first substrate anchor sequence, and adjacent to at least one of a plurality of barcoded substrate probes on the substrate, each barcoded substrate probe comprising, in a 5' to 3' orientation, a spatial barcode and a second substrate anchor sequence; and (b) contacting the tissue sample with a plurality of RNA capture probes that bind to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide sequence complementary to RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the substrate capture probes comprising, in a 5' to 3' orientation, a first domain and a first substrate anchor sequence, and adjacent to at least one of a plurality of barcoded substrate probes on the substrate, each barcoded substrate probe comprising, in a 5' to 3' orientation, a spatial barcode and a second substrate anchor sequence; (c) hybridizing an RNA capture oligonucleotide of the capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (d) capturing the RNA-RNA capture probe hybrid on a substrate by hybridizing a substrate capture oligonucleotide of the RNA-RNA capture probe hybrid to a first domain of the substrate capture probe; (d) extending the RNA capture oligonucleotide of the captured RNA-capture RNA capture probe hybrid using a reverse transcriptase to form a plurality of first strand cDNA molecules; and (e) ligating each of the first strand cDNA molecules to a proximal barcoded substrate probe, thereby forming spatially barcoded first strand cDNA molecules.

[0036] 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) distributing the tissue sample with a plurality of RNA capture probes that bind to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide complementary to an RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, the RNA capture oligonucleotide complementary to the RNA being blocked at its 3′ end, each of the substrate capture probes comprising, in a 5′ to 3′ orientation, a first domain and a first substrate anchor sequence, and adjacent to at least one of a plurality of barcoded substrate probes on the substrate, each barcoded substrate probe comprising, in a 5′ to 3′ orientation, a poly-T sequence, a spatial barcode, and a second substrate anchor sequence; Further contemplated are methods including: (a) contacting a substrate with a plurality of RNA capture probes, each of which comprises a substrate anchor sequence; (b) hybridizing the RNA capture oligonucleotides of the RNA capture probes to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (c) capturing the RNA-RNA capture probe hybrid on the substrate by hybridizing the substrate capture oligonucleotides of the RNA-RNA capture probe hybrid to a first domain of the substrate capture probe; (d) polyadenylating the RNA in the sample at the 3' end; and (e) using a reverse transcriptase to extend the RNA capture oligonucleotides of the captured RNA-RNA capture probe hybrid to form a plurality of first strand cDNA molecules.

[0037] In various embodiments, polyadenylation is carried out using polyA polymerase.

[0038] A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, each of the RNA capture probes having a hairpin structure and comprising a DNA capture oligonucleotide complementary to RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, the DNA capture oligonucleotide of the RNA capture probe comprising a single-stranded region, and each of the substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, a first domain, and a second domain, the second domain comprising at least one RNA nucleotide or nucleoside; and (b) hybridizing the RNA capture probes to the RNA in the tissue sample to form RNA-RNA capture probe hybrids, the RNA-RNA capture probe hybrids comprising at least one RNA nucleotide or nucleoside. Also provided is a method including: (a) forming RNA-RNA capture probe hybrids, each of the capture probe hybrids comprising a 5' single-stranded RNA terminal region; (c) capturing the substrate capture oligonucleotides of the RNA-RNA capture probe hybrids on a substrate by hybridizing the substrate capture oligonucleotides of the RNA-RNA capture probe hybrids with a first domain of the substrate capture probe; (d) phosphorylating the 5' single-stranded RNA terminal regions of the captured RNA-RNA capture probe hybrids and contacting the captured RNA-RNA capture probe hybrids with a 5' to 3' riboexonuclease to digest the phosphorylated 5' single-stranded RNA terminal regions; and (e) ligating the digested 5' RNA terminal regions of the captured RNA-RNA capture probe hybrids to the second domain of the substrate capture probe to form a plurality of DNA-RNA chimeras on the substrate.

[0039] In various embodiments, the ligation is performed using T4 ligase.

[0040] In various embodiments, the RNA of the captured RNA-RNA capture probe hybrid is 5' phosphorylated prior to ligation.

[0041] In various embodiments, the method further includes generating first strand cDNA from the plurality of DNA-RNA chimeras on the substrate. In various embodiments, the first strand cDNA can be hybridized from the surface and processed for sequencing.

[0042] In various embodiments, reverse transcription is performed using DNA random primers, optionally including P5 adapters.

[0043] In various embodiments, the cDNA extension template can be dehybridized from the RNA in the tissue by chemical, enzymatic, or thermal dehybridization. In various embodiments, the cDNA extension template can be dehybridized from the RNA on the substrate by chemical, enzymatic, or thermal dehybridization. In various embodiments, the dehybridization step occurs before or after the capture step.

[0044] In various embodiments, the RNA capture probe is selected from the group consisting of a poly-T sequence, a poly-U sequence, a randomer, a semi-random sequence, or a target-specific probe. In various embodiments, the RNA capture probe is a poly-T sequence.

[0045] In various embodiments, the RNA capture probe comprises at least 10 deoxythymidine residues. In various embodiments, the RNA capture probe comprises a plurality of different target-specific RNA capture probe sequences. In various embodiments, the RNA capture probe comprises at least 10 nucleotides complementary to the nucleotide sequence of the target RNA. In various embodiments, the RNA capture probe or surface capture probe is 8-80 nucleotides. In various embodiments, the RNA capture probe is 8-80 nucleotides or 10-50 nucleotides.

[0046] In various embodiments, the tissue sample is permeabilized before contacting the tissue sample with multiple RNA capture probes.In various embodiments, the tissue sample is treated with one or more blocking reagents before contacting the tissue sample with multiple RNA capture probes.In various embodiments, the tissue sample is permeabilized and treated with one or more blocking reagents before contacting the tissue sample with multiple RNA capture probes.

[0047] In various embodiments, the substrate is a bead, a bead array, a spot array, a substrate containing multiple wells, a flow cell, clustered particles disposed on the surface of a chip, a film, or a plate. In various embodiments, the substrate comprises multiple nanowells or microwells.

[0048] In various embodiments, the tissue sample is a fresh tissue sample, a fresh frozen tissue sample, or a formalin-fixed, paraffin-embedded (FFPE) tissue sample. In various embodiments, if the sample is an FFPE sample, the method can further include decrosslinking the FFPE sample, optionally using TE buffer, pH 9.

[0049] In various embodiments, the method further includes determining the spatial location of one or more of the spatially barcoded first strand cDNA molecules or copies thereof by correlating the spatial barcode sequence of the spatially barcoded first strand cDNA molecule or copies thereof with the spatial location of surface oligonucleotide molecules on the substrate containing the corresponding spatial barcode sequence.

[0050] In various embodiments, the method further comprises recovering the spatially barcoded first strand cDNA molecules and amplifying them to create a cDNA library.

[0051] In various embodiments, the spatially barcoded first strand cDNA molecules are recovered by contacting the spatially barcoded first strand cDNA on the substrate with a DNA polymerase and one or more primers to generate spatially barcoded second strand cDNA that is complementary to the spatially barcoded first strand cDNA, and removing the spatially barcoded second strand cDNA from the substrate.

[0052] In various embodiments, the one or more primers each comprise a random priming sequence, hi various embodiments, the random priming sequence comprises nine random nucleotides.

[0053] In various embodiments, each spatially barcoded second strand cDNA comprises a unique molecular identifier (UMI), wherein the UMI comprises an endogenous sequence and an exogenous sequence, wherein the exogenous sequence is a sequence complementary to a random priming sequence used to generate the second strand cDNA, and the endogenous sequence is a sequence complementary to a first strand cDNA template sequence used to generate the second strand cDNA.

[0054] In various embodiments, the one or more primers each comprise a molecular identifier barcode. In various embodiments, the one or more primers each comprise a UMI barcode.

[0055] In various embodiments, the spatially barcoded second strand cDNA is removed from the substrate by chemical or physical dehybridization.

[0056] In various embodiments, the anchor sequence comprises a cleavage site, and the spatially barcoded first and second strand cDNA hybrids are removed from the substrate by enzymatic cleavage at the cleavage site. In various embodiments, the cleavage site is a binding site for a restriction endonuclease. In various embodiments, the anchor sequence comprises a cleavage site, and the spatially barcoded first strand cDNA molecules are recovered by enzymatic cleavage at the cleavage site. In various embodiments, the cleavage site is a binding site for a restriction endonuclease.

[0057] In various embodiments, the method further comprises sequencing at least a portion of the cDNA library to determine the spatial barcode sequence of each molecule.

[0058] In various embodiments, the method further includes determining the spatial location of one or more cDNA molecules by correlating the spatial barcode sequences of the one or more cDNA molecules with the spatial location of surface oligonucleotide molecules on the substrate that contain corresponding spatial barcode sequences.

[0059] In various embodiments, the method further includes indexing and sequencing the spatially barcoded first strand cDNA, the method comprising: performing an extension reaction and PCR on the spatially barcoded first strand cDNA to obtain a PCR template comprising a first strand PCR product representing one or more RNA transcripts in the tissue sample; eluting the PCR template; and performing indexing PCR to generate a double-stranded PCR product comprising the first strand PCR product and a second strand complementary to the first strand PCR product. Includes:

[0060] In various embodiments, the method further comprises sequencing the PCR product and determining the location of the RNA transcript in the tissue based on the spatial barcode of the first strand cDNA.

[0061] In various embodiments, the double-stranded PCR product comprises a second clustered sequence on the second strand that is complementary to the PCR product of the first strand, and optionally an index sequence.

[0062] In various embodiments, the PCR products are further processed by tagging to create a spatial transcriptomics library. In various embodiments, tagging comprises tagging on a substrate. In some embodiments, tagging comprises tagging on beads, where the beads comprise a plurality of bead-linked transposomes (BLTs). In some embodiments, the BLTs comprise a plurality of oligonucleotides, including i) a plurality of oligonucleotides comprising a first clustered sequence (P7), a first index sequence, and a Read 1 sequencing primer (Rd1 SP), and ii) a plurality of oligonucleotides comprising a second clustered sequence (P5), a second index sequence, and a Read 2 sequencing primer (Rd2 SP).

[0063] In various embodiments, the RNA library is an mRNA library.

[0064] In various embodiments, the method uses a tissue sample to determine RNA expression in a single cell. In various embodiments, the method determines RNA expression in one or more subcellular components in the single cell. In various embodiments, the subcellular component is a cell nucleus, cytoplasm, or mitochondria.

[0065] In various embodiments, the substrate or substrate surface comprises a material selected from glass, silicon, poly-L-lysine coated material, nitrocellulose, polystyrene, cyclic olefin copolymer (COC), cyclic olefin polymer (COP), polyacrylamide, polypropylene, polyethylene, or polycarbonate.

[0066] Also provided is a method for identifying genetic variations in a subject having or at risk of having a disease, comprising generating a sample RNA library, e.g., an mRNA library, from a tissue sample from the subject according to the methods described herein, comparing the genetic information from the sample RNA library, e.g., the mRNA library, with a control RNA library, e.g., the mRNA library, or a sample from the subject before the disease, and identifying genetic variations in the sample RNA library, e.g., the mRNA library, that are associated with the disease. Optionally, the method includes treating the subject with a disease-specific therapy.

[0067] In various embodiments, the disease is a genetic defect, cancer, an autoimmune disease, a metabolic disorder, or any other disease described herein. Additional diseases or conditions are described in more detail in the detailed description.

[0068] It is understood that each feature or embodiment, or combination, described herein is a non-limiting illustrative example of one of the aspects of the invention and is therefore meant to be combinable with any other feature, embodiment, or combination described herein. For example, when a feature is described with words such as "one embodiment," "various embodiments," "some embodiments," "an embodiment," "further embodiment," "particular exemplary embodiment," and / or "another embodiment," each of these types of embodiments is a non-limiting example of the feature that is intended to be combined with any other feature or combination of features described herein, without necessarily listing every possible combination.

[0069] Any such feature or combination of features applies to any of the aspects of the invention. When example values ​​falling within ranges are disclosed, any of these examples are contemplated as possible endpoints of the range, and any and all values ​​between such endpoints are contemplated, with any and all combinations of upper and lower limits envisioned. [Brief explanation of the drawings]

[0070] [Figure 1] FIG. 1 is a schematic diagram of a method for capturing mRNA transcripts from tissue samples in situ using capture probes. [Figure 2] FIG. 1 is a schematic diagram of a method for capturing mRNA transcripts in situ from a tissue sample using capture probes, where the capture probes hybridized to the transcripts create a nucleotide gap between the hybridized sequences. [Figure 3] Schematic diagram of an exemplary RNA library preparation workflow described herein. [Figure 4A] Schematic diagram of alternative RNA library preparation workflows described herein. Figure 4A shows the general workflow, but methods for adding 3' oligonucleotides include oNTP-directed adapter processing or click chemistry (Figure 4B), template switching (Figure 4C), or template switching in which the template-switching primer is released from the substrate and spatially barcoded (Figure 4D). [Figure 4B] Schematic diagram of alternative RNA library preparation workflows described herein. Figure 4A shows the general workflow, but methods for adding 3' oligonucleotides include oNTP-directed adapter processing or click chemistry (Figure 4B), template switching (Figure 4C), or template switching in which the template-switching primer is released from the substrate and spatially barcoded (Figure 4D). [Figure 4C] Schematic diagram of alternative RNA library preparation workflows described herein. Figure 4A shows the general workflow, but methods for adding 3' oligonucleotides include oNTP-directed adapter processing or click chemistry (Figure 4B), template switching (Figure 4C), or template switching in which the template-switching primer is released from the substrate and spatially barcoded (Figure 4D). [Figure 4D]Schematic diagram of alternative RNA library preparation workflows described herein. Figure 4A shows the general workflow, but methods for adding 3' oligonucleotides include oNTP-directed adapter processing or click chemistry (Figure 4B), template switching (Figure 4C), or template switching in which the template-switching primer is released from the substrate and spatially barcoded (Figure 4D). [Figure 5A] Schematic of an alternative RNA library preparation workflow described herein, using 3' blocked oligonucleotides on the target probe. [Figure 5B] Schematic of an alternative RNA library preparation workflow described herein, using 3' blocked oligonucleotides on the target probe. [Figure 6A] Schematic of an alternative RNA library preparation workflow described herein. [Figure 6B] Schematic of an alternative RNA library preparation workflow described herein. [Figure 7] Schematic of an alternative RNA library preparation workflow described herein using hairpin probes. [Figure 8] The workflow based on the scheme in Figure 3 is shown. [Figure 9] A workflow based on the scheme of Figures 4A and 4B is shown. [Figure 10] The workflow is based on the scheme in Figure 4C. [Figure 11] A workflow based on the scheme in Figure 4D is shown. [Figure 12] A workflow based on the scheme in Figure 5A is shown. [Figure 13] A workflow based on the scheme in Figure 5B is shown. [Figure 14] A workflow based on the scheme in Figure 6A is shown. [Figure 15] A workflow based on the scheme in Figure 6B is shown. [Figure 16]The workflow based on the scheme in Figure 7 is shown. DETAILED DESCRIPTION OF THE INVENTION

[0071] To overcome the technical limitations of isolating mRNA transcripts from fresh-frozen or FFPE tissue samples, an in situ method is described herein to capture and create spatially barcoded libraries from such compromised tissue mRNA.

[0072] Various methods and compositions are described herein that allow for characterization of genetic profiles in tissues while preserving spatial information related to the origin of target genes or polynucleotides in the tissue. In various embodiments, the methods include a substrate on which multiple capture probes are immobilized, such that each capture probe occupies a different position on the array. Each capture probe contains, among other sequences and / or molecules, a unique positional nucleic acid tag (i.e., a spatial address or index sequence). Each spatial address corresponds to the position of the capture probe on the array. The position of the capture probe on the array can be correlated with a position in the tissue sample.

[0073] Examples of genes or polynucleotides in tissue samples include genomic DNA, methylated DNA, differentially methylated DNA sequences, messenger RNA (mRNA), polyA mRNA, fragmented mRNA, fragmented DNA, mitochondrial DNA, ribosomal RNA (rRNA), viral RNA, microRNA, in situ synthesized PCR products, and RNA / DNA hybrids. Non-coding RNA (ncRNA), small nucleolar RNA (snoRNA), and / or small nuclear RNA (snRNA) are also contemplated.

[0074] The nucleic acid tag encoding the location (i.e., spatial address or index sequence) can be attached to a nucleic acid capture region or any other molecule that binds to a target gene or polynucleotide. Examples of other molecules that can be attached to a nucleic acid tag include antibodies, antigen-binding domains, proteins, peptides, receptors, haptens, etc.

[0075] Various methods and compositions are described herein that allow for characterization of transcriptome and / or genomic mutations in tissue while preserving spatial information related to the origin of target nucleic acids in tissue.For example, the methods disclosed herein can allow for identifying the location of cells or cell clusters in tissue biopsies that carry abnormal mutations.Therefore, the methods provided herein can be useful for diagnostic purposes, such as diagnosing cancer, and possibly aid in the selection of targeted therapy.

[0076] The present disclosure is based in part on the realization that information related to the spatial origin of nucleic acids in a tissue sample can be encoded in the nucleic acid during the process of preparing the nucleic acid for sequencing. For example, nucleic acids from a tissue sample can be tagged with probes containing position-specific sequence information ("spatial addresses"). Spatially addressed nucleic acid molecules from the tissue sample can then be sequenced in bulk. Sequence-identical nucleic acid molecules from different regions in the tissue sample can be distinguished based on their spatial addresses and mapped to their regions of origin in the tissue sample. Additionally, spatial addressing of nucleic acids can increase the sensitivity of detecting single nucleotide variations (SNVs) or single nucleotide polymorphisms (SNPs) in a tissue sample.

[0077] In some methods described herein, probes for spatial tagging include, for example, a combination of a spatial address region and a gene-specific capture region. The spatially addressed gene-specific probes can be contacted with a tissue sample as immobilized probes on a capture array.

[0078] The present disclosure recognizes that spatial addressing of nucleic acids from tissue samples can involve two-dimensional spatial addressing, for example, to correlate the location of a nucleic acid on a two-dimensional capture array with the location of the nucleic acid in a two-dimensional tissue section. Spatial addressing can also be performed in additional dimensions. For example, a spatial address sequence can be added to a nucleic acid to describe the relative spatial location of the nucleic acid in a third or fourth dimension, for example, by describing the location of a tissue section in a tissue biopsy or the location of a tissue biopsy in an organ of interest. Temporal address sequences can be added to nucleic acids from tissue samples to indicate time points in time-course experiments, for example, experiments examining changes in gene expression in cells in response to physical or chemical stimuli, such as drug treatment in clinical trials.

[0079] definition Unless otherwise stated, the following terms used in this Application, including the specification and claims, have the definitions given below.

[0080] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to a "capture probe" includes mixtures of two or more capture probes, and so on.

[0081] The term "about," particularly with respect to a given quantity, is meant to encompass a deviation of plus or minus 5 percent.

[0082] As used herein, the terms "include," "including," "includes," "including," "contain," "containing," and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, product-by-process, or composition of matter that includes, includes, or contains an element or list of elements not only includes those elements, but may also include other elements not expressly listed in or inherent to such process, method, product-by-process, or composition of matter.

[0083] As used herein, "anchor" refers to a moiety that attaches a nanoscaffold to a substrate. Anchors include chemical moieties, peptides, or oligonucleotides. Polynucleotide anchors can be 4 to 20 nucleotides.

[0084] As used herein, a "splint oligonucleotide" refers to an oligonucleotide that includes a sequence complementary to a region on a surface probe on a nanostructure and another sequence complementary to a capture oligonucleotide (e.g., attached to a substrate). In various embodiments, the splint oligonucleotide is 10-25 nucleotides or 15-25 nucleotides. In various embodiments, the splint oligonucleotide is 20 nucleotides. In various embodiments, the splint oligonucleotide is 15, 16, 17, 18, 9, 20, 21, 22, 23, 24, or 25 nucleotides.

[0085] As used herein, "surface oligonucleotide" refers to an oligonucleotide that includes an anchor sequence for attaching the oligo to the surface of a substrate, a spatial barcode sequence, and a sequence that hybridizes with a splint oligonucleotide. In various embodiments, the surface oligonucleotide is 15-25 nucleotides. In various embodiments, the surface oligonucleotide is greater than 20 nucleotides. In various embodiments, the surface oligonucleotide is 15, 16, 17, 18, 9, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 or more nucleotides.

[0086] As used herein, the terms "address," "tag," or "index," when used in reference to a nucleotide sequence, are intended to mean a unique nucleotide sequence that is distinct from other indices and from other nucleotide sequences within polynucleotides contained within a sample. A nucleotide "address," "tag," or "index" can be a random or specifically designed nucleotide sequence. An "address," "tag," or "index" can be of any desired sequence length, so long as it is long enough to be a unique nucleotide sequence within multiple indices and / or multiple polynucleotides in the population being analyzed or interrogated. The nucleotide "addresses," "tags," or "indexes" of the present disclosure are useful, for example, for attaching to target polynucleotides to tag or mark specific species to identify all members of the tagged species within a population. Thus, indexes are useful as barcodes, where different members of the same molecular species can contain the same index, and different species within different polynucleotide populations can have different indexes.

[0087] As used herein, the terms "address," "tag," "index," or "barcode," when used in reference to a nucleotide sequence, are intended to mean a unique nucleotide sequence that is distinct from other indexes and from other nucleotide sequences within polynucleotides contained within a sample. A nucleotide "address," "tag," "index," or "barcode" can be a random or specifically designed nucleotide sequence. An "address," "tag," "index," or "barcode" can be of any desired sequence length, so long as it is long enough to be a unique nucleotide sequence within multiple indexes and / or multiple polynucleotides in the population being analyzed or interrogated. The nucleotide "addresses," "tags," "indexes," or "barcodes" of the present disclosure are useful, for example, for attaching to target polynucleotides to tag or mark specific species to identify all members of the tagged species within a population. Thus, an index is useful as a barcode, where different members of the same molecular species can contain the same index, and different species within different polynucleotide populations can have different indexes.

[0088] A tag / index / barcode sequence may be unique to a single nucleic acid species in a population, or may be shared by several different nucleic acid species in the population. For example, each nucleic acid probe in a population may contain a tag / index / barcode sequence that is different from all other nucleic acid probes in the population. Alternatively, each nucleic acid probe in a population may contain a tag / index / barcode sequence that is different from several or most other nucleic acid capture probes in the population. For example, each probe in a population may have a tag / index / barcode that is present in several different capture probes in the population, even if probes with a common tag / index / barcode differ from each other in other sequence regions along their length. In certain embodiments, one or more tag / index / barcode sequences used with a biological specimen are not present in the genome, transcriptome, or other nucleic acids of the biological specimen. For example, a tag / index / barcode sequence may have less than 80%, 70%, 60%, 50%, or 40% sequence identity to a nucleic acid sequence in a particular biological specimen.

[0089] As used herein, "spatial address," "spatial tag," "spatial barcode," "barcode sequence," or "spatial index," when used in reference to a nucleotide sequence, means an address, tag, barcode, or index that encodes spatial information related to the region or location of origin of the addressed, tagged, barcoded, or indexed nucleic acid in a tissue sample. The sequence can be a naturally occurring sequence or a sequence that does not naturally occur in the organism from which the barcoded nucleic acid is obtained.

[0090] As used herein, the term "substrate" is intended to mean a solid support or support structure. This term includes any material that can serve as a solid or semi-solid base for generating features such as wells for the deposition of biopolymers, including nucleic acids, polypeptides, and / or other polymers. Non-limiting examples of substrates include bead arrays, spot arrays, clustered particles arranged on the surface of a chip, films, multiwell plates, and flow cells. The substrates provided herein can be modified, for example, or modified to accommodate the attachment of biopolymers by various methods well known to those of skill in the art. Exemplary types of substrate materials include glass, modified glass, functionalized glass, inorganic glass, microspheres containing inert and / or magnetic particles, plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, optical fibers or fiber optic bundles, various polymers other than those exemplified above, and multiwell microtiter plates. Specific types of exemplary plastics include acrylic, polystyrene, copolymers of styrene with other materials, polypropylene, polyethylene, polybutylene, polyurethane, and TEFLON™. Specific types of exemplary silica-based materials include silicon and various forms of modified silicon.

[0091] Those skilled in the art will know or understand that the composition and shape of the substrates provided herein can vary depending on the intended use and user preference. Thus, while planar substrates such as slides, chips, wafers, or beads are useful for microarrays, those skilled in the art will understand that a wide variety of other substrates exemplified herein or known in the art can also be used in the methods and / or compositions herein.

[0092] In some embodiments, the solid support comprises one or more surfaces accessible to reagents, beads, or analytes. The surface may be substantially flat or planar. Alternatively, the surface may be rounded or contoured. Exemplary contours that may be included on the surface include wells (e.g., microwells or nanowells), depressions, posts, ridges, channels, and the like. Examples of materials that can be used as surfaces include glass, such as modified or functionalized glass; plastics, such as acrylic, polystyrene, or copolymers of styrene with another material, polypropylene, polyethylene, polybutylene, polyurethane, or TEFLON™; polysaccharides or cross-linked polysaccharides, such as agarose or Sepharose; nylon; nitrocellulose; resins; silica or silica-based materials, including silicon and modified silicon, carbon fiber; metals; inorganic glass; fiber optic bundles, or various other polymers. A single material or a mixture of several different materials can form a surface useful in the present invention. In some examples, the surface comprises a well (e.g., a microwell or nanowell). In some embodiments, the surface comprises wells of an array of wells (e.g., microwells or nanowells) on a glass, silicon, plastic, or other suitable solid support comprising a patterned, covalently linked gel such as poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication No. 2014 / 0079923 A1, incorporated herein by reference). In some examples, the support structure can comprise one or more layers.

[0093] Non-limiting examples of surfaces include bead arrays, spot arrays, clustered particles arranged on the surface of a chip, films, multi-well plates, and flow cells.

[0094] In some embodiments, the solid support comprises one or more surfaces of a flow cell. As used herein, the term "flow cell" refers to a chamber containing a solid surface through which one or more fluidic reagents can be passed. The flow cell can be an ordered or random flow cell. Examples of flow cells and associated fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, U.S. Patent No. 7,057,026, WO 91 / 06678, WO 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.

[0095] In some embodiments, the solid support comprises a patterned surface. A "patterned surface" refers to an arrangement of distinct regions within or on an exposed layer of a solid support. For example, one or more of the regions can be features in which one or more amplification primers are present. The features can be separated by interstitial regions in which amplification primers are absent. In some embodiments, the pattern can be an xy format of features in rows and columns. In some embodiments, the pattern can be a repeating arrangement of features and / or interstitial regions. In some embodiments, the pattern can be a random arrangement of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Patent Application No. 13 / 661,524 or U.S. Patent Application Publication No. 2012 / 0316086 or International Patent Publication No. 2017 / 019456, each of which is incorporated herein by reference.

[0096] As used herein, the term "immobilized," when used with respect to nucleic acids, is intended to mean direct or indirect attachment to a solid support via covalent or non-covalent bonds. "Immobilized" refers to the state in which two things are joined, fastened, adhered, attached, connected, or bound to one another. For example, an analyte, e.g., a nucleic acid, may be immobilized on a material, e.g., a bead, gel, or surface, by covalent or non-covalent attachment. In certain embodiments, covalent attachment may be used, but all that is required is that the nucleic acid remain immobilized or attached to the support under conditions under which the support is intended to be used, e.g., in applications requiring nucleic acid amplification and / or sequencing. Oligonucleotides used as capture primers or amplification primers can be immobilized so that the 3' end is available for enzymatic extension and at least a portion of the sequence is capable of hybridizing to a complementary sequence.

[0097] Immobilization can occur via hybridization to surface-attached oligonucleotides, in which case the immobilized oligonucleotide or polynucleotide can be in a 3' to 5' orientation. Alternatively, immobilization can occur by means other than base-pairing hybridization, such as covalent attachment as described above.

[0098] Exemplary covalent linkages include, for example, those resulting from the use of click chemistry techniques. Exemplary non-covalent linkages include, but are not limited to, non-specific interactions (e.g., hydrogen bonds, ionic bonds, van der Waals interactions, etc.) or specific interactions (e.g., affinity interactions, receptor-ligand interactions, antibody-epitope interactions, avidin-biotin interactions, streptavidin-biotin interactions, lectin-carbohydrate interactions, etc.). Exemplary linkages are described in U.S. Patent Nos. 6,737,236, 7,259,258, 7,375,234, and 7,427,678, and U.S. Patent Publication No. 2011 / 0059865(A1), each of which is incorporated herein by reference.

[0099] As used herein, the term "array" refers to a collection of sites that can be distinguished from one another according to their relative positions. Different molecules at different sites of an array can be distinguished from one another according to the site's position within the array. Each site of an array can contain one or more molecules of a particular type. For example, a site can contain a single target nucleic acid molecule having a particular sequence, or a site can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). The sites of an array can be different features disposed on the same substrate. Exemplary features include, but are not limited to, wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, or channels within a substrate. The sites of an array can be separate substrates, each with a different molecule. The different molecules attached to the separate substrates can be identified according to the position of the substrate on a surface to which the substrates are associated, or according to the position of the substrate within a liquid or gel. An exemplary array in which separate substrates are disposed on a surface includes, but is not limited to, beads in wells.

[0100] As used herein, the term "single molecular identifier" or "SMI" refers to a molecular tag that can be attached to a nucleic acid, either randomly, non-randomly, or semi-randomly. In various embodiments, the SMI is a unique molecular identifier (UMI). When incorporated into a nucleic acid, the SMI can be used to correct for subsequent amplification bias by directly counting the single molecular identifier (SMI) that is sequenced after amplification. The SMI (e.g., UMI) can be attached to similar nucleic acids, e.g., adapters, making each nucleic acid unique. The SMI (e.g., UMI) may also be used to uniquely tag individual molecules (e.g., individual mRNA molecules) in a sample (e.g., individual mRNA molecules in a tissue sample, cell sample, or sample library).

[0101] As used herein, "unique molecular index," "unique molecular identifier," or "UMI," when used in reference to a capture probe or other nucleic acid, is intended to refer to a portion of the probe that is useful as a molecular barcode for uniquely tagging each molecule in a sample library. A UMI may be depicted as "NNNN..." in a string of nucleic acid to designate that portion of the oligonucleotide as a UMI. A UMI may be 6-20 nucleotides or longer in length. In some embodiments, a UMI comprises a spatial barcode.

[0102] As used herein, the term "universal sequence" refers to a sequence of nucleotides common to two or more nucleic acid molecules, even if the molecules also have regions of different sequence. A universal sequence present in different members of a collection of molecules allows for the capture of multiple different nucleic acids using a population of universal capture nucleic acids that are complementary to the universal sequence. Similarly, a universal sequence present in different members of a collection of molecules allows for the replication or amplification of multiple different nucleic acids using a population of universal primers that are complementary to the universal sequence. Thus, a universal capture nucleic acid or universal primer comprises a sequence that can specifically hybridize to a universal sequence. Target nucleic acid molecules may be modified, for example, to attach universal adaptors to one or both ends of different target sequences. Universal capture oligonucleotides are applicable to interrogating multiple different oligonucleotides without necessarily distinguishing between different species, while target-specific capture sequences are applicable to distinguish between different species. A non-limiting example of a universal sequence is a poly-T nucleotide sequence.

[0103] As used herein, a "semi-random" nucleotide sequence comprises or consists of a partially predetermined nucleotide sequence combined with random nucleotide sequences.

[0104] As used herein, the term "adapter" generally refers to any linear nucleic acid molecule that can be added (e.g., by synthesis or ligation) to an oligonucleotide of the present disclosure. In some embodiments, the adapter is copied onto a library molecule using templated polymerase synthesis (e.g., second-strand cDNA synthesis as described herein). In some embodiments, the adapter is ligated to a first complementary strand of the present disclosure. In some embodiments, the oligonucleotide of the present disclosure comprises an adapter ("adapter oligonucleotide"). In some embodiments, the adapter oligonucleotide comprises, from 5' to 3', a third sequencing primer sequence (e.g., SBS3), a sequence complementary to a unique index sequence (e.g., i5'), and a second clustered primer sequence (e.g., P5). In some embodiments, the adapter comprises a sequence complementary to a primer. In further embodiments, the adapter comprises a sequence complementary to a P5 primer or a P5' primer. In some embodiments, the adapter comprises a sequence complementary to a P7 primer or a P7' primer. In some embodiments, the adaptor comprises a sequence complementary to the B15 primer or the B15' primer.

[0105] The terms "P5," "P7," "B15," "P5'" (P5 prime), "P7'" (P7 prime), "B15'" (B15 prime), "P15," and "P17" may be used when referring to examples of oligonucleotide sequences of primers, e.g., clustered primers, and / or oligonucleotide sequences complementary to primers. The terms "P5'" (P5 prime), "P7'" (P7 prime), and "B15'" (B15 prime) refer to the complements of P5, P7, and B15, respectively. Any suitable primers can be used in the methods presented herein, and it will be understood that the use of P5, P5', P7, P7', P15, P17, B15, and B15' are exemplary embodiments only. The use of primers such as P5, P5', P7, P7', P15, P17, B15, and B15' or their complements on a flow cell is known in the art, as exemplified by the disclosures of WO 2019 / 222264, WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957, each of which is incorporated by reference in its entirety. For example, any suitable forward amplification primer, whether immobilized or in solution, can be useful in the methods presented herein for hybridization to and amplification of complementary sequences. Similarly, any suitable reverse amplification primer, whether immobilized or in solution, can be useful in the methods provided herein for hybridization to and amplification of complementary sequences and sequences. Those skilled in the art will understand how to design and use suitable primer sequences for nucleic acid capture and / or amplification as provided herein. In some embodiments, the "first clustered primer" described herein is a P5 primer. In some embodiments, the "first clustered primer" described herein is a P7 primer. In some embodiments, the "first clustered primer" described herein is a P5' primer.In some embodiments, a "first clustered primer" described herein is a P7' primer. In some embodiments, a "second clustered primer" described herein is a P5 primer. In some embodiments, a "second clustered primer" described herein is a P7 primer. In some embodiments, a "second clustered primer" described herein is a P5' primer. In some embodiments, a "second clustered primer" described herein is a P7' primer. In some embodiments, P5 comprises or consists of the polynucleotide sequence 5' AAT GAT ACG GCG ACC ACC GA 3' (SEQ ID NO: 1), or a variant thereof. In some embodiments, P5 comprises or consists of the polynucleotide sequence 5' AAT GAT ACG GCG ACC ACC GAG ATC TAC AC 3' (SEQ ID NO: 2), or a variant thereof. In some embodiments, P7 comprises or consists of the polynucleotide sequence 5' CAA GCA GAA GAC GGC ATA CG 3' (SEQ ID NO:3), or a variant thereof. In some embodiments, P7 comprises or consists of the polynucleotide sequence 5' CAA GCA GAA GAC GGC ATA CGA GAT 3' (SEQ ID NO:4), or a variant thereof. In some embodiments, P5' comprises or consists of the polynucleotide sequence 5' TCG GTG GTC GCC GTA TCA TT 3' (SEQ ID NO:5), or a variant thereof. In some embodiments, P5' comprises or consists of the polynucleotide sequence 5' GTG TAG ATC TCG GTG GTC GCC GTA TCA TT 3' (SEQ ID NO:6), or a variant thereof. In some embodiments, P7' comprises the polynucleotide sequence 5' CGT ATG CCG TCT TCT GCT TG 3' (SEQ ID NO:7), or a variant thereof. In some embodiments, P7' comprises or consists of the polynucleotide sequence 5' ATC TCG TAT GCC GTC TTC TGC TTG 3' (SEQ ID NO: 8), or a variant thereof.In some embodiments, B15 comprises or consists of the polynucleotide sequence 5' GTCTCGTGGGCTCGG 3' (SEQ ID NO: 9), or a variant thereof. In some embodiments, B15' comprises or consists of the polynucleotide sequence 5' CCGAGCCCACGAGAC 3' (SEQ ID NO: 10), or a variant thereof. In some embodiments, P15 comprises or consists of the polynucleotide sequence 5' TTTTTTAATG ATACGGCGAC CACCGAGANC TACAC 3' (SEQ ID NO: 11), or a variant thereof. In some embodiments, P17 comprises or consists of the polynucleotide sequence 5' TTTTTTNNNC AAGCAGAAGA CGGCATACGA GAT 3' (SEQ ID NO: 12), or a variant thereof. The term "variant," as used herein with respect to any of the sequences recited herein, refers to a variant nucleic acid that is substantially identical, i.e., has only some nucleotide sequence variations relative to the non-variant sequence, for example. In some embodiments, variants have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% overall nucleotide sequence identity to the non-variant nucleic acid sequence. It is understood that references to P5 and P7 herein may refer to different primer sequences. Any suitable combination of primer sequences is encompassed by the present disclosure.

[0106] As used herein, the term "plurality" is intended to mean a population of two or more distinct members. Pluralities can range in size from small, medium, large, to very large. A small-sized plurality can range, for example, from a few members to tens of members. A medium-sized plurality can range, for example, from tens of members to about 100 or hundreds of members. A large plurality can range, for example, from about hundreds of members to about 1,000 members, thousands of members, and tens of thousands of members. A very large plurality can range, for example, from tens of thousands of members to about hundreds of thousands, millions, tens of millions, or hundreds of millions or more members. Thus, pluralities can range in size from 2 to 100 million or more, as well as all sizes measured by the number of members between and above the exemplary ranges above. An exemplary number of features in a microarray is 1.28 cm. 2 An exemplary plurality of nucleic acids may comprise, for example, about 1 x 10 5 , 5×10 5 and 1 x 10 6 or more distinct nucleic acid species. Thus, the definition of this term is intended to include all integer values ​​greater than 2. The upper limit of the plurality value can be set, for example, by the theoretical diversity of nucleotide sequences in a nucleic acid sample.

[0107] As used herein, the term "nucleic acid" is intended to be consistent with its use in the art and includes naturally occurring nucleic acids or functional analogs thereof. Particularly useful functional analogs are capable of hybridizing to nucleic acids in a sequence-specific manner or can be used as templates for replicating specific nucleotide sequences. Naturally occurring nucleic acids generally have backbones containing phosphodiester bonds. Analog structures can have alternative backbone linkages, including any of a variety known in the art. Naturally occurring nucleic acids generally have deoxyribose sugars (e.g., found in deoxyribonucleic acid (DNA)) or ribose sugars (e.g., found in ribonucleic acid (RNA)). Nucleic acids can contain any of a variety of analogs of these sugar moieties known in the art. Nucleic acids can include natural or unnatural bases. In this regard, natural deoxyribonucleic acids can have one or more bases selected from the group consisting of adenine, thymine, cytosine, or guanine, and ribonucleic acids can have one or more bases selected from the group consisting of uracil, adenine, cytosine, or guanine. Useful unnatural bases that can be included in nucleic acids are known in the art. The term "target," when used with respect to nucleic acids, is intended as a semantic identifier of the nucleic acid in the context of the methods or compositions described herein and does not necessarily limit the structure or function of the nucleic acid beyond what is otherwise expressly indicated. Specific forms of nucleic acids can include all types of nucleic acids found in living organisms, as well as synthetic nucleic acids, such as polynucleotides produced by chemical synthesis.

[0108] Specific examples of nucleic acids applicable for analysis by incorporation into microarrays produced by the methods provided herein include genomic DNA (gDNA), expressed sequence tags (ESTs), DNA copies (cDNA), RNA copies (cRNA), mitochondrial DNA or genomic RNA, messenger RNA (mRNA), ribosomal RNA (rRNA), and / or other RNA populations. Additional contemplated RNAs include microRNA, transfer RNA, non-coding RNA (ncRNA), small nucleolar RNA (snoRNA), and / or small nuclear RNA (snRNA). Fragments and / or portions of these exemplary nucleic acids are also included within the meaning of the term as used herein.

[0109] As used herein, the term "double-stranded," when used in reference to a nucleic acid molecule, means that substantially all of the nucleotides in the nucleic acid molecule are hydrogen bonded to complementary nucleotides. A partially double-stranded nucleic acid can have at least 10%, 25%, 50%, 60%, 70%, 80%, 90%, or 95% of its nucleotides hydrogen bonded to complementary nucleotides.

[0110] As used herein, the term "single-stranded," when used in reference to a nucleic acid molecule, means that essentially none of the nucleotides in the nucleic acid molecule are hydrogen bonded to a complementary nucleotide.

[0111] As used herein, the term "capture primer" or "capture probe" is intended to mean an oligonucleotide having a nucleotide sequence capable of specifically annealing to a single-stranded polynucleotide sequence being analyzed or subjected to nucleic acid interrogation under conditions encountered, for example, in the primer annealing step of an amplification or sequencing reaction. The terms "nucleic acid," "polynucleotide," and "oligonucleotide" are used interchangeably herein. The differences in terminology are not intended to indicate any specific differences in size, sequence, or other properties unless otherwise specified. For clarity of explanation, terms may be used to distinguish one species of nucleic acid from another when describing a particular method or composition that includes several nucleic acid species.

[0112] As used herein, the terms "gene-specific" or "target-specific," when used in reference to a capture probe or other nucleic acid, are intended to mean a capture probe or other nucleic acid that includes a nucleotide sequence specific to a targeted nucleic acid (e.g., a nucleic acid from a tissue sample), i.e., a sequence of nucleotides that can selectively anneal to an identified region of the targeted nucleic acid. A gene-specific capture probe may have a single species of oligonucleotide or may include two or more species with different sequences. Thus, a gene-specific capture probe may have two or more sequences, including 3, 4, 5, 6, 7, 8, 9, or 10 or more different sequences. A gene-specific capture probe may include a gene-specific capture primer sequence and a universal capture probe sequence. Other sequences, such as a sequencing primer sequence, may also be included in the gene-specific capture primer.

[0113] As used herein, "unique molecular index," "unique molecular identifier," or "UMI," when used in reference to a capture probe or other nucleic acid, is intended to refer to a portion of the probe that is useful as a molecular barcode for uniquely tagging each molecule in a sample library. A UMI may be depicted as "NNNN..." in a string of nucleic acid to designate that portion of the oligonucleotide as a UMI. A UMI may be 6-20 nucleotides or longer in length. In some embodiments, a UMI comprises a spatial barcode.

[0114] In comparison, the term "universal," when used with respect to a capture probe or other nucleic acid, is intended to mean a capture probe or nucleic acid that has a common nucleotide sequence among multiple capture probes. The common sequence can be, for example, a sequence complementary to the same adapter sequence. A universal capture probe is applicable to interrogate multiple different polynucleotides without necessarily distinguishing between different species, whereas a gene-specific capture primer is applicable to distinguish between different species.

[0115] In various embodiments, capture elements (e.g., capture primers or capture probes or other nucleic acid sequences) can be spaced to: A) spatially resolve nucleic acids within a single-cell geometry, i.e., within multiple capture sites per cell; B) spatially resolve nucleic acids at approximately the single-cell level, i.e., at about one capture site per cell. Additionally, capture elements can be spaced as in A or B above, I) spaced to sample nucleic acids from a sample at regular intervals, e.g., spaced in a grid or pattern such that about every other, fifth, or tenth cell is sampled, or about every other, fifth, or tenth group of 2, 3, 4, 5, 6, 7, 8, 9, 10, or more cells are sampled; II) spaced to capture sample from substantially all available cells in one or more regions of the sample; or III) spaced to capture sample from substantially all available cells in the sample.

[0116] As used herein, the term "amplicon," when used with reference to a nucleic acid, refers to the product of copying a nucleic acid, which product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be generated by any of a variety of amplification methods using a nucleic acid or its amplicon as a template, including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of a nucleotide sequence (e.g., a concatemeric product of RCA). A first amplicon of a target nucleic acid can be a complementary copy. Subsequent amplicons are copies made from the target nucleic acid or the first amplicon after the generation of the first amplicon. Subsequent amplicons can have a sequence that is substantially complementary to or substantially identical to the target nucleic acid.

[0117] The number of template copies or amplicons that can be generated can be modulated by appropriate modification of the amplification reaction, including, for example, varying the number of amplification cycles performed, using polymerases of different processivities in the amplification reaction, and / or varying the length of time the amplification reaction is performed, as well as modifying other conditions known in the art to affect amplification yield. The copy number of the nucleic acid template can be at least 1, 10, 100, 200, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, and 10,000 copies, and can vary depending on the particular application.

[0118] As used herein, the term "complementary" when used with respect to a polynucleotide is intended to mean a polynucleotide comprising a nucleotide sequence that can selectively anneal to an identified region of a target polynucleotide under specific conditions. As used herein, the term "substantially complementary" and grammatical equivalents are intended to mean a polynucleotide comprising a nucleotide sequence that can specifically anneal to an identified region of a target polynucleotide under specific conditions. Annealing refers to the nucleotide base pairing interaction between one nucleic acid and another, resulting in the formation of a duplex, triplex, or other higher-order structure. Primary interactions are typically nucleotide base specific, e.g., A:T, A:U, and G:C, via Watson-Crick and Hoogsteen hydrogen bonding. In certain embodiments, base stacking and hydrophobic interactions may also contribute to the stability of the duplex. Conditions under which a polynucleotide will anneal to a complementary or substantially complementary region of a target nucleic acid are well known in the art, as described, for example, in Nucleic Acid Hybridization, A Practical Approach, Hames and Higgins, eds., IRL Press, Washington, DC (1985) and Wetmur and Davidson, Mol. Biol. 31:349 (1968). Annealing conditions will depend on the particular application and can be routinely determined by one of ordinary skill in the art without undue experimentation.

[0119] As used herein, the term "hybridization" refers to the process by which two single-stranded polynucleotides non-covalently bind to form a stable double-stranded polynucleotide. The resulting double-stranded polynucleotide is a "hybrid" or "double-stranded." Hybridization conditions typically include a salt concentration of less than about 1 M, more usually less than about 500 mM, and can be less than about 200 mM. The hybridization buffer includes a buffered salt solution such as 5% SSPE or other such buffers known in the art. Hybridization temperatures can be as low as 5°C, but are typically greater than 22°C, more typically greater than about 30°C, and typically greater than 37°C. Hybridization is usually performed under stringent conditions, i.e., conditions under which a probe hybridizes to its target subsequence but not to other non-complementary sequences. Stringent conditions are sequence-dependent, vary in different circumstances, and can be routinely determined by one of skill in the art.

[0120] As used herein, the term "dNTP" refers to deoxynucleoside triphosphate. NTP refers to ribonucleotide triphosphate. Purine bases (Pu) include adenine (A), guanine (G), and their derivatives and analogs. Pyrimidine bases (Py) include cytosine (C), thymine (T), uracil (U), and their derivatives and analogs. Examples of such derivatives or analogs include, but are not limited to, those modified with reporter groups, biotinylated, amine-modified, radiolabeled, alkylated, and the like, including phosphorothioates, phosphites, and derivatives modified at ring atoms. Reporter groups can be fluorescent groups such as fluorescein, chemiluminescent groups such as luminol, terbium chelators such as N-(hydroxyethyl)ethylenediaminetriacetic acid, which allow detection by delayed fluorescence, and the like.

[0121] As used herein, the terms "ligation," "ligating," and their grammatical equivalents are intended to mean forming a covalent or covalent linkage between the ends of two or more nucleic acids, e.g., oligonucleotides and / or polynucleotides, typically in a template-driven reaction. The nature of the bond or linkage can vary widely, and ligation can be performed enzymatically or chemically. As used herein, ligation is typically performed enzymatically, forming a phosphodiester bond between the 5' carbon-terminal nucleotide of one oligonucleotide and the 3' carbon of another nucleotide. Template-driven ligation reactions are described in references such as U.S. Pat. Nos. 4,883,750, 5,476,930, 5,593,826, and 5,871,921, which are incorporated herein by reference in their entireties. The term "ligation" also encompasses the non-enzymatic formation of phosphodiester bonds, as well as the formation of non-phosphodiester covalent bonds between the ends of oligonucleotides, such as phosphorothioate bonds, disulfide bonds, and the like.

[0122] As used herein, the term "each," when used in reference to a set of items, is intended to identify an individual item in the set, but does not necessarily refer to every item in the set, unless the context clearly dictates otherwise.

[0123] As used herein, the term "extending," when used with respect to a nucleic acid, is intended to mean the addition of at least one nucleotide or oligonucleotide to a nucleic acid. In certain embodiments, one or more nucleotides can be added to the 3'-end of a nucleic acid, for example, via polymerase catalysis (e.g., DNA polymerase, RNA polymerase, or reverse transcriptase). Chemical or enzymatic methods can be used to add one or more nucleotides to the 3'- or 5'-end of a nucleic acid. One or more oligonucleotides can be added to the 3'- or 5'-end of a nucleic acid, for example, via chemical or enzymatic (e.g., ligase-catalyzed) methods. A nucleic acid can be extended in a template-directed manner, whereby the extension product is complementary to a template nucleic acid hybridized to the nucleic acid being extended.

[0124] Provided herein are arrays and methods for spatial detection and analysis of nucleic acids in tissue samples (e.g., mutation analysis or single nucleotide variation (SNV) detection and indel detection). The arrays described herein can include a substrate to which multiple capture probes are immobilized, such that each capture probe occupies a different position on the array. Some or all of the multiple capture probes can include a unique position tag (i.e., a spatial address or index sequence). The spatial address can describe the position of the capture probe on the array. The position of the capture probe on the array can be correlated to a position in the tissue sample.

[0125] As used herein, the terms "poly T" or "poly A," when used in reference to a nucleic acid sequence, are intended to mean a series of two or more thiamine (T) or adenine (A) bases, respectively. The poly T or poly A can contain at least about 2, 5, 8, 10, 12, 15, 18, 20, or more T or A bases, respectively. Alternatively, or additionally, the poly T or poly A can contain up to about 30, 20, 18, 15, 12, 10, 8, 5, or 2 T or A bases, respectively.

[0126] As used herein, the terms "poly T," "poly A," or "poly U," when used in reference to a nucleic acid sequence (e.g., a capture nucleotide sequence), are intended to mean a series of two or more thiamine (T), adenine (A), or uridine (U) bases, respectively. Poly T or poly A or poly U can contain at least about 2, 5, 8, 10, 12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, or more T or A bases, respectively. Alternatively or additionally, poly T or poly A or poly U can contain up to about 40, 38, 35, 32, 30, 28, 25, 22, 20, 18, 15, 12, 10, 8, 5, or 2 T or A bases. In some embodiments, the present disclosure contemplates the use of "TVN" sequences, where "T" is a capture nucleotide sequence, "V" is adenine (A), cytosine (C), or guanine (G), and "N" is adenine (A), cytosine (C), guanine (G), or thymine (T). TVN sequences are used in some embodiments to bias reverse transcription toward bases of a poly-A tail on an mRNA molecule.

[0127] As used herein, the terms "tagmentation," "tagment," or "tagmenting" refer to the conversion of nucleic acids, e.g., DNA, into adapter-modified templates in solution ready for clustering and sequencing using transposase-mediated fragmentation and tagging. This process often involves modification of the nucleic acid by a transposome complex containing a transposase enzyme complexed with adapters containing transposon end sequences. Tagging simultaneously results in fragmentation of the nucleic acid and ligation of adapters to the 5' ends of both strands of the double-stranded fragments. Following a purification step to remove the transposase enzyme, additional sequences are added to the ends of the adapted fragments by PCR.

[0128] "Transposase" refers to an enzyme capable of forming a functional complex containing a transposon end-containing composition (e.g., a transposon, a transposon end, a transposon end composition) and catalyzing the insertion or transposition of the transposon end-containing composition into a double-stranded target nucleic acid with which it is incubated, e.g., in an in vitro transposition reaction. Transposases provided herein can also include integrases from retrotransposons and retroviruses. Transposases, transposomes, and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of U.S. Patent Application Publication No. 2010 / 0120098, the entire contents of which are incorporated herein by reference. While many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it is understood that any transposition system capable of inserting transposon ends with sufficient efficiency to 5' tag and fragment the target nucleic acid for the intended purpose may be used in the present invention. In certain embodiments, a preferred transposition system can insert transposon ends in a random or near-random manner to 5' tag and fragment target nucleic acids.

[0129] As used herein, the term "transposition reaction" refers to a reaction in which one or more transposons are inserted into a target nucleic acid, for example, at random or near-random sites. The essential components of a transposition reaction are a transposase and a DNA oligonucleotide representing the nucleotide sequence of the transposon, including the transferred transposon sequence and its complement (the non-transferred transposon end sequence), as well as other components necessary to form a functional transposition or transposome complex. The DNA oligonucleotide may further include additional sequences (e.g., adapter or primer sequences) if needed or desired. In some embodiments, the methods provided herein are exemplified using transposition complexes formed by hyperactive Tn5 transposase and Tn5-type transposon ends (Goryshin and Reznikoff, 1998, J. Biol. Chem., 273:7367) or MuA transposase and Mu transposon ends containing R1 and R2 end sequences (Mizuuchi, 1983, Cell, 35:785; Savilahti et al., 1995, EMBO J., 14:4893). However, any transposition system capable of inserting transposon ends in a random or near-random manner with sufficient efficiency to 5'-tag and fragment target DNA for the intended purpose can be used in the present invention.Examples of transposition systems known in the art that can be used in the methods of the present invention include Staphylococcus aureus Tn552 (Colegio et al., 2001, J. Bacterid., 183:2384-8; Kirby et al., 2002, Mol. Microbiol., 43:173-86), TyI (Devine and Boeke, 1994, Nucleic Acids Res., 22:3765-72 and International Patent Application No. WO 95 / 23875), Transposon Tn7 (Craig, 1996, Science. 271:1512; Craig, 1996, Review in: Curr. Top Microbiol. Immunol., 204:27-48), TnIO and ISIO (Kleckner et al., 1996, Curr. Top Microbiol. Immunol., 204:27-48). Immunol, 204:49-82), mariner transposase (Lampe et al., 1996, EMBO J., 15:5470-9), Tci (Plasterk, 1996, Curr Top Microbiol Immunol, 204:125-43), P Element (Gloor, 2004, Methods Mol Biol, 260:97-114), TnJ (Ichikawa and Ohtsubo, 1990, J Biol Chem., 265:18829-32), bacterial insertion sequences (Ohtsubo and Sekine, 1996, Curr. Top. Microbiol. Immunol. 204:1-26), retrovirus (Brown et al., 1989, Proc Natl Acad Sci USA, 86:2525-9), and yeast retrotransposons (Boeke and Corces, 1989, Annu Rev Microbiol. 43:403-34). Methods for inserting transposon ends into target sequences can be performed in vitro using any suitable transposon system for which a suitable in vitro transposition system is available or which can be developed based on knowledge in the art.Generally, an in vitro transposition system suitable for use in the methods provided herein requires, at a minimum, a transposase enzyme of sufficient purity, sufficient concentration, and sufficient in vitro transposition activity, and transposon ends that form a functional complex with the respective transposase capable of catalyzing a transposition reaction. Suitable transposase transposon end sequences that can be used in the present invention include, but are not limited to, wild-type, derivative, or mutant transposon end sequences that form a complex with a transposase selected from wild-type, derivative, or mutant forms of the transposase. As used herein, the term "transposome complex" refers to a transposase enzyme that noncovalently binds to double-stranded nucleic acid. For example, the complex can be a transposase enzyme preincubated with double-stranded transposon DNA under conditions that support noncovalent complex formation. The double-stranded transposon DNA can include, but is not limited to, Tn5 DNA, a portion of Tn5 DNA, a transposon end composition, a mixture of transposon end compositions, or other double-stranded DNA that can interact with a transposase, such as a hyperactive Tn5 transposase.

[0130] As used herein, the term "random" can be used to refer to the spatial arrangement or composition of locations on a surface. For example, the arrays described herein have at least two types of order: one with respect to the spacing and relative positions of features (also called "sites"), and the second with respect to the identity or predetermined knowledge of specific molecular species present in a particular feature. Thus, the features of an array can be randomly spaced so that nearest neighboring features have variable spacing between each other. Alternatively, the spacing between features can be ordered to form a regular pattern, such as, for example, a rectilinear or hexagonal grid. In another aspect, the features of an array can be random with respect to the identity or predetermined knowledge of the gene of interest (e.g., nucleic acid of a particular sequence) occupying each feature, regardless of whether the spacing results in a random or regular pattern. The arrays described herein can be ordered in one respect and random in another respect. For example, in some embodiments described herein, a surface is contacted with a population of nucleic acids under conditions in which the nucleic acids are ordered with respect to their relative positions, but attach to sites that are "randomly arranged" with respect to knowledge of the sequence for the nucleic acid species present at any particular site. Reference to nucleic acids being "randomly distributed" at sites on a surface is intended to refer to a lack of knowledge or pre-determination as to which nucleic acids will be captured at which sites (whether or not the sites are arranged in an ordered pattern).

[0131] As used herein, a "biological sample" may include one or more biological or chemical substances, such as nucleic acids, oligonucleotides, proteins, cells, tissues, organisms, and / or biologically active chemical compounds, such as analogs or mimetics of the aforementioned species. As used herein, the term "tissue" is intended to mean an aggregate of cells and, optionally, intercellular material. Typically, cells in tissue are not free-floating in solution but are attached to each other to form multicellular structures. Exemplary tissue types include muscle, nerve, epidermis, and connective tissue. In some cases, a biological sample may include whole blood, lymph, serum, plasma, sweat, tears, saliva, sputum, cerebrospinal fluid, amniotic fluid, semen, vaginal discharge, serous fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, transudate, exudate, cystic fluid, bile, urine, gastric juice, intestinal fluid, fecal sample, fluid containing single or multiple cells, fluid containing cellular organelles, tissue fluid, organism fluid, viruses, including viral pathogens, fluid containing multicellular organisms, biological swabs, and biological washes. In further examples, the sample may be derived from an organ comprising, for example, an organ of the musculoskeletal system, such as muscle, bone, tendon, or ligament; an organ of the digestive system, such as the salivary gland, pharynx, esophagus, stomach, small intestine, large intestine, liver, gallbladder, or pancreas; an organ of the respiratory system, such as the larynx, trachea, bronchi, lungs, or diaphragm; an organ of the urinary system, such as the kidney, ureter, bladder, or urethra; a reproductive organ, such as the ovaries, fallopian tubes, uterus, vagina, placenta, testes, epididymis, vas deferens, seminal vesicles, prostate, penis, or scrotum; an organ of the endocrine system, such as the pituitary gland, pineal gland, thyroid gland, parathyroid gland, or adrenal gland; an organ of the circulatory system, such as the heart, arteries, veins, or capillaries; an organ of the lymphatic system, such as the lymphatic vessels, lymph nodes, bone marrow, thymus, or spleen; an organ of the central nervous system, such as the brain, brainstem, cerebellum, spinal cord, cranial nerves, or spinal nerves; a sensory organ, such as the eye, ear, nose, or tongue; or an organ of the integument, such as the skin, subcutaneous tissue, or mammary gland. In various embodiments, the tissue may be derived from a multicellular organism. In some embodiments, the tissue section may be contacted with the surface, for example, by placing the tissue on the surface.The tissue can be freshly excised from an organism, or the tissue may have been previously preserved, for example, by freezing (e.g., fresh-frozen tissue), embedding in a material such as paraffin (e.g., formalin-fixed paraffin-embedded (FFPE) samples), formalin fixation, infiltration, dehydration, etc. Optionally, the tissue section can be attached to a surface, for example, using the techniques and compositions described in U.S. Pat. No. 11,390,912, incorporated herein by reference in its entirety. In some embodiments, the tissue can be permeabilized, and the cells of the tissue can be lysed when the tissue contacts the surface. Any of a variety of treatments can be used, such as those described above for lysing cells. Target proteins and / or nucleic acids released from the permeabilized tissue can be captured by capture oligonucleotides on the surface. Thus, in various embodiments, the biological sample is a tissue sample. The thickness of the tissue or other biological sample contacted with the surface in the methods described herein can be any suitable thickness desired. In exemplary embodiments, the thickness is at least 0.1 μm, 0.25 μm, 0.5 μm, 0.75 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. Alternatively or additionally, the thickness of the biological sample in contact with the surface is 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, 0.25 μm, 0.1 μm, or less.

[0132] As used herein, the term "tissue sample" refers to a piece of tissue obtained from a subject, fixed, sectioned, and mounted on a planar surface, such as a microscope slide. The tissue sample may be a formalin-fixed, paraffin-embedded (FFPE) tissue sample, a fresh tissue sample, a frozen tissue sample, or the like. The methods disclosed herein may be performed before or after staining the tissue sample. For example, after hematoxylin and eosin staining, the tissue sample may be spatially analyzed according to the methods provided herein. The method may include analyzing the histology of the sample (e.g., using hematoxylin and eosin staining) and then spatially analyzing the tissue.

[0133] As used herein, the term "formalin-fixed paraffin-embedded (FFPE) tissue section" refers to a piece of tissue, e.g., a biopsy, obtained from a subject, fixed in formaldehyde (e.g., 3% to 5% formaldehyde in phosphate-buffered saline) or Bouin's solution, embedded in wax, cut into thin sections, and then mounted on a flat surface, e.g., a microscope slide.

[0134] As used herein, the term "subject" encompasses mammals and non-mammals. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates such as chimpanzees, and other ape and monkey species, cows, horses, sheep, goats, pigs, rabbits, dogs, cats, rodents, rats, mice, guinea pigs, etc. Examples of non-mammals include, but are not limited to, birds, fish, etc. The term does not denote a particular age or sex.

[0135] In some embodiments, nucleic acids in a tissue sample are transferred to an array and captured thereon. For example, a tissue section is placed in contact with the array, and nucleic acids are captured on the array and tagged by spatial address. The spatially tagged DNA molecules are released from the array and analyzed by high-throughput next-generation sequencing (NGS), such as sequencing-by-synthesis (SBS). In some embodiments, nucleic acids in a tissue section (e.g., a formalin-fixed, paraffin-embedded (FFPE) tissue section) are transferred to the array and captured thereon by hybridization to a capture probe. In some embodiments, the capture probe may be a universal capture probe that hybridizes to, for example, an adapter region in a nucleic acid sequencing library or the polyA tail of mRNA. Alternatively, the spatially tagged RNA or DNA molecules are released from the array and analyzed by high-throughput next-generation sequencing (NGS), such as sequencing-by-synthesis (SBS). In some embodiments, nucleic acids in tissue sections (e.g., formalin-fixed, paraffin-embedded (FFPE) tissue sections) are transferred to an array and captured on the array by hybridization to a capture probe. In some embodiments, the capture probe may be, for example, a gene-specific capture probe that hybridizes to a specifically targeted mRNA or cDNA in the sample, such as a TruSeq™ Custom Amplicon (TSCA) oligonucleotide probe (Illumina, Inc.). The capture probe may be multiple capture probes, e.g., multiple of the same or different capture probes.

[0136] In some embodiments, combinatorial indexing (addressing) systems are used to provide spatial information for the analysis of nucleic acids in tissue samples. Combinatorial indexing systems can involve the use of two or more spatial address sequences (e.g., two, three, four, five, or more spatial address sequences).

[0137] In some embodiments, two spatial address sequences are incorporated into nucleic acids during preparation of a sequencing library. The first spatial address can be used to define a specific location in the X dimension (i.e., a capture site) on the capture array, and the second spatial address sequence can be used to define a location in the Y dimension (i.e., a capture site) on the capture array. During library sequencing, both the X and Y spatial address sequences can be determined, and the sequence information can be analyzed to define a specific location on the capture array.

[0138] In some embodiments, three spatial address sequences are incorporated into nucleic acids during preparation of a sequencing library. The first spatial address can be used to define a specific location in the X dimension (i.e., a capture site) on the capture array, the second spatial address sequence can be used to define a location in the Y dimension (i.e., a capture site) on the capture array, and the third spatial address sequence can be used to define the location of a two-dimensional sample section (e.g., the location of a slice of a tissue sample) in a sample (e.g., a tissue biopsy) to provide positional spatial information in the third dimension (Z dimension) of the sample. During library sequencing, the X, Y, and Z spatial address sequences can be determined, and the sequence information can be analyzed to define a specific location on the capture array.

[0139] In some embodiments, a temporal address sequence (T) is optionally incorporated into the nucleic acid during preparation of a sequencing library. In some embodiments, the temporal address sequence can be combined with two or three spatial address sequences. The temporal address sequence can be used, for example, in the context of a time-course experiment to determine time-dependent changes in gene expression in a tissue sample. Time-dependent changes in gene expression can occur in a tissue sample, for example, in response to a chemical, biological, or physical stimulus (e.g., a toxin, drug, or heat). Nucleic acid samples obtained at different time points from comparable tissue samples (e.g., proximal slices of a tissue sample) can be pooled and sequenced in bulk. An optional first spatial address can be used to define a specific location in the X dimension (i.e., a capture site) on the capture array, an optional second spatial address sequence can be used to define a location in the Y dimension (i.e., a capture site) on the capture array, and an optional third spatial address sequence can define the location of a two-dimensional sample section (e.g., the location of a slice of a tissue sample) in a sample (e.g., a tissue biopsy) to provide positional spatial information in the third dimension (Z dimension) of the sample. During library sequencing, the T, X, Y, and Z address sequences are determined and the sequence information is analyzed to define a specific X, Y (and optionally Z) location on the capture array for each time point (T).

[0140] The address sequences X, Y, and optionally Z and / or T may be contiguous nucleic acid sequences, or the address sequences may be separated by one or more nucleic acids (e.g., 2 or more, 3 or more, 10 or more, 30 or more, 100 or more, 300 or more, or 1,000 or more). In some embodiments, the X, Y, and optionally Z and / or T address sequences may each individually and independently be combinatorial nucleic acid sequences.

[0141] In some embodiments, the length of an address sequence (e.g., X, Y, Z, or T) can each individually and independently be 100 nucleic acids or less, 90 nucleic acids or less, 80 nucleic acids or less, 70 nucleic acids or less, 60 nucleic acids or less, 50 nucleic acids or less, 40 nucleic acids or less, 30 nucleic acids or less, 20 nucleic acids or less, 15 nucleic acids or less, 10 nucleic acids or less, 8 nucleic acids or less, 6 nucleic acids or less, or 4 nucleic acids or less. The lengths of two or more address sequences in a nucleic acid can be the same or different. For example, if the length of address sequence X is 10 nucleic acids, the length of address sequence Y can be, for example, 8 nucleic acids, 10 nucleic acids, or 12 nucleic acids.

[0142] An address sequence (eg, a spatial address sequence such as X or Y) may be a partially or fully degenerate sequence.

[0143] In some embodiments, spatially addressed capture probes on the array may be released from the array onto tissue sections for the creation of spatially addressed sequencing libraries. In some embodiments, the capture probes comprise random primer sequences for in situ synthesis of spatially tagged cDNA from RNA in the tissue sections. In some embodiments, the capture probes are TruSeq™ Custom Amplicon (TSCA) oligonucleotide probes (Illumina, Inc.) for capturing and spatially tagging genomic DNA in tissue sections. Spatially tagged nucleic acid molecules (e.g., cDNA or genomic DNA) are recovered from the tissue sections and processed in a single-tube reaction to create a spatially tagged amplicon library.

[0144] In another embodiment, the present disclosure provides a substrate, e.g., a flow cell, nanoparticle, or bead, comprising the spatially addressable probes disclosed herein. In certain embodiments, the beads comprise the spatially addressable probes disclosed herein. In further embodiments, the substrate comprises streptavidin on the surface of the bead. In yet further embodiments, the beads comprise multiple oligos bound to the beads via linkages or reversible linkages. An example of a reversible linkage includes a biotin molecule, such as a ddBio molecule. The oligos bound to the substrate typically comprise an adapter sequence, such as a P5 sequence or a P7 sequence. As used herein, a P5 sequence includes a sequence comprising AAT GAT ACG GCG ACC ACC GA (SEQ ID NO: 1) or AAT GAT ACG GCG ACC ACC GAG ATC TAC AC (SEQ ID NO: 2), and a P7 sequence includes the sequence CAA GCA GAA GAC GGC ATA CG (SEQ ID NO: 3) or CAA GCA GAA GAC GGC ATA CGA GAT (SEQ ID NO: 4). In some embodiments, the P5 or P7 sequence can further comprise a spacer polynucleotide, which can be 1 to 20, e.g., 1 to 15, or 1 to 10, nucleotides, e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. In some embodiments, the spacer comprises 10 nucleotides. In some embodiments, the spacer comprises 10 nucleotides. In some embodiments, the spacer is a poly-T spacer, such as a 10T spacer. The spacer nucleotide can be included at the 5' end of the polynucleotide or can be attached to a suitable support via linkage to the 5' end of the oligo. Attachment can be achieved by a sulfur-containing nucleophile, such as a phosphorothioate, present at the 5' end of the polynucleotide. In some embodiments, the oligo comprises a poly-T spacer and a 5' phosphorothioate group. Thus, in some embodiments, the P5 sequence comprises 5' phosphorothioate-TTTTTTTTTTAATGATACGGCGACCACCGA-3' (SEQ ID NO: 17), and in some embodiments, the P7 sequence comprises 5' phosphorothioate-TTTTTTTTTTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 18).In certain embodiments, the oligo attached to the substrate comprises an address sequence that, when decoded, allows the x,y position of the oligo to be determined. In further embodiments, the address sequence is 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides in length, or a range including or between any two of the foregoing nucleotide lengths. In another embodiment, the oligo comprises a transposome hybridization region (Tsm hyb). In yet additional embodiments, the oligo comprises a sequencing primer site sequence. Examples of sequencing primer site sequences include sequences complementary to the R1 and R2 sequencing primers from Illumina™. In further embodiments, the oligo may further comprise one or more linker sequences. In yet further embodiments, the oligo may further comprise one or more index sequences. In certain embodiments, oligos may contain one or more unique molecular identifier (UMI) sequences. Unique molecular identifiers (UMIs) are a type of molecular barcoding that allows for error correction and improved accuracy during sequencing. These molecular barcodes are short sequences used to uniquely tag each molecule in a sample library. UMIs are used in a wide range of sequencing applications, often PCR replication in DNA and cDNA. UMI deduplication is also useful for RNA-seq gene expression analysis and other quantitative sequencing methods. As described above, oligos contain moieties or sequences that can specifically bind to polynucleotides derived from biological samples (e.g., tissue samples). Thus, oligos are spatially addressable probes for polynucleotides derived from biological samples. Moieties or sequences that can specifically bind to polynucleotides derived from biological samples may be selected for specific omic applications. For example, oligos may contain oligo d(T) sequences for transcriptomics or assays (e.g., RNA-seq assays). Alternatively, the oligos may contain sequences that bind to genomic DNA from a biological sample for genomic applications or assays (e.g., ATAC-seq assays).As provided in the examples presented herein, the substrate can include multiple types of oligos with different moieties or sequences such that the spatially addressable probes can specifically bind to two or more different types of polynucleotides from a biological sample. The use of multiple types of oligos is ideally suited for multi-omic or multi-assay applications.

[0145] In some embodiments, magnetic nanoparticles can be used to capture nucleic acids (e.g., in situ synthesized cDNA) in tissue samples for the creation of spatially addressed libraries.

[0146] In some embodiments, spatial detection and analysis of nucleic acids in tissue samples can be performed on a droplet actuator.

[0147] Described herein are improved methods and compositions for spatial omics applications that preserve spatial information related to the origin of RNA or DNA in tissues. Examples of spatial omics applications include, but are not limited to, spatial genomics applications, spatial proteomics applications, spatial transcriptomics applications, spatial agronomic applications, spatial epigenomics applications, spatial phenomics applications, spatial ligandomics applications, and spatial multiomics applications (e.g., transcriptome and genome applications).

[0148] Polynucleotide isolation In various embodiments, one or more samples contacted with a solid support can be lysed to release the target nucleic acid using methods known in the art, such as using one or more of chemical treatment, enzymatic treatment, electroporation, heat, hypotonic treatment, sonication, etc.

[0149] In some embodiments, tissue samples are treated to remove embedding material from the sample (e.g., remove paraffin or formalin) prior to nucleic acid release, capture, or modification. This can be accomplished by contacting the sample with an appropriate solvent (e.g., xylene and ethanol washes). Treatment can occur before contacting the tissue sample with a solid support described herein, or treatment can occur while the tissue sample is on the solid support. Exemplary methods for engineering tissue for use with solid supports to which nucleic acids are attached are described in U.S. Patent Application Publication No. 2014 / 0066318, incorporated herein by reference.

[0150] Preparation of polynucleotides The present disclosure is based in part on the recognition that the amount of RNA or DNA information that can be isolated from fresh or frozen tissue samples, as well as FFPE tissue samples, needs to be improved to provide information related to the genetic profile of the tissue sample. The present disclosure provides methods for improving the capture of genetic information by increasing the quantity and quality of RNA isolated from tissue samples that can be used in spatial transcriptomics analysis.

[0151] Total RNA can include ribosomal RNA (rRNA), messenger RNA (mRNA), transfer RNA (tRNA), microRNA (miRNA), non-coding RNA (ncRNA), small nucleolar RNA (snoRNA), and / or small nuclear RNA (snRNA). In various embodiments, the RNA is rRNA and / or mRNA.

[0152] In various embodiments, the RNA capture probe is selected from the group consisting of a poly-T sequence, a poly-U sequence, a randomer, a semi-random sequence, or a target-specific probe. In various embodiments, the target-specific probe comprises a plurality of different target-specific RNA capture probe sequences. In various embodiments, the RNA capture probe or surface capture probe is 8-80 nucleotides. In certain embodiments, the RNA capture probe or surface probe is between 10 and 80 nucleotides, between 10 and 70 nucleotides, between 10 and 60 nucleotides, between 10 and 50 nucleotides, between 10 and 40 nucleotides, between 10 and 30 nucleotides, between 10 and 20 nucleotides, between 20 and 80 nucleotides, between 20 and 70 nucleotides, between 20 and 60 nucleotides, between 20 and 50 nucleotides, between 20 and 40 nucleotides, or 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, or 80 nucleotides.

[0153] In various embodiments, the capture oligonucleotide comprises a clustered primer sequence and a capture nucleotide sequence configured to bind to a target nucleic acid of a biological sample. In some embodiments, the capture oligonucleotide comprises a clustered primer sequence (e.g., a P7 sequence), a spatial barcode (SBC) sequence, a sequencing primer sequence (e.g., a sequencing-by-synthesis (SBS) sequence such as SBS12), a single molecule identifier (SMI) sequence, a quality control sequence, and a TVN sequence, where "T" is the capture nucleotide sequence, "V" is adenine (A), cytosine (C), or guanine (G), and "N" is adenine (A), cytosine (C), guanine (G), or thymine (T). In various embodiments, the capture oligonucleotide is about 30 bases to about 100 bases in length, or about 30 bases to about 90 bases, or about 30 bases to 80 bases, or is about 30 to 70 bases, or about 30 to 60 bases, or about 30 to 55 bases, or about 30 to 50 bases in length, or 20 to 80 bases, or 10 to 80 bases. In further embodiments, capture oligonucleotides of the present disclosure are about 10 bases, 20 bases, 30 bases, 35 bases, 40 bases, 45 bases, 50 bases, 55 bases, 60 bases, 65 bases, 70 bases, 75 bases, 80 bases, 85 bases, 90 bases, 95 bases, or 100 bases in length. Capture nucleotide sequences that can hybridize to or otherwise associate with an analyte (e.g., a target nucleic acid) can be, for example, but not limited to, a universal sequence (e.g., a poly-T sequence, a random nucleotide sequence, or a semi-random nucleotide sequence) or a target-specific (e.g., a gene-specific) sequence. In various embodiments, the capture nucleotide sequence (e.g., a poly-T nucleotide sequence or a random nucleotide sequence) is , 2, 5, 8, 10, 12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, 45, 50 or more bases in length, or about 2, 5, 8, 10, 12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, 45, 50 or more bases in length, or at least about 2, 5, 8, 10, 12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, 45, 50 or more bases in length.Alternatively, or additionally, the capture nucleotide sequence can comprise no more than about 50, 45, 40, 38, 35, 32, 30, 28, 25, 22, 20, 18, 15, 12, 10, 8, 5, or 2 bases. The capture oligonucleotide can comprise additional elements, including, but not limited to, a single molecular identifier (SMI) (e.g., a unique molecular identifier (UMI)), an index sequence, a sequence complementary to a sequencing primer (e.g., SBS12), or combinations thereof. In some embodiments, the beads are packed on a solid support (e.g., a planar support or a flow cell), and the beads comprise a plurality of capture oligonucleotides immobilized thereon, wherein one or more of the plurality of capture oligonucleotides comprise, from 5' to 3', (a) a first clustered primer sequence, (b) a spatial barcode (SBC) sequence, (c) a first sequencing primer sequence, (d) a single molecule identifier (SMI) sequence, (e) a quality control sequence, and (f) a TVN sequence, wherein "T" is the capture nucleotide sequence, "V" is adenine (A), cytosine (C), or guanine (G), and "N" is adenine (A), cytosine (C), guanine (G), or thymine (T), and the spatial barcode sequence of the plurality of capture oligonucleotides is unique to each bead.

[0154] Oligonucleotides comprising a surface oligonucleotide (e.g., a poly-T sequence) may further comprise a spatial index sequence, including, but not limited to, one or more of a P7 sequence, an index sequence, and / or a Read 2 (Rd2) sequence. In various embodiments, the surface oligonucleotide comprises a P7 anchor sequence, a spatial barcode, and a sequence that hybridizes to a splint oligonucleotide.

[0155] In various embodiments, the sequence in the surface oligonucleotide that hybridizes to the splint oligonucleotide is a PZ (clustered) sequence. In various embodiments, the PZ sequence hybridizes to a splint oligonucleotide that includes a nucleotide sequence PZ' complementary to the PZ sequence and a PX' sequence complementary to the surface capture probe. In various embodiments, the PX sequence is a seeding sequence. In one embodiment, PX has the sequence AGGAGGAGGAGGAGGAGGAGGAGG (SEQ ID NO: 21).

[0156] In various embodiments, the cleavable linker that attaches the capture probe to the nanostructure is a cleavable polynucleotide, hi various embodiments, the cleavable polynucleotide is 5-25 nucleotides, or 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides.

[0157] In some embodiments, total RNA is released from a tissue sample. Release includes tissue lysis or tissue permeabilization. In various embodiments, one or more samples contacted with a solid support can be lysed to release the target nucleic acid. Lysis can be achieved using known techniques, such as using one or more of chemical treatment, enzymatic treatment, electroporation, heat, hypotonic treatment, sonication, etc. It is contemplated that the tissue sample is permeabilized prior to capture. In various embodiments, the tissue sample is treated with one or more blocking reagents prior to capture. In various embodiments, the tissue sample is permeabilized and treated with one or more blocking reagents prior to capture.

[0158] In some embodiments, tissue samples are treated to remove embedding material from the sample (e.g., remove paraffin or formalin) prior to nucleic acid release, capture, or modification. This can be accomplished by contacting the sample with an appropriate solvent (e.g., xylene and ethanol washes). Treatment can occur before contacting the tissue sample with a solid support described herein, or treatment can occur while the tissue sample is on the solid support. Exemplary methods for engineering tissue for use with solid supports to which nucleic acids are attached are described in U.S. Patent Application Publication No. 2014 / 0066318, incorporated herein by reference.

[0159] Formalin-fixed tissue samples may also be decrosslinked using known techniques. In various embodiments, decrosslinking is performed using, for example, Tris-EDTA (TE) buffer at pH 8, pH 9, or another suitable buffer at an appropriate pH. Decrosslinking may also be performed at elevated temperatures, for example, 70°C.

[0160] RNA from a sample can also be prepared by performing end repair of RNA using polynucleotide kinase before capturing RNA from a tissue sample, and / or by performing in situ polyadenylation using polyadenylate polymerase before capturing RNA from a tissue sample. Methods for end repair of RNA from a tissue sample are described in co-owned U.S. Provisional Application No. 63 / 477,730 (Attorney Docket No. 33080 / IP-2625-P), which is incorporated herein by reference.

[0161] The above methods are also useful for improving the capture efficiency of mRNA transcripts for in situ mRNA transcript library preparation and / or for improving the nucleotide length of polynucleotides used in generating an in situ transcriptome library (e.g., for improving the polynucleotide size of cDNA transcribed from mRNA isolated from a sample and used in generating an in situ transcriptome library).

[0162] Spatial detection and analysis of nucleic acids in tissue samples According to the methods described herein, spatial detection and analysis of nucleic acids in a tissue sample can be performed using a set of two or more capture probes (e.g., three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more capture probes). Typically, at least a first capture probe in a set of capture probes is immobilized on a capture array. In some embodiments, a second capture probe can be immobilized on the same capture array as the first capture probe, e.g., in close proximity to the first capture probe, e.g., at the same capture site. In some embodiments, the second capture probe can be immobilized on particles, such as magnetic particles or magnetic nanoparticles. In some embodiments, the second capture probe can be in a solution used to perform an in situ reaction with nucleic acids in a tissue sample, for example.

[0163] Typically, at least the first capture probe in the set of capture probes is immobilized on a capture array or nanostructure. In some embodiments, the second capture probe can be immobilized on the same capture array as the first capture probe, for example, in close proximity to the first capture probe, for example, at the same capture site. In some embodiments, the second capture probe can be immobilized on a nanostructure or particle, such as a magnetic particle or magnetic nanoparticle. In some embodiments, the second capture probe can be in a solution used to perform an in situ reaction with nucleic acids in a tissue sample, for example.

[0164] The capture probes in a capture probe set can individually and independently have a variety of different regions, such as a capture region (e.g., a first universal or gene-specific capture region or a first clustered region), a primer binding region (e.g., an SBS primer region such as an SBS3 or SBS12 region), or a second universal region / clustered sequence such as a P5 or P7 region, a spatial address region (e.g., a partial or combinatorial spatial address region), or a cleavable region.

[0165] "Sequencing by synthesis (SBS) technology" generally involves the enzymatic extension of a nascent nucleic acid chain by the repeated addition of nucleotides to a template chain.In the conventional method of SBS, a single nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in each delivery.However, in the method described herein, two or more types of nucleotide monomers can be provided to a target nucleic acid in the presence of a polymerase during delivery.

[0166] Briefly, SBS can be initiated by contacting the barcode with one or more labeled nucleotides, DNA polymerase, etc. These features, where a primer is extended using the barcode-containing sequence as a template, can incorporate a labeled nucleotide that can be detected. Optionally, the labeled nucleotide can further include a reversible termination feature that terminates further primer extension once the nucleotide is added to the primer. For example, a nucleotide analog with a reversible terminator moiety can be added to the primer so that further extension does not occur until a deblocking agent is delivered to remove the moiety. Thus, in embodiments using reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washing can be performed between the various delivery steps. The cycle is then repeated n times to extend the primer with n nucleotides, thereby allowing a sequence of length n to be detected. Exemplary SBS procedures, fluidic systems, and detection platforms that can be readily adapted for use with arrays generated by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, WO 91 / 06678, WO 07 / 123744, U.S. Patent Nos. 7,057,026, 7,329,492, 7,211,414, 7,315,019, or 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082(A1), each of which is incorporated herein by reference.

[0167] Exemplary sequences include the following Rd1 and Rd2 adapter sequences: Second universal adapter - Rd1 SBS3 (long):

[0168] [Table 1] (SEQ ID NO: 13); Second universal adaptor-Rd1 SBS3 (short chain): ACACTCTTTCCCTACACGAC (SEQ ID NO: 14); First universal adaptor-Rd2 SBS12 (long chain):

[0169] [Table 2] (SEQ ID NO: 15); First universal adapter - Rd2 SBS12 (short chain): GTGACTGGAGTTCAGACGTGT (SEQ ID NO: 16).

[0170] In some embodiments, only one capture probe in a set of capture probes comprises a capture region. In some embodiments, two or more capture probes in a set of capture probes comprise a capture region.

[0171] In some embodiments, only one probe in a set of capture probes comprises a spatial address region, such as a complete spatial address region that describes the location of a capture site on a capture array. In some embodiments, two or more probes in a set of capture probes can comprise a spatial address region, e.g., two or more probes can each comprise a partial spatial address region (i.e., combinatorial address region), where each partial address region describes the location of a capture site on a capture array, e.g., along the x-axis or y-axis.

[0172] In some embodiments, a set of capture probes (e.g., first and second capture probes) may include at least one capture probe that includes a capture region and a spatial-address region (e.g., a complete or partial spatial-address region). In some embodiments, a capture probe in a set of capture probes does not include both a capture region and a spatial-address region.

[0173] In some embodiments, the first capture probe is a 5' gene-specific probe that comprises a sequence complementary to the first universal adapter sequence and a 5' gene-specific primer. In some embodiments, the RNA capture probe is a 5' gene-specific or target-specific probe that comprises a sequence complementary to the first universal adapter sequence and a 5' gene-specific or target-specific primer.

[0174] In some embodiments, the second capture probe is a 3' gene-specific probe comprising a 3' gene-specific primer, a unique molecular index (UMI), and a second universal adapter sequence (e.g., an Rd1 adapter). In some embodiments, the second capture probe does not comprise a spatial address region. In some embodiments, the surface capture probe is a 3' gene-specific or target-specific probe comprising a 3' gene-specific or target-specific primer, a unique molecular index (UMI), and a second universal adapter sequence (e.g., an Rd1 adapter). In some embodiments, the surface capture probe does not comprise a spatial address region.

[0175] If the surface oligonucleotide molecules are randomly positioned on a substrate (e.g., a flow cell), the method further includes determining substrate positions of one or more surface oligonucleotide molecules by sequencing the spatial barcodes of the surface oligonucleotide molecules and assigning the spatial barcode sequences to positions on the substrate. Optionally, in some embodiments, the RNA capture probe is a 5' gene-specific or target-specific probe comprising a sequence complementary to a first universal adapter sequence and a 5' gene-specific or target-specific primer. The method further includes sequencing at least a portion of one or more spatially barcoded first strand cDNA molecules or copies thereof to identify spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules or copies thereof, and correlating the spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules or copies thereof with known positions of spatial barcode sequences of surface oligonucleotide molecules. In various embodiments, the sequences of the spatial barcodes are determined by next-generation sequencing.

[0176] If the surface oligonucleotide molecules are arranged in clusters on a substrate (e.g., a flow cell), the method further includes determining the substrate location of each cluster by sequencing the spatial barcode for at least one surface oligonucleotide molecule in each cluster and assigning the spatial barcode sequence to a location on the substrate before contacting the tissue sample with the substrate. Optionally, the method further includes sequencing at least a portion of one or more spatially barcoded first strand cDNA molecules or copies thereof to identify the spatial barcode sequence of the one or more spatially barcoded first strand cDNA molecules or copies thereof, and determining the spatial location of the RNA molecules within the tissue sample by correlating the spatial barcode sequence of the one or more spatially barcoded first strand cDNA molecules or copies thereof with the known locations of the spatial barcode sequences of the surface oligonucleotide molecules. In various embodiments, the sequence of the spatial barcode is determined by next-generation sequencing.

[0177] When the surface oligonucleotide molecules are arranged in a pattern on a substrate (e.g., a flow cell) such that the substrate positions and sequences of the spatial barcodes of the surface oligonucleotides on the substrate are known before contacting the tissue with the flow cell, the method further includes sequencing at least a portion of one or more spatially barcoded first strand cDNA molecules or copies thereof to identify the spatial barcode sequences of the spatially barcoded first strand cDNA molecules or copies thereof, and determining the spatial positions of RNA molecules within the tissue sample by correlating the spatial barcode sequences of the spatially barcoded first strand cDNA molecules or copies thereof with the known positions of the spatial barcode sequences of the surface oligonucleotide molecules. Optionally, the method further includes determining the spatial positions of RNA molecules within the tissue sample by sequencing at least a portion of the one or more spatially barcoded first strand cDNA molecules and correlating the spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules or copies thereof with one or more corresponding spatial barcode sequences of surface oligonucleotide molecules having predetermined positions on the substrate.

[0178] In some embodiments, the capture sites on the substrate are a plurality of capture sites, ie, 2 or more, 10 or more, 30 or more, 100 or more, 300 or more, 1,000 or more, 3,000 or more, 10,000 or more, 30,000 or more, 100,000 or more, 300,000 or more, 1,000,000 or more, 3,000,000 or more, or 10,000,000 or 1,000,000,000 or more capture sites.

[0179] In various embodiments, the capture array or substrate has an area of ​​1 square centimeter (cm 2 In various embodiments, the density may be about 100k / mm or greater, 2 or greater, 10 or greater, 30 or greater, 100 or greater, 300 or greater, 1,000 or greater, 3,000 or greater, 10,000 or greater, 100,000 or greater, 1,000,000 or greater, or 1,000,000 or greater capture sites per 1000k / mm. 2 ~about 1000k / mm 2, e.g., about 100k clusters / mm 2 , approximately 200k clusters / mm 2 , approximately 300k clusters / mm 2 , approximately 400k clusters / mm 2 , approximately 500k clusters / mm 2 , approximately 600k clusters / mm 2 , approximately 700k clusters / mm 2 , approximately 800k clusters / mm 2 , approximately 900k clusters / mm 2 , or about 1000k clusters / mm 2 is.

[0180] In various embodiments, the pair of capture probes at the capture site is a plurality of pairs of capture probes, hi some embodiments, the plurality of capture probes is 2 or more, 10 or more, 30 or more, 100 or more, 300 or more, 1,000 or more, 3,000 or more, 10,000 or more, 30,000 or more, 100,000 or more, 300,000 or more, 1,000,000 or more, 3,000,000 or more, or 10,000,000 or more, 100,000,000 or more, or 1,000,000,000 or more capture probes.

[0181] In some embodiments, the pair of capture probes in the capture site of the substrate is a plurality of pairs of capture probes. In some embodiments, each first capture probe in the plurality of pairs of capture probes in the same capture site comprises the same spatial address sequence. In some embodiments, each first capture probe in the plurality of pairs of capture probes in different capture sites comprises a different spatial address sequence.

[0182] In some embodiments, the surface of the capture array is a planar surface, such as a glass surface. In some embodiments, the surface of the capture array comprises one or more wells. In some embodiments, the one or more wells correspond to one or more capture sites. In some embodiments, the surface of the capture array is a bead surface.

[0183] In some embodiments, the capture region in the second capture probe is a gene-specific capture region.In some embodiments, the gene-specific capture region in the second capture probe comprises the sequence of TruSeq™ Custom Amplicon (TSCA) oligonucleotide probe (Illumina, Inc.).For example, the gene-specific capture region in the multiple second capture probes in the capture site can comprise multiple sequences of TSCA oligonucleotide probe.

[0184] In some embodiments, the capture region in the second capture probe is a gene-specific or target-specific capture region. In some embodiments, the gene-specific or target-specific capture region in the second capture probe comprises the sequence of a TruSeq™ Custom Amplicon (TSCA) oligonucleotide probe (Illumina, Inc.). For example, the gene-specific or target-specific capture region in multiple surface capture probes at a capture site can comprise multiple sequences of TSCA oligonucleotide probes.

[0185] mRNA library preparation The present disclosure provides improved methods for preparing mRNA transcript libraries from samples that provide a more complete spatial transcriptomics profile. The genetic profile of the sample can be used to diagnose and determine treatment for subjects who have or are at risk of having a disease determined by the genetic profile.

[0186] A method for preparing an mRNA transcript expression library from a tissue sample, e.g., a fixed tissue sample, comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, wherein the capture oligonucleotides comprise a first clustered sequence (e.g., P7), a spatial barcode sequence (SBC), and a first universal adaptor sequence (e.g., Rd2 adaptor); and b) coupling the tissue sample to: i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence and a 5' gene-specific primer; and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence (e.g., Rd1 adaptor), wherein one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample. c) contacting the tissue sample of (b) with a ligation reagent under conditions such that the 5' gene-specific probe and the 3' gene-specific probe hybridized to the mRNA transcripts adjacent to each other are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts hybridized to the ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; and e) capturing the ligated gene-specific probe pair oligonucleotides of (d) on a substrate by binding a sequence complementary to the first universal adaptor sequence in the 5' gene-specific probe to a first universal adaptor sequence (e.g., an Rd2 adaptor) of the capture oligonucleotide.

[0187] In various embodiments, the substrate is a glass slide, a bead, or a flow cell. In various embodiments, the flow cell is an ordered flow cell or a random flow cell.

[0188] In various embodiments, the 5' gene-specific probe and / or the 3' gene-specific probe are 10 to 50 nucleotides in length, or 20 to 40 nucleotides in length, or 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length.

[0189] In various embodiments, the 3' gene-specific probe comprises one or more ribobases. In some embodiments, the 3' gene-specific probe comprises 1, 2, 3, 4, 5 or more ribobases.

[0190] In various embodiments, the UMI comprises between 6 and 20 nucleotides, or 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0191] In another embodiment, the method includes hybridizing the transcripts as described above in step (a), wherein the hybridization leaves a nucleotide gap between the hybridized probes. Step (b) includes hybridizing the tissue sample with i) a plurality of 5' gene-specific probes comprising a 5' gene-specific primer and a sequence complementary to a first universal adaptor sequence, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence (e.g., an Rd1 adaptor), under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample, and hybridization of the 5' gene-specific probes and the 3' gene-specific probes on the mRNA transcripts. (c) contacting the tissue sample of (b) with a nucleotide base and a ligation reagent such that the nucleotide gap between the 5' and 3' gene-specific probes hybridized to the mRNA transcripts is filled with a nucleotide base complementary to the mRNA transcript, and the 5' and 3' gene-specific probes are ligated together to form one or more ligated gene-specific probe pairs. Steps (d) and (e) are as described above.

[0192] For gap-filling reactions, the gap can be 1 to 50 or more nucleotides, for example, 50 or more nucleotides, 1 to 50 nucleotides, 1 to 40 nucleotides, 1 to 30 nucleotides, 1 to 20 nucleotides, or 1 to 10 nucleotides.

[0193] In various embodiments, the 5' gene specific probe and / or the 3' gene specific probe comprises a locked nucleic acid (LNA) to reduce or prevent strand displacement.

[0194] The clustered sequences can be known index sequences. For example, in some embodiments, the first clustered sequence comprises a P7 sequence (e.g., CAAGCAGAAGACGGCATACG (SEQ ID NO: 3) or CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 4)), and the second clustered sequence comprises a P5 sequence (e.g., AATGATACGGGCGACCACCGA (SEQ ID NO: 1) or AATGATACGGGCGACCACCGAGATCTACAC (SEQ ID NO: 2)).

[0195] The universal primer also comprises a sequence known in the field of spatial transcriptomics. In some embodiments, the first universal primer sequence comprises GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 19). In some embodiments, the second universal primer sequence comprises the Rd1 sequence set forth in AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT (SEQ ID NO: 20).

[0196] To prevent unintended premature capture to the substrate, the 5' gene-specific probe and the 3' gene-specific probe anneal at different temperatures compared to the capture probe. In various embodiments, the 5' gene-specific probe and / or the 3' gene-specific probe have a melting temperature (Tm) of about 50-55°C. In various embodiments, the capture oligonucleotide has a melting temperature (Tm) of about 40-42°C.

[0197] Taking into account the desired melting temperature, it is contemplated that step (b) of the present method is carried out at approximately 50-55° C. It is further contemplated that step (e) is carried out at approximately 40-42° C.

[0198] For the ligation reaction, several reverse transcriptases (RT) or polymerases are useful in the present method. In various embodiments, the polymerase is T4 DNA ligase, T4 RNA ligase 2 (T4Rnl2), Splint® DNA ligase, E. coli DNA ligase, or R2D LIGASE. In various embodiments, the ligation reaction is performed at 37°C.

[0199] Prior to strand synthesis, mRNA transcripts can be removed from the reaction, for example, by enzymatic digestion. In various embodiments, mRNA is removed using RNase H or RNase A.

[0200] The methods herein further include indexing and sequencing the ligated gene-specific probe pairs, including: f) performing an extension reaction and PCR on the oligonucleotides of (e) to obtain PCR templates representing one or more mRNA transcripts in the tissue sample; g) eluting the PCR templates from the substrate; and h) performing indexing PCR to generate double-stranded PCR products comprising a first strand PCR product and a second strand complementary to the first strand PCR product.

[0201] In various embodiments, the PCR templates are eluted from the substrate using sodium hydroxide elution, hi various embodiments, the eluted PCR templates are placed into tubes for mRNA transcript library preparation.

[0202] In various embodiments, the method further includes sequencing the PCR product of (h) and determining the location of the mRNA transcript in the tissue based on the spatial barcode sequence of (a).

[0203] In various embodiments, the double-stranded PCR product comprises a second clustered sequence (eg, P5) on the second strand that is complementary to the PCR product of the first strand, and an index sequence.

[0204] The methods herein are intended to provide information regarding the location / position and expression level of specific genes in a tissue sample. For example, in the methods, contacting the tissue sample with a substrate allows the location of capture sites on the substrate to be correlated with locations in the tissue sample, the substrate comprising a plurality of capture sites comprising a plurality of capture probes immobilized on a surface, the capture probes comprising a spatial address region.

[0205] The present disclosure also provides an improved method for preparing a spatially barcoded RNA library from a tissue sample, which provides a more complete spatial transcriptomics profile.Previous methods for generating an RNA library from a tissue sample involve ligating a probe pair to the sample RNA and ligating the probes together, which provides little information about the RNA sequence itself.It is hypothesized herein that separating the hybridization and extension ligation steps will provide more robust sequence information from the initial capture of RNA from the sample.

[0206] Multiple methods have been proposed for copying or ligating portions of RNA with targeting probes, which can then be captured (and then ligated) onto spatially barcoded substrates. RNA includes ribosomal RNA (rRNA), messenger RNA (mRNA), non-coding RNA (ncRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and / or microRNA (miRNA).

[0207] In various embodiments, the substrate is a bead, a bead array, a spot array, a substrate containing multiple wells, a flow cell (e.g., a clustered flow cell), clustered particles disposed on the surface of a chip, a film, or a plate (e.g., a multi-well plate). In various embodiments, the substrate is a gel coating located in or on a flow cell.

[0208] In various embodiments, the substrate comprises a plurality of nanowells or microwells.

[0209] In various embodiments, the RNA capture probe is selected from the group consisting of a poly-T sequence, a randomer, or a target-specific probe. In various embodiments, the target-specific probe comprises a plurality of different target-specific RNA capture probe sequences. In various embodiments, the RNA capture probe or surface capture probe is 8-80 nucleotides. In certain embodiments, the RNA capture probe or surface probe is between 10 and 80 nucleotides, between 10 and 70 nucleotides, between 10 and 60 nucleotides, between 10 and 50 nucleotides, between 10 and 40 nucleotides, between 10 and 30 nucleotides, between 10 and 20 nucleotides, between 20 and 80 nucleotides, between 20 and 70 nucleotides, between 20 and 60 nucleotides, between 20 and 50 nucleotides, between 20 and 40 nucleotides, or 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, or 80 nucleotides.

[0210] In various embodiments, the target-specific probes and / or substrate-specific probes are 10 to 50 nucleotides in length, or 20 to 40 nucleotides in length, or 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length.

[0211] In various embodiments, the UMI comprises between 6 and 20 nucleotides, or 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0212] When clustered sequences are used, the clustered sequences can be known index sequences. For example, in some embodiments, the first clustered sequence comprises a P7 sequence (e.g., CAAGCAGAAGACGGCATACG (SEQ ID NO: 3) or CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 4)), and the second clustered sequence comprises a P5 sequence (e.g., AATGATACGGGCGACCACCGA (SEQ ID NO: 1) or AATGATACGGGCGACCACCGAGATCTACAC (SEQ ID NO: 2)).

[0213] The universal primer also comprises a sequence known in the field of spatial transcriptomics. In some embodiments, the first universal primer sequence comprises GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 19). In some embodiments, the second universal primer sequence comprises the Rd1 sequence set forth in AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT (SEQ ID NO: 20).

[0214] Prior to strand synthesis, RNA, e.g., mRNA transcripts, can be removed from the reaction by, e.g., enzymatic digestion. In various embodiments, RNA is removed using RNase H or RNase A.

[0215] The methods herein further include indexing and sequencing the ligated gene-specific or target-specific probe pairs, which includes performing an extension reaction and PCR on the oligonucleotides to obtain PCR templates representing one or more mRNA transcripts in the tissue sample, eluting the PCR templates from the substrate, and performing indexing PCR to generate double-stranded PCR products comprising a first strand PCR product and a second strand complementary to the first strand PCR product.

[0216] In various embodiments, the PCR templates are eluted from the substrate using sodium hydroxide elution, hi various embodiments, the eluted PCR templates are placed into tubes for mRNA transcript library preparation.

[0217] In various embodiments, the method further comprises sequencing the PCR product and determining the location of the mRNA transcript in the tissue based on the spatial barcode sequence.

[0218] In various embodiments, the double-stranded PCR product comprises a second clustered sequence (eg, P5) on the second strand that is complementary to the PCR product of the first strand, and an index sequence.

[0219] The methods herein are intended to provide information regarding the location / position and expression level of specific genes in a tissue sample. For example, in the methods, contacting the tissue sample with a substrate allows the location of capture sites on the substrate to be correlated with locations in the tissue sample, the substrate comprising a plurality of capture sites comprising a plurality of capture probes immobilized on a surface, the capture probes comprising a spatial address region.

[0220] Biological samples and methods of use The methods of the present invention are useful for determining genetic information or a genetic profile (i.e., specific genes or levels of gene expression) from a biological sample and detecting mutations or defects in genes or changes in genetic markers to aid in diagnosing individuals who have or are at risk of having a disease and to determine the effectiveness of a treatment. A genetic profile refers to the characteristic expression levels of one or more genes / genetic markers in a sample. In the present disclosure, the genetic profile can be measured before, during, and / or after administration of a therapeutic agent to treat a disease described herein, to determine whether the gene levels have changed, e.g., increased or decreased, in association with a particular disease, condition, or treatment regimen.

[0221] Biological samples for use in the present methods are obtained from a subject. In various embodiments, the subject is a mammal, e.g., a human, a non-human primate such as a chimpanzee, other ape and monkey species, cow, horse, sheep, goat, pig, rabbit, dog, cat, rodent, rat, mouse, guinea pig, etc. In various embodiments, the subject is a human.

[0222] The sample may be derived from organs or tissues including, for example, the musculoskeletal system, such as muscle, bone, tendon, or ligament; organs of the digestive system, such as salivary glands, pharynx, esophagus, stomach, small intestine, large intestine, liver, gallbladder, or pancreas; the respiratory system, such as larynx, trachea, bronchi, lungs, or diaphragm; the urinary system, such as kidneys, ureters, bladder, or urethra; reproductive organs / tissues, such as ovaries, fallopian tubes, uterus, vagina, placenta, testes, epididymis, vas deferens, seminal vesicles, prostate, penis, or scrotum; the endocrine system, such as pituitary gland, pineal gland, thyroid gland, parathyroid gland, or adrenal glands; the circulatory system, such as heart, arteries, veins, or capillaries; the lymphatic system, such as lymphatic vessels, lymph nodes, bone marrow, thymus, or spleen; the central nervous system, such as brain, brainstem, cerebellum, spinal cord, cranial nerves, or spinal nerves; the eye, ear, nose, or tongue; or from the integument, such as skin, subcutaneous tissue, or mammary gland.

[0223] When used in the present methods, samples from humans may be considered (or suspected) to be healthy or diseased. In some cases, two specimens may be used, with one specimen considered to be diseased and the second considered to be healthy (e.g., for use as a healthy control). Any of a variety of conditions may be assessed, including, but not limited to, autoimmune disease, cancer, cystic fibrosis, aneuploidy, pathogenic infection, psychological condition, hepatitis, metabolic disorder, diabetes, sexually transmitted disease, heart disease, stroke, cardiovascular disease, multiple sclerosis, or muscular dystrophy. In various embodiments, the disease or condition is cancer, a genetic condition, or a condition associated with a pathogen having an identifiable genetic signature.

[0224] It is contemplated that the methods herein are useful for detecting changes in genetic material compared to a control sample or a sample from a subject prior to the onset of disease, including mutations, deletions, insertions, single nucleotide polymorphisms (SNPs), combinations thereof, and other changes in the genetic profile.

[0225] The methods are also useful for determining whether initiation of therapy, e.g., cancer therapy, in a subject is necessary, and include: i) determining the subject's genetic profile using the methods described herein; ii) determining whether the genetic profile indicates that the subject has a disease or condition; and iii) initiating treatment of the disease or condition with an appropriate therapy.

[0226] Sequencing The methods described herein can be used in conjunction with various nucleic acid sequencing techniques. Particularly applicable techniques are those in which nucleic acids are attached to fixed positions within an array so that their relative positions do not change, and the array is repeatedly imaged. For example, embodiments in which images are obtained in different color channels corresponding to different labels used to distinguish one nucleotide base type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of a target nucleic acid can be an automated process. Preferred embodiments include sequencing-by-synthesis ("SBS") techniques.

[0227] SBS can utilize nucleotide monomers with terminator moieties or nucleotide monomers lacking any terminator moiety. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in more detail below. In methods using nucleotide monomers without terminators, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS techniques utilizing nucleotide monomers with terminator moieties, the terminators can be effectively irreversible under the sequencing conditions used, as in conventional Sanger sequencing using dideoxynucleotides, or the terminators can be reversible, as in the sequencing method developed by Solexa (now Illumina, Inc.).

[0228] SBS techniques can use nucleotide monomers with or without a label moiety. Therefore, incorporation events can be detected based on the properties of the label, such as the fluorescence of the label, the properties of the nucleotide monomer, such as molecular weight or charge, or by-products of nucleotide incorporation, such as the release of pyrophosphate. In embodiments in which two or more different nucleotides are present in the sequencing reagent, the different nucleotides can be distinguishable from one another, or alternatively, the two or more different labels can be distinguishable under the detection technique used. For example, the different nucleotides present in the sequencing reagent can have different labels, which can be distinguished using appropriate optical systems, as exemplified by the sequencing method developed by Illumina, Inc.

[0229] In various embodiments, the technique is pyrosequencing, which detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281(5375),363, U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, and U.S. Patent No. 6,274,320, the disclosures of which are incorporated herein by reference in their entireties. In pyrosequencing, released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of generated ATP is detected via luciferase-generated photons. Nucleic acids to be sequenced can be attached to features in an array, and the array can be imaged to capture chemiluminescent signals generated by nucleotide incorporation into the array features. Images can be obtained after treating the array with specific nucleotide types (e.g., A, T, C, or G). Images obtained after the addition of each nucleotide type differ in terms of which features in the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative positions of each feature remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after treating the array with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.

[0230] In another exemplary type of SBS, cycle sequencing is achieved by stepwise addition of reversible terminator nucleotides containing cleavable or photobleachable dye labels, as described, for example, in International Publication No. 04 / 018497 and U.S. Patent No. 7,057,026 (the disclosures of which are incorporated herein by reference). This approach has been commercialized by Illumina Inc. and is also described in International Patent Publication Nos. 91 / 06678 and 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators, both of which can be reversed and from which the fluorescent labels are cleaved, facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.

[0231] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be taken after incorporation of the label into the arrayed nucleic acid features. In certain embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally distinct label. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, with images of the array being obtained between each addition step. In such embodiments, each image shows nucleic acid features incorporating a particular type of nucleotide. Because the sequence content of each feature varies, different features are present or absent in different images. However, the relative positions of the features remain unchanged within the image. Images obtained from such reversible terminator SBS methods can be stored, processed, and analyzed as described herein. Following the image acquisition step, the label can be removed, and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removal of the label after detection in a particular cycle and before the subsequent cycle has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labeling and removal methods are described below.

[0232] In certain embodiments, some or all of the nucleotide monomers can contain reversible terminators. In such embodiments, the reversible terminator / cleavable fluorophore can comprise a fluorophore attached to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches separate the terminator chemistry from the cleavage of the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al. describe the development of a reversible terminator that uses a small 3' allyl group to block extension but can be easily deblocked by brief treatment with a palladium catalyst. The fluorophore was attached to the group via a photocleavable linker that can be easily cleaved by 30 seconds of exposure to long-wavelength UV light. Therefore, either disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is the use of a natural terminator followed by the placement of a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluorophore, effectively reversing the terminus. Examples of modified nucleotides are also described in U.S. Pat. Nos. 7,427,673 and 7,057,026, the disclosures of which are incorporated herein by reference in their entireties.

[0233] Additional exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Publication No. 2007 / 0166705, U.S. Patent Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Publication No. 2006 / 0240439, U.S. Patent Publication No. 2006 / 0281109, WO 05 / 065814, U.S. Patent Publication No. 2005 / 0100900, WO 06 / 064199, WO 07 / 010251, U.S. Patent Publication No. 2012 / 0270305, and U.S. Patent Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.

[0234] kit In a further aspect, the present disclosure includes kits containing one or more compounds or compositions packaged in a manner that facilitates their use in practicing the methods of the present disclosure. In one embodiment, such kits include a compound or composition described herein packaged in a container, such as a sealed bottle or container, with a label affixed to the container or included in the package that describes the use of the compound or composition in practicing the method. Preferably, the compound or composition is packaged in a unit dosage form. Preferably, the kit includes instructions that describe the use of the composition.

[0235] Kits and articles of manufacture are contemplated herein. Such kits may include a carrier, package, or container compartmentalized to receive one or more containers, such as vials, tubes, etc., each containing one of the separate elements used in the methods described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers may be formed from a variety of materials, such as glass or plastic. For example, the container can contain one or more spatially addressable probes disclosed herein, optionally in a composition or in combination with another agent disclosed herein (e.g., an array, a bead chip). The container optionally has a sterile access port (e.g., the container can be an intravenous solution bag or a vial with a stopper pierceable by a hypodermic needle). Such kits optionally include identifying descriptions or labels or instructions for their use in the methods described herein.

[0236] The kit typically includes one or more additional containers, each containing one or more of a variety of materials (such as reagents and / or devices, optionally in concentrated form) desirable from a commercial and user perspective for use with the spatially addressable probes described herein. Non-limiting examples of such materials include, but are not limited to, buffers, diluents, filters, needles, syringes, carriers, packages, containers, vials, and / or tube labels listing the contents and / or instructions for use, and inserts containing instructions for use. A set of instructions for use is also typically included.

[0237] A label can be on or associated with a container. A label can be present on a container when letters, numbers, or other symbols forming the label are attached, molded, or etched into the container itself, or can be associated with a container when the label is present in a receptacle or carrier that also holds the container, for example, as a package insert. A label can be used to indicate that the contents are to be used for a particular spatial omic application. A label can also indicate instructions for using the contents, such as the methods described herein.

[0238] Further aspects and details of the present disclosure will become apparent from the following examples, which are intended to be illustrative rather than limiting. [Example]

[0239] Example 1 - In situ method for capture of mRNA transcripts In order to improve the capture of mRNA transcripts from fixed or frozen tissue samples, improved methods for capturing mRNA transcripts from tissue samples were developed.

[0240] A schematic diagram of the first method is shown in Figure 1. In the first method, highly multiplexed oligonucleotide probes are hybridized to tissue mRNA, subsequently ligated, released, and captured on a solid surface containing spatially barcoded capture oligonucleotides. The captured, ligated products are eluted from the surface and PCR-amplified using universal adapter sequences to obtain a spatially barcoded library.

[0241] Two assay-specific oligonucleotides are designed to interrogate a single continuous mRNA sequence (≤50 nt). Each of these oligonucleotides consists of two parts: an upstream-specific oligonucleotide (USO) containing a 5' gene-specific sequence (5' GSP) with a terminal phosphate, and a 3' universal capture / partial Rd2' adapter sequence (Rd2'). Meanwhile, the downstream-specific oligonucleotide (DSO) contains a 3' gene-specific sequence (3' GSP) followed by a unique molecular index (UMI) (N=6) and a 5' Rd1 sequence (Rd1). The GSPs of the USO and DSO are each designed to have a Tm of approximately 55°C. Using this approach, oligonucleotide pairs can be designed and multiplexed (pooled) to target the entire transcriptome. The spatially barcoded substrate contains a 5' sequence for clustering (e.g., P7), followed by a spatial barcode (SBC), and a covalently attached surface capture oligonucleotide (SCO) containing a capture sequence (Rd2) complementary to the USO capture sequence. The SCO Rd2 sequence has a Tm of approximately 40°C.

[0242] Hybridization of the oligonucleotide pool occurs at elevated temperatures (approximately 50°C), favoring GSP-mediated hybridization while minimizing hybridization of USO to the capture oligonucleotide Rd2 sequence. Unbound 5' and 3' gene-specific probes are removed by a heated wash (approximately 50°C). The 3' end of the 3' gene-specific probe contains one or more ribonucleotides, which favors RNA ligase 2-mediated ligation. After removal of RNA with RNase H and permeabilization, the ligated cDNA is captured on the SCO by Rd2' / Rd2 hybridization.

[0243] Extension from the 3' end of the Rd2'-containing product releases the captured template from the surface, allowing off-surface indexed PCR (in solution) using an RNA-resistant PCR polymerase. For sequencing, Rd1 provides UMI and cDNA information, Rd2 yields spatial barcodes, and Rd3 enables demultiplexing of samples.

[0244] A schematic diagram of the second method is shown in Figure 2. The second method is similar to that described in the first method, with the important difference that it captures sequences corresponding to endogenous transcripts, thereby providing additional assay specificity.

[0245] For the second method, the GSPs of the 5' and 3' gene-specific probes are designed to have a gap of several nucleotides between the hybridized 3' end of the 3' gene-specific probe and the 5' end of the 5' specific probe (this gap corresponds to the endogenous mRNA sequence) to provide additional assay specificity. Optionally, a polymerase with reverse transcriptase activity but lacking strand displacement activity is used for gap filling, and the nick is then sealed with a ligase. The 5' terminal base of the 5' gene-specific probe contains several locked nucleic acid bases (LNA) to minimize any polymerase-derived strand displacement activity. Subsequent steps are identical to those described in the first method, except for the use of an LNA-resistant PCR polymerase.

[0246] A workflow was developed herein to exploit the ability to isolate and prepare mRNA libraries from formalin-fixed, paraffin-embedded (FFPE) tissue. The hybridization, ligation, surface capture, and transcript copying steps were performed on a substrate containing the tissue sample. To minimize off-target binding, probes with high melting temperatures and hybridizing temperatures were used, and initial hybridization of the probe to its target was performed at approximately 55°C. Once hydrogenation was complete, washes and ligation (37°C) were performed at the same high temperature. The capture reaction was performed at 40°C. The difference in reaction temperatures prevents the capture probe from hybridizing too quickly to the substrate surface, minimizing incomplete or premature capture of mRNA from the tissue sample.

[0247] Probe construction: As a first step, probes were designed to generate libraries using either RNA-mediated oligonucleotide annealing, selection, and ligation with next-generation sequencing (RASL-Seq) method (Illumina), spatial annealing, selection, and ligation with next-generation sequencing (SPASL-Seq), or TruSeq method (Illumina).

[0248] The probes for RASL-seq that bind to RNA contain a 5' index primer (Rd2 adapter) linked to a 3' target sequence (gene-specific probe) and a 5' target sequence (gene-specific probe) linked to a 3' P5 primer. TruSeq primers that bind to cDNA contain a 5' smRNA linked to a USO and DSO-linked 3' SBS3 sequence. Primers were designed to either leave gaps or no gaps upon hybridization to mRNA transcripts. If primer hybridization left gaps between probes on the target polynucleotide, an extension reaction was performed to fill the gaps. The probes were assayed using two library models: a human control library containing genes expressed with low CV across cell types, and an ERCC control model.

[0249] The probe was annealed to the target polynucleotide and ligated together to form a single strand complementary to the target polynucleotide. This was then captured on a substrate containing poly-T and a sequence (Rd2 adapter) complementary to the 5' end of the ligated product and extended by a polynucleotide extension reaction. The ligated strand was then eluted from the capture surface, and second-strand synthesis was carried out by PCR. The primer for second-strand synthesis contained the Rd1 adapter sequence, an index sequence (e.g., i5), and an Rd3 sequence (e.g., P5). The use of a 3' ribonucleotide in the upstream ligation oligonucleotide (USO) has been noted to improve the efficiency of ssDNA ligation in solution. The ribonucleotide reduces strand displacement during the ligation reaction.

[0250] Probes for the TruSeq method were designed in a similar manner to those for RASL-Seq, but contained a 5' Rd2 sequence, unique molecular identifier (UMI), and ULSO on the first probe, and a DLSO and 3' adapter sequence (Rd1) on the second probe. For reactions, we used a two-probe subpool (30 nM) of ERCC, 1 ul or 0.1 ul of ERCC (3 nM or 0.3 nM in a 10 ul mix), and titrated probe concentrations of Titration-ERCC Pool 8012 at 50 uM, 5 uM, or 0.5 uM. Annealing conditions were 5' at 65°C, 5' at 45°C, 10' at 37°C, and 10' at 25°C, with a gradient of 50 mM NaCl in IDTE. Results demonstrate that high probe concentrations appear to inhibit qPCR.

[0251] Ligation assays: Ligation assays for certain methods were also designed to ligate oligonucleotides in situ on tissue-containing substrates. Several different ligases were assayed for effectiveness in the reactions: T4 DNA ligase, T4 RNA ligase 2 (T4Rnl2), Splint® DNA ligase, E. coli DNA ligase, and R2D LIGASE™. The 9x enzyme condition performed best across conditions. T4 RNA ligase 2 performed similarly to the other enzymes at higher concentrations. R2D also appeared to have similar ligation efficiency.

[0252] During ligation assay analysis, ribonucleotides were added to the 3' end of DLSO to determine whether this improved ligation efficiency. Addition of 3' ribonucleotides to the ligation assay increased ligation efficiency for T4 RNA ligase 2 in both single-stranded and splint reactions, but not for T4 DNA ligase, E. coli DNA ligase, or Splint®.

[0253] We tested whether reversing tissue sample crosslinks (e.g., by-products of formalin fixation of the sample) prior to hybridization would improve hybridization and capture efficiency. Crosslink reversal was performed under different conditions using commercially available RNA extraction kits, RNAeasy or RNASTORM™. Fresh-frozen (FF) or FFPE tissue sections were collected, paraffin removed from the samples as needed, and tissue was lysed using standard protocols. Different RNA extraction conditions were used to determine the amount of RNA recovered. For Qiagen RNeasy (Qiagen): FF and FFPE (50 μl and 30 μl elution, respectively), a 15-minute reverse crosslinking step was included in the FFPE kit. The time courses used were: 0, 15, 30, 45, 1, 2, and 4 hours at 70°C overnight; 100 ng in 20 μl; or 0, 15, 30, and 60 minutes at 80°C during extraction. Under these conditions, the main increase in accessibility (RT-qPCR reduction in Cq) is 0-30°C at 80°C. mRNA reduction appears to correlate more strongly with RNA input to the reaction. RNASTORM™ (Cell Data Science) extraction was performed under the same conditions as above to determine the amount of RNA recovered. Reactions indicated that there appeared to be a slight increase in RNA recovery using RNAeasy extraction, but further experiments will be performed to confirm.

[0254] [Table 3] Note: Due to the difficulty of FFPE recovery from high laboratory temperatures, each tube contained 2–3.5 mouse kidney sections.

[0255] The results demonstrate that this method of capturing mRNA from FFPE tissue samples is effective in improving capture efficiency and transcript integrity, thereby providing a more robust spatial transcriptomics library. This improved library is useful, for example, to better characterize gene profiles at the cellular and locational level in samples from subjects with a disease or condition, and to aid in the diagnosis and treatment of such disease or condition.

[0256] Example 2 - Method for generating an RNA library To improve the capture of RNA transcripts from fixed or frozen tissue samples, improved methods for capturing RNA from tissue samples were developed.

[0257] A schematic of the first method is shown in FIG. 3, and an exemplary workflow is shown in FIG.

[0258] In a first exemplary method, an RNA capture probe is hybridized to RNA in a tissue and then extended with reverse transcriptase to form a first-strand cDNA molecule (Figure 3). The RNA capture probe includes a capture oligonucleotide sequence complementary to the RNA in the sample and a first substrate capture oligonucleotide complementary to a first domain of multiple splint oligonucleotides. In an optional step, the extended probe is then melted from the RNA (or the RNA is digested with RNase), and the probe is hybridized to a surface barcoded oligonucleotide on a substrate via the splint oligonucleotide. Each substrate capture probe includes a spatial barcode and a second substrate capture oligonucleotide complementary to a second domain of the splint oligonucleotide. The captured first-strand cDNA molecule is then ligated to the surface barcode oligonucleotide for the extended probe, e.g., using T4 ligase, to generate a spatially barcoded first-strand cDNA. The surface oligos may also include adapter sequences, e.g., P7 adapters, and the RNA capture probes may also contain read primer hybridization sites for reading the spatial barcode. Optionally, instead of dehybridizing the extended probes, if the RNA is mobile (i.e., not crosslinked to tissue or released by decrosslinking), the entire construct may be bound to substrate surface oligos and ligated, followed by reverse transcription on the surface. Optionally, instead of dehybridizing the extended probes, the RNA may be digested, releasing the DNA probes. Ligation may be by enzymatic or chemical methods.

[0259] As a variation of the above strategy, in a second approach, an oligonucleotide is added to the 3' end of the extended probe, which is complementary to a portion of the surface oligonucleotide (Figure 4A, Figure 9). In this method, the RNA capture probe contains an oligonucleotide sequence complementary to the RNA in the sample and a handle sequence. First, the RNA capture oligonucleotide of the RNA capture probe is hybridized to RNA in the tissue sample to form an RNA-RNA capture probe hybrid. The RNA-RNA hybrid is extended using RT to generate first-strand cDNA. Next, a 3' oligonucleotide sequence containing a substrate capture oligonucleotide complementary to the first domain of the substrate capture probe is added to the first-strand cDNA. The surface capture probe contains, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, and a first domain. Next, the substrate capture oligonucleotide of the first-strand cDNA molecule is hybridized to the first domain of the substrate capture probe, and extension of the first domain of the hybridized substrate capture probe is carried out, resulting in a spatially barcoded first-strand cDNA molecule.

[0260] The 3' oligonucleotide allows for extension from the surface capture probe, and the 5' end of the probe can then be used to introduce a P5 or other adaptor. The method of adding an oligo to the 3' end in Figure 4A is shown as tagging. Tn5 has some activity toward DNA / RNA hybrids and can be used to add a 3' OH by tagging (Figure 4A). 3' oligonucleotide addition can also be achieved by terminating the first extension step with a click-labeled nucleotide (e.g., azide or alkyne), or similarly, by oNTP-directed adaptor processing followed by chemical ligation with a reverse-functionalized surface-barcoded oligonucleotide (or its complementary sequence on the 3' end of the cDNA transcript) (Figure 4B).

[0261] It is contemplated that the added 3' oligonucleotide sequence can be captured by a surface oligo, such as a polyA tail or other capture sequence. These modified nucleotides (click and oNTP) can also be used to terminate the cDNA product to the appropriate insert length for sequencing. Another method is to polyadenylate the extended probe using TdT (or other mononucleotide addition) and attach this product to a polyT at the 3' end of the spatial barcode oligo.

[0262] Template switching can be another method for adding a polyA tail or other capture sequence to the 3' end of a first-strand cDNA molecule (Figure 4C, Figure 10). Similar to the method described above, the RNA capture oligonucleotide of an RNA capture probe hybridizes to RNA in a tissue sample to form an RNA-RNA capture probe hybrid. The RNA-RNA hybrid is extended using RT to generate a first-strand cDNA. To add the 3' oligonucleotide, the first-strand cDNA molecule is contacted with a reverse transcriptase (RT) and a template switch oligonucleotide (TSO), where the RT incorporates a non-template cytosine nucleotide at the 3' end of the first cDNA, the TSO contains a sequence capable of hybridizing to the non-template cytosine nucleotide, and the RT extends to generate a TSO complement. In this example, the 3'-terminal oligonucleotide comprises a substrate capture oligonucleotide complementary to a first domain of multiple substrate capture probes on a substrate, each of which contains, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, and a first domain.

[0263] Once the 3' oligonucleotide is added by template switching, the 3' terminal oligo is used to hybridize to the spatially barcoded surface oligo, and subsequent extension of the surface oligo links the spatial barcode to the first strand cDNA sequence. In one variation, DNA is dC-tailed by reverse transcription and used as a capture sequence, hybridizing to the dG terminal sequence on the spatially barcoded substrate capture oligo.

[0264] In another variation of template switching, spatially barcoded oligos on the substrate surface can be released to function as template-switching primers (Figure 4D, Figure 11). In this exemplary method, the 3'-terminal oligonucleotide can comprise a substrate capture oligonucleotide complementary to the first domain of a substrate capture probe on the substrate, each of which comprises, in a 5' to 3' orientation, a substrate anchor sequence, a second handle, a spatial barcode, and a first domain. The surface capture probes are released from the substrate and function as template-switching primers, which are then useful for spatially barcoding first-strand cDNA. Release of surface capture probes from the surface can be useful for spatially addressing tissues mounted on slides that do not have capture oligos.

[0265] Downstream library preparation steps on the ligated surface can include random priming of the second strand using random primers containing P5 sequences or similar handles, ligation of P5 adapters, TdT-mediated polyadenylation, followed by PCR to introduce the P5 adapters, etc. P5 ends can also be introduced using template switching during RT extension. SMIs can also be introduced during library preparation by any of the above approaches.

[0266] In another approach, random priming is performed from pulled-down RNA. RNA, e.g., mRNA, is bound to the surface of a substrate by hybridization between a blocked probe (e.g., containing a 3' phosphate) and a surface capture oligo (shown with the 5' end as the free end for hybridization, i.e., 5'-flying, but can also be 3'-flying) immobilized on the substrate (Figure 5A, Figure 12). A second barcoded oligo (5'-immobilized) is also located on the substrate in close proximity to the substrate capture oligo. The RNA capture oligonucleotide with a 3'-OH blocked probe hybridizes with RNA in the tissue sample to form an RNA-RNA capture probe hybrid with a 5' single-stranded RNA region. The substrate capture oligonucleotide of the RNA-RNA capture probe hybridizes to the first domain of the substrate capture probe, and the 5' single-stranded RNA region of the RNA-RNA capture probe hybrid anneals to the random priming sequence of the barcoded substrate probe. Extension of the random priming sequences hybridized to the 5' single-stranded RNA is carried out using RT to form spatially barcoded first strand cDNA molecules.

[0267] As can be seen in Figure 5A, barcoded oligos are used to randomly prime RNA, thereby ligating spatial barcodes to RNA transcripts. The barcode oligos also contain P7 or other adapter and barcode-reading primer sites. Downstream library preparation steps, as described above, are used to generate second-strand cDNA. If the capture oligo is poly-T, random priming of poly-A mRNA can also be performed. This allows for copying portions of the mRNA away from the poly-A tail 3' end (i.e., perhaps within the coding region instead of the 3' UTR). This is something missing from standard poly-A capture spatial approaches.

[0268] Another version of this scheme has the capture oligo and barcode oligo linked via a linker that cannot be read through by the polymerase (Figure 5B, Figure 13). The advantage of this scheme is that more space on the substrate surface is available, allowing more complex probe sets to be used to pull down RNA in the sample.

[0269] Probe extension followed by ligation to a surface spatial barcode oligo is also contemplated as a method herein (Figure 6A, Figure 14). In this method, RNA is bound to a surface using an unblocked RNA capture probe, which includes an RNA capture oligonucleotide complementary to the RNA in the sample and a substrate capture oligonucleotide complementary to the first domain of multiple substrate capture probes. The substrate capture probe can include a first domain and a first substrate anchor sequence in a 5' to 3' orientation, and is adjacent to a barcoded substrate probe on the substrate, which includes a spatial barcode and a second substrate anchor sequence in a 5' to 3' orientation. The RNA-RNA hybrid can also be used to prime RT. For example, the RNA capture oligonucleotide of the captured RNA-RNA capture probe hybrid is extended using RT to form a first-strand cDNA molecule. The first-strand cDNA is ligated to the spatial barcode oligo. The difference between this approach and previous approaches is that splint ligation is not required to enhance concentration by forcing localization to the ligation target, and a P5 adapter can be introduced 5' appended to the probe. Optionally, ligation can be performed chemically using a 3' click nucleotide in the cDNA or via an oNTP incorporated at the 3' end, which acts as its own splint for ligation to the barcode oligo. Other chemical ligation methods, such as 5' OH to 3' phos using EDC or oNTP incorporation followed by ligation, can also be used.

[0270] Another variation of this method uses a blocked RNA capture probe and 3' polyadenylation of the RNA (e.g., using PAP) to allow polyA addition and extension (Figure 6B, Figure 15). The barcode oligonucleotide can contain a polyT sequence that binds to the polyA, and the capture oligo on the substrate can be, but is not necessarily, 5'-flying. The RNA capture probe oligonucleotide binds to the capture probe, facilitating ligation of the extended RT to the barcode oligonucleotide on the substrate surface.

[0271] Another method involves direct ligation of RNA to a surface spatial barcoded oligo (Figures 7 and 16). This method uses an RNA capture probe with a hairpin structure, including a DNA capture oligonucleotide complementary to the RNA in the sample and a substrate capture oligonucleotide complementary to the first domain of the substrate capture probe. The DNA capture oligonucleotide contains a single-stranded region, and each of the substrate capture probes can contain, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, a first domain, and a second domain, with the second domain containing at least one RNA nucleotide or nucleoside. RNA-RNA capture probe hybrids are formed, each of which contains a 5' single-stranded RNA terminal region. The 5' RNA can be ligated to the 3' DNA using T4 ligase. The RNA-RNA capture probe hybrid is captured by the substrate capture oligo in the substrate capture probe. The RNA can be captured using a probe that also binds to the surface barcode oligo. Hairpin probes can prevent excess probes from occupying surface sites, otherwise requiring more stringent washes or higher Tm surface capture oligos. The 5' single-stranded RNA can then be 5' phosphorylated, allowing 5' to 3' riboexonuclease to digest the overhanging RNA. The digested 5' RNA terminal region of the captured RNA-RNA capture probe hybrid is ligated to the second domain of a substrate capture probe for a DNA-RNA chimera, for example, using T4 ligase. The surface oligo may have several ribonucleotides at the 3' end. The DNA-RNA chimera can be converted to DNA by reverse transcription using a DNA random primer that may contain P5. 3' polyadenylation of the chimera can also be used to enable priming using polyT. If performed on FFPE tissue, the RNA can be decrosslinked. If this method is performed using fresh-frozen tissue, the tissue is permeabilized to release the RNA.

[0272] It is understood, therefore, that the present invention is not limited to the particular embodiments disclosed, but is intended to cover all modifications described above and / or shown in the accompanying drawings that are within the spirit and scope of the present invention as defined by the appended claims. Accordingly, only such limitations as are set forth in the appended claims should be placed on the present disclosure.

Claims

1. 1. A method for preparing an mRNA transcript expression library from a tissue sample, comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, the capture oligonucleotides comprising a first clustered sequence, a spatial barcode sequence (SBC), and a first universal adaptor sequence; b) contacting the tissue sample with i) a plurality of 5' gene-specific probes comprising a 5' gene-specific primer and a sequence complementary to the first universal adaptor sequence, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence, under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample; c) contacting the tissue sample of (b) with a ligation reagent such that the 5' gene-specific probes and 3' gene-specific probes hybridized to the mRNA transcripts in close proximity to each other are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts hybridized to the ligated gene-specific probe pairs, leaving behind the ligated gene-specific probe pair oligonucleotide sequences; e) capturing the ligated gene-specific probe pair oligonucleotide of (d) on the substrate by binding the sequence complementary to the first universal adaptor sequence in the 5' gene-specific probe to the first universal adaptor sequence of the capture oligonucleotide.

2. 1. A method for determining mRNA transcript expression in a tissue sample, comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, the capture oligonucleotides comprising a first clustered sequence, a spatial barcode sequence (SBC), and a first universal adaptor sequence; b) contacting the tissue sample with i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence and a 5' gene-specific primer, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence, under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample; c) contacting the tissue sample of (b) with a ligation reagent such that the 5' gene-specific probes and 3' gene-specific probes hybridized to the mRNA transcripts in close proximity to each other are ligated together to form one or more ligated gene-specific probe pairs; d) removing the mRNA transcripts that hybridized to the ligated gene-specific probe pair, leaving behind the ligated gene-specific probe pair oligonucleotide sequences; e) capturing the ligated gene-specific probe pair oligonucleotide of (d) on the substrate by binding the sequence complementary to the first universal adaptor sequence in the 5' gene-specific probe to the first universal adaptor sequence of the capture oligonucleotide.

3. The method of claim 1 or 2, wherein the 3' gene-specific probe comprises one or more ribobases.

4. 1. A method for preparing an mRNA transcript expression library from a tissue sample, comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, the capture oligonucleotides comprising a first clustered sequence, a spatial barcode sequence (SBC), and a first universal adaptor sequence; b) contacting the tissue sample with i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence and a 5' gene-specific primer, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence, under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample; contacting under conditions such that hybridization of the 5' gene-specific probe and the 3' gene-specific probe on the mRNA transcript results in a nucleotide gap between the hybridized molecules; c) contacting the tissue sample of (b) with nucleotide bases and a ligation reagent such that the nucleotide gap between the 5' and 3' gene specific probes hybridized to the mRNA transcripts is filled with nucleotide bases complementary to the mRNA transcripts, and the 5' and 3' gene specific probes are ligated together to form one or more ligated gene specific probe pairs; d) removing the mRNA transcripts that hybridized to the ligated gene-specific probe pair, leaving behind the ligated gene-specific probe pair oligonucleotide sequences; e) capturing the ligated gene-specific probe pair oligonucleotide sequence of (d) on the substrate by binding the sequence complementary to the first universal adaptor sequence in the 5' gene-specific probe to the first universal adaptor sequence of the capture oligonucleotide.

5. 1. A method for isolating mRNA transcript expression in a tissue sample, comprising: a) placing the tissue sample on a substrate comprising a plurality of capture oligonucleotides, the capture oligonucleotides comprising a first clustered sequence, a spatial barcode sequence (SBC), and a first universal adaptor sequence; b) contacting the tissue sample with i) a plurality of 5' gene-specific probes comprising a sequence complementary to the first universal adaptor sequence and a 5' gene-specific primer, and ii) a plurality of 3' gene-specific probes comprising a 3' gene-specific primer, a unique molecular index, and a second universal adaptor sequence, under conditions such that one or more of the 5' gene-specific probes and one or more of the 3' gene-specific probes hybridize to one or more mRNA transcripts in the tissue sample; contacting under conditions such that hybridization of the one or more 5' gene-specific probes and one or more 3' gene-specific probes on the mRNA transcripts results in a nucleotide gap between the hybridized molecules; c) contacting the tissue sample of (b) with nucleotide bases and a ligation reagent such that the nucleotide gap between the 5' and 3' gene specific probes hybridized to the mRNA transcripts is filled with nucleotide bases complementary to the mRNA transcripts, and the 5' and 3' gene specific probes are ligated together to form one or more ligated gene specific probe pairs; d) removing mRNA transcripts hybridized to said ligated gene-specific probe pairs, leaving behind ligated gene-specific probe pair oligonucleotide sequences; e) capturing the ligated gene-specific probe pair oligonucleotide sequence of (d) on the substrate by binding the sequence complementary to the first universal adaptor sequence in the 5' gene-specific probe to the first universal adaptor sequence of the capture oligonucleotide.

6. The method of claim 4 or 5, wherein the nucleotide gap is 1 to 50 or more nucleotides.

7. f) performing an extension reaction and PCR on the oligonucleotides of (e) to obtain PCR templates representing one or more mRNA transcripts in the tissue sample; g) eluting the PCR template; h) performing indexing PCR to generate a double-stranded PCR product comprising a first strand PCR product and a second strand complementary to the first strand PCR product.

8. 8. The method of claim 7, further comprising: (h) sequencing the PCR product; and (a) determining the location of the mRNA transcript in the tissue based on the spatial barcode of (a).

9. 9. The method of claim 7 or 8, wherein the double-stranded PCR product comprises a second clustered sequence on the second strand that is complementary to the first-strand PCR product, and optionally an index sequence.

10. The method according to any one of claims 1 to 9, wherein the 5' gene-specific probe and / or the 3' gene-specific probe is 10 to 50 nucleotides.

11. The method of any one of claims 1 to 10, wherein the first clustered sequence comprises a P7 sequence.

12. 12. The method of any one of claims 1 to 11, wherein the first universal adaptor sequence comprises GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 19).

13. 13. The method of any one of claims 1 to 12, wherein the second universal adaptor sequence comprises AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTG (SEQ ID NO: 20).

14. The method according to any one of claims 1 to 13, wherein the 5' gene-specific probe and / or the 3' gene-specific probe has a melting temperature (Tm) of about 50 to 55°C.

15. The method of any one of claims 1 to 14, wherein the capture oligonucleotide has a melting temperature (Tm) of about 40-42°C.

16. 16. The method of any one of claims 1 to 15, wherein step (b) is carried out at approximately 50-55°C.

17. 17. The method of any one of claims 1 to 16, wherein step (e) is carried out at approximately 40-42°C.

18. 18. The method of any one of claims 1 to 17, wherein contacting the tissue sample with the substrate correlates the location of capture sites on the substrate with locations in the tissue sample, the substrate comprising a plurality of capture sites comprising a plurality of capture probes immobilized on a surface, the capture probes comprising a spatially addressable region.

19. The method of any one of claims 1 to 18, wherein the sample is from a mammal.

20. The method of any one of claims 1 to 19, wherein the sample is from a human.

21. The method of any one of claims 1 to 20, wherein the tissue sample is a tumor biopsy.

22. The method of any one of claims 1 to 21, wherein the tissue sample is formalin-fixed paraffin-embedded (FFPE) tissue or fresh-frozen (FF) tissue.

23. 1. A method for identifying a genetic variation in a subject having or at risk of having a disease, comprising: i) generating a sample mRNA library from a tissue sample from said subject according to the method of any one of claims 1 to 21; i) comparing the genetic information from the sample mRNA library with a control mRNA library; and iii) identifying genetic variations in said sample mRNA library that are associated with said disease.

24. 24. The method of claim 23, wherein the disease is a genetic defect, cancer, an autoimmune disease, or a metabolic disorder.

25. 25. The method of claim 23 or 24, wherein the disease is cancer.

26. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide sequence complementary to RNA in the sample and a first substrate capture oligonucleotide complementary to a first domain of a plurality of splint oligonucleotides; (b) hybridizing the RNA capture oligonucleotide of the RNA capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (c) extending the RNA capture oligonucleotide of the RNA-RNA capture probe hybrid using a reverse transcriptase to form a plurality of first strand cDNA molecules, each of the first strand cDNA molecules comprising the RNA capture oligonucleotide and the first substrate capture oligonucleotide; (d) capturing the first strand cDNA molecule onto a substrate, the substrate comprising a plurality of substrate capture probes, each comprising a spatial barcode and a second substrate capture oligonucleotide complementary to a second domain of the splint oligonucleotide, wherein the capturing comprises hybridizing the splint oligonucleotide to the first substrate capture oligonucleotide of the first strand cDNA molecule and the second substrate capture oligonucleotide of the substrate capture probe; (e) ligating the captured first strand cDNA molecule to the substrate capture probe, thereby forming a spatially barcoded first strand cDNA molecule.

27. 27. The method of claim 26, wherein the substrate capture probe further comprises a substrate anchor portion.

28. 28. The method of claim 26 or 27, wherein the surface oligonucleotides further comprise a P7 adaptor and an RNA capture probe primer for reading the spatial barcode sequence.

29. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, the RNA capture probes comprising an RNA capture oligonucleotide complementary to RNA in the sample and a handle sequence; (b) hybridizing the RNA capture oligonucleotide of the RNA capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (c) extending the RNA capture oligonucleotide of the RNA-RNA capture probe hybrid using a reverse transcriptase to form a plurality of first strand cDNA molecules, each of the first strand cDNA molecules comprising the RNA capture oligonucleotide and the handle sequence; (d) adding a 3' terminal oligonucleotide to the 3' end of each first strand cDNA molecule, wherein the 3' terminal oligonucleotide comprises a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on the substrate, each of the plurality of substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, and the first domain; (e) hybridizing the substrate capture oligonucleotide of the first strand cDNA molecule to the first domain of the substrate capture probe; (f) extending the first domain of the hybridized substrate capture probe to form a plurality of spatially barcoded first strand cDNA molecules.

30. 30. The method of claim 29, wherein the handle sequence is a PCR handle sequence, a molecular identifier, a UMI, or any combination thereof.

31. 31. The method of claim 29 or 30, wherein the handle sequence is a P5 adaptor sequence.

32. The method of any one of claims 29 to 31, wherein the 3' terminal oligonucleotide is added by tagging.

33. 33. The method of any one of claims 29 to 32, wherein the 3' terminal oligonucleotide is added by click chemistry or oNTP-directed adaptor processing.

34. 34. The method of claim 33, wherein the 3' OH is added by terminating the extension reaction with a click-labeled nucleotide.

35. 35. The method of claim 34, wherein the click-labeled nucleotide is an azide- or alkyne-labeled oligonucleotide.

36. 36. The method of claim 34 or 35, wherein the extension reaction adds a polyA sequence to the 3' extension sequence.

37. 37. The method of any one of claims 29 to 36, wherein the first strand cDNA is captured with a poly-T sequence on a surface capture oligonucleotide.

38. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, the RNA capture probes comprising an RNA capture oligonucleotide complementary to RNA in the sample and a handle sequence; (b) hybridizing the RNA capture oligonucleotide of the RNA capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (c) extending the RNA capture oligonucleotide of the RNA-RNA capture probe hybrid using a reverse transcriptase to form a plurality of first strand cDNA molecules, each of the first strand cDNA molecules comprising the RNA capture oligonucleotide and the handle sequence; (d) contacting the first strand cDNA molecules with a reverse transcriptase (RT) and a template switch oligonucleotide (TSO), wherein the RT incorporates a non-template cytosine nucleotide at the 3' end of the first cDNA, and the TSO comprises a sequence capable of hybridizing to the non-template cytosine nucleotide, and adding a 3' terminal oligonucleotide to the 3' end of each first strand cDNA molecule via template switching, wherein the RT incorporates a non-template cytosine nucleotide at the 3' end of the first cDNA, and the TSO comprises a sequence capable of hybridizing to the non-template cytosine nucleotide, and the RT extends to generate a TSO complement; adding a 3' terminal oligonucleotide, wherein the 3' terminal oligonucleotide comprises a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the plurality of substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, and the first domain; (e) hybridizing the substrate capture oligonucleotide of the first strand cDNA molecule to the first domain of the substrate capture probe; (f) extending the first domain of the hybridized substrate capture probe to form a plurality of spatially barcoded first strand cDNA molecules.

39. 39. The method of claim 38, wherein the first domain is a poly-T sequence.

40. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, the RNA capture probes comprising an RNA capture oligonucleotide complementary to RNA in the sample and a handle sequence; (b) hybridizing the RNA capture oligonucleotide of the RNA capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (c) extending the RNA capture oligonucleotide of the RNA-RNA capture probe hybrid using a reverse transcriptase to form a plurality of first strand cDNA molecules, each of the first strand cDNA molecules comprising the RNA capture oligonucleotide and the handle sequence; (d) contacting the first strand cDNA molecules with a reverse transcriptase (RT) and a template switch oligonucleotide (TSO), wherein the RT incorporates a non-template cytosine nucleotide at the 3' end of the first cDNA, and the TSO comprises a sequence capable of hybridizing to the non-template cytosine nucleotide, and adding a 3' terminal oligonucleotide to the 3' end of each first strand cDNA molecule via template switching, wherein the RT incorporates a non-template cytosine nucleotide at the 3' end of the first cDNA, and the TSO comprises a sequence capable of hybridizing to the non-template cytosine nucleotide, and the RT extends to generate a TSO complement; adding a 3' terminal oligonucleotide, wherein the 3' terminal oligonucleotide comprises a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the plurality of substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a second handle, a spatial barcode, and the first domain; (e) releasing the substrate capture probe from the substrate; (f) hybridizing the substrate capture oligonucleotide of the first strand cDNA molecule to the first domain of the substrate capture probe; (g) contacting the first strand with a second strand synthesis mix comprising a TSO primer and extending the TSO primer using the first strand as a template to produce a second strand complementary to the first strand, the second strand comprising the TSO, a second cDNA complementary to the first cDNA, and second strand barcode information comprising a spatial barcode sequence complement (SBC') complementary to a spatial barcode sequence (SBC).

41. 41. The method of claim 40, wherein the first domain is a poly-G sequence that hybridizes with a poly-C sequence on the TSO.

42. 42. The method of claim 40 or 41, wherein the handle is a P5 sequence and the second handle is a P7 sequence.

43. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that bind to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide complementary to an RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, the RNA capture oligonucleotides complementary to the RNA being blocked at their 3' ends, each of the substrate capture probes comprising, in a 5' to 3' orientation, the first domain and a first substrate anchor sequence, and adjacent to one or more barcoded substrate probes on the substrate, each of the barcoded substrate probes comprising, in a 5' to 3' orientation, a second substrate anchor sequence, a spatial barcode, and a random priming sequence; (b) hybridizing the RNA capture oligonucleotide of the RNA capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid having a 5' single-stranded RNA region; (c) hybridizing the substrate capture oligonucleotide of the RNA-RNA capture probe hybrid to the first domain of the substrate capture probe; (d) hybridizing the 5' single-stranded RNA region of the RNA-RNA capture probe hybrid with the random priming sequence of the barcoded substrate probe; (e) using a reverse transcriptase to extend the random priming sequence hybridized to the 5' single-stranded RNA region to form a plurality of spatially barcoded first strand cDNA molecules.

44. 44. The method of claim 43, wherein the nucleotide sequence complementary to RNA in the sample is a poly-T oligonucleotide, a randomer, a semi-randomer, or a target-specific sequence.

45. 45. The method of claim 43 or 44, wherein the nucleotide sequence complementary to RNA in the sample is a poly-T oligonucleotide.

46. 44. The method of claim 43, wherein the RNA is removed from the sample.

47. 47. The method of claim 46, wherein the RNA is removed from the sample after extension to form first strand cDNA.

48. 48. The method of claim 47, wherein the RNA is removed by enzymatic or thermal methods.

49. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide complementary to RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the substrate capture probes comprising, in 5' to 3' orientation, a substrate anchor sequence, the first domain, a linker, a spatial barcode, and a random priming sequence; (b) hybridizing the RNA capture probe to the RNA in the tissue sample to form an RNA-RNA capture probe hybrid having a 5' single-stranded RNA region; (c) hybridizing the substrate capture oligonucleotide of the RNA-RNA capture probe hybrid to the first domain of the substrate capture probe; (d) hybridizing the 5' single-stranded RNA region of the RNA-RNA capture probe hybrid with the random priming sequence of the substrate capture probe; (e) using a reverse transcriptase to extend the random priming sequence hybridized to the 5' single-stranded RNA region to form a plurality of spatially barcoded first strand cDNA molecules.

50. 50. The method of claim 49, wherein the linker is a linker that cannot be read through by a polymerase.

51. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that bind to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide complementary to RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, each of the substrate capture probes comprising, in a 5' to 3' orientation, the first domain and a first substrate anchor sequence, and adjacent to at least one of a plurality of barcoded substrate probes on the substrate, each barcoded substrate probe comprising, in a 5' to 3' orientation, a spatial barcode and a second substrate anchor sequence; (b) hybridizing the RNA capture oligonucleotide of the RNA capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (c) capturing the RNA-RNA capture probe hybrid on the substrate by hybridizing a substrate capture oligonucleotide of the RNA-RNA capture probe hybrid to the first domain of the substrate capture probe; (d) extending the RNA capture oligonucleotide of the captured RNA-RNA capture probe hybrid using a reverse transcriptase to form a plurality of first strand cDNA molecules; (e) ligating each of the first strand cDNA molecules to a proximal barcoded substrate probe, thereby forming spatially barcoded first strand cDNA molecules.

52. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that bind to RNA in the tissue sample, each of the RNA capture probes comprising an RNA capture oligonucleotide complementary to an RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, the RNA capture oligonucleotides complementary to the RNA being blocked at their 3' ends, each of the substrate capture probes comprising, in a 5' to 3' orientation, the first domain and a first substrate anchor sequence, and adjacent to at least one of a plurality of barcoded substrate probes on the substrate, each barcoded substrate probe comprising, in a 5' to 3' orientation, a poly-T sequence, a spatial barcode, and a second substrate anchor sequence; (b) hybridizing the RNA capture oligonucleotide of the RNA capture probe to RNA in the tissue sample to form an RNA-RNA capture probe hybrid; (c) capturing the RNA-RNA capture probe hybrid on the substrate by hybridizing a substrate capture oligonucleotide of the RNA-RNA capture probe hybrid to the first domain of the substrate capture probe; (d) polyadenylating the RNA in the sample at the 3' end; (e) using a reverse transcriptase to extend the RNA capture oligonucleotide of the captured RNA-RNA capture probe hybrid to form a plurality of first strand cDNA molecules.

53. 53. The method of claim 52, wherein the polyadenylation is carried out using polyA polymerase.

54. 1. A method for preparing a spatially barcoded RNA library from a tissue sample, comprising: (a) contacting the tissue sample with a plurality of RNA capture probes that hybridize to RNA in the tissue sample, each of the RNA capture probes having a hairpin structure and comprising a DNA capture oligonucleotide complementary to RNA in the sample and a substrate capture oligonucleotide complementary to a first domain of a plurality of substrate capture probes on a substrate, the DNA capture oligonucleotide of the RNA capture probe comprising a single-stranded region, each of the substrate capture probes comprising, in a 5' to 3' orientation, a substrate anchor sequence, a spatial barcode, the first domain, and a second domain, the second domain comprising at least one RNA nucleotide or nucleoside; (b) hybridizing the RNA capture probes to the RNA in the tissue sample to form RNA-RNA capture probe hybrids, each of the RNA-RNA capture probe hybrids comprising a 5' single-stranded RNA terminal region; (c) capturing the substrate capture oligonucleotide of the RNA-RNA capture probe hybrid onto the substrate by hybridizing the substrate capture oligonucleotide of the RNA-RNA capture probe hybrid to the first domain of the substrate capture probe; (d) phosphorylating the 5' single-stranded RNA terminal region of the captured RNA-RNA capture probe hybrid and contacting the captured RNA-RNA capture probe hybrid with a 5' to 3' riboexonuclease to digest the phosphorylated 5' single-stranded RNA terminal region; (e) ligating the digested 5' RNA terminal region of the captured RNA-RNA capture probe hybrid to the second domain of the substrate capture probe to form a plurality of DNA-RNA chimeras on the substrate.

55. 55. The method of claim 54, wherein the ligation is performed using T4 ligase.

56. 55. The method of claim 54, wherein the RNA of the captured RNA-RNA capture probe hybrid is 5' phosphorylated prior to ligation.

57. 57. The method of claim 56, further comprising generating first strand cDNA from the plurality of DNA-RNA chimeras on the substrate.

58. 58. The method of claim 57, wherein the first strand cDNA can be hybridized from a surface and processed for sequencing.

59. 59. The method of any one of claims 54 to 58, wherein the reverse transcription is carried out using DNA random primers, optionally comprising P5 adapters.

60. 60. The method of any one of claims 26 to 59, wherein the cDNA extension template can be dehybridized from the RNA in the tissue by chemical, enzymatic, or thermal dehybridization.

61. 60. The method of any one of claims 29 to 59, wherein the cDNA extension template can be dehybridized from the RNA on the substrate by chemical, enzymatic, or thermal dehybridization.

62. 62. The method of claim 60 or 61, wherein the dehybridization step occurs before or after the capture step.

63. 63. The method of any one of claims 26 to 62, wherein the tissue sample is formalin-fixed paraffin-embedded (FFPE) tissue or fresh-frozen (FF) tissue.

64. 64. The method of claim 63, further comprising decrosslinking the FFPE sample, optionally wherein the decrosslinking is performed using TE buffer, pH 9.

65. The method of any one of claims 26 to 64, wherein the RNA capture probe is selected from the group consisting of a poly-T sequence, a poly-U sequence, a randomer, a semi-random sequence, or a target-specific probe.

66. 66. The method of claim 65, wherein the RNA capture probe is a poly-T sequence.

67. 67. The method of claim 65 or 66, wherein the RNA capture probe comprises at least 10 deoxythymidine residues.

68. 68. The method of claim 67, wherein the target-specific probes comprise a plurality of different target-specific RNA capture probe sequences.

69. 69. The method of claim 68, wherein the target-specific probe comprises at least 10 nucleotides complementary to a nucleotide sequence of a target RNA.

70. 70. The method of claim 68 or 69, wherein the RNA capture probe or surface capture probe is 8 to 80 nucleotides.

71. 71. The method of any one of claims 26 to 70, wherein the targeting probe is 8 to 80 nucleotides or 10 to 50 nucleotides.

72. 72. The method of any one of claims 26 to 71, wherein the tissue sample is permeabilized prior to contacting the tissue sample with the plurality of RNA capture probes.

73. 73. The method of any one of claims 26 to 72, wherein the tissue sample is treated with one or more blocking reagents prior to contacting the tissue sample with the plurality of RNA capture probes.

74. 74. The method of any one of claims 26 to 73, wherein the tissue sample is permeabilized and treated with one or more blocking reagents prior to contacting the tissue sample with the plurality of RNA capture probes.

75. 75. The method of any one of claims 26 to 74, wherein the substrate is a bead, a bead array, a spot array, a substrate comprising multiple wells, a flow cell, clustered particles arranged on the surface of a chip, a film, or a plate.

76. 76. The method of claim 75, wherein the substrate comprises a plurality of nanowells or microwells.

77. 77. The method of any one of claims 26-76, wherein the spatially barcoded first strand cDNA molecules are recovered by contacting the spatially barcoded first strand cDNA on the substrate with a DNA polymerase and one or more primers to generate spatially barcoded second strand cDNA that is complementary to the spatially barcoded first strand cDNA, and removing the spatially barcoded second strand cDNA from the substrate.

78. 78. The method of claim 77, wherein the one or more primers each comprise a random priming sequence.

79. 79. The method of claim 78, wherein the random priming sequence comprises nine random nucleotides.

80. 80. The method of claim 78 or 79, wherein the spatially barcoded second strand cDNAs each comprise a unique molecular identifier (UMI), the UMI comprising an endogenous sequence and an exogenous sequence, the exogenous sequence being complementary to the random priming sequence used to generate the second strand cDNAs, and the endogenous sequence being complementary to a first strand cDNA template sequence used to generate the second strand cDNAs.

81. 78. The method of Claim 77, wherein the one or more primers each comprise a molecular identifier barcode.

82. 78. The method of claim 77, wherein the one or more primers each comprise a UMI barcode.

83. 83. The method of any one of claims 77 to 82, wherein the spatially barcoded second strand cDNA is removed from the substrate by chemical or physical dehybridization.

84. 84. The method of any one of claims 77 to 83, wherein the anchor sequence comprises a cleavage site and wherein the spatially barcoded first and second strand cDNA hybrids are removed from the substrate by enzymatic cleavage at the cleavage site.

85. 85. The method of claim 84, wherein the cleavage site is a binding site for a restriction endonuclease.

86. 85. The method of Claim 84, wherein said anchor sequence comprises a cleavage site and said spatially barcoded first strand cDNA molecules are recovered by enzymatic cleavage at said cleavage site.

87. 87. The method of claim 86, wherein the cleavage site is a binding site for a restriction endonuclease.

88. 88. The method of any one of claims 77-87, further comprising sequencing at least a portion of a cDNA library to determine said spatial barcode sequence of each molecule.

89. 89. The method of Claim 88, further comprising determining the spatial locations of one or more cDNA molecules by correlating the spatial barcode sequences of said one or more cDNA molecules with the spatial locations of surface oligonucleotide molecules on the substrate containing corresponding spatial barcode sequences.

90. performing an extension reaction and PCR on the spatially barcoded first strand cDNA to obtain a PCR template comprising a first strand PCR product representing one or more RNA transcripts in the tissue sample; eluting the PCR template; performing indexing PCR to generate a double-stranded PCR product comprising the first strand PCR product and a second strand complementary to the first strand PCR product; 90. The method of any one of claims 26-89, further comprising indexing and sequencing the spatially barcoded first strand cDNA comprising:

91. 91. The method of Claim 90, further comprising sequencing the PCR product and determining the location of the RNA transcript in the tissue based on the spatial barcode of the first strand cDNA.

92. 92. The method of Claim 90 or 91, wherein the double-stranded PCR product comprises a second clustered sequence on the second strand that is complementary to the first-strand PCR product, and optionally an index sequence.

93. 93. The method of claim 91 or 92, wherein the PCR products are further processed by tagging to create a spatial transcriptomics library.

94. 94. The method of claim 93, wherein said tagging comprises tagging onto a substrate.

95. 95. The method of any one of claims 26 to 94, wherein the method uses the tissue sample to determine RNA expression in a single cell.

96. 96. The method of claim 95, wherein said method determines RNA expression in one or more intracellular components in said single cell.

97. 97. The method of claim 96, wherein the intracellular component is a cell nucleus, cytoplasm, or mitochondria.

98. 98. The method of any one of claims 1 to 97, wherein the substrate or the surface of the substrate comprises a material selected from glass, silicon, poly-L-lysine coated material, nitrocellulose, polystyrene, cyclic olefin copolymer (COC), cyclic olefin polymer (COP), polyacrylamide, polypropylene, polyethylene, or polycarbonate.

99. The method of any one of claims 26 to 98, wherein the RNA library is an mRNA library.