Materials and methods for preparation of a spatial transcriptomics library
Nanostructures enhance RNA capture and synthesis in FFPE and frozen tissue samples by facilitating in situ polyadenylation and cDNA generation, addressing fragmentation and degradation issues, thereby improving spatial transcriptomics efficiency.
Patent Information
- Application Number
- PCT/US2024/062010
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
Current spatial transcriptomics workflows for FFPE and frozen tissue samples face challenges due to RNA fragmentation, degradation, and crosslinking, resulting in low capture and conversion efficiency of mRNA, leading to poor quality and quantity of RNA and DNA for library preparation.
The use of nanostructures comprising nano-scaffolds, oligonucleotides, and cleavable linkers for in situ polyadenylation and cDNA synthesis, enabling enhanced capture and conversion of fragmented RNA from tissue samples, with methods allowing for spatially barcoded RNA library generation.
Improves RNA capture and synthesis quality, generating multiple copies of cDNA from single RNA strands, enhancing the overall efficiency and resolution of spatial transcriptomics analysis.
Smart Images

Figure 00000061_0000 
Figure 00000062_0000 
Figure 00000063_0000
Abstract
Description
MATERIALS AND METHODS FOR PREPARATION OF A SPATIAL TRANSCRIPTOMICS LIBRARYCROSS REFEENCE TO RELATED APPLICATIONS
[0001] The present application claims the priority benefit of U.S. Provisional Patent Application No. 63 / 616,003, filed December 29, 2023, herein incorporated by reference in its entirety.INCORPORATION BY REFERENCE OF SEQUENCE DISCLOSURE
[0002] The Sequence Listing, which is a part of the present disclosure, is submitted concurrently with the specification as a computer readable file. The name of the file containing the Sequence Listing is “IP-2622_SeqListing.xml", which was created on December 23, 2024, and is 16,952 bytes in size. The subject matter of the Sequence Listing is incorporated herein in its entirety by reference.FIELD OF THE DISCLOSURE
[0003] The present disclosure relates, in general, to improved methods for preparing RNA from a tissue sample and preparation of a spatial transcriptomics library from the isolated RNA.BACKGROUND OF THE DISCLOSURE
[0004] Spatial transcriptomics enables highly multiplexed, spatially localized gene expression analysis from fresh frozen and formalin-fixed paraffin-embedded (FFPE) tissue samples. However, due to the freezing process or the fixation process of FFPE tissue, fragmentation, degradation, and crosslinking can alter the quality and quantity of RNA and DNA for transcriptomics library preparation. Current on-market spatial workflows capture and convert <1% mRNA within a tissue section.SUMMARY OF THE DISCLOSURE
[0005] Presented here are methods to generate higher RNA capture and spatial library conversion from preserved tissue samples, e.g., frozen or FFPE tissue samples. In situ polyadenylation can enable capture of fragmented FFPE RNA on oligo-dT surface. Also provided herein are improved methods to synthesize cDNA from isolated mRNA transcripts to improve the overall synthesis and alignment quality of the mRNA sequences and preparation of a spatial transcriptomics library.
[0006] In one aspect, the disclosure provides a method for preparing a spatially barcoded RNA library from a tissue sample comprising, (a) permeabilizing the tissue sample; (b) contacting the tissue sample with a plurality of nanostructures, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety; (c) capturing RNA from the tissue sample by hybridization of the RNA with RNA capture probes on the nanostructures; (d) capturing the plurality of nanostructures comprising captured RNA on a substrate, wherein the substrate comprises a plurality of capture sites and a plurality of surface oligonucleotide molecules, wherein each of the plurality of capture sites is capable of binding to the substrate anchor moiety of the nanostructure, and wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules; (e) contacting the captured RNA with a reverse transcriptase to generate one or more first strand cDNAs wherein each of the first strand cDNAs is contiguous with one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe; (f) cleaving the cleavable linkers to release the first strand cDNAs from the nanoscaffold; and (g) capturing the first strand cDNAs on the substate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
[0007] In various embodiments, the step of generating the one or more first strand cDNAs occurs in the tissue, and before capturing the plurality of nanostructures on the substrate. In various embodiments, the step of generating the one or more first strand cDNAs occurs after capturing the plurality of nanostructures on the substrate.
[0008] In various embodiments, the method further comprises treating the nanostructures with an RNase after the step of generating the one or more cDNAs.
[0009] In various embodiments, the method further comprises, prior to the step of capturing RNA from the tissue sample, the step of performing end repair of the RNA with polynucleotide kinase. In various embodiments, the method further comprises, prior to the step of capturing RNA from the tissue sample, the step of performing in situ polyadenylation with polyadenylate polymerase. In various embodiments, the method further comprises, prior to the step of capturing RNA from the tissue sample, the steps of performing end repairof the RNA with polynucleotide kinase followed by performing in situ polyadenylation with polyadenylate polymerase.
[0010] In various embodiments, the RNA comprises ribosomal RNA (rRNA), messenger RNA (mRNA), non-coding RNA (ncRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and / or microRNA (miRNA).
[0011] In another aspect, the disclosure provides a method for preparing a spatially barcoded RNA library from a tissue sample comprising, (a) contacting a tissue sample with a substrate comprising a plurality of nanostructures and a plurality of surface oligonucleotide molecules attached to the substrate, wherein each nanostructure comprises: (i) a nanoscaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; and; (iii) two or more cleavable linkers wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, wherein the nanoscaffold comprises two or more sites for attachment of the cleavable linkers; wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules; (b) permeabilizing the tissue sample to release RNA from the tissue sample; (c) capturing RNA from the permeabilized tissue sample by hybridization of the RNA with the RNA capture probe; (d) contacting the captured RNA with a reverse transcriptase to generate one or more first strand cDNAs wherein each of the first strand cDNAs is contiguous with one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe; (e) cleaving the cleavable linkers to release the first strand cDNAs from the nano-scaffold; and (f) capturing the first strand cDNAs on the substrate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
[0012] In a further aspect, contemplated herein is a method for preparing a spatially barcoded RNA library from a tissue sample comprising, (a) permeabilizing the tissue sample to release RNA from the tissue sample; (b) hybridizing the RNA with a plurality of RNA capture probes to generate a plurality of RNA-capture probe hybrids; (c) contacting the RNA-capture probe hybrids with a plurality of nanostructures, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, and (iv) a substrate anchor moiety, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavablelinkers; (d) covalently attaching the 5’ capture probe ends of the RNA-capture probe hybrids to the 3’ ends of the oligonucleotides; (e) contacting the covalently attached RNA-capture probe hybrids in step (d) with a reverse transcriptase to generate one or more first strand cDNAs, wherein each of the first strand cDNAs is contiguous with the one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe; (f) capturing the plurality of nanostructures on a substrate, wherein the substrate comprises a plurality of capture sites and a plurality of surface oligonucleotide molecules, wherein each of the plurality of capture sites is capable of binding to the substrate anchor moiety, and wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules; (g) cleaving the cleavable linkers to release the first strand cDNAs from the nano-scaffold; and (h) capturing the first strand cDNAs on the substrate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
[0013] In various embodiments, 5’ capture probe ends of the RNA-RNA capture probe hybrids are covalently attached to 3’ ends of the oligonucleotides on the nanostructure by ligation, splint ligation, or click chemistry.
[0014] Optionally, in various embodiments, the method comprises covalently attaching the 3’ capture probe ends of the RNA-capture probe hybrids to the 5’ ends of the oligonucleotides, or optionally 3’ capture probe ends of the RNA-RNA capture probe hybrids are covalently attached to 5’ ends of the oligonucleotides on the nanostructure by ligation, splint ligation, or click chemistry.
[0015] In various embodiments, the click chemistry uses azide-alkyne cycloaddition, a heterobifunctional linking group, a sulfo-NHS ester, a DBCO-NHS ester, a maleimide group, aldehydes with amines, hydrazides or aminooxy groups to form imines, hydrazones, or oximes. In various embodiments, the click chemistry covalent attachment is carried out by azide-alkyne cycloaddition.
[0016] In various embodiments, the attachment is by enzymatic ligation with DNA ligase and splint-assisted ligation, or by thermostable 5' App DNA / RNA ligase with a synthesized pre-adenylated single-stranded probe adapter. In various embodiments, the attachment is by 1 -ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC)-mediated ligation, e.g., in which probes with 3'phosphate groups are ligated to splinted adapters with 5' hydroxyl termini.
[0017] In various embodiments, the tissue sample is a fixed tissue sample. In various embodiments, the fixed tissue sample is a formalin-fixed paraffin embedded (FFPE) tissue sample. In various embodiments, the sample is a fresh frozen tissue sample.
[0018] In various embodiments, the RNA is released from the sample. In various embodiments, releasing comprises contacting the sample with a lysis buffer, a permeabilization buffer and / or a reagent to deparaffinize a FFPE sample. The method may further comprise decrosslinking the FFPE sample, optionally wherein the decrosslinking is carried out using TE buffer, pH 9.
[0019] In various embodiments, the tissue sample is treated with one or more blocking reagents prior to the permeabilization step. In various embodiments, the tissue sample is permeabilized and treated with one or more blocking reagents.
[0020] In various embodiments, the substrate is a bead, a bead array, a spotted array, a substrate comprising a plurality of wells, a flow cell (e.g., a clustered flow cell), clustered particles arranged on a surface of a chip, a film, or a plate (e.g., a multi-well plate). In various embodiments, the substrate is a gel coating located in or on a flow cell.
[0021] In various embodiments, the substrate comprises a plurality of nanowells or microwells. In various embodiments, the substrate is a gel coating located in or on a flow cell.
[0022] In various embodiments, the nano-scaffold is a dendrimer, a nanoparticle, a nanogel or a hyperbranched polymer. In various embodiments, the nano-scaffold comprises between 2 and 512 attachment sites or more. In various embodiments, the nano-scaffold comprises between 2 and 32, between 2 and 64, between 2 and 128 or between 2 and 256 attachment sites for the cleavable linkers.
[0023] In various embodiments, the nanostructure is a dendrimer. In various embodiments, the dendrimer is a peptide dendrimer comprising an alkyne end for anchoring to a substrate and an amine terminus. In various embodiments, the dendrimer comprises polylysine, branched lysine, or a polyamine.
[0024] In various embodiments, the RNA capture probe is selected from the group consisting of a poly-T sequence, a poly-U sequence, a randomer, a semi-random sequence, or a target-specific probe. In various embodiments, the RNA capture probe is a poly-T sequence. In various embodiments, the RNA capture probe comprises at least 10 deoxythymidine residues. In various embodiments, the target-specific probes comprise a plurality of different target-specific RNA capture probe sequences.
[0025] In various embodiments, the target-specific probes comprise a plurality of different target-specific RNA capture probe sequences. In various embodiments, the target-specific probes comprise at least 10 nucleotides complementary to a nucleotide sequence of a target RNA. In various embodiments, the RNA capture probe or surface capture probe is between 8 to 80 nucleotides. In certain embodiments, the RNA capture probe or surface probe is between 10 to 80 nucleotides, between 10 to 70 nucleotides, between 10 to 60 nucleotides, between 10 to 50 nucleotides, between 10 to 40 nucleotides, between 10 to 30 nucleotides, between 10 to 20 nucleotides, between 20 to 80 nucleotides, between 20 to 70 nucleotides, between 20 to 60 nucleotides, between 20 to 50 nucleotides, between 20 to 40 nucleotides, or is 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70 or 80 nucleotides.
[0026] In various embodiments, the cleavable linker is a cleavable polynucleotide. In various embodiments, the cleavable polynucleotide is between 5 to 25 nucleotides, or is 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, or 25 nucleotides.
[0027] In various embodiments, the surface oligonucleotide molecules further comprise a primer binding nucleotide sequence. In various embodiments, the primer binding sequence comprises a P7 nucleotide sequence.
[0028] In various embodiments, the sequence that hybridizes with the second domain of the splint oligonucleotide molecules comprises a PZ nucleotide sequence. In various embodiments, the second domain of the splint oligonucleotide molecules comprises a nucleotide sequence PZ’ that is complementary to the PZ sequence. In various embodiments, the sequence that hybridizes with the first domain of the splint oligonucleotide molecules comprises a PX nucleotide sequence. In various embodiments, the first domain of the splint oligonucleotide molecules comprises a nucleotide sequence PX’ that is complementary to the PX sequence. In various embodiments, the tissue sample is contacted with a mechanism to accelerate RNA diffusion from the sample. In various embodiments, the mechanism is magnetic or electrophoretic acceleration. If the surface is a 3D patterned surface, the mechanism may also include capture surfaces with pillars or other features extending outward from the surface. In various embodiments, additives to enhance crowding to surface are used to accelerate RNA diffusion. In various embodiments, the additives include PEG, ficoll, dextran sulfate, and the like.
[0029] In various embodiments, each of the two or more oligonucleotides further comprise a single molecular identifier (SMI) barcode. In various embodiments, each of the two or more oligonucleotides further comprise a unique molecular identifier (UMI) barcode.
[0030] In various embodiments, each of the surface oligonucleotide molecules further comprises a SMI barcode. In various embodiments, each of the surface oligonucleotide molecules further comprises a UM I barcode.
[0031] In various embodiments, the methods further comprise determining spatial locations of the spatial barcodes of the plurality of surface oligonucleotide molecules prior to the step of contacting the tissue with the substrate.
[0032] In various embodiments, the methods further comprise sequencing at least a portion of the spatially barcoded first strand cDNA molecules or copies thereof to determine the spatial barcode sequence for each molecule.
[0033] In various embodiments, the spatially barcoded first strand cDNA molecules are sequenced in situ.
[0034] In various embodiments, the methods further comprise determining the spatial location of one or more of the spatially barcoded first strand cDNA molecules or copies thereof by correlating the spatial barcode sequences of the spatially barcoded first strand cDNA molecules or copies thereof with the spatial locations of the surface oligonucleotide molecules on the substrate containing corresponding spatial barcode sequences.
[0035] In various embodiments, the methods further comprise recovering the spatially barcoded first strand cDNA molecules and amplifying them to generate cDNA libraries.
[0036] In various embodiments, the spatially barcoded first strand cDNA molecules are recovered by contacting the spatially barcoded first strand cDNAs on the substrate with a DNA polymerase and one or more primers to generate spatially barcoded second strand cDNAs complementary to the spatially barcoded first strand cDNAs and removing the spatially barcoded second strand cDNAs from the substrate.
[0037] In various embodiments, the one or more primers each comprise a random priming sequence. In various embodiments, the random priming sequences comprises nine random nucleotides.
[0038] In various embodiments, the spatially barcoded second strand cDNAs each comprise a unique molecular identifier (UMI), wherein the UMI comprises an intrinsic sequence and an extrinsic sequence, wherein the extrinsic sequence is a sequence complementary to the random priming sequence used to generate the second strand cDNA, and wherein the intrinsic sequence is a sequence complementary to the first strand cDNA template sequence used to generate the second strand cDNA.
[0039] In various embodiments, the one or more primers each comprise a molecular identifier barcode. In various embodiments, the one or more primers each comprise a UMI barcode.
[0040] In various embodiments, the spatially barcoded second strand cDNAs are removed from the substrate by chemical or physical dehybridization.
[0041] In various embodiments, the anchor sequence comprises a cleavage site, and hybrids of the spatially barcoded first and second strand cDNAs are removed from the substrate by enzymatic cleavage at the cleavage site. In various embodiments, the cleavage site is a binding site for a restriction endonuclease. In various embodiments, the anchor sequence comprises a cleavage site, and wherein the spatially barcoded first strand cDNA molecules are recovered by enzymatic cleavage at the cleavage site. In various embodiments, the cleavage site is a binding site for a restriction endonuclease.
[0042] In various embodiments, the methods further comprise sequencing at least a portion of the cDNA libraries to determine the spatial barcode sequence for each molecule.
[0043] In various embodiments, the methods further comprise determining the spatial location of one or more cDNA molecules by correlating the spatial barcode sequences of the one or more cDNA molecules with the spatial locations of the surface oligonucleotide molecules on the substrate containing corresponding spatial barcode sequences.
[0044] In various embodiments, the methods further comprise indexing and sequencing spatially barcoded first strand cDNAs, the method comprising, performing extension reactions and PCR on the spatially barcoded first strand cDNAs to yield a PCR template comprising a first strand PCR product representative of one or more RNA transcripts in the tissue sample; eluting the PCR template; carrying out an indexing PCR to generate a double stranded PCR product comprising the first strand PCR product and a second strand complementary to the first strand PCR product.
[0045] In various embodiments, the methods further comprise sequencing the PCR product and determining the location of the RNA transcript in the tissue based on the spatial barcode of first strand cDNA.
[0046] In various embodiments, the double stranded PCR product comprises a second clustering sequence on the second strand complementary to the first strand PCR product and, optionally, an index sequence.
[0047] In various embodiments, the PCR products are further processed by tagmentation to generate a spatial transcriptomics library. In various embodiments, the tagmentation comprises on substrate tagmentation. In some embodiments, the tagmentation comprises onbead tagmentation, wherein the bead comprises a plurality of bead-linked transposomes (BLT). In some embodiments, the BLT comprises i) a plurality of oligonucleotides comprising a first clustering sequence (P7), a first index sequence and a Read 1 sequencing primer (Rd1 SP) and ii) a plurality of oligonucleotides comprising a second clustering sequence (P5), a second index sequence and a Read 2 sequencing primer (Rd2 SP).
[0048] In various embodiments, the RNA library is an mRNA library.
[0049] In various embodiments, the methods determine RNA expression in a single cell with the tissue sample. In various embodiments, the methods determine RNA expression in one or more subcellular components in the single cell. In various embodiments, the subcellular component is a cell nucleus, cytoplasm, or mitochondria.
[0050] In various embodiments, the substrate or surface of the substrate comprises a material selected from glass, silicon, poly-L-lysine coated materials, nitrocellulose, polystyrene, cyclic olefin copolymers (COCs), cyclic olefin polymers (COPs), polyacrylamide, polypropylene, polyethylene, or polycarbonate.
[0051] The disclosure further provides a composition comprising a nanostructure, each nanostructure comprising: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nanoscaffold, and wherein the nanoscaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety.
[0052] Also provided is a composition comprising a nanostructure, wherein the nanostructure comprises: (i) a nanoscaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nanoscaffold, and wherein the nanoscaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety, and wherein the nanostructure is attached to a substrate.
[0053] Further contemplated is a composition comprising a nanostructure, wherein the nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nanoscaffold, and wherein the nano-scaffold comprises two or more sites for attachment of thecleavable linkers; and (iv) a substrate anchor moiety, and wherein the nanostructure is attached to a substrate comprising surface oligonucleotide molecules.
[0054] In various embodiments, the substrate contains between 1 x 109and 1 x 1011per mm2surface oligonucleotide molecules, e.g., capture oligonucleotides. In various embodiments, the substrate contains between 1 x 109and 1 x 1011per mm2, between 5 x 109and 5 x 1010per mm2, between 1 x 109and 1 x 1010per mm2, between 1 x 1010and 5 x 1 O10per mm2, or between 1 x 1 O10and 3 x 1 O10mm2surface oligonucleotide molecules. In various embodiments, the density is between about 100k / mm2to about 1000k / mm2, e.g., about 100k clusters / mm2, about 200k clusters / mm2, about 300k clusters / mm2, about 400k clusters / mm2, about 500k clusters / mm2, about 600k clusters / mm2, about 700k clusters / mm2, about 800k clusters / mm2, about 900k clusters / mm2, or about 1000k clusters / mm2. In various embodiments, the surface oligonucleotide molecules are arranged in patterns or clusters.
[0055] In various embodiments, the ratio of surface oligonucleotide molecules to capture sites is from about 1 :2 to about 1 :100 or more, e.g., about 1 :5, about 1 :10, about 1 :15, about 1 :20, about 1 :25, about 1 :30, about 1 :35, about 1 :40, about 1 :45, about 1 :50, about 1 :55, about 1 :60, about 1 :65, about 1 :70, about 1 :75, about 1 :80, about 1 :85, about 1 :90, about 1 :95, or about 1 :100.
[0056] In various embodiments, the diameter of each cluster is from about 500 nm to about 2,000 nm, e.g., about 500 nm, about 600 nm, about 700 nm, about 800 nm, about 900 nm, about 1000 nm, about 1 100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about 2000 nm.
[0057] In various embodiments, the surface oligonucleotide molecules comprise a cleavage domain. In various embodiments, the cleavage domain is a binding site for a restriction endonuclease or a site for chemical cleavage.
[0058] In another aspect, the disclosure provides a kit comprising a nanostructure, wherein the nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety, a substrate, and splint oligonucleotide molecules.
[0059] The disclosure further contemplates a kit comprising a nanostructure, wherein the nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with afirst domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nanoscaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety, and splint oligonucleotide molecules.
[0060] It is understood that each feature or embodiment, or combination, described herein is a non-limiting, illustrative example of any of the aspects of the invention and, as such, is meant to be combinable with any other feature or embodiment, or combination, described herein. For example, where features are described with language such as “one embodiment”, “various embodiments”, “some embodiments”, “certain embodiments”, “further embodiment”, “specific exemplary embodiments”, and / or “another embodiment”, each of these types of embodiments is a non-limiting example of a feature that is intended to be combined with any other feature, or combination of features, described herein without having to list every possible combination.
[0061] Such features or combinations of features apply to any of the aspects of the invention. Where examples of values falling within ranges are disclosed, any of these examples are contemplated as possible endpoints of a range, any and all numeric values between such endpoints are contemplated, and any and all combinations of upper and lower endpoints are envisioned.BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 . Schematic of an exemplary RNA library preparation workflow using nanostructures as described herein.
[0063] Figure 2. Schematic of an alternate RNA library preparation workflow using nanostructures as described herein.
[0064] Figure 3. Schematic of an alternate RNA library preparation workflow using nanostructures as described herein.
[0065] Figures 4A-4D. Schematic illustrating use of dendrimer nanostructures in capturing RNA from a tissue sample
[0066] Figure 5A shows different workflows for using a nanostructure as described herein. Figure 5B illustrates possible chemical structure for nano-scaffolds.
[0067] Figure 6A illustrates multivalent capture probes on a bead. Figure 6B shows binding efficiency and melting temperature for beads comprising different types of captures probes as in Fig. 6A.DETAILED DESCRIPTION
[0068] Isolating RNA (e.g., mRNA or rRNA) from preserved tissue samples and converting RNA to cDNA on a flat surface presents a number of problems, including lower quality RNA transcripts isolated from the tissue samples, shorter synthesized cDNA fragments (<450bp) in library preparation products and a high percentage of polyA presence in cDNA regions in the final sequencing products. These issues result in a subsequent low mapping rate to exonic mRNA transcript regions in RNA-seq alignment.
[0069] To solve this problem, it was hypothesized that an improved method to generate higher capture and spatial library conversion from FFPE tissue samples was needed.Definitions
[0070] Unless otherwise stated, the following terms used in this application, including the specification and claims, have the definitions given below.
[0071] As used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to "a capture probe" includes a mixture of two or more capture probes, and the like.
[0072] The term "about," particularly in reference to a given quantity, is meant to encompass deviations of plus or minus five percent.
[0073] As used herein, the terms "includes," "including," "includes," "including," "contains," "containing," and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, product-by-process, or composition of matter that includes, includes, or contains an element or list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, product-by-process, or composition of matter.
[0074] As used herein, a “nanoscaffold” refers to a multi-armed molecule containing multiple sites, i.e., is multivalent, for attachment of linkers and or oligonucleotide sequences. The nano-scaffold may also comprise a single anchor site for attachment of the nanoscaffold to a substrate. The nano-scaffold may comprise one or more anchor site(s) for attachment of the nano-scaffold to a substrate. Examples of nano-scaffolds contemplated herein include, but are not limited to, dendrimers, nanoparticles, nanogels, or hyperbranched polymers. The nano-scaffold comprises between 2 and 32 attachments sites for the cleavable linkers, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 or 32 attachment sites. It is contemplated that each siteprovides an attachment site for an oligonucleotide, optionally attached via a linker, which provides a hybridization site or binding site for a target nucleic acid.
[0075] As used herein a “nanostructure” refers to a complex of a nanoscaffold with two or more oligonucleotides optionally linked to the nano-scaffold via cleavable linkers. A nanostructure can range in size from about 10 nm to about 1000 nm, about 20 nm to 200 nm, or about 50 nm to about 500 nm, e.g., about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 150 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 700 nm, about 800 nm, about 900 nm, or about 1000 nm. In one embodiment, the nanostructure is about 50 nm.
[0076] As used herein a “linker” refers to a moiety that attached an oligonucleotide to the nano-scaffold. A linker may be cleavable or covalent. The linker may be a chemical moiety, peptide or cleavable polynucleotide. It is contemplated that the linker provides flexibility and length away from the nano-scaffold so the oligonucleotides can properly hybridize to a target sequence or other nucleotide. A polynucleotide linker may be between 4-20 or between 5- 25 nucleotides. In varus embodiments, the spacer is a polyethylene glycol (PEG) chain. In various embodiments, the PEG comprises between 3 to 50 or more repeating units. In various embodiments, the linker is a polymer of amino acids, e.g., polylysine. In various embodiments, the polypeptide is between 3 and 50 or more repeating units.
[0077] As used herein an “anchor” refers to a moiety that attaches a nano-scaffold to a substrate. An anchor includes a chemical moiety, peptide, or oligonucleotide. A polynucleotide anchor may be between 4-20 nucleotides.
[0078] As used herein a “splint oligonucleotide” refers to an oligonucleotide comprising a sequence complementary to a region on a surface probe on a nanostructure and another sequence complementary to a surface oligonucleotide, e.g., attached to a substrate. In various embodiments, the splint oligonucleotide is between 10-25 nucleotides or between 15-25 nucleotides. In various embodiments, the splint oligonucleotide is 20 nucleotides.
[0079] As used herein a “surface oligonucleotide” refers to an oligonucleotide comprising an anchor sequence for attaching the oligo to the surface of a substrate, a spatial barcode sequence and a sequence that hybridizes with a splint oligonucleotide. In various embodiments, the surface oligonucleotide is between 15-25 nucleotides. In various embodiments, the surface oligonucleotide is greater than or equal to 20 nucleotides. In various embodiment, the splint oligonucleotide is 15, 16, 17, 18, 9, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides or more.
[0080] As used herein, the terms "address," "tag," “barcode” or "index," when used in reference to a nucleotide sequence is intended to mean a unique nucleotide sequence that is distinguishable from other indices as well as from other nucleotide sequences within polynucleotides contained within a sample. A nucleotide "address," "tag," “barcode” or "index" can be a random or a specifically designed nucleotide sequence. An "address," "tag," “barcode” or "index" can be of any desired sequence length so long as it is of sufficient length to be unique nucleotide sequence within a plurality of indices in a population and / or within a plurality of polynucleotides that are being analyzed or interrogated. A nucleotide "address," "tag," “barcode” or "index" of the disclosure is useful, for example, to be attached to a target polynucleotide to tag or mark a particular species for identifying all members of the tagged species within a population. Accordingly, an index is useful as a barcode where different members of the same molecular species can contain the same index and where different species within a population of different polynucleotides can have different indices.
[0081] A tag / index / barcode sequence can be unique to a single nucleic acid species in a population or can be shared by several different nucleic acid species in a population. For example, each nucleic acid probe in a population can include different tag / index / barcode sequences from all other nucleic acid probes in the population. Alternatively, each nucleic acid probe in a population can include different tag / index / barcode sequences from some or most other nucleic acid probes in a population. For example, each probe in a population can have a tag / index / barcode that is present for several different probes in the population even though the probes with the common tag / index / barcode differ from each other at other sequence regions along their length. In particular embodiments, one or more tag / index / barcode sequences that are used with a biological specimen are not present in the genome, transcriptome or other nucleic acids of the biological specimen. For example, tag / index / barcode sequences can have less than 80%, 70%, 60%, 50% or 40% sequence identity to the nucleic acid sequences in a particular biological specimen.
[0082] As used herein, a "spatial address," "spatial tag", “spatial barcode”, “spatial barcode sequence” or "spatial index," when used in reference to a nucleotide sequence, means an address, tag, barcode or index encoding spatial information related to the region or location of origin of an addressed, tagged, barcoded or indexed nucleic acid in a tissue sample. The sequence can be a naturally occurring sequence or a sequence that does not occur naturally in the organism from which the barcoded nucleic acid was obtained.
[0083] As used herein, the term "substrate" is intended to mean a solid support or support structure. The term includes any material that can serve as a solid or semi-solid foundation for creation of features such as wells for the deposition of biopolymers, including nucleic acids, polypeptide and / or other polymers. Non-limiting examples of substrates include abead array, a spotted array, clustered particles arranged on a surface of a chip, a film, a multi-well plate, and a flow cell. A substrate as provided herein is modified, for example, or can be modified to accommodate attachment of biopolymers by a variety of methods well known to those skilled in the art. Exemplary types of substrate materials include glass, modified glass, functionalized glass, inorganic glasses, microspheres, including inert and / or magnetic particles, plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, an optical fiber or optical fiber bundles, a variety of polymers other than those exemplified above and multiwell microtiter plates. Specific types of exemplary plastics include acrylics, polystyrene, copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes and TEFLON™. Specific types of exemplary silica-based materials include silicon and various forms of modified silicon.
[0084] Those skilled in the art will know or understand that the composition and geometry of a substrate as provided herein can vary depending on the intended use and preferences of the user. Therefore, although planar substrates such as slides, chips wafers or beads are useful for microarrays, those skilled in the art will understand that a wide variety of other substrates exemplified herein or well known in the art also can be used in the methods and / or compositions herein.
[0085] In some embodiments, the solid support comprises one or more surfaces that are accessible to contact with reagents, beads, or analytes. The surface can be substantially flat or planar. Alternatively, the surface can be rounded or contoured. Example contours that can be included on a surface are wells (e.g., microwells or nanowells), depressions, pillars, ridges, channels or the like. Example materials that can be used as a surface include glass such as modified or functionalized glass; plastic such as acrylic, polystyrene or a copolymer of styrene and another material, polypropylene, polyethylene, polybutylene, polyurethane or TEFLON; polysaccharides or cross-linked polysaccharides such as agarose or Sepharose; nylon; nitrocellulose; resin; silica or silica-based materials including silicon and modified silicon, carbon-fiber; metal; inorganic glass; optical fiber bundle, or a variety of other polymers. A single material or mixture of several different materials can form a surface useful in certain examples. In some examples, a surface comprises wells e.g., microwells or nanowells). In some aspects, the surface comprises wells in an array of wells e.g., microwells or nanowells) on glass, silicon, plastic or other suitable solid supports with patterned, covalently-linked gel such as poly(N-(5-azidoacetamidylpentyl)acrylamide- coacrylamide) (PAZAM, see, for example, U.S. Pat. App. Pub. No. 2014 / 0079923 A1 , which is incorporated herein by reference). In some examples, a support structure can include oneor more layers. Non-limiting examples of a surface include a bead array, a spotted array, clustered particles arranged on a surface of a chip, a film, a multi-well plate, and a flow cell.
[0086] In some embodiments, the solid support comprises one or more surfaces of a flowcell. The term "flowcell" as used herein refers to a chamber comprising a solid surface across which one or more fluid reagents can be flowed. The flow cell can be an ordered or random flow cell. Examples of flowcells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), W004 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211 ,414; US 7,315,019; US 7,405,281 , and US 2008 / 0108082, each of which is incorporated herein by reference.
[0087] In some embodiments, the solid support includes a patterned surface. A "patterned surface" refers to an arrangement of different regions in or on an exposed layer of a solid support. For example, one or more of the regions can be features where one or more amplification primers are present. The features can be separated by interstitial regions where amplification primers are not present. In some embodiments, the pattern can be an x- y format of features that are in rows and columns. In some embodiments, the pattern can be a repeating arrangement of features and / or interstitial regions. In some embodiments, the pattern can be a random arrangement of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions set forth herein are described in US Ser. No. 13 / 661 ,524 or US Pat. App. Publ. No. 2012 / 0316086, or International Patent Publication WO 2017 / 019456, each of which is incorporated herein by reference.
[0088] As used herein, the term “immobilized” refers to the state of two things being joined, fastened, adhered, attached, connected, or bound to each other. For example, an analyte, such as a nucleic acid, can be immobilized on a material, such as a bead, gel, or surface, by a covalent or non-covalent bond. Immobilized in reference to a nucleic acid is intended to mean direct or indirect attachment to a solid support via covalent or non-covalent bond(s). In certain embodiments, covalent attachment can be used, but all that is required is that the nucleic acids remain stationary or attached to a support under conditions in which it is intended to use the support, for example, in applications requiring nucleic acid amplification and / or sequencing. Oligonucleotides to be used as capture primers or amplification primers can be immobilized such that a 3'-end is available for enzymatic extension and at least a portion of the sequence is capable of hybridizing to a complementary sequence.
[0089] Immobilization can occur via hybridization to a surface attached oligonucleotide, in which case the immobilized oligonucleotide or polynucleotide can be in the 3' -5' orientation. Alternatively, immobilization can occur by means other than base-pairing hybridization, such as the covalent attachment set forth above
[0090] Exemplary covalent linkages include, for example, those that result from the use of click chemistry techniques. Exemplary non-covalent linkages include, but are not limited to, non-specific interactions (e.g., hydrogen bonding, ionic bonding, van der Waals interactions etc.) or specific interactions (e.g., affinity interactions, receptor-ligand interactions, antibodyepitope interactions, avidin-biotin interactions, streptavidin-biotin interactions, lectincarbohydrate interactions, etc.). Exemplary linkages are set forth in U.S. Pat. Nos. 6,737,236; 7,259,258; 7,375,234 and 7,427,678; and US Pat. Pub. No. 2011 / 0059865 Al, each of which is incorporated herein by reference.
[0091] As used herein, the term "array" refers to a population of sites that can be differentiated from each other according to relative location. Different molecules that are at different sites of an array can be differentiated from each other according to the locations of the sites in the array. An individual site of an array can include one or more molecules of a particular type. For example, a site can include a single target nucleic acid molecule having a particular sequence or a site can include several nucleic acid molecules having the same sequence (and / or complementary sequence, thereof). The sites of an array can be different features located on the same substrate. Exemplary features include without limitation, wells in a substrate, beads (or other particles) in or on a substrate, projections from a substrate, ridges on a substrate or channels in a substrate. The sites of an array can be separate substrates each bearing a different molecule. Different molecules attached to separate substrates can be identified according to the locations of the substrates on a surface to which the substrates are associated or according to the locations of the substrates in a liquid or gel. Exemplary arrays in which separate substrates are located on a surface include, without limitation, those having beads in wells.
[0092] As used herein, the term “single molecular identifier” or “SMI” refers to a molecular tag, either random, non-random, or semi-random, that may be attached to a nucleic acid. In various embodiments, a SMI is a unique molecular identifier (UMI). When incorporated into a nucleic acid, a SMI can be used to correct for subsequent amplification bias by directly counting single molecular identifiers (SMIs) that are sequenced after amplification. A SMI e.g., a UMI) can be attached to similar nucleic acids, e.g., adapters, making each nucleic acid unique. SMIs e.g., UMIs) may also be used to uniquely tag individual molecules {e.g., individual mRNA molecules) in a sample {e.g., individual mRNA molecules in a tissue sample, cell sample, or sample library).
[0093] As used herein “unique molecular index”, “unique molecular identifier” or “IIMI”, when used in reference to a capture probe or other nucleic acid is intended to refer to a portion of a probe useful as a molecular barcode to uniquely tag each molecule in a sample library. A UMI may be denoted as “NNNN...” in a string of nucleic acids to designate that portion of the oligonucleotide as the UMI. A UMI may be from 6 to 20 nucleotides or more in length. In some aspects, the UMI comprises a spatial barcode.
[0094] As used herein, the term “universal sequence” refers to a series of nucleotides that is common to two or more nucleic acid molecules even if the molecules also have regions of sequence that differ from each other. A universal sequence that is present in different members of a collection of molecules can allow capture of multiple different nucleic acids using a population of universal capture nucleic acids that are complementary to the universal sequence. Similarly, a universal sequence present in different members of a collection of molecules can allow the replication or amplification of multiple different nucleic acids using a population of universal primers that are complementary to the universal sequence. Thus, a universal capture nucleic acid or a universal primer includes a sequence that can hybridize specifically to a universal sequence. Target nucleic acid molecules may be modified to attach universal adapters, for example, at one or both ends of the different target sequences. Universal capture oligonucleotides are applicable for interrogating a plurality of different oligonucleotides without necessarily distinguishing the different species whereas targetspecific capture sequences are applicable for distinguishing the different species. A nonlimiting example of a universal sequence is a polyT nucleotide sequence.
[0095] As used herein, a "semi-random" nucleotide sequence comprises or consists of a partially pre-determined nucleotide sequence combined with a random nucleotide sequence.
[0096] As used herein, the term “adapter” refers generally to any linear nucleic acid molecule that can be added (e.g., through synthesis or ligation) to an oligonucleotide of the disclosure. In some embodiments, adapters are copied onto the library molecules using templated polymerase synthesis e.g., second strand cDNA synthesis as described herein). In some embodiments, adapters are ligated to a first complementary strand of the disclosure. In some embodiments, oligonucleotides of the disclosure comprise adapters (“adapter oligonucleotides”). In some embodiments, an adapter oligonucleotide comprises from 5’ to 3’, a third sequencing primer sequence e.g., SBS3), a sequence complementary to a unique index sequence {e.g., i5’), and a second clustering primer sequence {e.g., P5). In some embodiments, an adapter comprises a sequence that is complementary to a primer. In further embodiments, an adapter comprises a sequence that is complementary to a P5 primer or a P5’ primer. In some embodiments, an adapter comprises a sequencecomplementary to a P7 primer or a P7’ primer. In some embodiments, an adapter comprises a sequence complementary to a B15 primer or a B15’ primer.
[0097] The terms “P5”, “P7”, “B15”, “P5”’ (P5 prime), “P7”’ (P7 prime), “B15”’ (B15 prime), “P15”, “P17”, “A14” and “A14”’ (A14 prime) may be used when referring to examples of oligonucleotide sequences of primers, e.g., clustering primers, and / or oligonucleotide sequences that are complementary to primers. The terms "P5"' (P5 prime), "P7"' (P7 prime), and “B15”’ (B15 prime) refer to the complement of P5, P7, and B15, respectively. It will be understood that any suitable primer can be used in the methods presented herein, and that the use of P5, P5’, P7, P7’, P15, P17, B15, and B15’ are exemplary embodiments only. Uses of primers such as P5, P5’, P7, P7’, P15, P17, B15, and B15’ or their complements on flow cells are known in the art, as exemplified by the disclosures of WO 2019 / 222264, WO 2007 / 010251 , WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941 , WO 1998 / 044151 , and WO 2000 / 018957, each of which is incorporated herein by reference in its entirety. For example, any suitable forward amplification primer, whether immobilized or in solution, can be useful in the methods presented herein for hybridization to a complementary sequence and amplification of a sequence. Similarly, any suitable reverse amplification primer, whether immobilized or in solution, can be useful in the methods presented herein for hybridization to a complementary sequence and amplification of a sequence. One of skill in the art will understand how to design and use primer sequences that are suitable for capture and / or amplification of nucleic acids as presented herein. In some embodiments, a “first clustering primer” as described herein is a P5 primer. In some embodiments, a “first clustering primer” as described herein is a P7 primer. In some embodiments, a “first clustering primer” as described herein is a P5' primer. In some embodiments, a “first clustering primer” as described herein is a P7' primer. In some embodiments, a “second clustering primer” as described herein is a P5 primer. In some embodiments, a “second clustering primer” as described herein is a P7 primer. In some embodiments, a “second clustering primer” as described herein is a P5' primer. In some embodiments, a “second clustering primer” as described herein is a P7' primer. In some embodiments, P5 comprises or consists of the polynucleotide sequence 5’ AAT GAT ACG GCG ACC ACC GA 3’ (SEQ ID NO: 1), or a variant thereof. In some embodiments, P5 comprises or consists of the polynucleotide sequence 5’ AAT GAT ACG GCG ACC ACC GAG ATC TAC AC 3’ (SEQ ID NO: 2), or a variant thereof. In some embodiments, P7 comprises or consists of the polynucleotide sequence 5’ CAA GCA GAA GAC GGC ATA CG 3’ (SEQ ID NO. 3), or a variant thereof. In some embodiments, P7 comprises or consists of the polynucleotide sequence 5’ CAA GCA GAA GAC GGC ATA CGA GAT 3’ (SEQ ID NO. 4), or a variant thereof. In some embodiments, P5' comprises or consists of the polynucleotide sequence 5’ TCG GTG GTCGCC GTA TCA TT 3’ (SEQ ID NO: 5), or a variant thereof. In some embodiments, P5' comprises or consists of the polynucleotide sequence 5’ GTG TAG ATC TCG GTG GTC GCC GTA TCA TT 3’ (SEQ ID NO: 6), or a variant thereof. In some embodiments, P7' comprises the polynucleotide sequence 5’ CGT ATG CCG TCT TCT GCT TG 3’ (SEQ ID NO. 7), or a variant thereof. In some embodiments, P7' comprises or consists of the polynucleotide sequence 5’ ATC TCG TAT GCC GTC TTC TGC TTG 3’ (SEQ ID NO. 8), or a variant thereof. In some embodiments, B15 comprises or consists of the polynucleotide sequence 5’ GTCTCGTGGGCTCGG 3’ (SEQ ID NO: 9), or a variant thereof. In some embodiments, B15’ comprises or consists of the polynucleotide sequence 5’ CCGAGCCCACGAGAC 3’ (SEQ ID NO: 10), or a variant thereof. In some embodiments, P15 comprises or consists of the polynucleotide sequence 5’ TTTTTTAATG ATACGGCGAC CACCGAGANC TACAC 3’ (SEQ ID NO: 11 ), or a variant thereof. In some embodiments, P17 comprises or consists of the polynucleotide sequence 5’ TTTTTTNNNC AAGCAGAAGA CGGCATACGA GAT 3’ (SEQ ID NO: 12), or a variant thereof. The term “variant” as used herein with reference to any of the sequences recited herein refers to a variant nucleic acid that is substantially identical, i.e., has only some nucleotide sequence variations, for example to the non-variant sequence. In some embodiments, a variant has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% overall nucleotide sequence identity to the nonvariant nucleic acid sequence. It will be understood that reference to P5 and P7 herein could refer to different primer sequences. Any suitable primer sequence combinations are encompassed by the present disclosure.
[0098] As used herein, the term "plurality" is intended to mean a population of two or more different members. Pluralities can range in size from small, medium, large, to very large. The size of small plurality can range, for example, from a few members to tens of members. Medium sized pluralities can range, for example, from tens of members to about 100 members or hundreds of members. Large pluralities can range, for example, from about hundreds of members to about 1000 members, to thousands of members and up to tens of thousands of members. Very large pluralities can range, for example, from tens of thousands of members to about hundreds of thousands, a million, millions, tens of millions and up to or greater than hundreds of millions of members. Therefore, a plurality can range in size from two to well over one hundred million members as well as all sizes, as measured by the number of members, in between and greater than the above exemplary ranges. An exemplary number of features within a microarray includes a plurality of about 500,000 or more discrete features within 1 .28 cm2. Exemplary nucleic acid pluralities include, for example, populations of about 1 x 105, 5 x 105and 1 x 106or more different nucleic acidspecies. Accordingly, the definition of the term is intended to include all integer values greater than two. An upper limit of a plurality can be set, for example, by the theoretical diversity of nucleotide sequences in a nucleic acid sample.
[0099] As used herein, the term "nucleic acid" is intended to be consistent with its use in the art and includes naturally occurring nucleic acids or functional analogs thereof. Particularly useful functional analogs are capable of hybridizing to a nucleic acid in a sequence specific fashion or capable of being used as a template for replication of a particular nucleotide sequence. Naturally occurring nucleic acids generally have a backbone containing phosphodiester bonds. An analog structure can have an alternate backbone linkage including any of a variety of those known in the art. Naturally occurring nucleic acids generally have a deoxyribose sugar (e.g., found in deoxyribonucleic acid (DNA)) or a ribose sugar (e.g., found in ribonucleic acid (RNA)). A nucleic acid can contain any of a variety of analogs of these sugar moieties that are known in the art. A nucleic acid can include native or non-native bases. In this regard, a native deoxyribonucleic acid can have one or more bases selected from the group consisting of adenine, thymine, cytosine or guanine and a ribonucleic acid can have one or more bases selected from the group consisting of uracil, adenine, cytosine or guanine. Useful non-native bases that can be included in a nucleic acid are known in the art. The term "target," when used in reference to a nucleic acid, is intended as a semantic identifier for the nucleic acid in the context of a method or composition set forth herein and does not necessarily limit the structure or function of the nucleic acid beyond what is otherwise explicitly indicated. Particular forms of nucleic acids may include all types of nucleic acids found in an organism as well as synthetic nucleic acids such as polynucleotides produced by chemical synthesis.
[0100] Particular examples of nucleic acids that are applicable for analysis through incorporation into microarrays produced by methods as provided herein include genomic DNA (gDNA), expressed sequence tags (ESTs), DNA copied messenger RNA (cDNA), RNA copied messenger RNA (cRNA), mitochondrial DNA or genome, RNA, messenger RNA (mRNA), ribosomal RNA (rRNA) and / or other populations of RNA. Fragments and / or portions of these exemplary nucleic acids also are included within the meaning of the term as it is used herein.
[0101] As used herein, the term "double-stranded," when used in reference to a nucleic acid molecule, means that substantially all of the nucleotides in the nucleic acid molecule are hydrogen bonded to a complementary nucleotide. A partially double stranded nucleic acid can have at least 10%, 25%, 50%, 60%, 70%, 80%, 90% or 95% of its nucleotide’s hydrogen bonded to a complementary nucleotide.
[0102] As used herein, the term "single-stranded," when used in reference to a nucleic acid molecule, means that essentially none of the nucleotides in the nucleic acid molecule are hydrogen bonded to a complementary nucleotide.
[0103] As used herein, the term "capture primers" or “capture probe” is intended to mean an oligonucleotide having a nucleotide sequence that is capable of specifically annealing to a single stranded polynucleotide sequence to be analyzed or subjected to a nucleic acid interrogation under conditions encountered in a primer annealing step of, for example, an amplification or sequencing reaction. The terms "nucleic acid," "polynucleotide" and "oligonucleotide" are used interchangeably herein. The different terms are not intended to denote any particular difference in size, sequence, or other property unless specifically indicated otherwise. For clarity of description the terms can be used to distinguish one species of nucleic acid from another when describing a particular method or composition that includes several nucleic acid species.
[0104] As used herein, the term "gene-specific" or "target specific" when used in reference to a capture probe or other nucleic acid is intended to mean a capture probe or other nucleic acid that includes a nucleotide sequence specific to a targeted nucleic acid, e.g., a nucleic acid from a tissue sample, namely a sequence of nucleotides capable of selectively annealing to an identifying region of a targeted nucleic acid. Gene-specific or target-specific capture probes can have a single species of oligonucleotide, or can include two or more species with different sequences. Thus, the gene-specific or target-specific capture probes can be two or more sequences, including 3, 4, 5, 6, 7, 8, 9 or 10 or more different sequences. The gene-specific or target-specific capture probes can comprise a gene-specific or target-specific capture primer sequence and a universal capture probe sequence. Other sequences such as sequencing primer sequences and the like also can be included in a gene-specific or target-specific capture primer.
[0105] As used herein, the term "amplicon," when used in reference to a nucleic acid, means the product of copying the nucleic acid, wherein the product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be produced by any of a variety of amplification methods that use the nucleic acid, or an amplicon thereof, as a template including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of the nucleotide sequence (e.g., a concatameric product of RCA). A first amplicon of a target nucleic acid can be a complementary copy. Subsequent amplicons are copies that are created, after generation of the first amplicon, from the targetnucleic acid or from the first amplicon. A subsequent amplicon can have a sequence that is substantially complementary to the target nucleic acid or substantially identical to the target nucleic acid.
[0106] The number of template copies or amplicons that can be produced can be modulated by appropriate modification of the amplification reaction including, for example, varying the number of amplification cycles run, using polymerases of varying processivity in the amplification reaction and / or varying the length of time that the amplification reaction is run, as well as modification of other conditions known in the art to influence amplification yield. The number of copies of a nucleic acid template can be at least 1 , 10, 100, 200, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 and 10,000 copies, and can be varied depending on the particular application.
[0107] As used herein, the term “complementary” when used in reference to a polynucleotide is intended to mean a polynucleotide that includes a nucleotide sequence capable of selectively annealing to an identifying region of a target polynucleotide under certain conditions. As used herein, the term "substantially complementary" and grammatical equivalents is intended to mean a polynucleotide that includes a nucleotide sequence capable of specifically annealing to an identifying region of a target polynucleotide under certain conditions. Annealing refers to the nucleotide base-pairing interaction of one nucleic acid with another nucleic acid that results in the formation of a duplex, triplex, or other higher-ordered structure. The primary interaction is typically nucleotide base specific, e.g., A:T,A:ll, and G:C, by Watson-Crick and Hoogsteen-type hydrogen bonding. In certain embodiments, base-stacking and hydrophobic interactions can also contribute to duplex stability. Conditions under which a polynucleotide anneals to complementary or substantially complementary regions of target nucleic acids are well known in the art, e.g., as described in Nucleic Acid Hybridization, A Practical Approach, Hames and Higgins, eds., IRL Press, Washington, D.C. (1985) and Wetmur and Davidson, Mol. Biol. 31 :349 (1968). Annealing conditions will depend upon the particular application and can be routinely determined by persons skilled in the art, without undue experimentation.
[0108] As used herein, the term "hybridization" refers to the process in which two singlestranded polynucleotides bind non-covalently to form a stable double-stranded polynucleotide. A resulting double-stranded polynucleotide is a "hybrid" or "duplex." Hybridization conditions will typically include salt concentrations of less than about 1 M, more usually less than about 500 mM and may be less than about 200 mM. A hybridization buffer includes a buffered salt solution such as 5% SSPE, or other such buffers known in the art. Hybridization temperatures can be as low as 5°C, but are typically greater than 22°C, and more typically greater than about 30°C, and typically in excess of 37°C. Hybridizationsare usually performed under stringent conditions, i.e., conditions under which a probe will hybridize to its target subsequence but will not hybridize to the other, uncomplimentary sequences. Stringent conditions are sequence-dependent and are different in different circumstances, and may be determined routinely by those skilled in the art.
[0109] As used herein, the term “dNTP” refers to deoxynucleoside triphosphates. NTP refers to ribonucleotide triphosphates. The purine bases (Pu) include adenine (A), guanine(G) and derivatives and analogs thereof. The pyrimidine bases (Py) include cytosine (C), thymine (T), uracil (U) and derivatives and analogs thereof. Examples of such derivatives or analogs, by way of illustration and not limitation, are those which are modified with a reporter group, biotinylated, amine modified, radiolabeled, alkylated, and the like and also include phosphorothioate, phosphite, ring atom modified derivatives, and the like. The reporter group can be a fluorescent group such as fluorescein, a chemiluminescent group such as luminol, a terbium chelator such as N-(hydroxyethyl) ethylenediaminetriacetic acid that is capable of detection by delayed fluorescence, and the like.
[0110] As used herein, the terms "ligation," “ligating,” and grammatical equivalents thereof are intended to mean to form a covalent bond or linkage between the termini of two or more nucleic acids, e.g., oligonucleotides and / or polynucleotides, typically in a template-driven reaction. The nature of the bond or linkage may vary widely, and the ligation may be carried out enzymatically or chemically. As used herein, ligations are usually carried out enzymatically to form a phosphodiester linkage between a 5' carbon terminal nucleotide of one oligonucleotide with a 3' carbon of another nucleotide. Template driven ligation reactions are described in the following references: U.S. Patent Nos. 4,883,750; 5,476,930;5,593,826; and 5,871 ,921 , incorporated herein by reference in their entireties. The term “ligation” also encompasses non-enzymatic formation of phosphodiester bonds, as well as the formation of non-phosphodiester covalent bonds between the ends of oligonucleotides, such as phosphorothioate bonds, disulfide bonds, and the like.
[0111] As used herein, the term "each," when used in reference to a collection of items, is intended to identify an individual item in the collection but does not necessarily refer to every item in the collection unless the context clearly dictates otherwise.
[0112] As used herein, the term "extend," when used in reference to a nucleic acid, is intended to mean addition of at least one nucleotide or oligonucleotide to the nucleic acid. In particular embodiments one or more nucleotides can be added to the 3' end of a nucleic acid, for example, via polymerase catalysis (e.g., DNA polymerase, RNA polymerase or reverse transcriptase). Chemical or enzymatic methods can be used to add one or more nucleotide to the 3' or 5' end of a nucleic acid. One or more oligonucleotides can be added tothe 3' or 5' end of a nucleic acid, for example, via chemical or enzymatic (e.g., ligase catalysis) methods. A nucleic acid can be extended in a template directed manner, whereby the product of extension is complementary to a template nucleic acid that is hybridized to the nucleic acid that is extended.
[0113] Provided herein are arrays for and methods of spatial detection and analysis (e.g., mutational analysis or single nucleotide variation (SNV) detection as well as indel detection) of nucleic acid in a tissue sample. The arrays described herein can comprise a substrate on which a plurality of capture probes is immobilized such that each capture probe occupies a distinct position on the array. Some or all of the plurality of capture probes can comprise a unique positional tag (i.e., a spatial address or indexing sequence). A spatial address can describe the position of the capture probe on the array. The position of the capture probe on the array can be correlated with a position in the tissue sample.
[0114] As used herein, the term "poly T", "poly A," or “poly II” when used in reference to a nucleic acid sequence e.g., a capture nucleotide sequence), is intended to mean a series of two or more thiamine (T), adenine (A) or uridine (U) bases, respectively. A poly T or poly A can include at least about 2, 5, 8, 10, 12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, or more of the T or A bases, respectively. Alternatively, or additionally, a poly T or poly A can include at most about 40, 38, 35, 32, 30, 28, 25, 22, 20, 18, 15, 12, 10, 8, 5, or 2 of the T or A bases, respectively. In some embodiments, the disclosure contemplates use of a "TVN" sequence, wherein “T” is a capture nucleotide sequence, “V” is adenine (A), cytosine (C), or guanine (G), and “N” is adenine (A), cytosine (C), guanine (G), or thymine (T). The TVN sequence is used, in some embodiments, to bias reverse transcription to the base of the poly A tail on a mRNA molecule.
[0115] As used herein, the term “tagmentation,” “tagment,” or “tagmenting” refers to transforming a nucleic acid, e.g., a DNA, into adaptor-modified templates in solution ready for cluster formation and sequencing by the use of transposase mediated fragmentation and tagging. This process often involves the modification of the nucleic acid by a transposome complex comprising transposase enzyme complexed with adaptors comprising transposon end sequence. Tagmentation results in the simultaneous fragmentation of the nucleic acid and ligation of the adaptors to the 5' ends of both strands of duplex fragments. Following a purification step to remove the transposase enzyme, additional sequences are added to the ends of the adapted fragments by PCR.
[0116] A “transposase” refers to an enzyme that is capable of forming a functional complex with a transposon end-containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposonend-containing composition into the double-stranded target nucleic acid with which it is incubated, for example, in an in vitro transposition reaction. A transposase as presented herein can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of US Pat. Publ. No. 2010 / 0120098, the content of which is incorporated herein by reference in its entirety. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon end with sufficient efficiency to 5'-tag and fragment a target nucleic acid for its intended purpose can be used in the present invention. In particular embodiments, a preferred transposition system is capable of inserting the transposon end in a random or in an almost random manner to 5'-tag and fragment the target nucleic acid.
[0117] As used herein, the term “transposition reaction” refers to a reaction wherein one or more transposons are inserted into target nucleic acids, e.g., at random sites or almost random sites. Essential components in a transposition reaction are a transposase and DNA oligonucleotides that exhibit the nucleotide sequences of a transposon, including the transferred transposon sequence and its complement (the non- transferred transposon end sequence) as well as other components needed to form a functional transposition or transposome complex. The DNA oligonucleotides can further comprise additional sequences (e.g., adaptor or primer sequences) as needed or desired. In some embodiments, the method provided herein is exemplified by employing a transposition complex formed by a hyperactive Tn5 transposase and a Tn5-type transposon end (Goryshin and Reznikoff, 1998, J. Biol. Chem., 273: 7367) or by a MuA transposase and a Mu transposon end comprising Rland R2 end sequences (Mizuuchi, 1983, Cell, 35: 785; Savilahti et al., 1995, EMBO J., 14:4893). However, any transposition system that is capable of inserting a transposon end in a random or in an almost random manner with sufficient efficiency to 5'- tag and fragment a target DNA for its intended purpose can be used in the present invention. Examples of transposition systems known in the art which can be used for the present methods include but are not limited to Staphylococcus aureus Tn552 (Colegio et al., 2001 , J Bacterid., 183: 2384-8; Kirby et al., 2002, Mol Microbiol, 43: 173-86), Tyl (Devine and Boeke, 1994, NucleicAcids Res., 22: 3765-72 and International Patent Application No. WO 95 / 23875), TransposonTn7 (Craig, 1996, Science. 271 : 1512; Craig, 1996, Review in: Curr Top Microbiollmmunol, 204: 27-48), TnlO and ISIO (Kleckner et al., 1996, Curr Top Microbiol Immunol, 204: 49-82), Mariner transposase (Lampe et al., 1996, EMBO J., 15: 5470-9), Tci (Plasterk,1996, Curr Top Microbiol Immunol, 204: 125-43), P Element (Gloor, 2004, Methods Mol Biol, 260: 97-114), TnJ (Ichikawa and Ohtsubo, 1990, J Biol Chem. 265: 18829-32),bacterial insertion sequences (Ohtsubo and Sekine, 1996, Curr. Top. Microbiol. Immunol. 204:1 -26), retroviruses (Brown et al., 1989, Proc Natl Acad Sci USA, 86: 2525-9), and retrotransposon of yeast (Boeke and Corces, 1989, Annu Rev Microbiol. 43: 403-34). The method for inserting a transposon end into a target sequence can be carried out in vitro using any suitable transposon system for which a suitable in vitro transposition system is available or that can be developed based on knowledge in the art. In general, a suitable in vitro transposition system for use in the methods provided herein requires, at a minimum, a transposase enzyme of sufficient purity, sufficient concentration, and sufficient in vitro transposition activity and a transposon end with which the transposase forms a functional complex with the respective transposase that is capable of catalyzing the transposition reaction. Suitable transposase transposon end sequences that can be used in the invention include but are not limited to wild-type, derivative or mutant transposon end sequences that form a complex with a transposase chosen from among a wild-type, derivative or mutant form of the transposase. As used herein, the term “transposome complex” refers to a transposase enzyme non-covalently bound to a double stranded nucleic acid. For example, the complex can be a transposase enzyme pre-incubated with double-stranded transposon DNA under conditions that support non-covalent complex formation. Double-stranded transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon end composition, a mixture of transposon end compositions or other doublestranded DNAs capable of interacting with a transposase such as the hyperactive Tn5 transposase.
[0118] As used herein, the term "random" can be used to refer to the spatial arrangement or composition of locations on a surface. For example, there are at least two types of order for an array described herein, the first relating to the spacing and relative location of features (also called "sites") and the second relating to identity or predetermined knowledge of the particular species of molecule that is present at a particular feature. Accordingly, features of an array can be randomly spaced such that nearest neighbor features have variable spacing between each other. Alternatively, the spacing between features can be ordered, for example, forming a regular pattern such as a rectilinear grid or hexagonal grid. In another respect, features of an array can be random with respect to the identity or predetermined knowledge of the gene of interest (e.g., nucleic acid of a particular sequence) that occupies each feature independent of whether spacing produces a random pattern or ordered pattern. An array set forth herein can be ordered in one respect and random in another. For example, in some embodiments set forth herein a surface is contacted with a population of nucleic acids under conditions where the nucleic acids attach at sites that are ordered with respect to their relative locations but 'randomly located' with respect to knowledge of the sequencefor the nucleic acid species present at any particular site. Reference to "randomly distributing" nucleic acids at locations on a surface is intended to refer to the absence of knowledge or absence of predetermination regarding which nucleic acid will be captured at which location (regardless of whether the locations are arranged in an ordered pattern or not).
[0119] As used herein, a "biological sample" may include one or more biological or chemical substances, such as nucleic acids, oligonucleotides, proteins, cells, tissues, organisms, and / or biologically active chemical compound(s), such as analogs or mimetics of the aforementioned species. As used herein, the term "tissue" is intended to mean an aggregation of cells, and, optionally, intercellular matter. Typically the cells in a tissue are not free floating in solution and instead are attached to each other to form a multicellular structure. Exemplary tissue types include muscle, nerve, epidermal and connective tissues. In some instances, the biological sample may include whole blood, lymphatic fluid, serum, plasma, sweat, tear, saliva, sputum, cerebrospinal fluid, amniotic fluid, seminal fluid, vaginal excretion, serous fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, transudates, exudates, cystic fluid, bile, urine, gastric fluid, intestinal fluid, fecal samples, liquids containing single or multiple cells, liquids containing organelles, fluidized tissues, fluidized organisms, viruses including viral pathogens, liquids containing multi-celled organisms, biological swabs and biological washes. In further examples, the sample can be derived from an organ, including for example, an organ of the musculoskeletal system such as muscle, bone, tendon or ligament; an organ of the digestive system such as salivary gland, pharynx, esophagus, stomach, small intestine, large intestine, liver, gallbladder or pancreas; an organ of the respiratory system such as larynx, trachea, bronchi, lungs or diaphragm; an organ of the urinary system such as kidney, ureter, bladder or urethra; a reproductive organ such as ovary, fallopian tube, uterus, vagina, placenta, testicle, epididymis, vas deferens, seminal vesicle, prostate, penis or scrotum; an organ of the endocrine system such as pituitary gland, pineal gland, thyroid gland, parathyroid gland, or adrenal gland; an organ of the circulatory system such as heart, artery, vein or capillary; an organ of the lymphatic system such as lymphatic vessel, lymph node, bone marrow, thymus or spleen; an organ of the central nervous system such as brain, brainstem, cerebellum, spinal cord, cranial nerve, or spinal nerve; a sensory organ such as eye, ear, nose, or tongue; or an organ of the integument such as skin, subcutaneous tissue or mammary gland. In various embodiments, the tissue can be derived from a multicellular organism. In some embodiments, a tissue section can be contacted with a surface, for example, by laying the tissue on the surface. The tissue can be freshly excised from an organism, or it may have been previously preserved for example by freezing (e.g., fresh frozen tissue),embedding in a material such as paraffin (e.g., formalin fixed paraffin embedded (FFPE) samples), formalin fixation, infiltration, dehydration or the like. Optionally, a tissue section can be attached to a surface, for example, using techniques and compositions described in, for example, U.S. Patent No. 11 ,390,912, incorporated by reference herein in its entirety. In some embodiments, a tissue can be permeabilized and the cells of the tissue lysed when the tissue is in contact with a surface. Any of a variety of treatments can be used such as those set forth above in regard to lysing cells. Target proteins and / or nucleic acids that are released from a tissue that is permeabilized can be captured by capture oligonucleotides on the surface. Thus, in various embodiments, the biological sample is a tissue sample. The thickness of a tissue sample or other biological sample that is contacted with a surface in a method set forth herein can be any suitable thickness desired. In representative embodiments, the thickness will be at least 0.1 pm, 0.25 pm, 0.5 pm, 0.75 pm, 1 pm, 5 pm, 10 pm, 50 pm, 100 pm or thicker. Alternatively or additionally, the thickness of a biological sample that is contacted with a surface will be no more than 100 pm, 50 pm, 10 pm, 5 pm, 1 pm, 0.5 pm, 0.25 pm, 0.1 pm or thinner.
[0120] As used herein, the term "tissue sample" refers to a piece of tissue that has been obtained from a subject, optionally fixed, sectioned, and mounted on a planar surface, e.g., a microscope slide. The tissue sample can be a formalin-fixed paraffin-embedded (FFPE) tissue sample or a fresh tissue sample or a frozen tissue sample, etc. The methods disclosed herein may be performed before or after staining the tissue sample. For example, following hematoxylin and eosin staining, a tissue sample may be spatially analyzed in accordance with the methods as provided herein. A method may include analyzing the histology of the sample (e.g., using hematoxylin and eosin staining) and then spatially analyzing the tissue.
[0121] As used herein, the term "formalin-fixed paraffin embedded (FFPE) tissue section" refers to a piece of tissue, e.g., a biopsy that has been obtained from a subject, fixed in formaldehyde (e.g., 3%-5% formaldehyde in phosphate buffered saline) or Bouin solution, embedded in wax, cut into thin sections, and then mounted on a planar surface, e.g., a microscope slide.
[0122] As used herein, the term “subject” encompasses mammals and non-mammals. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates such as chimpanzees, and other apes and monkey species, cattle, horses, sheep, goats, swine, rabbits, dogs, cats, rodents, rats, mice, guinea pigs, and the like. Examples of non-mammals include, but are not limited to, birds, fish, and the like. The term does not denote a particular age or gender.
[0123] In some embodiments, nucleic acids in a tissue sample are transferred to and captured onto an array. For example, a tissue section is placed in contact with an array and nucleic acid is captured onto the array and tagged with a spatial address. The spatially- tagged DNA molecules are released from the array and analyzed, for example, by high throughput next generation sequencing (NGS), such as sequencing-by-synthesis (SBS). In some embodiments, a nucleic acid in a tissue section (e.g., a formalin-fixed paraffin- embedded (FFPE) tissue section) is transferred to an array and captured onto the array by hybridization to a capture probe. In some embodiments, a capture probe can be a universal capture probe hybridizing, e.g., to an adaptor region in a nucleic acid sequencing library, or to the poly-A tail of an mRNA. In some embodiments, the capture probe can be a genespecific or target-specific capture probe hybridizing, e.g., to a specifically targeted RNA or cDNA in a sample, such as a TruSeq™ Custom Amplicon (TSCA) oligonucleotide probe (Illumina, Inc.). A capture probe can be a plurality of capture probes, e.g., a plurality of the same or of different capture probes.
[0124] In some embodiments, a combinatorial indexing (addressing) system is used to provide spatial information for analysis of nucleic acids in a tissue sample. The combinatorial indexing system can involve the use of two or more spatial address sequences (e.g., two, three, four, five or more spatial address sequences).
[0125] In some embodiments, two spatial address sequences are incorporated into a nucleic acid during preparation of a sequencing library. A first spatial address can be used to define a certain position (i.e., capture site) in the X dimension on a capture array and a second spatial address sequence can be used define a position (i.e., a capture site) in the Y dimension on the capture array. During library sequencing, both X and Y spatial address sequences can be determined and the sequence information can be analyzed to define the specific position on the capture array.
[0126] In some embodiments, three spatial address sequences are incorporated into a nucleic acid during preparation of a sequencing library. A first spatial address can be used to define a certain position (i.e., capture site) in the X dimension on a capture array, a second spatial address sequence can be used define a position (i.e., a capture site) in the Y dimension on the capture array, and a third spatial address sequence can be used to define a position of a two-dimensional sample section (e.g., the position of a slice of a tissue sample) in a sample (e.g., a tissue biopsy) to provide positional spatial information in the third dimension (Z dimension) of a sample. During library sequencing, X, Y, and Z spatial address sequences can be determined and the sequence information can be analyzed to define the specific position on the capture array.
[0127] In some embodiments, a temporal address sequence (T) is optionally incorporated into a nucleic acid during preparation of a sequencing library. In some embodiments, the temporal address sequence can be combined with two or three spatial address sequences. The temporal address sequence can, for example, be used in the context of a time-course experiment for determining time-dependent changes in gene-expression in a tissue sample. Time-dependent changes in gene-expression can occur in a tissue sample, for example, in response to a chemical, biological or physical stimulus (e.g., a toxin, a drug, or heat). Nucleic acid samples obtained at different timepoints from comparable tissue samples (e.g., proximal slices of a tissue sample) can be pooled and sequenced in bulk. An optional first spatial address can be used to define a certain position (i.e., capture site) in the X dimension on a capture array, a second optional spatial address sequence can be used to define a position (i.e., a capture site) in the Y dimension on the capture array, and a third optional spatial address sequence can be used to define a position of a two-dimensional sample section (e.g., the position of a slice of a tissue sample) in a sample (e.g., a tissue biopsy) to provide positional spatial information in the third dimension (Z dimension) of the sample. During library sequencing, T, X, Y, and Z address sequences are determined and the sequence information is analyzed to define the specific X, Y (and optionally Z) position on the capture array for each timepoint (T).
[0128] The address sequences X, Y, and, optionally, Z and / or T, can be consecutive nucleic acid sequences or the address sequences can be separated by one or more nucleic acids (e.g., 2 or more, 3 or more, 10 or more, 30 or more, 100 or more, 300 or more, or 1 ,000 or more). In some embodiments, the X, Y, and optionally Z and / or T address sequences can each individually and independently be combinatorial nucleic acid sequences.
[0129] In some embodiments, the length of the address sequences (e.g., X, Y, Z, or T) can each individually and independently be 100 nucleic acids or less, 90 nucleic acids or less, 80 nucleic acids or less, 70 nucleic acids or less, 60 nucleic acids or less, 50 nucleic acids or less, 40 nucleic acids or less, 30 nucleic acids or less, 20 nucleic acids or less, 15 nucleic acids or less, 10 nucleic acids or less, 8 nucleic acids or less, 6 nucleic acids or less, or 4 nucleic acids or less. The length of two or more address sequences in a nucleic acid can be the same or different. For example, if the length of address sequence X is 10 nucleic acids, the length of address sequence Y can be, e.g., 8 nucleic acids, 10 nucleic acids, or 12 nucleic acids.
[0130] Address sequences, e.g., spatial address sequences such as X or Y, can be either partially or fully degenerate sequences.
[0131] In some embodiments, spatially addressed capture probes on an array can be released from the array onto a tissue section for generation of a spatially addressed sequencing library. In some embodiments, a capture probe comprises a random primer sequence for in situ synthesis of spatially-tagged cDNA from RNA in the tissue section. In some embodiments, a capture probe is a TruSeq™ Custom Amplicon (TSCA) oligonucleotide probe (Illumina, Inc.) for capturing and spatially tagging genomic DNA in the tissue section. The spatially-tagged nucleic acid molecules (e.g., cDNA or genomic DNA) are recovered from the tissue section and processed in single tube reactions to generate a spatially-tagged amplicon library.
[0132] In some embodiments, magnetic nanoparticles can be used to capture nucleic acid (e.g., in situ synthesized cDNA) in a tissue sample for generation of a spatially addressed library.
[0133] In some embodiments, spatial detection and analysis of nucleic acid in a tissue sample can be performed on a droplet actuator.
[0134] Described herein are improved methods and compositions for spatial-omics applications that preserve spatial information related to the origin of RNA or DNA in the tissue. Examples of spatial omics applications include, but are not limited to, spatial genomic applications, spatial proteomic applications; spatial transcriptomic applications; spatial agrigenomic applications; spatial epigenomics s applications; spatial phenomic applications;spatial ligandomic applications; and spatial multiomic applications (e.g., transcriptomic and genomic applications).Nanostructures in RNA Capture
[0135] Contemplated herein are methods for improved capture of RNA from a tissue sample comprising contacting the sample with a nanostructure comprising RNA capture probes that hybridize to the target nucleotide sequence and surface capture probes that are linked to the nanostructure. Different workflows are provided that may be used in different RNA purification and library preparation situations (Figures 1-3).
[0136] In a first workflow (Figure 1), a top-down method is applied wherein following tissue fixation and permeabilization, the nanostructures are added to the surface of the fixed tissue and allowed to diffuse to the substrate while simultaneously capturing RNA (Figure 4). Following anchoring to the surface, a reverse transcriptase treatment is applied, generating cDNA from the captured RNA. A key advantage of this approach compared to the standard workflow is that because single RNA strands can be captured by multiple probes, when a strand displacing reverse transcriptase is used, multiple copies of cDNA are generated from a single RNA strand (Figure 4). After RT, cleavage of the dendrimer oligos is initiated andthrough the implementation of a splint oligo and ligation step, the cDNA strands are attached to the barcoded surface clusters (Figure 4D). It is provided that first strand cDNA synthesis occurs in the tissue or on the substrate in this method. Optionally, once the first strand cDNA is formed, the sample may be treated with an RNase to release the cDNA attached to the capture probes on the nanostructure.
[0137] In this method there is a benefit over surface clustered strand-based capture in that the capture probes are directly introduced to the tissue such that the capture probes and RNA are co-localized, significantly enhancing capture efficiency. This may be particularly useful for target-based spatial transcriptomics, where a researcher can selectively choose which regions of tissue to introduce the nanostructures. In the “top-down” approach, some loss in resolution may occur, as the nanostructure must diffuse to the surface and captured RNA travels with it. However, since the nanostructures are small relative to the clusters (~50 nm vs. ~1 pm), a large degree of lateral diffusion should not result in significant loss of resolution. The top-down approach as the added benefit of being modular and the underlying surface chemistry is not so relevant (provided nanostructure anchoring sites are implemented in the standard workflow). Following mRNA capture and diffusion to the surface, an anchoring site will secure the nanostructure to the surface, and reverse transcriptase is initiated. The oligos, now containing cDNA, are cleaved from the surface and undergo hybridization with a split oligo and then are ligated to the surface cluster strands where the subsequent steps are based on standard library preparation workflow
[0138] For example, the top down method of preparing a spatially barcoded RNA library from a tissue sample comprises,
[0139] (a) permeabilizing the tissue sample;
[0140] (b) contacting the tissue sample with a plurality of nanostructures, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nanoscaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety;
[0141] (c) capturing RNA from the tissue sample by hybridization of the RNA with RNA capture probes on the nanostructures;
[0142] (d) capturing the plurality of nanostructures comprising captured RNA on a substrate, wherein the substrate comprises a plurality of capture sites and a plurality of surface oligonucleotide molecules, wherein each of the plurality of capture sites is capable ofbinding to the substrate anchor moiety of the nanostructure, and wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules;
[0143] (e) contacting the captured RNA with a reverse transcriptase to generate one or more first strand cDNAs wherein each of the first strand cDNAs is contiguous with one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe;
[0144] (f) cleaving the cleavable linkers to release the first strand cDNAs from the nanoscaffold; and
[0145] (g) capturing the first strand cDNAs on the substate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
[0146] In a “bottom-down” approach (Figure 2) (Figure 5A, center), the nanostructures are anchored to the substrate surface prior to introduction of tissue. The workflow is largely unchanged as from described above, but due to crowding of neighboring surface clustered strands, capture quantities could be reduced. In both workflows, the nanostructures could be the sole capture species or work in conjunction with the barcoded clusters also containing capture sites.
[0147] For example, a bottom down method for preparing a spatially barcoded RNA library from a tissue sample comprises,
[0148] (a) contacting a tissue sample with a substrate comprising a plurality of nanostructures and a plurality of surface oligonucleotide molecules attached to the substrate, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; and; (iii) two or more cleavable linkers wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules;
[0149] (b) permeabilizing the tissue sample to release RNA from the tissue sample;
[0150] (c) capturing RNA from the permeabilized tissue sample by hybridization of the RNA with the RNA capture probe;
[0151] (d) contacting the captured RNA with a reverse transcriptase to generate one or more first strand cDNAs wherein each of the first strand cDNAs is contiguous with one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe;
[0152] (e) cleaving the cleavable linkers to release the first strand cDNAs from the nanoscaffold; and
[0153] (f) capturing the first strand cDNAs on the substrate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
[0154] A third approach, an “oligo-first” approach (Figure 3) (Figure 5A, right), relies on first allowing the capture probes to hybridize to the mRNA after tissue permeabilization. Following capture probe hybridization, the nanostructure is then introduced and via a click reaction (e.g., azide-alkyne, thiol-Michael, or the like) or ligation where the capture probes are linked to the nanostructure. The subsequent workflow follows the steps outlined above in the “top-down” approach. This approach is the most modular and adaptable to commercial substrates as the nanostructure is completely independent of the capture probes.
[0155] An exemplary oligo-first method for preparing a spatially barcoded RNA library from a tissue sample comprises,
[0156] (a) permeabilizing the tissue sample to release RNA from the tissue sample;
[0157] (b) hybridizing the RNA with a plurality of RNA capture probes to generate a plurality of RNA-capture probe hybrids;
[0158] (c) contacting the RNA-capture probe hybrids with a plurality of nanostructures, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, and (iv) a substrate anchor moiety, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers;
[0159] d) covalently attaching the 5’ capture probe ends of the RNA-capture probe hybrids to the 3’ ends of the oligonucleotides;
[0160] (e) contacting the covalently attached RNA-capture probe hybrids in step (d) with a reverse transcriptase to generate one or more first strand cDNAs, wherein each of the first strand cDNAs is contiguous with the one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe;
[0161] (f) capturing the plurality of nanostructures on a substrate, wherein the substrate comprises a plurality of capture sites and a plurality of surface oligonucleotide molecules, wherein each of the plurality of capture sites is capable of binding to the substrate anchor moiety, and wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules;
[0162] (g) cleaving the cleavable linkers to release the first strand cDNAs from the nanoscaffold; and
[0163] (h) capturing the first strand cDNAs on the substrate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
[0164] These nanostructures can be applied to any standard surface, provided there are anchoring sites present, thus allowing for one type of flow cell but limitless capture specificity. Random, polyT oligos and any number of target-specific probes could be used for whole transcriptome sequencing. This is a significant advantage from previous workflows in that target-specific spatial transcriptomics would only require a custom nanostructure compared to requiring a custom flow cell as would be required in the standard workflow.
[0165] In addition to modularity in the design of the capture probes, there is a significant degree of tunability. The length of the spacer connecting the capture oligo to the dendrimer core can be tuned over a wide range of lengths. Here there is expected to be a trade-off in capture efficiency (longer spacers providing additional reach for mRNA material, particularly in the bottom-up strategy) and resolution (shorter spacers can only hybridize to mRNA in close proximity).
[0166] Advantages of the various methods include enhancing the hybridization of mRNA through multiple binding sites afforded by dendritic probes to address the challenge of low quality or damaged mRNA from FFE samples. Further-more, gene specific probes can be used to target specific mRNA species. When randomer probes are employed, multiple binding can occur on single RNA strands, allowing for RNA which has been degraded and is missing a polyA tail to still be captured.
[0167] Additionally, total mass of cDNA generated in standard workflow is very low. Not only should more cDNA be generated in the present methods due to enhanced mRNA capture when using nanostructures as described above, but multiple copies of cDNA can be generated from single RNA strands. When multiple capture probes are hybridized to a single piece of RNA, the use of strand displacing reverse transcriptase will create multiple copies of that strand (equal to the number of hybridized capture probes), amplifying the total RNA material.
[0168] The methods also address poor data quality resulting from single cDNA strand synthesis per captured RNA and improved bioinformatics will result from the generation of multiple copies of RNA strands by the present methods.
[0169] Many types of morphologies that may be used in this concept are dendrimers, nanoparticles, nanogels, or hyperbranched polymers (Figure 2B). While dendrimers are precisely structured molecules with specific numbers of end groups and predictable molecular weights, they are more complex to synthesize than hyperbranched polymers or nanogels which also offer multiple end groups and can be functionalized in a manner similar to dendrimers. Nanoparticles may also be envisioned for this concept and may offer unique advantages such as magnetic pulldown to the substrate surface upon introduction to the tissue sample.Preparation of polynucleotides
[0170] The present disclosure is based, in part, on the realization that the amount of RNA or DNA information isolatable from fresh or frozen tissue samples as well as FFPE tissue samples needs to be improved to provide information related to the genetic profile of the tissue sample. The present disclosure provides methods for improved capture of genetic information by increasing the amount and quality of RNA isolated from tissue samples that can be used in spatial transcriptomics analysis.
[0171] The total RNA can comprise ribosomal RNA (rRNA), messenger RNA (mRNA), transfer RNA (tRNA), microRNA (miRNA), non-coding RNA (ncRNA), small nucleolar RNA (snoRNA), and / or small nuclear RNA (snRNA). In various embodiments, the RNA is rRNA and / or mRNA.
[0172] In various embodiments, the RNA capture probe is selected from the group consisting of a poly-T sequence, a poly-U sequence, a randomer, a semi-random sequence, or a target-specific probe. In various embodiments, the target-specific probes comprise a plurality of different target-specific RNA capture probe sequences. In various embodiments, the RNA capture probe or surface capture probe is between 8 to 80 nucleotides. In certain embodiments, the RNA capture probe or surface probe is between 10 to 80 nucleotides,between 10 to 70 nucleotides, between 10 to 60 nucleotides, between 10 to 50 nucleotides, between 10 to 40 nucleotides, between 10 to 30 nucleotides, between 10 to 20 nucleotides, between 20 to 80 nucleotides, between 20 to 70 nucleotides, between 20 to 60 nucleotides, between 20 to 50 nucleotides, between 20 to 40 nucleotides, or is 8, 9, 10, 11 , 12, 13, 14,15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70 or 80 nucleotides.
[0173] In various embodiments, a capture oligonucleotide comprises a clustering primer sequence and a capture nucleotide sequence that is configured to bind to target nucleic acids of a biological sample. In some embodiments, a capture oligonucleotide comprises a clustering primer sequence (e.g., a P7 sequence), a spatial barcode (SBC) sequence, a sequencing primer sequence e.g., a sequencing by synthesis (SBS) sequence such as SBS12), a single molecule identifier (SMI) sequence, a quality control sequence, and a TVN sequence, wherein “T” is a capture nucleotide sequence, “V” is adenine (A), cytosine (C), or guanine (G), and “N” is adenine (A), cytosine (C), guanine (G), or thymine (T). In various embodiments, a capture oligonucleotide is between about 30 bases to about 100 bases in length, or between about 30 bases to about 90 bases, or between about 30 bases and 80 bases, or between about 30 bases and 70 bases, or between about 30 bases and 60 bases, or between about 30 bases and 55 bases, or between about 30 bases and 50 bases in length or between 20 bases to 80 bases, or between 10 bases to 80 bases. In further embodiments, a capture oligonucleotide of the disclosure is about 10 bases, 20 bases, 30 bases, 35 bases, 40 bases, 45 bases, 50 bases, 55 bases, 60 bases, 65 bases, 70 bases, 75 bases, 80 bases, 85 bases, 90 bases, 95 bases, or 100 bases in length. The capture nucleotide sequence capable of hybridizing or otherwise associating with an analyte e.g., a target nucleic acid) is, for example and without limitation, a universal sequence {e.g., a poly T sequence, a random nucleotide sequence, or a semi-random nucleotide sequence), or a target-specific {e.g., a gene-specific) sequence. In various embodiments, a capture nucleotide sequence {e.g., a poly T nucleotide sequence or a random nucleotide sequence) is, is about, or is at least about 2, 5, 8, 10, 12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, 45, 50, or more bases in length. Alternatively or additionally, a capture nucleotide sequence can include less than or equal to about 50, 45, 40, 38, 35, 32, 30, 28, 25, 22, 20, 18, 15, 12, 10, 8, 5, or 2 bases. A capture oligonucleotide can comprise additional elements, including but not limited to a single molecule identifier (SMI) {e.g., a unique molecular identifier (UMI)), an index sequence, a sequence that is complementary to a sequencing primer {e.g., SBS12), or a combination thereof. In some embodiments, beads are packed onto a solid support {e.g., a planar support or flow cell), wherein the beads comprise a plurality of capture oligonucleotides immobilized thereon, wherein one or more of the plurality of capture oligonucleotides comprises, from 5’ to 3’: (a) a first clustering primer sequence; (b) a spatialbarcode (SBC) sequence; (c) a first sequencing primer sequence; (d) a single molecule identifier (SMI) sequence; (e) a quality control sequence; and (f) a TVN sequence, wherein “T” is a capture nucleotide sequence, “V” is adenine (A), cytosine (C), or guanine (G), and “N” is adenine (A), cytosine (C), guanine (G), or thymine (T), and wherein the spatial barcode sequence of the plurality of capture oligonucleotides is unique to each bead
[0174] The oligonucleotides comprising a surface oligonucleotide (e.g., poly T sequences) can further comprise spatial index sequences, including, but not limited to, one or more of a P7 sequence, an index sequence, and / or a Read 2 (Rd2) sequence. In various embodiments, the surface oligonucleotide comprises a P7 anchor sequence, a spatial barcode and a sequence that hybridizes with a splint oligonucleotide.
[0175] In various embodiments, the sequence in the surface oligonucleotide that hybridizes with a splint oligonucleotide is a PZ (clustering) sequence. In various embodiments, the PZ sequence hybridizes to a splint oligonucleotide comprising a nucleotide sequence PZ’ complementary to the PZ sequence and a PX’ sequence that is complementary to the surface capture probe. In various embodiments, the PX sequence is a seeding sequence. In one embodiment, PX has the sequence AGGAGGAGGAGGAGGAGGAGGAGG.
[0176] In various embodiments, the cleavable linker that attached a capture probe to the nanostructure is a cleavable polynucleotide. In various embodiments, the cleavable polynucleotide is between 5 to 25 nucleotides, or is 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, or 25 nucleotides.
[0177] In some embodiments, the total RNA is released from the tissue sample. Release includes lysis of tissue or permeabilization of the tissue. In various embodiments, one or more samples that have been contacted with a solid support can be lysed to release target nucleic acids. Lysis can be carried out using known techniques, such as those that employ one or more of chemical treatment, enzymatic treatment, electroporation, heat, hypotonic treatment, sonication or the like. It is contemplated that the tissue sample is permeabilized prior to step (a) of the methods. In various embodiments, the tissue sample is treated with one or more blocking reagents prior to step (a) of the methods. In various embodiments, the tissue sample is permeabilized and treated with one or more blocking reagents prior to step (a) of the methods.
[0178] In some embodiments, a tissue sample will be treated to remove embedding material (e.g., to remove paraffin or formalin) from the sample prior to release, capture or modification of nucleic acids. This can be achieved by contacting the sample with an appropriate solvent (e.g., xylene and ethanol washes). Treatment can occur prior tocontacting the tissue sample with a solid support set forth herein or the treatment can occur while the tissue sample is on the solid support. Exemplary methods for manipulating tissues for use with solid supports to which nucleic acids are attached are set forth in US Pat. App. Publ. No. 2014 / 0066318, which is incorporated herein by reference.
[0179] A formalin-fixed tissue sample may also be decrosslinked using known techniques. In various embodiments, decrosslinking is carried out using Tris-EDTA (TE) buffer, e.g., at pH 8, pH 9, or another appropriate buffer at an appropriate pH. Decrosslinking may also be carried out at high heat, e.g., 70° C.
[0180] In various embodiments, the tissue sample is contacted with a mechanism to accelerate RNA diffusion from the sample. In various embodiments, the mechanism is magnetic or electrophoretic acceleration. If the surface is a 3D patterned surface, the mechanism may also include capture surfaces with pillars or other features extending outward from the surface. In various embodiments, additives to enhance crowding to surfaces are used to accelerate RNA diffusion. In various embodiments, the additives include PEG, ficoll, dextran sulfate, and the like. For example, the nanostructure, e.g., a dendrimer, may comprise a charge such that it can be pulled down via a mechanism that exploits the charged moiety. In various embodiments, the acceleration mechanism is applied after the mRNA capture step. In various embodiments, the acceleration mechanism is applied after the cDNA conversion step. It is contemplated that acceleration mechanism facilitates mRNA / cDNA binding the surface barcodes more quickly, which limits lateral diffusion and provides better resolution of the final tissue gene heat map.
[0181] A variety of different chemical reactions can be employed for the nanostructures both in terms of surface anchoring and oligo conjugation. For dendrimer synthesis, contemplated are peptide-based dendrimers with single alkyne functionality for surface anchoring and amine groups at the arm termini. To link oligos to the amine-terminated arms, a heterobifunctional linking molecule containing a sulfo-NHS ester (reactive towards amines), and a maleimide group (reactive towards thiols) can be used (in conjunction with thiol-terminated oligos). Alternatively, a DBCO-NHS ester compound could be used with azide-terminated oligos. In one method, dendrimers are linked to the surface using azidealkyne click chemistry with the azides being provided by the functionality on PAZAM. In various embodiments, the nanostructure is linked with the capture probe by azide-alkyne cycloaddition, a heterobifunctional linking group, a sulfo-NHS ester, a DBCO-NHS ester or a maleimide group.
[0182] RNA from the sample may also be prepared by performing end repair of the RNA with polynucleotide kinase prior to the step of capturing RNA from the tissue sample, and / orby performing in situ polyadenylation with polyadenylate polymerase, prior to the step of capturing RNA from the tissue sample. Methods of carrying out end repair of RNA from a tissue sample are described in co-owned US Provisional Application No. 63 / 477,730 (Docket No. 33080 / IP-2625-P) (herein incorporated by reference).
[0183] The methods above are also useful for improving capture efficiency of mRNA transcripts for in situ mRNA transcript library preparation, and / or for improving the nucleotide length of polynucleotides used in generating an in situ transcriptome library (e.g., improving the polynucleotide size of cDNA transcribed from mRNA isolated from a sample and used in generating an in situ transcriptome library).Spatial Detection and Analysis of Nucleic Acids in a Tissue Sample
[0184] According to the methods described herein, spatial detection and analysis of nucleic acids in a tissue sample can be performed using sets of two or more capture probes (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more capture probes). Typically, at least a first capture probe in a set of capture probes is immobilized on a capture array or a nanostructure. In some embodiments, a second capture probe can be immobilized on the same capture array as the first capture probe, e.g., in proximity to the first capture probe, e.g., in the same capture site. In some embodiments, a second capture probe can be immobilized on a nanostructure or a particle, such as a magnetic particle or a magnetic nanoparticle. In some embodiments, a second capture probe can be in solution, e.g., to be used to perform in situ reactions with a nucleic acid in a tissue sample. The capture probes in the capture probe sets individually and independently can have a variety of different regions, e.g., a capture region (e.g., a first universal or genespecific capture region or first clustering region), a primer binding region (e.g., a SBS primer region, such as a SBS3 or SBS12 region), or a second universal region / clustering sequence, such as a P5 or P7 region, a spatial address region (e.g., a partial or combinatorial spatial address region), or a cleavable region.
[0185] Briefly, SBS can be initiated by contacting the barcodes with one or more labeled nucleotides, DNA polymerase, etc. Those features where a primer is extended using the sequences comprising the barcode as a template will incorporate a labeled nucleotide that can be detected. Optionally, the labeled nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analog having a reversible terminator moiety can be added to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detectionoccurs). Washes can be carried out between the various delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluidic systems and detection platforms that can be readily adapted for use with an array produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Pat. Nos. 7,057,026; 7,329,492; 7,211 ,414; 7,315,019 or 7,405,281 , and US Pat. App. Pub. No. 2008 / 0108082 A1 , U.S. Patent Pub. No.2007 / 0166705, U.S. Patent Pub. No. 2006 / 0188901 , U.S. Patent Pub. No. 2006 / 0240439, U.S. Patent Pub. No. 2006 / 0281109, International Patent Pub. No. WO 05 / 065814, U.S. Patent Pub. No. 2005 / 0100900, International Patent Pub. No. WO 06 / 064199, International Patent Pub. No. WO 07 / 010251 , U.S. Patent Pub. No. 2012 / 0270305 and U.S. Patent Pub. No. 2013 / 0260372, each of which is incorporated herein by reference.
[0186] Exemplary sequences include the following Rd1 and Rd2 adaptor sequences.Second Universal Adapter - Rd1 SBS3 (long): ACACTCTTTCCCTACACGACGCTCTTCCGATCT ( SEQ ID NO : 13 ) ; Second Universal Adapter - Rd1 SBS3 (short): ACACTCTTTCCCTACACGAC ( SEQ ID NO : 14 ) ; First Universal Adapter - Rd2 SBS12 (long): GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT ( SEQ ID NO : 15 ) ; First Universal Adapter - Rd2 SBS12 (short): GTGACTGGAGTTCAGACGTGT ( SEQ ID NO : 16 ) .
[0187] In some embodiments, only one capture probe in a set of capture probes comprises a capture region. In some embodiments, two or more capture probes in a set of capture probes comprise as capture region.
[0188] In some embodiments, only one probe in a set of capture probes comprises a spatial address region, e.g., such as a complete spatial address region describing the position of a capture site on a capture array. In some embodiments, two or more probes in a set of capture probes can comprise a spatial address region, e.g., two or more probes can each comprise a partial spatial address region (i.e., combinatorial address region), wherein each partial address region describes the position of a capture site on a capture array, e.g., along the x-axis or the y-axis.
[0189] In some embodiments, a set of capture probes (e.g., a RNA and surface capture probe) can comprise at least one capture probe comprising a capture region and a spatial address region (e.g., a complete or a partial spatial address region). In some embodiments, no capture probe in a set of capture probes comprises both a capture region and a spatial address region.
[0190] In some embodiments, the RNA capture probe is a gene specific probe or a universal capture sequence, comprising a sequence complementary to the first universal adapter sequence and a gene specific primer or a universal capture sequence.
[0191] In some embodiments, the surface capture probe comprises one two or three of a unique molecular index (UMI), a spatial address region, and an universal adapter sequence (e.g., a second universal adapter sequence, such as a Rd1 adapter). In some embodiments, the surface capture probe does not comprise a spatial address region.
[0192] It is contemplated that the methods herein are carried out on substrates, e.g., flow cells, containing surface oligonucleotide molecules arranged randomly on the substrate, arranged in clusters, or arranged in patterns.
[0193] When surface oligonucleotide molecules are arranged randomly on the substrate (e.g., a flow cell), the method further comprises determining the substrate location of one or more surface oligonucleotide molecules by sequencing the spatial barcodes of the surface oligonucleotide molecules and assigning the spatial barcode sequences to locations on the substrate. Optionally, the method further comprises sequencing at least a portion of one or more spatially barcoded first strand cDNA molecules, or copies thereof, to identify the spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules, or copies thereof, and correlating the spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules, or copies thereof, with the known locations of spatial barcode sequences of the surface oligonucleotide molecules. In various embodiments, the sequence of the spatial barcodes is determined by next generation sequencing.
[0194] When surface oligonucleotide molecules are arranged in clusters on the substrate (e.g., a flow cell), the method further comprises, prior to contacting the tissue sample with the substrate, determining the substrate location of each cluster by sequencing the spatial barcode for at least one surface oligonucleotide molecule in each cluster and assigning the spatial barcode sequence to a location on the substrate. Optionally, the method further comprises determining the spatial location of the RNA molecules within the tissue sample by sequencing at least a portion of one or more spatially barcoded first strand cDNA molecules, or copies thereof, to identify the spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules, or copies thereof, and correlating the spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules, or copies thereof, with the known locations of spatial barcode sequences of the surface oligonucleotide molecules. In various embodiments, the sequence of the spatial barcodes is determined by next generation sequencing.
[0195] When surface oligonucleotide molecules are arranged in a pattern on the substrate (e.g., a flow cell), such that the substrate locations and sequences of the spatial barcodes of the surface oligonucleotides on the substrate are known prior to contacting the tissue with the flow cell, the method further comprises determining the spatial location of the RNA molecules within the tissue sample by sequencing at least a portion of one or more spatially barcoded first strand cDNA molecules, or copies thereof, to identify the spatial barcode sequences of the spatially barcoded first strand cDNA molecules, or copies thereof, and correlating the spatial barcode sequences of the spatially barcoded first strand cDNA molecules, or copies thereof, with the known locations of spatial barcode sequences of the surface oligonucleotide molecules. Optionally, the method further comprises determining the spatial locations of RNA molecules within the tissue sample by sequencing at least a portion of one or more spatially barcoded first strand cDNA molecules and correlating the spatial barcode sequences of the one or more spatially barcoded first strand cDNA molecules, or copies thereof, with one or more corresponding spatial barcode sequences of surface oligonucleotide molecules having predetermined locations on the substrate.
[0196] In some embodiments, the capture site on the substrate is a plurality of capture sites. In some embodiments, the plurality of capture sites is 2 or more, 10 or more, 30 or more, 100 or more, 300 or more, 1 ,000 or more, 3,000 or more, 10,000 or more, 30,000 or more, 100,000 or more, 300,000 or more, 1 ,000,000 or more 3,000,000 or more, or 10,000,000 or 1 ,000,000,000 or more capture sites.
[0197] In various embodiments, the capture array or substrate comprises a capture site density of 1 or more, 2 or more, 10 or more, 30 or more, 100 or more, 300 or more, 1 ,000 or more, 3,000 or more, 10,000 or more, 100,000 or more, 1 ,000,000 or more, capture sites per square centimeter (cm2). In various embodiments, the density is between about 100k / mm2to about 1000k / mm2, e.g., about 100k clusters / mm2, about 200k clusters / mm2, about 300k clusters / mm2, about 400k clusters / mm2, about 500k clusters / mm2, about 600k clusters / mm2, about 700k clusters / mm2, about 800k clusters / mm2, about 900k clusters / mm2, or about 1000k clusters / mm2.
[0198] In various embodiments, the pair of capture probes in a capture site is a plurality of pairs of capture probes. In some embodiments, the plurality of capture probes is 2 or more, 10 or more, 30 or more, 100 or more, 300 or more, 1 ,000 or more, 3,000 or more, 10,000 or more, 30,000 or more, 100,000 or more, 300,000 or more, 1 ,000,000 or more 3,000,000 or more, or 10,000,000 or more, 100,000,000 or more, or 1 ,000,000,000 or more capture probes.
[0199] In some embodiments, the pair of capture probes in a capture site of a substrate is a plurality of pairs of capture probes. In some embodiments, each RNA capture probe in the plurality of pairs of capture probes within the same capture site comprises the same spatial address sequence. In some embodiments, each RNA capture probe in the plurality of pairs of capture probes in different capture sites comprises a different spatial address sequence.
[0200] In some embodiments, the surface of the capture array is a planar surface, e.g., a glass surface. In some embodiments, the surface of the capture array comprises one or more wells. In some embodiments, the one or more wells correspond to one or more capture sites. In some embodiments, the surface of the capture array is a bead surface.
[0201] In some embodiments, the capture region in the surface capture probe is a genespecific or target -specific capture region. In some embodiments, the gene-specific or target -specific capture region in the surface capture probe comprises the sequence of a TRLISEQ™ Custom Amplicon (TSCA) oligonucleotide probe (Illumina, Inc.). For example, the gene-specific or target -specific capture regions in a plurality of second capture probes in a capture site can comprise a plurality of sequences of TSCA oligonucleotide probes.
[0202] In another embodiment, the disclosure provides for a surface or substrate that comprises the spatially addressable probes disclosed herein. In a particular embodiment, the surface or substrate comprise the spatially addressable probes disclosed herein. In a further embodiment, the surface or substrate comprises streptavidin. In yet a further embodiment, the surface or substrate comprises a plurality of oligos bound to the surface or substrate via a linkage or a reversible linkage. Examples of reversible linkages include biotin molecule(s), such as ddBio molecules. The oligos bound the bead typically comprise an adaptor sequence, such as P5 sequence or a P7 sequence. As used herein a P5 sequence comprises a sequence defined by AAT GAT ACG GCG ACC ACC GA (SEQ ID NO: 1) or AAT GAT ACG GCG ACC ACC GAG ATC TAC AC (SEQ ID NO: 2) and a P7 sequence comprises a sequence defined by CAA GCA GAA GAC GGC ATA CG (SEQ ID NO: 3) or CAA GCA GAA GAC GGC ATA CGA GAT (SEQ ID NO: 4). In some embodiments, the P5 or P7 sequence can further include a spacer polynucleotide, which may be from 1 to 20, such as 1 to 15, or 1 to 10, nucleotides, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. In some embodiments, the spacer includes 10 nucleotides. In some embodiments, the spacer includesl O nucleotides. In some embodiments, the spacer is a polyT spacer, such as a 10T spacer. Spacer nucleotides may be included at the 5' ends of polynucleotides, which may be attached to a suitable support via a linkage with the 5' end of the oligo.Attachment can be achieved through a sulfur-containing nucleophile, such as phosphorothioate, present at the 5' end of the polynucleotide. In some embodiments, the oligos will include a polyT spacer and a 5'phosphorothioate group. Thus, in someembodiments, the P5 sequence comprises 5'phosphorothioate- TTTTTTTTTTAATGATACGGCGACCACCGA-3' (SEQ ID NO: 17), and in some embodiments, the P7 sequence comprises 5' phosphorothioate- TTTTTTTTTTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 18). In certain embodiments, the oligos attached to the beads comprise an address sequence that allows for determining the x, y position of the oligo / bead when decoded. In further embodiments, the address sequence is 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides in length, or a range that includes or is between any two of the foregoing nucleotides in length. In another embodiment, the oligos attached to beads comprising a transposome hybridization region (Tsm hyb). In yet additional embodiments, the oligos comprise sequencing primer(s) site sequence(s). Examples of sequencing primer site sequences include sequences that are complementary to R1 and R2 sequencing primers from Illumina™. In further embodiments, the oligos may further comprise one or more linker sequences. In yet further embodiment, the oligos may further comprise one or more index sequences. In certain embodiments, the oligos may comprise one or more unique molecular identifier (UMI) sequences. Unique molecular identifiers (UMIs) are a type of molecular barcoding that provides error correction and increased accuracy during sequencing. These molecular barcodes are short sequences used to uniquely tag each molecule in a sample library. UMIs are used for a wide range of sequencing applications, many around PCR duplicates in DNA and cDNA. UMI deduplication is also useful for RNA- seq gene expression analysis and other quantitative sequencing methods. As noted previously, the oligos comprise moieties or sequences that can bind with specificity to polynucleotides from a biological sample (e.g., a tissue sample). As such, the oligos attached to the beads are spatially addressable probes for polynucleotides from a biological sample. The moieties or sequences that can bind with specificity to polynucleotides from a biological sample can be selected for a particular “-omic” application. For example, the oligos can comprise an oligo d(T)sequence for transcriptomics or for assay (e.g., RNA-seq assays). Alternatively, the oligos can comprise sequences to bind with genomic DNA from a biological sample for genomic applications or for assays (e.g., ATAC-seq assays). As provided in the Examples presented herein, the nanostructures can comprise multiple types of oligos that have different moieties or sequences so that the spatially addressable probes can bind specifically to two or more different types of polynucleotides from a biological sample. The use of multi types of oligos is ideally suited for multiomic or multiple assay applications.Kits
[0203] Kits and articles of manufacture are also contemplated herein. Such kits can comprise a carrier, package, or container that is compartmentalized to receive one or morecontainers such as vials, tubes, and the like, each of the container(s) comprising one of the separate elements to be used in a method described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers can be formed from a variety of materials such as glass or plastic. For example, the container(s) can comprise one or more spatially addressable probes disclosed herein, optionally in a composition or in combination with another agent (e.g., an array, a beadchip) as disclosed herein. The container(s) optionally have a sterile access port (for example the container can be an intravenous solution bag or a vial having a stopper pierceable by a hypodermic injection needle). Such kits optionally comprise an identifying description or label or instructions relating to its use in the methods described herein.
[0204] A kit contemplated herein comprises a nanostructure as described herein, and a substrate as described herein. In another embodiment, the kit comprises a nanostructure as described herein, and a splint oligonucleotide molecule as described herein Also provided is a kit comprising a nanostructure as described herein, a substrate as described herein, and splint oligonucleotides molecules as described herein. In various embodiments, the substrate comprises attached splint oligonucleotides.
[0205] A kit will typically comprise one or more additional containers, each with one or more of various materials (such as reagents, optionally in concentrated form, and / or devices) desirable from a commercial and user standpoint for use with the spatially addressable probes described herein. Non-limiting examples of such materials include, but are not limited to, buffers, diluents, filters, needles, syringes; carrier, package, container, vial and / or tube labels listing contents and / or instructions for use, and package inserts with instructions for use. A set of instructions will also typically be included.
[0206] A label can be on or associated with the container. A label can be on a container when letters, numbers or other characters forming the label are attached, molded or etched into the container itself, a label can be associated with a container when it is present within a receptacle or carrier that also holds the container, e.g., as a package insert. A label can be used to indicate that the contents are to be used for a specific spatial -omic applications. The label can also indicate directions for use of the contents, such as in the methods described herein.
[0207] The following examples are intended to illustrate but not limit the disclosure. While they are typical of those that might be used, other procedures known to those skilled in the art may alternatively be used.EXAMPLESExample 1
[0208] The genetic profile of a tissue sample may be used to diagnose and determine treatment for a subject having or at risk of having a disease as determined by the genetic profile
[0209] In order to improve capture of RNA from a tissue sample, a nanostructure that comprises multiple attachment sites for RNA released from a tissue sample is used to capture RNA.
[0210] In a proof of concept experiment, the utility of probes with multiple capture sites was demonstrated. Beads were linked with either identical 10-mer capture probes, identical 50-mer capture probes, or 5 unique 10-mer capture probes (Figure 6A). When these beads were introduced to fluorescently tagged target molecules a doubling in capture efficiency was observed in the 50-mer versus identical 10-mer beads (Figure 6B, left).
[0211] While an increase in capture efficiency is expected for a longer probe, the 5 unique 10-mer beads displayed a ~4x increase over the 50-mer and nearly an order of magnitude increase over the identical 10-mer beads (Figure 6B, left). The strength of these binding interactions are also demonstrated in the melt curves of the target hybridized capture beads (Figure 6B, right), where ~10 °C increase in melt temp is observed for the 5 unique 10-mer beads versus the identical 10-mer beads.
[0212] The increase in melt temperature when multiple probes are used has important ramifications for binding specificity and purification. mRNA capture can be carried out at elevated temperatures which decreases the chances of off-target binding. Furthermore, more stringent washes can be utilized increasing purity of the substrates which is especially important during reverse transcription.
[0213] It is understood, therefore, that this invention is not limited to the particular embodiments disclosed, but is intended to cover all modifications which are within the spirit and scope of the invention as defined by the appended claims; the above description, and / or shown in the attached drawings. Consequently only such limitations as appear in the appended claims should be placed on the disclosure.
Claims
What is claimed is:1 . A method for preparing a spatially barcoded RNA library from a tissue sample comprising,(a) permeabilizing the tissue sample;(b) contacting the tissue sample with a plurality of nanostructures, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers , wherein each of the cleavable linkers links one of the oligonucleotides to the nanoscaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety;(c) capturing RNA from the tissue sample by hybridization of the RNA with RNA capture probes on the nanostructures;(d) capturing the plurality of nanostructures comprising captured RNA on a substrate, wherein the substrate comprises a plurality of capture sites and a plurality of surface oligonucleotide molecules, wherein each of the plurality of capture sites is capable of binding to the substrate anchor moiety of the nanostructure, and wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules;(e) contacting the captured RNA with a reverse transcriptase to generate one or more first strand cDNAs wherein each of the first strand cDNAs is contiguous with one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe;(f) cleaving the cleavable linkers to release the first strand cDNAs from the nanoscaffold; and(g) capturing the first strand cDNAs on the substate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
2. The method of claim 1 , wherein the step of generating the one or more first strand cDNAs occurs in the tissue, and before capturing the plurality of nanostructures on the substrate.
3. The method of claim 1 , wherein the step of generating the one or more first strand cDNAs occurs after capturing the plurality of nanostructures on the substrate.
4. The method of any one of claims 1 to 3, further comprising f treating the nanostructures with an RNase after the step of generating the one or more cDNAs.
5. The method of any one of claims 1 to 4, further comprising, prior to the step of capturing RNA from the tissue sample, the step of performing end repair of the RNA with polynucleotide kinase.
6. The method of any one of claims 1 to 4, further comprising, prior to the step of capturing RNA from the tissue sample, the step of performing in situ polyadenylation with polyadenylate polymerase.
7. The method of any one of claims 1 to 4, further comprising, prior to the step of capturing RNA from the tissue sample, the steps of performing end repair of the RNA with polynucleotide kinase followed by performing in situ polyadenylation with polyadenylate polymerase.
8. The method of any one of claims 1 to 7, wherein the RNA comprises ribosomal RNA (rRNA), messenger RNA (mRNA), non-coding RNA (ncRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and / or microRNA (miRNA).
9. A method for preparing a spatially barcoded RNA library from a tissue sample comprising,(a) contacting a tissue sample with a substrate comprising a plurality of nanostructures and a plurality of surface oligonucleotide molecules attached to the substrate, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; and (iii) two or more cleavable linkers wherein each of the cleavable linkers links one of the oligonucleotides to the nanoscaffold, wherein the nanoscaffold comprises two or more sites for attachment of the cleavable linkers; wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules;(b) permeabilizing the tissue sample to release RNA from the tissue sample;(c) capturing RNA from the permeabilized tissue sample by hybridization of RNA with the RNA capture probe;(d) contacting the captured RNA with a reverse transcriptase to generate one or more first strand cDNAs wherein each of the first strand cDNAs is contiguous with one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe;(e) cleaving the cleavable linkers to release the first strand cDNAs from the nanoscaffold; and(f) capturing the first strand cDNAs on the substrate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
10. The method of claim 9, further comprising treating the nanostructures with an RNase after the step of generating the one or more cDNAs.11 . The method of any one of claims 9 or 10, further comprising, prior to the step of capturing RNA from the permeabilized tissue sample, the step of performing end repair of the RNA with polynucleotide kinase.
12. The method of any one of claims 9 or 10, further comprising, prior to the step of capturing RNA from the permeabilized tissue sample, the step of performing in situ polyadenylation with polyadenylate polymerase.
13. The method of any one of claims 9 or 10, further comprising, prior to the step of capturing RNA from the permeabilized tissue sample, the steps of performing end repair of the RNA with polynucleotide kinase followed by performing in situ polyadenylation with polyadenylate polymerase.
14. The method of any one of claims 9 to 13, wherein the RNA comprises rRNA, mRNA, ncRNA, snRNA, snoRNA, and / or miRNA.
15. A method for preparing a spatially barcoded RNA library from a tissue sample comprising,(a) permeabilizing the tissue sample to release RNA from the tissue sample;(b) hybridizing the RNA with a plurality of RNA capture probes to generate a plurality of RNA-RNA capture probe hybrids;(c) contacting the RNA-RNA capture probe hybrids with a plurality of nanostructures wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotideseach comprising a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety;(d) covalently attaching the 5’ capture probe ends of the RNA-RNA capture probe hybrids to the 3’ ends of the oligonucleotides on the nanostructure;(e) contacting the covalently attached RNA-RNA capture probe hybrids in step (d) with a reverse transcriptase to generate one or more first strand cDNAs, wherein each of the first strand cDNAs is contiguous with one of the oligonucleotides on the nanostructure and comprises the RNA capture probe and the surface capture probe;(f) capturing the plurality of nanostructures on a substrate wherein the substrate comprises a plurality of capture sites and a plurality of surface oligonucleotide molecules, wherein each of the plurality of capture sites is capable of binding to the substrate anchor moiety, and wherein each of the plurality of surface oligonucleotide molecules comprises an anchor sequence, a spatial barcode, and a sequence that hybridizes with a second domain of the splint oligonucleotide molecules;(g) cleaving the cleavable linkers to release the first strand cDNAs from the nanoscaffold; and(h) capturing the first strand cDNAs on the substrate by hybridizing the first strand cDNAs and the surface oligonucleotide molecules with the splint oligonucleotide molecules and ligating the first strand cDNAs to the surface oligonucleotide molecules, thereby generating spatially barcoded first strand cDNAs.
16. The method of claim 15, wherein the 5’ capture probe ends of the RNA-RNA capture probe hybrids are covalently attached to the 3’ ends of the oligonucleotides on the nanostructure by ligation, splint ligation, or click chemistry.
17. The method of claim 16, wherein the click chemistry uses azide-alkyne cycloaddition, a heterobifunctional linking group, a sulfo-NHS ester, a DBCO-NHS ester or a maleimide group.
18. The method of any one of claims 14-17, wherein the click chemistry covalent attachment is carried out by azide-alkyne cycloaddition.
19. The method of any one of claims 1-18, wherein the sample is a fresh frozen tissue sample or a formalin-fixed paraffin embedded (FFPE) sample.
20. The method of any one of claims 1-19, wherein the tissue sample is a fixed tissue sample.21 . The method of claim 20, wherein the fixed tissue sample is a formalin-fixed paraffin embedded (FFPE) tissue sample.
22. The method of claim 21 , further comprising decrosslinking the FFPE sample, optionally wherein the decrosslinking is carried out using TE buffer, pH 9. Move to correct placement.
23. The method of any one of claims 1-22, wherein the tissue sample is treated with one or more blocking reagents prior to the permeabilization step.
24. The method of any one of claims 1-23, wherein the substrate is a bead, a bead array, a spotted array, a substrate comprising a plurality of wells, a flow cell, clustered particles arranged on a surface of a chip, a film, or a plate.
25. The method of claim 24, wherein the substrate comprises a plurality of nanowells or microwells.
26. The method of any one of claims 1-25, wherein the substrate is a gel coating located in or on a flow cell.
27. The method of any one of the preceding claims wherein the nanoscaffold is a dendrimer, a nanoparticle, a nanogel or a hyperbranched polymer.
28. The method of any one of the preceding claims wherein the nanoscaffold comprises between 2 and 32 attachments sites for the cleavable linkers.
29. The method of claim 27 or claim 28, wherein the nanostructure is a dendrimer.
30. The method of claim 29, wherein the dendrimer is a peptide dendrimer comprising an alkyne end for anchoring to a substrate and an amine terminus.31 . The method of claim 29 of 30, wherein the dendrimer comprises polylysine, branched lysine, or a polyamine.
32. The method of any one of claims 1 -31 , wherein the RNA capture probe is selected from the group consisting of a poly-T sequence, a poly-U sequence, a randomer, a semi-random sequence, or a target-specific probe.
33. The method of claim 32, wherein the RNA capture probe is a poly-T sequence.
34. The method of claim 32 or 33, wherein the RNA capture probe comprises at least 10 deoxythymidine residues.
35. The method of claim 32, wherein the target-specific probes comprise a plurality of different target-specific RNA capture probe sequences.
36. The method of claim 35, wherein the target-specific probes comprise at least 10 nucleotides complementary to a nucleotide sequence of a target RNA.
37. The method of claim 34 or 35, wherein the RNA capture probe or surface capture probe is between 8 to 80 nucleotides.
38. The method of any one of claims 1-37, wherein the cleavable linker is a cleavable polynucleotide.
39. The method of any one of the preceding claims, wherein the surface oligonucleotide molecules further comprise a primer binding nucleotide sequence.
40. The method of claim 39, wherein the primer binding nucleotide sequence comprises a P7 nucleotide sequence.41 . The method of any one of the preceding claims, wherein the sequence that hybridizes with the second domain of the splint oligonucleotide molecules comprises a PZ nucleotide sequence.
42. The method of claim 41 , wherein the second domain of the splint oligonucleotide molecules comprises a nucleotide sequence PZ’ that is complementary to the PZ sequence.
43. The method of any one of the preceding claims, wherein the sequence that hybridizes with the first domain of the splint oligonucleotide molecules comprises a PX nucleotide sequence.
44. The method of claim any one of the preceding claims, wherein the first domain of the splint oligonucleotide molecules comprises a nucleotide sequence PX’ that is complementary to the PX sequence.
45. The method of any one of the preceding claims, wherein the tissue sample is contacted with a mechanism to accelerate RNA diffusion from the sample.
46. The method of claim 45, wherein the mechanism is magnetic or electrophoretic.
47. The method of any one of the preceding claims, wherein each of the two or more oligonucleotides further comprise a single molecular identifier (SMI) barcode.
48. The method of any one of the preceding claims, wherein each of the two or more oligonucleotides further comprise a unique molecular identifier (UMI) barcode.
49. The method of any one of the preceding claims, wherein each of the surface oligonucleotide molecules further comprises a SMI barcode.
50. The method of any one of the preceding claims, wherein each of the surface oligonucleotide molecules further comprises a UMI barcode.51 . The method of any one of claims 1 to 50, further comprising determining spatial locations of the spatial barcodes of the plurality of surface oligonucleotide molecules prior to the step of contacting the tissue with the substrate.
52. The method of claim 51 , further comprising sequencing at least a portion of the spatially barcoded first strand cDNA molecules or copies thereof to determine the spatial barcode sequence for each molecule.
53. The method of claim 52, wherein the spatially barcoded first strand cDNA molecules are sequenced in situ.
54. The method of claim 52 or 53, further comprising determining the spatial location of one or more of the spatially barcoded first strand cDNA molecules or copies thereof by correlating the spatial barcode sequences of the spatially barcoded first strand cDNA molecules or copies thereof with the spatial locations of the surface oligonucleotide molecules on the substrate containing corresponding spatial barcode sequences.
55. The method of claim 54, further comprising recovering the spatially barcoded first strand cDNA molecules and amplifying them to generate cDNA libraries.
56. The method of claim 55, wherein the spatially barcoded first strand cDNA molecules are recovered by contacting the spatially barcoded first strand cDNAs on the substrate with a DNA polymerase and one or more primers to generate spatially barcoded second strand cDNAs complementary to the spatially barcoded first strand cDNAs and removing the spatially barcoded second strand cDNAs from the substrate.
57. The method of claim 56, wherein the one or more primers each comprise a random priming sequence.
58. The method of claim 57, wherein the random priming sequences comprises nine random nucleotides.
59. The method of claim 57 or 58, wherein the spatially barcoded second strand cDNAs each comprise a unique molecular identifier (UMI), wherein the UMI comprises an intrinsic sequence and an extrinsic sequence, wherein the extrinsic sequence is a sequence complementary to the random priming sequence used to generate the second strand cDNA, and wherein the intrinsic sequence is a sequence complementary to the first strand cDNA template sequence used to generate the second strand cDNA.
60. The method of claim 56, wherein the one or more primers each comprise a molecular identifier barcode.61 . The method of claim 56, wherein the one or more primers each comprise aUMI barcode.
62. The method of any one of claims 56-61 , wherein the spatially barcoded second strand cDNAs are removed from the substrate by chemical or physical dehybridization.
63. The method of claim 56, wherein the anchor sequence comprises a cleavage site, and hybrids of the spatially barcoded first and second strand cDNAs are removed from the substrate by enzymatic cleavage at the cleavage site.
64. The method of claim 63, wherein the cleavage site is a binding site for a restriction endonuclease.
65. The method of claim 55, wherein the anchor sequence comprises a cleavage site, and wherein the spatially barcoded first strand cDNA molecules are recovered by enzymatic cleavage at the cleavage site.
66. The method of claim 65, wherein the cleavage site is a binding site for a restriction endonuclease.
67. The method of any one of claims 55-66, further comprising sequencing at least a portion of the cDNA libraries to determine the spatial barcode sequence for each molecule.
68. The method of claim 67, further comprising determining the spatial location of one or more cDNA molecules by correlating the spatial barcode sequences of the one or more cDNA molecules with the spatial locations of the surface oligonucleotide molecules on the substrate containing corresponding spatial barcode sequences.
69. The method of any one of claims 1 to 68 further comprising indexing and sequencing spatially barcoded first strand cDNAs, comprising, performing extension reactions and PCR on the spatially barcoded first strand cDNAs to yield a PCR template comprising a first strand PCR product representative of one or more RNA transcripts in the tissue sample; eluting the PCR template;carrying out an indexing PCR to generate a double stranded PCR product comprising the first strand PCR product and a second strand complementary to the first strand PCR product.
70. The method of claim 69 further comprising sequencing the PCR product and determining the location of the RNA transcript in the tissue based on the spatial barcode of first strand cDNA.71 . The method of claim 69 or 70, wherein the double stranded PCR product comprises a second clustering sequence on the second strand complementary to the first strand PCR product and, optionally, an index sequence.
72. The method of claim 70 or 71 , wherein the PCR products are further processed by tagmentation to generate a spatial transcriptomics library.
73. The method of claim 72 wherein the tagmentation comprises on substrate tagmentation.
74. The method of any one of claims 1 to 73, wherein the methods determine RNA expression in a single cell with the tissue sample.
75. The method of claim 74, wherein the methods determine RNA expression in one or more subcellular components in the single cell.
76. The method of claim 75, wherein the subcellular component is a cell nucleus, cytoplasm, or mitochondria.
77. The method of any one of the preceding claims, wherein the substrate or surface of the substrate comprises a material selected from glass, silicon, poly-L-lysine coated materials, nitrocellulose, polystyrene, cyclic olefin copolymers (COCs), cyclic olefin polymers (COPs), polyacrylamide, polypropylene, polyethylene, or polycarbonate.
78. A composition comprising a nanostructure, wherein each nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein eachof the cleavable linkers links one of the oligonucleotides to the nanoscaffold, and wherein the nanoscaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety.
79. A composition comprising a nanostructure, wherein the nanostructure comprises: (i) a nanoscaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nanoscaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety, wherein the nanostructure is attached to a substrate.
80. A composition comprising a nanostructure, wherein the nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety, wherein the nanostructure is attached to a substrate comprising surface oligonucleotide molecules.81 . The composition of claims 80, wherein the substrate contains between 1 x 109and 1 x 1011per mm2surface oligonucleotide molecules.
82. The method of claim 80 or 81 , wherein the surface oligonucleotide molecules are arranged in patterns or clusters.
83. The method of any one of claims 80 to 82, wherein the ratio of surface oligonucleotide molecules to capture sites is about 1 :2 to about 1 :100.
84. The method of any one of claims 80 to 83, wherein the diameter of each cluster is from about 500 nm to about 2,000 nm.
85. The method of any one of claims 1 to 84, wherein the surface oligonucleotide molecules comprise a cleavage domain.
86. The method of claim 85, wherein the cleavage domain is a binding site for a restriction endonuclease or a site for chemical cleavage.
87. A kit comprising a nanostructure, wherein the nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety, a substrate, and splint oligonucleotide molecules.
88. A kit comprising a nanostructure, wherein the nanostructure comprises: (i) a nano-scaffold; (ii) two or more oligonucleotides each comprising an RNA capture probe and a surface capture probe capable of hybridizing with a first domain of a plurality of splint oligonucleotide molecules; (iii) two or more cleavable linkers, wherein each of the cleavable linkers links one of the oligonucleotides to the nano-scaffold, and wherein the nano-scaffold comprises two or more sites for attachment of the cleavable linkers; and (iv) a substrate anchor moiety, and splint oligonucleotide molecules.
Citation Information
Patent Citations
Determining Antigen Recognition through Barcoding of MHC Multimers
US20170343545A1
Methods and kits using nucleic acid encoding and / or label
US20210355483A1
Spatial analysis of analytes
WO2021102039A1
Methods for capturing library DNA for sequencing
WO2023069927A1
Analysis of antigen and antigen receptor interactions
WO2023215612A1