METHOD FOR OBTAINING SPATIAL INFORMATION AND SEQUENCING INFORMATION OF m-RNA FROM TISSUE
Patent Information
- Application Number
- JP2022151013
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-23
- Filing Date
- 2022-09-22
- Publication Date
- 2025-09-30
AI Technical Summary
Existing methods for sequencing and spatial decoding of tissue m-RNA molecules struggle to achieve subcellular resolution due to limitations in spatial identifier resolution, which is currently at the multicellular level, and single-molecule decoding remains a challenging task.
A method involving a solid surface with fiducial markers, photocleavable linkers, and scaffold molecules to generate and decode DNA barcodes, incorporating nucleobases with photodetectable units, and using sequencing-by-synthesis to determine spatial positions and sequences simultaneously.
Enables high spatial resolution sequencing of m-RNA molecules at 50-300 nm, generating unique spatial identifiers for each molecule, allowing for precise spatial and sequence information detection.
Smart Images

Figure 00000008_0000
Abstract
Description
[Technical Field]
[0001] background The present invention relates to a method for sequencing and spatial decoding of mRNA molecules in tissues.
[0002] By maintaining spatial information of these genes and subsequent alignment for potential variants, it has been difficult to detect all genes expressed at the subcellular level on tissues, and it is particularly difficult to reach subcellular resolution.
[0003] For example, the development of a sequencing system for spatial decoding of DNA barcode molecules is described by Yusuke Oguchi in “Single-molecule resolution” (Nature, Dec 2020).
[0004] Commercial techniques are available in which tissue RNA molecules are tagged with spatial identifiers pre-spotted onto arrays, preserving spatial information. The resulting libraries are subsequently sequenced by standard in vitro NGS (next-generation sequencing) sequencing techniques. The spotting process ensures that the locations of the spatial identifiers on the array are known before the sequencing process is performed. After sequencing of the spatial identifiers, the concatenated RNA sequences of interest can be assigned to tissue locations. One major limitation of this technique is the resolution of the technique, as it currently relies on the feature size of the spots on the array, which is only at the multicellular level.
[0005] Techniques for obtaining sequencing and spatial information are published, for example, in U.S. Patent Application Publication No. 20150344942, WO 2016162309, and WO 2012 / 140224.
[0006] Single-cell transcriptome analysis has been revolutionized by DNA barcodes that index cDNA libraries, enabling highly multiplexed analyses. Furthermore, DNA barcodes have been utilized for spatial transcriptome analysis. While spatial resolution depends on the method used to decode the DNA barcode, achieving single-molecule (mRNA) decoding remains a challenging task.
[0007] overview It was therefore an object of the present invention to provide a method for sequencing and spatial decoding of m-RNA molecules in tissues with simultaneous generation and decoding of DNA barcodes.
[0008] The subject of the present invention is therefore a method for obtaining spatial location and sequence information of mRNA target sequences on a tissue sample, comprising the following steps: a. Providing a solid surface having at least one fiducial marker b. Attaching a plurality of anchor molecules comprising a photocleavable linker and an adaptor unit to a solid surface c. Binding a scaffold molecule to the anchor molecule, the scaffold molecule comprising a unit capable of binding to the adaptor unit of the anchor molecule, a polyinosine unit having 5 to 30 inosine bases, and a polyadenine unit having 10 to 50 adenine bases. d. Optionally, during sequencing-by-synthesis (SBS), randomly incorporating adenine, guanine, cytosine, and thymine as nucleobases into the anchor molecule complementary to the inosine bases of the scaffold molecule, thereby generating barcodes on the anchor molecule, wherein the nucleobases comprise optically detectable units, and the sequence of the barcodes (and optionally their generation) and their spatial location relative to the fiducial markers are simultaneously detected as spatial information. e. incorporating thymine into the anchor molecule complementary to the polyadenine unit of the scaffold molecule, thereby generating a polyT unit. f. Removing the scaffold molecule from the anchor molecule g. Providing a tissue sample containing at least one mRNA strand, wherein at least one mRNA strand of the sample binds to a polyT unit of at least one anchor molecule. reverse transcription of the hm-RNA strand to generate a c-DNA strand attached to a solid surface; i. Removing the c-DNA strand from the solid surface by cleaving the photocleavable linker of the anchor molecule. A process of obtaining sequence information of jc-DNA strands and linking the spatial information with sequence information of c-DNA strands. The method includes: [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 illustrates the general workflow of the present invention. [Figure 2] FIG. 1 shows the construction of a barcode portion in an anchor molecule.
[0010] Detailed Description The method described herein has been used to detect mRNA on tissues with very high spatial resolution of 50-300 nm.
[0011] The method of the invention is explained in more detail below according to the process steps which must be carried out in the order a) to j).
[0012] Step a) Providing a solid surface having at least one fiducial marker.
[0013] The solid substrate preferably comprises a functionalized surface, which may be functionalized with commonly used chemistries available for solid support attachment of oligonucleotides containing photocleavable linkers, and optionally assembled with other moieties to form a microfluidic flow cell having an inlet and an outlet.
[0014] Functionalized surfaces can include, for example, amine-modified oligomers covalently attached to activated carboxylate groups or succinimidyl esters, thiol-modified oligomers covalently attached via alkylating reagents such as iodoacetamide or maleimide, Acrydite™-modified oligomers covalently attached via thioethers, digoxigenin NHS ester, or biotin-modified oligomers captured by immobilized streptavidin capable of interacting with biotin.
[0015] Optionally, the anchor molecules are randomly distributed on the solid substrate at a density corresponding to the molecule concentration.
[0016] The anchor molecule may further comprise at least one primer sequence for a polymerase capable of rolling circle amplification.
[0017] Step b) attaching a plurality of anchor molecules comprising a photocleavable linker to the functionalized solid surface.
[0018] Prefabricated anchor molecules are then loaded onto the solid surface in a buffer solution, where they interact with the functionalized surface and become randomly spatially distributed throughout the area.
[0019] Preferably, anchor molecules are generated with photocleavable linkers, adapter units such as P5 or P7 adapters, or other short sequences. Billions of such single molecules can be produced. The density of anchor molecules can be controlled by the loading concentration.
[0020] Step c) binding a scaffold molecule to the anchor molecule, the scaffold molecule comprising a unit capable of binding to the adaptor unit of the anchor molecule, a polyinosine unit having 5 to 30 inosine bases, and a polyadenine unit having 10 to 50 adenine bases.
[0021] The scaffold molecule is shown in Figures 1a and 1b and is attached to the anchor molecule via an adaptor unit. Suitable adaptors such as P5 or P7 and their counterparts are known to those skilled in the art.
[0022] The poly-inosine unit serves as a template for the generation of randomized barcodes and may preferably have between 5 and 30 inosine bases.
[0023] Step d) optionally during sequencing-by-synthesis (SBS), randomly incorporating adenine, guanine, cytosine and thymine as nucleobases into the anchor molecule that are complementary to the inosine bases of the scaffold molecule, thereby generating barcodes on the anchor molecule, wherein the nucleobases are equipped with optically detectable units, and the sequence of the barcodes (and optionally their generation) and their spatial position relative to the reference markers are simultaneously detected as spatial information.
[0024] Preferably, the step of randomly incorporating adenine, guanine, cytosine and thymine as nucleobases into the anchor molecule is carried out by supplying a mixture of A, T, G and C and a polymerase.
[0025] During this process, random barcodes are generated on the anchor molecules, which are used to simultaneously identify and assign their spatial location relative to the fiducial markers.
[0026] In this process, the inosine bases of the scaffold molecule are randomly complemented with adenine, guanine, cytosine, and thymine (A, C, G, T) to generate a barcode sequence. For this purpose, the nucleobases are equipped with optically detectable units. Such nucleobases are known as sequencing-by-synthesis and are commercially available. The random sequence is established by supplying a mixture of A, T, G, and C to the sample along with polymerase.
[0027] With a photodetectable unit, the generation of the barcode can be easily monitored by taking appropriate images.
[0028] Spatial information is preferably obtained by detecting optically detectable units of the nucleobases and / or by sequencing-by-synthesis.
[0029] Images are taken after each extension cycle to determine the incorporated base, as shown in Figure 2. The images may have super-resolution to identify each barcode.
[0030] The bases A, C, G, and T are randomly incorporated into each individual single molecule in each cycle, thus generating a unique spatial single-molecule identifier. The solid substrate also contains at least two independent fiducial marks, which may be fluorescent or autofluorescent markers. Images of these fiducial marks are also taken, and the XY positions of the incorporated bases are determined relative to the markers. In this way, the sequence of each individual single molecule is recorded, and the spatial positions of the incorporated bases relative to the fiducial marks are stored as special information (XY distances).
[0031] Step e) incorporating thymine into the anchor molecule complementary to the polyadenine unit of the scaffold molecule, thereby generating a polyT unit.
[0032] After sequencing, the bases of the polyA tail are filled with unlabeled T nucleotides, which allows the formation of a stable double-stranded DNA molecule.
[0033] Of course, the mixture of bases from the previous step must be removed, and optionally after washing, thymine and polymerase are provided.
[0034] Step f) Removing the scaffold molecule from the anchor molecule
[0035] In the next step, the solid substrate with the double-stranded DNA complex is heat denatured to remove the scaffolding molecules and form single-stranded oligonucleotides on the solid surface.
[0036] Step g) providing a tissue sample containing at least one m-RNA strand, wherein at least one m-RNA strand of the sample binds to the polyT unit of at least one anchor molecule.
[0037] The tissue sample is contacted with a solid substrate on which every single molecule is placed and permeabilized to release the mRNA molecules. The 30-50 base poly-T tail of the adjacent oligonucleotide then hybridizes to the mRNA expressed in the tissue via canonical DNA / mRNA interactions.
[0038] Optionally, the tissue sample is permeabilized after being delivered to the surface.
[0039] Step h) reverse transcription of the m-RNA strand to produce a c-DNA strand attached to a solid surface.
[0040] In the next step, all anchor molecules bearing mRNA are reverse transcribed into cDNA, after which the tissue can be removed from the solid substrate using an enzyme or via a chemical reaction.
[0041] The c-DNA strand may be circularized and then amplified into multiple DNA concatemers by a polymerase capable of rolling circle amplification.
[0042] Optionally, only a subset of mRNA (region of interest or ROI) can be reverse transcribed if only a predetermined subset of anchors is maintained by removing unwanted regions using targeted light prior to cDNA generation.
[0043] Regions of interest (ROIs) in tissue can be defined using an external laser UV light source with a digital micromirror device (DMD). Each ROI receives a specific dose of UV laser light, resulting in the cleavage and removal of all photocleavable bonds in single molecules bound to the solid substrate.
[0044] Step i) Removing the c-DNA strand from the solid surface by cleaving all photocleavable linkers of the anchor molecule.
[0045] Optionally, as described above, an external laser UV light source targeted to the region of interest on the solid surface can be used to remove only a subset of the cDNA.
[0046] Step j) obtaining sequence information of the c-DNA strand and linking the spatial barcode sequence information of the c-DNA strand to a solid surface barcode.
[0047] Further processes may be performed to prepare cDNA for a sequencing library. In the final step, each single molecule with a unique spatial anchor molecular identifier is sequenced in vitro and linked to its original, determined location on the tissue. Since the barcode sequences and their spatial locations relative to the fiducial markers are known from step d), the spatial information of the mRNA can be easily determined. Since the sequence information of the c-DNA strand contains the barcode sequence, and the sequence information of the barcode generated on the solid surface is known from steps d and i), the two barcode sequences can be matched to identify the origin / spatial information of the mRNA on the tissue.
[0048] The method of the present invention is ideal for sequencing the 3' end of captured mRNA.
[0049] In many applications of mRNA sequencing, it is also necessary to identify the 5' end of the mRNA. The following workflow (also applicable to solid supports) solves this problem, as visualized in Figure 2. B reverses to a universal base such as inosine.
[0050] Optionally, the tissue may be further characterized by identification of the different proteins expressed on the tissue, for example using antibody-conjugated dyes, preferably using MACSima technology.
[0051] The method of the present invention can theoretically produce 30 different DNA barcode molecules with an average read length of approximately 20 nt, with an error rate of less than 5% per nucleotide, which is sufficient for spatial identification. Furthermore, the spatially identified DNA barcode molecules bound to antibodies can be detected with single molecule resolution.
[0052] Example workflow: An oligonucleotide identical to the anchor molecule consisting of a photocleavable linker with a P5 adaptor having 20-30 bases of inosine and 30-50 bases of polyA base tail is loaded onto a functionalized solid surface (e.g., a standard coverslip 25 x 75 x 1 mm).
[0053] The anchor molecules become randomly immobilized on the functionalized surface, and the density of the anchor molecules on the surface is controlled by the concentration.
[0054] The anchor molecules may have a minimum distance of about 50 to 300 nm from each other.
[0055] The 25-30 bases of inosine allow for the theoretical 1.2 x 10^15 - 1.5 x 10^18 possible combinations by randomly incorporating labeled single nucleotides consisting of A, C, T, and G when performing sequencing by synthesis (SBS).
[0056] Sequencing-by-synthesis is performed for 25-30 cycles to determine each base incorporated into each single molecule of the inosine's complementary base. Figure 1 shows a sequencing cycle. This shows the extension step in which the first round of labeled nucleotides is incorporated into the inosine base.
[0057] Sequencing-by-synthesis can be performed with dedicated super-resolution optics. The positional information for each single molecule on the tissue slice holder is written to a database, which can be downloaded by the user (website, cloud) or shipped with consumables.
[0058] After sequencing all the bases of inosine, each single molecule will have a spatially unique molecular identifier.
[0059] The location of each single molecule, bearing a spatially unique molecular identifier, can be determined relative to a relative position (e.g., a fiducial marker) on the functionalized solid surface.
[0060] The solid substrate also contains two independent fiducial marks, whose images are taken and their XY positions are determined. Each sequence of each individual single molecule is recorded and its spatial position relative to the fiducial marks is stored as an XY coordinate.
[0061] After 25-30 bases of sequencing, the polyA tail (approximately 30 bases) is filled with unlabeled T nucleotides. This process results in the formation of a stable double-stranded DNA molecule:
[0062] As a next step, the solid substrate with the double-stranded molecules is denatured, whereby the double-stranded DNA molecules form single-stranded oligonucleotides covalently attached to the surface.
[0063] The tissue sample is contacted with a solid substrate containing single molecules and permeabilized. mRNA expressed within the tissue is released from the tissue after permeabilization, and their poly(A) tails interact with 30-50 base poly(T) tails. In the next step, the tissue is removed from the solid substrate.
[0064] The mRNA will be reverse transcribed using the single molecule anchored on the surface as a primer.
[0065] The uniquely generated, sequenced, and pre-decoded barcode-containing cDNA, linked to the solid surface via a photocleavable linker, is released using a targeted external laser UV light source from the solid surface for further processing.
[0066] Further processing is performed to prepare the cDNA for sequencing libraries using approaches commonly known as end-repair, A-tailing, and adapter ligation.
[0067] In the final step, each cDNA with a unique spatial single-molecule identifier is sequenced in vitro and ligated to its original, determined location on the tissue.
[0068] Optionally, one can also further characterize the tissue sample and identify several different proteins expressed on the tissue, for example, using MACSima technology.
[0069] Furthermore, regions of interest (ROIs) on tissues can be defined by using an external UV laser light source with a digital micromirror device (DMD) to remove unwanted portions available for mRNA hybridization, removing all photocleavable bonds of undesired single molecules bound to the solid substrate. Conversely, cDNAs of interest can also be released via the same principle, but only from predetermined ROIs after cDNA generation.
Claims
1. A method for obtaining spatial location and sequence information of mRNA target sequences on a tissue sample, comprising the steps of: a. Providing a solid surface having at least one fiducial marker thereon; b. Attaching a plurality of anchor molecules comprising a photocleavable linker and an adaptor unit to the solid surface. c) Binding a scaffold molecule to the anchor molecule, the scaffold molecule comprising a unit capable of binding to the adaptor unit of the anchor molecule, a polyinosine unit having 5 to 30 inosine bases, and a polyadenine unit having 10 to 50 adenine bases. d. Randomly incorporating adenine, guanine, cytosine, and thymine as nucleobases into the anchor molecule complementary to the inosine bases of the scaffold molecule, thereby generating a barcode on the anchor molecule, wherein the nucleobases are provided with optically detectable units, and the sequence of the barcode and its spatial position relative to the fiducial marker are simultaneously detected as spatial information. e. incorporating thymine into the anchor molecule complementary to the polyadenine unit of the scaffold molecule, thereby generating a polyT unit. f. Removing the scaffold molecule from the anchor molecule g. Providing a tissue sample containing at least one mRNA strand, wherein at least one mRNA strand of said sample binds to a polyT unit of at least one anchor molecule. h. reverse transcribing the mRNA strand to produce a c-DNA strand attached to the solid surface. i. Removing the c-DNA strand from the solid surface by cleaving the photocleavable linker of the anchor molecule. j) Obtaining sequence information of the c-DNA strand and linking the spatial information to the sequence information of the c-DNA strand A method comprising:
2. 2. The method according to claim 1, wherein the step of randomly incorporating adenine, guanine, cytosine and thymine as nucleic acid bases into the anchor molecule is carried out by supplying a mixture of A, T, G and C and a polymerase.
3. 3. The method according to claim 1, wherein the spatial information is obtained by detecting the optically detectable units of the nucleic acid bases.
4. 3. The method according to claim 1, wherein the spatial information is obtained by sequencing by synthesis.
5. 3. The method of claim 1, further comprising removing the tissue sample from the surface.
6. 3. The method according to claim 1, wherein the c-DNA strand is circularized and then amplified into a plurality of DNA concatemers by a polymerase capable of rolling circle amplification.
7. 3. The method according to claim 1, wherein the anchor molecules are randomly distributed on the solid substrate at a density corresponding to the molecular concentration.
8. 3. The method according to claim 1, wherein the tissue sample is permeabilized after being applied to the surface.
9. 3. The method according to claim 1, wherein the anchor molecule comprises at least one primer sequence for a polymerase capable of rolling circle amplification.