UNIT-DNA composition for spatial barcode typing and sequencing
Patent Information
- Application Number
- JP2023579848
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-01
- Filing Date
- 2022-06-30
- Publication Date
- 2025-07-08
AI Technical Summary
Existing spatial sequencing methods face limitations in the number of mRNA sequences that can be measured in situ due to high density of rolonies, which compromises single-cell analysis and RNA capture efficiency, and large capture spots limit resolution.
The method employs spatial barcoding of nucleic acids using universal template-directed DNA synthesis (UNIT-DNA) with photocleavable blocking groups, allowing high-resolution optical coding and decoding of mRNA distribution in tissues, enabling unique barcodes for each cell.
This approach significantly enhances the number of mRNA sequences measurable in situ, achieving single-cell analysis with high resolution and efficient RNA capture, overcoming the limitations of existing methods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the technique of spatial sequencing. The aim of the present invention is to determine the distribution of mRNA in tissue regions or in individual cells within a tissue.
[0002] Spatial sequencing is a general term for methods that allow direct sequencing of the mRNA content of cells within tissues. These methods are useful, on the one hand, for analyzing the mRNA expression profiles of cells in a type of highly multiplexed fluorescent in situ hybridization (FISH) assay.
[0003] On the other hand, in situ sequencing can also read mRNA sequence information using specific mRNA-binding probes that can incorporate copies of predefined portions of specific mRNA or cDNA sequences ("Gap-fill padlock probes", Ke et al., Nature Methods 2013, doi:10.1038 / nmeth.2563). Recently, in situ genome sequencing (IGS) approaches have also been presented ("In situ genome sequencing resolves DNA sequence and structure in intact biological samples", AC Payne et al., Science 10.1126 / science.aay3446 (2020)). All in situ sequencing methods require a signal amplification step, which is most often performed by circularization of mRNA or cDNA-bound probes or gDNA inserts by hairpin ligation and subsequent rolling circle amplification (RCA), creating DNA molecules containing multiple copies of the probe and / or target sequence, so-called nanoballs, Rolonys or rolling circle amplification products (RCPs). As these are large molecules with sizes on the nm or μm scale, the number of Rolonys that can form in one cell is strictly limited by the size of this cell.
[0004] Furthermore, if the density of LOLONY in cells is too high, the discrimination of single mRNA signals during the optical detection step of the sequencing procedure is greatly impaired. This is a major drawback of this technique, so to avoid this, various techniques have been developed, such as the design of smaller LOLONY or the generation and removal of tissue-hydrogel complexes (Asp et al., BioEssays 2020, DOI:10.1002 / bies.201900221), or extending the cellular target, called extended sequencing (Alon et al., Science 371, eaax2656(2021)).
[0005] However, these methods still cannot completely circumvent the inherent spatial limitations of in situ sequencing. Other approaches avoid signal amplification in situ: in situ capture relies on the transfer of mRNA molecules from tissue onto a surface coated with spots of barcoded primers, allowing backtracking of sequence information obtained ex situ to the specific tissue area where the sequenced mRNA was extracted. Nevertheless, this method is also limited by the relatively large size of the barcoded capture spots, which limits the RNA capture efficiency and the low resolution (not allowing single-cell analysis) (Asp et al., BioEssays 2020, D01:10.1002 / bies.201900221).
[0006] Summary of the Invention The present invention relates to a method for inserting barcodes into polynucleotides, preferably DNA sequences, using optical methods, which can be used to obtain the location where the coding was performed. The aim here is to largely overcome the limitations of existing in situ sequencing methods regarding the number and expression dynamics of mRNA sequences that can be measured in situ. Optical coding can have a resolution in the range of 1 μm, and a variability of the code that is sufficient for each cell in a typical size tissue section to receive a unique code. The coding and decoding workflow provided is shown in Figure 1.
[0007] The basic principle disclosed herein is based on spatial barcoding of nucleic acids by universal template directed DNA synthesis. The method is subsequently referred to as UNIT-DNA ( UNI Versal T emplate DNA )
[0008] An object of the present invention is therefore a method for providing a polynucleotide comprising a first strand and a second strand carrying a barcode nucleotide sequence, characterized in that the first strand provides at least one universal base overhang at its 5' end and the corresponding recessed 3' end of the second strand of the polynucleotide carrying at least one nucleotide provides a blocking group, wherein the blocking group is removed from the incorporated nucleotide by irradiation with light.
[0009] Removal of blocking groups may be achieved by providing the blocking groups with suitable photocleavable units or by adding cleavage reagents which are activated by irradiation with light or provided in their activated form.
[0010] In a first variant of the method, the blocking group is removed from the incorporated nucleotide by irradiation with light by providing a cleavage reagent, where the cleavage reagent is provided by irradiating a precursor of the cleavage reagent with light.
[0011] In a second variant of the method, the nucleotide is provided with a photocleavable blocking group which is removed from the incorporated nucleotide by irradiation with light. Such reagents are known, for example "Cy5-TECP", which is cleaved by irradiation with light into the active cleavage reagent "Cy5".
[0012] In the following, the term "polynucleotide" refers to double-stranded nucleic acids, such as DNA, RNA, DNA-RNA, c-DNA, ssDNA and the like, such as PNA or LNA. [Brief description of the drawings]
[0013] [Figure 1] FIG. 1 shows the general workflow of the present invention, including sample preparation, tissue staining and imaging (transmission, fluorescence), segmentation or cluster analysis, and calculation of masks for structured illumination and single cell sequencing. [Diagram 2] FIG. 2 shows the general structure of a double-stranded polynucleotide molecule having at least one 5′ overhang with at least one universal base. [Diagram 3] FIG. 3 shows the basic steps of the barcoding method of the present invention. [Figure 4] Structured illumination of three cells with light. [Diagram 5] FIG. 5 shows embodiments of the present invention designated AH. [Figure 6] FIG. 6 shows the combination of a template switch oligonucleotide (TSO) with embodiment H. [Figure 7] FIG. 7 shows the general workflow of embodiment H of the present invention for in vitro sequencing with tagging molecules used directly for sequencing. [Figure 8] FIG. 8 shows the combination of circular ssDNA with embodiment H of UNIT-DNA. [Figure 9] FIG. 9 shows the combination of target DNA amplification with embodiment H of UNIT-DNA.
[0014] Detailed Description The UNIT-DNA workflow provided is shown in Figure 1 for the coding workflow according to embodiment A. Sample preparation, tissue staining and imaging (transmission, fluorescence) are followed by segmentation or cluster analysis and calculation of masks for structured illumination. For the decoding workflow of Figure 1 with UNIT-DNA, spatial barcoding of cells is provided in tissue sections induced by structured illumination and periodic incorporation of nucleotides (fixation of tissue with encoded UNIT-DNA before cell release is not shown). Release of cells from tissue sections and encapsulation of cells for single cell sequencing.
[0015] For the decoding workflow shown in Figure 7 according to embodiment H in UNIT-DNA, spatial barcoding of target molecules induced by structured illumination and periodic incorporation of nucleotides is provided to tissue sections. The sequences and spatial barcodes obtained for the target mRNAs are annotated to the images. Bioinformatics analysis and correlation of the results with the original sample source completes the workflow.
[0016] The UNIT-DNA method comprises providing a double-stranded DNA molecule comprising a first strand having a barcode nucleotide sequence and a second strand having at least one 5' overhang, the 5' overhang comprising at least one universal base and a recessed 3' end having a free 3'-OH.
[0017] Figure 2 shows an example of a composition obtained by the method of the invention for a 9 bp double-stranded nucleic acid (N) and a 6 universal base (B) 5' overhang. The number of bp for the double-stranded nucleic acid, or the number of universal bases for the single-stranded 5' overhang, varies in length.
[0018] The term "universal base" refers to a nucleotide that can bind to all natural nucleotides. The design of such universal bases has been described in the literature, mainly as a part of degenerate primers or probes, due to their property of pairing with all natural bases (e.g., Loakes, Nucleic Acid Research, 2001, Vol. 29, No. 122437-2447).
[0019] The UNIT-DNA composition obtained by the method of the present invention is shown in Figure 2 and consists of a double-stranded nucleic acid with at least one 5' overhang, where the 5' overhang contains at least one universal base and the recessed 3' end has a free 3'-OH. Here, as an example, a UNIT-DNA composition for a 9bp double-stranded nucleic acid and 5 universal base 5' overhang is shown. N represents a natural base (G or C or T or A) and B represents a universal base.
[0020] Figure 3 shows the basic principle of coding by the UNIT-DNA method. With the support of a polymerase (not shown), a provided nucleotide (G) is incorporated into the recessed free 3'-OH of the second strand, thereby extending the recessed 3'-OH of the second strand and encoding it with the first nucleotide. Any provided nucleotide will pair with the composition as long as the 5' overhang of the first strand contains a universal base that supports universal template-dependent DNA synthesis. The recessed 3'-OH is extended by one nucleotide because the 3' end of the nucleotide is blocked. As a next step, the blocked 3'-OH is unblocked by a cleavage agent, allowing the next coding cycle to occur.
[0021] The incorporation of optionally fluorescently labeled 3'-OH blocked nucleotides, which are subsequently unblocked by cleavage reagents, is also known from Sequencing by Synthesis (Chen et al., Genomics, Proteomics & Bioinformatics, Volume 11, Issue 1, February 2013, Pages 34-40). In contrast to Sequencing by Synthesis, the UNIT-DNA process is used to write the DNA code rather than to read it.
[0022] To write spatial polynucleotide barcodes, structured illumination may be used as part of the coding workflow already conceptually introduced by Figure 1. Figure 4 details how structured illumination of individual cells with light (e.g., UV light) spatially chemically releases a cleavage reagent (e.g., TCEP). The chemistry that releases TCEP after irradiation of a Cy5-TCEP complex with UV light has been described previously (Vaughan et al., J Am Chem Soc., 2013 Jan 30, 135(4), 1197-1200).
[0023] In summary, the sequence of the spatial code written by the UNIT-DNA method depends on the order of the nucleotides provided and the spatial activation of the cleavage reagent by light. The total number of spatial codes that can be written by UNIT-DNA depends on the number of universal bases in the 5' overhang that allow the incorporation of nucleotides (e.g., 10 universal bases would allow for ~1 million codes (4 10 ). The spatial resolution of the coding principle depends on the resolution of the light used for illumination (~300 nm for UVB) and on the local reaction kinetics of the released cleavage reagents, and therefore can easily reach cellular (~10 μm) or subcellular (~1 μm) resolution levels.
[0024] After coding is complete, whole cells (or nuclei and organelles) may be isolated from the tissue sample and subjected to single-cell sequencing. In principle, the method has no limit to the number of cells examined simultaneously. The number of cells examined simultaneously and individually depends on the number of universal bases in the 5' overhangs that provide a unique spatial barcode. The only practical limitation is ultimately the capacity and throughput of the sequencer.
[0025] An embodiment of UNIT-DNA spatial barcoding The embodiments of the method of the present invention for spatial barcoding are summarized in Figure 5 and are referred to as embodiments A to H. The core functional elements of the UNIT-DNA method are maintained, as visualized by the dotted boxes. Additional functionality is introduced by 3'- and 5'-end modifications, and further nucleic acid manipulation workflows are combined with the core UNIT-DNA composition for spatial barcoding.
[0026] The embodiments are described in more detail below.
[0027] Figure 6 shows an example of embodiment H (Figure 5) of UNIT-DNA used as part of the template-switched oligonucleotide process within a single-cell sequencing workflow (Picelli et al., Nat Methods 2013 Nov;10(ll):1096-8. doi:10.1038 / nmeth.2639. Epub 2013 Sep 22). After generation of spatial DNA barcodes by UNIT-DNA, the resulting nucleic acids can be further analyzed by sequencing using unique molecular identifiers (UMIs) for error correction.
[0028] Note that the TSO shown in Figure 6 does not include a cell identifier. The UNIT-DNA composition provides a spatial barcode that functions as a cell identifier when the resolution of the structured illumination is selected to match the cellular resolution level.
[0029] Figure 7 is a diagram updating the coding and decoding workflow for the use of UNIT-DNA derivatives in embodiment H. After coding, a sequencing library is prepared, and the spatial barcode and target nucleic acid are sequenced. Since the spatial barcode is physically bound to the target nucleic acid, the spatial information of the target sequence is provided by in vitro sequencing, providing a relationship between the result and the original sample source.
[0030] The UNIT-DNA composition H for spatial barcoding of target nucleic acids can also be used in the padlock workflow leading to circularized ssDNA (see FIG. 8) or in the target DNA workflow (see FIG. 9). After coding, the resulting nucleic acid is sequenced to determine the spatial barcode and the bound target nucleic acid.
[0031] Depending on the molecular workflow, different UNIT-DNA embodiments may be used to combine spatial coding with sequencing and decoding workflows. The UNIT-DNA embodiment shown in Figure 5 may also be used for multimodal targeted RNA and DNA workflows or solid support workflows (not shown).
[0032] The UNIT-DNA process may be carried out in a cyclic process triggered by structured illumination of the tissue sample by treating the tissue sample with another spatially structured light pattern in each cycle. "Structured illumination" and "spatially structured light pattern" refer to illuminating only a portion or selected area of the sample.
[0033] Figure 4 shows the structured illumination of three cells with light (indicated by a lightning symbol), leading to the release of the cleavage reagent TCEP (tris(2-carboxyethyl)phosphine) in the illuminated cells. The Cy5-TECP complex is used as a substrate for the light-induced cleavage reagent release.
[0034] 1 further shows how a tissue section 002 (optionally stained) is obtained from a tissue donor 001 and subjected to imaging 100, allowing segmentation or clustering, i.e., selection of parts of the sample to be further investigated by the method of the present invention. Such segmentation / clustering / selection allows the computation of masks for structured illumination and / or spatially structured light patterns.
[0035] Further downstream in the method of the invention, the information obtained about the structured illumination and / or the spatially structured pattern of light is utilized during light processing for UNIT-DNA code generation (102), which results in the encoded UNIT-DNA that is encapsulated together with the cellular mRNA and single cell indexing reagent (202) for subsequent sequencing (104). With the help of structured illumination, only selected regions / cells of the sample 106 with spatial barcodes are spatially decoded by next generation sequencing (104) and sequence analysis (106).
[0036] Sequencing One step in the method of the present invention aims to determine the sequence of the nucleotides coded on the UNIT-DNA probe that are read by sequencing (104). One method for sequencing is sequencing by synthesis (SBS). To increase the read signal, amplification of the UNIT-DNA probe sequence can be performed. One method of clonal amplification is rolling circle amplification (RCA) of the coded UNIT-DNA probe, which is performed before starting the sequencing process of Lorony.
[0037] In a variant of the present invention, the sequence of the UNIT-DNA code may be read separately from the sequence of the target gene. This can be achieved by splitting the sequencing procedure into two runs with two different sequencing primers.
[0038] EMBODIMENTS OF THE PRESENT DISCLOSURE Eight embodiments (A-H) of the UNIT-DNA method are shown in Figure 5. The core elements of the UNIT-DNA composition are maintained for all embodiments (indicated by the dotted box). Additional elements added to the 5' and 3' ends of the UNIT-DNA are:
[0039] (A) A second 5' overhang with a blocked 3' end within the overhang. In embodiments A and B, the first strand further comprises a blocking group at its 3' end.
[0040] (B) The 5' end of the duplex is joined to the 3' end of the duplex.
[0041] (C) The 5' end of the 5' universal base overhang is extended with a natural base. In embodiment C, the first strand overhang comprises a first oligonucleotide at its 5' end.
[0042] (D) The 5' end of a 5' universal base overhang extended with natural bases is attached to the 3' end of the duplex. In embodiment D, a first oligonucleotide is linked to the 3' end of the first strand directly or via an oligonucleotide bridge, thereby forming a circle. The oligonucleotide bridge may have a length of 5 to 100 nucleotides.
[0043] (E) The 5' end of the 5' universal base overhang extended by a natural base forms a double strand of natural bases. In embodiment E, a first oligonucleotide is hybridized with a corresponding nucleotide, thereby obtaining a polynucleotide having a blunt end and a gap at at least one universal base position.
[0044] (F) The 5' end of the 5' universal base overhang extended with natural bases forming a natural base duplex is joined to the 3' end of the adjacent duplex.
[0045] In embodiment F, a first oligonucleotide is hybridized with a corresponding nucleotide, thereby obtaining a polynucleotide having blunt ends and a gap at at least one universal base position, where the blunt ends are linked to each other directly or via an oligonucleotide bridge, which may have a length of 5 to 100 nucleotides.
[0046] (G) The 5' end of a 5' universal base overhang extended by a natural base forming a duplex of natural bases with a blocked 3' end, while the 5' end is bound to the 3' end of the opposite strand of the duplex to form a circle. In embodiment G, a first oligonucleotide is hybridized with a corresponding nucleotide, thereby obtaining a polynucleotide with a blunt end and a gap at the position of at least one universal base, where the first oligonucleotide is linked to the 3' end of the first strand directly or via an oligonucleotide bridge, thereby forming a circle, and the 3' end of the hybridized corresponding nucleotide contains a non-cleavable blocking group. The oligonucleotide bridge may have a length of 5 to 100 nucleotides.
[0047] (H) The 5' end of the duplex is linked to the 3' end of the opposite duplex to form a padlock-like structure with the 3' end of the duplex blocked. In embodiment G, a first oligonucleotide hybridizes with a corresponding nucleotide, thereby obtaining a polynucleotide with blunt ends and a gap at at least one universal base, where the 3' end of the hybridized corresponding nucleotide is linked to the 5' end of a second strand, either directly or via an oligonucleotide bridge, thereby forming a circle, and the 3' end of the first strand contains a non-cleavable blocking group. The oligonucleotide bridge may have a length of 5 to 100 nucleotides.
[0048] Embodiment H may be implemented in the first variant shown in Figure 6, where the UNIT-DNA method of the invention is combined with the template switch oligonucleotide (TSO) process.
[0049] The variants shown in Figure 6 include (A) reverse transcription (in situ) of mRNA with an oligo-dT primer to generate a cDNA with a triple C at the 3' end. (B) A TSO with a 5' PCR handle (shown as N), a unique molecular identifier (UMI) and a triple G at the 3' end. (C) In situ template switching of the cDNA (from A) with the TSO (from B) results in a captured cDNA with PCR handles at both ends (optional in situ PCR amplification to increase sensitivity is not shown). (D) Hybridization of an oligonucleotide containing a universal base with the PCR handle results in the UNIT-DNA derivative H (shown as a dotted box).
[0050] A further variation of embodiment H is shown in Figure 8, where the UNIT-DNA method of the invention is combined with circular ssDNA.
[0051] The variants shown in FIG. 8 include (A) circular ssDNA (containing the captured target sequence, a unique molecular identifier (UMI) and a PCR handle (depicted as N)). The UMI is added before amplification and is often used within NGS workflows to reduce errors and quantitative bias caused by amplification. The circular ssDNA may be generated in situ by the Padlock method (ungapped), the Padlock method (gapped) or the Direct method (FISSEQ) as described by Chen et al. (Nucleic Acids Research, 2018, Vol. 46, No. 4e22). (B) Oligohybridization generates double strands, allowing restriction endonuclease digestion as shown in (B), generating linear ssDNA (C) (optional in situ PCR amplification to increase sensitivity is not shown). (D) Hybridization of an oligonucleotide containing a universal base with the PCR handle results in UNIT-DNA derivative H (depicted by a dotted box).
[0052] A further variant of embodiment H is shown in Figure 9. Here, the UNIT-DNA method of the present invention is combined with target DNA amplification, for example by the target DNA workflow by QIAGEN disclosed in the QIAseq Targeted DNA Panel Handbook 03 / 2021. Figure 2 has been modified to combine with UNIT-DNA derivative H for space (B). The adapter contains a unique molecular identifier (UMI) and a PCR handle (N). (C) For enrichment, the ligated target sequence undergoes several cycles of in situ targeted PCR with a gene-specific primer (GSP with PCR handle) and a universal primer (shown as N). (D) Hybridization of an oligonucleotide containing a universal base with the PCR handle results in UNIT-DNA derivative H (shown in a dotted box).
[0053] Explanation of terms in Figures 1 and 7 001 Tissue Donor 002 Stained tissue section 003 Cell 004 Cell nucleus 005 Cytoplasmic mRNA 006 mRNA (bound to UNIT-DNA composition) 100 Imaging 101 Segmentation or cluster analysis, calculation of masks for structured illumination 102 UNIT-Optical processing for DNA code generation 103 Single Cell Encapsulation 104 Sequencing 105 Circular Barcoding 106 Sequence Analysis 200 UNIT-DNA composition before coding 201 UNIT-DNA composition after coding 202 Single Cell Indexing Reagent 203 UNIT-DNA Composition H 204 Linearized template-switched cDNA with spatial barcodes 205 cDNA-derived sequencing library
Claims
**Claim 1** A method for providing a polynucleotide comprising a first strand and a second strand having a barcode nucleotide sequence, wherein the first strand has an overhang of at least one universal base at its 5'-end, and the corresponding recessed 3'-end of the second strand of the polynucleotide having at least one nucleotide is provided with a blocking group, and the blocking group is removed from the incorporated nucleotide by light irradiation, characterized in that the method. **Claim 2** The method according to claim 1, wherein the blocking group is removed from the incorporated nucleotide by light irradiation by providing a cleavage reagent, and the cleavage reagent is provided by irradiating light on a precursor of the cleavage reagent. **Claim 3** The method according to claim 1, wherein the nucleotide is provided with a photocleavable blocking group that is removed from the incorporated nucleotide by light irradiation. **Claim 4** The method according to claim 1, wherein the first strand further comprises a blocking group at its 3'-end. **Claim 5** The method according to claim 1, wherein the overhang of the first strand comprises a first oligonucleotide at its 5'-end. **Claim 6** The method according to claim 5, wherein the first oligonucleotide is linked directly or via an oligonucleotide bridge to the 3'-end of the first strand, thereby forming a ring. **Claim 7** The method according to claim 5, wherein the first oligonucleotide is hybridized with the corresponding nucleotide, thereby obtaining a polynucleotide strand having a blunt end and a gap at the position of at least one universal base. **Claim 8** The method according to claim 5, wherein the first oligonucleotide is hybridized with the corresponding nucleotide, thereby obtaining a polynucleotide having a blunt end and a gap at the position of at least one universal base, and the blunt ends are linked to each other directly or via an oligonucleotide bridge. **Claim 9** The method according to claim 5, wherein the first oligonucleotide is hybridized with the corresponding nucleotide to obtain a polynucleotide having blunt ends and a gap at the position of at least one universal base, and the first oligonucleotide is linked directly or via an oligonucleotide bridge to the 3'-end of the first strand, thereby forming a loop, and the 3'-end of the hybridized corresponding nucleotide contains a blocking group that cannot be cleaved.
10. The method according to claim 5, wherein the first oligonucleotide hybridizes with the corresponding nucleotide to obtain a polynucleotide chain having blunt ends and a gap at the position of at least one universal base, and the 3'-end of the hybridized corresponding nucleotide is linked directly or via an oligonucleotide bridge to the 5'-end of the second strand, thereby forming a loop, and the 3'-end of the first strand contains a blocking group that cannot be cleaved.
11. The method according to any one of claims 1 to 10, wherein the polynucleotide chain is provided by template switching of the m-RNA strand by the following steps: a) synthesizing a first strand by reverse transcription of mRNA by oligo-dT priming, resulting in a 3'-terminal C nucleotide added to the captured target sequence, and subsequently b) template-switching the cDNA by hybridization of a template-switching oligo with the corresponding 3'-terminal G nucleotide, and c) hybridizing with the corresponding nucleotide of the first strand to generate a free 3' OH and a gap at the position of at least one universal base.
12. The following steps: a) capturing circular ssDNA having a target sequence by padlock probe hybridization (including a gap filling reaction for a gap filling padlock probe) and ligation; b) enabling dsDNA restriction by oligonucleotide hybridization to the circular ssDNA to obtain linear ssDNA, which is c) hybridized to the nucleotides of the corresponding first strand to generate a free 3′ OH and a gap at the position of at least one universal base, the method according to any one of claims 1 to 10, characterized in that the DNA strand is provided by a padlock workflow resulting in a circular ssDNA template. **Claim 13** The following steps: a) fragmenting DNA and ligating an adapter; b) enriching a ligated target sequence by PCR with gene-specific primers (with PCR handles) and universal primers to obtain linear ssDNA, which is c) hybridized to the corresponding first strand nucleotides to generate a free 3′ OH and a gap at the position of at least one universal base, the method according to any one of claims 1 to 10, characterized in that the DNA strand is provided by target DNA amplification.