Methods, uses and kits for the construction of spatial networks
By randomly immobilizing exogenous nucleic acids with unique identifiers in a sample, the method constructs a spatial network that addresses the limitations of current spatial transcriptomics techniques, offering high-resolution spatial and transcriptomic mapping without optical imaging or beacon reliance.
Patent Information
- Application Number
- PCT/EP2025/061354
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-04-25
- Publication Date
- 2025-10-30
Smart Images

Figure IMGF000068_0001 
Figure IMGF000072_0001 
Figure IMGF000076_0001
Abstract
Description
[0001] METHODS, USES AND KITS FOR THE CONSTRUCTION OF SPATIAL NETWORKS
[0002] FIELD OF THE INVENTION
[0003] The present application relates to methods, systems and kits for the construction of a spatial network in a sample containing cells using nucleic acids and for determining the spatial arrangement of targetable molecules in a sample. The methods generally involve randomly immobilizing exogenous nucleic acids in a sample comprising cells, wherein each nucleic acid comprises a unique sequence identifier and identifies a node in the spatial network. More particularly, the methods of the invention may generally comprise the steps of providing a sample comprising targetable molecules; labelling each targetable molecule with a unique sequence identifier to form a seed; forming a plurality of seed-specific amplicons, which activate at least one type of sub-seeds to form at least one type of concatemer; and determining the spatial arrangements of the targetable molecules based on the at least one types of concatemer.
[0004] BACKGROUND TO THE INVENTION
[0005] The spatial organisation of cells within tissues plays an important role in the function of tissues in both health and disease [1-3]. However, while spatial architecture is an important factor, the gene expression and identity of these cells are also important in their functionality, as seen in the tissue functionality of tumour heterogeneity [4-6]. Over the years, techniques have been designed to simultaneously investigate the transcriptomic profile and location of cells. This led to the creation of the field of spatial transcriptomics. Techniques such as SABER-FISH [7], osmFISH [8], MERFISH [9], In Situ Sequencing
[0010] , and seqFISH
[0011] aim to add transcriptomic information to microscopy data to obtain a dual readout of gene expression and cell location. These techniques have good spatial resolution for samples but have limitations based on the density of molecules and optics [12, 13]. Additionally, one of the main challenges of these methods is that they rely on targeted approaches.
[0006] Alternative means to assess the spatial organisation of cells within tissues include microdissection methods. Techniques using this strategy try to separate the tissue in discrete chunks into many tubes, while keeping track of their original location and sequence them separately. The methods can be considered a brute force approach to solving the problem, i.e. to know the transcriptome of a particular area, just physically isolate that area. Methods in this group have been refined over the years from early laser capture microdissection to Tomo-Seq (where several identical samples are sliced in different directions) and pseudo-spatial partial dissociation techniques. Some newer techniques involve in situ tagging and selective capture of transcripts using photoactivation of specific tissue locations, like TIVA or Niche-Seq. However, while these techniques typically allow for deep sequencing, they are often labour intensive and hard to scale. Several of the techniques are also hard to generalise from model systems to patient samples, largely due to the requirements for things like multiple identical samples (Tomo-Seq), genetic tagging (Niche-Seq), or the use of live tissue (TIVA).
[0007] Further alternatives include certain forms of sequencing in situ, which utilise padlock probes that bind to cDNA and are circularised. These can then produce rolling circle products (RCPs) that amplify the number of sequences locally, which can then be visualised in situ using sequencing-by-ligation or sequencing-by-synthesis. However, these methods are targeted (barcoded padlock probes are directed to specific preselected RNAs). Additionally, a common problem for this type of method is the iterative sensitive imaging protocols during sequencing and the fact that the number of reads per cell / per area is limited by optical crowding and spectral overlap.
[0008] On the other hand, techniques such as Spatial Transcriptomics
[0014] , DBiT-seq
[0015] , Slide-seq
[0016] , and SeqScope
[0017] rely on sequencing and aim to add spatial information to the transcriptomic datasets. In these techniques, this is accomplished by creating a pattern of known DNA barcodes corresponding to coordinates on the pattern, adding these barcodes to the cDNA, and sequencing them. Because the location of the barcode is known beforehand, the capture location of the cDNA molecule is also known. Although these techniques have good transcriptomic resolution, they require patterning and characterisation of the surface beforehand.
[0009] A relatively newer sequencing-based approach, called DNA microscopy, has since been developed [18, 19] (e.g. Figure 17). DNA microscopy is a PCR-based method that allows for the reconstruction of a sample layout together with transcriptomic information without the need for microscopic analysis [18, 19]. The basis of this technique revolves around the creation of polymerase colonies (polonies)
[0020] of barcoded DNA strands from target cDNA molecules in a hydrogel. Polonies grow during PCR and can exchange barcode information with neighbouring polonies, resulting in a map of neighbouring polonies without requiring a priori characterisation beforehand. However, this work relies on the presence of RNA and requires the careful design of each target to be incorporated into the reaction, making it difficult to map space with little to no RNA sequences, such as between cells. Thus, current techniques in the art have several drawbacks. For example, the methods are targeted, due to relying on diffusion between strands of defined length, and all involved sequences must be generated from primers that produce known, same-size amplicons. The methods also rely on highly expressed 'beacons', by using specific sets of genes (such as GAPDH and ACTB) to create beacon locations that should provide a uniform network, but this creates problems both in generating a large number of wasted reads during sequencing, while also leading to collapse if the density of transcripts of these genes drops. Additionally, empty space cannot be reconstructed, given the reliance on endogenous beacons and a constant diffusion speed.
[0010] SUMMARY OF THE INVENTION
[0011] Against this background, the present inventors have developed an advantageous approach that overcomes many of the problems and difficulties associated with prior art methods. As discussed in detail herein, the inventors' methods remove the need for optical imaging, and instead offload the spatial workload to a computer, which drastically reduces the experimental protocol. Additionally, the inventors' methods do not require beacons as attachment points in samples, and instead allow for the targeting of polynucleotide sequences irrespective of their sequences. Including a 3' block on sub-seeds also advances on known methods by reducing smearing and the amount of side products that may form; while methods incorporating a two-layer approach provide the ability to tune the technique for higher spatial resolution versus higher read depth.
[0012] More generally, the inventors have determined that populating a sample with exogenous nucleic acids (e.g. oligonucleotides) that are randomly immobilized in the sample enables the generation of spatially-restricted or delineated synthetic nodes in the sample that can be used to construct a spatial network using the relative positions of the nodes. Advantageously, as the spatial network is not reliant on endogenous nucleic acids, the network can map space with little to no endogenous nucleic acids, such as between cells. The nodes can be used to generate a spatial network that can provide both geometric (i.e. spatial) information and transcriptomic information.
[0013] Thus, one aspect provided herein is a method for preparing a sample containing one or more cells for the construction of a spatial network using nucleic acids, the method comprising:
[0014] (A) providing a sample containing one or more cells; and (B) randomly immobilizing (fixing) exogenous nucleic acids (seeds as defined herein) in the sample, wherein each nucleic acid comprises a unique sequence identifier, wherein each nucleic acid identifies a node in the spatial network.
[0015] A further aspect provided herein is a sample containing one or more cells for the construction of a spatial network using nucleic acids, the sample comprising randomly immobilized (fixed) exogenous nucleic acids (seeds as defined herein) in the sample, wherein each nucleic acid comprises a unique sequence identifier and identifies a node in the spatial network.
[0016] Another aspect provided herein is a system for the construction of a spatial network in a sample comprising cells, the system comprising:
[0017] (A) a plurality nucleic acids (seeds and / or sub-seeds as defined herein), wherein each nucleic acid comprises a unique sequence identifier; and
[0018] (B) means for randomly immobilizing (fixing) the plurality of nucleic acids (seeds as defined herein) in the sample, such each nucleic acid identifies a node in the spatial network.
[0019] Yet another aspect provided herein is the use of the system described above for the construction of a spatial network in a sample comprising cells.
[0020] A further aspect provided herein is method for constructing of a spatial network in a sample containing one or more cells using nucleic acids, the method comprising:
[0021] (A) preparing a sample containing cells as described herein or providing a sample as described herein;
[0022] (B) forming nucleic acid nodes (e.g. polonies) from the nucleic acids immobilized in the sample; and
[0023] (C) deducing the location of the nucleic acid nodes based on interactions between nucleic acids in adjacent nodes thereby constructing a spatial network.
[0024] Still another aspect provided herein is a method for mapping transcriptomic and / or proteomic information in a sample containing one or more cells, comprising:
[0025] (A) constructing a spatial network in the sample using nucleic acids as described herein;
[0026] (B) forming nucleic acids (e.g. concatemers) from interactions between nucleic acid nodes and target nucleic acids in the sample that contain transcriptomic or proteomic information; and (C) using the nucleic acids (e.g. concatemers) formed in (b) to generate a map of transcriptomic and / or proteomic information in the sample.
[0027] It will be understood that the methods, sample and system above also may comprise immobilizing (fixing) other nucleic acids in the sample, e.g. sub-seeds, target nucleic acids. The other nucleic acids may contain unique sequence identifiers.
[0028] In another aspect, the invention provides a method for determining the spatial arrangement of two or more targetable molecules comprising: a. providing a sample comprising at least two targetable molecules, wherein each targetable molecule is at a distinct location in the sample; b. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule; c. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein:
[0029] (i) at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and / or
[0030] (ii) at least one of the plurality of seed-specific amplicons hybridises with and extends on a measurement sub-seed to form a measurement monomer, and wherein each measurement monomer is complementary to a marker monomer, and the monomers hybridise and extend on each other to form a measurement concatemer; d. determining the spatial arrangement of the two or more targetable molecules on the basis of the geometry concatemers and / or measurement concatemers formed in step (c).
[0031] In a further aspect, the invention provides a method for recording or determining cellular co-localisation and / or spatial distributions of polynucleotide molecules, comprising: a. providing a sample comprising a cell or population of cells, wherein the cell or population of cells comprises or comprise at least two targetable molecules, and wherein each targetable molecule is at a distinct location in the sample; b. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule; c. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein:
[0032] (i) at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and / or
[0033] (ii) at least one of the plurality of seed-specific amplicons hybridises with and extends on a measurement sub-seed to form a measurement monomer, and wherein each measurement monomer is complementary to a marker monomer, and the monomers hybridise and extend on each other to form a measurement concatemer; d. sequencing the geometry concatemers and / or measurement concatemers formed in step (c), thereby recording or determining cellular co-localisation of the polynucleotides molecules in the sample.
[0034] In a still further aspect, the invention provides a method for single cell mapping comprising: a. providing a sample of cells comprising at least two targetable molecules, wherein each targetable molecule is at a distinct location in the sample; b. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule; c. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein:
[0035] (i) at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and / or
[0036] (ii) at least one of the plurality of seed-specific amplicons hybridises with and extends on a measurement sub-seed to form a measurement monomer, and wherein each measurement monomer is complementary to a marker monomer, and the monomers hybridise and extend on each other to form a measurement concatemer; d. generating a spatially resolved physical map of each cell, optionally by isolation and sequencing, on the basis of the geometry concatemers and / or measurement concatemers formed in step (c).
[0037] In yet another aspect, the invention provides a method of identifying a disease, disorder, or condition in a patient comprising: a. providing a sample from the patient, wherein the sample comprises at least two targetable molecules, and wherein each targetable molecule is at a distinct location in the sample; b. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule; c. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein:
[0038] (i) at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and / or
[0039] (ii) at least one of the plurality of seed-specific amplicons hybridises with and extends on a measurement sub-seed to form a measurement monomer, and wherein each measurement monomer is complementary to a marker monomer, and the monomers hybridise and extend on each other to form a measurement concatemer; d. determining the spatial arrangement of the two or more targetable molecules on the basis of the geometry concatemers and / or measurement concatemers formed in step (c), e. and identifying the disease, disorder, or condition on the basis of the spatial arrangement.
[0040] In another aspect, the invention provides a method for determining the spatial arrangement of two or more targetable molecules comprising: i. providing a sample comprising at least two targetable molecules, wherein each targetable molecule is at a distinct location in the sample; ii. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule; iii. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, wherein the geometry sub-seed comprises a 3' end that is blocked from extension, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and iv. determining the spatial arrangement of the two or more targetable molecules on the basis of the geometry concatemers formed in step (c).
[0041] DETAILED DESCRIPTION
[0042] By "targetable molecule" we include the meaning of a molecule that is at a location or point of origin for a polony to form. The two or more targetable molecules may be a cellular, sub-cellular, and / or extracellular target in a sample (for example, to allow for a cellular map to be determined). The two or more targetable molecules may comprise polynucleotides and / or biomolecules (e.g. polypeptides and / or proteins), regardless of the organism of origin, and can be specifically targeted by another molecule. For example, polynucleotides may be targeted by using a complementary polynucleotide. Alternatively, a biomolecule (e.g. polypeptide or protein) may be targeted with an antibody or antigen-binding fragment thereof with binding specificity for an epitope within the biomolecule (e.g. polypeptide or protein). A targetable molecule may be added to the sample and / or may already be present within the sample. For example, the two or more targetable molecules may comprise polynucleotides and / or biomolecules (e.g. polypeptides and / or proteins) that are endogenous to a cell within a sample. Alternatively, or additionally, the two or more targetable molecules may comprise polynucleotides and / or biomolecules (e.g. polypeptides and / or proteins) residing within a gel that is applied to the sample. Thus, since the molecule can be targeted and act as a point of origin for a polony, it is a targetable molecule. In some aspects, it is preferred that the targetable molecule is an exogenous nucleic acid (i.e. polynucleotide, oligonucleotide etc.) that contains unique sequence identifier, i.e. a nucleic acid that is added to the sample.
[0043] By "cellular" we include the meaning that the molecule (e.g. a targetable molecule, seed or sub-seed) is intracellular and / or on a cell surface (e.g. facing intracellularly). By "sub-cellular" we include the meaning that the molecule (e.g. a targetable molecule, seed or sub-seed) is within an intracellular compartment, for example, a nucleus, a nuclear membrane, a nucleoplasm, a cytosol, a cytoplasm, a mitochondrion, an endoplasmic reticulum, a Golgi apparatus, and a vesicle (including endosomes, lipid droplets, lysosomes, and peroxisomes). By "extracellular" we include the meaning that the molecule (e.g. a targetable molecule, seed or sub-seed) is on a cell surface (e.g. facing extracellularly), a vesicle (including endosomes, lipid droplets, lysosomes, peroxisomes, and apoptotic bodies), or any space in the sample that is not surrounded by a cellular membrane. The term "intracellular" encompasses the term "sub-cellular".
[0044] In some embodiments, the sample comprises at least 3 targetable molecules, for example at least 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 2,000,000, 3,000,000 or more targetable molecules. In some embodiments, the sample comprises from 2 to 3,000,000 targetable molecules, for example, from 10 to 3,000,000, from 100 to 3,000,000, from 1,000 to 3,000,000, from 10,000 to 3,000,000, from 100,000 to 3,000,000, from 1,000,000 to 3,000,000 targetable molecules. Preferably, the number of targetable molecules is from about 30,000 to about 3,000,000.
[0045] By "seed" we include the meaning of a polynucleotide molecule comprising a unique sequence identifier, upon which elongation may occur to create one or more amplicon comprising the unique sequence identifier (i.e. a seed-specific amplicon). The polynucleotide molecule may be introduced to the sample (for example, as a component of a gel) or may already be present in the sample. By placing or having a polynucleotide comprising a unique sequence identifier in the sample, the seed can be traced to a particular location (optionally relative to a further seed or seeds) or point of origin. In some embodiments, one or more seed is fixed to a random location in the sample. For example, a gel comprising one or more seed can be applied to a sample (such as a fixed sample) that results in a spread of one or more of the seeds at multiple points in the sample in a non-targeted approach. The location may be cellular, sub- cellular and / or extracellular to cells within a sample. In some embodiments, the seed further comprises a domain or portion that can hybridise with a sub-seed.
[0046] By "sub-seed" we include the meaning of a polynucleotide molecule with a domain that hybridises with one or more seed-specific amplicon. Following hybridisation, elongation may continue for the seed-specific amplicon, which may append additional nucleotides (for example, domains or portions as described herein) to form a monomer. In some embodiments, the additional domain introduced is an activation domain. In some embodiments, the additional domain allows the monomer to hybridise to other monomers, as described herein. By "labelling" we include the meaning that the unique sequence identifier is introduced to the two or more targetable molecules, thereby acting as a label or recognition point for the two or more targetable molecules. Methods of labelling polynucleotides and / or biomolecules (e.g. polypeptides and / or proteins) are known to the skilled person. In some embodiments, the biomolecules (e.g. polypeptides and / or proteins) are labelled with a polynucleotide tag, optionally comprising a unique sequence identifier. Since the two or more targetable molecules may be polynucleotides, an active labelling step may not be necessary if the two or more targetable molecules already comprise a unique sequence identifier and / or are unique sequence identifiers. For example, at least one seed, which comprises (or is) a unique sequence identifier, may be added to a sample. In this case, each seed is already a targetable molecule. Alternatively viewed, a targetable molecule may be a seed as defined herein. The labelling in step (b) may occur prior to step (a) of the method. For example, a targetable molecule may be labelled with a unique sequence identifier, or with a polynucleotide comprising a unique sequence identifier, and subsequently introduced to a sample. Thus, a seed may be formed (as per step (b)), and then the seed applied to the sample. In some embodiments, labelling occurs prior to step (a). In some embodiments, labelling occurs after step (a). In some embodiments, labelling occurs prior to and after step (a), such that at least one targetable molecule is labelled prior to step (a) and at least one targetable molecule is labelled after step (a).
[0047] Accordingly, it will be understood that in some aspects, steps (a) and (b) of the methods provided herein may be performed as sample preparation steps such that the method involves: providing a sample comprising at least two targetable molecules (e.g. polynucleotides, such as exogenous nucleic acids) comprising a unique sequence identifier, wherein each targetable molecule is at a distinct location in the sample thereby forming a seed. As noted above, the targetable molecules (e.g. polynucleotides, such as exogenous nucleic acids) comprising a unique sequence identifier, may be applied to the sample via a gel and randomly fixed (immobilized) in the sample. Thus, in a representative aspect, the method for determining the spatial arrangement of two or more targetable molecules may comprise: a. providing a sample comprising at least two targetable molecules each comprising a polynucleotide comprising a unique sequence identifier, wherein each targetable molecule is at a distinct location in the sample thereby forming a seed; b. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein: (i) at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and / or
[0048] (ii) at least one of the plurality of seed-specific amplicons hybridises with and extends on a measurement sub-seed to form a measurement monomer, and wherein each measurement monomer is complementary to a marker monomer, and the monomers hybridise and extend on each other to form a measurement concatemer; c. determining the spatial arrangement of the two or more targetable molecules on the basis of the geometry concatemers and / or measurement concatemers formed in step (b).
[0049] It will be evident that the combination of steps (a) and (b) to provide step (a) as defined above may be applied to the other methods provided herein, i.e. the method for recording or determining cellular co-localisation and / or spatial distributions of polynucleotide molecules, the method for single cell mapping and the method of identifying a disease, disorder, or condition in a patient.
[0050] In some embodiments, one or more of the targetable molecules comprises a messenger RIMA (mRNA); and / or comprises a molecule labelled with a polynucleotide tag. In some embodiments, the mRNA is reverse transcribed to cDNA, optionally prior to the labelling in step (b). In some embodiments, one or more of the targetable molecules comprises a biomolecule (e.g. a polypeptide or protein) target, optionally wherein the biomolecule target is targeted with an antibody or antigen-binding fragment thereof with binding specificity for the biomolecule target. Preferably, the antibody or antigen-binding fragment thereof comprises a polynucleotide tag, or is labelled with a polynucleotide tag.
[0051] In some embodiments, the two or more targetable molecules are positioned and / or immobilised in the sample, optionally wherein the two or more targetable molecules are positioned and / or immobilised in a gel. In some embodiments, PEG-based gelling components are used to form a hydrogel, for example by using acrylate-PEG and / or thio-PEG gelling components (optionally as part of a buffer). Preferably, gelation occurs after fixation and / or permeabilisation of the sample (or the cells in the sample). Preferably, the gel is dissolvable, cleavable, thermostable, and / or diffusion restrictive. By "dissolvable" we include the meaning that the gel can be removed by converting it to a solution state, while maintaining the integrity of any polynucleotide products (preferably at least the concatemers). By "cleavable" we include the meaning that components within the gel can be split apart, thereby liberating any polynucleotide products (preferably at least the concatemers). By "thermostable" we include the meaning that the gel structure is maintained during thermal cycling (e.g. thermal cycling for PCR, as described herein). By "diffusion restrictive" we include the meaning that the movement (i.e. diffusion) of amplicons within a polony is impeded or restricted.
[0052] It will be appreciated that preferences and options for the targetable molecules may be combined. For example, one or more of the targetable molecules may comprise a mRNA and / or a molecule labelled with a polynucleotide tag, with the two or more targetable molecules being positioned and / or immobilised in the sample. I.e. in some embodiments, one or more of the targetable molecules may be an mRNA that is positioned and / or immobilised in the sample.
[0053] In some embodiments, the sample comprises one or more cell. For example, the sample may be a population of cells, a tissue, or an organ. It is preferred that the cell, or one or more cell in the population of different cells, is selected from the following: a eukaryotic cell (for example, from an animal, a plant, or a fungus), a bacterial cell (for example, from Eubacteria), or an archaeal cell (for example, from Archaebacteria). It will be appreciated that the methods of the invention can be adapted to a cell, tissue, or organ of any genomic origin. Preferably, the cell, or one or more cell in the population of cells to be analysed, is provided in a biological sample, for example, a biopsy of a tissue, a whole tissue, blood, serum, urine, or saliva. Methods for obtaining such samples are well known in the art, as are methods for preparing cells in such samples for in situ amplification methods. A skilled person would therefore be able to obtain such samples and analyse them using the methods of the invention without difficulty. As mentioned above, the targetable molecules in the sample may be cellular, sub-cellular and / or extracellular targets.
[0054] In some embodiments, the sample is prepared as a tissue slice, i.e. as a section of tissue cut from larger pieces or extracted organs. Techniques for preparing tissue slices are known to the skilled person. The thickness of the tissue slices may vary. The methods of the invention are conveniently not limited by the thickness of the tissue slices. The targetable molecules in a tissue slice may be cellular, sub-cellular and / or extracellular targets, as described above. Additionally, or alternatively, the targetable molecules in a tissue slice may be in an acellular space (i.e. a portion of the tissue lacking cells within the tissue slice).
[0055] In some embodiments, the sample is cleared, or the method includes a step of clearing the sample. By "cleared" or "clearing" we include the meaning that any opaque aspects (e.g. opaque tissue) has been treated to render it translucent for microscopy. In some embodiments the sample is fixed prior to clearing. For example, the sample may be fixed, and an acrylamide-based hydrogel may be cast inside the sample to preserve its shape. The sample may then be treated to remove the lipids (i.e. in a delipidation step), and then treated with antibodies. Alternatively, the sample may be dehydrated in methanol to fix and preserve its shape.
[0056] Methods for getting various component molecules deep into tissue (i.e. 'clearing methods'), such as by diffusion initially, are known to the skilled person. For example, CLARITY and 3DISCO / iDISCO+ are two established clearing methods that are compatible with imaging, such as 3D imaging (Azaripour et al., 2016, Progress in Histochemistry and Cytochemistry, 51:9-23). Alternatively, or additionally, the sample may be cleared by dehydration and / or delipidation (i.e. removal of lipid content in the sample). Additional details on suitable clearing techniques are included in, for example: (i) Chung et al., 2013, Nature, 497:332-337; (ii) Ertiirk et al., 2012, Nature Protocols, 7: 1983-1995; and (iii) Renier et a / ., 2014, Cell, 4:896-910.
[0057] In some embodiments, the invention permits analyses of multiple individual cells from a population of cells. In one embodiment, every cell in the population is analysed. A skilled person would appreciate that there is no theoretical limit of how many cells can be analysed with the method of invention. For example, the method may analyse at least a single cell, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 or more cells. Alternatively, the method may analyse from 1 to 1,000,000 cells, for example, from 10 to 1,000,000 cells, from 100 to 1,000,000 cells, from 1,000 to 1,000,000 cells, from 10,000 to 1,000,000 cells, from 100,000 to 1,000,000 cells, from 1 to 100,000 cells, from 1 to 10,000 cells, from 1 to 1,000 cells, from 1 to 100 cells, from 10 to 100,000 cells, from 100 to 100,000 cells, from 1,000 to 100,000 cells, from 10,000 to 100,000 cells, or any other range therebetween. In some embodiments, the population of cells comprises: cells derived from embryogenesis; cells derived from a particular organ (for example, whole blood, liver, kidney, heart, lung, bone, muscle, stomach, gallbladder, intestine, bladder, brain, pancreas, adrenal gland, thymus, parathyroid, thyroid, spinal cord, skin, bone marrow, spleen, thymus, tonsil, ovary, testis), or cells obtained from cultured cells.
[0058] In some embodiments, the population of cells comprise cells derived from a non- healthy or diseased tissue, such as from cancerous tissue. Cancers include: bladder cancer, breast cancer, colon and rectal cancer, endometrial cancer, kidney cancer, leukaemia, liver cancer, lung cancer, melanoma, non-Hodgkin lymphoma, pancreatic cancer, prostate cancer, and thyroid cancer.
[0059] The method described herein is conveniently not limited by cellular and tissue settings with empty spaces and fewer cells per volume, such as with germinal centres.
[0060] The terms "nucleotide sequence" or "nucleic acid" or "polynucleotide" or "oligonucleotide" refer to a heteropolymer of nucleotides or the sequence of these nucleotides. These phrases also refer to DNA or RIMA of genomic or synthetic origin which may be single-stranded or double-stranded and may represent the sense or the antisense strand, to peptide nucleic acid (PNA) or to any DNA-like or RNA-like material. In the sequences herein, A is adenine, C is cytosine, T is thymine, G is guanine and N is A, C, G or T (U). It is contemplated that where the polynucleotide is RNA, the T (thymine) in the sequences provided herein is substituted with U (uracil). It will be appreciated that the method of the present invention will perform well on any polynucleotide, regardless of the organism of origin. Similar polynucleotides of any origin can be applied to biomolecules from any other origin. It will be understood from the disclosures herein that the term "monomer" may be used interchangeably with the terms "nucleotide sequence" or "nucleic acid" or "polynucleotide" or "oligonucleotide" where the context admits, e.g. a marker monomer alternatively may be termed a marked polynucleotide, a geometry monomer alternatively may be termed a geometry polynucleotide etc.
[0061] In some embodiments of the methods disclosed herein, the size range of the concatemers is 50 to 1,000 base pairs, such as 100 to 1,000 base pairs, such as 200 to 1,000 base pairs, such as 300 to 1,000 base pairs, such as 400 to 1,000 base pairs, such as 500 to 1,000 base pairs, such as 600 to 1,000 base pairs, such as 700 to 1,000 base pairs, such as 800 to 1,000 base pairs, such as 900 to 1,000 base pairs. In some embodiments of the methods disclosed herein, the size range of the concatemers is 50 to 900 base pairs, such as 50 to 800 base pairs, such as 50 to 700 base pairs, such as 50 to 600 base pairs, such as 50 to 500 base pairs, such as 50 to 400 base pairs, such as 50 to 300 base pairs, such as 50 to 200 base pairs, such as 50 to 100 base pairs. In some embodiments, the concatemers are about 400 base pairs, for example 408 base pairs. In some embodiments, the concatemers are about 250 base pairs.
[0062] In some embodiments, the sample is fixed. Techniques for fixing samples (or cells within a sample) are known to the skilled person. Cells can be fixed with chemical and / or physical methods. Chemical methods of fixing cells include cross-linking agents, such as formaldehyde, glutaraldehyde and succinimide esters as well as solvents such as acetone and methanol, which precipitate proteins. Physical methods of fixing cells include freezing the sample (e.g. cells or tissues) and air drying. Fixed samples are preferable to reduce or prevent the movement of targets within a sample.
[0063] In some embodiments, the sample is permeabilised. Permeabilisation is advantageous when the marker monomer forms based on an intracellular target to allow access to the target. For example, the marker monomer may form from an intracellular polynucleotide (such as an mRNA, cDNA of the mRNA, or DNA (for example, genomic DNA), or from an intracellular biomolecule (e.g. polypeptide or protein). However, in cases where the marker monomer forms based on an extracellular and / or cell surface target, permeabilisation may not be necessary, as the target is already accessible. The target upon which marker monomers form may be a mix of intracellular and extracellular (e.g. cell surface) targets, in which case permeabilisation may facilitate targeting of the intracellular targets, while not impacting the extracellular and / or cell surface targets. In some embodiments, a detergent is used to permeabilise the cells, for example Triton X-100, Proteinase K, Tween-20; preferably Triton X-100.
[0064] In some embodiments, the step of forming seed-specific amplicons comprises extension of the polynucleotide molecule, e.g. using a polymerase enzyme, optionally wherein extension is performed using a DNA polymerase. Preferably, the DNA polymerase is a blunt end polymerase. In some embodiments, the DNA polymerase is a B-family DNA polymerase, such as KAPA HiFi polymerase. In some embodiments, the DNA polymerase has 5'->3' polymerase and 3'->5' exonuclease (proofreading) activity, but no 5'->3' exonuclease activity. The terms "extension" and "elongation" are used interchangeably herein. I.e. an "extension step" may also be considered an "elongation step". It will be appreciated that extension and elongation as used herein includes the use of the polynucleotide (e.g. of the targetable molecule) to template the extension or elongation of another polynucleotide, e.g. the seed functions to template the extension of primers to generate amplicons. It also will be appreciated the method is not limited to a particular type of polymerase, nor to blunt end polymerases. Alternative polymerases may include Taq polymerases that have been adapted with an A-overhanging portion (i.e. 3'-dA overhangs), such as Platinum Taq I. Additionally, it may not be necessary to use a polymerase with proofreading activity.
[0065] By "amplicon" we include the meaning of a polynucleotide (e.g. DNA molecule) that has been amplified, or is capable of being amplified, from a polynucleotide (e.g. DNA) template, for example, a PCR product. By "seed-specific" we include the meaning that the amplicon formed by elongation on a particular seed can be specifically traced to the location of that particular seed. By "seed-specific amplicon" we include the meaning that it is capable of being amplified into multiple copies. Thus, multiple seeds can be distinguished from each other based on its sequence (e.g. based on the unique sequence identifier) being attributed to a location in a sample. Additionally, seedspecific amplicons, following amplification, can form polonies. Seed-specific amplicons may also be referred to as seed-specific monomers.
[0066] Preferably, the step of forming seed-specific amplicons from each polynucleotide molecule occurs in situ, for example during an in situ thermal cycling protocol that enables reverse transcription and / or amplification cycles. Amplicons spool from the seeds and the cycle repeats, which forms a polony (also referred to herein as a 'cloud') of seed-specific amplicons. Hence, the formation of a plurality of seed-specific amplicons in step (c) may also be referred to as a polony formation stage. A polony may define a node of a spatial network as described herein.
[0067] Each seed comprises a unique sequence identifier. In some embodiments, the unique sequence identifier comprises a "barcode domain", which is a uniquely identifiable domain, or the reverse complement thereof. In some embodiments, each seed further comprises a primer domain (also referred to as a "primer site"), which is a domain to which a primer may hybridise to initiate elongation. The primer domain may be part of the unique sequence identifier, a separate domain added to the unique sequence identifier, or a domain that overlaps with the unique sequence identifier. In some embodiments, each seed further comprises an activation domain, which has homology for a site on a sub-seed (optionally, the activation domain has homology for a geometry sub-seed and / or a measurement sub-seed, i.e. in some embodiments, the homologous region of the activation domain may be identical or shared for a geometry sub-seed and measurement sub-seed). The activation domain may be comprised of two parts, a pre-activation portion ("Pre-act.") and an activator portion. In some embodiments, the pre-activation portion has homology to a further polynucleotide in the sample that contains the reverse complement of the activator portion, such that continued elongation of the seed-specific amplicon appends the activator portion to the amplicon, thereby forming the activator domain. The activator domain may be part of the unique sequence identifier, a separate domain added to the unique sequence identifier, or a domain that overlaps with the unique sequence identifier.
[0068] Thus, a seed-specific amplicon may comprise multiple domains, as follows (it will be appreciated that options ii and iii, below, may collectively be considered the unique sequence identifier): i. 5' - [unique sequence identifier] - 3'; ii. 5' - [primer domain] - [barcode domain] - [activator domain] - 3'; iii. 5' - [primer domain] - [barcode domain] - [pre-activation portion] - [activation portion] - 3'.
[0069] In the case of multiple seeds, it will be readily understood that the seeds may be attributed with a numbering system. For example, a first seed (Seed 1), a second seed (Seed 2), a third seed (Seed 3), and so on. The same applies for a first seed-specific amplicon (or first domains thereof, e.g. Barcode 1), a second seed-specific amplicon (or second domains thereof, e.g. Barcode 2), a third seed-specific amplicon (or third domains thereof, e.g. Barcode 3), and so on. Alternatively, a lettering system may be used (e.g. Primer Y, Primer R, etc) to denote distinct sequences for the domains. Preferably, the seed-specific amplicons do not comprise a homology domain for each other (i.e. preferably, the seed-specific amplicons do not hybridise to each other). The same may be applied to sub-seeds and targets / markers.
[0070] In step (b) of the methods described herein, at least two distinct seeds are formed, i.e. a first seed and a second seed each with a unique sequence identifier. By "distinct" we include the meaning that each seed can be attributed with a unique location (e.g. as based on a unique sequence). The term "distinct" may be used interchangeably herein with "different". I.e. a distinct location for one thing versus another may be described as being a different location.
[0071] In some embodiments, a first seed comprises a first pre-activation portion, a first barcode domain, and a first primer domain, while a second seed comprises a second pre-activation portion, a second barcode, and a second primer domain. The first seed may be distinguishable from the second seed based on any one or more of the aforementioned domains or portions. In such embodiments, in step (c), copies are made of each seed based on the respective primers. Optionally, copies of the seedspecific amplicons undergo exponential PCR, which may add a new sequence at the 3' end of the molecule (also referred to herein as an activation / activator sequence or activation / activator domain).
[0072] In some embodiments, a first seed comprises: a primer domain complementary to the primer 'PR_Yel_5N', Barcode 1, optionally Pre-act. Y, and Activator Y; and a second seed comprises: a primer domain complementary to PR_Red_5N, Barcode 2, optionally Pre-act. R, and Activator R. The sequences for these domains can be seen in Table 1. In some embodiments, the seeds comprise one or more of the sequences as shown in Table 1. For example, a first seed may be SD_N30R, and a second seed may be SD_N30Y. In some embodiments, the seeds comprise one or more of the sequences as shown in Table 3. For example, a first seed may be BCSD_R1, and a second seed may be BCSD_Y1.
[0073] In some embodiments, Barcode 1 and Barcode 2 comprise the following sequences, respectively:
[0074] Barcode 1 : NNNNNNNNNNGGTGGNNNNNNNNNNAATTGNNNNNNNNNN (SEQ ID NO: 1); and / or
[0075] Barcode 2: NNNNNNNNNNAGTGGNNNNNNNNNNCACATNNNNNNNNNN (SEQ ID NO: 2).
[0076] However, it will be appreciated that the preparation of barcodes (or UMIs) of random nucleotides is routine to the skilled person.
[0077] In some embodiments, primers are added to the sample (e.g. as part of a solution applied to the sample). The primers may comprise an activation domain and / or a domain that hybridises to seed-specific amplicons. For example, the activation domain may be selected from AAATCGTTTCATCGC (SEQ ID NO: 3) (as used in EX_Red) or ATGCTGAGTATTCCC (SEQ ID NO: 4) (as used in EX_Yel). Preferably, the same activation domain is also present in one or more sub-seed. Thus, in some embodiments, the sub-seeds comprise an activation domain that matches the activation domain of a primer used in the method. For example, the activation domain for a first primer, a first geometry sub-seed, and a first measurement sub-seed may be AAATCGTTTCATCGC (SEQ ID NO: 3), as used in EX_Red and OE_Geo_R_15 (geometry sub-seed) and OE_Mea_R_8 (measurement sub-seed); wherein the activation domain for a second primer, a second geometry sub-seed, and a second measurement sub-seed may be ATGCTGAGTATTCCC (SEQ ID NO: 4), as used in EX_Yel and OE_Geo_Y_15 (geometry sub-seed) and OE_Mea_Y_8 (measurement sub-seed). By having identical activation domains in a primer and a sub-seed, and a domain that hybridises to a seed-specific amplicon, the seed-specific amplicons are capable of hybridising to the sub-seed and continuing elongation, which appends the activation domain to the seed-specific amplicon.
[0078] Alternatively, or additionally, the activation domain may already be present in the seed-specific amplicon, in which case the primers (e.g. EX primers described herein) are not required.
[0079] Preferably, the activation domain is added via the use of primers, as this allows multiple copies of the seed-specific amplicons to be produced before they become involved in the geometry layer or measurement layer. This approach conveniently reduces or prevents potential PCR bias. The alternative of designing the activation domain to be present in seed-specific amplicons may result in a reaction that is inefficient or prone to bias.
[0080] In some embodiments, the step of forming seed-specific amplicons from each polynucleotide molecule comprises linear spooling amplification (for example, see Figure 14(a)). In some embodiments, the formation of the geometry monomer, measurement monomer and / or marker monomer comprises linear spooling amplification (optionally as a continuation of ongoing linear spooling in the method). An advantage of using linear spooling is that the pre-activation domains and / or activation domains may be excluded, as this method does not require exponential amplification.
[0081] Thus, in the methods disclosed herein, step (c) may comprise linear amplification of the target. By "linear amplification", "linear spooling" or"linear spooling amplification", we include the meaning of increasing target complementary sequence copies in a linear fashion. Linear amplification reactions are performed using a single oligonucleotide primer and produce a single complementary polynucleotide (e.g. DNA) copy of the target sequence with each cycle, rather than the exponential amplification that occurs with traditional polymerase chain reaction (PCR).
[0082] In some embodiments, seed strands may comprise primer sites where primers land, elongation occurs, and the elongated products spool off to yield room for new primers to land. As temperature cycling proceeds, single-strand copies are generated from the initial seeds.
[0083] Linear amplification reaction reagents, ingredients and concentrations may be similar to PCR and well known in the art.
[0084] The terms "oligonucleotide primer" or "primer" refer to a nucleic acid molecule having a sequence of nucleotide residues of at least about 5 nucleotides, more preferably at least about 7 nucleotides, more preferably at least about 9 nucleotides, more preferably at least about 11 nucleotides, and even more preferably at least about 17 nucleotides. In a preferred embodiment, the oligonucleotide primer is preferably between 5-50 nucleotides in length, more preferably between 10-40 nucleotides in length, and even more preferably between 18-30 nucleotides in length. In the context of the present invention, oligonucleotide primers are used to form monomers (polynucleotides) from seeds, i.e. to generate polonies that form nodes of a spatial network.
[0085] A skilled person would understand that the method of the invention is not restricted to any particular oligonucleotide primer sequences. The oligonucleotide primers may be designed to have specificity (i.e. preferentially hybridise) on a particular seed or portion thereof (e.g. the unique sequence identifier and / or a primer domain of a seed). In some embodiments, the oligonucleotide primer used in step (c) is "specific" for the seed (or a domain thereof) and / or the marker (or a domain thereof), in that it must specifically bind to the seed (or a domain thereof) and / or the marker (or a domain thereof). By "specifically bind" we include the meaning that the oligonucleotide primer binds preferentially to the seed (or a domain thereof) and / or the marker (or a domain thereof), and does not bind to sequences that are different to the seed (or a domain thereof) and / or the marker (or a domain thereof); (for example, sequences that have less than 50% or 40% or 30% or 20% or 10% sequence identity). In some embodiments, each primer has specificity to a single type of seed or specificity to the marker (optionally via a unique sequence identifier or barcode being added to the marker). In some embodiments, the primers share specificity for a region present in each type of seed. In some embodiments, the primers share specificity to one type of seed and to the marker. In some embodiments, the primers share specificity to each type of seed and to the marker. Preferably, each primer is specific to a single type of seed (e.g. a first primer with specificity to a first seed, but reduced or no specificity to a second seed), as this reduces the risk of producing unwanted side products. The primers may be used for polony initiation and for final concatemer amplification. After a specific hybridisation to a complementary region of the seed (or a domain thereof), the oligonucleotide primer will provide the 3' hydroxyl end by which DNA polymerase mediated synthesis proceeds.
[0086] By "complementary" or "complementarity" we refer to the ability of a polynucleotide to form hydrogen bond(s) with another polynucleotide sequence, for example by Watson-Crick interactions or non-traditional types of interactions. A percent complementarity indicates the percentage of residues in a first polynucleotide molecule that can form hydrogen bonds with a second polynucleotide molecule. A first polynucleotide may have perfect complementarity to a second polynucleotide, in which case all the contiguous residues of the first polynucleotide will hydrogen bond with the same number of contiguous residues in the second polynucleotide. By "substantially complementary" we refer to a degree of complementarity that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more over a region of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, or more nucleotides. For example, in some embodiments, the term "complementary" simply means that a region or domain of one polynucleotide (i.e. not the entire polynucleotide) is complementary to a region or domain of a further polynucleotide. Preferably, the complementarity for a region or domain is based on overlapping 3' ends of the polynucleotides.
[0087] Sequence identity is related to sequence homology. Computer programs for calculating percent (%) homology or percent identity between two or more sequences are readily available to the skilled person. Examples of other software than may perform sequence comparisons include, but are not limited to, the BLAST package (see Ausubel et al., 1999 ibid - Chapter 18), FASTA (Atschul et al., 1990, J. Mol. Biol., 403-410) and the GENEWORKS suite of comparison tools. Both BLAST and FASTA are available for offline and online searching (see Ausubel et al., 1999 ibid, pages 7-58 to 7-60). Alternatively, % homology may be calculated using multiple alignment features, for example in CLUSTAL (Higgins DG & Sharp PM (1988), Gene 73(1), 237-244) or DNASISTM (Hitachi Software). Once the software has produced an optimal alignment, it is possible to calculate % homology, preferably % sequence identity. The software typically does this as part of the sequence comparison and generates a numerical result.
[0088] By "hybridises" or "hybridisation" we include the meaning of a reaction in which one of more polynucleotides react to form a complex that is stabilised via hydrogen bonding between the bases of the nucleotide residues, for example by Watson Crick base pairing. The complex may comprise two strands forming a duplex structure, three or more strands forming a multi-stranded complex, a single strand that hybridises with itself (e.g. a shRNA), or any combination thereof. A hybridisation reaction may comprise a step in a more extensive process, such as the initiation of PCR, or enzyme- mediated cleavage of a polynucleotide. A sequence capable of hybridising with a given sequence is referred to as the "complement" of the given sequence.
[0089] Hybridisation can be performed under conditions of various stringency. Conditions that increase the stringency of a hybridisation reaction are widely known and published in the art. See, for example, (Sambrook, et al., (1989); Nonradioactive In Situ Hybridization Application Manual, Boehringer Mannheim, second edition).
[0090] The one or more oligonucleotide primer may be modified, for example to improve recognition of the target sequence and / or performance in an amplification reaction. In a preferred embodiment, the oligonucleotide is 5' biotinylated, and / or used together with LNA (locked nucleic acid) modification. This combination increases the melting temperature of the oligonucleotide primer which creates higher specificity and protects against exonuclease activity of polymerases.
[0091] A skilled person would appreciate that oligonucleotide primers used in the method of the invention should have appropriate GC content, for example between 40% and 60%, and that it is more advantageous to use a primer with G or C base on its 3'-OH extremity, which creates a GC clamp and facilitates efficient binding at the target sequence. Preferably, primers should not contain repeats of a given base (PCR Primer: A Laboratory Manual. New York: Cold Spring Harbor Press, 1995).
[0092] In some embodiments, the primers comprise a sequence selected from the group consisting of PR_Red_5N, PR_Yel_5N, and combinations thereof, as shown in Table 1.
[0093] In some embodiments, the step of forming seed-specific amplicons from each polynucleotide molecule comprises PCR amplification (for example, see Figure 14(b)). In some embodiments, the formation of the geometry monomer, measurement monomer and / or marker monomer comprises PCR amplification (optionally as a continuation of ongoing PCR amplification in the method).
[0094] In PCR, cycling through different temperature stages duplicates or "amplifies" the target sequence (e.g. the monomers and / or concatemers) many times over. This doubling is facilitated by short synthetic primer oligonucleotides specific to the target of interest. Once the primers are used up, the reaction stops. In quantitative PCR, a synthetic "probe" sequence is included as well to generate a fluorescent signal with each duplication of the target. The probe is typically designed with a fluorophore at one end and a quencher at the other. While the probe remains intact, the quencher absorbs the light emitted by the fluorophore, preventing it from being detected. As the reaction proceeds, however, the polymerase degrades the probe. This separates the fluorophore from the quencher, leading to a detectable fluorescent signal. Standard PCR or asymmetric PCR-like method where an additional primer pair is added at a lower concentration (necessary to add adapters for the blue and blue striped regions on the sub-seed strands) may be used.
[0095] In the Examples, amplification was carried out for 10 cycles (98 °C 10 s, 67 °C 30 s, 72 °C 30 s), activation for two cycles (98 °C 10 s, 50 °C 30 s, 72 °C 30 s), and subsequent amplification for 23 cycles; however, the number of amplification cycles can be adjusted to achieve the desired level of amplification in a polony - for example, 20, 30, 40, 50 or 60 or more amplification cycles may be used.
[0096] In some embodiments, the seed strands may have additional primer regions in their 5' ends, which can be used at a lower concentration than the primers used to prepare seed-specific amplicons. Following temperature cycling, a PCR process initiates around the seed strands, generating double stranded copies. However, if asymmetrical concentrations of primers are used, a large portion of single stranded copies will also be generated.
[0097] In some embodiments, the step of forming seed-specific amplicons from each polynucleotide molecule comprises rolling circle amplification, for example via the production of single-stranded DNA oligonucleotides using a monoclonal stoichiometric (MOSIC) (for example, see Figure 14(c)). In some embodiments, the formation of the geometry monomer, measurement monomer and / or marker monomer comprises MOSIC (optionally as a continuation of ongoing MOSIC in the method). MOSIC is described in more detail in, for example, Ducani C., Kaul C., Moche M., Shih W. M. and Hdgberg B. Nature Methods 10, p. 647-652 (2013); and Bernardinelli, 2019, Single- Stranded DNA: Methods and Applications in Nanotechnology, Karolinska Institutet. MOSIC uses rolling circle amplification (RCA) and digestion to generate a pool of seed copies from a circle containing the seeds. The MOSIC method may be adapted for use in situ, for example by attaching the RCAs to a gel matrix. For example, single stranded anchor strands may be distributed in the sample instead of seed strands, with the seed sequences put on small circular DNA produced by circularising DNA oligonucleotides complementary to the intended seed strand (i.e. forming a circularised seed). Optionally, the circularised DNA further comprises a hairpin sequence encoding a digestion site for a restriction enzyme. The circularised seeds may be distributed in the sample (e.g. with a gel) and can hybridise with compatible anchors. Using a strand displacing polymerase, such as Phi29, these complexes will generate a long rolling circle product around the original anchor. By subsequently adding and perfusing the sample (e.g. the gel) with the restriction enzyme, many copies of the seed will be released locally in single-stranded form.
[0098] Seed-specific amplicons interact with sub-seed strands, which adds an overlap sequence to the 3' end of the seed-specific amplicon. This reaction may be referred to herein as an activation step. Thus, sub-seeds may also be referred to herein as "activators". I.e. a measurement sub-seed may be considered a measurement activator, and a geometry sub-seed may be considered a geometry activator. Following the activation step, PCR may continue in an amplification step, in which geometry monomers and measurement monomers begin to form.
[0099] Preferably, following activation by a geometry sub-seed in step (c), a first seed-specific amplicon receives a sequence that is complementary to the sequence received by a second seed-specific amplicon (i.e. the two geometry monomers formed from each sub-seed will have complementary 3' ends, which may be referred to herein as a "geometry overlap", "geometry overlapping region" or "geometry overlapping domain"), which allows the two distinct geometry monomers to interact with each other (e.g. Figures 18 and 20). Following interaction of the geometry monomers, they anneal and extend on each other to create a single double-stranded geometry concatemer. Thus, the unique sequence identifier (or portion thereof e.g. the barcode) combinations from this reaction of two geometry monomers can be used to identify the neighbourhoods of polonies and may be used to construct a spatial network. By "activation by a geometry sub-seed" we include the meaning that a portion of the seedspecific amplicon (for example, an activator sequence, an activator domain, and / or a geometry overlapping domain) has hybridised with a portion of a geometry sub-seed, thereby allowing elongation to continue along the geometry sub-seed. Preferably, no elongation occurs on geometry sub-seeds until a seed-specific amplicon hybridises to the geometry sub-seed. In some embodiments, the geometry sub-seed comprises the sequence selected from the group consisting of OE_Geo_R_15, OE_Geo_Y_15, and combinations thereof, as shown in Table 1.
[0100] Preferably, following activation by a measurement sub-seed in step (c), a first seedspecific amplicon and a second seed-specific amplicon receive identical strands from the measurement sub-seeds (i.e. both measurement monomers have the same 3' end and do not interact / hybridise with one another). The identical strands received may be referred to herein as a "measurement overlap", "measurement overlapping region" or "measurement overlapping domain". Preferably, the measurement overlapping domain does not have complementarity to a geometry overlapping domain, but has complementarity to a marker overlapping domain, such that a measurement monomer and marker monomer may hybridise to interact with each other, while excluding hybridisation of a measurement monomers to each other or to geometry monomers (e.g. Figures 19 and 20). Following interaction of a measurement monomer and marker monomer, they anneal and extend on each other to create a single doublestranded measurement concatemer. By "activation by a measurement sub-seed" we include the meaning that a portion of the seed-specific amplicon (for example, an activator sequence, an activator domain, an overlapping domain) has hybridised with a portion of a measurement sub-seed, thereby allowing elongation to continue along the measurement sub-seed. Thus, no elongation occurs on measurement sub-seeds until a seed-specific amplicon hybridises to the measurement sub-seed.
[0101] In some embodiments, the measurement sub-seed comprises the sequence selected from the group consisting of OE_Mea_R_8, OE_Mea_Y_8, and combinations thereof, as shown in Table 1.
[0102] Marker monomers are also referred to herein as "target monomers" or "target polynucleotides". Within the sample there are targets from which a monomer may form, i.e. the marker monomer. Thus, the term "target" may be used interchangeably with the term "marker". For example, the target may be a DNA strand (such as cDNA formed from mRNA present in a cell), which may be referred to herein as a "marker polynucleotide". The target within the sample behaves as a further seed (e.g. an endogenous seed), but is only capable of participating in the measurement reaction (i.e. marker monomers do not hybridise with geometry monomers). Thus, in some embodiments, the marker monomer interacts with and is activated by measurement sub-seeds, which adds a 3' sequence complementary to the measurement monomer from any type of seed (e.g. complementary to monomers from the first seed and the second seed). Following interaction of the marker monomer and measurement monomer, they anneal and extend on each other to create a single double-stranded measurement concatemer, which contains target (e.g. transcriptomic or proteomic) information on one side, and the polony barcode on the other side. Thus, measurement concatemers represent the content of the polonies and are used to deduce transcriptomic and / or proteomic information to the constructed geometry network (since a target is associated with a location). By "activation of a measurement subseed" we include the meaning that a portion of the seed-specific amplicon (for example, an activator sequence) has hybridised with a portion of a measurement sub-seed, thereby allowing elongation to continue along the measurement sub-seed.
[0103] Alternatively, the method may further comprise a third type of sub-seed, a marker sub-seed. I.e. the method comprises one or more geometry sub-seed, one or more measurement sub-seed, and one or more marker sub-seed. The marker sub-seeds (also referred to as "target sub-seeds", "marker activators" or "target activators") are specific for the markers (or targets), and add an overlapping region that is complementary to the measurement overlap (i.e. adds an [overlap mea']) to the 3' end of marker monomers. Thus, the marker monomers hybridise with measurement monomers via the complementary 3' ends of each monomer. In cases where there are multiple measurement monomers (e.g. from a first seed and a second seed, etc), it is preferable that the monomer overlap is the same for all measurement monomers. Accordingly, all options and preferences described herein for geometry sub-seeds and / or measurement sub-seeds may apply to marker sub-seeds.
[0104] Preferably, during activation in step (c), a unique sequence identifier is added to a marker or target within the sample, such as an mRNA from a cell (or a cDNA corresponding to the mRNA), optionally by targeting the poly-A tail of an mRNA, to form a "marker monomer". For example, the poly-A tail may be used as a primer for reverse transcription, which conveniently allows the method to target all mRNA irrespective of its sequence (thus, a transcriptomic analysis can be achieved). The cDNA will then extend all the way through to the 5' end of the mRNA. Optionally, a template-switching oligonucleotide (TSO) can then be used to add a sequence to the 3' end of the cDNA by adding an additional template, which is on the 5' end of the mRNA. The TSO may be used to create adapters for amplification of a cDNA in situ that corresponds to the mRNA target (i.e. in cases where the mRNA is reverse transcribed). Alternatively, or additionally, a marker monomer may be formed based on a polynucleotide tag on an antibody or antigen-binding fragment thereof, which comprises the unique sequence identifier. Since the sequence of the polynucleotide tag can be controlled, a TSO approach may not be required for such embodiments. In some embodiments, the unique sequence identifier comprises a "target domain" or "marker domain" (e.g. as shown in the figures as Tar ID', which is the reverse complement of Tar ID), which is a domain with complementarity to, for example, the mRNA, cDNA or polynucleotide tag target. In some embodiments, the marker monomer further comprises a primer domain, which is a domain to which a primer may hybridise to initiate elongation. The primer domain may be part of the unique sequence identifier, a separate domain added to the unique sequence identifier, or a domain that overlaps with the unique sequence identifier. In some embodiments, the marker monomer further comprises an activation domain, which has homology for a site on a sub-seed. The activation domain may be comprised of two parts, a pre-activation portion ("Pre-act.") and an activator portion. In some embodiments, the pre-activation portion has homology to a further polynucleotide in the sample that contains the reverse complement of the activator portion, such that continued elongation of the marker monomer appends the activator portion to the marker monomer, thereby forming the activator domain. The activator domain may be part of the unique sequence identifier, a separate domain added to the unique sequence identifier, or a domain that overlaps with the unique sequence identifier. In some embodiments, the marker monomer further comprises a "monomer overlap" or "monomer overlapping" domain. Preferably the monomer overlap domain is not complementary to a geometry overlapping domain, but is complementary to a measurement overlapping domain, such that a marker monomer and measurement monomer may hybridise to interact with each other.
[0105] In some embodiments, a measurement overlap is added to the target mRNA through the TSO. For example, a reverse transcription (RT) primer may be designed to have a poly-T sequence (for hybridising with the poly-A tail of an unknown mRNA target), and a measurement overlap that acts as a further primer domain, which results in a cDNA molecule that behaves like an activated monomer (i.e. XXXXX l l l l l , wherein XXXXX is for the measurement overlap that acts as a further primer domain). In this approach, a sub-seed for the marker monomers is not required (i.e. for its activation), as the marker monomers are already able to interact with measurement monomers. For example, the resulting sequence of the marker monomer may be (primer T being a poly-T primer that will hybridise to the mRNA poly-A tail):
[0106] 5’ [primer T - cDNA sequence - measurement overlap] 3' A variation of this method can be seen in Figure 15, in which TSO is used to add preact. T and / or activator T. In this case, the XXXXX region remains a further primer domain, but may also be a pre-activation domain (Pre-act. T). Thus, the resulting sequence of the marker monomer may be:
[0107] 5' [Primer T - cDNA sequence - Pre-act. T] 3'
[0108] The above product resembles an amplicon of a target (i.e. a marker-specific amplicon, equivalent to a seed-specific amplicon), which can be activated through a sub-seed. Thus, all options and preferences described herein for the seed-specific amplicons may also apply for such marker-specific amplicons, and vice versa.
[0109] A further variation of this method can be seen in Figure 16, in which the order of domains for the monomer may be changed. Here, the primer domain may be added through TSO, and the XXXXX region of the reverse transcription primer adds a preactivation domain (pre-act. T). In this case, the resulting sequence of the marker monomer may be:
[0110] 5' [pre-act T - cDNA sequence - primer T'] 3'
[0111] The TSO can be used to add any sequence to the 3' end of a cDNA, and would be used to add pre-activation domains, or activator domains. Alternatively, TSO may be used to add a barcode and / or primer domain.
[0112] In some embodiments, the marker monomers are introduced using artificially created targets. For example, the artificial target may be selected from the group consisting of TR_ArT_l, TR_ArT_2, TR_ArT_3, TR_ArT_4, and combinations thereof, as shown in Table 1.
[0113] In some embodiments, the method includes a step of amplification of the geometry concatemers and measurement concatemers, for example by using exponential PCR. In some embodiments, exponential primers are also used, such as an exponential primer comprising the sequence of EX_Gap, as shown in Table 1.
[0114] In some embodiments, step (c) and / or step (d) further comprises the step of exponentially amplifying the population of concatemers to generate one or more amplicon of each concatemer in the population, optionally wherein the step comprises PCR amplification. Exponential amplification reaction reagents, ingredients and concentrations for PCR are well known in the art.
[0115] In some embodiments, the seed and sub-seeds are positioned and / or immobilised in the sample, optionally wherein the seed and sub-seeds are positioned and / or immobilised in a gel, preferably a hydrogel. In some embodiments, PEG-based gelling components are used to form a hydrogel, for example by using acrylate-PEG and / or thio-PEG gelling components (optionally as part of a buffer). Preferably, gelation occurs after fixation and / or permeabilisation of the sample (or the cells in the sample). Preferably, the gel is dissolvable, cleavable, thermostable, and / or diffusion restrictive, as described herein.
[0116] The number and / or concentration of the seeds affects the size of the polonies that can form in the sample, which can correspond to the resolution obtained by DNA microscopy. In some embodiments, the number and / or concentration of seeds in the sample is lower than the number and / or concentration of sub-seeds in the sample. The concentration may be expressed as a molar concentration, (i.e. M, mol / L, mol / dm3, or mol / m3). In some embodiments, the concentration of seeds in the sample is from about 100 fM to about 1 pM. For example, the concentration of seeds may be from about 200 fM to about 1 pM, from about 300 fM to about 1 pM, from about 400 fM to about 1 pM, from about 500 fM to about 1 pM, from about 600 fM to about 1 pM, from about 700 fM to about 1 pM, from about 800 fM to about 1 pM, from about 900 fM to about 1 pM, from about 950 fM to about 1 pM, from about 960 fM to about 1 pM, from about 970 fM to about 1 pM, from about 980 fM to about 1 pM, or from about 990 fM to about 1 pM. Preferably, the concentration of seeds in the sample is 1 pM.
[0117] In some embodiments, the number of seeds in the sample is from about 10,000 to about 10,000,000. For example, the number of seeds may be from about 20,000 to about 9,000,000, about 30,000 to about 8,000,000, about 40,000 to about 7,000,000, about 50,000 to about 6,000,000, about 60,000 to about 5,000,000, about 70,000 to about 4,000,000, about 80,000 to about 3,000,000, about 90,000 to about 2,000,000, about 100,000 to about 1,000,000, about 200,000 to about 900,000, about 300,000 to about 800,000, about 400,000 to about 700,000, about 500,000 to about 600,000, or any range in between. Preferably, the number of seeds is from about 30,000 to about 3,000,000. In some embodiments, the seeds are spaced uniformly in the sample. In some embodiments, the seeds are clustered within a particular volume of a sample, for example clustered in l / 10thof the volume a sample.
[0118] The diameter of polonies formed from seeds may be calculated by dividing the volume of the sample (or of the gel, such as the hydrogel, comprising the seeds) by the total number of seed strands, and calculating the diameter of that volume as a sphere. In some embodiments, the diameter of the polonies that form is from about 7 pm to about 68 pm, such as about 7 pm to about 32 pm, about 7 pm to about 14 pm, about 14 pm to about 68 pm, about 14 pm to about 32 pm, or about 32 pm to about 68 pm. Although it is possible to sequence libraries derived from polonies with a diameter of 68 pm or larger, the resulting sequencing may have more noise and side reactions. Accordingly, in a preferable embodiment, the polony diameter is smaller than 68 pm, for example >0 pm and <68 pm, preferably between 7 pm and 32 pm, which advantageously reduces background noise.
[0119] Preferably, each sub-seed is at a distinct location to each other, the seeds and the markers / targets. In some embodiments, the sub-seeds are spaced uniformly in the sample.
[0120] In some embodiments, the concentration of sub-seeds in the sample is from about 10 nM to 800 nM. For example, the number and / or concentration of seeds may be from about 30 nM to 800 nM, about 100 nM to 800 nM, from about 200 nM to 700 nM, from about 300 nM to 600 nM, from about 350 nM to 500 nM, from about 400 nM to 500 nM, from about 400 nM to 450 nM, from about 10 nM to 400 nM, from about 30 nM to 400 nM. Preferably, the number and / or concentration of sub-seeds in the sample is about 400 nM.
[0121] In some embodiments, the number of sub-seeds in the sample is from about 10,000 to about 10,000,000. For example, the number of sub-seeds may be from about 20,000 to about 9,000,000, about 30,000 to about 8,000,000, about 40,000 to about 7,000,000, about 50,000 to about 6,000,000, about 60,000 to about 5,000,000, about 70,000 to about 4,000,000, about 80,000 to about 3,000,000, about 90,000 to about 2,000,000, about 100,000 to about 1,000,000, about 200,000 to about 900,000, about 300,000 to about 800,000, about 400,000 to about 700,000, about 500,000 to about 600,000, or any range in between. Preferably, the number of sub-seeds is from about 30,000 to about 3,000,000. In some embodiments, the ratio of geometry sub-seeds to measurement sub-seeds in the sample is 0.1 : 1, 0.2: 1, 0.3: 1, 0.4: 1, 0.5: 1, 0.6: 1, 0.7: 1, 0.8: 1, 0.9: 1, 1: 1, 1:0.9, 1 :0.8, 1:0.7, 1:0.6, 1:0.5, 1:0.4, 1:0.3, 1 :0.2, or 1:0.1. Preferably, the ratio of geometry sub-seeds to measurement sub-seeds in the sample is 1 : 1.
[0122] In some embodiments, the number and / or concentration of measurement sub-seeds in the sample is greater than the number and / or concentration of geometry sub-seeds in the sample. By increasing the number and / or concentration of measurement subseeds in the sample, over that of the geometry sub-seeds, the number and / or concentration of interactions with marker monomers increases (i.e. a greater number and / or concentration of measurement concatemers form over that of geometry concatemers). Thus, the analysed sample provides additional 'measurement' or 'target' information based on what polynucleotides are in the sample. As the measurement concatemers contain a single polony barcode and transcriptomic and / or proteomic information, the determination of the spatial arrangement may then include additional information on what targets (e.g. genes, transcripts, polypeptides and / or proteins) are associated with what location.
[0123] In some embodiments, the number and / or concentration of geometry sub-seeds in the sample is greater than the number and / or concentration of measurement sub-seeds in the sample. By increasing the number and / or concentration of geometry sub-seeds in the sample, over that of the measurement sub-seeds, the number and / or concentration of geometry concatemers increases (i.e. a greater number and / or concentration of geometry concatemers form over that of measurement concatemers). As the geometry concatemers form from the interactions between products from two separate polonies, each geometry concatemer contains polony barcodes of two polonies, and so the analysed sample provides additional 'geometry' or 'location' information, which further improves spatial resolution.
[0124] The use of the sub-seeds not only enables the creation of a precise geometry layer on top of genome wide-compatible DNA microscopy, but it also conveniently enables the ability to tune the technique for higher spatial resolution versus higher read depth (e.g. Figure 20). By tuning the concentration of sub-seeds associated with either the geometry layer or the measurement layer, one can force the reaction to create more information on one versus the other. The end result is a highly flexible technology that can look at variable samples and / or requirements, from small samples of dense epithelial tissue with subcellular resolution to larger samples with empty spaces and fewer cells per volume like germinal centres. Accordingly, the method advantageously allows for the two layers ('geometry' and 'measurement') to be independently tuned. This may be referred to as Two Layer DNA Microscopy (e.g. Figure 22).
[0125] In some embodiments, the geometry concatemers and measurement concatemers are the same size. The advantage of the concatemers being the same size is that they may be isolated simultaneously, for example by isolating as part of the same band in a gel, which is quicker and easier than if the two types of concatemers differ in size and settle at distinct points in a gel. On the other hand, it can be useful to separate the two types of concatemers and sequence them separately. Thus, in some embodiments, the geometry concatemers and measurement concatemers are of different sizes, such that they can be isolated as separate fractions from a gel.
[0126] In some embodiments, the geometry sub-seed comprises a unique sequence identifier. The geometry sub-seed may optionally comprise a 3' end that is blocked from extension. By blocking the 3' end, it prevents the geometry sub-seed (also referred to as a geometry activator) from being extended, which reduces smearing and reduces the amount of side products. In some embodiments, the geometry sub-seed comprises a barcode, and annealing and extension of two geometry monomers forms a geometry concatemer comprising a pair of corresponding barcodes on each strand.
[0127] Methods of blocking the 3' end of polynucleotide sequences are known to the skilled person. For example, a 3' inverted T base may be used. Alternatively, or additionally, blocking can be achieved with 3' phosphate, 3' fluorophores, 3' fluorophore quenchers, non-canonical based means, ablating a base in the dNTPs mixture, or modification that replaces the 3' OH group with something that would prevent elongation.
[0128] In some embodiments, the geometry monomers are not activated by primers and / or need to interact with a sub-seed (i.e. a geometry sub-seed) to obtain an activation domain in order to activate (i.e. the seed-specific amplicon interacts with a geometry sub-seed to form a geometry monomer, such that elongation may continue on a monomer). As the seed-specific amplicons land and extend on the sub-seed strands, they acquire a homology region (and, optionally, a barcode or half-fusion barcode). Due to the homology region, the geometry monomers can extend on other geometry monomers and fuse to created geometry concatemers. Preferably, each seed has its own associated geometry sub-seeds (i.e. a first seed that is distinct to a second seed comprises first geometry sub-seeds that are distinct from the second geometry subseeds). This option is advantageous, as it precludes the creation of concatemers with themselves and creates a bipartite network (i.e. first geometry monomers do not hybridise with other first geometry monomers). By reading out the sequence of the geometry concatemers, it can be inferred which seed is proximate to another seed. This topological information alone can be used to computationally infer geometry (i.e. topological distance - geometrical distance).
[0129] In some embodiments, the measurement sub-seed comprises a unique sequence identifier. The measurement sub-seed may optionally comprise a 3' end that is blocked from extension. By blocking the 3' end, it prevents the measurement sub-seed (also referred to as a measurement activator) from being extended, which reduces smearing and reduces the amount of side products. In some embodiments, the measurement sub-seed comprises a barcode, and annealing and extension of the measurement monomers and marker monomer forms a measurement concatemer comprising the newly introduced barcode (i.e. in addition to the barcode present in the seed-specific amplicon). Optionally, a barcode may also be added to the marker monomer (for example, by using a template-switching oligonucleotide that incorporates the barcode). Thus, the measurement concatemer may form with a pair for newly-introduced barcodes, one from the measurement monomer and one from the marker monomer.
[0130] A unique sequence identifier is a means to identify a particular location (also referred to herein as a node or seed) of a network, and may also be referred to as a Unique Molecular Identifier (UMI). By "unique sequence identifier", "UMI", "barcode", or "molecular barcode", we include the meaning of a pool or string of nucleic acid sequences that is added to, or forms part of, a particular RIMA or DNA molecule (or populations thereof) and can act as a means to uniquely identify the polynucleotide, thereby allowing the grouping or identification of concatemers. By including a unique sequence identifier, each transcript is assigned a traceable origin.
[0131] Measurement concatemers may be distinguishable based on a UMI, concatenated seed barcodes, concatenated sub-seed barcodes, and / or the polynucleotide sequence elongated from the target of the marker monomer (e.g. an endogenous mRNA sequence), independent of subsequent amplification. The length of the UMI can be scaled to match the scale of the total expected number of transcripts sequence around each seed (for example, by the method outlined in Islam et al., 2014, Nat Methods, 11: 163-166). In some embodiments, it is the combination of a UMI, the RNA sequence, and its concatenated seed barcodes that together make the measurement concatemer traceable and attributable to a specific location. In some embodiments, the 3' UMI, if present, may be discarded. In some embodiments, the addition of further barcodes following extension on sub-seeds is sufficient to uniquely determine a traceable origin.
[0132] In some embodiments, the unique sequence identifier comprises at least one barcode. Barcodes may further be defined by the molecule for which they are attached. For example, a barcode being added to a monomer forming from a seed may be described as a "seed barcode" or a "polony barcode"; a barcode being added to a monomer forming from a geometry sub-seed may be described as a "geometry barcode"; a barcode being added to a monomer forming from a measurement sub-seed may be described as a "measurement barcode"; and a barcode being added to a monomer forming from a marker (i.e. the barcode within a marker monomer) may be described as a "marker barcode".
[0133] In some embodiments, a barcode is added to the marker monomer. For example, a barcode may be added using TSO, for example by using oligo-T and a template switching ribo-G oligonucleotide, as described herein.
[0134] In some embodiments, the methods use at least two types of barcodes. For example, a first type of barcode may identify a point in space (e.g. a node in the network, or a point of origin for a polony), which may be referred to as a "seed barcode", while a second type of barcode may be used to track products formed within each polony, which may be referred to as a "geometry barcode", "measurement barcode", "marker barcode" or "fusion barcode". Geometry barcodes, measurement barcodes and marker barcodes may also be referred to as fusion barcodes.
[0135] By "fusion barcode" (also described herein as "event barcode"), we include the meaning that the barcode derives from two or more separate barcodes of different origin. For example, a geometry monomer forming from first geometry sub-seed may comprise one half of a fusion barcode (i.e. a "half-fusion barcode" or "first half-fusion barcode"), which joins with a geometry monomer forming from a second geometry sub-seed that comprises the other half of the fusion barcode (alternatively referred to as a "second half-fusion barcode"). Thus, two half-fusion barcodes are present, which forms a fusion barcode in the geometry concatemer. The same applies for barcodes added to measurement monomers forming from measurement sub-seeds (i.e. a first half-fusion barcode) and barcodes added to marker monomers (i.e. a second halffusion barcode), in which two half-fusion barcodes form the fusion barcode in the measurement concatemer. With respect to geometry monomers comprising a half-fusion barcode, the combination of two geometry monomers forms full fusion barcodes. Preferably, the completed fusion barcodes are unique to each molecular fusion event, and so the number of unique fusion events (which can be inferred by sequencing the fusion barcodes) determines the proximity of amplicon clouds (i.e. proximity of polonies or seeds to each other). Clouds close to each other generate more unique fusion barcodes than clouds further away from each other. This information can be used to further refine the distance estimation between transcripts during geometrical reconstruction.
[0136] Barcodes may be a stretch of 4 to 40 random nucleotides, but may be longer. In some embodiments, the barcode is from 4 to 40 nucleotides: for example, from 5 to 39 nucleotides, from 6 to 38 nucleotides, from 7 to 37 nucleotides, from 8 to 36 nucleotides, from 9 to 35 nucleotides, from 10 to 34 nucleotides, from 11 to 33 nucleotides, from 12 to 32 nucleotides, from 13 to 31 nucleotides, from 14 to 30 nucleotides, from 15 to 29 nucleotides, from 16 to 28 nucleotides, from 17 to 27 nucleotides, from 18 to 26 nucleotides, from 19 to 25 nucleotides, from 20 to 24 nucleotides, or from 21 to 23 nucleotides.
[0137] In some embodiments, one or more of the barcodes is at least 4 nucleotides in length, at least 5 nucleotides in length, at least 6 nucleotides in length, at least 7 nucleotides in length, at least 8 nucleotides in length, at least 9 nucleotides in length, at least 10 nucleotides in length, at least 11 nucleotides in length, at least 12 nucleotides in length, at least 13 nucleotides in length, at least 14 nucleotides in length, at least 15 nucleotides in length, at least 16 nucleotides in length, at least 17 nucleotides in length, at least 18 nucleotides in length, at least 19 nucleotides in length, at least 20 nucleotides in length, at least 21 nucleotides in length, at least 22 nucleotides in length, at least 23 nucleotides in length, at least 24 nucleotides in length, at least 25 nucleotides in length, at least 26 nucleotides in length, at least 27 nucleotides in length, at least 28 nucleotides in length, at least 29 nucleotides in length, at least 30 nucleotides in length, at least 31 nucleotides in length, at least 32 nucleotides in length, at least 33 nucleotides in length, at least 34 nucleotides in length, at least 35 nucleotides in length, at least 36 nucleotides in length, at least 37 nucleotides in length, at least 38 nucleotides in length, at least 39 nucleotides in length, at least 40 nucleotides in length, or longer.
[0138] Since the unique sequence identifier typically needs to track individual molecules in a sample, it is preferable that the unique sequence identifier is long enough to minimise the risk of two seed barcodes being indistinguishable, which can cause errors in reconstruction. Preferably, the seed barcodes are at least 30 nucleotides in length.
[0139] As mentioned above, fusion barcodes may form from two separate half-fusion barcodes. Since the fusion barcodes form from two separate half-fusion barcodes, they may be shorter than seed barcodes, while still maintaining uniqueness. Additionally, as these barcodes are typically used to track molecules within a polony (and not in the full sample), they may be shorter in length, since the chance of finding an identical sequence is relatively smaller within the same polony, even though that chance is higher if they were identified in the full sample. Thus, in some embodiments, an individual barcode (i.e. a single half-fusion barcode) added from a sub-seed is shorter in length relative to the barcode added from a seed. Preferably, the half-fusion barcodes used for geometry monomers are 15 nucleotides in length, such that the fusion barcode formed in a geometry concatemer is a total of 30 nucleotides in length. Preferably, the half-fusion barcodes used for measurement monomers and marker monomers are 8 nucleotides in length, such that the fusion barcode formed in a measurement concatemer is a total of 16 nucleotides in length.
[0140] If the barcodes are too long, it increases the chance of forming sequences that resemble primers. This would mean that primers could land at the barcode region and create shorter products using the wrong part of the molecule. These are considered off-products on a reaction level. Although there are strategies to reduce such off- products (also referred to as "side products"), they have conveniently not been necessary in the methods described herein.
[0141] On a reconstruction level, if the barcodes are too short, it may increase the chance of finding the same or very similar sequences across the sample (off-products / side products on a reconstruction level). If this happens, an identifier is no longer unique, meaning that (for example) point A that interacts with points B, C, and D cannot be distinguished from point W that interacts with points X, Y, and Z. Instead, it would be interpreted as a single point (A / W) that interacts with points B, C, D, X, Y, and Z, which is incorrect and risks warping the reconstruction of the sample. Although there are also ways to reduce or remove this risk (as described in Kloosterman, et al., 2024, Nature Computational Science, 4: 119-127), they have conveniently not been necessary in the methods described herein.
[0142] Preferably, the methods of the invention have reduced, or no, off-products, on a reaction level and / or a reconstruction level, optionally in comparison with a suitable control (for example, in comparison with a method using barcodes long enough that they act as primer sites).
[0143] In some embodiments, the unique sequence identifier comprises at least one primer sequence. The primer may contribute towards the unique sequence identifier. However, it is also possible to strip off the primers, e.g. bioinformatically, such that the primer sequences do not contribute towards the unique sequence identifier (i.e. since the sequence of the primer is known, it can be excluded from the total sequenced concatemer). In such cases, the primers may only be necessary for reactions (i.e. elongation and / or amplification) to occur, and are not necessary for identification of a unique sequence identifier or a concatemer. Accordingly, in some embodiments, the primer(s) are excluded from the sequenced concatemer(s).
[0144] In some embodiments, the unique sequence identifier comprises at least one preactivation sequence. In some embodiments, the unique sequence identifier comprises at least one activator sequence. The pre-activation sequence is also referred to herein as "pre-activation portion" or "Pre-act.", and the activator sequence is also referred to herein as "activator portion". These two portions may collectively be called an activator domain. The sequences of the activator domain (or the portions thereof) vary depending on the type of monomer. For geometry monomers, the sequence of the activator domain (or the portions thereof) in one geometry monomer (e.g. from a first geometry sub-seed) is complementary to that of another geometry monomer (e.g. from a second geometry sub-seed). For measurement monomers and marker monomers, the sequence of the activator domain (or the portions thereof) in one monomer is complementary that of the other monomer (and, preferably, is not complementary to a geometry monomer or to monomers of the same type, i.e. a first measurement monomer is not complementary to a second measurement monomer, and a first marker monomer is not complementary to a second marker monomer).
[0145] The marker monomer may form from a polynucleotide target in the sample (e.g. an endogenous nucleic acid), or a polynucleotide target introduced to the sample (for example, using an antibody or antigen-binding fragment thereof that comprises a polynucleotide tag), which may be an unknown sequence. Cells within the sample comprise polynucleotide targets (for example, genomic DNA, mRNA, or cDNA transcribed from mRNA) and protein targets (for example, intracellular proteins and cell surface proteins). Cells may also secrete proteins extracellularly. Targeting means for polynucleotides and proteins are known to the skilled person. Accordingly, intracellular targets, cell surface targets, and extracellular targets are contemplated. For example, an unknown mRNA could be tagged on its 5' end using a reverse transcriptase enzyme with terminal transferase activity (i.e. by using a templateswitching oligonucleotide (TSO)). Template switching permits ligation-free incorporation of a 5' adapter during reverse transcription. The reverse transcriptase adds non-templated nucleotides to the 3' end of the cDNA when reaching the 5' end of the RIMA template, and the non-templated nucleotides can then anneal to a TSO with a known sequence of choice. The TSO can prompt the reverse transcriptase to switch from the RNA template to the TSO, resulting in a cDNA with a universal sequence complementary to the TSO at the 3' end. Thus, it is possible to add a targetable site to a target of unknown sequence, from which a marker monomer can form.
[0146] Alternatively, or additionally, a protein (irrespective of being intracellular, extracellular, and / or on the cell surface) may be targeted with an antibody or antigen-binding fragment thereof. The antibody or antigen-binding fragment thereof may have a polynucleotide tag, e.g. an oligonucleotide (such as DNA or cDNA) sequence conjugated thereto, from which a marker monomer can form. By targeting a protein in this way, the method may also provide a proteomic analysis.
[0147] Accordingly, in some embodiments, the marker monomer comprises a marker polynucleotide. The marker polynucleotide may be selected from the group consisting of: mRNA, cDNA, and DNA. The marker polynucleotide (for example, the mRNA, cDNA or DNA) may be conjugated to an antibody or antigen-binding fragment thereof. The marker polynucleotide may be endogenous to the sample or may be exogenously introduced to the sample. For example, the marker polynucleotide may be added as a cDNA (e.g. Green Fluorescent Protein (GFP) cDNA).
[0148] The term "exogenous" is intended to include that the marker polynucleotide is heterologous to the genomic DNA of a cell (or all cells) in the sample.
[0149] By "heterologous to the genomic DNA of the cell" we mean that the marker polynucleotide is foreign to the genomic DNA of the cell. For example, the sequence may differ from the "normal" or "natural" sequence of the genomic DNA of the cell, and / or it may comprise or consist of sequence that has been exogenously introduced or inserted into the genomic DNA of the cell. Such insertions may include, for example, identifiable elements, transposable elements, polymorphic germline transposable elements, somatic transposable elements and / or DNA sequences which do not naturally occur in a host genomic DNA (for example, retroviruses, synthetic constructs, whole genes, mutated sequences).
[0150] Advantageously, the methods described herein can be used for samples in which the target for a marker monomer is unknown. Thus, in some embodiments, the polynucleotide sequence of the marker polynucleotide has not been determined. This may be referred to herein as a non-targeted approach, since the marker monomers form without any knowledge of what polynucleotide sequences may be present. Instead, the target is not identified until after the measurement concatemers have been sequenced. Thus, in some embodiments, the method determines the identity of the target after a sequencing step in the method. By "determines the identity" we include the meaning that the target was unknown prior to commencing the method, or becomes known following sequencing and comparison with a database of known sequences.
[0151] In some embodiments, the marker polynucleotide is not GAPDH and / or ACTB, which means that the target for forming the marker monomer was not determined, prior to the formation of the monomers and / or concatemers in the method, to be GAPDH and / or ACTB. Since the method allows for a non-targeted approach, it is possible that GAPDH and / or ACTB will be identified (for example, following sequencing) and associated with a particular seed or location. However, such a determination occurs after formation of the concatemers in the method, and the targets are not used to initiate the methods (i.e. GAPDH and / or ACT are not used as pre-defined nucleic acid sequences in the method). Thus, GAPDH and / or ACTB are conveniently unnecessary as beacon molecules in the sample. Indeed, no beacon molecules are required by the claimed methods.
[0152] In some embodiments, the sub-seeds are incapable of extending until a seed-specific amplicon hybridises with the sub-seed. By "incapable of extending" we include the meaning that the sub-seed sequences cannot elongate (e.g. by PCR) during the method until a seed-specific amplicon interacts with the sub-seed, for example via an activator domain. Preferably, the sub-seed sequences cannot elongate during the method irrespective of a seed-specific amplicon hybridising with it, for example by blocking the 3' end of the sub-seed from extension.
[0153] In the methods described herein, step (c), part (i), may be referred to as a 'geometry layer', while step (c), part (ii), may be referred to as a 'measurement layer'. Thus, the geometry layer forms geometry concatemers from the coupling of two geometry monomers; and the measurement layer forms measurement concatemers from the coupling of a measurement monomer with a marker monomer, as follows:
[0154] Geometry layer: [first geometry monomer] + [second geometry monomer] = [geometry concatemer] .
[0155] Measurement layer: [measurement monomer] + [marker monomer] = [measurement concatemer].
[0156] In some embodiments, the step of determining the spatial arrangement of the targetable molecules comprises determining the polynucleotide sequence of one or more geometry concatemer and / or one or more measurement concatemer formed in step (c). Preferably, the polynucleotide sequences of one or more geometry concatemer and one or more measurement concatemer is determined. In some embodiments, a plurality of geometry concatemers is determined, for example at least 2 copies, at least 3 copies, at least 4 copies, at least 5 copies, at least 6 copies, at least 7 copies, at least 8 copies, at least 9 copies, at least 10 copies, at least 20 copies, at least 30 copies, at least 40 copies, at least 50 copies, at least 60 copies, at least 70 copies, at least 80 copies, at least 90 copies, at least 100 copies, at least 200 copies, at least 300 copies, at least 400 copies, at least 500 copies, at least 600 copies, at least 700 copies, at least 800 copies, at least 900 copies, at least 1,000 copies, at least 2,000 copies, at least 3,000 copies, at least 4,000 copies, at least 5,000 copies, at least 6,000 copies, at least 7,000 copies, at least 8,000 copies, at least 9,000 copies, at least 10,000 copies, at least 20,000 copies, at least 30,000 copies, at least 40,000 copies, at least 50,000 copies, at least 60,000 copies, at least 70,000 copies, at least 80,000 copies, at least 90,000 copies, at least 100,000 copies, at least 200,000 copies, at least 300,000 copies, at least 400,000 copies, at least 500,000 copies, at least 600,000 copies, at least 700,000 copies, at least 800,000 copies, at least 900,000 copies, at least 1,000,000 copies or more are determined. In some embodiments, a plurality of measurement concatemers is determined, for example at least 2 copies, at least 3 copies, at least 4 copies, at least 5 copies, at least 6 copies, at least 7 copies, at least 8 copies, at least 9 copies, at least 10 copies, at least 20 copies, at least 30 copies, at least 40 copies, at least 50 copies, at least 60 copies, at least 70 copies, at least 80 copies, at least 90 copies, at least 100 copies, at least 200 copies, at least 300 copies, at least 400 copies, at least 500 copies, at least 600 copies, at least 700 copies, at least 800 copies, at least 900 copies, at least 1,000 copies, at least 2,000 copies, at least 3,000 copies, at least 4,000 copies, at least 5,000 copies, at least 6,000 copies, at least 7,000 copies, at least 8,000 copies, at least 9,000 copies, at least 10,000 copies, at least 20,000 copies, at least 30,000 copies, at least 40,000 copies, at least 50,000 copies, at least 60,000 copies, at least 70,000 copies, at least 80,000 copies, at least 90,000 copies, at least 100,000 copies, at least 200,000 copies, at least 300,000 copies, at least 400,000 copies, at least 500,000 copies, at least 600,000 copies, at least 700,000 copies, at least 800,000 copies, at least 900,000 copies, at least 1,000,000 copies or more are determined. In some embodiments, the same number of geometry concatemers and measurement concatemers is determined. In some embodiments, a different number of geometry concatemers and measurement concatemers is determined. By "determined" we include the meaning of "obtained", i.e. as a result of the method a particular copy number of concatemers could be obtained.
[0157] In some embodiments, the copy number of the geometry concatemers and / or measurement concatemers obtained by the method is at a copy number of 1-10 copies, 1-20 copies, 1-30 copies, 1-40 copies, 1-50 copies, 1-60 copies, 1-70 copies, 1-80 copies, 1-90 copies, 1-100 copies, 1-125 copies, 1-150 copies, 1-175 copies, 1-200 copies, 1-225 copies, 1-250 copies, 1-275 copies, 1-300 copies, 1-400 copies, 1-500 copies, 1-600 copies, 1-700 copies, 1-800 copies, 1-900 copies, 1-1,000 copies, 1- 2,000 copies, 1-3,000 copies, 1-4,000 copies, 1-5,000 copies, 1-10,000 copies, 1- 25,000 copies, 1-50,000 copies, 1-75,000 copies, 1-100,000 copies, 1-200,000 copies, 1-300,000 copies, 1-400,000 copies, 1-500,000 copies, 1-1,000,000 copies, or 1,000,000 or more copies. Preferably the target sequence is present at a copy number of 1-500,000 copies, more preferably 1-250,000 copies, yet more preferably 1- 100,000 copies, yet more preferably 1-50,000 copies, most preferably 1-10,000 copies.
[0158] In some embodiments, the method further comprises one or more of the following steps (in any order):
[0159] 1. gel extraction, optionally an automated gel extraction;
[0160] 2. dissolving the gel, optionally using an alkaline buffer with DTT;
[0161] 3. isolation and / or purification of the concatemers;
[0162] 4. tagmentation;
[0163] 5. a circularisation protocol; and / or
[0164] 6. sequencing, preferably high-throughput sequencing.
[0165] In some embodiments, the method further comprises the step of gel extraction (also referred to as "gel purification"), for which techniques are known to the skilled person. For example, a portion of interest in the gel (i.e. a band representative of the geometry concatemers and / or measurement concatemers) may be excised from the gel. Preferably, the gel is extracted using automated gel extraction (such as a Pippin™ Prep method), which advantageously creates a tight range surrounding the peak / band of interest in the gel. This automated approach provides a cleaner extraction of the target of interest (i.e. the geometry concatemers and / or measurement concatemers) to the exclusion of other polynucleotides that may be present in the gel, in comparison with non-automated methods.
[0166] Following excision of the gel, the method may further comprise the step of dissolving the gel, such as by using an alkaline buffer with dithiothreitol (DTT). This step liberates the geometry concatemers and / or measurement concatemers from the gel and other components that may affect downstream applications, such as sequencing.
[0167] In some embodiments, the method further comprises the step of isolating and / or purifying the geometry concatemers and / or measurement concatemers. Methods of DNA isolation (i.e. extraction) and purification (including from gels, for example) are known in the art. For example, the concatemers may be purified using paramagnetic beads that selectively bind nucleic acids, such as the AMPure DNA purification method (Beckman Coulter), or by ethanol purification.
[0168] The term "isolated", or "isolating", as used herein includes nucleic acid (such as the geometry concatemers and / or measurement concatemers formed in step (c)) separated from at least one other component (e.g. polypeptide), such as those present with the nucleic acid in its natural source. It will be appreciated that approaches for isolating, purifying and storing polynucleotides are routine in the art.
[0169] The term "purified", or "purifying", as used herein includes the meaning that the indicated polynucleotides (such as geometry concatemers and / or measurement concatemers) are present in the substantial absence of other biological macromolecules (biomolecules), e.g. polypeptides, proteins, and the like. In one embodiment, the polynucleotide is purified such that it constitutes at least 70% or 80% or 85% or 90% or 95% by weight, more preferably at least 99% by weight, of the indicated biological macromolecules present (but water, buffers, and other small molecules, especially molecules having a molecular weight of less than 1000 Daltons, can be present). Impurities may include enzymes, buffer components, proteins, lipids, RNA, and intermediate products, for example. Intermediate products may include the seeds, sub-seeds, monomers that are formed prior to the concatemer formation in step (c), and / or off-products / side-products. In some embodiments, the method further comprises the step of dissolving the gel. The gel may be dissolved using an alkaline buffer with DTT. After dissolving the gel, the concatemers may be harvested and / or isolated in bulk and prepared for sequencing, for example by using adapter ligation.
[0170] In some embodiments, the method further comprises the step of tagmentation. By "tagmentation" we include the meaning of a process for fragmenting of double stranded DNA and integrating of a known polynucleotide sequence (sometimes called a "sequencing adapter") into DNA using a transposase. Those skilled in the art will understand that when performing tagmentation, transposases randomly fragment the double-stranded DNA into smaller fragments and add sequencing adapters simultaneously, thereby generating double-stranded DNA having sequencing adapters at each end.
[0171] Any transposase may be used in the methods of the invention. In some embodiments, MuA orTnY transposase could be used for tagmentation (Liscovitch-Brauer etal, 2021. Nat. Biotechnol., 39: 1270-1277). In a preferred embodiment, a hyperactive mutant transposase could be used for tagmentation. In a preferred embodiment, the transposase is the Tn5 transposase. It will be appreciated that transposase, such as Tn5 transposase, can be engineered to introduce any known polynucleotide sequence or not to insert any sequence into DNA.
[0172] Tagmentation precludes sequencing of both ends of the polynucleotide sequence, in which case the 3' UMI must be dropped, or a circularisation protocol may be used to pair the ends for sequencing those regions.
[0173] In some embodiments of the methods disclosed herein, the size range of the fragments following tagmentation is about 50 base pairs to about 2000 base pairs in length, about 50 base pairs to about 1900 base pairs in length, about 50 base pairs to about 1800 base pairs in length, about 50 base pairs to about 1700 base pairs in length, about 50 base pairs to about 1600 base pairs in length, about 50 base pairs to about 1500 base pairs in length, about 50 base pairs to about 1400 base pairs in length, about 50 base pairs to about 1300 base pairs in length, about 50 base pairs to about 1200 base pairs in length, about 50 base pairs to about 1100 base pairs in length, about 50 base pairs to about 1000 base pairs in length, about 50 base pairs to about 950 base pairs in length, about 50 base pairs to about 900 base pairs in length, about 50 base pairs to about 850 base pairs in length, about 50 base pairs to about 800 base pairs in length, about 50 base pairs to about 750 base pairs in length, about 50 base pairs to about 700 base pairs in length, about 50 base pairs to about 650 base pairs in length, about 50 base pairs to about 600 base pairs in length, about 50 base pairs to about 550 base pairs in length, about 50 base pairs to about 500 base pairs in length, about 50 base pairs to about 450 base pairs in length, about 50 base pairs to about 400 base pairs in length, about 50 base pairs to about 350 base pairs in length, about 50 base pairs to about 300 base pairs in length, about 50 base pairs to about 250 base pairs in length, about 50 base pairs to about 200 base pairs in length, about 50 base pairs to about 150 base pairs in length, about 50 base pairs to about 100 base pairs in length, about 100 base pairs to about 1500 base pairs in length, about 150 base pairs to about 1400 base pairs in length, about 200 base pairs to about 1300 base pairs in length, about 250 base pairs to about 1200 base pairs in length, about 300 base pairs to about 1100 base pairs in length, about 350 base pairs to about 1000 base pairs in length, about 400 base pairs to about 1000 base pairs in length, about 450 base pairs to about 950 base pairs in length, about 500 base pairs to about 900 base pairs in length, about 550 base pairs to about 850 base pairs in length, about 600 base pairs to about 800 base pairs in length, about 650 base pairs to about 750 base pairs in length, about 700 base pairs to about 1500 base pairs in length, about 750 base pairs to about 1500 base pairs in length, about 800 base pairs to about 1500 base pairs in length, about 850 base pairs to about 1500 base pairs in length, about 900 base pairs to about 1500 base pairs in length, about 950 base pairs to about 1500 base pairs in length, about 1000 base pairs to about 1500 base pairs in length, about 1100 base pairs to about 1500 base pairs in length, about 1200 base pairs to about 1500 base pairs in length, about 1300 base pairs to about 1500 base pairs in length, or about 1400 base pairs to about 1500 base pairs in length. Preferably the size range of exponentially amplified fragments is about 100 base pairs to about 1200 base pairs in length, more preferably 200 base pairs to 800 base pairs in length, most preferably 400 base pairs to 600 base pairs in length.
[0174] In some embodiments, the method further comprises a circularisation protocol. Methods and protocols of circularising nucleotide sequences are known to the skilled person. Accordingly, in some embodiments, the concatemers (i.e. the geometry concatemers and / or measurement concatemers) are circularised. Circularised products may be referred to as plasmids or vectors (which may be used interchangeably). For example, the method may further comprise the preparation of geometry concatemer plasmids and / or measurement concatemer plasmids. By circularising the concatemers, the ends are paired, which may allow the 3' end of a unique sequence identifier to be sequenced. In some embodiments, the method further comprises sequencing, for example sequencing the concatemers formed in step (c). By "sequencing" we include the meaning of determining the nucleic acid sequence, i.e. the order of nucleotides in a nucleic acid sequence. DNA sequencing is widely applied in determining the sequence of synthetic DNA sequences, whole genes or fragments of genes, larger genetic regions (i.e. clusters of genes or operons), full chromosomes, or entire genomes of any organism.
[0175] In some embodiments, the sequencing is high-throughput sequencing, preferably short-read high throughput sequencing. Longer read sequencing is also compatible with the methods described herein. In some embodiments, the short-read sequencing method is selected from the list consisting of: massive parallel short-read sequencing; DNA nanoball sequencing (Drmanac et al, 2010. Science, 327(5961):78-81); Illumina dye Sequencing (Solexa sequencing), (Meyer and Kircher, 2010. Cold Spring Harb Protoc); 454 pyrosequencing (Nyren and Lundin, 1985. Analytical Biochemistry, 151(2):504-509); SOLiD sequencing (Shendure and Ji, 2008. Nature Biotechnology, 26: 1135-1145); Helicos single molecule fluorescent sequencing (Thompson and Steinmann, 2010. Curr Protoc Mol Biol, 7 Unit 7.10); combinatorial probe anchor synthesis (cPAS), (Fehlmann et al, 2016. Clinical Epigenetics, 8(123)); polony sequencing (Porreca, Shendure and Church, 2006. Curr Protoc Mol Biol, 7 Unit 7.8); electrical sequencing chips (e.g. GenapSys) (Lahens et al, 2017. BMC Genomics. 18(602)); or combinations thereof. In some embodiments, long-read sequencing methods can be used, for example, PacBio (Eid et al, 2009. Science, 323(5910): 133- 138).
[0176] In some embodiments, a sequencing library (or "library") is prepared prior to sequencing. By "sequencing library", or "library", we include the meaning of a plurality of double stranded concatemers obtained in step (c).
[0177] Thus, in some embodiments of the methods disclosed herein, the method further comprises the step of library preparation. By "library preparation" we include the meaning of pooling of the concatemers generated in step (c), which allows multiple concatemers from different seeds or polonies within a sample to be sequenced simultaneously. Advantageously, since the concatemers already have unique barcodes (or combinations of barcodes), there is no need for an indexing step, i.e. to add a unique identifier (i.e. an index) to the concatemers during library preparation. In some embodiments of the methods disclosed herein, the library comprises sequences which are at least 150 base pairs in length, at least 160 base pairs in length, at least 170 base pairs in length, at least 180 base pairs in length, at least 190 base pairs in length, at least 200 base pairs in length, at least 210 base pairs in length, at least 220 base pairs in length, at least 230 base pairs in length, at least 240 base pairs in length, at least 250 base pairs in length, at least 260 base pairs in length, at least 270 base pairs in length, at least 280 base pairs in length, at least 290 base pairs in length, at least 300 base pairs in length, at least 310 base pairs in length, at least 320 base pairs in length, at least 330 base pairs in length, at least 340 base pairs in length, at least 350 base pairs in length, at least 360 base pairs in length, at least 370 base pairs in length, at least 380 base pairs in length, at least 390 base pairs in length, at least 400 base pairs in length, at least 420 base pairs in length, at least 440 base pairs in length, at least 460 base pairs in length, at least 480 base pairs in length, at least 500 base pairs in length, at least 520 base pairs in length, at least 540 base pairs in length, at least 560 base pairs in length, at least 580 base pairs in length, at least 600 base pairs in length, at least 620 base pairs in length, at least 640 base pairs in length, at least 680 base pairs in length, at least 700 base pairs in length, at least 720 base pairs in length, at least 740 base pairs in length, at least 760 base pairs in length, at least 780 base pairs in length, at least 800 base pairs in length, at least 820 base pairs in length, at least 840 base pairs in length, at least 860 base pairs in length, at least 880 base pairs in length, at least 900 base pairs in length, at least 920 base pairs in length, at least 940 base pairs in length, at least 960 base pairs in length, at least 980 base pairs in length, or at least 1000 base pairs in length. Preferably, the library comprises sequences which are about 200 base pairs to about 500 base pairs in length, even more preferably 250 base pairs in length, 408 base pairs in length, or a mixture thereof.
[0178] The methods described herein may use sequencing to determine spatial organisation. These methods conveniently remove the need for optical imaging (i.e. optical microscopy), and the spatial workload may be offloaded to a computer, which drastically reduces the experimental protocol compared with other spatial sequencing techniques. Additionally, these methods remove the need for preliminary slide characterisation of the sample, which broadens the application of the method to various settings (such as those with low cell densities and empty space).
[0179] Computational methods for spatial transcriptomics are known to the skilled person. For example, the mathematical analysis may be performed as described in Hoffecker et al., 2019, Kloosterman et al., 2024, and / or Qian et al., 2023 (Volumetric imaging of an intact organism by a distributed molecular network, bioRxiv, Preprint).
[0180] In some embodiments, determining the spatial arrangement of the two or more targetable molecules in step (d) comprises an algorithmic method of analysis, preferably using the Minimum Indirect Path (MiniPath) analyses outlined in Kloosterman et al., 2024 and Example 4.
[0181] The methods described herein may be used to identify a disease, disorder, or condition in a patient. For example, the disease, disorder, or condition may be associated with altered geometry of cells within the sample. Alternatively, or additionally, the disease, disorder, or condition may be associated with the identification of a particular marker (e.g. gene) at a particular location in the sample. This may be the case for cancer, in which the interactions and / or mapping of tumour cells with healthy cells may provide information for diagnostic and / or treatment options. For example, the readout for a geometry layer and / or measurement layer may provide a diagnosis of cancer; or may provide information on responsiveness to a cancer treatment. In some embodiments, such methods may compare a test sample with an appropriate control (e.g. a healthy control in which there is no cancer or no suspicion of cancer). Thus, the methods described herein may be a method of diagnosis. In some embodiments, the method further comprises the step of treating the identified disease, disorder, or condition.
[0182] In some embodiments, the cancer is selected from the group consisting of: bladder cancer, breast cancer, colon and rectal cancer, endometrial cancer, kidney cancer, leukaemia, liver cancer, lung cancer, melanoma, non-Hodgkin lymphoma, pancreatic cancer, prostate cancer, and thyroid cancer.
[0183] Upon identification (i.e. diagnosis) of a particular disease, disorder, or condition, a suitable therapeutic agent may be selected that is suitable for treatment.
[0184] In a further aspect, the invention provides a kit. In certain embodiments, kits are provided that may be used to carry out the methods of the invention as described herein, for example for determining the spatial arrangement of two or more targetable molecules, using the polynucleotides, method steps and / or reagents described herein. The kit may comprise one or more reagents for performing any method as disclosed herein. For example, the kit may comprise at least two seeds, as described herein, such as at least two seeds each comprising a polynucleotide molecule comprising a unique sequence identifier. Alternatively, the kit may comprise the components required to create a suitable seed. Preferably, the polynucleotide molecules are DNA strands.
[0185] In some embodiments, the kit comprises a measurement sub-seed, wherein the measurement sub-seed as described herein, and / or a geometry sub-seed as described herein. The measurement sub-seed and / or geometry sub-seed may comprise a 3' end that is blocked from extension.
[0186] In some embodiments, the kit comprises a gelling agent as described herein. The kit may also comprise an agent for dissolving a gel, such as an alkaline buffer with DTT.
[0187] In some embodiments, the kit comprises a fixation agent, for example a chemical fixation agent. For example, the chemical fixation agent may be selected from the group consisting of: a cross-linking agent, such as formaldehyde, glutaraldehyde and succinimide esters; and a solvent, such as acetone and methanol. In some embodiments, the kit comprises instructions for performing fixation by using a physical fixation method, such as by freezing the sample (e.g. cells or tissues) and air drying.
[0188] In some embodiments, the kit comprises a permeabilisation agent, such as Triton X- 100, Proteinase K, Tween-20; preferably Triton X-100.
[0189] In some embodiments, the kit comprises a buffer. For example, the buffer may be suitable for performing a PCR, such as a Taq buffer (e.g. 10 mM Tris-HCI pH 8.0, 50 mM KCI, and 1.5 mM MgC ).
[0190] In some embodiments, the kit comprises deoxynucleotide triphosphates (dNTPs). There are four types of dNTP, with each using a different DNA base: adenine (dATP), cytosine (dCTP), guanine (dGTP), and thymine (dTTP). Using dNTP during the extension phase provides single bases ready to go into DNA and double it, like building blocks. Preferably, the kit comprises all four types of dNTP.
[0191] In some embodiments, the kit comprises at least one pair of primers, wherein the primers have specificity for a primer site within the seeds. The kit may further comprise exponential primers suitable for amplifying concatemers formed by the methods described herein, and / or PCR primers as described herein. In some embodiments, the kit comprises at least one enzyme. For example, the enzyme may be at least one reverse-transcriptase, at least one polymerase, at least one restriction enzyme, or combinations thereof.
[0192] In some embodiments, the kit comprises instructions for use. In some embodiments, the kit comprises software (e.g. analysis software), such as software for obtaining a readout based on any of the methods described herein. For example, the software may be for determining the spatial arrangement of two or more targetable molecules based on sequencing data.
[0193] In another aspect, the invention provides a method, a use, or a kit, substantially as described herein with reference to the accompanying claims, description, examples and / or figures.
[0194] Preferences and options for a given aspect, feature or parameter of the invention should, unless the context indicates otherwise, be regarded as having been disclosed in combination with any and all preferences and options for all other aspects, features and parameters of the invention.
[0195] Further particular embodiments of the invention include:
[0196] 1. A method for preparing a sample containing one or more cells for the construction of a spatial network using nucleic acids, the method comprising:
[0197] (a) providing a sample containing one or more cells; and
[0198] (b) randomly immobilizing exogenous nucleic acids in the sample, wherein each nucleic acid comprises a unique sequence identifier, wherein each nucleic acid identifies a node in the spatial network.
[0199] 2. The method of embodiment 1, wherein step (b) comprises applying a gel comprising the nucleic acid molecules to a sample that results in immobilization of the nucleic acids at multiple points in the sample.
[0200] 3. The method of embodiment 2, wherein the gel is a hydrogel.
[0201] 4. The method of embodiment 3, wherein PEG-based or acrylamide-based gelling components are used to form the hydrogel. 5. The method of embodiment 4, wherein PEG-based gelling components are acrylate-PEG and / or thio-PEG gelling components and / or wherein the hydrogel is an acrylamide-based hydrogel.
[0202] 6. The method of any one of embodiments 1 to 5 comprising a step of forming polonies from the nucleic acids immobilized in the sample.
[0203] 7. The method of embodiment 6, wherein each polony forms a node in the spatial network.
[0204] 8. The method of any one of embodiments 1 to 7, wherein the nucleic acids are immobilized (e.g. crosslinked) to endogenous biomolecules in the sample.
[0205] 9. The method of embodiment 8, wherein the biomolecule is a protein, nucleic acid or lipid.
[0206] 10. The method of any one of embodiments 2 to 7, wherein the nucleic acids are immobilized to the gel.
[0207] 11. The method of any one of embodiments 1 to 10, wherein the nucleic acids are immobilized via a covalent bond.
[0208] 12. The method of any one of embodiments 1 to 11, wherein the nucleic acids are immobilized via a reaction with an amine or thiol group in the sample or gel.
[0209] 13. The method of any one of embodiments 1 to 12, wherein the nucleic acids are acrydite-modified nucleic acids.
[0210] 14. The method of embodiment 13, wherein acrydite-modified nucleic acids comprise a methacryl group at their 5' ends.
[0211] 15. A sample containing one or more cells for the construction of a spatial network using nucleic acids, the sample comprising randomly immobilized exogenous nucleic acids in the sample, wherein each nucleic acid comprises a unique sequence identifier and identifies a node in the spatial network. 16. The sample of embodiment 15, wherein the nucleic acid molecules a provided in the sample via a gel that results in immobilization of the nucleic acids at multiple points in the sample.
[0212] 17. The sample of embodiment 16, wherein the gel is a hydrogel.
[0213] 18. The sample of embodiment 17, wherein PEG-based or acrylamide-based gelling components are used to form the hydrogel.
[0214] 19. The sample of embodiment 18, wherein PEG-based gelling components are acrylate-PEG and / or thio-PEG gelling components and / or wherein the hydrogel is an acrylamide-based hydrogel.
[0215] 20. The sample of any one of embodiments 15 to 19, wherein the nucleic acids immobilized in the sample are used to form polonies.
[0216] 21. The sample of embodiment 20, wherein each polony forms a node in the spatial network.
[0217] 22. The sample of any one of embodiments 15 to 21, wherein the nucleic acids are immobilized (e.g. crosslinked) to endogenous biomolecules in the sample.
[0218] 23. The sample of embodiment 22, wherein the biomolecule is a protein, nucleic acid or lipid.
[0219] 24. The sample of any one of embodiments 16 to 21, wherein the nucleic acids are immobilized to the gel.
[0220] 25. The sample of any one of embodiments 15 to 24, wherein the nucleic acids are immobilized via a covalent bond.
[0221] 26. The sample of any one of embodiments 15 to 25, wherein the nucleic acids are immobilized via a reaction with an amine or thiol group in the sample or gel.
[0222] 27. The sample of any one of embodiments 15 to 26, wherein the nucleic acids are acrydite-modified nucleic acids. 28. The sample of embodiment 27, wherein acrydite-modified nucleic acids comprise a methacryl group at their 5' ends.
[0223] 29. A system for the construction of a spatial network in a sample comprising cells, the system comprising:
[0224] (A) a plurality nucleic acids, wherein each nucleic acid comprises a unique sequence identifier; and
[0225] (B) means for randomly immobilizing the plurality of nucleic acids in the sample, such each nucleic acid identifies a node in the spatial network.
[0226] 30. The system of embodiment 29, wherein the means for randomly immobilizing the plurality of nucleic acids in the sample comprises a cross-linking agent.
[0227] 31. The system of embodiment 29 or 30, wherein the means for randomly immobilizing the plurality of nucleic acids in the sample comprises a gel that results in immobilization of the nucleic acids at multiple points in the sample.
[0228] 32. The system of embodiment 31, wherein the gel is a hydrogel.
[0229] 33. The system of embodiment 32, wherein PEG-based or acrylamide-based gelling components are used to form the hydrogel.
[0230] 33. The system of embodiment 33, wherein PEG-based gelling components are acrylate-PEG and / or thio-PEG gelling components and / or wherein the hydrogel is an acrylamide-based hydrogel.
[0231] 34. The system of any one of embodiments 29 to 33, wherein the nucleic acids immobilized in the sample are used to form polonies.
[0232] 35. The system of embodiment 34, wherein each polony forms a node in the spatial network.
[0233] 36. The system of any one of embodiments 29 to 35, wherein the nucleic acids can be immobilized (e.g. crosslinked) to endogenous biomolecules in the sample.
[0234] 37. The system of embodiment 36, wherein the biomolecule is a protein, nucleic acid or lipid. 38. The system of any one of embodiments 31 to 35, wherein the nucleic acids can be immobilized to the gel.
[0235] 39. The system of any one of embodiments 29 to 38, wherein the nucleic acids can be immobilized via a covalent bond.
[0236] 40. The system of any one of embodiments 29 to 39, wherein the nucleic acids can be immobilized via a reaction with an amine or thiol group in the sample or gel.
[0237] 41. The system of any one of embodiments 29 to 40, wherein the nucleic acids are acrydite-modified nucleic acids.
[0238] 42. The system of embodiment 41, wherein acrydite-modified nucleic acids comprise a methacryl group at their 5' ends.
[0239] 43. Use of the system of any one of embodiments 29 to 42 for the construction of a spatial network in a sample comprising cells.
[0240] 44. A method for constructing of a spatial network in a sample containing one or more cells using nucleic acids, the method comprising:
[0241] (a) preparing a sample containing cells according to any one of embodiments 1 to 14 or providing a sample according to any one of embodiments 15 to 28;
[0242] (b) forming nucleic acid nodes (e.g. polonies) from the nucleic acids immobilized in the sample; and
[0243] (c) deducing the location of the nucleic acid nodes based on interactions between nucleic acids in adjacent nodes thereby constructing a spatial network.
[0244] 45. The method of embodiment 44, wherein interactions between nucleic acids in adjacent nodes involves the formation of nucleic acids (e.g. concatemers) comprising unique sequence identifiers from adjacent nodes.
[0245] 46. The method of embodiment 45, wherein deducing the location of the nucleic acid nodes determining which node identifier sequences have been combined in the nucleic acids formed by the interactions between nucleic acids in adjacent nodes.
[0246] 47. A method for mapping transcriptomic and / or proteomic information in a sample containing one or more cells, comprising: (a) constructing a spatial network in the sample using nucleic acids according to any one of embodiments 44 to 46;
[0247] (b) forming nucleic acids (e.g. concatemers) from interactions between nucleic acid nodes and target nucleic acids in the sample that contain transcriptomic or proteomic information; and
[0248] (c) using the nucleic acids (e.g. concatemers) formed in (b) to generate a map of transcriptomic and / or proteomic information in the sample.
[0249] 48. The method of embodiment 47, wherein the target nucleic acids are selected from the group consisting of: mRNA, cDNA, and DNA.
[0250] 49. The method of embodiment 47, wherein the target nucleic acids are polynucleotides conjugated to an antibody or antigen-binding fragment thereof, wherein the antibody or antigen-binding fragment thereof is bound to a biomolecule.
[0251] 50. The method of embodiment 47, wherein the target nucleic acids are endogenous to the sample.
[0252] FIGURES
[0253] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying figures, in which:
[0254] Figure 1: Overview of the general method of DNA microscopy. A) Cells are fixed, and RIMA is reverse transcribed. The cells are then permeabilized with a gel containing PCR buffer, barcoded seed strands, primers, and activators. The seed strands will form polonies that can interact with each other and with transcripts within the cells. B) These interactions form products that either contain the polony barcodes of two polonies, or the product contains a single polony barcode and transcriptomic information. These products are sequenced and used to create a map of where each polony is located and what is located within them.
[0255] Figure 2: Overview of the PCR reaction that takes place in this method. A) Barcoded seed strands in a PEG hydrogel are amplified by PCR, forming polonies. During thermocycling these polonies grow and start overlapping. The copies of seed strands then get activated to allow them to participate in one of two reactions. Approximately half of these activated monomers participate in the geometry reaction, showing neighbouring polonies, which is used to create the spatial network. The other half of the activated monomers will participate in the measurement reaction that appends polony barcodes to transcripts, which shows what transcripts are present in the polony, allowing the spatial network to be populated with transcriptomic information. NGS = Next Generation Sequencing.
[0256] Figure 3: Verification of the DNA microscopy PCR reaction. A) 10% UREA PAGE gel for reactions obtained from DNA microscopy with target. By adding the corresponding activator strands, products corresponding to the respective reactions appear, and the two expected sizes appear in the full reaction. B) Results of library preparation of DNA microscopy samples. The final library produced was 408 bp in length and contained the product of the geometry layer.
[0257] Figure 4: Reconstruction of space with artificial targets with a MiniPath cutoff of 10. A) Experimental setup. In this experiment, four droplets were placed on a glass slide, each containing 1 pM seed strand and 1 pM target (1, 2, 3, or 4), and all components required for PCR. The droplets were placed in a line from positions 1 to 4. This gel was then encased in a 3pL gel containing 1 pM seed and PCR components, but no target strands. Once gelled, this was then encased in 40 pL of gel containing the PCR components but no seed strands or targets. This gel prevented the depletion of PCR components and evaporation. The signals of these four targets in reconstruction are shown in B-D. What can be seen, is that each target seems to inhabit a distinct location in the reconstruction and follows the order of 1 (B), 2(C), 3(D), and 4 (E). For B-E, the x-axis and the y-axis = -15, -10, -5, 0, 5, 10, 15. For B, the shaded bar scale is 10°, 101, 102, 103. For C-E, the shaded bar scale is 10°, 101, 102, 103, 104.
[0258] Figure 5: Reconstruction of four separate samples, each containing four targets placed in a line. A-D Targets 1 - 4 from sample 1, with a MiniPath cutoff of 1. E-H Targets 1- 4 from sample 2, MiniPath cutoff of 10. I-L Targets 1-4 from sample 3, Minipath cutoff of 100. M-P Targets 1-4 from sample 4, with no MiniPath filtering. In all cases, the four targets were mostly present in different regions. Furthermore, the targets follow a logical order, meaning that target 1 is followed by targets 2, 3, and 4. For A-D and I- L, the x-axis and the y-axis = -10, -5, 0, 5, 10. For E-H, the x-axis and the y-axis = -20, -15, -10, -5, 0, 5, 10, 15, 20. For M-P, the x-axis and the y-axis = -15, -10, -5, 0, 5, 10, 15. For B, the shaded bar scale is 10°, 101, 102, 103. For A-P, the shaded bar scale is 10°, 101, 102, 103, 104.
[0259] Figure 6: Detailed scheme of the polony formation stage of the seed strands. A) The DNA strands used in this scheme, with the name as appears in Table 1. B) Both seed strand types create copies of the seed strand using their respective primers. C) Seed copies are amplified using exponential PCR. During amplification, they gain a new sequence at the 3' end of the molecule, which is later used for activation. D) The top strand of each of these reactions (the inactivate monomers) can be activated for either geometry (Figure 7) or measurement (Figure 9) reactions, which happens at random for each strand.
[0260] Figure 7: Detailed scheme on the activation and amplification of the geometry reaction. A) The names of the strands that appear in this reaction are shown in Table 1. B) Inactive monomers from the amplification reactions (Figure 6) that land on geometry activator strands, are extended with a geometry 3' sequence. C) This 3' sequence is complementary between red and yellow types. When two activated geometry monomers of different seed types come in proximity, they can anneal to their 3' ends and extend, creating a double stranded geometry product that now contains the polony barcodes of both the red and yellow monomers. D) As these sequences end in their original primers, they can be exponentially amplified in subsequent cycles of PCR. Figure 8: Scheme for the target strands. A) The names of the sequences used in this scheme as they are shown in Table 1. B) During the polony formation stage of the reaction, target strands behave in a similar way to seed strands. Primers land on the target strands and create copies. C) Copies of target strands are amplified by exponential PCR and receive a new 3' end sequence that allows them to be activated. D) Unlike seed strands, targets do not participate in the geometry reaction, instead they can only be activated for the measurement reaction, thus creating only one type of activated monomer. E) The activated target monomers can only participate in the measurement reaction and will thus only interact with measurement activated seed monomers.
[0261] Figure 9: Detailed scheme on the measurement reaction for both seed and target monomers. A) The names of the sequences used in this scheme as they are shown in Table 1. B) Similar to the geometry reaction, inactive seed monomers (Figure 6) can land on a measurement activator strand and will receive a 3' end measurement sequence. C) The sequence on the seed monomers is identical between the two seed types but is complementary to the sequence received by the target monomer (Figure 8). When a measurement monomer of either seed strand type finds a target monomer, they anneal on their 3' end and extend, forming a double stranded measurement product that contains the polony barcode on one side, and the target information on the other. D) This double-stranded product ends with the original primers and starts exponentially amplifying in the next cycle.
[0262] Figure 10: Gradient PCR of the reaction without targets. The geometry reaction band intensity starts decreasing between 66.4 °C and 69.1 °C. Therefore, an annealing temperature of 67 °C was used.
[0263] Figure 11: Effective polony size estimation using spacing. In this setup, gels are cast as a single gel containing a uniform distribution of seed strands (A) or are cast as a small gel with uniform distribution, surrounded by a gel without seed (B). The difference between these samples is that, while the total amount of molecules and volumes remained the same, seed strands in the samples with smaller gels had a small cluster of seed strands that were relatively closer to one another. C) Three experimental setups were tested, 30 000, 300 000 and 3 000 000 total amount of seed strands. The total number of seed was kept the same between pairs, but the distance between the polonies was changed by clustering them in the smaller gel. D-E) Seed strands with 68 pm distance produced a library that did not have clear peaks (D). In contrast, when the distance decreased to 32 pm, the library pattern changed and started to produce clearer peaks. F-I) Distances lower than 32 pm maintain clear peaks. However, higher concentrations also resulted in an increased peak signal. For D, x-axis = 0, 50, 100, 150, 200 FU; for E, x-axis = 0, 50, 100, 150 FU; for F, x-axis = 0, 500, 1000 FU; for G, x-axis = 0, 500, 1000 FU; for H, x-axis = 0, 500, 1000, 1500 FU; for I, x-axis = 0, 500, 1000, 1500 FU. For D-I, y-axis = 35, 100, 200, 300, 400, 600, 2000, 10380 [bp].
[0264] Figure 12: Activation of strands was compared between blocking the 3' end on the activator with an inverted T (A) and leaving the 3' end unblocked (B). The gel shows that the reaction creates many side products if the 3' end is not blocked.
[0265] Figure 13: Purification of the final DNA microscopy library creates more useful product. Shown is the fragment size estimated from paired end reads with a percentage of full products, known side products, or other unknown side products. A) DNA microscopy that is prepared into a library, but not ran through a Pippin™ Prep produces a library with side reactions. Only a small fraction of products (7%) is used in reconstruction and network population, while the rest is not used in reconstruction and is discarded. A large portion of the side products, such as incomplete monomers (13%), are precursors to the final product. The remaining sequences obtained from this library did not fall into a category of known products and could be PCR or ligation artifacts of either library prep or the DNA microscopy reaction. B) The library after library prep produces a larger band of full products (27.3% of reads), suggesting that the Pippin™ Prep reduces side products and increases sequencing yield.
[0266] Figure 14: Alternative methods for initial seed copies production in situ, (a) Spooling of copies in a linear fashion via thermal cycling, (b) PCR or asymmetric PCR-like where an additional primer pair is added at a lower concentration (necessary to add adapters for the 'lower concentration' primer regions on the sub-seed strands), (c) Using rolling circle amplification and digestion to generate a pool of seed copies from a circle containing the seeds.
[0267] Figure 15: An alternative method for creating marker monomers.
[0268] Figure 16: A further alternative method for creating marker monomers.
[0269] Figure 17: Getting an image of RNA using sequencing alone - A simplified sketch, (a) By fixing tissue, the mRNA will remain localized at their original positions. Conversion to cDNA and in-situ amplification creates 'clouds' of amplicons inside the tissue, (b), around the original RIMA sites. By using barcoding and seeded oligos, and allowing amplicons to overlap extend by polymerase, (c), one creates fused amplicons that contain information about which mRNA are neighbours. By harvesting and sequencing all these fused amplicons, one can mathematically reconstruct a geometrical network of all neighbours, (d), and consequently provide an image, (e), of what RNA was sequenced, and where the amplicons were located in the tissue. Key advantages compared with competing strategies involving microscopy (f), this technique defer most of the complicated experiments to computer work (g).
[0270] Figure 18: Schematic overview of how seed polynucleotides define geometry.
[0271] Figure 19: The measurement reaction and its connection to the geometry layer, (a) The seeds send out two types of monomers depending on if the amplicons extend on sub-seeds that are either: geometry associated (top, conferring the grey adapter) or measurement associated (bottom, conferring the orange adapter), (b) By using oligo- T and a template switching ribo-G oligo, adapters for amplification of cDNA in situ are created. These 'measurement monomers' will then obtain an adapter (in orange) at the site corresponding to the 5' end of the original transcript. These can then concatenate with measurement monomers originating from the seeds, like the ones in (a)(bottom) - forming the measurement reaction. Schematically: (c) Each seed can create two types of sub-seed polonies. This makes it possible to create two types of concatemers, (d). One type connects seeds to mRNA reads (measurement layer), the other connects seeds to other seeds (geometry layer). By increasing the concentration of measurement associated sub-seeds - the reaction will favour measurements. By increasing the concentration of geometry associated sub-seeds - the reaction will favour spatial resolution.
[0272] Figure 20: Positioning and merging of the layers. The diffusion of amplicons from unknown transcripts will be slow due to the large size, and vary depending on the exact lengths of the unknown cDNA. By using a two-layer strategy, one can create a very well-defined set of seed to seed amplicons, (a). One computational challenge is to rapidly convert the raw graph from the sequencing data, into a laid out graph that define geometry. Force directed graph layout using the fusion event barcodes as proximity indicators works. Notably, the two layer approach allows for tuning the concentration of the sub-seed strand sets (measurement- vs. geometry-associated) so that the reaction can be tuned for more geometrical precision, or for higher read depth of transcripts. Something that is unique in the field of imaging by sequencing and spatial transcriptomics in general, (b) At the same time, unknown transcripts are allowed to fuse to the seed data. This simultaneously provides a well-defined geometry as well as genome-wide transcriptomics when unknown amplicons are anchored to the correct geometry by their fusions to well defined seed locations.
[0273] Figure 21: (a) Data from earlier computational work showing mathematical reconstruction (right) of image data (left) using simulated sequence information alone. Here we showed that topological distance (as resulting from graphs from interconnected sequence reads), is a good approximation for Euclidian distance in large graphs, (b) Simulating TLDM, a simulated distribution of seed-nodes (Original, left) is used to reconstruct a simulated network with 5 (middle) or 10 (right) connections between each node (seed) on average. This illustrates the need for a minimum number of seed-seed connections for adequate reconstructions, (c) TLDM experiment in mouse brain and reconstructing a geometry network was seeded in a mouse brain tissue sample (photo to the left) TLDM performed and followed by Illumina sequencing. The reconstruction of 90k nodes (seeds) in total is shown in two views with a 90 degree clockwise rotation between top and bottom views. The field-of views show the layout of nodes, and the zoom-ins (right) shows the node types (red- or yellow- primer associated) and connections between them corresponding to sequenced geometry concatemers. Loose nodes shown without edges only have connections outside the zoomed view. Initial testing of the measurement layer together with the geometry layer, (d) PAGE gel showing DNA from TLDM reactions inside a gel matrix. All lanes are from reactions containing all seeds, sub-seeds and mock cDNA targets. Lane 1: Only primers. Lane 2: With the geometry-associated sub-seeds. Lane 3: With the measurement-associated sub-seeds. Lane 4: With both geometry- and measurement sub-seeds (Full reaction). The full reaction produces both geometry concatemers ('G') and measurement concatemers ('M') simultaneously. 'Mon' is monomers, 'SubS' is sub-seeds and 'Pr' primers. The results were verified by sequencing, (e) Test of the TLDM protocol in cell culture showing the read distribution of measurement concatemers vs geometry concatemers after sequencing.
[0274] Figure 22: Schematic of Two Layer DNA Microscopy. EXAMPLES
[0275] Example 1
[0276] The topological organisation of cells in tissues is important for their functions. Defining the location of cells and their transcriptomic profiles is therefore paramount for understanding the physiological function and pathological dysregulation of these tissues. This study uses DNA microscopy, a PCR-based method that allows for the reconstruction of spatial information using sequencing without the need for a priori characterization or optics, which was achieved by creating polymerase colonies (polonies), each with a unique polony barcode. These polonies diffuse and interact with neighbouring polonies and transcripts. Once sequenced, neighbour information is used to reconstruct the relative position of the polonies and create a spatial map of where each polony is located. The transcripts also contain these polony barcodes, allowing them to be added to the nodes in the network. In this work, the design and verification of PCR and the reconstruction of an artificially created pattern of targets is presented.
[0277] Results and Discussion
[0278] Design of the DNA microscopy reaction. This method is based on using polonies to reconstruct spatial information using PCR without a priori characterization. This is done by using polonies in a PEG-based hydrogel. The gel contains DNA strands, PCR buffer, dNTPs, and enzymes to facilitate the DNA microscopy reaction. The PCR in the diffusion-restrictive gel will form polonies, which grow throughout the PCR program. Once the polonies are large enough and overlap, the strands in the polonies will create products containing information on which polonies are in proximity and which transcripts are within the area of each polony (Figure 1A). These products are sequenced and are used to create a spatial network and populate it with transcriptome information (Figure IB).
[0279] The polonies in this system are formed from uniquely barcoded seed strands, singlestranded DNA oligonucleotides that are immobilized in the gel and act as a point of origin for a polony. There are two types of seeds: a first type (e.g. the red type) and a second type (e.g. the yellow type), which differ in sequence, but not in function, meaning they work identically but do so independently. The PCR starts with a polony formation step (Figure 2A, Figure 6), in which the polonies are formed with copies of their respective seed strands (monomers). To add information to the reconstruction, two types of information are needed: the spatial layout of the polonies, and the transcriptomic content of these polonies. This is facilitated by two reactions occurring in parallel: the geometry reaction to create the spatial network (Figure 2B) and the measurement reaction to append transcriptomic information to this network (Figure 2C).
[0280] At this point, the monomers cannot interact with any other monomers and do not participate in any geometry or measurement reactions yet. This changes after the activation step, in which monomers bind to activator strands, adding an overlap sequence to the 3' end of the monomer, activating it and allowing it to participate in one of the two reactions (Figure 7, Figure 9). Two activator strand types are used per polony type: one that adds a sequence for the geometry reaction (Figure 7B) and one that adds a sequence for the measurement reaction (Figure 9B). As these activator strands are added at equimolar concentrations, roughly half of the monomers should be activated to participate in the geometry reaction and half activated for participation in the measurement reaction. Once the monomers are activated, PCR continues in the amplification step, in which the geometry and measurement reactions take place.
[0281] Monomers that are activated by the geometry activator strand receive a sequence that is complementary to the sequence received by the opposite type of seed strand (that is, red geometry monomers and yellow geometry monomers have complementary 3' ends) (Figure 2B, Figure 7B). This allows geometry-activated red monomers to interact with geometry-activated yellow monomers. When these two monomers meet, they anneal and extend one another, creating a single double-stranded geometry product that contains both yellow and red polony barcodes (Figure 7C). The barcode combinations from this reaction can be used to identify the neighbourhoods of polonies and are used for the construction of a spatial network.
[0282] Monomers that land on the measurement activator strand will receive a sequence that is identical between the two types (i.e. both red and yellow measurement monomers will have the same 3' end and do not interact with one another) (Figure 9B). These monomers will instead interact with target monomers (Figure 9C). In this setup, a target is defined as a DNA strand that behaves like a seed strand, but only participates in the measurement reaction, populating the spatial map with information, while not involved in the creation of the map itself. These strands can be used to flag places of interest, for example by transforming cDNA into a target strand, the location of a transcript can be determined. Target strands will form a polony together with seed strands, but will only be activated by a measurement activator, which grants these target monomers a 3' sequence complementary to the measurement sequence of the seed strands (i.e. both red and yellow measurement monomers will interact with target monomers) (Figure 8). A monomer of either type and a target can anneal and extend each other, creating a double-stranded measurement product that contains target (transcriptomic) information on one side and the polony barcode on the other (Figure 9C). Measurement products represent the content of the polonies and are used to append transcriptomic information to the constructed geometry network.
[0283] The geometry and measurement products are also amplified in this step using exponential PCR (Figure 7D, Figure 9D). Once the PCR is complete, the gel is dissolved, the geometry and measurement products are isolated using AMPure DNA purification, turned into a sequencing library, and sequenced (Figure 2D).
[0284] Verification of the PCR scheme. As the setup relies on seed strands in a PEG-based hydrogel, the design could be verified in vitro without cells (Figure 3). Furthermore, because the geometry and measurement reactions are activated by different activator strands, it is possible to selectively remove one of the two reactions by removing the associated activator strand from the mix. To test the measurement reaction, a target strand was made from Gfp cDNA to mimic cDNA in a cellular setting. The target was designed to be of a size that would be slightly larger than seed strands, forming a measurement product that is larger than geometry product to be able to distinguish them on a PAGE gel. The number of seed strands in this gel was assessed to work best in the 100 fM - 1 pM range (Figure 11), and the concentration of 1 pM was used throughout this study. We also show that activator strands require to have a 3' end blocked from extension, to generate the cleanest product (Figure 12).
[0285] A DNA microscopy reaction was set up in a gel without cells, containing one of the following activator mixes: no activators, only geometry activators, only measurement activators, or all activators. A gel was cast for each sample, containing PCR components (dNTPs, buffer, and enzyme), seed and target strands, primers, and their respective mix of activator strands. These gels were then subjected to a PCR program that was previously optimized for annealing temperature (Figure 10). After PCR, the gel was dissolved, and the solution was placed on a PAGE gel for verification (Figure 3A). To verify that the reaction could be prepared for sequencing, a separate DNA microscopy reaction was set up containing all activators, but no targets. The resulting product was isolated from the gel and was subjected to library preparation. The final library was further purified to remove intermediate products from the reaction (Figure 13) and run on a Bioanalyzer for verification (Figure 3B). The Bioanalyzer produced one main band of ~408 base pairs (bp), which corresponds to the size of the geometry product with Illumina adapters. This shows that the resulting PCR reaction can be successfully prepared into a relatively pure sequencing library.
[0286] Reconstruction of space with an artificially patterned target. To test the ability of DNA microscopy reconstructing a spatial layout, a fully artificial setup was constructed, as shown in Figure 4A, to create a pattern of targets strands. Four different target strands were designed, each containing a stretch of DNA that could correspond to the identity of each respective target. The measurement product generated with these targets strands was designed to be of the same size as the geometry product for ease of purification. Small gels (~0.25 pL) were cast in a line, each of which contained the full DNA microscopy mix (see Materials and Methods in Example 3), including 1 pM of seed strands and 1 pM of target strands. Each droplet contained a different target, ordered from target 1 to target 4. The droplets would be placed in a line, but distant enough from each other that they would not touch. Once all gel droplets were placed and gelled, the area was fully encased in a DNA microscopy gel (3 pL) containing the full DNA microscopy mix, including 1 pM of seed strands, but no targets. As such, the reaction reconstructs the full space of this 3 pL gel and the four droplets, but the target signal should only be seen in the droplets. This gel was then further encased in a gel with DNA microscopy mix, but excluding any seed strands or targets, to prevent evaporation.
[0287] After sequencing, all space containing seed strands was reconstructed using the previously described ptMLE method
[0018] , the graph was then further filtered using MiniPath
[0021] , and the presence of the targets was mapped onto the reconstruction. As shown in Figure 4B-E, each target inhabits a different area in the reconstructed space. Furthermore, the order of targets is the same as cast on the slide (1-2-3-4). It was noted that while the content of the reconstruction does follow the order of targets, the shape of the reconstruction is not fully representative of the original shape of the gel. One potential source of error could be false connections that warp the reconstruction. It was shown that the data herein has few events of false connections due to library preparation, but did have a small amount of contamination of products that likely originate from different sequencing libraries (Table 2). Contamination of different libraries is unlikely to cause warping, as these will likely contain barcodes that are otherwise not present in the reconstruction, reducing the chance of false connections being formed between existing nodes in the network. This data was also subject to connection filtering on a graph level
[0021] , which would reduce the number of false connections as well, though the sequencing depth might not be deep enough to fully remove all false connections. A larger sequencing depth and sequencing without other DNA microscopy libraries might reduce this warping effect.
[0288] To assess the degree of error and reproducibility of the technique, four additional samples were prepared following the same setup and reconstruction. The MiniPath cutoff values were assessed per sample to produce a reconstruction that was the least warped. The results of these filtered reconstructions are shown in Figure 5. These four reconstructions each show the same result as before, where all targets follow the same order as they were placed in. From these reconstructions, there does not appear to be a consistent error in the reconstructions. For example, target 1 was not always well represented (Figure 4A and Figure 5E), but then was present in other samples, suggesting that some of the error is not systemic, but more sample related. Overall, these results show that the method can reproducibly reconstruct the spatial arrangement of DNA strands in the DNA microscopy samples.
[0289] Conclusion. In this study, a novel approach to spatial transcriptomics is introduced that utilises a PCR-based technique to reconstruct spatial arrangements in conjunction with expression data. This method eliminates both the need for preliminary slide characterisation and the use of optics. By generating polymerase colonies with unique barcodes, we can map these polonies spatially in relation to one another. This research presents a specialised PCR protocol designed to minimise side reactions and optimise product yield. This protocol can reconstruct the spatial configuration of artificial targets and this outcome can be replicated. The reconstructions can produce a slightly warped reconstruction, which might suggest that some erroneous connections are present in the network, but these connections do not impact the overall neighbour information of the polonies. Based on these results, it is plausible that the method can be applied in biological samples, where the target strand is designed to contain transcriptomics or potentially proteomic information of the sample.
[0290] Example 2
[0291] This Example provides a more detailed explanation of the DNA microscopy reaction as performed in this work.
[0292] Seed strands and target strands are present in and attached to the gel. In the first ten cycles of the reaction, during the polony formation stage, copies of the seed strands (Figure 6B) and copies of target strands (Figure 8B) are produced. These copies are exponentially amplified and given a 3' end that allows them to be activated for the geometry and measurement reactions (Figure 6C, Figure 8C), making them inactive monomers (i.e. monomers that can be, but are not yet activated).
[0293] This is followed by the activation stage of the reaction, in which inactive monomers are activated. Inactive seed monomers can be activated for the geometry (Figure 7B) or measurement (Figure 9B) reactions using activator strands. These activators add a sequence to the 3' end of the now activated monomers, respective to the activator. As all activator strands are added in equal amount, an inactive seed monomer has a 50% chance to become a geometry or 50% chance to become a measurement monomer. However, targets can only be activated as measurement monomers (Figure 8D) as they do not participate in the geometry reaction.
[0294] During the activation stage, geometry monomers gain a 3' end sequence that is complementary to the 3' end sequence of geometry monomers of the other seed type (i.e. red geometry monomers interact with yellow geometry monomers). When geometry monomers come in proximity during the amplification stage, they hybridise their 3' ends and extend each other (Figure 7C). This results in a double stranded geometry product that contains the polony barcodes of both monomers and can be further amplified in the reaction (Figure 7D). The geometry product is used to reconstruct the network of the space the seed strands inhabited.
[0295] Measurement monomers from seed strands gain a 3' end sequence that is identical between the two seed types (Figure 9B). The activated target monomers are only activated as measurement monomers and will gain a 3' end sequence that is complementary to measurement monomers of both seed types (Figure 8D). Once a target monomer and either type of seed measurement monomer meets, they will hybridize and extend each other (Figure 9C). This creates a measurement product that will on one side contain the polony barcode, and on the other contain information about the target. This product is amplified in the reaction (Figure 9D). The measurement product is used to populate the network, by appending the target information to the nodes of the network that correspond to the polony barcode it was found with. primers (PR), exponential primers (EX), activator strands (OE), and artificial targets (TR). To obtain the annealing temperature for the reactions, a gradient PCR was performed (Figure 10) and the highest temperature was taken that did not affect reaction efficiency.
[0296] The size of the polonies can correspond to the resolution obtained by DNA microscopy. This was investigated with an approach in which smaller gels containing seed strands were cast in gels that had PCR components, but no seed strands. This setup had two types of samples : a sample where the seed strands were uniformly distributed in a single gel (Figure 11A) or a sample where the seed strands are placed in a smaller gel, which is surrounded by a second larger gel without seed strands (Figure 11B). This approach will keep the total amount of molecules and total volume of the gel the same, but still affects the spacing of the seed strands. As illustrated in Figure 11C, three conditions were tested: 30,000 total seed strands, 300,000 total seed strands, and 3,000,000 total seed strands either spaced uniformly or clustered in l / 10thof the volume. The diameter of the polonies was calculated by dividing the volume of the gel with seed strands by the total number of seed strands, and calculating the diameter of that volume as a sphere. When looking at the uniformly spaced 30,000 strands sample (Figure 11D), the resulting library looks much like a bell curve with no clear peaks. However, if the seed strands are spaced lOx closer, and the diameter of the polonies is decreased ~2 fold, the resulting library starts showing clearer peaks with less background noise (Figure HE). If this is then compared with the sample where spacing is the same, but the total number of molecules is increased to 300,000 (Figure 11F), the clear peaks and overall profile remains the same, but the peaks themselves have a stronger signal.
[0297] In the remaining libraries (Figure 11G-I), the clear peaks remain the same, changing mainly in intensity as the total number of seed strands increases. This suggests that polonies between 68 pm and 32 pm in diameter can interact sufficiently to create clear DNA microscopy libraries. However, this diameter does not correspond to the actual diameter of the polony, as they must overlap to produce the product. Rather, the calculated diameter can be interpreted as the maximum diameter of a polony before the reaction itself is affected. It is still possible to sequence the libraries even with a diameter of 68 pm; however, the resulting sequencing will have more noise and side reactions.
[0298] The initial PCR was also optimized to produce as few side products as possible. To this end, a 3' modification was added to the activator strands to prevent them from extending. This blocking reaction shows that allowing the activator to not be extended reduces smearing and reduces the amount of side products (Figure 12).
[0299] The DNA Microscopy reaction relies on only two types of products: the geometry product, and the measurement product. However, as this reaction relies on production and activation of strands, precursors to these products will also be present in the mix, seen in the bioanalyzer results of Figure 11. Some of these products are double stranded and will also be prepared in a library and will be sequenced. As most products are incomplete precursors, they are smaller in size than the main products used for reconstruction.
[0300] To remove these side products, an automated gel extraction was performed using a Pippin™ Prep instrument to create a tight range surrounding the peak of interest. The data were sequenced, and the insert size of each read was calculated using paired-end reads. An insert size of 250 bp is expected in this setup for both measurement and geometry product. The library without purification (Figure 13A) produced relatively many side products and relatively little product of interest. Several of these products were identified as known side products and incomplete precursors. However, much of the data is made up of various side reactions, not part of the main reaction. When purification is performed through Pippin™ Prep (Figure 13B), both known and unknown side products decrease, and the main products increase approximately 4-fold. Notably, there are still many side reactions that appear around the area of interest, and further purification might be required to increase the yield.
[0301] Combination Combination Combination End Adapter Library 1 / 2 only 3 / 4 only 1 / 2 / 3 / 4 Repair Ligation amplification
[0302] Table 2: Barcode swapping assessment. In this experiment, DNA microscopy was performed on an empty gel containing seed with known barcodes. To assess the amount of barcode swapping, two samples with two different barcodes (1 / 2 and 3 / 4) were mixed in different stages of library preparation. This would form combinations of barcodes that would not be present in the original sample (1 / 4 and 3 / 2), but instead formed because of the library prep. Shown in this table are sequenced combinations compared to the percentage of barcode combinations found. Marked in bold are combinations of barcodes that should not be present in the library and could indicate barcode or index swapping. Barcode combinations 1 / 2 and 3 / 4 were sequenced once separately, forming libraries that only should contain those respective combinations. The "Combination 1 / 2 / 3 / 4" sample contained all seed strands, forming all possible barcode combinations. In the remaining samples, two separate barcode combinations were mixed at different steps of the library preparation: End repair, Adapter ligation, and library amplification. Indicated are also the total percentage of expected and unexpected combinations. As the products in this setup are generated through monomers interacting and creating products through PCR, it is important to check to what extent false data can be generated after DNA microscopy during library preparation. One possible source is barcode swapping, in which incomplete PCR products or monomers interact with other monomers during the library prep and create new full products. If this happens, the resulting product will be identical to a main product generated in DNA microscopy, but the barcode combination will have no spatial significance, contributing only to false connections that distort the reconstruction.
[0303] To test if barcode swapping occurs, DNA microscopy reactions using known barcodes were set up. Reactions were set up to produce products with known combinations, in this case Combination 1 / 2 and 3 / 4. At different stages of the library preparation, these samples would be mixed, and the resulting library would be sequenced, to check if barcode swapping is more likely to occur in one of these steps. Combinations between these samples (i.e. combination 1 / 4 or 3 / 2) would indicate that monomers or incomplete products interacted and created false data. The results of this experiment are shown in Table 2. As controls, the pure combinations (1 / 2 and 3 / 4) and a sample containing the 1 / 4 and 3 / 2 combinations natively were sequenced as well.
[0304] No clear trend is seen in the number of barcode swaps when mixing at different stages of library preparation, all containing a similar ratio of correct to incorrect products. A notable observation is that even the pure libraries (libraries containing only 1 / 2 or 3 / 4 that were never mixed during library prep) sequenced barcodes that should not be present (e.g. finding products with barcode 1 in a sample containing only barcodes 3 and 4). This could indicate that there is swapping happening on the sequencing chip itself, or that other libraries are added to this sample by sequencing error or index hopping. It was also observed that most of the incorrect products fell into the "Other" category, which are sequences that are found to be geometry products, but do not contain barcodes from this experiment. These may have originated from libraries sequenced in parallel to these samples on the chip. If this limitation originates from the sequencer running several libraries, it is possible error can be reduced by either switching to a sequencer that does not rely on bridge PCR (for example MGI), as the spots on the flow cell will not run the risk of exchanging information.
[0305] Alternatively, it could also be possible to run samples on larger Illumina sequencers (e.g. Novaseq) in tandem with other libraries that are not DNA microscopy. As then many other libraries are spiked in and interaction between different libraries of DNA microscopy is less likely and incorrect libraries are easier identified and removed. However, while some errors were found here, most of the errors were completely different products, not containing any of the barcodes from this experiment and likely originate from another library. This might not contribute much to false connections, as these barcodes from other libraries are likely not represented in the network and would therefore not create false connections between existing nodes.
[0306] Table 3: Seed strand sequences used in the Barcode Swapping experiment.
[0307] Example 3
[0308] Materials and Methods
[0309] Oligo design and preparation. Sequences for the PCR scheme of the DNA microscopy reactions were designed using NUPACK
[0022] and ordered from Integrated DNA Technologies (IDT). Annealing temperatures were verified using gradient PCR and a 10% 8M UREA PAGE gel.
[0310] Gelling and PCR. Gelling was performed by creating two mixes that, when added together, formed a 12.5% PEG hydrogel in lx Taq buffer (10 mM Tris-HCI pH 8.0, 50 mM KCI, and 1.5 mM MgC ). Mix 1 was a 2x mix containing 2x Taq buffer and 192 pg / pL 4arm-PEG-Acrylate, MW 10,000 (Broad Pharm, BP-25163). Mix 2 was a 4x mix containing 2 mg / mL BSA (NEB, B9000S), 32% Glycerol, and 116.46 pg / pL Thiol-PEG- Thiol, MW 1,500 (Sigma, JKA4105). The final gel had a concentration of lx Taq buffer, 96 pg / pL 4arm-PEG-Acrylate, MW 10,000, 29.12 pg / pL Thiol-PEG-Thiol, MW 1,500, 0.5 mg / mL BSA, 8% glycerol. The gel would also be supplemented with 400 nM of each activator strand (OE_Geo_R_15, OE_Mea_R_8, OE_Geo_Y_15, OE_Mea_Y_8, and OE_Mea_T_8 ), 30 nM of each exponential primer (EX_Red, EX_Yel, and EX_Gap), 300 nM of each main primer (PR_Red_5N, PR_Yel_5N, and PR_Tar_5N), 0.5 pM of each seed strand (SD_N30R and SD_N30Y), 1 mM dNTPs, and 0.5 U KAPA HIFI polymerase (Roche, KK2502). Unless stated otherwise, 20 pL of reaction volume was used in PCR tubes. The samples were placed in a Gene Explorer™ 48 well Dual Block Thermal Cycler (Techtum, 48-GE-48DS). Samples were subjected to a 1 h 22 °C step to gel, an initial heating step of 95 °C for 1 min, initial amplification of lOx (98 °C for 10 s, 67 °C for 30 s, 72 °C for 30 s), activation of 2x (98 °C for 10 s, 50 °C for 30 s, 72 °C for 30 s), amplification of 23x (98 °C for 10 s, 67 °C for 30 s, 72 °C for 30 s), and a final step of 72 °C for 2 min and a final temperature of 4 °C.
[0311] Target Containing gel. For the target experiments, a slightly modified version of gelling was used. Four gels were prepared, all containing the same mix as described above but supplemented with 1 pM of a target (TR_ArT_l, TR_ArT_2, TR_ArT_3, or TR_ArT_4). A droplet of each gel (~ 0.25 pL) was placed on a plain microscopy slide (Sigma, S8902) in a row. Each gel was placed on glass and left to incubate for 15 min in a humidified chamber before the next droplet was placed. Droplets were placed such that they could not touch. When all droplets were placed on the glass, a mixture of 3 pL gel mix containing seed strands but no targets (same as described above) was added to fully encapsulate the four droplets. The gel was left for 15 min in a humidified chamber. After this, a hybridisation chamber was placed over the gel (Sigma, GBL621502), and 40 pL of a new gel containing all components, but no targets or seed strands, was added to the slide. The slide was then placed on a GeneTouch Thermal Cycler (Techtum, 48-TC-EA) with an In Situ plate block (Techtum, 48-B-41A). Samples underwent a PCR program of 1 h 22 °C step to gel, an initial heating step of 95 °C for 1 min, initial amplification of 10 x (98 °C 10 s, 67 °C 30 s, 72 °C 30 s), activation of 2 x (98 °C 10 s, 50 °C 30 s, 72 °C 30 s), amplification of 23 x (98 °C 10 s, 67 °C 30 s, 72 °C 30 s), and a final step of 72 °C for 2 min, and a final temperature of 4 °C.
[0312] Strand Purification and Library Preparation. After completion of the PCR reaction, dissolution buffer (100 mM EDTA, 460 mM KOH, and 40 mM DTT) was added to the gel in equal volumes and incubated at 4 °C overnight. The resulting liquid was transferred to a 1.5 mL DNA Lo-Bind tube. This liquid was used for PAGE gel assays. To this tube, a neutralisation buffer of 150 mM HCI, and 2x Taq buffer was added in equal volumes to the dissolution buffer, resulting in a total volume of 3x the gel volume. The resulting mixture was cleaned using AMPure beads (Beckman Coulter, A63881) at 1.5x excess. The cleaned product was end-repaired using the NEBNext Ultra II End Repair / dA-Tailing Module (NEB, E7546S), and IDT xGen-UDI-UMI adapters (IDT, 10005903) were ligated using the NEBNext Ultra II Ligation Module (NEB, E7595S). The ligated library was cleaned with an AMPure ratio of 0.8x and amplified for 10 cycles using the KAPA Library Amplification Kit with Primer Mix (Roche, KK2620), and cleaned again with a 0.8x volume of AMPure. To further clean the library and remove side reactions, samples were run on a BluePippin Instrument (BioNodika, SASBLU0001) with 2% agarose, dye-free gel with internal standards 100 bp - 600 bp (BioNordika, SASBDF2010). Samples were collected using a tight collection set of 408 bp. The resulting library was collected and quantified using the Invitrogen quant iT Qubit dsDNA HS Assay Kit (ThermoFisher, Q32851), and its quality was checked using High Sensitivity Bioanalyzer chips (Agilent, 5067-4626). The library was loaded onto a NextSeq 550 High Output Kit v2.5 (300 cycles) (Illumina, 20024908).
[0313] Data processing. Raw FASTQ data was first trimmed using fastp [23, 24] using default settings with "cut-right" and "cut-front" enabled. The filtered reads were aligned to the reference concatemers with Bowtie2
[0025] using the default settings plus settings the maximum number of allowed N's to 60 ("--n-ceil 60"), the seed length to 12 ("-L 12"), the seed interval to 0.5 + Vx, where x is the length of the read ("-i S, 0.5,1"), the minimum score to -60 - 0.5 * x, where x is the length of the read ("— score-min L, -60, -0.5"), and enabling the xeq" option to display matches and mismatches in the SAM record. Barcodes were extracted using an in-house built pipeline written in Python 3, using packages Pysam [26, 27], Numpy
[0028] , Scipy
[0029] , Matplotlib
[0030] , Numba
[0031] , Pandas
[0032] , NetworkX
[0033] and configargparse. Only aligned reads with an average PHRED quality score of at least 20 were considered, from which the barcodes were extracted using the aligned positions of sections flanking the barcodes. Extracted barcodes were only kept if they did not contain insertions, deletions, or ambiguous bases. Barcodes were also discarded if their flanking sequences had an edit distance to their reference sequence of more than 1 + O.lx, where x is the length of the section. Also, if insertions were directly adjacent to the barcodes, or insertions in flanking sequences could be placed next to the barcode without changing the alignment score, the barcode was discarded. Extracted barcodes were then fused together per paired read to obtain the two types of polony barcodes and the two types of event barcodes. Fused event barcodes were created by fusing the two barcodes together. Paired reads without a full set of extracted barcodes were discarded.
[0314] Barcode error correction. Sequencing errors were corrected in each of the different barcode types, using principles described earlier in UMI-tools
[0034] and EASL
[0018] . Per barcode type, barcodes were unidirectionally linked to other barcodes if they were within a Hamming distance of 1, and if their total read count was equal or higher. From the resulting tree-like structures, only the most abundant barcode was kept, while the rest was discarded, to minimize potential barcode fusions. This procedure was first done for each of the two types of polony barcodes, using all polony barcodes of both the geometry reaction and the measurement reaction. For the geometry reaction, the fused event barcode was used for standalone error correction. Each unique event barcode was then assigned to the polony barcode pair with which it was most frequently associated (based on read count), to minimize errors arising out of chimeric products. Paired reads from which barcodes were filtered were discarded completely. The paired reads of the geometry reaction were then converted to a graph, where the two polony barcodes were used as node IDs, and the number of event barcodes indicated the edge weight. For the measurement layer, the fused event barcodes were too short for standalone pooling. These were instead first grouped per associated polony barcode, and then pooled to remove the sequencing errors as described above. The remaining number of event barcodes per polony barcode was then used as the measurement count for each targeted measurement product.
[0315] Graph filtering and reconstruction. The generated adjacency graphs obtained from the geometry layer were then pruned, by first taking the largest connected component, and then iteratively removing nodes that contained only a single edge, until no more nodes were removed by this process. The resulting graphs were analysed using inhouse scripts, and were either reconstructed as is, or first filtered by the edge filter implemented by MiniPath
[0021] . After MiniPath filtering, the largest connected component of the graphs was again iteratively pruned before being used for reconstruction.
[0316] Spatial 2D reconstructions were obtained from the graphs using the previously described point Maximum Likelihood Embedding (ptMLE) or spectral Maximum Likelihood Embedding (sMLE) methods described earlier
[0018] . The number of eigenvectors used for ptMLE initialization or for sMLE was set to 100. The previously described pipeline was slightly adapted to change the API and allow parsing of different file formats. Resulting reconstructions were displayed by creating 2D histograms using the matplotlib library, using either the number of nodes or the number of measurement products as weights.
[0317] Barcode swapping assessment. For samples that contained a smaller gel inside a bigger one, 0.5 pL of seed-containing gel was placed at the bottom of a PCR tube and left to gel for 10 min. This was done by creating a mix for a 20 pL reaction following the standard DNA microscopy protocol described earlier and transferring 0.5 pL of this to a new tube before the gel was solidified. After this incubation, 4.5 pL of seed-free gel was added in a similar manner and was then subjected to the PCR program as described earlier. This process was also performed for the uniform gels. Here, 20 pL of gel mix containing seed strands was made and 5 pL was transferred to a new tube and subjected to PCR as described before. Samples were dissolved and prepared into a library and run on a Bioanalyzer as described before. The volume of the polonies was calculated as: volume of polony (V) = (volume of gel) I (number of polonies). The diameter of the polonies was calculated using the formula:
[0318] Barcode swapping assessment. A PCR reaction was set up as described before, with no target and known barcodes instead of random seeds. Samples would be made that would contain a single combination of seeds (e.g. seed combination 1 / 2). In samples where barcode swapping would be assessed, the barcode combinations are kept separate during DNA microscopy. The dissolution and purification happen as described before. During library prep, at each step of the library prep, two separate barcode combinations would be mixed into one sample containing 1 / 2 and 3 / 4 combinations. The library prep itself happens as described before. The resulting library was collected and quantified using the Invitrogen quant iT Qubit dsDNA HS Assay Kit (ThermoFisher, Q32851), and its quality was checked using High Sensitivity Bioanalyzer chips (Agilent, 5067-4626). The library was loaded onto a NextSeq 550 High Output Kit v2.5 (300 cycles) (Illumina, 20024908).
[0319] Example 4
[0320] Algorithm implementation
[0321] Two methods were implemented in python to calculate indirect paths. The first method starts from the asymmetric adjacency matrix A(i,;) = w(i,;), then calculating the three- step adjacency matrix A3(i,j) = A * AT* A, then, for each node pair with an edge, subtracting the paths that use their direct edge, while setting other node pairs to 0:
[0322] Calculating A3 becomes computationally too challenging for large datasets, but since it calculates all paths of length three, not just those between nodes that were originally connected. The calculation of only the indirect paths of length three between nodes connected in the original dataset was therefore implemented, using sparse matrices. The method first calculates A2(i,;) =A*AT, which contains all two-step paths between all nodes of one partition (e.g., LT). Then it iterates over every edge to find all indirect paths of length three.
[0323] For node a in V:
[0324] For every node b connected to a (i.e. , A(ib, ia) > 0) : common nodes = set (nodes connected to b in two steps) U set (nodes connected to a) (i.e. where (A2 (ib, : ) > 0) U where (A( : , ia) > 0) ) For node c in common node indices:
[0325] A3 (ia, ic) += (A2 (ib, ic) - A(ic, ib) * A(ib, ia) ) * A(ic, ia)
[0326] Here ixdenotes the index of node x. The second method is implemented in python with Numba16acceleration to allow for the use of multiple threads.
[0327] For node splitting, the graph Gs was extracted and partitioned using spectral graph partitioning tool from the scikit-learn package (Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: Machine Learning in {P}ython. Journal of Machine Learning Research. 2011;12:2825-2830). Nodes were not considered for splitting if they had either fewer than four connections, or if Gs consisted of more than two components. Normalized cuts were calculated first for all nodes, then nodes were selected for splitting according to the applied cutoff. If two nodes with an edge were both split, the edge was removed.
[0328] Simulated data processing
[0329] Polony locations were reconstructed using the sMLE method described previously (original pipeline found at https: / / github.com / jaweinst / dnamic). A slightly adapted version of the pipeline was used to process large amounts of files more easily. In contrast to experimental data, simulated data was not subjected to the iterative minimum UEI filter prior to reconstruction. Reconstructions from graphs where the largest connected component was smaller than 80% of all nodes were not considered for further analysis. Global reconstruction quality was assessed using the Procrustes disparity, while local reconstruction quality was assessed by the overlap the k-nearest neighbours for each node in the original and reconstruction positions, as suggested in Fernandez Bonet D, Hoffecker IT, Image recovery from unknown network mechanisms for DNA sequencing-based microscopy, Nanoscale, 2023;15(18):8153-8157, doi: 10.1039 / D2NR05435C. To evaluate node splitting accuracy, we paired each set of nodes Sai and Sa2 to their closest match in Sb and Sc, and calculated the overlap:
[0330] Experimental data processing
[0331] Raw sequencing data was downloaded from the Sequencing Read Archive (project number PRJNA487001, sample 3), and processed as described previously, without a minimum read count. After applying a read count filter or indirect path filter, the remaining nodes were filtered as described earlier, first by iteratively removing nodes with less than two associated products (UEIs) to remove possible uncorrected sequencing errors, then by selecting the largest connected component.
[0332] Example 5
[0333] Previous work has demonstrated the feasibility of image reconstructing using DNA sequencing microscopy (e.g. Figure 21(a)). Both the geometry reactions and the measurement reactions have been successfully performed in a cell-culture sample inside the gel matrix with anchored seed strands and sub-seed strands, as shown in, for example, Figure 21(d) and (e).
[0334] Moreover, geometric data has also been obtained from mouse brain on the geometry layer alone (i.e. geometry reconstruction without measurement layer). This was done by seeding a fixed and permeabilised slice of mouse brain and running PCR cycles as described herein. The sequencing of the network revealed a rich interconnected network of the seed strands that can be laid out in 3D, for example in Figure 21(c). A simulation pipeline has also been implemented, both for in silico PCR simulations and reconstructions with one of the runs shown in Figure 21(b). REFERENCES
[0335] The listing or discussion of an apparently prior-published document in this specification should not necessarily be taken as an acknowledgement that the document is part of the state of the art or is common general knowledge. All references listed below and throughout this application are incorporated by reference.
[0336] 1. Lewis, S. M. et al. Spatial omics and multiplexed imaging to explore cancer biology. Nature Methods 2021 18:9 18, 997-1012 (2021).
[0337] 2. Crosetto, N., Bienko, M. & Van Oudenaarden, A. Spatially resolved transcriptomics and beyond. Nature Reviews Genetics 2014 16: 1 16, 57-66 (2014).
[0338] 3. Moffitt, J. R., Lundberg, E. & Heyn, H. The emerging landscape of spatial profiling technologies. Nature Reviews Genetics 2022 1-19 (2022) doi: 10.1038 / s41576-022-00515-3.
[0339] 4. Kashima, Y. et al. Single-cell sequencing techniques from individual to multiomics analyses. Exp Mol Med 52, 1419 (2020).
[0340] 5. Stuart, T. & Satija, R. Integrative single-cell analysis. Nature Reviews Genetics 2019 20:5 20, 257-272 (2019).
[0341] 6. Elmentaite, R., Dominguez Conde, C., Yang, L. & Teichmann, S. A. Single-cell atlases: shared and tissue-specific cell types across human organs. Nature Reviews Genetics 2022 23:7 23, 395-410 (2022).
[0342] 7. Kishi, J. Y. et al. SABER amplifies FISH: enhanced multiplexed imaging of RNA and DNA in cells and tissues. Nature Methods 2019 16:6 16, 533-544 (2019).
[0343] 8. Codeluppi, S. et al. Spatial organization of the somatosensory cortex revealed by osmFISH. Nature Methods 2018 15: 11 15, 932-935 (2018).
[0344] 9. Chen, K. H., Boettiger, A. N., Moffitt, J. R., Wang, S. & Zhuang, X. Spatially resolved, highly multiplexed RNA profiling in single cells. Science (1979) 348, (2015).
[0345] 10. Ke, R. et al. In situ sequencing for RNA analysis in preserved tissue and cells. Nature Methods 2013 10:9 10, 857-860 (2013).
[0346] 11. Lubeck, E., Coskun, A. F., Zhiyentayev, T., Ahmad, M. & Cai, L. Single cell in situ RNA profiling by sequential hybridization. Nat Methods 11, 360 (2014).
[0347] 12. Wang, G., Moffitt, J. R. & Zhuang, X. Multiplexed imaging of high-density libraries of RNAs with MERFISH and expansion microscopy. Scientific Reports 2018 8: 1 8, 1-13 (2018).
[0348] 13. Eng, C. H. L. et al. Transcriptome-scale super-resolved imaging in tissues by RNA seqFISH + . Nature 2019 568:7751 568, 235-239 (2019).
[0349] 14. Stahl, P. L. et al. Visualization and analysis of gene expression in tissue sections by spatial transcriptomics. Science (1979) 353, 78-82 (2016). 15. Liu, Y. et al. High-Spatial-Resolution Multi-Omics Sequencing via Deterministic Barcoding in Tissue. Cell 183, 1665-1681. el8 (2020).
[0350] 16. Rodriques, S. G. et al. Slide-seq: A scalable technology for measuring genomewide expression at high spatial resolution. Science (1979) 363, 1463-1467 (2019).
[0351] 17. Cho, C. S. et al. Microscopic examination of spatial transcriptome using Seq- Scope. Cell 184, 3559-3572. e22 (2021).
[0352] 18. Lewis, S. M. et al. Spatial omics and multiplexed imaging to explore cancer biology. Nature Methods 2021 18:9 18, 997-1012 (2021).
[0353] 19. Crosetto, N., Bienko, M. & Van Oudenaarden, A. Spatially resolved transcriptomics and beyond. Nature Reviews Genetics 2014 16: 1 16, 57-66 (2014).
[0354] 20. Moffitt, J. R., Lundberg, E. & Heyn, H. The emerging landscape of spatial profiling technologies. Nature Reviews Genetics 2022 1-19 (2022) doi: 10.1038 / S41576-022-00515-3.
[0355] 21. Kashima, Y. et al. Single-cell sequencing techniques from individual to multiomics analyses. Exp Mol Med 52, 1419 (2020).
[0356] 22. Stuart, T. & Satija, R. Integrative single-cell analysis. Nature Reviews Genetics 2019 20:5 20, 257-272 (2019).
[0357] 23. Elmentaite, R., Dominguez Conde, C., Yang, L. & Teichmann, S. A. Single-cell atlases: shared and tissue-specific cell types across human organs. Nature Reviews Genetics 2022 23:7 23, 395-410 (2022).
[0358] 24. Kishi, J. Y. et al. SABER amplifies FISH: enhanced multiplexed imaging of RNA and DNA in cells and tissues. Nature Methods 2019 16:6 16, 533-544 (2019).
[0359] 25. Codeluppi, S. et al. Spatial organization of the somatosensory cortex revealed by osmFISH. Nature Methods 2018 15: 11 15, 932-935 (2018).
[0360] 26. Chen, K. H., Boettiger, A. N., Moffitt, J. R., Wang, S. & Zhuang, X. Spatially resolved, highly multiplexed RNA profiling in single cells. Science (1979) 348, (2015).
[0361] 27. Ke, R. et al. In situ sequencing for RNA analysis in preserved tissue and cells. Nature Methods 2013 10:9 10, 857-860 (2013).
[0362] 28. Lubeck, E., Coskun, A. F., Zhiyentayev, T., Ahmad, M. & Cai, L. Single cell in situ RNA profiling by sequential hybridization. Nat Methods 11, 360 (2014).
[0363] 29. Wang, G., Moffitt, J. R. & Zhuang, X. Multiplexed imaging of high-density libraries of RNAs with MERFISH and expansion microscopy. Scientific Reports 2018 8: 1 8, 1-13 (2018).
[0364] 30. Eng, C. H. L. et al. Transcriptome-scale super-resolved imaging in tissues by RNA seqFISH + . Nature 2019 568:7751 568, 235-239 (2019).
[0365] 31. Stahl, P. L. et al. Visualization and analysis of gene expression in tissue sections by spatial transcriptomics. Science (1979) 353, 78-82 (2016). 32. Liu, Y. et al. High-Spatial-Resolution Multi-Omics Sequencing via Deterministic Barcoding in Tissue. Cell 183, 1665-1681. el8 (2020).
[0366] 33. Rodriques, S. G. et al. Slide-seq: A scalable technology for measuring genomewide expression at high spatial resolution. Science (1979) 363, 1463-1467 (2019).
[0367] 34. Cho, C. S. et al. Microscopic examination of spatial transcriptome using Seq- Scope. Cell 184, 3559-3572. e22 (2021).
[0368] 35. Weinstein, J. A., Regev, A. & Zhang, F. DNA Microscopy: Optics-free Spatio- genetic Imaging by a Stand-Alone Chemical Reaction. Cell 178, 229-241. el6 (2019).
[0369] 36. Hoffecker, I. T., Yang, Y., Bernardinelli, G., Orponen, P. & Hdgberg, B. A computational framework for DNA sequencing microscopy. Proc Natl Acad Sci U S A 116, 19282-19287 (2019).
[0370] 37. Mitra, R. D. & Church, G. M. In situ localized amplification and contact replication of many individual DNA molecules. Nucleic Acids Res 27, e34-e39 (1999).
[0371] 38. Kloosterman, A., Baars, I. & Hdgberg, B. An error correction strategy for image reconstruction by DNA sequencing microscopy. Nature Computational Science 2024 4:2 4, 119-127 (2024).
[0372] 39. Zadeh, J. N. et al. NUPACK: Analysis and design of nucleic acid systems. J Comput Chem 32, 170-173 (2011).
[0373] 40. Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884-i890 (2018).
[0374] 41. Chen, S. Ultrafast one-pass FASTQ data preprocessing, quality control, and deduplication using fastp. iMeta 2, el07 (2023).
[0375] 42. Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357-359 (2012).
[0376] 43. Andreas Heger. Pysam. https: / / github.com / pysam-developers / pysam.
[0377] 44. Li, H. et al. The Sequence Alignment / Map format and SAMtools. Bioinformatics 25, 2078-2079 (2009).
[0378] 45. Harris, C. R. et al. Array programming with NumPy. Nature 585, 357-362 (2020).
[0379] 46. Virtanen, P. et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods 17, 261-272 (2020).
[0380] 47. Hunter, J. D. Matplotlib: A 2D graphics environment. Comput Sci Eng 9, 90-95 (2007).
[0381] 48. Lam, S. K., Pitrou, A. & Seibert, S. Numba: A llvm-based python jit compiler, in Proceedings of the Second Workshop on the LLVM Compiler Infrastructure in HPC 1- 6 (2015).
[0382] 49. Team, T. pandas development. pandas vl.4.4. (2022) doi: doi.org / 10.5281 / zenodo.7037953. 50. Hagberg, A. A., Schult, D. A. & Swart, P. J. Exploring Network Structure, Dynamics, and Function using NetworkX. in Proceedings of the 7th Python in Science Conference (eds. Varoquaux, G., Vaught, T. & Millman, J.) 11-15 (Pasadena, CA USA, 2008). 51. Smith, T., Heger, A. & Sudbery, I. UMI-tools: modeling sequencing errors in
[0383] Unique Molecular Identifiers to improve quantification accuracy. Genome Research 27, 491-499 (2017).
Claims
CLAIMS1. A method for preparing a sample containing one or more cells for the construction of a spatial network using nucleic acids, the method comprising:(a) providing a sample containing one or more cells; and(b) randomly immobilizing exogenous nucleic acids in the sample, wherein each nucleic acid comprises a unique sequence identifier, wherein each nucleic acid identifies a node in the spatial network.
2. The method of claim 1, wherein step (b) comprises applying a gel comprising the nucleic acid molecules to a sample that results in immobilization of the nucleic acids at multiple points in the sample.
3. The method of claim 2, wherein the gel is a hydrogel.
4. The method of claim 3, wherein PEG-based or acrylamide-based gelling components are used to form the hydrogel.
5. The method of claim 4, wherein PEG-based gelling components are acrylate- PEG and / or thio-PEG gelling components and / or wherein the hydrogel is an acrylamide- based hydrogel.
6. The method of any one of claims 1 to 5 comprising a step of forming polonies from the nucleic acids immobilized in the sample.
7. The method of claim 6, wherein each polony forms a node in the spatial network.
8. The method of any one of claims 1 to 7, wherein the nucleic acids are immobilized (e.g. crosslinked) to endogenous biomolecules in the sample.
9. The method of claim 8, wherein the biomolecule is a protein, nucleic acid or lipid.
10. The method of any one of claims 2 to 7, wherein the nucleic acids are immobilized to the gel.
11. The method of any one of claims 1 to 10, wherein the nucleic acids are immobilized via a covalent bond.
12. The method of any one of claims 1 to 11, wherein the nucleic acids are immobilized via a reaction with an amine or thiol group in the sample or gel.
13. The method of any one of claims 1 to 12, wherein the nucleic acids are acrydite- modified nucleic acids.
14. The method of claim 13, wherein acrydite-modified nucleic acids comprise a methacryl group at their 5' ends.
15. A sample containing one or more cells for the construction of a spatial network using nucleic acids, the sample comprising randomly immobilized exogenous nucleic acids in the sample, wherein each nucleic acid comprises a unique sequence identifier and identifies a node in the spatial network.
16. The sample of claim 15, wherein the sample and / or nucleic acids are as defined in any one of claims 2 to 14.
17. A system for the construction of a spatial network in a sample comprising cells, the system comprising:(A) a plurality nucleic acids, wherein each nucleic acid comprises a unique sequence identifier; and(B) means for randomly immobilizing the plurality of nucleic acids in the sample, such each nucleic acid identifies a node in the spatial network.
18. The system of claim 17, wherein the means for randomly immobilizing the plurality of nucleic acids in the sample comprises a cross-linking agent.
19. The system of claim 17 or 18, wherein the sample and / or nucleic acids are as defined in any one of claims 2 to 14.
20. Use of the system of any one of claims 17 to 19 for the construction of a spatial network in a sample comprising cells.
21. A method for constructing of a spatial network in a sample containing one or more cells using nucleic acids, the method comprising:(a) preparing a sample containing cells according to any one of claims 1 to 14 or providing a sample according to claim 15 or 16;(b) forming nucleic acid nodes (e.g. polonies) from the nucleic acids immobilized in the sample; and(c) deducing the location of the nucleic acid nodes based on interactions between nucleic acids in adjacent nodes thereby constructing a spatial network.
22. A method for mapping transcriptomic and / or proteomic information in a sample containing one or more cells, comprising:(a) constructing a spatial network in the sample using nucleic acids according to claim 21;(b) forming nucleic acids (e.g. concatemers) from interactions between nucleic acid nodes and target nucleic acids in the sample that contain transcriptomic or proteomic information; and(c) using the nucleic acids (e.g. concatemers) formed in (b) to generate a map of transcriptomic and / or proteomic information in the sample.
23. A method for determining the spatial arrangement of two or more targetable molecules comprising: a. providing a sample comprising at least two targetable molecules, wherein each targetable molecule is at a distinct location in the sample; b. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule; c. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein:(i) at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and / or(ii) at least one of the plurality of seed-specific amplicons hybridises with and extends on a measurement sub-seed to form a measurement monomer, and wherein each measurement monomer is complementary to a marker monomer, and the monomers hybridise and extend on each other to form a measurement concatemer;d. determining the spatial arrangement of the two or more targetable molecules on the basis of the geometry concatemers and / or measurement concatemers formed in step (c).
24. A method of identifying a disease, disorder, or condition in a patient comprising: a. providing a sample from the patient, wherein the sample comprises at least two targetable molecules, and wherein each targetable molecule is at a distinct location in the sample; b. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule; c. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein:(i) at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and / or(ii) at least one of the plurality of seed-specific amplicons hybridises with and extends on a measurement sub-seed to form a measurement monomer, and wherein each measurement monomer is complementary to a marker monomer, and the monomers hybridise and extend on each other to form a measurement concatemer; d. determining the spatial arrangement of the two or more targetable molecules on the basis of the geometry concatemers and / or measurement concatemers formed in step (c), e. and identifying the disease, disorder, or condition on the basis of the spatial arrangement.
25. A method for determining the spatial arrangement of two or more targetable molecules comprising: a. providing a sample comprising at least two targetable molecules, wherein each targetable molecule is at a distinct location in the sample; b. labelling each targetable molecule with a polynucleotide molecule comprising a unique sequence identifier, thereby forming a seed at each targetable molecule;c. forming a plurality of seed-specific amplicons from each polynucleotide molecule, wherein at least one of the plurality of seed-specific amplicons hybridises with and extends on a geometry sub-seed to form a geometry monomer, wherein the geometry sub-seed comprises a 3' end that is blocked from extension, and wherein each geometry monomer is complementary to an adjacent geometry monomer, and the monomers hybridise and extend on each other to form a geometry concatemer; and d. determining the spatial arrangement of the two or more targetable molecules on the basis of the geometry concatemers formed in step (c).
26. The method of any preceding claim, wherein one or more of the targetable molecules comprises an mRNA; and / or comprises a molecule labelled with a polynucleotide tag.
27. The method of any preceding claim, wherein the two or more targetable molecules are positioned and / or immobilised in the sample, optionally wherein the two or more targetable molecules are positioned and / or immobilised in a gel, preferably a dissolvable and / or cleavable gel.
28. The method of any preceding claim, wherein the sample: i. comprises one or more cell; ii. is fixed; and / or iii. is permeabilised.
29. The method of any preceding claim, wherein the step of forming seed-specific amplicons comprises extension of the polynucleotide molecule, optionally wherein extension is performed using a DNA polymerase.
30. The method of any preceding claim, wherein the step of forming seed-specific amplicons from each polynucleotide molecule comprises: i. linear spooling amplification; ii. PCR amplification; and / or iii. rolling circle amplification, for example using MOSIC.
31. The method of any preceding claim, wherein the formation of the geometry monomer, measurement monomer and / or marker monomer comprises: i. linear spooling amplification;ii. PCR amplification; and / or iii. rolling circle amplification, for example using MOSIC.
32. The method of any preceding claim, wherein the seed and sub-seeds are positioned and / or immobilised in the sample, optionally wherein the seed and subseeds are positioned and / or immobilised in a gel, preferably a hydrogel.
33. The method of any preceding claim, wherein the number and / or concentration of seeds in the sample is: i. lower than the number and / or concentration of sub-seeds in the sample; and / or ii. from about 100 fM to about 1 pM, preferably 1 pM, and optionally wherein the number and / or concentration of sub-seeds in the sample is about 400 nM.
34. The method of any preceding claim, wherein: i. the ratio of geometry sub-seeds to measurement sub-seeds in the sample is 1 : 1; and / or ii. the geometry concatemers and measurement concatemers are the same size.
35. The method of any preceding claim, wherein: i. the number and / or concentration of measurement sub-seeds in the sample is greater than the number and / or concentration of geometry sub-seeds in the sample; and / or ii. the number and / or concentration of geometry sub-seeds in the sample is greater than the number and / or concentration of measurement sub-seeds in the sample.
36. The method of any preceding claim, wherein the geometry sub-seed comprises: i. a unique sequence identifier; ii. optionally, a 3' end that is blocked from extension; and / or iii. a barcode, and annealing and extension of two geometry monomers forms a geometry concatemer comprising a pair of complementary barcodes on each strand.
37. The method of any preceding claim, wherein the measurement sub-seed comprises a unique sequence identifier.
38. The method of any preceding claim, wherein the unique sequence identifier comprises: i. at least one barcode; ii. at least one primer sequence; iii. at least one pre-activation sequence; and / or iv. at least one activator sequence.
39. The method of any preceding claim, wherein the marker monomer comprises a marker polynucleotide, optionally wherein: i. the marker polynucleotide is selected from the group consisting of: mRNA, cDNA, and DNA; optionally wherein the mRNA, cDNA or DNA is conjugated to an antibody or antigen-binding fragment thereof; ii. the marker polynucleotide is endogenous to the sample, or is exogenously introduced to the sample; iii. the polynucleotide sequence of the marker polynucleotide has not been determined; and / or iv. the marker polynucleotide is not GAPDH and / or ACTB.
40. The method of any preceding claim, wherein a barcode is added to the marker monomer, optionally by using oligo-T and a template switching ribo-G oligonucleotide.
41. The method of any preceding claim, wherein the sub-seeds are incapable of extending until a seed-specific amplicon hybridises with the sub-seed.
42. The method of any preceding claim, wherein the step of determining the spatial arrangement of the targetable molecules comprises determining the polynucleotide sequence of one or more geometry concatemer and / or one or more measurement concatemer formed in step (c).
43. The method of any preceding claim, further comprising one or more of the following steps (in any order): i. gel extraction, optionally an automated gel extraction; ii. dissolving the gel, optionally using an alkaline buffer with DTT; iii. isolation and / or purification of the concatemers; iv. tagmentation; v. a circularisation protocol; and / or vi. sequencing, preferably high-throughput sequencing.
44. A kit comprising: a. at least two populations of polynucleotides, wherein each polynucleotide comprises a unique sequence identifier; and b. a gelling agent and / or a cross-linking agent.
45. The kit of claim 44, wherein the polynucleotides are acrydite-modified polynucleotides.
46. The kit of claim 45, wherein acrydite-modified nucleic acids comprise a methacryl group at their 5' ends.
47. The kit of any one of claims 44 to 46, wherein the gelling agent is PEG-based or acrylamide-based.
Citation Information
Patent Citations
DNA microscopy
WO2017044893A1
Method of in SITU gene sequencing
WO2019199579A1
Barcode diffusion-based spatial omics
WO2024020395A2