Spatial perturb-SEQ: multiplexed screens with spatial localization in intact tissues
The use of polynucleotides with unique barcodes for gene perturbation in intact tissues addresses scalability and spatial context limitations of existing techniques, enabling efficient and cost-effective analysis of gene expression changes.
Patent Information
- Application Number
- PCT/SG2025/050446
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-07-02
- Publication Date
- 2026-01-08
AI Technical Summary
Existing functional genomic techniques, such as CRISPR-based pooled genetic screens and Perturb-Seq, lack scalability, versatility, and single-cell resolution, and require tissue dissociation, limiting the understanding of cellular interactions in native tissue environments.
A method involving polynucleotides with unique barcodes that are delivered to a population of cells to perturb specific genes, allowing for spatial and transcriptomic analysis in intact tissue samples, using a library of polynucleotides with specific sequence characteristics to enable spatial localization and barcode detection.
Enables high-throughput, scalable, and cost-effective analysis of gene expression changes with preserved spatial context, providing detailed insights into cellular interactions without the need for extensive technical expertise or specialized setups.
Smart Images

Figure SG2025050446_08012026_PF_FP_ABST
Abstract
Description
SPATIAL PERTURB-SEQ: MULTIPLEXED SCREENS WITH SPATIAL LOCALIZATION IN INTACT TISSUESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority of Singapore application No. 10202401946S filed on 2 July 2024, the contents of it being hereby incorporated by reference in its entirety for all purposes.FIELD OF THE INVENTION
[0002] The present invention relates to the field of molecular and cell biology. In particular, the present invention relates to polynucleotides comprising barcodes detectable by optical and sequencing methods, and methods of obtaining spatial and transcriptomic information of cells within intact tissue samples.SEQUENCE LISTING
[0003] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created 5 May 2025, is named 89216SGl_Sequence Listing.xml and is 91 kilobytes in size.BACKGROUND OF THE INVENTION
[0004] Recent advancements in functional genomic techniques address the limitations of the traditional approaches such as human genome-wide association studies to determine causality and function of genes associated with diseases. Such approaches such as CRISPR- based pooled genetic screens, Perturb-Seq and Perturb-Map allow high throughput transcriptomic analysis of the effects of genetic perturbations on gene expression in vivo.
[0005] The CRISPR / Cas9 system as applied in CRISPR-based pooled genetic screens enables targeted genetic perturbation in a population of cells to systematically interrogate gene function. However, the system lacks scalability, versatility for use with sequencing readouts, and single-cell resolution. To overcome these challenges, Perturb-Seq was developed, utilizing the combination of pooled CRISPR screens with single cell transcriptomics to investigate theeffects of genetic perturbation on gene expression changes in individual cells. Although Perturb-Seq has been shown to increase the scale of functional genomics studies, this approach requires tissue dissociation into individual cells and lacks spatial context, limiting the ability to capture and understand the interactions between cells in a native tissue environment. Another approach is Perturb-map, a spatial functional genomics technique that addresses the limitations of Perturb-Seq, which integrates profiling of gene expression and spatial-mapping technologies, relying on protein barcodes that can be detected by imaging methods to preserve the spatial context of the cells. While this approach allows the understanding of the effects of genetic perturbations on cellular communication in their native tissue architecture through the incorporation of spatial information with transcriptomic information, the processes are laborious and costly with poor scalability due to the requirement of extensive technical expertise and specialized microscopic setups.
[0006] Therefore, there is a need to provide a method of obtaining transcriptomic and spatial information that overcomes, or at least ameliorates, one or more of the disadvantages described above. Furthermore, other desirable features and characteristics will become apparent from the subsequent detailed description and the appended claims, taken in conjunction with the accompanying drawings referred to herein.SUMMARY
[0007] In one aspect, provided herein is a polynucleotide comprising: a) a polynucleotide sequence complementary to a polynucleotide sequence encoding a gene of interest or a segment thereof, and b) a barcode that is unique to the gene of interest comprising a polynucleotide sequence, wherein i. the polynucleotide sequence of the barcode is between about 20 to about 500 nucleotides, ii. wherein the barcode comprises an edit distance of more than 16 nucleotides between segments, iii. wherein the barcode is devoid of a polynucleotide sequence encoding a stop codon, and iv. wherein the barcode is devoid of a polynucleotide sequence comprising4 or more identical consecutive nucleotides.
[0008] In another aspect, provided herein is a library comprising a plurality of polynucleotides as described herein.
[0009] In another aspect, provided herein is a method of obtaining spatial and transcriptomic information of a population of cells, wherein the method comprises: a) delivering a plurality of polynucleotides as described herein or the library as described herein to the population of cells, wherein each of the plurality of polynucleotides perturbs a distinct gene of interest in the population of cells; b) obtaining an intact tissue sample comprising the population of cells; and c) detecting the plurality of polynucleotides in the population of cells in the intact tissue sample, thereby obtaining spatial and transcriptomic information of the population of cells.DEFINITIONS
[0010] The following words and terms used herein shall have the meaning indicated:
[0011] As used herein, the term “gRNA” or “guide RNA” refers to a synthetic polynucleotide sequence complementary to a polynucleotide sequence of a target gene or gene of interest which hybridizes with the polynucleotide sequence of the gene. The terms “target gene” and “gene of interest” may be used interchangeably in the context of this disclosure. It should be generally understood that the extent of complementarity of the polynucleotide sequence of the gRNA with the polynucleotide sequence of the targeted gene may result in deletion of the gene entirely or removal or insertion of one or more nucleotides of the gene, leading to a reduction or loss of genetic and protein function.
[0012] As used herein, the terms “perturbed” or “edited” in the context of a gene refer to the alteration in the polynucleotide sequence of the gene, resulting in a change in genetic and protein function of the gene. A change in genetic function of a gene can be a gain of genetic function or a loss of genetic function. A gain of genetic function refers to the expression of a gene that was not previously expressed prior to the perturbation, or expression of a gene at a higher level than what was expressed prior to the perturbation. A loss of genetic function refers to a gene that is not being expressed after the perturbation, or expression of a gene at a lower level than what was expressed prior to the perturbation. A change in protein function can be a gain of protein function or a loss of protein function. A gain of protein function refers to the expression of a protein that was not previously expressed prior to the perturbation, expressionof a protein at a higher level than what was expressed prior to the perturbation. A loss of protein function refers to a protein that is not being expressed after the perturbation, or expression of a protein at a lower level than what was expressed prior to the perturbation. With respect to a cell, the terms “perturbed” or “edited” refer to the alteration of the polynucleotide sequences of one or more genes in the cell The gene may be perturbed or edited to completely remove the polynucleotide sequence of a gene, add a polynucleotide sequence of a gene, delete one or more nucleotides to a gene or add one or more nucleotides to a gene. The polynucleotide sequences of one or more genes in the cells may be altered, resulting in changes in protein activity, protein function, or protein structure. Examples of protein activity may comprise but are not limited to enzymatic or signaling activity, protein interactions and protein localization.
[0013] As used herein, the term “barcode” refers to a tag or a label that allows the identification of one or more genes targeted in a cell. Barcodes may be but are not limited to polynucleotide sequences, such as DNA or RNA sequences, or polypeptide sequences and protein barcodes The barcode may be inserted into a vector or a plasmid comprising one or more polynucleotide sequences that are complementary to a target gene or gene of interest. The barcode allows the mapping of the location of the gene in the cells on a tissue. Barcode may be exogenously introduced into the cells on a tissue.
[0014] As used herein, the term “tolerated” with respect to barcode refers to the state of a cell or a population of cells to which the barcode is delivered to. The parameters that may be used to measure the state of the cell or the population of cells comprise the health and / or viability of cells The state of the cell or the population of cells to which the barcode is delivered to may be measured relative to the state of the cell or population of cells in which the barcode is absent. One example of a cell health and / or viability parameter is the percentage of mitochondrial genes expressed as a percentage of the whole-cell transcriptome, wherein a lower percentage of mitochondrial genes expressed relative to a reference cell indicates better health and / or higher viability relative to the reference cell. Other suitable parameters of cell health and / or viability and methods of measuring said parameters will be apparent to those skilled in the art and can also be used.
[0015] As used herein, the terms “deleterious effect” and “unwanted biological effect” with reference to the cells refer to a negative impact on the state of a cell or a population of cells. The state of a cell or population of cells includes but is not limited to viability, morphology and function.
[0016] As used herein, the term “tropism” with respect to viral vectors refers to the specificity of a viral vector delivering genetic material that targets a cell or a population of cells. It should be generally understood that tropism differs in viral vectors and each viral vector may be one or more serotypes. The commonly used viral vectors for gene delivery are lenti virus and adeno-associated virus (AAV). It is well known in the art that lentivirus and AAV target different cells. For example, lentivirus targets immune cells comprising but not limited to T cells and B cells while AAV targets cells of different parts of the body comprising but not limited to brain cells and liver cells.
[0017] As used herein, the term “multiplicity of infection” refers to the ratio of viral particles to target cells in the pre-determined region of a sample.
[0018] As used herein, the term “sparse editing” with reference to the cells refers to a population of cells comprising a perturbed or an edited cell and a plurality of cells adjacent to the perturbed or edited cell which are not perturbed or edited. It should be understood that the criteria of a population of cells that is considered to be sparsely edited may differ in the types of experiments performed. For example, the number of cells that are edited or perturbed may differ in an experiment that studies direct cell-cell communication as compared to another experiment that studies long range signalling between the cells.
[0019] As provided herein, the term “neighbouring cells” refers to cells within a specified radius of a cell of interest or cells that are adjacent to a cell of interest. The cells may interact directly or indirectly with the cell of interest.
[0020] As used herein, the term “edit distance” with reference to polynucleotide sequences refers to the measurement of the minimum number of single-nucleotide edits that can be made to modify a first polynucleotide sequence to a second polynucleotide sequence. The editing of the nucleotides may be by means of insertions, deletions or substitutions. The measurement of the minimum number of single-nucleotide edits may be computed by more than one method. Example of the methods may comprise but are not limited to Hamming and Levenshtein. The Hamming distance refers to the number of nucleotides that are mismatched by substitutions at the same positions between two strings of equal length while the Levenshtein distance refers to the minimum number of nucleotides that are mismatched by insertions, deletions, substitutions between two strings of same or differing length. It should be understood that the higher the edit distance between two polynucleotide sequences, the lower the similarity in nucleotides between the two polynucleotide sequences. In contrast, the lower the edit distancebetween two polynucleotide sequences, the higher the similarity in the nucleotides between the two polynucleotide sequences. Edit distance is measured based on the number of nucleotides. For example, an edit distance of 15 in the context of two sequences indicates that there are 15 nucleotides that are different between the two sequences and 15 nucleotide edits are to be made to modify a first polynucleotide sequence to a second nucleotide sequence.
[0021] As used herein, the term “unique” in the context of a polynucleotide sequence refers to a segment of a polynucleotide sequence or a polynucleotide sequence that is present only once in the nucleic acid library. The extent of the differences between each segment of the polynucleotide sequence or each polynucleotide sequence in the nucleic acid library may be measured by the edit distance. The higher the edit distance between two polynucleotide sequences, the higher the number of nucleotides that is different between the two polynucleotide sequences. The term “unique” may also mean the percentage of sequence identity between two or more nucleotide sequences. For example, two or more nucleotide sequences having 100% sequence identity with each other will be considered as “not unique” from each other. On the other hand, two or more nucleotide sequences having less than 100% sequence identity will be considered as being “unique” from each other.
[0022] As used herein, the term “gene module” refers to a plurality of genes that shows alteration in expression levels within a specific cell type. For example, a first gene module refers to genes in oligodendrocytes, a second gene module refers to genes in neurons and a third gene module refers to genes in the choroid plexus.
[0023] As used herein, the term “single-cell RNA sequencing” or “scRNA-seq” refers to the state-of-the-art sequencing approach which allows the detection of expression profiles of individual cells. Single-cell RNA sequencing uncovers the heterogeneity and complexity of RNA transcripts within single cells, as well as revealing the composition of different cell types and functions within highly organized tissues / organs / organisms.
[0024] As used herein, the term “spatial transcriptomics” refers to methods that measure all the gene activity (transcriptome) in a cell or tissue and map the gene activity spatially to a location within the cell or tissue
[0025] As used herein, the term “probe-based spatial phenotyping” refers to a method that utilizes probes to measure all the gene activity in a tissue and map of the location of the activity. Examples of probes include but are not limited to nucleotides, polynucleotides, oligonucleotides and proteins. These nucleotides, polynucleotides, oligonucleotide andproteins may further comprise one or more detectable labels which allow characterization of cells at a single-cell level. The methods commonly used in probe-based spatial phenotyping are fluorescent in situ hybridization (FISH), CosMx, Multiplexed Error-Robust Fluorescence In Situ Hybridization (MERFISH), and Xenium. The present disclosure provides polynucleotides that may be used as targets for probes, for example as a target for a polynucleotide FISH probe, for probe-based spatial phenotyping.
[0026] As used herein, the term “sequence-based spatial phenotyping” refers to a method that allows the measurement of all the gene activity in a tissue and mapping of the location of the activity. Examples of the method may comprise but are not limited by Visium, Stereo-seq, and GeoMx. These methods integrate sequence information including gene expression profiles and spatial context including tissue architecture to define cellular phenotypes.
[0027] As used herein, the term “about”, is used in the context of, but not limited to, concentrations of components and percentages of compounds, typically refers to + / - 10% of the stated value, to + / - 9% of the stated value, to + / - 8% of the stated value, to + / - 7% of the stated value, to + / - 6% of the stated value, to + / - 5% of the stated value, + / - 4% of the stated value, more typically + / - 3% of the stated value, more typically, + / - 2% of the stated value, even more typically + / - 1% of the stated value, and even more typically + / - 0.5% of the stated value. Throughout this disclosure, certain embodiments may be disclosed in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosed ranges. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The invention will be better understood with reference to the detailed description when considered in conjunction with the non-limiting examples and the accompanying drawings, in which:
[0029] Fig. 1 provides an in vivo spatial perturb-seq of the adult mouse brain by intracranial injection of a barcoded AAV library. Fig. 1A depicts a schematic of the in vivo spatial Perturb- seq pipeline, and depiction of components in the AAV-PHP.eB library, showing 3 gRNAs, TdTomato and 1 barcode in each member. Fig. IB shows the confocal imaging of a coronal section of a mouse brain following stereotaxic injection into the hippocampus, showing that transduction spread was limited to the hippocampus. Scale bar: 1000 pm. Fig. 1C depicts the Uniform Manifold Approximation and Projection (UMAP) projection of cells from chip 1 following Stereo-seq, showing recovery of neurons, oligodendrocytes, choroid plexus cells, red blood cells (RBCs), microglia, endothelial cells, and astrocytes. Left: identified by cluster number. Right: identified by cell type. Fig. ID shows that similar proportions of cell types were recovered for each chip, with the majority cell type being neurons as expected. Fig. IE depicts the Dot Plot showing expression of cell-type enriched genes for each cell type. Fig. IF depicts the spatial scatter of annotated mouse brain cells. Fig. 1G shows the spatial gene plot of canonical cell-type markers major cell populations (Microglia: Cl qa; Oligodendrocytes: Mbp, Endothelial cells: Fltl, Neurons: Snap25; Astrocytes: Gfap, Choroid Plexus: Ttr; RBC: Hba-al; Immature oligodendrocytes: Birc5). Fig. Ill depicts the ssDNA image for cellular localization on the Stereo-seq chip. Fig. II depicts the spatial scatter following BANKSY clustering with lamba 0.2, showing clusters recovered, which showed strong alignment between the two sections. Fig. 1J shows the selection of cornu ammoni (CA) and dentate gyrus (DG) neurons in the hippocampus by focusing on BANKSY clusters 3 and 10. Fig. IK shows the subset of barcode-positive cells from Fig. 1J. Fig. IL depicts the pie chart representation of the number of unique gRNA barcodes among all gRNA barcode positive cells, showing a small proportion of barcode-positive cells with multiple barcodes detected. Fig. IM depicts the Dot Plot showing percentage of barcode-positive neurons among hippocampal neurons in Fig. 1J for each chip. Fig. IN shows the total number of cells for each gRNA across all 3 chips. Min: 29, max: 330, mean: 122. Fig. 1F-1L: Data from chip 1.
[0030] Fig. 2 depicts the in vivo Perturb-Seq of the adult mouse brain, with AAV-PHP.eB showing strongest tropism for neurons, astrocytes and oligodendrocytes. Fig. 2A shows the UMAP projection of cells isolated from the mouse brain injected with AAV-PHP.eB, followed by 10X library prep and sequencing and stacked bar plot on the right shows cell type composition, showing recovery of microglia, oligodendrocytes, astrocytes, endothelial cells, fibroblasts, neurons and choroid plexus (CP) cells. Fig. 2B displays the Dot plot showingexpression of 4 cell-type enriched genes for each cell type. Fig. 2C shows the UMAP representation showing the expression of well-known, canonical cell-type markers for 7 major cell populations (Oligodendrocytes: Mog; Microglia: Clqa; Endothelial cells: Cldn5; Astrocytes: Slcla2; Choroid Plexus: Ttr; Fibroblasts: Vtn; Neurons: Snap25). Fig. 2D shows the UMAP projection showing cells with any detected gRNA barcodes in black, and cells with no detected barcodes in grey. Fig. 2E displays the percentage of cells with detected gRNA barcodes for each annotated cell type. Each point represents a batch. Graph depicts mean and SEM. Fig. 2F shows the representation of the number of unique gRNA barcodes among all gRNA barcode positive cells. Fig. 2G shows the violin plots showing representation of all 18 members of the AAV library, split by batches. Fig. 2H shows the Dot Plot showing percentage of barcode-positive oligodendrocytes for each experimental batch. The size of the dots represents the percentage of barcode-positive oligodendrocytes. Fig. 21 shows the total number of oligodendrocytes for each gRNA across all 3 experiments. Min: 9, max: 58, mean: 26. Fig. 2J displays the ridge plot showing distribution of Mixscale scores for each barcode. “Neg”, the control used, denotes cells with no detected barcodes. Fig. 2K displays the ridge plot showing distribution of Mixscale scores for each barcode. “mSafe.gRNA” is used as the NT (nontargeting) control.
[0031] Fig. 3 depicts the parameters of intracranial injection: Spread and biodistribution of AAV. Fig. 3A depicts the percentage of barcode-positive cells in the brain following 3 weeks and 8 weeks incubation post intracranial injection. Fig. 3B shows the expression levels of each of the 18 barcodes 3 weeks and 8 weeks post intracranial injection, showing intended sparse transduction. Fig. 3C shows that AAV cargo was not readily detected in the liver, following intracranial delivery into the brain. PCR amplification of the TdTomato-Barcode region of the AAV cargo showed no significant accumulation in the liver. Fig. 3D shows the representative FACS plot images showing ~7% TdTomato+ cells in each injected mouse brain coronal section.
[0032] Fig. 4 depicts the results showing the presence of Cas9 activity and genetic perturbations in transduced cells Fig. 4A shows the Tndel% by ICE analysis, following electroporation of RNPs, using single guides and two guides in C2C12 cell line, primary microglia and primary astrocytes, showing higher than expected indel% with two guides. Fig. 4B shows the 91% indels by ICE analysis in BMDMs from Cas9 mice electroporated with a single sgRNA, confirming Cas9 presence and activity. Fig. 4C displays the Crispresso2 resultsshowing good alignment of reads from control cells, and absence of indels. Fig. 4D shows the percentage of indels from TdTomato+ cells (with percentage of indels from WT reads subtracted). Fig. 4E shows the mutation profde shown for two guide RNAs, with histograms showing the distribution of indels, based on >30,000 aligned reads. No indels were detected in the corresponding un-injected control.
[0033] Fig. 5 depicts the analysis of perturbations in oligodendrocytes from 10X scRNA- Seq. Fig. 5A shows the UMAP representation of oligodendrocytes sub-clustering, showing division into several different clusters, and recovery of different oligodendrocyte sub-types: M0L1, M0L2, MOL5 / 6, DA lfn, with the majority being MOL5 / 6. Fig. 5B shows the expression of module scores of oligodendrocytes sub-type markers (M0L1 : Fos, EgrE, MOL5 / 6: Ptgds, Opalm,' M0L2: Klk6, Hopx,' DA Ifn: Ifltl, Ifit3, Stall, Ir ). Fig. 5C shows the UMAP representation of oligodendrocytes, showing cells with any detected gRNA barcodes in black and cells with no detected barcodes in grey. Fig. 5D shows the percentage of cells with detected gRNA barcodes for each oligodendrocyte sub-type. Each point represents a batch. Graph depicts mean and SEM. Fig. 5E shows the violin plots showing myelination gene scores (average expression level of Mobp, Mog, Opalin, Plpl, Mbp, Cnp, Mag, Mai) for 6 different perturbations compared to control, showing a reduced myelination program in RragaA&O and FlcnAAO. Fig. 5F shows the violin plots showing myelination gene scores (average expression level of Srebfl, Hmgcr, Scdl, Scd.2, Acaca, Dgatl, Cptla, Elovl.6) for 6 different perturbations compared to control, showing a reduced lipid metabolism program in Fasn- O. NS, not significant; * p < 0.05; ** p < 0.01.
[0034] Fig. 6 depicts the detection of all 18 injected exogenous barcodes in the Stereo-seq workflow. Fig. 6A shows the FOV of a subsection of Stereo-seq data, highlighting individual cells, with cells negative for barcodes in grey, positive cells in black. The vertical axis (i.e., y- axis) and horizontal axis (i.e., x-axis) both represent spatial coordinates in the tissue sample. Fig. 6B shows the detection of exogenous barcodes through PCR amplification of cDNA synthesized 51 from RNA captured from chips 1, 2, and 3. All 18 barcodes were successfully detected, demonstrating effective RNA capture of exogenous sequences. Fig. 6C shows the expression levels (counts) of all 18 barcodes for each chip, following the whole Stereo-seq workflow.
[0035] Fig. 7 depicts the spatially-resolved transcriptomics of the mouse brain. Fig. 7A shows the molecular identifier (MID) counts map and gene counts map following cellsegmentation using the DeepCell model, for all 3 chips. Chip 1 contains 2 sections from the same mouse. Chip 2 contains 3 sections from 2 different mice. Chip 3 contains 2 sections from the same mouse. Fig. 7B shows the entropy and scatteredness score of each annotated cell type using Stereopy’s Cell Community Detection function. Fig. 7C shows the sequencing saturation curves, for all 3 chips. Fig. 7D shows the cell segmentation statistics using three different segmentation algorithms. Deepcell model was used for analysis. Fig. 7E shows the spatial hotspot analysis reveals nine gene expression modules that were consistent with both brain sections. Spatial distribution of module scores for each of the modules. Fig. 7F shows the heatmap depicting gene expression profiles of the nine modules. Fig. 7D- Fig. 7F: Data from chip 1.
[0036] Fig. 8 shows the parameter sweep of used for BANKSY embedding.
[0037] Fig. 9 shows the spread of perturbations in a cell niche from a BANKSY cluster. The data showed that 62% of domains contained no perturbations and 27% contained 1 perturbation, which demonstrates that sparse gene editing was achieved.
[0038] Fig. 10 shows that the majority of perturbed cells had no perturbed neighbours as depicted in the frequency histogram of barcode-positive cells in chip 1, chip 2, and chip 3, showing the number of perturbed neighbours, with 5, 10, 15, or 20 neighbours called (“k geom” = 5, 10, 15, or 20).
[0039] Fig. 11 depicts the power analysis with powsimR. Fig. 11A shows the quality control metrics showing sequencing depth (left), library size factors (middle) and detected genes (right). Black line denotes the median Fig. 11B shows the marginal distribution of gene mean, gene dispersion and the dropout rate. Fig. 11C shows the number of genes and samples (cells) provided for the modeling. Detected: The number of genes and samples with one or more count. All: The number of genes for which the mean, dispersion, and dropout rate could be estimated, excluding outliers. Filtered: The number of genes above the filtering threshold, for which mean, dispersion, and dropout rate could be estimated, excluding outliers. Dropout Genes: The number of genes excluded. Fig. 11D shows the local polynomial regression showing the relationship between the mean and dispersion with variability indicated by the grey band. The common dispersion estimate was represented by the black dashed line. Fig. HE shows the proportion of dropouts against the estimated mean expression for each gene. Fig. HF shows the marginal error rates (False Discovery Rate(FDR) and True Positive Rate(TPR)) per sample size. Fig. 11G shows the marginal FDR and TPR, with dotted line indicating nominal alphalevel (type I error) at 0.1, and nominal 1-beta level (type II 85 error) at 0.8. Fig. 11H shows the error rates stratified by dispersion. Conditional FDR and TPR per sample size per stratum. Fig. Ill shows the number of equally (EE) and differentially expressed (DE) genes per stratum. For Fig. IIF-Fig. Ill, modeling was performed with Ifc = 2 for DE genes.
[0040] Fig. 12 depicts the compatibility of Spatial Perturb-Seq with FISH Fig. 12A depicts the percentage of mitochondria reads in cells positive and negative for barcodes, showing barcodes were tolerated in cells and did not cause major deleterious effects Fig. 12B depicts the confocal images of the same coronal mouse brain section showing FISH staining of barcode 22 (top), TdTomato protein (middle), DAPI (bottom). Fig. 12C depicts confocal images of the same coronal mouse brain section showing FISH staining of TdTomato, Actb, TdTomato protein and DAPI Fig. 12D depicts the correlation scatter plots of TdTomato fluorescent intensity against number of FISH barcode spots, FISH TdTomato spots and FISH Actb spots, showing correlation with barcode FISH and TdTomato FISH, but not Actb.
[0041] Fig. 13 depicts the analysis of cell-autonomous and non-cell autonomous effects of perturbations and cell-cell communication in hippocampal neurons. Fig. 13A depicts a schematic showing two different comparisons that can be made - cell autonomous and noncell autonomous, i.e. 1) transcriptomes of perturbed cells (own). 2) transcriptomes of the wildtype neighbours of perturbed cells (neighbours). Fig. 13B depicts Top: Number of differentially-expressed (DE) genes (Ifc > 0.5, pval < 0 05) for each perturbation (own, neighbours) compared to the control group. Bottom: Average Ifc for the DE genes. Fig. 13C depicts the Dot Plot showing top DE genes for each of perturbation, compared to all other perturbations. Top: Own transcriptome. Bottom: Neighbour cells transcriptome. Fig. 13D depicts the volcano plots showing DE genes (Ifc > 0.5, fdr < 0.05) for Cfap410-KO neighbours (left-most), Lrrk-KO and Srf-KO own and neighbour cells. Only Ifc > 0.5 genes are shown. Downregulated genes are left of the centre (i.e., Iog2 fold change negative) and upregulated genes are right of the centre (i.e., Iog2 fold change positive). Genes with no change in expression have a log2 fold change of approximately = 0. Fig. 13E depicts the heatmap visualizing ligand-receptor expression for top 20 differently-expressed ligand-receptor pairs, compare communication scores between Lrrk2-KO or Srf-KO neurons and their wildtype neighbours communication between control neurons and their wildtype neighbours. Fig. 13F shows the spatial plot of each source (ligand-expressing) cell (marked with x) and their 15 wildtype neighbour (receptor-expressing) cells, showing expression of Epha4 ligand inperturbed cells (marked in x) and Efnb3 receptor in neighbour cells. Top: Srf-KO. Botom: control.
[0042] Fig. 14 depicts Spatial Perturb-Seq with probe-based Xenium. Fig. 14A depicts the transcript density map showing transcripts per bin (bin size: 20pm). Fig. 14B shows the zoomed in area showing cell segmentation boundaries and individual barcode molecules Three molecules plotted: Barcodel-mSafe. Fig. 14C shows the spatial plot showing TdTomato molecules in expected positions in the neuronal region of the injected hippocampus. Fig. 14D depicts that similar proportions of cell types were recovered for each slide, with the majority cell type being neurons as expected. Fig. 14E shows UMAP representation showing the expression of canonical cell-type markers (Excitatory neurons: Slcl7a6, Neurod6; Inhibitory neurons: Sst, Gadl; Microglia: Trem2; Oligodendrocytes: Opalin; OPCs: Pdgfra; Astrocytes: Gfap; Endothelial cells: Pecaml; Fibroblasts: Den). Fig. 14F depicts that each dot is a cell, showing average expression of all transcripts (left) and gRNA target (right) against barcode expression (counts). X and Y axis are scaled where 100% represents the highest expressing cell as shown in Fig. 16. Fig. 14G depicts a Dot Plot showing DE genes for each of perturbation, compared to all other perturbations. Not all perturbations had detectable DEGs with the gene panel used. Each box highlights DEGs specific to the perturbation indicated by the column. Fig. 14H depicts the Violin plots showing expression levels (counts) of selected transcripts in the 15 closest neighbours of cells with the indicated perturbation. NS, not significant; * p < 0.05; ** p < 0.01; *** p < 0.001. Scale bar represents 1mm in Fig. 14A and Fig. 14C, and 50um in Fig. 14B.|0043| Fig. 15 depicts the detection of barcodes in the mouse brain by Xenium. Fig. 15A shows the expression levels (counts) of all 18 barcodes for each Xenium slide, following cell segmentation. Fig. 15B shows the UMAP projection labelled by annotated cell type, showing recovery of excitatory neurons, inhibitory neurons, astrocytes, microglia, oligodendrocytes, OPCs, endothelial cells and fibroblasts. Fig. 15C shows the UMAP projection colored by barcode-positive or barcode-negative status (positive: count >1). Fig. 15D shows the percentage of barcode-positive cells for each recovered cell type. Each dot represents a slide. Fig. 15E shows the Dot plot depicting percentage of each gRNA-associated barcode in neurons.
[0044] Fig. 16 shows that the target gene expression in neurons decreases with increasing expression of associated barcodes, while the average global transcript levels are not affected.
[0045] Fig. 17 shows that the neurons in DG and CAI regions of the hippocampus show overlapping transcriptional responses to AAV transduction. Fig. 17A shows the Xenium spatial plot showing Proxl (black) and Neurod6 (grey) molecules, markers of DG and CA neurons respectively. DG and CAI regions are manually circled and highlighted. Fig. 17B shows the violin plots of Proxl and Neurod6 expression levels following manual sub-setting of DG and CAI neurons, like in Fig. 17A, showing Proxl is being enriched in DG neurons, and Neurod6 in CAI neurons. Fig. 17C shows the volcano plots comparing barcode-positive (AAV- transduced) and barcode-negative neurons in both DG and CAI regions. Most DEGs overlap between DG and CAI neurons, but there are some region-specific differences. DEGs: p-value < 0.05 and log2 fold change > 0.5 (upregulated) or < -0.5 (downregulated). Fig. 17D shows the Venn diagrams showing the overlap of upregulated and downregulated genes in CAI and DG neurons, demonstrating a substantial overlap.
[0046] Fig. 18 depicts the verification and detection of full-length gRNA expression cassette. Fig. 18A shows the whole plasmid sequencing of the gRNA expression vector using long-read sequencing followed by assembly and automatic annotation, showing the expected full-length construct with no evidence of major rearrangements. The mSafe plasmid is shown in Fig. ISA. Fig. 18B shows the gel image of PCR products from lyzed 293T cells transduced with AAVPhP.eB-gRNAs which showed a distinct -1313 bp band corresponding to the full- length sgRNA cassette This band was absent in cells transduced with AAVPhP eB-GFP and in non-transduced cells. This indicates that the full-length construct is present post-AAV packaging and transduction, suggesting that recombination is not a significant issue under these conditions. Fig. 18C shows the densitometry analysis of the gel image showing 73.7% full- length product (1313 bp), 16.8% expected size 894 bp product, and 9.5% 475 bp product. Densities are normalized to GFP and non-transduced lanes, followed by normalization by length.
[0047] Fig. 19 shows the comparison between Perturb-seq and Spatial Perturb-Seq.DETAILED DESCRIPTION OF THE PRESENT INVENTION
[0048] In a first aspect, provided herein is a polynucleotide comprising: a) a polynucleotide sequence complementary to a polynucleotide sequence encoding a gene of interest or a segment thereof; and b) a barcode that is unique to the gene of interest comprising a polynucleotide sequence, wherein i. the polynucleotide sequence of the barcode is between about 20 to about500 nucleotides, ii. wherein the barcode comprises an edit distance of more than 16 nucleotides between segments, iii. wherein the barcode is devoid of a polynucleotide sequence encoding a stop codon, and iv. wherein the barcode is devoid of a polynucleotide sequence comprising 4 or more identical consecutive nucleotides.
[0049] In one example, the polynucleotide further comprises a polynucleotide sequence of a poly- A tail.
[0050] The polynucleotide may further comprise a detectable label to enable the location of insertion of the polynucleotide to be determined. In one example, the polynucleotide further comprises a polynucleotide sequence of a detectable label, such as a fluorescence tag. Examples of fluorescence tags include but are not limited to mCherry, mRbuy2, GFP and RFP. In a specific example, the fluorescence tag is TdTomato. It will generally be understood that the detectable label allows visualisation of the location of insertion of the polynucleotide into the cell and the polynucleotide would also function without the detectable label.
[0051] The complementary polynucleotide sequences of the polynucleotide of the invention may comprise polynucleotide sequences that are capable of modifying or perturbing the cell or gene of interest. Examples of such polynucleotide sequences include but are not limited to open reading frame (ORF) of a gene of interest, modified, engineered or synthetic genes, small interference RNA (siRNA), anti-sense oligonucleotide (ASO), polynucleotide sequences that may be folded into three dimensional structures or guide RNA.
[0052] In one example, the complementary polynucleotide sequence of the polynucleotide may comprise one or more guide RNA (gRNA). For example, the complementary polynucleotide sequence may comprise one, two, three, four, five or more gRNAs targeted to a gene of interest. It will generally be understood that a gRNA that is targeted to a gene of interest is able to hybridize to the gene of interest by complementary binding. The one or more gRNA may be complementary to a polynucleotide sequence encoding the gene of interest, or may be complementary to a segment of a polynucleotide sequence encoding the gene of interest. The one or more gRNA may be complementary to the same or different polynucleotide sequences or segments of the gene of interest.
[0053] Where the polynucleotide comprises a polynucleotide sequence that comprises more than one gRNA, the gRNAs may be located in tandem or adjacent to each other. For example, where the polynucleotide comprises three gRNAs arranged in tandem, this would be understood to mean that the polynucleotide sequence of one gRNA is located adjacent to thepolynucleotide sequence of a second gRNA and the polynucleotide sequence of a third gRNA is located adjacent to the polynucleotide sequence of the second gRNA. The polynucleotide sequences may be arranged in tandem immediately adjacent to each other, that is to say there are no additional nucleotides or spacer sequences between the polynucleotide sequences. The polynucleotides sequences may also be arranged in tandem with additional non-coding nucleotides or non-coding spacer sequences between the polynucleotide sequences.
[0054] Where the polynucleotide comprises a polynucleotide sequence that comprises more than one gRNA, the gRNAs may be located not in tandem with each other. This would be understood to mean that the gRNAs are separated by one or more polynucleotide sequences that encode one or more genes, such as a detectable label. For example, where the polynucleotide comprises three gRNAs that are not arranged in tandem with each other, the polynucleotide sequence of a detectable label such as a fluorescence tag is between the polynucleotide sequences of a first gRNA and a second gRNA, and the polynucleotide sequence of a barcode is between the polynucleotide sequences of a second gRNA and a third gRNA.
[0055] In one example, the polynucleotide may comprise three gRNAs targeted to the gene of interest. In one example, the three gRNAs targeted to the gene of interest is distinct.
[0056] In one example, the gRNA comprises the polynucleotide sequences selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3 , SEQ ID NO: 5 , SEQ ID NO: 6, SEQ ID NO: 7 SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11 , SEQ ID NO: 13 , SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 17 SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 21 , SEQ ID NO: 22 SEQ ID NO: 23 , SEQ ID NO: 25 , SEQ ID NO: 26, SEQ ID NO: 27 SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 41 , SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 45 , SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 53 , SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59 SEQ ID NO: 61 , SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 73, SEQ ID NO: 74 and SEQ ID NO: 75.
[0057] The polynucleotide sequence of the one or more gRNA of the polynucleotide of the invention is targeted to the gene of interest by complementary base pairing which allows the gRNA to bind or hybridize to the gene of interest. The one or more gRNAs may becomplementary to different sections of the gene of interest or segment of the gene of interest thereof. It would be understood that a certain percentage complementarity is acceptable as long as the gRNA is able to bind to the target gene of interest. In some examples, each of the one or more gRNA in the polynucleotide may be about 90% to 100% complementary to the polynucleotide sequence encoding the gene of interest to minimize mismatching to the gene of interest. In some examples, each of the one or more gRNA in the polynucleotide may be 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% complementary to the polynucleotide sequence encoding the gene of interest.
[0058] In one example, where the polynucleotide comprises the polynucleotide sequences of three gRNAs, the polynucleotide sequence of a first gRNA in the polynucleotide may be 95% complementary to the polynucleotide sequence encoding the gene of interest, the polynucleotide sequence of a second gRNA in the polynucleotide may be 90% complementary to the polynucleotide sequence encoding the gene of interest and the polynucleotide sequence of a third gRNA in the polynucleotide may be 100% complementary to the polynucleotide sequence encoding the gene of interest. In another example, wherein the polynucleotide comprises the polynucleotide sequence of four gRNAs, the polynucleotide sequence of a first gRNA in the polynucleotide may be 100% complementary to the polynucleotide sequence encoding the gene of interest, the polynucleotide sequence of a second gRNA in the polynucleotide may be 90% complementary to the polynucleotide sequence encoding the gene of interest, the polynucleotide sequence of a third gRNA in the polynucleotide may be 95% complementary to the polynucleotide sequence encoding the gene of interest and the polynucleotide sequence of a fourth gRNA in the polynucleotide may be 94% complementary to the polynucleotide sequence encoding the gene of interest. It is to be understood that these examples are not exhaustive and are meant to illustrate that the complementarity of the polynucleotide sequence of each gRNA in the polynucleotide to the polynucleotide sequence encoding the gene of interest may be different or identical as long as the complementarity of the polynucleotide sequence of each gRNA in the polynucleotide to the polynucleotide sequence encoding the gene of interest is between 90%-100%.
[0059] In one example, the one or more gRNA is a CRISPR-gRNA. It should be generally understood that the CRISPR-gRNA is configured to mediate site-specific editing of a gene in a subject that expresses a Cas9 protein. The expression of Cas9 protein may be constitutive or inducible.
[0060] The polynucleotide sequence comprising the one or more gRNAs in the polynucleotide of the invention may be operably linked to one or more promoters. In some examples, the one or more gRNA in the polynucleotide is operably linked to one promoter. In other examples, where there are more than one gRNA in the polynucleotide, the polynucleotide of each of the gRNA in the polynucleotide is operably linked to one promoter, which may be the same promoter or may be different promoters. In yet another example, the first and second gRNA may be operably linked to a single promoter, and the third gRNA may be operably linked to another promoter, which may be the same promoter or may be a different promoter. In one example, the one or more promoters is an RNA polymerase promoter. In one example, the RNA polymerase is an RNA polymerase III promoter. Examples of the promoters may include but are not limited to the U6 promoter and U3 promoter. The promoters may be a mouse or a human promoter or a combination thereof. In a specific example, the RNA polymerase promoter is a mouse U6 promoter (mU6) promoter. For example, where there are three gRNAs in the polynucleotide, each of the three gRNAs may be linked to a U6 promoter or all three gRNAs may be linked to a single U6 promoter. In another example, the first gRNA and the second gRNAs may be linked to a U6 promoter and the third gRNA may be linked to a U3 promoter. It is to be understood that these examples are not exhaustive and are meant to illustrate that the gRNA may be operably linked to the same or different promoters.
[0061] The one or more gRNA targeting a specific gene of interest is tagged to a barcode that is unique to that gene of interest. The gRNAs that are targeted to a specific gene of interest are tagged to a unique barcode, while gRNAs that are targeted to another gene of interest are tagged to a different unique barcode. In some examples, more than one gRNA that are targeted to a specific gene of interest is tagged to a single barcode that is unique to said specific gene of interest. In other examples, each gRNA that is targeted to a specific gene of interest is tagged to a single barcode that is unique to said specific gene of interest.
[0062] The barcode of the present invention is designed to have several features. The polynucleotide sequence of the barcode is designed to comprise a sequence that is distinct from endogenous human and mouse sequences. It would be understood that in the context of this disclosure, the term “endogenous sequence” refers to a polynucleotide sequence that is native to the tissue sample e.g., a human tissue sample or a mouse tissue sample. The design of the barcode minimizes or prevents alignment with endogenous human and mouse transcripts, ambiguous detection via probe-based spatial technologies and sequencing based methods, andunwanted biological effects by having no genetic function. In some examples, the polynucleotide sequence of the barcode or each segment of the barcode has an edit distance of more than 16 nucleotides relative to the endogenous human and mouse sequences. In one example, the polynucleotide sequence of the barcode has an edit distance of about 17 nucleotides to about 500 nucleotides relative to the endogenous human and mouse sequences. The polynucleotide sequence of each segment of the barcode has an edit distance of about 17 to about 100 nucleotides relative to the endogenous human and mouse sequences. In some examples, the polynucleotide sequence of the barcode is about 35% or less similar to endogenous human and mouse sequences. In one example, the polynucleotide sequence of the barcode or each segment of the barcode may comprise a sequence identity of about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 33%, about 34% or about 35% to endogenous human and mouse sequences. In one example, the polynucleotide sequence is about 33% to endogenous human and mouse sequences.
[0063] Another feature of the barcode is that the number of nucleotides in the polynucleotide sequence of the barcode does not induce deleterious effects in the cell or population of cells that the polynucleotides are delivered to. The deleterious effects in the cell or population of cells may be determined by the sequencing reads from the mitochondrial genome. Sequencing reads from the mitochondrial genome provides an assessment of cell quality and the percentage of mitochondrial reads relative to total reads of less than 20% indicates that the cell or the population of cells is healthy. In one example, the polynucleotide sequence of the barcode is between about 20 to about 500 nucleotides. The polynucleotide sequence of the barcode may be about 20 nucleotides, about 50 nucleotides, about 100 nucleotides, about 150 nucleotides, about 200 nucleotides, about 250 nucleotides, about 300 nucleotides, about 350 nucleotides, about 400 nucleotides, about 450 nucleotides and about 500 nucleotides. In a preferred example, the polynucleotide sequence of the barcode is about 450 nucleotides.
[0064] Another feature of the barcode is that the polynucleotide sequence of the barcode may be segmented and has an edit distance of more than 16 nucleotides between each segment of the barcode. In some examples, the barcode may comprise more than 1 segment. The barcode may comprise 2 segments, 3 segments, 4 segments, 5 segments, 6 segments, 7 segments, 8 segments, 9 segments and 10 segments. In one example, the barcode may comprise 9 segments. Each segment may comprise about 20 to about 100 nucleotides. In a preferred example, eachsegment comprises about 50 nucleotides. In some examples, the polynucleotide sequence of the barcode has an edit distance of about 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides and 50 nucleotides between each segment. Where the barcode comprises 450 nucleotides, the barcode is engineered by combining the polynucleotide sequences of 9 segments with each segment comprising 50 nucleotides and each segment has an edit distance of more than 16. This allows each segment of the barcode to be detected individually and unambiguously for the differentiation of the perturbed cells. Each segment of the barcode may bind to one or more additional encoding probes in downstream assays such as fluorescent in situ hybridization (FISH), CosMx, Multiplexed Error -Robust Fluorescence In Situ Hybridization (MERFISH), and Xenium. Without being bound by theory, a higher edit distance between each segment of the barcode allows binding of the barcode to multiple additional encoding probes which in turn allows signal amplification while reducing cross-reactivity. In some examples, the polynucleotide sequence of each encoding probe may be about 20 to 100 nucleotides. The polynucleotide sequence of each encoding probe may be about 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 52 nucleotides, 54 nucleotides, 56 nucleotides, 58 nucleotides, 60 nucleotides, 65 nucleotides, 70 nucleotides, 75 nucleotides, 80 nucleotides, 85 nucleotides, 90 nucleotides or 100 nucleotides. In one example, the polynucleotide sequence of each encoding probe is about 52 nucleotides. The polynucleotide sequence of each encoding probe may comprise more than one polynucleotide sequence that binds to a read-out probe and a polynucleotide sequence of a segment that is complementary to the barcode. In one example, the polynucleotide sequence of each encoding probe comprises a first polynucleotide sequence that binds to the read-out probe at the 3’ end, a second polynucleotide sequence that binds to the read-out probe at the 5’ end and the polynucleotide sequence of the segment between the first and the second polynucleotide sequences. In some examples, each of the polynucleotide sequences that binds to the read-out probe is about 15-20 nucleotides and the polynucleotide sequence of the segment that is complementary to the barcode is about 10-30 nucleotides. Tn a preferred example, each of the polynucleotide sequences that binds to the read-out probe is about 16 nucleotides and the polynucleotide sequence of the segment that is complementary to the barcode is about 20 nucleotides. The polynucleotide sequence of each encoding probe may be about 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides or 30 nucleotides complementary to thepolynucleotide sequence of the barcode. In a preferred example, the polynucleotide sequence of each encoding probe is about 20 nucleotides complementary to the polynucleotide sequence of the barcode.
[0065] Yet another feature of the barcode is that it is devoid of a polynucleotide sequence encoding a stop codon to minimize or eliminate nonsense mediated decay. It would generally understood that a stop codon is a sequence of three nucleotides that signals the termination of protein synthesis. The polynucleotide sequences of a stop codon include but are not limited to TAG, TAA, TGA, UAA, UAG, and UGA.
[0066] Yet another feature of the barcode is that it is devoid of a polynucleotide sequence comprising 4 or more identical consecutive nucleotides. The barcode may be devoid of 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 or 60 identical consecutive nucleotides. Examples of the sequence include but are limited to AAAA, TTTT, CCCC, GGGG, AAAAAA, TTTTTTTT, CCCCCCC, and GGGGGGGG.
[0067] In one example, the barcode comprises the polynucleotide sequences selected from the group consisting of SEQ ID NO: 8, SEQ ID NO: 12, SEQ ID NO: 16, SEQ ID NO: 20, SEQ ID NO: 24, SEQ ID NO: 28, SEQ ID NO: 32, SEQ ID NO: 36, SEQ ID NO: 40, SEQ ID NO: 44, SEQ ID NO: 48, SEQ ID NO: 52, SEQ ID NO: 56, SEQ ID NO: 60, SEQ ID NO: 64, SEQ ID NO: 68, SEQ ID NO: 72 and SEQ ID NO: 76.
[0068] It should be understood that the polynucleotide sequences of the one or more complementary polynucleotide sequence, the barcode, the poly-A tail and fluorescence tag in the polynucleotide of the invention may be arranged in various configurations relative to one another as long as the complementary polynucleotide sequence is able to bind to and perturb the gene of interest, and the barcode in the polynucleotide of the invention is able to be detected by probe- and / or sequencing- based methods. In one example, the polynucleotide sequence of the poly-A tail located at the 5’ of the polynucleotide is followed by the polynucleotide sequence of the barcode, the polynucleotide sequence of the barcode is followed by the polynucleotide sequence encoding the one or more gRNA and the polynucleotide sequence encoding the one or more gRNA is followed by the polynucleotide sequence of the fluorescence tag located at the 3’ end of the polynucleotide. In another example, the polynucleotide sequence of the barcode located at the 5’ of the polynucleotide is followed by the polynucleotide sequence encoding the one or more gRNA, the polynucleotide sequence encoding the one or more gRNA is followed by the polynucleotide sequence of thefluorescence tag and the polynucleotide sequence of the fluorescence tag is followed by the polynucleotide sequence of the poly-A tail located at the 3’ end of the polynucleotide. In another example, the polynucleotide sequence of the fluorescence tag at the 5’ end of the polynucleotide is followed by the polynucleotide sequence encoding the one or more gRNA, the polynucleotide sequence encoding the one or more gRNA is followed by the polynucleotide sequence of the barcode, the polynucleotide sequence of the barcode is followed by the nucleotide sequence of the poly-A tail at the 3’ end of the polynucleotide In a preferred example, the polynucleotide sequence encoding the one or more gRNA at the 5’ of the polynucleotide is followed by the polynucleotide sequence of the fluorescence tag, the polynucleotide sequence of the fluorescence tag is followed by the polynucleotide sequence of the barcode, the polynucleotide sequence of the barcode is followed by the nucleotide sequence of the poly-A tail at the 3’ end of the polynucleotide.
[0069] In one example, in the polynucleotide of the present invention, the polynucleotide sequence encoding three gRNA sequences located at 5’ end is followed by the polynucleotide sequence of a fluorescence tag, the polynucleotide sequence of a fluorescence tag is followed by the polynucleotide sequence of the barcode, and the polynucleotide sequence of the barcode is followed by the polynucleotide sequence of the poly-A tail at 3’ end, wherein the polynucleotide sequence of the barcode is about 450 nucleotides, wherein the barcode comprises an edit distance of more than 16 nucleotides between each segment of 50 nucleotides, wherein the barcode is devoid of a polynucleotide sequence encoding a stop codon, and wherein the barcode is devoid of a polynucleotide sequence comprising 4 or more identical consecutive nucleotides.
[0070] In another aspect, provided herein is a library comprising a plurality of polynucleotides as described herein.
[0071] The library may comprise any number of polynucleotides that allow genetic perturbation as low as one cell of a population of cells, a tissue sample, an organ or a subject, and as high as all the cells in a population of cells, a tissue sample, an organ or a subject. It should be understood the number of polynucleotides in the library is dependent on the number of cells in a pre-determined region of a tissue sample or an organ or the size of a tissue sample or an organ. In some examples, the library may comprise any number of polynucleotides that allow genetic perturbation in approximately about 10% or less of a population of cells. The number of polynucleotides in the library may perturb a gene of interest in about 1%, about 2%,about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9% and about 10% of the population of cells. For example, the library may comprise more than 2 polynucleotides. The library may comprise 2 polynucleotides, 3 polynucleotides, 4 polynucleotides, 5 polynucleotides, 6 polynucleotides, 7 polynucleotides, 8 polynucleotides, 9 polynucleotides,10 polynucleotides, 11 polynucleotides, 12 polynucleotides, 13 polynucleotides, 14 polynucleotides, 15 polynucleotides, 16 polynucleotides, 17 polynucleotides, 18 polynucleotides, 19 polynucleotides, 20 polynucleotides, 21 polynucleotides, 22 polynucleotides, 23 polynucleotides, 24 polynucleotides, 25 polynucleotides, 26 polynucleotides, 27 polynucleotides, 28 polynucleotides, 29 polynucleotides, 30 polynucleotides, 35 polynucleotides, 40 polynucleotides, 45 polynucleotides, 50 polynucleotides, 100 polynucleotides, 200 polynucleotides, 500 polynucleotides, 1000 polynucleotides, 1500 polynucleotides, 2000 polynucleotides, 4000 polynucleotides, 6000 polynucleotides, 8000 polynucleotides or 10000 polynucleotides. In a preferred example, the library comprises 18 polynucleotides In some examples, the polynucleotides in the library may also allow perturbation of any genomic sequence in a cell of the population of cells, including genetic sequences and intergenic sequences. Some examples of intergenic sequences include enhancer sequences / regions and promoter sequences / regions.
[0072] It would be understood that the library comprises at least one polynucleotide comprising a polynucleotide sequence complementary to a polynucleotide sequence encoding a control gene and at least one polynucleotide comprising a polynucleotide sequence complementary to a polynucleotide sequence encoding a gene of interest. Tn one example, the library may comprise a total of seven polynucleotides, wherein two polynucleotides comprise polynucleotide sequences that are complementary to polynucleotide sequences encoding two separate control genes, and wherein the five remaining polynucleotides comprise polynucleotide sequences complementary to polynucleotide sequences encoding five separate genes of interest. In one example, the library comprises a total of ten polynucleotides, wherein one polynucleotide comprises a polynucleotide sequence that is complementary to a control gene, and wherein the nine remaining polynucleotides comprise polynucleotide sequences that are complementary to polynucleotide sequences encoding nine separate genes of interest. In yet another example, the library comprises a total of one thousand polynucleotides, wherein one polynucleotide comprises a polynucleotide sequence that is complementary to a control gene, and wherein the nine hundred and ninety-nine remaining polynucleotides comprisepolynucleotide sequences that are complementary to polynucleotide sequences encoding nine hundred and ninety-nine separate genes of interest. It is to be understood that the above examples are not exhaustive and are merely meant to illustrate that the library may comprise a plurality of polynucleotides in various combinations.
[0073] In one example, the library may comprise a total of eighteen polynucleotides, wherein one polynucleotide comprises a polynucleotide sequence complementary to a polynucleotide sequence encoding a control gene, and wherein the remaining seventeen polynucleotides comprise polynucleotide sequences that are complementary to the polynucleotide sequences of seventeen separate genes of interest.
[0074] The edit distance between each barcode of the plurality of the polynucleotides may be determined by the number of nucleotides or the percentage of nucleotides. Each of the plurality of the polynucleotides in the library has a barcode comprising more than one segment with each segment having an edit distance of more than 16 nucleotides relative to each segment of the barcode of each of the other polynucleotides in the library. The polynucleotide sequence of each segment of the barcode in each of the plurality of polynucleotides may have an edit distance of about 17 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides and 50 nucleotides. For example, wherein there are eighteen polynucleotides in the library and the polynucleotide sequence of each segment of the barcode in each polynucleotide has an edit distance of more than 16 nucleotides, the difference in the number of nucleotides between each segment of the barcode of each polynucleotide is more than 16 nucleotides or the number of nucleotides that is required to edit a segment of a first barcode to a segment of a second barcode is more than 16. Each of the plurality of the polynucleotides in the library has a barcode comprising more than one segment with each segment having a sequence identity that is 34% or less similar relative to each segment of the barcode in each of other polynucleotides in the library. In some examples, each segment of the barcode in each of the polynucleotides in the library may comprise a sequence identity that may be about 2%, 5%, 10%, 15%, 20%, 25%, 30% or 34% similar to each segment of the barcode in each of the other polynucleotides in the library. It should be generally understood that the higher the edit distance or the lesser the similarity of the polynucleotide sequences of the polynucleotides, the likelihood of each barcode being able to hybridize to one or additional encoding probes used in downstream assays such as optical assays is decreased.
[0075] It would generally be understood that the genes targeted by one or more gRNA may be any gene of interest. In one example, the gene of interest may be a control gene. The control gene includes but is not limited to mSafe, Fasti, Flcn, Gfap, Olig2, Rraga and Srf In one example, the genes of interest are associated with diseases, including but not limited to neurological diseases. Examples of neurological diseases are Alzheimer’s Disease, Amyotrophic Lateral Sclerosis and Parkinson’s Disease. The genes of interest include but are not limited to CLU, LRRK2 NDUFAF2, RBFOX1, TREM2, C9oif72, CFAP410, DPP6, TBK1, SH3GL2, STK39.
[0076] The genes of interest may be part of any cellular pathway or intracellular network that is associated with diseases. Examples of a pathway or intracellular network may comprise but not limited to Wnt7a-Frizzled-LRP5 / 6 and WNT / p-catenin pathways.
[0077] In another aspect, the present invention refers to a method of obtaining spatial and transcriptomic information of a population of cells, wherein the method comprises: a) delivering a plurality of polynucleotides as described herein or the library as described herein to the population of cells, wherein each of the plurality of polynucleotides perturbs a distinct gene of interest in the population of cells; b) obtaining an intact tissue sample comprising the population of cells; and c) detecting the plurality of polynucleotides in the population of cells in the intact tissue sample, thereby obtaining spatial and transcriptomic information of the population of cells.
[0078] In one example, the plurality of polynucleotides or the library of polynucleotides is delivered via one or more viral vectors. It would be generally understood that the viral vectors may be transduced in different cell types. The cell types may comprise but not limited to neuronal and non-neuronal cell types. The neuronal cell types may comprise but not limited to pyramidal cells, granule cells interneurons. The non-neuronal cell types may comprise but are not limited to microglia, oligodendrocytes, astrocytes, endothelial cells and oligodendrocyte precursor cells (OPCs). Examples of a viral vector include but are not limited to an adeno- associated virus (AAV) vector, a lentivirus vector and vectors known in the art. In one example, the viral vector is an adeno-associated virus (AAV) vector.
[0079] The AAV vector may be one or more serotypes. Examples of the one or more serotypes include but are limited to AAV-PHP.eB, AAV2, AAV8 and AAV9. It would be understood that the serotypes of the AAV vector used for delivery to the population of cellsdepend on the location of the cells in the subject. For example, the AAV-PHP.eB is a serotype specific for the brain and the AAV8 vector is a serotype specific for the liver. In a preferred example, the serotype of AAV vector is AAV-PHP.eB. The library of polynucleotides delivered via one or more viral vectors may comprise one or more serotypes.
[0080] Each of the polynucleotides in the plurality of polynucleotides or library delivered to the viral vector targets in the methods of the invention perturbs a distinct gene of interest in the population of cells. In some examples, perturbation of the gene of interest in the population of cells comprises one or more of deletion of the gene, reduction in expression of the gene, introduction of the gene, deletion of one or more nucleotides in the gene of interest, substitution of one or more nucleotides in the gene of interest and insertion of one of more nucleotides in the gene of interest. In one example, perturbation of the gene of interest in the population of cells comprises substitution, deletion, or insertion of one or more nucleotides in the polynucleotide sequence of the gene of interest. Substitution, deletion, or insertion of one or nucleotides in the gene of interest may lead to a partial loss of genetic function or a complete loss of genetic function. In one example, perturbation of the gene of interest of the population of cells comprises deletion of about 50% to about 100% of the polynucleotide sequence of the gene of interest. Perturbation of the gene of interest may comprise deletion of about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, about 95% and about 100% of the polynucleotide sequence of the target gene. In a preferred example, perturbation of the gene of interest comprises substitution, deletion, or insertion of about 90% of the polynucleotide sequence of the gene of interest.
[0081] The plurality of polynucleotides or the library may be delivered to a population of cells in a pre-determined region of one or more organs of a subject. Examples of one or more organs may comprise but are not limited to the brain, liver, heart or lungs. In one example, the plurality of polynucleotides or the library may be delivered to cells in the hippocampus of the brain.
[0082] The one or more gRNA of each of the polynucleotides in the plurality of polynucleotides or the library of polynucleotides may be delivered via a viral vector to the population of cells in the methods of the invention at a low multiplicity of infection (MOI). The MOI may be an average of less than about five viral particles per cell, less than about four viral particles per cell, less than about three viral particles per cell, less than about two viral particles per cell and less than about one viral particle per cell. In a preferred example, the MOIis about less than about one viral particle per cell. In some examples, the plurality of polynucleotides or the library of polynucleotides is delivered to about 100% or less of the cells in the population of cells. In other examples, the plurality of polynucleotides or the library of polynucleotides is delivered to about 50% or less of the cells in the population of cells. In another example, the plurality of polynucleotides or the library of polynucleotides is delivered to at least one cell in the population of cells. The plurality of polynucleotides or the library of polynucleotides may be delivered to between about 5% to 100% or about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50% and about 100% of the cells in the population of cells. In one example, the plurality of polynucleotides or the library of polynucleotides is delivered to about 10% of the cells in the population of cells. For example, the population of cells comprises 10% of the cells having genetic perturbations and 90% of the cells having no genetic perturbation. In another example, the population of cells comprises 20% of the cells having genetic perturbations and 80% of the cells having no genetic perturbation
[0083] Each cell having a genetic perturbation may be surrounded by cells having no genetic perturbation, or may be surrounded by cells having no genetic perturbation and cells having genetic perturbation, or may be surrounded by cells having genetic perturbation. These surrounding (neighbouring) cells may be defined as any cell within a radius of between about 10 pm to 100 pm from a cell having a genetic perturbation. The surrounding (neighbouring) cells may also be defined as any cell within a radius of within 0 pm (i.e., immediately adjacent to) to about a few centimeters of a cell having a genetic perturbation. The surrounding (neighbouring) cells may also be defined as any number of cell(s) surrounding a cell having a genetic perturbation. It should be generally understood that the radius for defining the surrounding (neighbouring) cells or number of cells for defining surrounding (neighbouring) cells of a cell having a genetic perturbation can vary depending on factors such as tissue type, cell size, the gene and / or conditions being examined by the polynucleotide of the invention, etc. For example, the radius for defining surrounding (neighbouring) cells can be from about 1 pm to about 2 cm from a cell having a genetic perturbation In another example, the number of cells for defining surrounding (neighbouring) cells can range from 1 to as many as 20,000 cells. In one example, the cell having a genetic perturbation is surrounded by cells having no genetic perturbation. In another example, the cell having a genetic perturbation is surrounded by more than one cell having genetic perturbation. The cell having a genetic perturbation may besurrounded by 1, 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 cells having no genetic perturbation. In one example, the cell having a genetic perturbation is surrounded by 15 cells having no genetic perturbation.
[0084] The gene of interest in the population of cells comprises a polynucleotide sequence that is complementary to the polynucleotide sequence of the one or more gRNA in each of the plurality of polynucleotides or the library. In one example, the gene of interest comprises a polynucleotide sequence that is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% complementary to the polynucleotide sequence of the one or more gRNA in the plurality of polynucleotides or in the library. In a preferred example, the gene of interest comprises a polynucleotide sequence that is about 90% complementary to the polynucleotide sequence of the one or more gRNA in each of the plurality of polynucleotides or the library.
[0085] The plurality of polynucleotides may be detected optically or by sequencing, or both. Examples of optical methods may comprise but are not limited to fluorescence in situ hybridisation (FISH), multiplexed error-robust fluorescence in situ hybridization (MerFISH), Xenium, and CosMx. The Xenium platform utilizes one or more oligonucleotide probes that target a specific gene of interest, wherein the one or more oligonucleotide probe hybridizes to a complementary target RNA sequence in a tissue sample and the hybridized probes further bind to one or more oligonucleotides linked with one or more fluorescence tags (Marco Salas S, et al. Optimizing Xenium In Situ data utility by quality assessment and. best-practice analysis workflows,' Nature Methods. 2025;22(4):813-823). The one or more fluorescence signals allow detection, identification and localization of each gene. CosMx platform utilizes one or more oligonucleotide probes or antibodies conjugated with one or more fluorescence tags that bind to one or more RNA or protein target in a population of cells in a tissue sample, allowing visualization and mapping of the targets.
[0086] Examples of sequencing methods include but are not limited to single-cell RNA sequencing, Stereo-Seq, Visium, and GeoMx. Visium utilizes barcoded oligonucleotides arrays on glass slides, wherein the barcoded oligonucleotides bind to mRNA of a population of cells in a tissue sample for in situ reverse transcription. The cDNA obtained from the in situ reverse transcription is sequenced for quantification of expression and localization of one or more genes in a tissue section. GeoMx utilizes oligonucleotide barcodes and antibody or RNA probes, wherein each of the barcode is linked to an antibody RNA probe via a linker that maybe cleaved by a light source such as UV light, for the analysis of expression of one or more genes in a pre-determined region of the tissue sample.
[0087] The barcode in each polynucleotide in the plurality of polynucleotides or the library as described herein delivered to the population of cells in the method of the invention may be able to bind to one or more additional encoding probes for detection by FISH. Each encoding probe comprises a polynucleotide sequence located at one end or both ends that allows binding to a read-out probe. Detection of the read-out probe allows spatial and transcriptomic information to be obtained. In one example, the polynucleotide sequence of each encoding probe for detection by FISH is between about 20 nucleotides and about 100 nucleotides. The polynucleotide sequence of each encoding probe may comprise about 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 52 nucleotides, 54 nucleotides, 56 nucleotides, 58 nucleotides, 60 nucleotides, 65 nucleotides, 70 nucleotides, 75 nucleotides, 80 nucleotides, 85 nucleotides, 90 nucleotides, 95 nucleotides and 100 nucleotides. In one example, the polynucleotide sequence of each encoding probe is about 52 nucleotides. The polynucleotide sequence of each encoding probe may comprise more than one polynucleotide sequence that binds to a read-out probe and a polynucleotide sequence of a segment that is complementary to the barcode. In one example, the polynucleotide sequence of each encoding probe comprises a first polynucleotide sequence that binds to the read-out probe at the 3’ end, a second polynucleotide sequence that binds to the read-out probe at the 5’ end and the polynucleotide sequence of the segment between the first and the second polynucleotide sequences. In some examples, each of the polynucleotide sequences that binds to the read-out probe is about 15-20 nucleotides and the polynucleotide sequence of the segment that is complementary to the barcode is about 10-30 nucleotides. In a preferred example, the polynucleotide sequence of each of the polynucleotide sequences that binds to the read-out probe is about 16 nucleotides and the polynucleotide sequence of the segment that is complementary to the barcode is about 20 nucleotides. In one example, the polynucleotide sequence of each encoding probe is about 5 nucleotides to about 30 nucleotides complementary to the polynucleotide sequence of the barcode. The polynucleotide sequence of each encoding probe may be about 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides or 30 nucleotides complementary to the polynucleotide sequence of the barcode. In a preferred example, the polynucleotide sequence of each encoding probe is about 20 nucleotides complementary to the polynucleotide sequence of the barcode.
[0088] In one example, the polynucleotide sequence that binds to a read-out probe is SEQ ID NO: 2. In one example, the encoding probe comprises a polynucleotide sequence selected from the group consisting of SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84 and SEQ ID NO: 85.
[0089] The optical and sequencing detection of the plurality of polynucleotides may be performed on different tissue samples. For example, FISH is performed on a first tissue sample, such as a section of a tissue sample and Stereo-seq is performed on a second tissue sample, such as another section of a tissue sample, to obtain spatial and transcriptomic information. In one example, the first and second tissue samples are serial tissue sections. In another example, the first and second tissue samples are not serial tissue sections. This would be understood to mean that the first tissue sample may comprise one end of a specific region of the tissue and the second tissue sample may comprise another end of the same specific region of the tissue. For example, the first tissue sample may comprise a dorsal hippocampus of the brain and the second tissue sample may comprise a ventral hippocampus of the brain.
[0090] The spatial and transcriptomic information may be obtained simultaneously or asynchronously. In one example, the spatial and transcriptomic information may be obtained simultaneously.
[0091] In one example, the spatial information comprises location, arrangement and distance of one or more cells in the population of cells in the intact tissue sample. In one example, the transcriptomic information comprises gene expression of one or more cells in the population of cells in the intact tissue sample For example, the transcriptomic information comprises identification of cell types and measurements of gene expression in a single cell or a plurality of cells.
[0092] The spatial and transcriptomic information may be further analysed using algorithmbased modelling. Examples of algorithm-based modelling may comprise but are not limited to Building Aggregates with a Neighborhood Kernel and Spatial Yardstick (BANKSY) (Singhal, V., Chou, N., Lee, J. et al. BANKSY unifies cell typing and tissue domain segmentation for scalable spatial omics data analysis. Nat Genet 56, 431-441 (2024)) and Ligand-Receptor Analysis (LIANA) (Dimitrov, D., Tiirei, D., Garrido-Rodriguez, M. et al. Comparison of methods and resources for cell-cell communication inference from single-cell RNA-Seq data. Nat Commun 13, 3224 (2022)). For example, the BANKSY is a computational framework for modelling of transcriptome of the perturbed cell and the neighbouring cells to quantify thedifferentially expressed genes and the effect sizes between each group of perturbed cells relative to the control cells. BANKSY is used for 1) cell clustering 2) cell embedding; and 3) cell classification. Cell clustering refers to identification of the cell types and grouping of cells based on the gene expression, cell embedding refers to transformation of the gene expression data into graphical representation to identify one or more groups of cells and cell classification refers to labelling of cells to a cell type based on the gene expression to allow for the prediction of cell type of one or more unlabelled cells. BANKSY is performed by adjusting the lambda parameter for the analysis of cell clustering, cell embedding and cell classification. The lambda parameter may be 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or 1. In one example, the lambda parameter is 0.2. BANKSY may be used to analyse non-cell autonomous function of genes in the perturbed cell and the neighbouring cells. In one example, BANKSY may be used to analyse genetic perturbation on the gene expression and function in the neighbouring cells. In another example, BANKSY may be used to analyse the neighbouring cells having no genetic perturbation on gene expression and function in the cell having a genetic perturbation.
[0093] In another example, LIANA allows characterisation of ligand-receptor interactions in the population of cells using single-cell transcriptomic data. The LIANA framework integrates the gene expression data in a cell type with ligand-receptor databases to compute the SCA score for the identification and ranking of the ligand-receptor interaction relative to a specific signalling pathway. The SCA score refers to the quantification of the activity of one or more genes in a signalling pathway or an intracellular network.
[0094] In one example, the method as described herein is performed in vitro, in vivo or ex vivo. The plurality of polynucleotides as described herein may be delivered to an organ, organoid, cell culture, or subject. In some examples, the subject expresses a Cas protein.
[0095] In another aspect, provided herein is a use of the plurality of polynucleotides as described herein or library as described herein in the manufacture of a medicament for obtaining spatial and transcriptomic information of a population of cells, a) wherein the plurality of polynucleotides or the library is to be delivered to the population of cells, wherein each of the plurality of polynucleotides perturbs a distinct gene of interest in the population of cells, b) wherein an intact tissue sample comprising the population of cells is to be obtained and c) wherein the plurality of polynucleotides in the population of cells is to be detected in the intact tissue sample, thereby obtaining spatial and transcriptomic information of the population of cells.
[0096] In another aspect, provided herein is the plurality of polynucleotides as described herein or library as described herein for use in obtaining spatial and transcriptomic information of a population of cells, a) wherein the plurality of polynucleotides or the library is to be delivered to the population of cells, wherein each of the plurality of polynucleotides perturbs a distinct gene of interest in the population of cells, b) wherein an intact tissue sample comprising the population of cells is to be obtained and c) wherein the plurality of polynucleotides in the population of cells is to be detected in the intact tissue sample, thereby obtaining spatial and transcriptomic information of the population of cells.
[0097] In one example, the intact tissue sample in step b) is obtained from a subject. In one example, the subject is an animal. In one example, the animal is a mammal. Examples of animals may include but are not limited to a primate, a mouse, a rat, a guinea pig or a rabbit. In a preferred example, the subject is a mouse. The plurality of polynucleotides or the library comprising a plurality of polynucleotides is delivered to an intact tissue sample comprising the population of cells. It should be understood that the intact tissue sample refers to a biological sample that retains the cellular architecture and arrangement within a tissue microenvironment. The intact tissue sample may be a whole organism, a whole tissue or a section of an organ or a tissue. Examples of a whole tissue may comprise but are not limited to an intact organ, an organoid and a tissue biopsy. A section of an organ or a tissue may be generated by slicing biological tissues using techniques known in the art such as cryosectioning and vibratome sectioning. Examples of a section of an organ or a tissue may comprise but are not limited to brain, skin or liver tissue. In one example, the intact tissue sample is a brain sample.
[0098] The plurality of polynucleotides or the library of polynucleotides may also be delivered to a specific region or regions in the intact tissue sample. In one example, the plurality of polynucleotides or the library of polynucleotides may be delivered to the hippocampus of the brain.
[0099] The invention illustratively described herein may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, for example, the terms "comprising", "including", "containing", etc. shall be read expansively and without limitation. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scopeof the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by preferred embodiments and optional features, modification and variation of the inventions embodied therein herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention.
[0100] The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the invention. This includes the generic description of the invention with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein.
[0101] Other embodiments are within the following claims and non- limiting examples. In addition, where features or aspects of the invention are described in terms of Markush groups, those skilled in the art will recognize that the invention is also thereby described in terms of any individual member or subgroup of members of the Markush group.EXPERIMENTAL SECTION
[0102] Non-limiting examples of the invention and comparative examples will be further described in greater detail by reference to specific Examples, which should not be construed as in any way limiting the scope of the invention.
[0103] Materials and Methods
[0104] Mice
[0105] All animal work was performed under the guidelines of the Institutional Animal Care and Use Committee (1ACUC protocol #211646). Mice were housed on a standard light cycle under pathogen-free conditions. Female Rosa26-CAG-Cas9 mice (JAX 024858) aged 8-16 weeks were used.
[0106] Stereotaxic Injections
[0107] Animals were anesthetized with ketamine / xylazine (injected intraperitoneally) and placed on a thermostatically controlled heating pad Animals were checked for sedation. Hair was removed from the incision site on the scalp with a shaver and the area was cleaned with an alcohol wipe. Corneas were protected with an ophthalmic lubricating ointment. Animals were then positioned in a stereotaxic head holder with the skull firmly barred. An incision on the scalp was made and the scalp reflected over the brain region of interest. For all experiments,the left hippocampus was targeted with the following stereotaxic coordinates relative to the bregma: AP -1.7, ML +1.6, DV -1.8. A handheld dental drill was used to drill a <lmm diameter hole in calcified bone in the AP, ML region of interest. The cannula was inserted into the region of interest and the 5 x 10A8viral particles of AAVs were injected at a rate of 2- 5nl / second in a maximal volume of 1.5ul, so as not to cause tissue injury. Following injection, the cannula was withdrawn after 5 minutes. The incision on the scalp was closed with Vetbond Tissue Adhesive. Atipamezole was administered intraperitoneally for anaesthesia reversal and Buprenorphine was administered subcutaneously for pain relief. Animals were subsequently placed on a thermostatically controlled heating pad for recovery and monitoring. For 3 days post-surgery, Buprenorphine was administered for pain relief, and animals were monitored for signs of distress and wound inflammation. Mice were kept for 3 weeks under standard conditions before being sacrificed by CO2 asphyxiation for tissue extraction and processing.
[0108] gRNA library design
[0109] Three guides were designed each for 18 loci - 17 genes and a safe harbour locusmSafe) as a control (54 guides in total). The three guides for the same locus were linked to the same barcode on one vector, to facilitate detection of the genetic knockout via a single unique barcode. Each vector contained TdTomato, with 18 vectors in total. The gRNA library was focused to target genes associated by GWAS. From the NHGRLEBI Catalog of human genome-wide association studies, the GWAS hits associated with Alzheimer's Disease were chosen (https: / / www.ebi.ac.uk / gwas / efotraits / MONDO_0004975) (CLU, LRRK2 NDUFAF2, RBFOX1, TREM2), Amyotrophic Lateral Sclerosis(https: / / www.ebi.ac.uk / gwas / efotraits / MONDO_0004976) (C9orf72, CFAP410, DPP6, TBK1), and Parkinson’s disease (SH3GL2, STK39) (https: / / www.ebi.ac.uk / gwas / efotraits / MONDO_0005180). These genes have mouse orthologs and are expressed in the mouse brain. 6 positive control genes (Fasn, Flcn, Gfap, Olig2, Krciga, Srf) and a safe harbour site were included. With these 18 genes / loci of interest, three gRNAs each were designed using the online tool CHOPCHOP (https: / / chopchop.cbu.uib.no / ). Bestscoring guides within the first few exons were selected.
[0110] Cloning of gRNAs plasmid pool
[0111] The plasmid pAAV-CAG-tdTomato (codon diversified) was ordered from Addgene (#59462). The plasmid backbone (2ug) was digested with Hindlll / Sall (NEB) for Ih at 37°C followed by an inactivation step for 20 mins at 80°C. DNA fragments consisting of both theoverlapping sequences with the digested pAAV-CAG-tdTomato backbone and the 450 bp barcodes were ordered from IDT. PCR reactions were used to amplify the DNA fragments. The Gibson assembly reaction was set as follows: 100 ng of digested plasmid backbone, lOng of amplified DNA fragment, 10 pl NEBuilder HiFi DNA Assembly Master Mix (NEB, E2621L) and H2O up to 20 pl total reaction. The reaction was incubated for 1 h at 50°C. 5ul of the Gibson reaction was used for transformation using Stbl3 cells (Thermo Fisher, Cat#C737303). Sequencing reactions (Biobasic) were used to confirm the correct insertion of the barcodes. Following barcodes cloning, three gRNAs were designed for each locus-of- interest to be inserted into each barcoded plasmid. Golden gate cloning strategy was employed to arrange the three gRNAs in tandem, each driven by a U6 promoter. The gRNAs were then inserted into the Ndel-digested pAAV-CAG-tdTomato using Gibson assembly.
[0112] AAV production and purification
[0113] AAVs were generated in-house. Briefly, AAVs were packaged via a triple transfection of AAV cell line (AAV-100, Cell Biolabs, San Diego, CA, USA). The cells were plated in a HYPERFlask ‘M’ (Coming) in growth media which consisted of DMEM + glutaMax + pyruvate + 10% FBS (Thermo Fisher), that was supplemented with IX MEM nonessential amino acids (Gibco). Transfection was done when the cells were between 70% and 90% confluent. Media was replaced with fresh pre-warmed growth media before transfection For each HYPERFlask ‘M’, 200 pg of pHelper (Cell Biolabs), 100 pg of pRepCap encoding capsid proteins for serotype DJ or 2], and 100 pg of pZac-CASI-GFP or pZac-CMV- CasRx gRNAs were mixed in 5 ml of DMEM, and 2 mg of PET “MAX” (Polysciences) (40 kDa, 1 mg / ml in H2O, pH 7.1) was added for PEI: DNA mass ratio of 5: 1. The mixture was incubated for 15 min before being transferred to the cell media drop-wise. The day after transfection, the media was changed to DMEM + glutamax + pyruvate + 2% FBS. 48-72 h after transfection, the cells were harvested by scrapping or dissociation with IX PBS (pH7.2) + 5 mM EDTA and then pelleted at 1500 g for 12 min. Cell pellets were resuspended in 1-5 ml of lysis buffer (Tris HC1 pH 7.5 + 2 mM MgCl + 150 mM NaCl), before freeze-thawing thrice using a dry-ice-ethanol bath and a 37°C water bath. Cell debris was clarified via centrifuging at 4000 g for 5 min, and the supernatant was collected. To remove the unpackaged nucleic acids, the collected supernatant was treated with 50 U / ml of Benzonase (Sigma- Aldrich) and 1 U / ml of RNase cocktail (Invitrogen) for 30 min at 37°C. Next, the lysate was loaded on top of a discontinuous density gradient that consisted of 6 ml each of 15%, 25%,40%, and 60% Optiprep (Sigma-Aldrich) in a 29.9 ml Optiseal polypropylene tube (Beckman- Coulter). The tubes were then ultra-centrifuged at 54,000 rpm, for 1.5 h at 18°C, using a Type 70 Ti rotor. The 40% fraction was extracted and dialyzed with IX PBS (pH 7.2) supplemented with 35 mM NaCl, using Amicon Ultra-15 (100 kDa MWCO) (Millipore). qPCR was then carried out with the ITR-sequence-specific primers and probes, using the ATCC reference standard material 8 (ATCC) as a standard, to determine the titres of the purified AAV vector stocks.
[0114] Guide vector recombination assay
[0115] HEK293 cells were seeded in 48-well plates and transduced with recombinant adeno-associated viruses (AAVs) for 48 hours at multiplicity of infection (MOI) of 10000. The transduction conditions included transduction with either AAVPhP.eB-Tdtomato-gRNAs or AAVPhP.eB-GFP and a non-transduced control. At 48 hours post-transduction, the cells were washed three times with phosphate-buffered saline (PBS), followed by cell lysis using 50 pL of QuickExtract™ DNA Extraction Solution (Lucigen) per well. The lysates were incubated according to the manufacturer’s instructions, and 1 pL of the resulting lysate was used as template DNA for downstream PCR analysis. PCR was performed using a primer pair designed to detect recombination events at the target locus. The primers were synthesized based on a calculated melting temperature (Tm) of 55 °C: Forward primer: 5'- CACTTGGCAGTACATCAAGTGT-3' (SEQ ID NO: 92), Reverse primer: 5'- GCCATTTACCGTCATTGACGT-3' (SEQ ID NO: 93). Each 25 pL PCR reaction contained IX high-fidelity PCR buffer, 0 4 pM of each primer, 0.5 U of high-fidelity Q5 DNA polymerase (NEB), and 1 pL of lysate template. The thermal cycling conditions were as follows: initial denaturation at 98°C for 1 minute; 35 cycles of denaturation at 95°C for 15 seconds, annealing at 54°C for 15 seconds, and extension at 70°C for 1 minute; followed by a final extension at 70°C for 5 minutes. PCR products were analyzed on a 1% agarose gel stained with Gel Red. Electrophoresis was performed at 120 V for approximately 40 minutes. Bands were visualized using a gel documentation system, and densitometry measurements were done in Image!
[0116] Tissue dissection, dissociation and FACS
[0117] After mice were sacrificed by CO2 asphyxiation, intracardial perfusion with ice-cold PBS was carried out to remove blood cells. Brains were dissected and either embedded in OCT for sectioning or dissociated into single cells. Dissociation of brain tissue was done with theNeural Dissociation Kit (Miltenyi Biotec) following manufacturer’s instructions. Briefly, brain tissue of interest was minced with a razor blade followed by enzymatic digestion at 37°C and triturated. The cell suspension was then strained through a 70um cell strainer and centrifuged at 300g for lOmin. The cell pellets were then re-suspended in Hibernate A with lOuM Calcien- AM (Invitrogen) and lug / ml Propidium Iodide (PI) (Stemcell), for FACS sorting. 275 Live, single cells were selected by positive selection of Calcien-AM and negative selection of PI. Immediately after FACS sorting, cells were used for the generation of 10X 3’ chromium libraries and sequenced on Novaseq (Illumina).[00118 J Coverslip functionalization and sample preparation for FISH
[0119] Coverslip functionalization was carried out before use for tissue sectioning. Coverslips (Warner Instruments, cat. no. 64-1500) were cleaned in IM KOH for 1 h before being rinsed thrice in Mili Q water. The coverslips were rinsed with 100% methanol before being functionalised in an amino-silane solution (3% vol / vol (3 -aminopropyl) triethoxysilane (Merck cat no. 440140), 5% vol / vol acetic acid (Sigma, cat. no. 537020) for 2 min at room temperature. Following which, the coverslips were rinsed thrice with Mili Q water before being dried overnight at 47°C in an oven. The mouse brain sample was frozen in optimal cutting temperature compound and sectioned into 7 pm sections onto functionalized coverslips using a cryostat. The sections were fixed using 4% vol / vol paraformaldehyde in lx PBS for 15 minutes, then rinsed with lx PBS and stored at -80°C.
[0120] FISH
[0121] The encoding probes were diluted in a 10% hybridization solution that was composed of 10% deionized formamide (Ambion™ Cat: AM9342) (vol / vol), 1 mg ml-1 yeast tRNA (Life Technologies, cat. no. 15401-011) and 10% dextran sulfate (Sigma, cat. no. D8906) (wt / vol) in 2x SSC. A final concentration of 10-50 nmol per probe was used. Mouse brain sections were permeabilized with 70% ethanol overnight at 4°C. After permeabilization, samples were rinsed twice with 2x SSC. Next, encoding probe staining was performed and samples were stained for 16 h at 37°C. After encoding probe staining, samples were washed in 10% formamide wash buffer (10% deionized formamide in 2x SSC) at 37°C for 15 min, twice. The samples were rinsed thrice with 2x SSC, before DAPI (Sigma, cat. no. D9564) staining was carried out.
[0122] Microscopy for FISH
[0123] Microscopy and acquisition of fluorescent images was carried out on a custom-built microscope, constructed using a Nikon Ti2-E body. A Marzhauser SCANplus IM 130 mm x 85 mm motorized X-Y stage, a pco.edge 4.2 BI-USB Back Illuminated sCMOS camera, a custom fiber-coupled laser box from CNI laser for illumination, and a Nikon CFI Plan Apo Lambda 60x 1.4-n.a oil-immersion objective (MRD01605) was used. The Nikon Perfect Focus system was used to maintain focus while imaging, and in each imaging cycle, one Z position was imaged for each field of view. Custom in-house software was used for acquiring these images. Imaging was carried out on a microscope constructed around a Nikon Ti2-E body. A Marzhauser SCANplus IM 130 mm * 85 mm motorized X-Y stage, and an Andor Sona 4.2B-11 sCMOS camera were used, and imaging was done with a Nikon CFI Plan Apo Lambda 60* 1.4- n.a. oil-immersion objective. Lasers used for illumination were as follows: 2RU-VFL-P-500-592-B1R (500 mW), 2RU-VFL-P-1000-647-BlR (1000 mW), 2RU-VFL-P- 500-750-B1R (500 mW) (MPB Communications), and Coherent Obis 405 100-mW laser. An exposure time of 1 s was used for imaging, and focus was kept using the Nikon Perfect Focus System (PFS).
[0124] Samples were mounted onto the microscope stage via use of a flow chamber (Bioptechs, cat. no. FCS2). This allowed buffer exchange to be carried out via use of a computer-controlled fluidics system (Cite2). The sample was stained at room temperature for 15 min with readout probe solution via buffer exchange, prior to imaging. The readout probe solution consisted of 10 nM of each fluorescently labelled readout probe, 10% deionized formamide (vol / vol) and 10% dextran sulfate (wt / vol) in 4* SSC. After hybridization, 2* SSC was flowed in before a rinse with 10% formamide wash buffer. 2* SSC flowed again before imaging buffer. The imaging buffer contained 2* SSC, 10% glucose, 50 mM Tris-HCl pH 8, 2 mM Trolox (Sigma, cat. no. 238813), 40 pg / ml catalase (Sigma, cat. no. C30), and 0.5 mg / ml glucose oxidase (Sigma, cat. no. G2133). Signal removal was done by washing the samples with a 55% formamide wash buffer containing 0.1% TritonX-100.
[0125] IPX library preparation and sequencing
[0126] Libraries from dissociated single cells were generated using Chromium Next GEM Single Cell 3’ Reagent Kits v3.2 (Dual Index) (10X Genomics). Cell numbers were counted by FACS and 20,000 cells per mouse were used to generate each library. Libraries were generated according to the manufacturer’ s instructions, with 12 cycles of cDNA amplifications, 25ng of cDNA input and 14 cycles of sample index PCR. Barcode sequences were notamplified to avoid introduction of biasness and overamplification. Libraries generated had a fragment length of 400-450bp and were sequenced on an Illumina Novaseq (PEI 50).
[0127] iSeq library preparation and sequencine
[0128] In order to confirm the presence of indels in brain tissue after AAV transduction, 50k TdTomato+ cells from 4 brains were lysed with QuickExtract (Lucigen) according to manufacturer’s instructions. Uninjected brain tissue was used as WT control. Two rounds of PCR were used. For PCR1, genotyping primers were designed using Primer3 (https: / / primer3.ut.ee / ) to amplify a 150-300bp region around the Cas9 cut site with 34 PCR cycles. The following sequences were added to the forward and reverse genotyping primers respectively, for the second round of amplification:CTTTCCCTACACGACGCTCTTCCGATCTNNNNNN (SEQ ID NO: 1), GGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 2). NNNNN is used to introduce diversity at the start of the iSeq read to improve read quality. For PCR2 used to index samples, 6 PCR cycles were used.
[0129] For both PCR1 and PCR2, Q5 High Fidelity Master Mix (NEB) was used according to manufacturer’s recommended PCR temperatures and parameters. Indexed PCR products from PCR2 were pooled, ran on a gel, and gel purified using Wizard SV Gel and PCR Clean- Up System (Promega). DNA concentration was measured with Qubit HS dsDNA kit (Vazyrne). The library was spiked with 2% PhiX library (Illumina) and sequenced on an iSeq 100 system.
[0130] Stereo-seq chip preparation
[0131] OCT blocks were stored at -80°C and equilibrated at -20°C for 2 h prior to sectioning. OCT blocks were cryosectioned at a thickness of 10 pm using a CM 1950 cryostat. Following satisfactory quality control (QC) results (RIN = 8.5-8.6), mouse brain coronal sections were collected on the Stereo-seq chip surface. Tissue sections were adhered to a Stereo-seq chip surface and were incubated at 37°C for 5 min. The tissues were then fixed in methanol and incubated at -20°C for 30 min. The tissue sections were then permeabilized at 37°C for 10 min and washed with 0. l x SSC buffer. The RNA released from permeabilized tissues was captured using DNB probes and reverse-transcribed at 42°C for 3 hrs. After in situ reverse transcription, tissues were removed by tissue removal buffer. The chips were then incubated with 400 pL cDNA release buffer overnight at 55°C, and then washed once with 400 pL of 0.1 x SSC buffer. The released cDNA was then collected and purified using 0.8 Ampure XP Beads. The cDNA was then amplified using the following PCR conditions 95°C for 5 min,15 cycles of 98°C for 20 s, 58°C for 20 s, then 72°C for 3 min, and a final incubation at 72 °C for 5 min. PCR products were purified using 0.6* Ampure XP Beads and the concentrations of cDNA were quantified using a Qubit™ dsDNA Assay Kit.
[0132] Stereo-seq library preparation and sequencing
[0133] STOMICS libraries were constructed with the STOMICS Gene Expression kit. The cDNA was first checked for presence of barcodes by PCR. Briefly, 20ng of cDNA were fragmented at 55°C for 10 min The fragmented products were then amplified using the following PCR conditions one cycle at 95°C for 5 min, 13 cycles at 98°C for 20 s, 58°C for 20 s and 72°C for 30 s, and lastly one cycle at 72°C for 5 min. The PCR products were purified 374 using Ampure XP Beads for DNB generation and were finally sequenced (100 bp PE) using T7 sequencer.
[0134] Preparation of slides for Xenium VI assay
[0135] The Xenium vl assay workflow was carried out according to the manufacturer’s instructions (1 Ox Genomics, Xenium In Situ Gene Expression with Cell Segmentation Staining User Guide, Rev B, CG000749). Briefly, fresh frozen tissue sections (10 pm) placed on Xenium slides were fixed and permeabilized with minor modification to the manufacturer’s protocol (lOx Genomics, Xenium In Situ for Fresh Frozen Tissues - Fixation and Permeabilization User Guide, Rev D, CG000581). Before immersing sections in 1% SDS, sections were photobleached in lx PBS for 1 h at 4°C. Following fixation and permeabilization, probe hybridization with a combination of both the pre-designed mouse brain panel (247 genes) and an add-on custom panel of barcode sequences was carried out Probes that were bound to the target RNA were ligated at both ends to generate a circular DNA probe. Rolling circle amplification of the circularized DNA probe then generated multiple copies of the genespecific barcode for each RNA target. This was followed by morphology-based cell segmentation staining (boundary stain used: ATP 1 Al + CD45 + E-Cadherin) before continuing with autofluorescence quenching and nuclei staining. Slides were then loaded onto the Xenium Analyzer for high throughput and automated in situ imaging and analysis. Fluorescence-labeled oligos bound to the amplified DNA probes and tissues underwent multiple rounds of fluorescent probe hybridization, imaging, and probe removal to generate a codeword specific for each barcode. Each codeword was then converted into a gene identity.
[0136] Computational methods
[0137] Barcode design
[0138] The nucleotide barcodes were designed to be long enough to be probed against, that are tolerated in cells. Barcodes with these features were designed: 1) having an appropriate length between 20 nucleotides - 500 nucleotides- long enough for probing against, but short enough to avoid inducing deleterious effects and causing unwanted biological effects in the cell (450nt); 2) not containing stop codons, to avoid nonsense mediated decay; 3) having a high edit distance between each barcode; having a high edit distance between different segments of the same barcode, to allow design of a greater diversity of probes against the same barcode (with each probe being 20-100nt long), 4) having a high edit distance between the barcode and the endogenous human and mouse transcriptome, so that there is no alignment to endogenous transcripts, no unambiguous detection via probe-based spatial technologies and sequencing based methods and no genetic function for the prevention of unwanted biological effects, and 5) not containing 4 or more identical consecutive nucleotides i.e. AAAA, TTTT, CCCC, GGGG With these criteria, a set of 22 barcodes was generated and 18 barcodes were used for the library, with one barcode per gene (Fig. 1 A).
[0139] FISH encoding probe design
[0140] 14-16 encoding probes were designed against each of the 450bp barcodes using the Stellaris Probe Designer software (LGC Biosearch Technologies). Parameters used were: Masking level = 5 (mouse), Oligo length = 19-20nt, Min. spacing length = 2nt. Each encoding probe sequence was flanked on both sides by the readout probe sequence 5’- TCTGTTTGACGCGCT-3’ (SEQ ID NO: 3) with a spacer nucleotide (A) in between the readout and the encoding probe region. The concatenated sequences, and the readout probe ( / 5IRD800CWN / AGCGCGTCAAACAGA) (SEQ ID NO: 4) were purchased from Integrated DNA Technologies (IDT).
[0141] IPX sequencing data processing
[0142] Sequencing results from IPX libraries were processed and demultiplexed with CellRanger- 7.1.0 pipeline version (10X Genomics). An edited mouse reference genome (mm 10-2020) with the 450bp barcodes added was used for alignment and generating the unique molecular identifiers (UMI) count matrices allowing the barcode counts to be included in count matrices. Sequencing saturation of all reads and of barcode reads were calculated for all generated libraries and were >0.6 and >0.5 respectively. Filtered count matrices were analysed using the Seurat package (version 4.3.0.1) in R (version 4.3.1). Briefly, quality control (QC) was performed to keep only cells with percentage mitochondrial reads below 15%, andnFeature_RNA between 200 and 7000, to exclude empty droplets, multiplets and dying cells. Counts were then normalized, scaled to 10000 transcripts per cell using the NormalizeData() and ScaleData() function respectively. Following which, Principal Component Analyasis (PCA) was run on the top 2000 most variable features, using the RunPCAQ function. Clusters were identified with the FindNeighbors() function by generating a K-nearest neighbour graph with 10-16 dimensions, and clustered using the Louvain algorithm with a resolution of 0.3, using the FindClusters() function The cells were then represented by a 2-dimensional Uniform Manifold Approximation and Projection (UMAP) graph. Lastly, coarse cell-type annotation was performed using a combination of analyzing the most highly expressed marker genes of each Seurat cluster, and expression of known canonical marker genes for each brain cell type (neurons, oligodendrocytes, astrocytes, endothelial cells, microglia). As each cluster was represented by all batches, no batch correction or integration methods were applied. For the analysis of oligodendrocyte perturbation phenotypes, oligodendrocytes expressing barcodes linked to mSafe, I)pp6, Lrrk2, Gfap, Cfap410, Rbfox3 are grouped as “control”, as these genes are not expected to impact oligodendrocyte function.[00143 J Assignment of barcode / gRNA status for IPX and Stereo-seq sequencing data
[0144] For assignment of barcode-positive and barcode-negative status, negative status was assigned to cells with no expression of any barcodes, and all other cells were positive. For assignment of gRNA status, cells were classified by the expression of barcodes as follows: No barcodes expressed = Neg; more than 1 unique barcode expressed = Multiple, exactly 1 unique barcode expressed = the gene corresponding to that barcode. For oligodendrocytes, C9orf72, Clu, Lrrk2, mSafe were grouped as “control” as these genes were very lowly expressed in oligodendrocytes and / or their KO were expected not to have an effect in oligodendrocytes. For neurons, Trem2, Olig2, Gfap, Stk39 and mSafe were grouped as “control”.
[0145] Gene editing analysis
[0146] Following preparation and sequencing of iSeq libraries as described above, quantification of percentage of indels was calculated with CRISPresso2 (v2.0.20) with default parameters (default min aln score 60, quantification window center -3, exclude bp from left 15, exclude bp from right 15, quantification window size 1, plot_window_size 20, min_bp_quality_or_N 0). % Indel for each guide was calculated as % Modified (edited) - % Modified (WT). For all loci, there were at least 20k aligned reads.
[0147] Stereo-seq data processing
[0148] For processing of fastq files, SAW pipeline version V7.0.0 and Image Studio version 3.0.0 was used. Genome and gtf files used for alignment were generated by combining mouse reference GRCm39 and the 450bp barcodes. CID (Coordinate Identity) sequences were mapped to the coordinates on the Stereo-seq chip, allowing 1 mismatch. Reads were filtered based on Q30 phred quality score, indicating less than 1 in 1000 chance of base calling error. Following alignment to the reference, a gene count file was generated quantifying deduplicated, annotated reads. This gene count matrix was then aligned to the image (image registration). Cell segmentation was conducted using nuclear-stained tissue images and gene expression data, with the DeepCell algorithm. Following generation of a Stereopy object from the cellbin GEF file, a h5ad file and Seurat object was generated. For quality control (QC), cells with percent.mito > 20, and number of counts between <100 or > 6000 were discarded. After QC, cells were processed through the standard Seurat pipeline, as described under the “10X sequencing data and processing” methods section. Neurons expressing barcodes linked to mSafe, Gfap, Stk39, Trem2, and Olig2 were grouped as control, as these genes were lowly expressed in neurons and / or are not expected to exert a strong phenotyping change.
[0149] Selection of hippocampal neurons with BANKSY
[0150] BANKSY was run on each chip separately, after data scaling, and before PCA using npcs = 30. Lambda 0.2 was used to achieve spatially -informed cell type embeddings. Higher lambda values resulted in spatial domain segmentation with separate sections on the same Stereo-seq chip being clustered separately. The following parameters were used: spatial mode = “knn r”, ndim = 2, k geom = 15. Clusters corresponding to hippocampal CAI , CA3 and DG neurons, the hippocampal neuronal niche, were kept for analysis.
[0151] Differential gene expression (DGE) and analysis
[0152] DGE analysis was performed with Seurat’ s FindMarkers function with the following parameters were: min.pct = 0.1, logfc.threshold = 0.5, test.use = "wilcox". For reporting of the number of DEGs, the same number of cells for each perturbation was used, to ensure comparability. For calculation of neighbour DEGs, 15 neighbours were identified with BANKSY (spatial mode = “knn r”) for each perturbed cell in the group, and barcode-positive neighbours were excluded. The wildtype neighbours of the 2 groups (case and control), were then compared. Genes with p < 0.05 and Ifc > 0.5 were counted as DEGs. For volcano plots, p values were adjusted with the p. adjust function to obtain fdr values. Genes with fdr < 0.05 and Ifc > 0.5 were marked as significant.
[0153] Cell-cell communication (CCC) analysis
[0154] The LIANA (Ligand-Receptor Analysis) tool, a computational pipeline for prioritizing ligand-receptor interactions based on databases like CellPhoneDB, connectomeDB2020, and CellChat, to elucidate cell -cell communication between neurons was used. As cells within the same microenvironment were more likely to be interacting, the communication between perturbed cells and their immediate neighbours was focused on. For each of these 3 perturbed groups, Srf-KO, Lrrk2-KO, and control, the perturbed cells and their wildtype / unperturbed neighbours were identified. 15 neighbours were identified for each perturbed cell with BANKSY (spatial mode = “knn_r”), and barcode-positive neighbours were excluded. For each pair (Srf-KO + neighbours, Lrrk2-KO + neighbours, control -KO + neighbours), LIANA with statistical methods “sea” and “natmi” was applied to calculate and prioritize top interactions, as these were most likely to be biologically relevant. The perturbed cells were modelled as the source cells and the wildtype neighbours as the target cells. The common ligand receptor pairs were then identified from the prioritized list between Lrrk2-KO and control, and Srf-KO and control, and showed the top 20 ligand-receptor pairs with biggest differences in sca.LRscore.
[0155] Processing of Xenium spatial transcriptomics data
[0156] Xenium spatial transcriptomics data were processed following the Seurat v5 Spatial Vignette (https: / / satijalab.org / seurat / articles / seurat5_spatial_vignette_2). Xenium output files were first converted to Seurat objects with the LoadXenium() function with the following parameters: mols.qv threshold = 20, cell. centroids = TRUE, molecule. coordinates = TRUE, segmentations = “cell”, flip.xy = FALSE. Gene expression data were normalized and scaled using default parameters with the SCTransform() function, followed by dimensionality reduction with the RunPCA() and RunUMAP() with dims = 1: 10. Clustering was performed with the FindNeighborsQ with dims = 1:10, and FindClustersQ functions. Images were produced with the ImageDimPlot() function. Cell types were annotated based on expression of canonical marker genes. Cells with barcode count >1 for were classified as barcode-positive.
[0157] Calcul tion of R2 values
[0158] The Cell Profiler software was used to calculate the coefficient of determination (R2 value) between TdTomato protein intensity and number of FISH spots of TdTomato RNA, Barcode, Actb control, from confocal images of TdTomato protein and FISH staining. A cropped image was first analyzed using Identify PrimaryObjects (for analysis of TdTomatoprotein fluorescent intensity or FISH spots), then gridded using the DefineGrid module (200 grid squares). Next, the IdentifyObjectsInGrid module was done, followed by the RelateObjects module to relate parent objects (grid squares) to child objects (TdTomato protein fluorescent intensity or FISH spots). The pearson R2 coefficient was then computed (one point per grid area) with a fixed intercept at origin, to assess correlation between TdTomato protein fluorescent intensity and number of FISH spots.
[0159] Power analysis with powsimR
[0160] First, the noise in the dataset was modelled, as this affected the sensitivity of the platform. To do this, the mean-dispersion relationship was estimated using a negative binomial model, using counts data from neurons from chip2 randomly down-sampled to 10000 cells, and normalized with scran. These distribution statistics were then used to set up simulations, using the following parameters: nsims = 25, p.DE = 0.1, pLFC = 2, LibSize = “equal”. Lastly, marginal and conditional FDR and TPR were evaluated in the evaluateDE function with the following parameters: alpha.type = 'adjusted', MTC = 'BH', alpha.nominal = 0.1, stratify by = 'dispersion', filter.by = 'dispersion', strata. filtered = 1, target.by = 'Ifc', delta = 0. As the average number of cells that was obtained per perturbation was 122, marginal and conditional 506 FDR and TPR were shown with the following number of cells: 5 vs 5, 25 vs 25, 125 vs 125, 625 vs 625.
[0161] Data availability
[0162] Raw and processed sequencing data was deposited at NCBI’s Gene Expression Omnibus (GEO) with accession numbers GSE274447 (Stereo-seq) and GSE274058 (scRNA- seq).
[0163] Code availability
[0164] Analysis of the processed data in this study was done using open-source R packages Seurat (https: / / github.com / satijalab / seurat), BANKSY(https : / / github . com / prabhakarlab / Banksy), powsimR (https : / / github . com / bvieth / powsimR), scCustomize (https: / / github.com / samuel-marsh / scCustomize), and open-source Python packages Stereopy (https: / / stereopy.readthedocs.io / en / latest / ) and SAW (https: / / github.com / STOmics / SAW).
[0165] Results
[0166] Example 1
[0167] In vivo Spatial Perturb-Seq pipeline
[0168] The Spatial Perturb-Seq platform was designed to be compatible with 1) single-cell transcriptomics, 2) sequencing-based spatial transcriptomics and 3) FISH-based spatial transcriptomics. The Spatial Perturb-Seq was demonstrated by simultaneously interrogating 18 genes within the same animals, using a pooled barcoded CRISPR-gRNA AAV library (Table 1) which consists of AAV-delivered CRISPR gRNAs, barcoded at the 3’ end (Fig. 1A), and injected intracranially into the hippocampus (Fig.lB) of Cas9-expressing mice at low multiplicity of infection (MOI). The goal was to sparsely edit cells, so that individual edited cells are surrounded by non-edited cells, allowing us to isolate cell autonomous and non-cell autonomous (i.e. neighbourhood) effects of the genetic perturbation without confounding influences from the surrounding cellular environment. Two NGS assays, 10X scRNA-seq and Stereo-seq, were used for profiling the perturbations. Most neighbours of perturbed cells were barcode-negative as designed. (Fig. 12).
[0169] Table 1. Polynucleotide sequences of the guide RNA (gRNA) and barcodes of the 18 genes.
[0170] The injected mouse brains with conventional single-cell in vivo Perturb-seq using 10X scRNA-seq (Fig. 2A - Fig. 2C) were profiled. For 10X scRNA-seq, the major cell types in the brain: neurons, oligodendrocytes, microglia, astrocytes, endothelial cells, choroid plexus, epithelial cells were recovered (Fig. 2A - Fig. 2C). Unlike in single-nucleus sequencing, neurons were underrepresented in single-cell sequencing as neurons were prone to dying during cell sorting and isolation as known in the art, as intended for the objective of studying glial cell types in this study. 41,667 cells over 3 batches (5 mice total) were obtained, with 2.38% of cells being positive for barcodes, consistent with a low MOI used. For the spatial component of the Spatial Perturb-seq pipeline, genetic edits were localized within the native hippocampal tissue using Stereo-seq as described in the Section “Stereo-seq chip preparation” under materials and methods (Fig. 1F-1H).
[0171] Among the different cell types, oligodendrocytes were predominantly recovered, with 11% of oligodendrocytes being positive for barcodes and an average of 26 oligodendrocytes per perturbation (Fig. 2D, 2E, 21 and 2J) Neurons were underrepresented, as the neurons are fragile during the cell dissociation process used.
[0172] Spatial Perturb-Seq provides spatial data that reveals biological interactions beyond conventional single-cell Perturb-Seq. For the spatial component of our Spatial Perturb-seq pipeline, genetic edits within the native hippocampal tissue using Stereo-seq were localized (Fig. 1F-1H) Similar cell types: neurons, RBCs, oligodendrocytes, immature oligodendrocytes, microglia, astrocytes, choroid plexus, endothelial cells, T cells, with expected spatial domains were recovered (Fig. 1G) As expected, majority of cells were neurons (Fig. IF). Spatial sequencing data was obtained for 39,865 cells over two sections (Fig. 1H), with 0.44% of cells being positive for barcodes. This is lower than observed in the dissociated single-cell sequencing experiments due to the lack of enrichment of transduced cells by FACS, as well as different efficiencies in mRNA capture in the different technologies.
[0173] Spatial Perturb-Seq provides spatial data that reveals biological interactions beyond conventional single-cell Perturb-Seq. Applying in vivo Spatial Perturb-Seq with Stereo-seq (Fig. 1A), spatial sequencing data for 229,775 cells were obtained over 3 experiments and 4 mice (Fig. 7A, Fig. 7B and Table 2). The libraries were prepared without barcode-specific PCR-enrichment to avoid quantification bias. Nonetheless, intentional barcode enrichment did not drastically increase the number of barcode-positive cells, but increased the detected levels of barcodes, indicating that the sequencing saturation is sufficient (Fig. 7C). The dataset is then processed with DeepCell for cell segmentation (Fig. 7D), where an average of 462 and 773 expressed genes were detected per cell and per neuron, respectively (Table 2). Neurons, RBCs, oligodendrocytes, microglia, astrocytes, choroid plexus and endothelial cells (Fig. IB - IF) were recovered with spatial patterns consistent with gene expression modules (Fig. 7E and 7F). The consistency in the manual annotation and spatial hotspot analyses of gene modules, provides confidence in the data quality and cell segmentation. A subsection of stereo-seq data highlights individual cells with barcode positive cells in black, and negative cells in grey (Fig. 6A). All 18 barcodes are represented in the sequencing data (Fig. 6B and Fig. 6C). The libraries were prepared without targeted, barcode-specific PCR-enrichment.
[0174] Table 2. Summary of Stereo-seq metrics|00175| Example 2
[0176] AAV serotype and tropism
[0177] AAV was chosen to deliver the library, over the more commonly used lentivirus to avoid possible genotoxic integration of the lentiviral cargo into the genome , as well as to avoid the previously observed undesired template switching events that uncouples barcodes from the respective gRNAs Despite this, lentiviral libraries are often used. This is because the cargo capacity of AAVs is limited. This issue was circumvented by using mice that constitutively express Cas9, bypassing the need to include Cas9 in the AAV vector. Because each of the three gRNAs is driven by the mU6 promoter and prone to recombination, the presence of the full- length gRNA expression cassette before and after AAV transduction was confirmed (Fig. 18). Tropism of the AAV-PHP.eB serotype used here was indeed highest in neurons, with high tropism also observed for astrocytes and oligodendrocytes, but not for microglia and endothelial cells (Fig. 2A-2E).
[0178] To select a broad brain-tropic AAV serotype, a pooled mix of 3 AAV serotypes(AAV-PHP.eB, AAV2 and AAV9) containing a unique serotype-specific barcode and GFP cargo was injected into the adult mouse hippocampus and detection of each of the 3 barcodesby scRNA-seq was assessed. Briefly, the names and barcodes of each AAV serotypes are manually included into both files that will be used for the execution of the Cell Ranger Count command pipeline in order to include the AAV barcode transcripts into the read alignment, UMI counting, and clustering. To include the AAV barcode representation in the genome reference file, the command line “>GFP1 TAAATCGATCGNNNNNNNN (SEQ ID NO: 91)” is included for each barcode, where the 8Ns represent a unique 8 nucleotide barcode sequence. The command line “GFP me exon 1 19 - + - gene id “GFP1”; transcript ! d “GFP1”” was included in the genome transcript file for each AAV barcode representation added to the genome reference file (Supplementary Data SI in Keng et al. Multiplex viral tropism assay in complex cell populations with single-cell resolution. Gene Therapy 29, 555-565 (2022)). After processing by Cell Ranger to obtain the gene count matrices, further computations were carried out in R (version 4.0.4). Quality control, normalization, PCA, clustering and downstream analysis were performed with Seurat (version 4.0.1). Briefly, cells with less than 200 RNA features or more than 10% mitochondrial reads were removed from analysis The dimensions of the final data matrices are 17,022 features across 5741 cells (Ocular dataset) and 20281 features across 15, 5 cells (Cerebral dataset). After normalization, PCA was performed with the top 2000 most variable features, followed by Louvain clustering with a resolution of 0.3- 0.4 to achieve reasonable distinction of cell type clusters, and dimensional reduction using tSNE. Plots and figures were assembled using Seurat and the ggpubr (version 0.4.0) package. To determine the transduction specificity of a specific viral vector against a specific cell niche relative to other cell niches, we calculated the frequencies with which the presence of that specific viral vector is detected in the cells of the specific cell niche, against the frequencies with which the presence of the same specific viral vector is detected in the cells of other specific cell niches.
[0179] As these serotypes were already known to have strong tropism in neurons, the transduction of these 3 serotypes in 5 non-neuronal cell types: microglia, oligodendrocytes, astrocytes, endothelial cells and oligodendrocyte precursor cells (OPCs) was compared. There was no transduction observed with AAV2 AAV-PHP.eB transduced oligodendrocytes, microglia, astrocytes and endothelial cells, while AAV9 transduced oligodendrocytes, microglia and astrocytes. Based on this, the AAV-PHP.eB serotype was chosen for all subsequent experiments.
[0180] The specific tropism of AAV-PHP.eB for non-neuronal CNS cell types such as microglia and endothelial cells were not clear. To determine the tropism of AAV-PHP.eB in the different cell types, the mouse brain was injected with AAV-PHP.eB alone, quantified the proportion of transduced cells in each cell type by scRNA-seq. It was shown that AAV-PHP.eB has broad spectrum tropism and transduces the main cell types in the brain - neurons, microglia, astrocytes, oligodendrocytes, endothelial cells. Tropism was highest in neurons, astrocytes and oligodendrocytes, and lowest for microglia and endothelial cells (Fig. 2A-2E). This potentially reflects the similar lineages of neurons, astrocytes and oligodendrocytes being derived from neural stem cells, and microglia and endothelial cells arising from the different developmental lineages of myeloid progenitors and mesoderm-derived angioblasts, respectively.
[0181] Example 3
[0182] Spread and biodistribution of AAV
[0183] While the spread of AAV transduction across the brain following intravenous injection had been characterized and shown to be even and homogenous, the spread of AAV vectors following intracranial injection into a specific and localized region was not well- established. To assess the specificity and spread of AAV across different brain regions and structures when injected directly into the hippocampus, confocal imaging of coronal sections of the mouse brain were carried out. Transduction was observed in both the ipsilateral and contralaterally hippocampus, but not in brain structures such as the cortex and corpus callosum (Fig. IB). This indicates that the intracranial injection of AAV allowed specific targeting of the hippocampus, the region associated with memory and the structure that is most severely affected during neurodegenerative diseases such as Alzheimer’s Disease.
[0184] In addition, unlike the liver accumulation of delivery vectors following intravenous administration, the results showed that the intracranial delivery route did not lead to accumulation of viral particles in the liver, suggesting a localised transduction into the brain (Fig. 3C).
[0185] Based on a hypothesis that promoters with strong activity may drive higher RNA levels of transgene expression over time, the levels of detected barcodes were investigated by scRNA-seq 3 weeks or 8 weeks after intracranial injection (Fig. 3A). The results demonstrated that expression levels of each barcode were similar in both durations (Fig. 3B), suggesting thatthe transgene of AAV cargo was stable for at least 2 months, but did not increase beyond 3 weeks. An intracranial delivery route and a 3 week incubation time period was chosen.
[0186] Following dissection of the injected, ipsilateral side of the mouse brain, ~7% are TdTomato positive by FACS (Fig. 3D) and -2.38% of those were barcode-positive, reflecting the difference in levels and detection sensitivity of TdTomato fluorescent protein by FACS, and levels of detection sensitivity of RNA by scRNA-seq.
[0187] Example 4
[0188] Gene editing efficiency
[0189] Each AAV delivered 3 gRNAs targeting the same gene, because gene editing efficiency was increased using multiple guides compared to single guides (Fig. 4A), and to reduce perturbation failure due to non-functional guides. In silico off-target prediction for all 54 guides identified only 6 sites in the mouse genome with <3 mismatches.
[0190] To verify CRISPR-Cas9 activity in the Cas9 mice, gRNAs alone were transfected into BMDMs derived from Cas9 mice and and achieved high editing rates, showing robust Cas9 activity and viability of the gRNA design approach (Fig. 4B). Next, to verify in vivo editing in the experimental tissue, iSeq sequencing was performed on sorted TdTomato+ cells from the injected hippocampus and confirmed presence of indels, validating our gene editing architecture (Fig. 4C and 4D). An average of 0.15% indels was observed for each locus genotyped in the single cell lysate (Fig. 4C and 4D), close to the observed mean of 0.13% of sequenced single cells containing each barcode.
[0191] Mutational profiling showed distinct indels between 1 -10 bp (Fig. 4E), consistent with precise on-target CRISPR / Cas9-mediated editing at the target sites. This suggests that cells delivered with barcoded CR1SPR gRNAs were largely edited at the genomic loci sequenced, consistent with editing efficiencies of single gRNAs in Cas9-mice. The consistency between the intentional low fraction of cells harbouring individual barcodes and the fraction of cells exhibiting targeted genetic perturbations, together indicate that barcode detection by sequencing is a reasonable proxy for inferring linked genetic perturbations.
[0192] Example s
[0193] Description and comparison of transcriptomic phenotypes of AAV transduction in different cell types by scRNAseq
[0194] Research on AAV applications has been focused on improving transduction efficiency and cell type specificity. However, understanding the effects of AAV transductionand repairing genomic lesions beyond the intended genetic edits, is also crucial. Therefore, using scRNAseq, these effects were profiled in different cell types in the brain, by the analysis of DEGs between cells with and without detected barcodes.
[0195] The results showed that different brain cell types activated different transcriptional programs in response to AAV-CRISPR-gRNA transduction. Transduced microglia upregulated immune response and inflammation related genes, as well as programs related to chemotaxis and migration, consistent with their role as innate immune cells, and downregulated genes related to transcription and biosynthesis. Transduced neurons upregulatde genes related to immune response and antigen processing, and downregulated genes associated with cilium movement and assembly, cell projection, axonogenesis and axon guidance. This is consistent with reports demonstrating that neurons exhibted characteristics of antigen presenting cells (APCs) in pathological conditions, despite not being primary players of antigen presentation. Transduced astrocytes upregulated genes associated with cell adhesion and downregulated transcriptional programs related to signalling pathways and transduction such as ERK and Wnt signalling pathways. Together with recent reports underscoring the role of Wnt signalling in the astrocytic function of blood brain barrier (BBB) integrity maintenance, the present data suggested potential mechanistic links between viral infection and BBB breakdown. Transduced oligodendrocytes upregulated genes related to transcription and downregulated programs related to signal transduction and ion transport (Fig. 2H). Further investigations of host changes following viral transduction will pave the way for improved efficacy and safety profiles of AAV-based gene therapy.
[0196] Example 6
[0197] Analysis of oligodendrocyte phenotypes by scRNAseq
[0198] Oligodendrocyte and myelin dysfunction underlie demyelinating and dysmyelinating disorders such as multiple sclerosis and leukodystrophies, a vast majority of which have no cure. The glial cells, specifically oligodendrocytes, in the dissociated single cell RNA-seq data were studied.
[0199] First, the oligodendrocytes were sub-clustered and showed recovery of known oligodendrocyte subtypes: healthy oligodendrocytes subtype M0L1, M0L2, MOL5 / 6, and the rare disease associated interferon-related oligodendrocyte sub-type (DA_lfn) (Fig. 5A and 5B) The proportion of gRNA+ M0L1 oligodendrocytes is significantly higher than gRNA+ MOL5 / 6 oligodendrocytes (Fig. 5C and 5D), suggesting that injection injury and AAVtransduction could steer oligodendrocytes to a less mature, M0L1 state, consistent with reports showing M0L1 enrichment at injury sites. MOL1, MOL2 and MOL5 / 6 are distinct oligodendrocyte sub-types, expressing different molecular markers and playing different roles in neuronal trophic support and myelination. It has also been suggested that these oligodendrocyte subtypes represent different maturation stages, with MOL5 / 6 being the most progressive. The data suggested that AAV-PHP.eB could have higher tropism for M0L1 oligodendrocytes, or that the effects of AAV transduction and / or CRISPR-Cas9-induced genomic lesions steered oligodendrocytes to a less mature, M0L1 state; the latter being consistent with reports observing that M0L1 is enriched at injury sites. As there was a spatial preference for M0L2 to reside in white matter areas of the brain, and MOL5 / 6 in grey matter regions, it is unlikely that the observed difference in tropism was due to the hippocampus (injection site) harbouring a lower fraction of MOL5 / 6 oligodendrocytes compared to surrounding regions like the corpus callosum.
[0200] Oligodendrocytes are responsible for making and maintaining the myelin sheath, regulating its composition and structure. This maintenance involves lipid metabolism and synthesis, a process crucial in oligodendrocytes. Dysregulation of lipid metabolism in oligodendrocytes leads to myelin dysfunction, impairments in neuronal function, and possible neurodegeneration. The expression of two suites of genes in oligodendrocytes: myelin genes (Mobp, Mog, Opalin, Plpl, Mbp, Cnp, ag, Mad) and lipid metabolism genes (Srebfl, Hmgcr, Scdl, Scd2, Acaca, Dgall, Cptla, Elovl6) was investigated. Oligodendrocytes expressing barcodes linked to mSafe, Dpp6, Lrrk2, Gfap, Cfap410, Rbfox3 were grouped as “Control”, as these genes were not expected to impact oligodendrocyte function. Olig2-KO did not significantly reduce the expression of canonical mature myelin genes (Fig. 5E), consistent with its role in oligodendrocyte specification but not maintenance. Rraga and Flcn perturbations cause slight downregulation of myelination genes in mice (Fig. 5C), suggesting that the Rraga- Flcn-Tfeb pathway could be a promising target for promoting remyelination. This is consistent with Flcn and Rraga being required for myelination in zebrafish, even though their function in mammalian myelination has not been recognized Fasn is a key regulator of fatty acid metabolism. Using Fasn as a positive control, the results confirmed that Fasn was required for lipid metabolism in oligodendrocytes (Fig. 5F). In summary, the results demonstrated that the perturb-seq platform is applicable for direct, in vivo Perturb-seq.
[0201] Example 7
[0202] Barcode design and compatibility with FISH
[0203] Spatial transcriptomics holds huge significance in discovery biology as it enables the precise mapping of gene expression patterns within the context of tissue architecture and microenvironments. There are several branches of spatial phenotyping - the two main branches being sequencing-based such as Visium, Stereo-seq, and GeoMx and probe-based such as CosMx, MERFISH, and Xenium. FISH-based spatial transcriptomics offer several advantages, including single molecule sensitivity and resolution.
[0204] Spatial Perturb-Seq was presented, which enables the interrogation of genetic perturbations on cells and their microenvironment directly in the tissue architecture with singlecell resolution and whole transcriptome coverage. This new ability to segment a genetic perturbation’s autonomous and non-cell autonomous effects reveals insights into cell-cell communication that would otherwise be lost via conventional dissociated single cell sequencing.
[0205] Firstly, the results showed that the Spatial Perturb-Seq system was compatible with FISH, by using nucleotide barcodes that can be probed against. The advantages of nucleotide barcodes over protein barcodes are its scalability, versatility, and compatibility with sequencing readouts. While detection of exogenous barcodes from AAV cargo was shown in dissociated single-cell sequencing, detection of barcodes by FISH probes is not routine, chiefly because parameters of barcode design are not established
[0206] Spatial Perturb-Seq was presented, which enables the interrogation of genetic perturbations on cells and their microenvironment directly in the tissue architecture with singlecell resolution and whole transcriptome coverage. This new ability to segment a genetic perturbation’s autonomous and non-cell autonomous effects reveals insights into cell-cell communication that would otherwise be lost via conventional dissociated single cell sequencing. To assess compatibility of Spatial Perturb-Seq with an alternative nucleotide- based detection system, Xenium was performed and the results showed robust detection of the barcodes in this present study and reduction of target genes with increasing barcode counts (Fig. 14 and Fig. 16). RNA FISH was also performed against the nucleotide barcodes and the results showed high specificity FISH spots (R2 = 0.78) (Fig. 12), demonstrating compatibility with probe-based spatial transcriptomics platforms such as MERFISH or CosMx. Further introduction of cell-type specific barcode expression or sequential injections of different barcodes could endow capabilities for spatial lineage tracing and spatial molecular timing. Theprimary bottlenecks of Spatial Perturb-Seq are cost of sequencing as both perturbed cells and their unperturbed neighbours must be sequenced, and the naturally limited cell numbers in a tissue niche. However, the steady decline in sequencing costs will enable larger sample sizes to be sequenced to increase statistical power.
[0207] Using this script, a set of 22 barcodes (using 18 of them for the library in this report) was generated. The results showed that the features of the uniquely designed barcodes were tolerated in cells and did not cause overt deleterious effects (Fig. 12A) and were detectable not only by 10X 3’ single-cell RNAseq and Stereo-seq, but also by FISH without amplification (Fig. 12B). As there was the possibility of non-specific probe binding in FISH staining, the specificity of the detection was verified by the correlation of FISH spots (Barcode 22 RNA, TdTomato RNA) with TdTomato protein, but not with the Actb control (Fig. 12C and 12D).
[0208] Probe-based spatial transcriptomics technologies tend to have lower plex and scalability as each unique RNA species require a working probe. Therefore, Stereo-seq, a high- resolution sequencing-based spatial technology was utilized to meet the needs of the Spatial Perturb-seq pipeline.
[0209] Power analysis done with powsimR to assess the statistical robustness of the platform showed FDR <10% and TPR >60-80% for 125-625 cells per perturbation (Fig. 11), demonstrating suitability of multiplexed experiments using this platform. As is the case with most spatial technologies, the accuracy of cell segmentation underpins the accuracy of Spatial Perturb-Seq, though significant advancements are being made in the field. Spatial Perturb-Seq unlocks a strategy to functionally interrogate genes at single-cell resolution within intact tissues, enabling the profiling of perturbation effects in both the target cells and their wildtype microenvironment.
[0210] Example 8
[0211] Analysis of neuronal phenotypes by Stereo-seq
[0212] For Spatial Perturb-seq, having a minimum single-cell resolution is requisite, as each genome / cell must be parsed individually. Being able to plex the whole transcriptome is advantageous as it provides an unbiased and comprehensive overview of gene expression, facilitating discovery which is important for phenotyping Perturb-seq experiments. Based on these two factors, Stereo-seq was chosen for the Spatial Perturb-seq platform. Stereo-seq is a sequencing-based spatial technology that utilizes DNA nanoball (DNB)-patterned arrays on lithographically etched chips. Tissue sections were placed on these chips, followed by in situRNA capture and sequencing. Each DNB had a coordinate identity (CID) barcode which is sequenced to obtain XY coordinates. Compared to other reported technologies, Stereo-seq had a higher spots number per area (400 spots per 100 pm2), achieving sub-cellular resolution. Another merit of Stereo-seq was its large field of view (1cm2chip), which afforded scal ability in the number of cells sequenced, better ensuring adequate sampling of each perturbation in Perturb-seq experiments.
[0213] As a first-of-its-kind study, an in vivo spatial perturb-seq platform was developed using scalable nucleotide barcodes with Stereo-seq, which allowed the study of non-cell autonomous functions of target genes, and the discovery of novel cell-cell communication within tissues. Whole transcriptome analysis of neuronal phenotypes was performed, and all 18 barcodes were detected with most of barcode-containing cells containing only one barcode. The MID and gene counts map were also shown (Fig. 7).
[0214] A key technical feature of the system was the detection of exogenous gRNA-linked barcodes by Stereo-seq, which has a different mRNA capture technology compared to dissociated scRNAseq. These barcodes were unlikely to be very highly expressed, several measures were implemented to ensure optimal barcode detection. Firstly, after RNA capture using the DNB probes and reverse-transcription, the cDNA was checked for presence of the barcodes by PCR. This ensures that the barcodes will be present in the sequencing libraries. Secondly, sequencing was performed to a saturation of >90%, increasing the likelihood of lowly-expressed genes being represented in the data. Thirdly, the barcodes were not enriched in the libraries prior to sequencing, as the results did not show a drastic increase in the number of barcode-positive cells, indicating that the sequencing saturation was sufficient. Enrichment could also introduce biasness and artefacts in the libraries. Lastly, as accurate cell segmentation is paramount for Spatial Perturb-seq platforms to maintain the genotype-phenotype relationship in the analysis of each cell, several different cell segmentation algorithms were employed. The DeepCell segmentation model was used, which used nuclear labelling and deep learning (with the network trained on the data obtained in the laboratory to increase accuracy due to laboratory-to-laboratory differences) to obtain robust cell segmentation with an average of 380 genes detected per cell.
[0215] In the Stereo-seq data, using known cell marker genes, 9 broad cell types were identified, with expected spatial domains (Fig. IF, Fig, 1H, Fig. 7B). Analysis of spatial hotspots that identify gene modules, revealed nine different modules, with module 1resembling oligodendrocytes, module 2 resembling choroid plexus cells in the ventricles, and module 3 resembling hippocampal neurons (Fig. 7E and 7F). The consistency in the spatial localization of our manually annotated cell types and the spatial localization of gene modules, provides confidence in the quality and segmentation of the data.
[0216] Among the recovered cell type, neurons were the most abundant. Neuronal cell-cell communication underlies proper functioning of the nervous system. The results demonstrated that the Spatial Perturb-Seq platform dissected these complex phenotypes in the context of functional genomics, focusing on cell-cell communication (CCC), and spatial gene regulatory network analysis (GRN).
[0217] Spatial Perturb-Seq enables the interrogation of cell-autonomous versus microenvironment effects of each gene perturbation within native intact tissues. To cluster and organize the spatial sequencing data, the Building Aggregates with a Neighborhood Kernel and Spatial Yardstick (BANKSY) framework was employed for embedding cells using their own transcriptome as well as that of their local neighbourhood, representing a cell’s state and its microenvironment, respectively. BANKSY was employed to select the region of interest, through adjusting the lambda parameter to vary the cells’ embedding. Lambda [0,1] is a numeric parameter in the BANKSY algorithm that adjusts the relative contribution of a cell’s own transcriptome and its neighbours’ transcriptome for its embedding. At low lambda values, BANKSY functioned in single cell typing mode, whereas at high lambda values, it identified spatial domains (Fig. 8). The cells were clustered using lambda 0.2 and 14 clusters were obtained (Fig. 8). This value of lambda allowed the achievement of spatially-informed celltyping and the goal of spatially informed cell-typing to identify hippocampal neurons (Fig. 1G- 11)
[0218] Neurons in BANKSY clusters 3 and 12 (CAI, CA2, CA3 and dentate gyrus neurons) that were in the same spatial domain (3720 cells) were focused on. Encouragingly, at lambda 0.2, embedding labels in both brain sections aligned well, meaning that clustering was driven largely by cell-type biology and not the spatial location of the cell on the chip (Fig. 8). Focusing on this niche, confirmation was required that perturbations are sparse and spread out in the tissue. To do this, BANKSY was performed with lambda 0.9 and a resolution of 50 to obtain 236 spatial domains with an average of 16 cells per domain. The results showed that 62% of these domains had no perturbations and 27% contain only 1 perturbation (Fig. 9). This ensuresthat subsequent analyses will not be confounded by multiple perturbations in the same spatial neighbourhood.
[0219] Among hippocampal neurons, all 18 barcodes were detected, with most barcodecontaining cells containing only 1 unique barcode as designed (Fig. 1J - IM). Cells with multiple barcodes (12.5% on average) were excluded from the analysis. Among all cells profiled by Stereo-seq, 2.1% are barcode-positive (Table 2), closely aligning with the 2.4% observed in the 10X dataset, suggesting comparable barcode detection sensitivity.
[0220] The spatial localization afforded by Spatial Perturb-Seq was used to examine loss- of function of Rbfox3, Sh3gl2, Clu, Rraga, Flcn, Ndufaf2 and Cfap410, Lrrk2, Srf n clusters 3 and 12 neurons. Neurons expressing barcodes linked to mSafe, Olig2, Trem2, Gfap, Stk39 were grouped as “Control”, as these genetic targets were not expressed in neurons and hence their perturbations were not expected to impact neuronal function.
[0221] The BANKSY framework allows modelling of both the cell’s own transcriptome and neighbours’ transcriptome (ie the cell’s microenvironment), doubling the number of features for each cell. Making use of this, the number of DEGs and the effect sizes between each group of perturbed neurons and the control group was quantified, to illustrate the effect of each genetic KO on itself and its microenvironment. There were no significant DEGs (defined as Ifc > 0.5 and fdr < 0.05) in Clu-KO neurons, consistent with the known function of Clu in mediating astrocytic function, even though it is also expressed in neurons There were also no significant DEGs in Rraga-KO neurons, potentially due to redundancy of Rraga with Rragb in the neuronal context. Compared to Ndufaf2-KO, Sh3gl2-KO neurons had a greater proportion of neighbouring DEGs compared to own transcriptome DEGs, suggesting that Sh3gl2 exerts a more significant influence on intercellular phenotypes, which can be expected giving its roles in synaptic vesicle dynamics and neurotransmitter release, as opposed to Ndufaf2’s role in metabolism. GSEA analysis of DEGs of neighbouring genes of Sh3gl2 -KO showed “synaptic transmission” as the top hit (fdr = 0.0018). This did not come up in analysis of Sh3gl2-K.O own transcriptome genes. These results showed viability of the platform in teasing apart cell autonomous effects and non-autonomous effects of a gene, which cannot be done in dissociated scRNA-seq. In addition, Egfr, shown to physically interact with Sh3gl2, was upregulated in Sh3gl2-KO (own transcriptome). This is consistent with observations in human tumour cells where SH3GL2 is frequently deleted, causing an upregulation of EGFR signalling and subsequent tumour growth. Similarly, GSEA analysis of DEGs of neighbouringgenes of Rbfox3-KO showed “neuron projection development”, suggestive of a role of Rbfox3 in the growth of neighbouring neurons.
[0222] To compare the effects of the different perturbations, the number of DEGs (p < 0.05, Ifc > 0.5) was quantified for each perturbation compared to control, and their average effect size (Fig. 13B) Further, as expected, perturbations caused a stronger cell autonomous effect, compared to their influence on the microenvironment (defined as 15 closest neighbouring cells). A caveat of this analysis is the assumption that the niche is uniform, whereas there may be spatial and microdomain-specific differences. Cfap410 perturbation caused the highest number of DEGs in its non-perturbed cellular neighbourhood, consistent with its role as a player in cilia function and synaptic plasticity in neurons. Overall, there were fewest DEGs in Rraga-KO neurons, potentially due to redundancy between Rraga and Rragb in neurons. The top 5 significantly upregulated genes for each perturbation were distinct among the 18 perturbations (Fig. 13C), highlighting the specificity of transcriptomic changes induced by each genetic KO.
[0223] Knock out of Lrrk2 (n = 29 cells) led to 213 DEGs observed (p < 0.05, Ifc > 0.5), the highest number amongst all perturbed genes (Fig. 13B - 13D). Lrrk2 is a key player in neuronal signalling and Parkinson’s disease pathology . Bel, a IncRNA found in dendrites that regulates translation of specific mRNAs in synapses, was downregulated upon Lrrk2 KO. Among the microenvironmental changes induced by Lrrk2 KO: downregulation of Sparc; upregulation of Vps351 (Vps35 is known to functionally interact with Lrrk2), DocklO (involved in dendritic spine formation), and Gpr37 (known to be associated with PD). The microenvironmental effects stemming from Lrrk2 KO can potentially underlie Lrrk2 -mediated pathology in neuronal signalling and Parkinson’s disease.
[0224] Another striking finding was that from Srf (n = 50 cells), a transcription factor regulating genes involved in neuronal growth and synaptic plasticity. Among the genes dysregulated upon Srf- KO, most were target genes identified through ChlP-seq data from the ENCODE Transcription Factors Target dataset (Ma'ayan Lab), though some did not reach the fdr threshold of < 0 05. Tn the Srf-KO microenvironment, Arhgapl2 (involved in cytoskeletal and actin dynamics) and Ssrpl (co-activator of Srf) were downregulated, Gadl (involved in the synthesis of neurotransmitter GABA) was upregulated.
[0225] The analysis showed that using BANKSY, the Spatial Perturb-seq pipeline can uncover the effect of a perturbation not only on a cell’s own transcriptome, but also on itsmicroenvironment. As many biological processes involve interactions between cells, such analyses bring us a step forward in understanding complex biological processes.
[0226] Example 9
[0227] Neuronal cell-cell communication
[0228] Cell-cell communication was further investigated, as it is essential in the function of neurons. To ensure accuracy of cell-cell communication analysis, cells were studied by keeping to the same neuronal and spatial niches, as ligand-receptor communication occur among neighbouring cells in proximity. The LIANA receptor-ligand framework was used to interrogate the dataset, and prioritize biologically important ligand-receptor pairs (Fig. 13E) and look at the signalling pathways of each perturbed neuron group to its unperturbed neighbours.
[0229] A 19% reduction in the SC A score for Lrpl signaling in the cellular neighbours of Lrrk2-KO neurons was observed, compared to control neurons. Lrpl mediates important neuronal functions including a- synuclein uptake, suggesting a possible mechanistic link between Lrrk2 activity and PD pathogenesis. The data showed a reduction in Nlgnl signaling in Srf-KO neighbours, the importance of Srf in synaptic function. To visualize these cell communication networks, individual cells showing Efnb3-Epha4 signaling in Srf-KO and control-KO microenvironments were shown on a spatial plot (Fig. 13F). Source cells and their 15 closest neighbours (target cells), were shown according to their XY coordinates, with expression levels of the ligand represented by “X” (source cells) and expression levels of the receptor surrounding the ligand (target cells). These analyses highlighted the ability of Spatial Perturb-Seq to interrogate genetic determinants of intercellular cell communication signaling pathways within the native tissue.
[0230] To test the compatibility of Spatial Perturb-Seq in an orthogonal, probe-based platform, the Xenium platform with custom probes against the 450bp barcodes (Table 3) was utilized. Transcript density maps showed expected cell densities across the tissue (Fig. 14C). Fig. 14B shows visualization of cell segmentation boundaries along with individual molecules of detected barcodes. TdTomato transcript localized to the expected region of the neuronal cell bodies in the injected side of the hippocampus (Fig. 14C). All 18 barcodes were robustly detected using the Xenium platform (Fig. 15), demonstrating their dual functionality for sequencing-based and imaging-based spatial platforms. Among recovered cell types, neurons were the most abundant (Fig. 14D), with cell-type identities annotated via canonical markerexpression (Fig. 14E). Excitatory neurons showed the highest tropism (Fig. 15B-15E). Because sequencing-based RNA-seq captures the 3 ' end of transcripts, probe-based methods hybridizing to multiple regions along the transcript may offer higher sensitivity for detecting CRISPR-based knock-out. Using the Xenium platform, the results showed that target gene expression in neurons decreased with increasing expression of associated barcodes, with the average global transcript levels not affected (Fig. 14F and Fig. 16). DEGs were investigated for each barcode compared to the rest of the barcodes Compared to probe panels with -200 genes, whole transcriptome profding is better suited for unbiased discovery of perturbation effects. Among up-regulated genes, only nine out of 18 perturbations led to DEGs with the probe panel used (Fig. 14G). For Lrrk2-KO compared to all other perturbations, Cacna2d2, Kctd8, Tmem255a, Syt6, Calb2, Pdella, Tmeml63, and Anol were upregulated. Lrrk2-KO is known to disrupt calcium signaling in neurons, leading to dysregulation of genes involved in calcium homeostasis. Consequently, genes such as Cacna2d2 (a subunit of voltage-dependent calcium channels), Calb2 (calbindin 2, a calcium-binding protein), and Anol (a calcium- activated chloride channel) may exhibit altered expression in LRRK2 KO cells. Next, the transcriptomic effects of Lrrk2-KO and Srf-KO on their neighbouring cells were investigated to validate the top four observed neighbour DEGs in Stereo-seq. While a similar magnitude in fold change was not observed perhaps due to different dynamic ranges of probe-based and sequencing-based platforms, a similar trend of up or down-regulation was seen (Fig. 14).
[0231] Table 3. Xenium metrics
[0232] Spatial Perturb-Seq not only enables us to assess how a perturbation influences its cellular neighbours and microenvironment, but also how the microenvironment influences the response to perturbation. By comparing barcode-positive and barcode-negative cells within distinct hippocampal regions, the dentate gyrus (DG) and CAI neurons, a substantial overlap was observed in both upregulated and downregulated genes, indicating shared transcriptional responses in those two regions. However, some region-specific differences were also observed, suggesting that the local tissue context can affect the transcriptional outcome of a given perturbation (Fig. 17).
[0233] Several recent technologies represent important steps toward enabling spatially resolved and multiplexed CR1SPR screens, each with distinct strengths and limitations (Table 4). Perturb-FISH offers high spatial resolution but due to its reliance on targeted probes, it lacks whole-transcriptome coverage, which limits its ability for unbiased screens. Perturb-map relies on the antibody-based Pro-Code system, limiting scalability and requiring imaging-based readouts and complex instrumentation. Perturb-DBiT, which is microfluidics-based, requires more extensive technical optimization and does not offer single-cell resolution. These methodsunderscore growing momentum in the field and the critical need for functional genomics with spatial and microenvironmental context.
[0234] Table 4. Comparison of methods
[0235] Cell-cell communication scores were calculated by using the cells’ spatial coordinates and the co-expression of the ligand and receptor genes. Each perturbation was examined for its impact on 15 most significant ligand receptor (LR) pairs. As expected, there were no overt differences in these LR signalling pairs in Clu and Fasn knockout neurons with their neighbours, as those genes are not expected to play a key role in neuronal signalling. The results showed that Srf-KO increased NLGN2->NRXN2 synaptic communication with wildtype neighbours and decreased NRXN2->NLGN2communication. This is consistent with the role of Srf as a transcription factor that controls activity-driven gene expression in neurons by both repression and activation.
[0236] Focussing on Wnt signalling, the cell-cell communication analysis showed that Lrrk2 knockout neurons display increased Wnt7a-Frizzled-LRP5 / 6 signalling with their neighbours compared to controls. Wnts regulate many aspects of neural circuitry, includingcircuit formation, axon guidance, spine growth, and synaptic function48. The importance of Lrrk2 in neurons is also underscored by the fact that Lrrk2 are among the most common genetic risk factors for Parkinson’s disease. Consistent with this data, Lrrk2 kinase is known to affect a range of cellular processes, but most significantly Wnt signalling. Loss of Lrrk2 increased Wnt signalling, with pathogenic LRRK2 variants in humans found to be gain-of-function mutations that impair Wnt signalling in neurons, leading to the development of Parkinson’ s disease This data thus added to the evidence that Wnt signalling is part of the mechanistic link between Lrrk2 and PD, showing that loss of Lrrk2 kinase in neurons caused increased Wnt7a signalling between itself and its wildtype neighbours. It is unlikely that Lrrk2 is the only protein controlling Wnt signalling in neurons. The analysis also showed that loss of Sh3gl2 led to an increase in Wnt5b signalling between Sh3gl2 neurons and its wildtype neighbours. Unlike Wnt7a, Wnt5b signals via the non-canonical WNT / p-catenin pathway. While a link between Sh3gl2 and Wnt signalling is not characterized, loss of Sh3gl2 function is known to increase EGFR signalling43, which in turn known to transactivate Wnt signalling. This data thus suggests a mechanistic link underlying the role of Sh3gl2 as a tumour suppressor in the nervous system: through the suppression of non-canonical Wnt signalling. These data and analyses highlight the ability of Spatial Perturb-Seq to interrogate genetic determinants of intercellular signalling pathways within the native spatial context
[0237] Example 10
[0238] Neuronal gene regulatory networks
[0239] The spatial gene regulatory networks (GRNs) were subsequently investigated, making use of the spatially-resolved single-cell gene expression profdes, combined with available databases on transcription factor-target regulatory relationships. The dataset against databases of pre-defined transcription factors (TF) was interrogated, predicted transcription factors binding sites and their target genes, and motif to transcription factors databases (Aerts, S. 2002 cisTarget: Integrative Motif Discovery and Regulatory Network Analysis Resources. Retrieved from https: / / resources.aertslab.org / cistarge and Dhainaut et al. Spatial CRISPR genomics identifies regulators of the tumor microenvironment. Cell, 185(7), 1223-1239. e20, 2022). With this, the regulatory network inference was calculated for each of the 3720 neurons in clusters 3 and 12, identifying 15 regulons: Bcl6, Cebpd, Dbp, Egr3, Foxol, Jun, Junb, Lhx2, Mef2c, Mef2d, Pou3fl, Rfx3, Smad3, Srebf2, Thra. For each regulon and each cell, the Area Under the Curve (AUC) enrichment score represents the expression of genes in that signature.With this, the relative AUC enrichment score at an intracellular level and neighbourhood level was plotted. For neighbourhood level, scores were the average of the perturbed cell’s 15 neighbours, as identified by BANKSY.
[0240] In general, genes have more robust cell autonomous than non-cell autonomous functions. Therefore, more subtle changes in TF activity at a domain level compared to within individual perturbed cells were seen. The data showed that Lrrk2-KO resulted in lowered Foxol scores intracellularly, consistent with well -documented evidence that Lrrk2 directly phosphorylates and activates Foxol, which in turn affects survival of dopaminergic neurons. Evidence was added that Foxol is a mechanistic link for Lrrk2-related Parkinson’s Disease. In addition, Flcn-KO reduced the effect of Srebf2, in line with reports showing that loss of Flcn may affect SREBF2 -mediated transcriptional regulation of genes involved in cholesterol biosynthesis and lipid metabolism. Consistent with an observed lack of significant DEGs in Clu-KO neurons, the results showed Clu-KO does not affect GRNs. Interestingly, loss of function of Cfap410 in neurons caused an increase in Rfx3 signalling in its neighbours but not itself. Cfap410 is required formaintenance of cilia function in neurons, whereas Rfx3 promotes ciliogenesis, suggesting that loss of ciliary signalling could stimulate cilia growth in the microenvironment. Lastly, consistent with its function as a transcription factor, ,S> / -KO led to significant changes in gene regulatory networks intracellularly and in its neighbours.
[0241] The above analysis revealed transcription factor networks that were dysregulated in a perturbed cell or its microenvironment. The relationship between CCC and GRNs is bidirectional interconnected: signalling pathways activated by extracellular signals often converge on transcription factors, while certain transcription factors may regulate the expression of genes involved in extracellular signalling. Therefore, future analysis with the integration of CCC and spatial GRNs will shed light on L-R-TF interactions, and how L-R-TF interactions govern the behaviour of cells within a multicellular tissue niche.
[0242] Table 5. Summary of sequence listing.
[0243] Equivalents
[0244] The foregoing examples are presented for the purpose of illustrating the invention and should not be construed as imposing any limitation on the scope of the invention. It will readily be apparent that numerous modifications and alterations may be made to the specific embodiments of the invention described above and illustrated in the examples without departing from the principles underlying the invention. All such modifications and alterations are intended to be embraced by this application.
Claims
Claims1. A polynucleotide comprising: a) a polynucleotide sequence complementary to a polynucleotide sequence encoding a gene of interest or a segment thereof; and b) a barcode that is unique to the gene of interest comprising a polynucleotide sequence, wherein i. the polynucleotide sequence of the barcode is between about 20 to about 500 nucleotides, ii. wherein the barcode comprises an edit distance of more than 16 nucleotides between segments, iii. wherein the barcode is devoid of a polynucleotide sequence encoding a stop codon, and iv. wherein the barcode is devoid of a polynucleotide sequence comprising 4 or more identical consecutive nucleotides.
2. The polynucleotide according to claim 1, further comprising a poly-Atail.
3. The polynucleotide of claim 1 or 2, further comprising a fluorescence tag, optionally wherein the fluorescence tag is TdTomato.
4. The polynucleotide of any one of claims 1-3, wherein the complementary polynucleotide sequence comprises one or more guide RNA (gRNA), optionally wherein the one or more gRNA is CRTSPR-gRNA.
5. The polynucleotide of any one of claims 1-4, wherein the complementary polynucleotide sequence comprises three distinct gRNA targeted to the gene of interest.
6. The polynucleotide of any one of claims 1-5, wherein the complementary polynucleotide sequence is operably linked to one or more promoters.
7. The polynucleotide of claim 6, wherein the one or more promoters is an RNA polymerase promoter, optionally wherein the RNA polymerase promoter is a U6 promoter.
8. The polynucleotide of any one of claims 1-7, wherein each segment of the barcode comprises about 50 nucleotides.
9. The polynucleotide of any one of claims 1-8, wherein the polynucleotide sequence of each segment of the barcode has an edit distance of more than 16 relative to endogenous human and mouse sequences.
10. A library comprising a plurality of polynucleotides of any one of claims 1-9.
11. The library of claim 10, wherein the polynucleotide sequence of each segment of the barcode in each of the plurality of polynucleotides has an edit distance of more than 16 nucleotides12. A method of obtaining spatial and transcriptomic information of a population of cells, wherein the method comprises: a) delivering a plurality of polynucleotides according to claims 1-9 or the library according to claim 10 or 11 to the population of cells, wherein each of the plurality of polynucleotides perturbs a distinct gene of interest in the population of cells; b) obtaining an intact tissue sample comprising the population of cells; and c) detecting the plurality of polynucleotides in the population of cells in the intact tissue sample, thereby obtaining spatial and transcriptomic information of the population of cells.
13. The method of claim 12, wherein the plurality of polynucleotides or the library is delivered via a viral vector.
14. The method of claim 13, wherein the viral vector is an adeno-associated virus (AAV) vector, optionally wherein the AAV vector is AAV-PHP.eB15. The method of any one of claims 12-14, wherein perturbation of the gene of interest in the population of cells comprises one or more of deletion of the gene, reduction in expression of the gene, introduction of the gene, deletion of one or more nucleotides in the gene of interest, substitution of one or more nucleotides in the gene of interest and insertion of one of more nucleotides in the gene of interest.
16. The method of any one of claims 12-15, wherein perturbation of the gene of interest in the population of cells comprises substitution, deletion, or insertion of one or more nucleotides in the polynucleotide sequence of the gene of interest.
17. The method of any one of claims 12-16, wherein the plurality of polynucleotides or library is delivered to about 100% or less of the cells in the population of cells;optionally about 50% or less of the cells in the population of cells; optionally at least one cell in the population of cells.
18. The method of any one of claims 12-17, wherein the gene of interest comprises a polynucleotide sequence that is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% complementary to the polynucleotide sequence of the one or more gRNA in the plurality of polynucleotides or in the library.
19. The method of any one of claims 12-18, wherein the plurality of polynucleotides or the library is detected optically or by sequencing.
20. The method of claim 19, wherein the optical and sequencing detection of the plurality of polynucleotides or the library is on different tissue samples.
21. The method of any one of claims 12-20, wherein the spatial information comprises location, arrangement and distance of one or more cells in the population of cells in the intact tissue sample.
22. The method of any one of claims 12-21, wherein the transcriptomic information comprises gene expression of one or more cells in the population of cells in the intact tissue sample.
23. The method of any one of clams 12-22, further comprising analysing the spatial and transcriptomic information using algorithm-based modelling.
24. The method of any one of claims 12-23, wherein the gene of interest is a gene in a pathway or intracellular network.
25. The method of any one of claims 12-23, wherein the method is performed in vitro, in vivo or ex vivo.
26. The method of any one of claims 12-24, wherein the intact tissue sample is obtained from a subject, optionally wherein the subject is an animal, optionally wherein the subject is a mouse.
27. The method of claim 25, wherein the subject expresses a Cas protein.
Citation Information
Patent Citations
Robust quantification of single molecules in next-generation sequencing using non-random combinatorial oligonucleotide barcodes
US11661597B2