High resolution spatial DNA chip and methods of manufacturing thereof

High-resolution spatial DNA chips with 2 x 2 μm features address the limitations of current spatial transcriptomics by providing detailed spatial gene expression mapping with high throughput and efficiency, achieving 97.33% decoding and 69.37% unique reads.

WO2025151445A1PCT designated stage expired Publication Date: 2025-07-17CENTRILLION TECHNOLOGY HOLDINGS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/010614
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2025-01-07
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Current spatial transcriptomics technologies face limitations in resolution and efficiency, with commercially available microarrays having feature spacing of 10-55 μm, necessitating additional single-cell sequencing for detailed resolution and capture efficiency, while imaging-based methods are low throughput.

Method used

Development of high-resolution spatial DNA chips with feature sizes of 2 x 2 μm, manufactured using photolithographic methods on silicon wafers and transferred to hydrogels, incorporating barcode probes for precise spatial mapping of nucleic acids, enabling high-throughput sequencing without the need for additional single-cell sequencing.

Benefits of technology

The high-resolution spatial DNA chips provide detailed spatial distribution of gene expression with 100% surface coverage and high sequencing saturation, capturing over 8000 genes per region with 97.33% decoding rate and 69.37% unique reads, facilitating accurate spatial transcriptome analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025010614_17072025_PF_FP_ABST
    Figure US2025010614_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Spatial transcriptomics has showcased its efficacy in deciphering the intricate relationships between individual cells and tissues. We present spatial transcriptomics data using a novel high-resolution DNA chip with a capture region feature size of 2 x 2pm. Feature-to- feature gap space is zero, maximizing the capture area. Chips are manufactured at wafer scale using photolithography and are transferred to hydrogels, making them compatible with existing sample preparation and analysis workflows for fresh frozen or paraffin-embedded samples. For this report, we examined a fresh frozen sample from adult mouse liver. Using a bin size of 10, representing a 20 pm x 20 pm capture area, at 69% sequencing saturation, we obtained over 600 million unique mapped reads, the median number of genes captured was over 8000 per region, which demonstrated potential for obtaining additional unique reads with deeper sequencing. This high-resolution mapping of liver cell types and visualization of gene expression patterns demonstrates significant advances in spatial sequencing technology.
Need to check novelty before this filing date? Find Prior Art

Description

HIGH RESOLUTION SPATIAL DNA CHIP AND METHODS OF MANUFACTURING THEREOFCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 618,894, filed January 8, 2024, the disclosure of which is incorporated by reference herein in its entirety.SEQUENCE LISTING

[0002] This application contains an electronic Sequence Listing which has been submitted in XML file format with this application, the entire content of which is incorporated by reference herein in its entirety. The Sequence Listing XML file submitted with this application is entitled “14791-040-228_SEQ_LISTING.xml”, was created on January 6, 2025, and is 7,132 bytes in size.FIELD

[0003] The present disclosure relates generally to methods and systems for analyzing the spatial distribution of nucleic acids using high-resolution DNA chips.BACKGROUND

[0004] The high-throughput exploration of gene expression profiles has transformed our understanding of cellular processes and molecular mechanisms. Traditional transcriptomics methods, such as bulk RNA sequencing, provide invaluable insights into global gene expression patterns within a tissue or cell population. However, they often overlook the intricate spatial organization of cells within complex biological systems. The advent of spatial transcriptomics technologies has revolutionized our ability to capture gene expression data within its native spatial context, enabling us to delve deeper into the heterogeneity and interactions of cells within tissues.

[0005] Spatial transcriptomics bridges the gap between high-resolution spatial information and transcriptomic data, allowing researchers to unravel the spatial distribution of gene expression across tissue sections. Molecular spatial analysis methods are divided into twoapproaches: imaging-based approaches and sequencing-based approaches. Imaging-based approaches suffer from low throughput, only capable of capturing the expression of a handful of transcripts per imaging run, while sequencing-based approaches suffer from low resolution. The instant application demonstrates a significant improvement in sequencing-based spatial sequencing.

[0006] Microarrays for spatial trancriptomics, also called spatial DNA chips, are manufactured as DNA microarrays with each “spot” containing a poly-T capture sequence for mRNA and a unique barcode sequence to indicate the feature to which the downstream sequencing reads correspond. Due to limitations in the resolution of commercially available microarrays employed in spatial transcriptomics (10-55 pm), additional single-cell sequencing data has been required in order to provide detailed resolution and capture efficiency. Photolithographic microarray fabrication methods are currently used for the production of commercial microarrays with feature spacing in the range of 4-10 pm (Hoff, K. et al. Langmuir 2021, 37 (16), 4763-4771). We have previously demonstrated technological advances in waferscale manufacturing of 3’-up functional DNA chips in hydrogels with sub-micrometer resolution (Costa, J. A. et al. ACS App Mater. Interfaces 2019). Using photoresist technology (proprietary) that uses photoacid generated chemistry in a polymer matrix to spatially deblock DMT protecting groups (Costa, J.A. et al. 2019), incredibly high-resolution, completely confluent DNA chips can be manufactured at semi-conductor scale.

[0007] The instant application demonstrates the use of these chips with tissue sections generated from fresh frozen adult mouse liver.SUMMARY

[0008] The present disclosure provides a system for analyzing a sample for the spatial distribution of at least one target molecule. In some embodiments, the system may comprise a chip comprising a plurality of features, wherein each feature of the plurality of features comprises a plurality of barcode probes, and wherein each feature of the plurality of features is at a distinct location on the chip associated with two or more coordinates; and a computer processorcoupled to the chip and programmed to (i) measure at least one signal while the chip is in contact with the sample; and (ii) determine, based on the at least one signal, the spatial distribution of said at least one target molecule at the plurality of features.

[0009] The present disclosure also provides a system for analyzing a sample for the spatial distribution of at least one target molecule comprising a chip comprising a plurality of features, wherein each feature of the plurality of features is at a distinct location on the chip associated with two or more coordinates; and a plurality of barcode probes, wherein each barcode probe is attached to a feature, wherein the plurality of barcode probes are capable of binding to at least one target molecule, and wherein the location of each of the barcode probes is identified by the two or more coordinates.

[0010] In some embodiments, the plurality of barcode probes are nucleic acid probes. In some embodiments, the system comprises between approximately 5 million and approximately 10 million features, between approximately 10 million and approximately 15 million features, between approximately 15 million and approximately 20 million features, between approximately 20 million and approximately 25 million features, between approximately 25 million and approximately 30 million features, between approximately 30 million and approximately 35 million features, between approximately 35 million and approximately 40 million features, between approximately 40 million and approximately 45 million features, and between approximately 45 million and approximately 50 million features. In some embodiments, the system comprises approximately 25 million features.

[0011] In certain embodiments, each feature comprises approximately 1 million to approximately 3 million barcode probes, approximately 3 million to approximately 5 million barcode probes, approximately 5 million to approximately 7 million barcode probes, approximately 7 million to approximately 9 million barcode probes, or approximately 9 million to approximately 11 million barcode probes. In some embodiments, each feature comprises approximately 5 million barcode probes.

[0012] In certain embodiments, each of the features is a square of approximately 0.25 pm by 0.25 pm to approximately 0.75 pm by 0.75 pm in size, approximately 0.75 pm by 0.75 pm to approximately 1.25 pm by 1.25 pm in size, approximately 1.25 pm by 1.25 pm to approximately 1.75 pm by 1.75 pm in size, approximately 1.75 pm by 1.75 pm to approximately 2.25 pm by 2.25 pm in size, approximately 2.25 pm by 2.25 pm to approximately 2.75 pm by 2.75 pm in size, approximately 2.75 pm by 2.75 pm to approximately 3.25 pm by 3.25 pm in size, approximately 3.25 pm by 3.25 pm to approximately 4.75 pm by 4.75 pm in size, or approximately 4.75 pm by 4.75 pm to approximately 5.25 pm by 5.25 pm in size. In certain embodiments, each of the features is a square of approximately 2 pm by about 2 pm.

[0013] In certain embodiments, at least approximately 50%, at least approximately 60%, at least approximately 70%, at least approximately 80%, at least approximately 90%, or at least approximately 100% of the surface area of the chip is covered by at least one barcode probe. In certain embodiments, at least approximately 90% of the surface area of the chip is covered by at least one barcode probe.

[0014] In certain embodiments, at least two of the plurality of distinct locations are adjacent. In certain embodiments, there is no interstitial space between adjacent features.

[0015] The present disclosure also provides a method for detecting the spatial distribution of a plurality of target molecules within a sample, comprising: (a) providing a substrate comprising a plurality of features, wherein each of the plurality of features comprises two or more coordinates, and wherein at least 90% of the surface area of the plurality of features is covered by at least one barcode probe; (b) contacting the surface with the sample, wherein the target molecules are capable of coupling with the barcode probes; and (c) determining spatial distribution of the target molecules on the substrate.

[0016] The present disclosure also provides a spatial distribution pattern of target molecules from a sample, wherein the spatial distribution pattern comprises two or more coordinates for each of a plurality of features on a chip, wherein at least 90% of the surface area of the chip iscovered by at least one barcode probe, and wherein the barcode probe is capable of binding to the target molecules.

[0017] The present disclosure also provides a method of forming a pattern of oligonucleotides on a substrate having a plurality of functional groups, comprising: forming a vertical region with an X barcode sequence comprising a first nucleotide by at least coupling the first nucleotide to a portion of the plurality of functional groups in a first vertical exposed region of the substrate; and forming a second vertical region having an X barcode sequence comprising the first nucleotide and a second nucleotide by at least coupling the second nucleotide to a portion of the plurality of functional groups in a second vertical exposed region of the substrate that overlaps with the first vertical region.

[0018] In some embodiments, the method further comprises forming a third vertical region having an X barcode sequence comprising the first nucleotide, the second nucleotide, and a third nucleotide by at last coupling the third nucleotide to a portion of the plurality of functional groups in a third vertical exposed region of the substrate that overlaps with the second vertical region. In some embodiments, the method further comprises forming, from a non-overlapping region between the third vertical exposed region and the second vertical region, a fourth vertical region having an X barcode sequence that comprises the first nucleotide and the second nucleotide but not the third nucleotide. In some embodiments, a location within each vertical region is identified by a corresponding X barcode sequence.

[0019] In some embodiments, the method further comprises forming a first horizontal region with a Y barcode sequence comprising a third nucleotide by at least coupling the third nucleotide to a portion of the plurality of functional groups in a first horizontal exposed region of the substrate; and forming a second horizontal region with a Y barcode sequence comprising the third nucleotide and a fourth nucleotide by at least coupling the fourth nucleotide to a portion of the plurality of functional groups in a second horizontal exposed region of the substrate that overlaps with the first horizontal region.

[0020] In some embodiments, the method further comprises forming a third horizontal region having a Y barcode sequence comprising the third nucleotide, the fourth nucleotide, and a fifth nucleotide by at least coupling the fifth nucleotide to a portion of the plurality of functional groups in a third horizontal exposed region of the substrate that overlaps with the second horizontal region.

[0021] In some embodiments, the method further comprises forming, from a nonoverlapping region between the third horizontal exposed region and the second horizontal region, a fourth horizontal region having a Y barcode sequence that includes the third nucleotide and the fourth nucleotide but not the fifth nucleotide. In some embodiments, a location within each horizontal region of the substate is identified by a corresponding Y barcode sequence.

[0022] In some embodiments, an intersection between a vertical region and a horizontal region comprises at least one feature whose location is identified by a corresponding X barcode sequence and Y barcode sequence.

[0023] In some embodiments, according to the method disclosed herein, the coupling comprises: (i) forming a first photoresist layer by applying a photoresist composition onto an underlying layer of a substrate comprising the plurality of functional groups, wherein the plurality of functional groups are protected by protective groups; (ii) exposing a dose of light through a patterned mask onto the substrate; (iii) removing the protective groups on a section of the plurality of functional groups within at least one exposed region of the substrate; and (iv) contacting the functional groups within the at least one exposed region of the substrate with a nucleotide reagent; thereby coupling a fraction of the functional groups within the at least one exposed region of the substrate with a nucleotide.

[0024] In some embodiments, adjacent spatially-defined vertical and / or horizontal regions in the substrate have barcode sequences that differ by a single nucleotide. In some embodiments, at least 90% of the surface area of the substrate is covered with at least one barcode probe having a unique spatially-defined X barcode sequence and a unique spatially-defined Y barcode sequence.BRIEF DESCRIPTION OF THE DRAWINGS

[0025] FIGS. 1A-1F. Chip Design and Performance. (FIG. 1A) Manufacturing of individual chips using photolithography. Vertical manufacturing of the zipcode design is shown, giving X positional information. Y positional information is achieved through the same methods with horizontal light-directed synthesis. (FIG. IB) Chip Design. High-resolution spatial DNA chips are first manufactured using light-directed synthesis on silicon wafers then transferred to hydrogels as previously described (Costa, J.A. et al., 2019). Capture regions are confluent 2 x 2 pm squares. (FIG. 1C) Assay workflow. First, the sample is permeabilized and RNA is hybridized to the chips. The next steps of the assay workflow are reverse transcription using a template-switching oligonucleotide (TSO), and second strand synthesis. The second strand is stripped from the chip and used in library preparation and downstream sequencing. (FIG. ID) Structure of final sequencing library. (FIG. IE) Chip performance. Number of UMI and number of genes captured for binning sizes of 1, 5, 10, and 16. (FIG. IF) Sequence saturation. Duplicated reads, defined as those sharing identical spatial coordinates, UMI, and gene identifiers with prior reads, are excluded to calculate the number of unique reads. The plot shown displays the saturation rate — the ratio of duplicated reads to the total number of reads — emphasizing the incremental accumulation of redundancy with the progression of sequencing depth.

[0026] FIG. 2. Spatial expression of cell-specific markers. For comparison, an H&E staining of a neighboring section is shown in panel A and clustering results from FIGS. 3A-3C are shown in panel B. The key on the right displays cell types determined through clustering.Transcriptomics results are shown for markers for periportal hepatocytes (CPS1, CYP2F2), pericentral hepatocytes (GLUL, OAT), and midlobular hepatocytes (HAMP) are shown in panels C through H. Panel C offers a composite view, merging the information from Panels D through H. Additionally, Panels I through L display markers found in other liver cells. SPP1 is a marker found in biliary epithelial cells (BECs), Kupffer cells, and activated hepatic stellate cells (HSCs). JCHAIN is expressed in immune cells and liver endothelial cells. Immunoglobulin IGKC is expressed by immune cells. Panel l is a merged representation of Panel I through L.

[0027] FIGS. 3A-3C Successful capture of the mouse liver transcriptome. (FIG. 3 A) Reconstructed map of spatial expression based on unsupervised clustering data. Hash marks represent barcode positions with each barcode being 0.2pm in height and width. (FIG. 3B) UMAP plot for clustering data. (FIG. 3C) Differential metabolic profile analysis of cell clusters. Full expression levels are provided in Supplementary materials.

[0028] FIGS 4A-4D. Cell type composition and U-CIE. (FIG. 4A) Cell2Location analysis of cell types. Cholangeocytes, endothelial cells, fibroblasts, hepatocytes, Kuppfer cells, macrophages and monocytes, and T cells are color-coded and plotted using the spatial zipcode. (FIGS. 4B-4D) Clustering data using UMAP and U-CIE.

[0029] FIG. 5. The bioinformatics pipeline. Custom bioinformatics pipeline for the processing of spatial transcriptomics data obtained from paired-end FASTQ files.

[0030] FIGS 6A-6B Flowcharts illustrating methods of forming a pattern of oligonucleotides on a substrate. (FIG. 6A) Method of forming a pattern of oligonucleotides on a substrate by forming vertical regions with X barcode sequences. (FIG. 6B) Method of forming a pattern of oligonucleotides on a substrate by forming vertical regions with Y barcode sequences.DETAILED DESCRIPTION

[0031] The present disclosure provides systems and methods for analyzing a sample for the spatial distribution of at least one target molecule. Also provided herein is a spatial distribution pattern of target molecules from a sample. The systems disclosed herein may comprise a chip comprising a plurality of features, wherein each feature of the plurality of features comprises a plurality of barcode probes, and a computer processor coupled to the chip and programmed to measure at least one signal while the chip is in contact with the sample; and determine, based on the at least one signal, the spatial distribution of said at least one target molecule at the plurality of features. The systems disclosed herein may have no interstitial space between adjacent features. The systems may comprise chips with approximately at least 50% of the surface area of the chip is covered by at least one barcode probe.

[0032] Also disclosed herein are methods for detecting the spatial distribution of a plurality of target molecules within a sample, and the spatial distribution pattern of target molecules from a sample. Further disclosed herein are methods of forming a pattern of oligonucleotides on a substrate having a plurality of functional groups, comprising forming vertical regions with an X barcode sequence and / or horizontal regions with a Y barcode sequence.1. Definitions

[0033] The term “nucleotide,” as used herein, generally refers a molecule that can serve as the monomer, or subunit, of a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). A nucleotide can be a deoxynucleotide triphosphate (dNTP) or an analog thereof, e.g., a molecule having a plurality of phosphates in a phosphate chain, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10 phosphates. A nucleotide can generally include adenosine (A), cytosine (C), guanine (G), thymine (T) and uracil (U), or variants thereof. A nucleotide can include any subunit that can be incorporated into a growing nucleic acid strand. Such subunit can be an A, C, G, T, or U, or any other subunit that is specific to one or more complementary A, C, G, T or U, or complementary to a purine (i.e., A or G, or variant thereof) or a pyrimidine (i.e., C, T or U, or variant thereof). A subunit can enable individual nucleic acid bases or groups of bases (e.g., AA, TA, AT, GC, CG, CT, TC, GT, TG, AC, CA, or uracil-counterparts thereof) to be resolved. A nucleotide may be labeled or unlabeled.

[0034] The term “oligonucleotide” as used herein generally refers to a nucleotide chain. In some cases, an oligonucleotide is less than 200 residues long, e g., between 15 and 100 nucleotides long. The oligonucleotide can comprise at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 bases. The oligonucleotides can be from about 3 to about 5 bases, from about 1 to about 50 bases, from about 8 to about 12 bases, from about 15 to about 25 bases, from about 25 to about 35 bases, from about 35 to about 45 bases, or from about 45 to about 55 bases. The oligonucleotide (also referred to as “oligo”) can be any type of oligonucleotide (e.g., a primer). Oligonucleotides can comprise natural nucleotides, non-natural nucleotides, or combinations thereof.

[0035] As used herein, the term “polymer” generally refers to any kind of natural or nonnatural large molecules, composed of multiple subunits. Polymers may comprise homopolymers, which contain a single type of repeating subunits, and copolymers, which contain a mixture of repeating subunits. In some cases, polymers are biological polymers that are composed of a variety of different but structurally related subunits, for example, polynucleotides such as DNA or RNA composed of a plurality of nucleotide subunits.

[0036] The terms “about” or “approximately” as used herein generally refers to + / - 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the designated amount.

[0037] The term “immobilization” as used herein generally refers to forming a covalent bond between two reactive groups. For example, polymerization of reactive groups is a form of immobilization. A carbon-to-carbon covalent bond formation is an example of immobilization. Genetic information can be utilized in a myriad of ways with the advent of rapid genome sequencing and large genome databases. One of such applications is oligonucleotide arrays. The general structure of an oligonucleotide array, or commonly referred to as a DNA microarray or DNA array or a DNA chip (or RNA microarrays, RNA arrays or RNA chip), is a well-defined array of spots or addressable locations on a surface. Each spot can contain a layer of relatively short strands of DNA or RNA called a “probe” or “capture probe” or “barcode probe” (e.g., Schena, ed., “DNA Microarrays A Practical Approach,” Oxford University Press; Marshall et al. (1998) Nat. Biotechnol. 16:27-31 ; each incorporated herein by reference). There are at least two technologies for generating arrays. One is based on photolithography (e.g. Affymetrix) while the other is based on robot-controlled inkjet (spotbot) technology (e.g., Arrayit.com). Other methods for generating microarrays are known and any such known method may be used herein.

[0038] As used herein, open terms, for example, “comprise”, “contain”, “include”, “including”, “have”, “having” and the like refer to comprising unless otherwise indicates.

[0039] As used herein, the term “embedding” and “a string of synthetic steps” generally refer to a series of active and inactive steps designed for forming an individual polymer on the substrate and can be used interchangeably. For example, in cases where light-directed syntheticmethods are employed, the “embedding” refer to a series of exposure and non-exposure steps.The term “embedding” herein refers to the binary sequence that specifies when the barcode's position on the chip is exposed to light during the chip synthesis via photolithography“Embedding” is described in Raman Kumar, M.; Naveen Kumar, V. A Numerical Representation Method for a DNA Sequence Using Gray Code Method. In Soft Computing for Problem Solving,' Das, K. N., Bansal, J. C., Deep, K., Nagar, A. K., Pathipooranam, P., Naidu, R. C., Eds.; Advances in Intelligent Systems and Computing; Springer: Singapore, 2020; pp 645-654. https: / / doi org / 10.1007 / 978-981 - 15-0184-5_55.

[0040] As used herein, the term “edit distance” generally refers to the minimum number of changes (such as insertions, deletions, substitutions and translocations) to a nucleotide sequence needed to convert one polymer (e.g., polynucleotide) into another. Edit distance is described, for example, in US Application No. 16 / 614,677.

[0041] As used herein, the term “adjacent” or “adjacent to,” includes “next to,” “adjoining,” and “abutting.” In one example, a first location is adjacent to a second location when the first location is in direct contact and shares a common border with the second location (e.g., there is no space between the two locations). In some cases, the adjacent is not diagonally adjacent.

[0042] As used herein, the term “zipcode” or “zipcode sequence” generally refers to a known, determinable, and / or decodable sequence, such as, for example, a nucleic acid sequence (e.g., a DNA sequence or RNA sequence), a protein sequence, and a polymer sequence (including synthetic polymers, carbohydrates, lipids, etc.), that allows the identification of a specific location of the sequence, e.g., the nucleic acid, in one, two or multiple dimensional spaces. A zipcode can encode the decodable sequence's own location. For example, each of the zipcodes may be a nucleic acid and there may be many copies in a spatially defined location (e g., feature) such as a square feature of any size from about 10 nm to about 1 cm, including for example, no larger than 0.1 pm, no larger than 0.2 pm, no larger than 0.5 pm, no larger than 1 pm, no larger than 2 pm, no larger than 5 pm, no larger than 10 pm, no larger than 20 pm, no larger than 30 pm, no larger than 40 pm, no larger than 50 pm, no larger than 100 pm, no largerthan 200 pm, no larger than 500 gm, no larger than 1 mm, no larger than 2 mm, and no larger than 5 mm. In some embodiments, an array may comprise an array of zipcodes that can be used to detect the distribution of ribonucleic acid (RNA), protein, deoxyribonucleic acid (DNA) or other molecules distribution in two or three dimensional space. These biomolecules can be detected in tissue, cell, organism or non-living systems. If a nucleic acid sequence is a zipcode, the complementary sequence of the nucleic acid sequence can also be a zipcode. In some embodiments, a zipcode and its complementary copy can encode the same position / location on a zipcode array (e.g., coordinates).

[0043] As used herein, the term “barcode probe” may comprise an “oligonucleotide” or a “nucleotide.” In some embodiments, a barcode probe may comprise one or more zipcode sequences. In some embodiments, a barcode probe may comprise a vertical “X-barcode.” In some embodiments, a zipcode may comprise a horizontal “Y-barcode.” In some embodiments, a zipcode comprises a X-barcode (e.g., a 15-16 nucleoside sequence corresponding to the X coordinate), a 'GGG' separator, and a Y-barcode (e.g., a 15-16 nucleoside sequence corresponding to the Y coordinate). In some embodiments, zipcodes can be designed for precision sequence performance, e.g., GC content between 40% and 60%, no homo-polymer runs longer than two, no self-complementary stretches longer than 3, and be comprised of sequences not present in a human genome reference. Zipcodes can be of sufficient length and comprise sequences that can be sufficiently different to allow the identification of each nucleic acid (e.g., oligonucleic acids) or peptides based on zipcode(s) with which each nucleic acid or peptides is associated.2. Chip Manufacture

[0044] Methods for fabricating high-resolution chips are described for example, in U.S. Patent Application No.: 16 / 959,578 and Costa, J. A. et al., ACS Appl. Mater. Interfaces 2019. In some embodiments, barcode probes with spatially-defined zipcode sequences are synthesized photolithographically using high-efficiency 5 ’-(2-nitronaphth-l-yl)benzyloxy carbonyl (“NNBOC”) phosphoramidite reagents with standard coupling, masking, alignment and exposureprotocols (Costa, J. A. et al., ACS Appl. Mater. Interfaces 2019; McGall, G. H. and Fidanza, J. A.Methods Protoc . 2001, 71-101; and Hoff, K. et al. Langmuir 2021, 37 (16), 4763-4771).3. Features

[0045] In one aspect, a substrate (e.g., a chip) may comprise a plurality of features. In some embodiments, each feature of the plurality of features is at a distinct location on the chip associated with two or more coordinates.

[0046] In some embodiments, a chip may comprise a square of approximately 15 x 15mm, approximately 10 x 10mm or approximately 5 x 5mm. In some embodiments, a chip may comprise a square of approximately 6.5 x 6.5mm.

[0047] In some embodiments, a chip may comprise one or more features, wherein each feature is a square of approximately 0.25 pm by 0.25 pm to approximately 0.75 pm by 0.75 pm in size, approximately 0.75 pm by 0.75 pm to approximately 1.25 pm by 1.25 pm in size, approximately 1.25 pm by 1.25 pm to approximately 1.75 pm by 1.75 pm in size, approximately1.75 pm by 1.75 pm to approximately 2.25 pm by 2.25 pm in size, approximately 2.25 pm by2.25 pm to approximately 2.75 pm by 2.75 pm in size, approximately 2.75 pm by 2.75 pm to approximately 3.25 pm by 3.25 pm in size, approximately 3.25 pm by 3.25 pm to approximately4.75 pm by 4.75 pm in size, or approximately 4.75 pm by 4.75 pm to approximately 5.25 pm by5.25 pm in size. In some embodiments, each feature is a square of approximately 2 pm x 2 pm.

[0048] In one aspect, the system described herein comprises between approximately 5 million and approximately 10 million features, between approximately 10 million and approximately 15 million features, between approximately 15 million and approximately 20 million features, between approximately 20 million and approximately 25 million features, between approximately 25 million and approximately 30 million features, between approximately 30 million and approximately 35 million features, between approximately 35 million and approximately 40 million features, between approximately 40 million and approximately 45 million features, and between approximately 45 million and approximately 50million features. As a non-limiting example, in some embodiments, a chip may comprise a square of approximately 10 x 10mm with 5,000 x 5,000 (25 million total) features.4. Barcode Probes

[0049] In some embodiments, a barcode probe comprises an oligonucleotide (e.g., with a predefined sequence) that is attached or immobilized to the surface of a substrate and is capable of binding a target molecule. In some embodiments, the substrate can be a bead, flat substrate, a chip, a flow cell or other suitable surface. In some embodiments, the substrate is a chip. In some embodiments, a barcode probe is a nucleic acid. In some embodiments, barcode probes can be of various lengths, such as from approximately 18 nucleotides to approximately 50 nucleotides, from approximately 50 nucleotides to approximately 100 nucleotides, or from approximately 100 nucleotides to approximately 200 nucleotides or more.

[0050] In one aspect, each feature of the system described herein comprises one or more barcode probes. In some embodiments, each barcode probe comprises a partial Illumina sequencing primer (e.g., for amplification). In some embodiments, each barcode probe comprises a Unique Molecular Identifier (UMI). In some embodiments, each barcode probe comprises a zipcode sequence. In some embodiments, each barcode probe comprises a poly-T sequence. In some embodiments, the poly-T sequence is a 30 nucleotide poly-T sequence followed by 3 ’-VMS’ (e.g., SEQ ID NO: 1). In some embodiments, each barcode probe is capable of capturing the polyA region of an mRNA molecule via hybridization with the poly-T sequence. In some embodiments, each barcode probe comprises a unique predefined sequence, comprising a partial Illumina sequencing primer (e.g., primer 1), a Unique Molecular Identifier (UMI), a zipcode, and a 30 nucleotide poly-T sequence followed by 3’-VN-5’. In some embodiments, a barcode probe is a nucleic acid as depicted in FIG. 1C.

[0051] In some embodiments, a UMI can have a length of at least, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. In some embodiments, a zipcode can have a length of at least, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27,28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. In some embodiments, a zipcode comprises both an X coordinate sequence (e.g., a X-zipcode) and a Y coordinate sequence (e.g., a Y-zipcode). In some embodiments, a zipcode comprises a 15-16 nucleoside sequence (e.g., corresponding of the X coordinate), a 'GGG' separator, and another 15-16 nucleoside sequence (e.g., corresponding to the Y coordinate).

[0052] In one aspect, each feature of the system described herein comprises approximately 100 to approximately 500 barcode probes, approximately 500 to approximately 1000 barcode probes, approximately 1000 to approximately 10,000 barcode probes, approximately 10,000 to approximately 50,000 barcode probes, approximately 50,000 to approximately 100,000 barcode probes, approximately 100,000 to approximately 250,000 barcode probes, approximately 250,000 to approximately 500,000 barcode probes, approximately 750,000 to approximately 1 million barcode probes, approximately 1 million to approximately 5 million barcode probes, approximately 5 million to approximately 10 million barcode probes, or approximately 10 million to approximately 20 million barcode probes.

[0053] In some embodiments, a system (e.g., comprising a chip) described herein may comprise at least 100, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 150,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 10,000,000, 20,000,000, 30,000,000, 40,000,000, 50,000,000, 60,000,000, 70,000,000, 80,000,000, 90,000,000, 100,000,000, 200,000,000, 300,000,000, 400,000,000, 500,000,000, 600,000,000, 700,000,000, 800,000,000, 900,000,000, 1,000,000,000, 2,000,000,000, 3,000,000,000, 4,000,000,000, 5,000,000,000 barcode probes.

[0054] In some embodiments, the barcode probe can be attached or immobilized onto the substrate at either the 5' end or the 3' end. In some embodiments, the barcode probe can be attached to the substrate at the 5' end. In some embodiments, the barcode probe is attached to the substrate at the 5' end, and the 3' end of the barcode probe can be extended by the incorporation of nucleotides as described herein.

[0055] Methods for manufacturing chips by laying probes are disclosed herein and known in the art. In some embodiments, in situ oligonucleotide synthesis is employed, in which the barcode probes are in known geographic locations in the X-Y coordinate plane. In one embodiment, the oligonucleotide barcode probe is synthesized on the surface. Examples of technologies that allow on-surface oligo synthesis include but are not limited to photolithography and inkjet. In another embodiment, the pre-synthesized oligonucleotide probes are spotted onto the surface. Various microarray protocols, for example, protocol for Agilent inkjet-deposited presynthesized oligo arrays are known to one skilled in the art.

[0056] Molecules that can be immobilized in the array include nucleic acids such as DNA or RNA and analogues and derivatives thereof, such as PNA. Nucleic acids can be obtained from any source, for example genomic DNA or cDNA or synthesized using known techniques such as step-wise synthesis.5. Feature-to-Feature Gap Space

[0057] In one aspect, the substrate (e.g., the chip) disclosed herein comprises a plurality of features. In some embodiments, each of the features may comprise one or more sites that is capable of attaching a subunit of a polymer (e.g., a barcode probe) onto the substrate. Each location may be adjacent to at least one, two, three, four, five, or six other locations.

[0058] In some aspect, adjacent barcode probes have an edit distance of 1. In some embodiments, pairs of DNA barcodes on the chip have a minimum edit distance between barcodes that are sufficiently far apart. In some embodiments, DNA barcodes that are 1, 2, 3, or 4 units apart have an edit distance equal to the number of units apart, whereas DNA barcodes that are at least 5 units apart have an edit distance of at least 5.

[0059] In one aspect, at least two of the plurality of features are adjacent with no interstitial space between adjacent features (i.e., no gap space). In some embodiments, the feature-to-feature gap space is zero. A person of skill in the art would appreciate that by minimizing gap space, the capture area is increased. Exemplary methods of manufacturing wafer-scale 3 ’-up functional DNA chips in hydrogels with sub-micrometer resolution are described in U.S. Patent ApplicationNo.: 16 / 959,578 and Costa, J. A., et al., ACS Appl. Mater. Interfaces 2019, the entire contents of which are hereby incorporated by reference. Exemplary methods of using photoresist technology (proprietary) that uses photoacid generated chemistry in a polymer matrix to spatially deblock DMT protecting groups are described in U.S. Patent Application No.: 16 / 982,349; Epstein, J. R., et al. Biosens. Bioelectron. 2003, 18 (5), 541-546; Costa, J. A. et al., 2019, the entire contents of which are hereby incorporated by reference.6. Targets

[0060] In one aspect, the present invention provides a method for sequencing a target nucleic acid molecule. By “target”, “target nucleic acid molecule”, “target molecule”, “target polynucleotide”, “target polynucleotide molecule” or grammatically equivalent thereof, herein is meant a nucleic acid of interest.

[0061] In one aspect, the systems and methods disclosed herein can be used to analyze a sample for the spatial distribution of at least one target molecule. In some embodiments, the systems and methods disclosed herein are for detecting the spatial distribution of a plurality of one or more target molecules within a sample. In some embodiments, a target nucleic acid is an RNA (e.g., an mRNA) derived from the genetic material of a particular organism.

[0062] In some embodiments, target nucleic acids include naturally occurring or genetically altered or synthetically prepared nucleic acids (such as mRNA from a mammalian disease model). Target nucleic acids can be obtained from virtually any source and can be prepared using methods known in the art and / or disclosed herein. For example, in some embodiments target nucleic acids can be directly isolated without amplification using methods known in the art to obtain target nucleic acids. In some embodiments the target nucleic acid may be from a fresh frozen or paraffin-embedded sample. In another example, target nucleic acids can also be isolated by amplification using methods known in the art, including without limitation polymerase chain reaction (PCR) technologies, whole genome amplification (WGA), multiple displacement amplification (MDA), rolling circle amplification (RCA), rolling circle replication(RCR) and other amplification methodologies. In some embodiments, the target nucleic acid molecule comprises ribonucleic acid (RNA).

[0063] In some embodiments, the polynucleotide target to be detected can be unmodified or modified. Exemplary modifications include, without limitation, radioactive and fluorescent labels as well as anchor ligands such as biotin or digoxigenin. The modification(s) can be placed internally or at either the 5' or 3' end of the targets.7. Methods of Forming a Pattern of Oligonucleotides

[0064] In some aspects, provided herein are methods of forming a pattern of oligonucleotides on a substrate. FIG. 6A depicts a flowchart illustrating an example of a process 600 for forming a pattern of oligonucleotides on a substrate having a plurality of functional groups. In some cases, the process 600 may be performed to form, on the substrate, multiple vertical regions. In some cases, each vertical region may be associated with an X barcode sequence, or a sequence of nucleotides, that uniquely identifies the location of each vertical region. In some cases, the process 600 may be performed to form vertical regions with different X barcode sequences. In some cases, the X barcode sequence of one vertical region and the X barcode sequence of another vertical region may have an edit distance (or a total number of different nucleotides) that is proportional to the distances between the two vertical regions. For example, the X barcode sequences of two adjacent vertical regions may have an edit distance of at least one (or differ by at least one nucleotide) while the X barcode sequences of two vertical regions separated by one vertical region may have an edit distance of at least two (or differ by at least two nucleotides). One advantage of the process 600 is that no interstitial spaces (or gaps) are present between adjacent vertical regions, thus maximizing coverage of the surface area of the substrate and the resolution of the spatial transcriptomics derived from signals captured from the oligonucleotide pattern on the substrate.

[0065] At 601, a vertical region with an X barcode sequence comprising a first nucleotide is formed by at least coupling the first nucleotide to a portion of the plurality of functional groups in a first vertical exposed region of the substrate. In some example embodiments, the firstnucleotide may be coupled to the portion of functional groups in the first vertical exposed region of the substrate by at least forming a first photoresist layer, which may include applying a photoresist composition onto an underlying layer of the substrate comprising the plurality of functional groups. The plurality of functional groups may be protected by protective groups. A dose of light may be exposed through a patterned mask onto the substrate before the protective groups on a section of the plurality of functional groups within the first vertical exposed region of the substrate are removed. In some cases, the functional groups within the first vertical exposed region of the substrate may be contacted with a nucleotide reagent in order to couple a fraction of the functional groups within the first vertical exposed region of the substrate with the first nucleotide. The resulting first vertical region may have an X barcode sequence that includes at least the first nucleotide. To further illustrate, in the example shown in FIG. 1 A, the first nucleotide may be adenine (denoted A in FIG. 1 A). A first patterned mask denoted “Mask 1” in FIG. 1 A may be used form the first vertical region (which includes two separate portions in FIG. 1A) in which the adenine (A) is coupled with the functional groups therein to form an X barcode sequence with adenine (A).

[0066] At 602, a second vertical region having an X barcode sequence comprising the first nucleotide and a second nucleotide is formed by at least coupling the second nucleotide to a portion of the plurality of functional groups in a second vertical exposed region of the substrate that overlaps with the first vertical region. In some example embodiments, the second nucleotide may be coupled to the portion of functional groups in the second vertical exposed region of the substrate by at least exposing a dose of light through a different patterned mask onto the substrate before the protective groups on a section of the plurality of functional groups within the second vertical exposed region of the substrate are removed. In some cases, the functional groups within the second vertical exposed region of the substrate may be contacted with a nucleotide reagent in order to couple a fraction of the functional groups within the second vertical exposed region of the substrate with the second nucleotide. In some cases, the second vertical exposed region may overlap, at least partially, with the second vertical exposed region. As such, thefunctional groups that are in a second vertical region that is the overlap between the first vertical exposed region and the second vertical exposed region may include the first nucleotide as well as the second nucleotide. To further illustrate, FIG. 1 A shows a second patterned mask denoted“Mask 2” being applied to form the second vertical exposed region. The second vertical exposed region may overlap at least partially with the first vertical regions formed in operation 601. The functional groups in the second vertical exposed region are coupled with cytosine (C) while the functional groups in the first vertical exposed region are coupled with adenine (A). Accordingly, the X barcode sequence (or zipcodes) in the second vertical region, which is overlap regions between the first exposed vertical region and the second exposed vertical region, includes adenine (A) and cytosine (C) (e.g., zipcode AC in FIG. 1A).

[0067] At 603, a third vertical region having an X barcode sequence comprising the first nucleotide, the second nucleotide, and a third nucleotide is formed by at last coupling the third nucleotide to a portion of the plurality of functional groups in a third vertical exposed region of the substrate that overlaps with the second vertical region. In some example embodiments, one or more additional vertical exposed regions may be formed on the substrate. In some cases, the one or more additional vertical exposed regions may be formed by patterned masks whose patterns are configured to enable the coupling of additional nucleotides to certain portions of the substrate. In some cases, the one or more additional vertical exposed regions may overlap with one or more existing vertical regions on the substrate, thereby adding additonal nucleotides to the X barcode sequences associated with those existing vertical regions. In FIG. 1A, for example, a third mask denoted “Mask 3” is applied in order to couple guanine (G) to a third exposed vertical region on the substrate. The X barcode sequence of a third vertical region in which the third exposed vertical region overlaps with the second vertical region, guanine (G) is added to the X barcode sequence of the second vertical region. In FIG. 1 A, for example, where the X barcode sequence of the second vertical region includes adenine (A) and cytosine (C), the X barcode sequence of the third vertical region may further include guanine (G) (e.g., zipcode ACG in FIG. 1A). It should be appreciated that operation 603 may be optional. Moreover, in some cases,operation 603 may be repeated any number of times in order to achieve a desired oligonucleotide pattern on the substrate.

[0068] At 604, a fourth vertical region having an X barcode sequence that comprises the first nucleotide and the second nucleotide but not the third nucleotide is formed from a nonoverlapping region between the third vertical exposed region and the second vertical region. In some example embodiments, non-overlapping regions between the additional vertical exposed regions (e.g., formed in operation 603) and the existing vertical regions on the substrate may retain the X barcode sequence of the existing vertical regions. For example, in FIG. 1 A, the third vertical exposed region formed by the third mask denoted “Mask 3” may not overlap with at least a portion of the second vertical region with the X barcode sequence adenine (A) and cytosine (C) (e.g., zipcode AC in FIG. 1A). This non-overlapping region may form a fourth vertical region that retains the X barcode sequence adenine (A) and cytosine (C) (e g., zipcode AC in FIG. 1 A). In other words, guanine (G) is not added to the X barcode sequence of the nonoverlapping region between the second vertical region and the third vertical exposed region, thus forming the fourth vertical region with the X barcode sequence adenine (A) and cytosine (C) (e.g., zipcode AC in FIG. 1A). It should be appreciated that operation 604 may be optional.Moreover, in some cases, operation 604 may be repeated any number of times in order to achieve a desired oligonucleotide pattern on the substrate.

[0069] In some embodiments, the pattern of oligonucleotides on the substrate may include X barcode sequences and Y barcode sequences. FIG. 6B depicts an example of a process 610 for forming a pattern of Y barcode sequences on the substrate. In some cases, the process 610 may be performed to form, on the substrate, multiple horizontal regions, each of which having a unique Y barcode sequence to identify the location of each horizontal region. In some cases, the process 610 may be performed in addition to the process 600 shown in FIG. 6A in order to form, on the substrate, features having unique combinations of X barcode sequences and Y barcode sequences. For example, in some cases, for horizontal Y barcode sequences, the same masks used for the vertical X barcode sequences may be rotated 90 degrees (e.g., clockwise or counter-clockwise) and the process 610 is performed to create an array of adjacent features, each of which having a unique combination of X- and Y barcode sequences. It should be appreciated that the process 610 is advantageous at least because the horizontal regions are formed with no interstitial space (or gaps) therebetween, thus maximizing coverage of the surface area of the substrate and the resolution of the spatial transcriptomics derived from signals captured from the oligonucleotide pattern on the substrate.

[0070] In some embodiments, the methods disclosed herein are those depicted in FIGS. 1A- 1C and / or FIGS. 6A-6B. As depicted in FIG. 6B, in some embodiments, the method of forming a pattern of oligonucleotides on a substrate having a plurality of functional groups, comprises: (605) forming a first horizontal region with a Y barcode sequence comprising a third nucleotide by at least coupling the third nucleotide to a portion of the plurality of functional groups in a first horizontal exposed region of the substrate, and (606) forming a second horizontal region having a Y barcode sequence comprising the third nucleotide and a fourth nucleotide by at least coupling the fourth nucleotide to a portion of the plurality of functional groups in a second horizontal exposed region of the substrate that overlaps with the first horizontal region.

[0071] In some embodiments, the method depicted in FIG. 6B may optionally further comprise (607) forming a third horizontal region having a Y barcode sequence comprising the third nucleotide, the fourth nucleotide, and a fifth nucleotide by at last coupling the fifth nucleotide to a portion of the plurality of functional groups in a third horizontal exposed region of the substrate that overlaps with the second horizontal region.

[0072] In some embodiments, the method depicted in FIG. 6B may optionally comprise (607) and may further comprise (608) forming, from a non-overlapping region between the third horizontal exposed region and the second horizontal region, a fourth horizontal region having a Y barcode sequence that comprises the third nucleotide and the fourth nucleotide but not the fifth nucleotide.

[0073] In some embodiments, the first, second, and optional additional nucleotides are coupled to generate a Y barcode sequence. In some embodiments, a location within each horizontal region is identified by a corresponding Y barcode sequence.

[0074] In some embodiments, a set of masks may be provided. Each mask of the set may be used for defining a different subset of regions (e.g., horizontal regions or vertical regions) on the substrate. Each mask may comprise a plurality of openings, which define a pattern of active regions and inactive regions on the substrate. During polymer synthesis, subunits can be added onto the polymers within the active regions.

[0075] In some embodiments, the mask openings may take various shapes, regular or irregular (e g., square, rectangular, triangular, diamond, hexagonal, and circle). In some embodiments, each mask may have its own design of openings, which defines a distinct pattern of active and inactive regions on the substrate. In some embodiments, the openings may or may not be aligned in a single direction (e.g., horizontal or vertical). In some embodiments, each opening may cover an integer number of regions on the substrate (e g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more). For each mask, the openings may or may not be of the same shape. For each region on the substrate, the set of masks collectively may define a unique string of synthetic steps or embedding (i.e., a sequence of nucleotide subunits to be introduced onto the substrate) used to form the polymers in that location. In some embodiments, each mask may be used for at least one synthetic step for forming the polymers.

[0076] In some aspects, any individual mask contains polygons (i.e., exposure areas) much larger than the minimum feature size. In some embodiments, overlay of masks creates the minimum feature size, and, as such, the limiting factor is the overlay accuracy of the mask alignment equipment used (e.g., ±0.3 pm at 3-sigma). In some embodiments, neighboring zipcodes only differ by one base (gray-codes), but their spatial extent covers the full chip area eliminating any unused border regions. In some embodiments, at least approximately 50%, at least approximately 60%, at least approximately 70%, at least approximately 80%, at least approximately 90%, or at least approximately 100% of the surface area of the chip is covered byat least one probe. In some embodiments, at least approximately 90% of the surface area of the chip is covered by at least one probe. In some embodiments, at least approximately 95% of the surface area of the chip is covered by at least one probe. In some embodiments, at least approximately 100% of the surface area of the chip is covered by at least one probe.

[0077] In some embodiments, the set of masks are designed such that each pair of strings of synthetic steps (or embeddings) used to form the polymers at two adjacent locations differ from each other by a maximum number of synthetic steps. In some embodiments, barcode that are 1, 2, 3, or 4 units apart have an edit distance equal to the number of units apart, whereas DNA barcodes that are at least 5 units apart have an edit distance of at least 5. This property holds true for both sections of the DNA barcodes (e.g., X and Y). In some embodiments, if the X and Y coordinates of a pair of barcodes each differ by at least 5, the edit distance of the first sections will be at least 5, and the edit distance of the second sections will be at least 5. Methods

[0078] In some cases, two strings of synthetic steps used to form polymers at two adjacent locations differ from each other by one and only one synthetic step. For example, each pair of embeddings used to synthesize neighboring polymers in two adjacent locations differs by one and only one exposure / non-exposure step.

[0079] For each mask, a certain percentage (e.g., 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% or more) or all of the openings may have the same length and / or width. In some cases, the length of the openings may be the same as the substrate. In some cases, the length of the openings may be less than that of the substrate such that one mask is only capable of masking a portion of the substrate. In cases where all of the openings have the same length, their widths may vary and one or more of the openings may or may not have the same width. For example, the width of the openings may be greater than or equal to about 1 nm, 10 nm, 50 nm, 100 nm, 250 nm, 500 nm, 750 nm, 1 pm, 2 pm, 3 pm, 4 pm, 5 pm, 6 pm, 7 pm, 8 pm, 9 pm, 10 pm, 20 pm, 40 pm, 60 pm, 80 pm, 100 pm, 200 pm, 300 pm, 400 pm, 500 pm, 600 pm, 700 pm, 800 pm, 900 pm, 1,000 pm, or more. In some cases, the width of the openings may be smaller than or equal to about 50 mm, 10 mm, 1,000m, 900 pm, 800 pm, 700 pm, 600 pm, 500 pm, 400 pm, 300 pm, 200 pm, 100 pm, 90 pm, 80 pm, 70 pm, 60 pm, 50 pm, 40 pm, 30 pm, 20 pm, 10 pm, 8 pm, 6 pm, 4 pm, 2 pm, 1 pm, or less. In some cases, the width of the openings may be between any of the two values described herein, for example, 12 pm.

[0080] The length of the openings may vary. In some cases, each of the openings has a length of greater than or equal to about 1 pm, 10 pm, 25 pm, 50 pm, 75 pm, 100 pm, 200 pm, 400 pm, 600 pm, 800 pm, 1,000 pm, 2,000 pm, 3,000 pm, 3,500 pm, 4,000 pm, 4,500 pm, 5,000 pm, 5,500 pm, 6,000 pm, 7,000 pm, 8,000 pm, 9,000 pm, 10,000 pm, or more. In some cases, the length of the opening may be smaller than or equal to about 50,000 pm, 25,000 pm, 10,000 pm, 8,000 pm, 7,000 pm, 6,500 pm, 6,000 pm, 5,500 pm, 5,000 pm, 4,500 pm, 4,000 pm, 3,000 pm, 2,000 pm, 1,000 pm, 800 pm, 600 pm, 400 pm, 200 pm, 100 pm or less. In some cases, the length of the openings may be between any of the two values described herein, for example, 4,900 pm.

[0081] The mask can be formed of various materials known in the art, such as glass, silicon- based (e.g., silica nitrides, silica), polymeric, semiconductor, or metallic materials. In some cases, the mask comprises lithographic masks (or photomasks). Thickness of the mask may vary. In some cases, the mask may have a thickness of greater than or equal to 1 pm, 10 pm, 50 pm, 100 pm, 250 pm, 500 pm, 750 pm, 1 millimeter (mm), 2 mm, 3 mm, 4 mm, 5 mm, 6 mm, 7 mm, 8 mm, 9 mm, 10 mm, 15 mm, 20 mm, 25 mm, 30 mm, 35 mm, 40 mm, 45 mm, 50 mm, or more. In some cases, the mask may have a thickness of less than or equal to about 500 mm, 250 mm, 100 mm, 50 mm, 40 mm, 30 mm, 20 mm, 10 mm, 8 mm, 6 mm, 4 mm, 2 mm, 1 mm, 900 pm, 800 pm, 700 pm, 600 pm, 500 pm, 400 pm, 300 pm, 200 pm, 100 pm, or less. In some cases, thickness of the mask may be between any of the two values described herein, e.g., about 7.5 mm.

[0082] In one aspect, methods are provided to detect the distribution of a biomolecule in a two dimensional space. In some embodiments, the biomolecule may be made to react with the barcode probes. The barcode probes that comprise zipcodes that have reacted with thebiomolecule may then be sequenced or otherwise detected. Tn some embodiments, because the zipcodes encode their own locations, by detecting zipcodes, the biomolecule's spatial distribution can then be determined accordingly.

[0083] Transcriptomics

[0084] Those of skill in the art will appreciate that, in some embodiments, the systems and methods disclosed herein can be used in spatial transcriptomics to provide information of biological, medical, and / or therapeutic significance. In some embodiments, the systems and methods disclosed herein can enable the identification or quantitation of one or more targets. In some embodiments, the targets may be biomarkers for diagnostic, prognostic, and / or efficacy determination related to a disease or disorder or treatment thereof. In some embodiments, the biomarkers may provide information for the identification of drug target candidates, monitoring disease progression or treatment, and / or identification of patient populations or subjects for clinical trials.

[0085] In some embodiments, the target nucleic acids may be RNAs (e.g., mRNAs), and the barcode probes may capture RNAs having specific nucleotide sequences by hybridization. In some embodiments, for example, the barcode probes may comprise a poly-T capture sequence for mRNA that can hybridize a polyA tail of a mRNA. In some embodiments, reverse transcription of the captured mRNA can be initiated using added primers, and cDNA can be produced using the barcode probe as a template. In some embodiments, following hybridization, RNA sequences are copied onto the chip using a reverse transcriptase and a template-switching oligonucleotide. Methods related to reverse transcriptase and a template-switching oligonucleotide are described, for example, in Zhu, Y. Y. et al. Biotechniques 2001, 30 (4), 892- 897; and Hudson, W. H. and Sudmeier, L. J. STARProtoc. 2022, 3 (2), 101391.

[0086] In some embodiments, the original RNA strand is removed and a second strand is synthesized, stripped from the chip, and used in downstream library preparation and sequencing.

[0087] In some embodiments, the resultant cDNA that is synthesized incorporates the sequences from the barcode probe (e.g., zipcode sequences). In some embodiments, the cDNAsmay be amplified and / or purified using methods known in the art. In some embodiments, the cDNAs may be further processed using methods known in the art (e.g., fragmentation, end repair, A-tailing, and / or ligation). In some embodiments, adaptors are added to the cDNAs.

[0088] In some embodiments, a library of the cDNAs / amplified cDNAs is prepared and nucleotide sequences of the libraries are obtained. In some embodiments, nucleotide sequences of the barcode probes (e.g., barcode sequences) can provide data to map a target oligonucleotide (e.g., an mRNA transcript) to its location on the support. In some embodiments, the location information (e.g., X barcode and Y barcode coordinates) can be compared to other data (e.g., an image of tissues and / or cells overlaid onto the support (e.g., chip)) and mRNA transcripts may thereby be mapped to the location in the overlaid tissue. In some embodiments, the image is a Hematoxylin and Eosin (H&E) staining.

[0089] In some embodiments, any methods known in the art and / or described herein can be used to analyze mRNA transcript in the context of the overlaid tissues. In some embodiments, computation analysis (e.g., data binning, clustering) methods can be employed in connection with the systems and methods disclosed herein. In some embodiments, any methods known in the art and / or described herein can be used for spatial transcriptomics data analysis.EXAMPLESExample 1: High-Resolution Spatial DNA chips Manufacture

[0090] High-resolution spatial DNA chips were manufactured as previously described (Costa, J. A. et al., 2019). Briefly, semiconductor photolithographic manufacturing techniques were used in conjunction with highly efficient light-directed DNA synthesis to prepare high- density DNA arrays on silicon wafer substrates. Positional information is encoded in the DNA sequences on the silicon chips photolithographically (Costa, J. A. et al., 2019). High-resolution chips are fabricated by sequential exposures of lithography masks, each one directing a single base addition. Oligonucleotides are subsequently transferred from the donor silicon substrate to a thin polyacrylamide gel covalently cast on a silanized glass slide (Costa, J. A. et al., 2019).

[0091] Chips are manufactured as 6.5 x 6.5mm or 10 x 10mm with 5,000 x 5,000 (25 million total) features on the 10 x 10mm chips. Each wafer produces 111 10 x 10mm chips. Each feature is approximately square, with a size of 2pm x 2pm. Features are juxtaposed adjacently without interstitial spaces (FIGS. 1A-1F). Embedded within each feature is a unique predefined sequence, comprising of the partial Illumina sequencing primer 1 for amplification, a Unique Molecular Identifier (UMI), a zipcode, and a 30 nucleotide poly-T sequence followed by 3’-VN- 5’, which is optimized to capture the polyA region of an mRNA molecule just before the coding region (FIGS. 1A-1F). The zipcode contains X and Y coordinates that are decoded during bioinformatics analysis following sequencing using custom PostMaster™ software (Centrillion Tech, Palo Alto, CA). For the work presented in this analysis, 6.5 x 6.5 mm synthesized chips were used.Example 2: DNA Zipcode Synthesis

[0092] Zipcodes are designed to reduce errors during DNA zipcode synthesis. A long-range minimum edit distance, which dictates for some given D, if 2 zipcodes are > D units apart on the chip, their edit distance must be D. For this particular mask set, D = 5. FIGS. 1A-1F illustrate the method of producing vertical X-zipcodes on the array. For horizontal Y-zipcodes, the masks are rotated 90 degrees counter-clockwise and the process is repeated again, creating a quilt of unique X-Y zipcodes. Each zipcode border is delineated once and only once, alleviating the need for exacting overlay. In effect, neighboring zipcodes only differ by one base (gray-codes), but their spatial extent covers the full chip area eliminating any unused border regions. Such a scheme allows for 100% surface coverage. In terms of resolution, high-resolution chips used in this work have individual features of unique zipcodes that are 2 x 2 pm in size, however we have previously demonstrated this technology at 1 m resolution (Costa, J. A. et al., 2019) and have been able to fabricate 0.66 x 0.66 pm features with additional layers of light directed synthesis (data not shown). Notably, any individual mask contains polygons (exposure areas) much larger than the minimum feature size. Their overlay creates the minimum feature size, and, as such, thelimiting factor is the overlay accuracy of the mask alignment equipment used (±0.3 pm at 3- sigma).Example 3: High-Resolution DNA Chip Workflow

[0093] High-resolution spatial DNA chips are compatible with existing workflows for fresh frozen or paraffin-embedded samples. The data presented here uses a fresh frozen sample from mouse liver. The assay workflow is outlined in FIGS. 1A-1F and FIGS. 6A-6B. Briefly, tissues are fixed, permeabilized, and decrosslinked, and RNA from the tissue is hybridized to the chip. Following hybridization, RNA sequences are copied onto the chip using a reverse transcriptase and a template-switching oligonucleotide. The original RNA strand is removed and a second strand is synthesized, stripped from the chip, and used in downstream library preparation, which is then sequenced using an Illumina system.Example 4: Spatial Transciptome Library Sequencing

[0094] The spatial transcriptome library was sequenced in four separate batches using the Illumina Novaseq 6000 system. This yielded a total of 3,890,659,083 reads from counts of 354,336,561, 501,709,129, 1,437,827,348, and 1,596,786,045 in each batch, respectively. Of this grand total, 3,478,542,340 reads contained a "zipcode" that was decoded using PostMaster software (Centrillion Tech, Inc., Palo Alto, CA). Briefly, the PostMaster decoder compares putative zipcodes extracted from the reads and designed zipcodes and allows the desired edit distance not larger than 3], In the spatial transcriptome data analysis, among the 3,563,068,585 read sequences identified with a zipcode structure, a total of 3,467,777,970 were successfully decoded, resulting in a decoding rate of 97.33%. Within this decoded subset, the proportion of sequences with an edit distance of 3 was 1.18% for X zipcodes and 1.29% for Y zipcodes. Upon decoding, 2,399,941,684 of these reads were successfully mapped to a unique location of the mouse genome GRCm39 using the STAR, with 2,100,341,576 of these aligning to a specific gene by featureCounts. Subsequent to this mapping, deduplication was undertaken. Reads that shared the same Ensembl gene ID, X position, Y position, and UMI were classified as duplicates; only the first occurrence was retained. Post-deduplication, a total of 643,142,336 unique readsremained. This implies a sequencing saturation rate of 69.37%, suggesting the potential for obtaining a higher number of unique reads with additional sequencing. Capture efficiency (Tables 1 and 2) is comparable to previously-reported scRNA-seq methods (Ziegenhain, C., et al. Mol. Cell ' 2017, 65 (4), 631 -643. e4).

[0095] Despite the advancements associated with the high-resolution chip, challenges persist. Delineating precise cell boundaries and achieving accurate single-cell genetic mapping remain areas of active research (Fang, S. et al. Genomics Proteomics Bioinformatics. 2023, 21 (1), 24- 47). Potential solutions may involve the deployment of advanced algorithms, possibly integrating machine learning techniques, to enhance spatial transcriptomics data interpretation (Kleshchevnikov, V. et al. Nat. Biotechnol. 2022, 40 (5), 661-671; Littman, R. et al. Mol. Syst. Biol. 2021, 17 (6), el0108; Chen, A. Cell. 2022 May 12;185(10):1777-1792.e21). Additionally, utilizing markers from specific subcellular compartments, such as pre-mRNA, and merging high- resolution microscopy might further refine cellular boundary determinations in tissues with intricate architectures(Kleshchevnikov, V. et al. Nat. Biotechnol. 2022, 40 (5), 661-671).Example 5: Spatial Expression of Cell-Specific Markers

[0096] The liver performs a vast array of metabolic functions that are regulated in spatial, temporal, and injury-related manners as determined by the microenvironment of each hepatocyte in the liver acinus. Zonation in the liver has been studied for over 80 years. Liver lobules are hexagonally shaped with portal arteries and veins supplying blood at each of the 6 points of the hexagon and a central vein, through which blood is drained, at the center of the lobule. Each of the approximately 12 concentric layers of hepatocytes has a microenvironment slightly different from its neighbors’, defined by gradients of oxygen, glucose, pancreatic hormones such as insulin, morphogens, etc.

[0097] This structure enables production line patterns of cells, as with the synthesis of bile salts. The first two enzymes in this cascade, CYP7A1 and HSD3B7 are expressed in the most pericentral cell layers where cholesterol, the starting material in the pathway, is most abundant. The next-two enzymes in the cascade, CYP8B1 and CYP27A1, are expressed in the next layer,just periportal to those cells in layer 1, indicating different steps of the enzymatic cascade occur in a spatially-regulated fashion. Researchers have further delineated the intricate zonation of opposing pathways in the liver including hepatic glucose metabolism (glycolysis and gluconeogenesis), glutamine synthesis and ureagenesis, lipogenesis and P-oxidation, protein secretion, and defense against pathogens and xenobiotics. These distinct markers and intricate zonal distributions of cells emphasize the specialized roles of hepatocytes across these separate hepatic zones.

[0098] A large number of markers have been reported for cells at various hepatocyte layers as well as for other liver cells such as endothelial cells and Kupffer cells. The dynamic expression patterns of several canonical marker genes can be observed in FIG. 2 GLUL is expressed by pericentral hepatocytes. It encodes glutamine synthetase, which converts glutamate into glutamine. OAT encodes organic anion transporters, mainly involved in the uptake of bile acids and exogenous drugs including cancer medication. It is also expressed by pericentral hepatocytes. HAMP is expressed by midlobule hepatocytes. It encodes hepcidin, which blocks the iron export protein, ferroportin, thus controlling iron storage in the liver and systemic iron availability. CYP2F2 and CPS1 are periportal hepatocyte markers. There are 450 human cytochrome P450 proteins, which are involved in xenobiotic metabolism of which CYP2F2 encodes one. CPS1 is involved in ammonia metabolysis into urea. JCHAIN, as previously discussed, is expressed by immune cells and liver epithelial cells. Expression analysis demonstrates the ability to capture and identify genes that are representative of distinct regions of the mouse liver.

[0099] From FIGS. 3A-3C, it can be observed that the gene expression pattern closely follows the shape of the tissue section. This would indicate the capture of mRNA was specifically restricted to the area where the tissue section was present. Unsupervised clustering identified seven distinct cell types (FIGS. 3A-3C): periportal hepatocytes, pericentral hepatocytes, three distinct midlobular hepatocyte cell types, and two clusters of cells expressing unique markers, the first expressing JCHAIN and the second expressing RSAD2 and CMPK2.While midlobular clusters 2 and 3 have markers solely expressed in these clusters, genes expressed in midlobular cluster 1 are not exclusive to midlobular cluster 1 (FIGS. 3A-3C).

[0100] JCHAIN encodes a small polypeptide (joining chain) that governs multimerization of secretory immunoglobulins M and A. While it is an immune cell marker, it is also expressed by liver epithelial cells. The JCHAIN+ cluster also shows strong expression of CD74, the cognate receptor for macrophage migration inhibitory factor (MIF). CD74 is expressed by hepatocytes and hepatic stellate cells. CD74 plays a role in antigen presentation in the liver in models of fibrosis following toxic liver injury and both MIF and CD74 are overexpressed in liver during hepatocarinogenesis. In the liver, CD74 is expressed on hepatocytes as well as hepatic stellate cells (HSC). Specifically, the MIF / CD74 axis was identified to exert hepatoprotective effects in a model of NASH, as well as liver fibrosis induced by chronic toxic liver injury.

[0101] RSAD2 (also known as Viperin) and CMPK2 are interferon (IFN) inducible genes. IFNs are produced in response to viral infection. However, viperin has been demonstrated to be highly expressed in liver in the absence of IFN-stimulation, suggesting a need for further research to better understand the role of these IFN-inducible genes within the context of liver. Further investigation is required to determine whether the tissue can be used to generate additional meaningful clusters. For reference, a bin size of 5, corresponding to 10 x 10 pm regions, was used to generate the images in FIGS. 2 and 3A-3C. Hepatocytes are approximately 80000 pm3or 20 x 20 x 20 pm in diameter in adult mice.

[0102] Table 1 depicts the UMI captured at various bin sizes. Table 2 depicts the number of gene captured at various bin sizes.Table 1. UMI captured at various bin sizes.Table 2: Number of genes captured at various bin sizes.Example 6: Generation of Cell Type Spatial and Cluster Maps

[0103] In order to better visualize the spatial composition of the tissue, Cell21ocation29 and U-CIE were used to generate cell type spatial and cluster maps. Cell21ocation is a spatial deconvolution method used here to interpret gene expression data within liver tissue sections, as referenced from the Liver Cell Atlas. By applying Cell21ocation, we quantify the cell type composition at each tissue location, integrating this information to form a detailed cellular map. The analysis translates high-dimensional transcriptomic data into a resolved spatial context, revealing the distribution and abundance of cell types across the liver section. The result is a color-coded representation that provides insights into the tissue's cellular architecture (FIG. 4A), with the intensity of each color corresponding to the presence of specific cell types. This approach highlights the spatial heterogeneity of the liver and can be used to identify regions with distinct cellular functions or compositions, offering a window into the intricate structure and functionality of liver tissue at the cellular level.

[0104] U-CIE (Koutrouli, M. et al. Protein Sci. 2022, 31 (9), e4388) is utilized for visualizing gene expression with UMAP (Mclnnes, L. et al. Journal of Open Source Software, 3(29), 861) condensing high-dimensional data into three dimensions (Becht, E. et al. Nat.Biotechnol. 2019, 37 (1), 38-44). This maintains the spatial integrity of gene expression patterns. Post-reduction, U-CIE translates the spatial data into the CIELAB color space, highlighting gene expression variations. The visualization turns gene expression at each liver location into a color (FIG. 4B), demonstrating the heterogeneity of the liver's cellular landscape. Further, single-cell data from the Liver Cell (Guilliams, M. et al. Cell. 2022, 185 (2), 379-396. e38) is projected inthe three UMAP dimensions of the spatial data, with U-CIE coloring applied to depict cell types (FIGS. 4C and 4D). This combination reveals the distribution of cell types across the liver, providing a color-coded map of cellular organization.Example 7: Materials and Methods

[0105] Chip Fabrication and Decoding. The high-resolution chips, fabricated by Centrillion Technologies, are synthesized on 150mm silicon wafers using aligner model CENAL1000 (Centrillion Taiwan), generating 111 10mm x 10mm chips per wafer. Chips are singulated and subsequently inverted and transferred to a thin polyacrylamide gel (10-20 pm) cast on a silanized glass slide. This process has been previously described (Costa, J. A. ACS Appl. Mater. Interfaces 2019). Each of the 10mm x 10mm chips contains 5,000 x 5,000 features that are 2pm x 2pm each and juxtaposed adjacently without interstitial spaces. Embedded within each feature is a predefined sequence, comprising of a fragment of the Illumina read primer 1 for amplification, a Unique Molecular Identifier (UMI), a zipcode, and a 3' sequence, 3’-TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTVN 5, (SEQID N0:1),WHICH IS OPTIMIZED TOcapture the polyadenylated region of an mRNA molecule closest to the coding region. The zipcode is constructed from a 15-16 nucleoside sequence indicative of the X coordinate, a 'GGG' separator, and another 15-16 nucleoside sequence corresponding to the Y coordinate. This zipcode can be effectively decoded using the PostMaster software. The sequences common to all barcode probes (poly-T, UMI & primer), which required no spatial definition, were synthesized using standard 5’-DMT deoxynucleoside synthesis reagents & coupling protocols (Costa, J.A. ACS Appl. Mater. Interfaces 2019; McGall, G. H. et al. DNA Arrays Methods Protoc . 2001, 71- 101). The spatially-defined zipcode sequences were synthesized photolithographically using high-efficiency 5’-(2-nitronaphth-l-yl)benzyloxycarbonyl (“NNBOC”) phosphoramidite reagents with standard coupling, masking, alignment and exposure protocols. Vacuum contact lithography was performed using chrome on quartz masks. Two-micron features achieved by employing a modified masking protocol as described in FIGS. 1 A-1F.

[0106] The barcodes were designed such that the “embeddings” of adjacent barcodes have an edit distance of 1. The term "embedding" refers to the binary sequence that specifies when the barcode's position on the chip is exposed to light during the chip synthesis via photolithography (Raman Kumar, M. et al. Eds.; Advances in Intelligent Systems and Computing; Springer: Singapore, 2020; pp 645-654). The embeddings of all the chip barcodes form a two-dimensional gray code. Another important property of pairs of DNA barcodes on the chip is the minimum edit distance between barcodes that are sufficiently far apart. DNA barcodes that are 1, 2, 3, or 4 units apart have an edit distance equal to the number of units apart, whereas DNA barcodes that are at least 5 units apart have an edit distance of at least 5. This property holds true for both sections of the DNA barcodes, meaning that if the X and Y coordinates of a pair of barcodes each differ by at least 5, the edit distance of the first sections will be at least 5, and the edit distance of the second sections will be at least 5. The combination of using gray codes and maintaining a minimum edit distance for distant DNA barcodes enables the chip to be fully covered by DNA barcodes, ensuring efficient transcript capture, and allowing for error correction during both chip synthesis and subsequent decoding.

[0107] Bioinformatics. The spatial sequencing data processing pipeline was constructed utilizing proprietary software, PostMaster, developed by Centrillion Technologies, Palo Alto. PostMaster parses Unique Molecular Identifiers (UMIs), and X, Y coordinates from the raw fastq files. It then decodes these coordinates, comparing them to the theoretical sequence of the decoded zipcodes. This decoding information, including edit distances and UMI, is stored within the sequence identifiers of a new R2.fastq.gz file. Only coordinates with an edit distance less than or equal to three (compared to the theoretical sequence) are retained for further analysis. Subsequently, the trimmed reads are mapped to the reference genome using the STAR aligner. These mapped reads are then processed by FeatureCounts to assign them to genomic features, thus generating a count matrix.

[0108] Our custom script 'bamreader.sh' parses the decoding information stored in the sequence identifiers, creating a CSV file with gene ID, UMI, and spatial coordinates.

[0109] Finally, an in-house Python script converts this data into an h5ad fde, a format suitable for spatial transcriptomic analysis. This script also removes duplicates, identified by UMI, X-Pos, Y-Pos, and Gene-id, ensuring the output consists of unique spatial gene expression data. The script constructs a sparse matrix, effectively integrating the gene expression and spatial data.

[0110] The bioinformatics pipeline. FIG. 5 depicts the bioinformatics pipeline. This study implemented a custom bioinformatics pipeline for the processing of spatial transcriptomics data obtained from paired-end FASTQ files. Our method primarily consisted of five key steps: 1. Sequence Parsing and Decoding with PostMaster: Our initial preprocessing step involved the utilization of PostMaster software. This tool was specifically designed to parse sequencing data from read 1 sequence, extract unique molecular identifiers (UMIs), and decode x and y zipcodes. PostMaster then provides the x and y coordinates as well as the edit distance of each zipcode relative to the theoretical sequences. The output of this step is a new R2.fastq.gz file with decoding information stored in the sequence ID (SeqID). 2. Sequence Trimming with seqtk: Following decoding, we applied the seqtk tool to trim the sequences in the newly formed R2.fastq.gz file. This process is vital in the removal of sequencing adapters and other potential sequence contaminants. 3. Read Alignment with STAR Aligner: Subsequent to trimming, we aligned the cleaned sequences to a reference genome using the Spliced Transcripts Alignment to a Reference (STAR) software. The alignment parameters were set to ensure only unique matches to the reference genome were retained, thereby eliminating multi-mapping reads. This step resulted in a Binary Alignment Map (BAM) file. 4. Read Counting with featureCounts: After alignment, we employed the featureCounts tool to perform a gene-level quantification.Specifically, this step involved counting the number of reads that were successfully mapped to each gene in the reference genome. 5. Data Parsing with bamreader.sh: After featureCounts, we utilized a custom script, bamreader.sh, to parse the decoding information embedded within the sequence IDs in the BAM file. The output of this step was a CSV file containing detailed information such as gene id, x and y coordinates, edit distance, and UMI, which can then beused for subsequent downstream analyses. 6. Converting CSV fde to Anndata: The final step of the data processing pipeline converts the CSV output into a h5ad file, which is a binary format used to efficiently store large, multi-dimensional arrays. This is accomplished through a custom Python script using the Scanpy library, named h5ad.py.

[0111] In this step, any duplicate gene expression data, based on Unique Molecular Identifiers (UMI), X-Pos, Y-Pos, and Gene-id, are removed to ensure only unique spatial gene expression information is retained. The script then transforms this data into a sparse matrix representation, which helps to efficiently store and process the large amount of zero entries typical of gene expression data. Each unique gene and its corresponding spatial coordinates are recorded in this matrix, which is organized by gene symbols. Finally, an AnnData object is created with the sparse matrix, enabling the integration of gene expression and spatial data. This h5ad file provides a convenient and efficient format for subsequent data analysis tasks.

[0112] Spatial Library Construction. The spatial transcriptomics procedure is initiated with tissue fixation, permeabilization using pepsin, and decrosslinking. mRNA was captured through hybridization then copied using reverse transcription with a template-switching oligo. Templating RNA molecules were removed and the cDNA was copied using second-strand synthesis. The second strand was denatured from the chip and used in downstream library preparation. The library was subsequently purified and size-selected.

[0113] Pre-processing of Tissue

[0114] Tissue slides, retrieved from a -80°C freezer, were briefly warmed at 37°C for 1 minute. Slides were fixed using lOOpl of 4% formaldehyde for 30 minutes at room temperature. Subsequently, decrosslinking was performed at 70°C over 20 minutes in TE Buffer (lOmM Tris- HCl;lmM EDTA). The slides were then swiftly rinsed in a I xPBS-filled staining jar. For permeabilization, after removing excess water, the slides were treated with lOOpL of 1 / 1000 dilution of 0.1% Pepsin and incubated at 37°C for 5 minutes.

[0115] Hybridization

[0116] For hybridization, a 20pL solution was formulated on ice with 9.4pL nuclease-free water, 5 pL 20* SSC, 2.4 pL 25mM MgC12, 0.2pL 0.1M DTT, 2pL lOmM dNTPs, and IpL RNase Inhibitor A (NEB, M0314L). A new Sequoia 2 chip was retrieved from the original container. After drying the surrounding of the chip's gel area, any residue liquid was quickly aspirated. Both the chip gel and the designated tissue area on the slide received lOpL of Hybridization Buffer. The slide was subsequently inverted onto the chip, ensuring a snug, bubble-free contact. This configuration was incubated at 42°C for 4 hours in a humidity chamber.

[0117] Reverse Transcription

[0118] For reverse transcription on the chip, a 40 pL RT MIX was prepared on ice, containing 4pL PEG8000, 0.4pL Triton X, 1.5pL TSO Oligo with sequence 5‘- AAGCAGTGGTATCAACGCAGAGTACATrGrGrG (SEQ ID NO:2), 2pL dNTPs, 1.6pL Tris HC1 (pH 8.3), 1.2pL NaCl, 4pL MgC12, 0.53pL GTP, 3.2pL DTT, IpL RNase inhibitor (NEB, M0314L), 0.8pL Maxima H RT Enzyme (Thermo Fisher, EP0751), and 19.77pL nuclease free Water. After hybridization, the chip was rinsed in a staining dish with 0.1 x SSC until the tissue slide released. A frame-seal (Biorad, SLF0201) was then aligned with the dashed region of the chip gel. The holes of the frame were filled with the RT Mix. The prepared chip underwent incubation in a pre-heated preservation box at 42°C for 90 minutes and subsequently at 53°C for 30 minutes.

[0119] Second Strand Synthesis

[0120] For second strand synthesis, a 50pL reaction mix was prepared on ice, using 24pL nuclease-free water, 5pL Thermopol Buffer, 4pL Bst Pol Full Length (NEB, M0328S), 7pL lOmM dNTPs, and lOpL of a primer with the sequence AAGCAGTGGTATCAACGCAGAG (SEQ ID NO:3). A 20pL Wash Buffer SSB: 13.2pL nuclease-free water, IpL Thermol Buffer, 2.8pL lOmM dNTPs, and 4pL primer.

[0121] The chip was sequentially cleaned: thrice using 75pL nuclease-free water, incubated for 2 minutes; twice with 75pL 0.08M KOH for 5 minutes; and twice with 75pL Tris (pH=7.5) for 2 minutes.

[0122] For synthesis, chips were washed in 20pL Wash Buffer SSB and 50pL Second Strand Synthesis Mix was applied. The chip underwent a 30-minute incubation in a 65°C humidity chamber. The chip was then rinsed twice with lOOpL Buffer EB and treated with 20pL 0.08M KOH for 10 minutes. After mixing, 20pL was moved to a 0.2ml PCR tube. This KOH process was repeated, combined in the tube, adjusted with 5pL of IM Tris (pH 7.0), and stored on ice. The chip was finally preserved in 4xSSC at 4°C.

[0123] cDNA Amplification and Purification

[0124] For cDNA pre-amplification, a mixture consisting of 50pL KAPA HiFi Hot Start Ready Mix (Roche, KK2601), 35pL cDNA, and 15pL of primers (Primer-f: CTACACGACGCTCTTCCGATCT (SEQ ID NO:4); Primer-r:AAGCAGTGGTATCAACGCAGAG (SEQ ID NO:3)) was prepared in a total volume of lOOpL. The PCR regimen included: initial denaturation at 98°C for 3 minutes, followed by 14 cycles at 98°C for 15 seconds, 63°C for 20 seconds, and 72°C for 1 minute, concluding with a 5- minute final extension at 72°C. Purification of cDNA was performed using 0.7X Clean NGS-R Beads (CleanNA, CNGS-0050). The final elution yielded 20pL of size-selected cDNA, which was transferred to a fresh PCR tube for further analysis.

[0125] Fragmentation, End Repair, A-tailing and Ligation

[0126] For fragmentation, end repair, and A-tailing, an FS DNA Mix was prepared, consisting of 20pL Input DNA, lOpL Smearase® Mix (Yeasen, 12619ES24), and 30pL ddH2O. After gentle mixing and brief centrifugation, the mixture was subjected to a PCR program with a lid temperature set at 85°C. The regimen began with fragmentation at 30°C for 2 minutes, followed by end repair and A-tailing at 72°C for 30 minutes, and then held at 4°C.

[0127] For ligation, a Ligation Mix consisting of 60pL dA-tailed DNA, 20pL 5xNovel Ligation Buffer, 5pL Novel T4 DNA Ligase (Yeasen, 12626ES24), 3pL DNA Adapter (6pM,annealed from adapter A ( / 5'Phosph / GATCGGAAGAGCACACGTCTGAACTCCAGTCAC (SEQ ID NO:5)) and adapter B (5’-GCTCTTCCGATCT (SEQ ID NO:6))), and 12pL ddH2O was incubated at 25°C for 15 minutes.

[0128] Following ligation, the products underwent purification with Clean NGS-R Beads. Briefly, beads were mixed with the product, washed, and the DNA eluted. A size selection targeting 300-700 bp fragments was achieved using a stepwise application of 0.6* and then 0.8* bead volumes. The final purified DNA was resuspended, with 25 pL of the resultant solution transferred to a new PCR tube for subsequent analyses.

[0129] Amplification and Adapter Addition.

[0130] For library construction, a 50pL reaction mix was prepared using 25pL KAPA HiFi Hot Start Ready Mix, 2.5pL NEBNext Index primer (NEB, E7335S), 2.5pL NEB Universal Primer, and 20pL of purified ligated product. The PCR amplification was executed with an initial denaturation at 98°C for 45 seconds, followed by 14 cycles of 98°C for 20 seconds, 67°C for 30 seconds, and 72°C for 20 seconds. This was finalized with a 72°C extension for 1 minute and then held at 4°C indefinitely. Post-amplification, the library underwent size selection using Clean NGS-R Beads. Briefly, beads (0.6x and 0.8* volumes) were applied in successive steps, with washes and magnetic separation. The purified library was eluted in 30.5pL nuclease-free water, with 30pL transferred to a new PCR tube.

[0131] Animal Husbandry and tissue preparation. Male C57BL / 6 mice, eight weeks postnatal (PN), were used in this study. All procedures involving animals were conducted in strict accordance with the ethical guidelines established by the Institutional Animal Care and Use Committee of Zhengzhou University and were in compliance with the current laws regarding animal use in scientific research in China. Following perfusion, liver was harvested, embedded in OCT, stored at -80°C, and sectioned at a thickness of 10pm.

[0132] Hematoxylin and Eosin (H&E) Staining. A section 20pm away from the section used for spatial transcriptomics was selected for H&E staining. Using the Beyotime C0105S Hematoxylin and Eosin Staining Kit, tissue samples were fixed with 4% paraformaldehyde atroom temperature for 10 minutes. They were then stained with hematoxylin for 5 minutes, rinsed in water for 2 minutes, differentiated in 0.5% hydrochloric acid ethanol for 1 minute, followed by a 1-minute soak in IxPBS. Finally, eosin staining was applied for 1 minute, and excess dye was removed with a 2-minute water rinse.

[0133] Computational analysis: Data binning. Cho et al. showed that 10pm sided grids produced a much less noisy UMAP relative to 5pm square grids and was simultaneously able to discover clusters with single cell resolution (Cho, C.-S., et al. Cell. 2021, 184 (13), 3559- 3572. e22). Therefore, we combined counts of 5 consecutive features, performed convolution using Gaussian kernel with a width of 1 feature resulting in lOpm-sided binned features.

[0134] Identifying regions within mouse liver tissue using unsupervised clustering. A threshold for total transcript count was applied to include only those binned features that were overlaid by the tissue. This binned count matrix was processed using Seurat v4 R package to identify regions present within the tissue. Briefly, the binned count matrix was normalized using Seurat’s SCTransform function. Clusters using the normalized binned count matrix were identified using Seurat’s shared nearest neighbor graph-based algorithm implemented in the FindNeighbors and FindClusters functions. A resolution of 0.2 was used to perform clustering.

[0135] Visualization and annotation of clusters. A custom python script was used to visualize clusters in the context of the tissue section. Marker genes for each cluster were identified using Seurat’s FindAllMarkers function. Genes that have a positive logFC >0.25, p-adj < 0.05 and were present in at least 20% of cells were classified as marker genes. The marker genes identified in this study were compared to other studies to annotate clusters generated by unsupervised clustering.

[0136] Spatial Transcriptomics Data Analysis Using Cellllocation. The analysis was performed using the Cell21ocation library, specifically employed for spatial transcriptomics data deconvolution. The spatial transcriptomics data were loaded using Scanpy, targeting the H5AD file that contains the preprocessed spatial gene expression matrix. Similarly, annotated singlecell RNA-seq data from the Liver Cell Atlas was loaded as a reference dataset. Preliminaryexamination and preprocessing were performed on the spatial transcriptomics data. This included inspecting variable features and filtering out mitochondrial genes to clean the data for subsequent analysis. The single-cell RNA-seq data was prepared similarly, ensuring it was in an appropriate format for model input.

[0137] The Cell21ocation model was applied to deconvolute the spatial transcriptomics data, estimating cell type abundances at each spatial location. This involved integrating the spatial data with the reference single-cell RNA-seq data to create a comprehensive spatial map of cell distribution. Model parameters, fitting procedures, and diagnostic checks were carried out following best practices as recommended in Cell21ocation documentation and tutorials. The function cell21ocation.models.RegressionModel() prepares the AnnData object for compatibility with the Cell21ocation model, ensuring that the input data conforms to the expected format and structure. The function RegressionModel() initializes the RegressionModel object with the preprocessed and formatted AnnData. This step is crucial for specifying the model to be fitted to the spatial transcriptomics data. The rest of the parameters were the default model configurations.

[0138] Upon successful model convergence and estimation, the cell type abundance results were visualized using Matplotlib, with particular attention to the spatial distribution of different cell types within the liver sections. This step is critical for interpreting the cellular composition and architecture of the liver tissue, as well as understanding the heterogeneity and spatial organization of cell types.Example 8. Conclusions

[0139] This work demonstrates significant advances in the photolithographic manufacturing of spatial DNA chips with a highly optimized capture surface. High-resolution spatial DNA chips are manufactured at semi-conductor scale using photolithography. This production technique is suitable for large-scale, highly efficient manufacturing, leveraging the principles of Moore's Law. The development of high-resolution chips represents a pivotal advancement in spatial transcriptomic technologies, merging positional information with extensive tissuecoverage and efficient mRNA capture. To achieve these results, we utilized a combination of chemistry, chip design, and bioinformatics development.

[0140] The mammalian liver is a well-characterized, complex tissue containing many different cell types that perform coordinated metabolic functions. This makes it an excellent tissue to demonstrate the capabilities of new spatial sequencing technologies. We were able to validate the use of the chips with mouse liver, achieving sequencing results comparable to previously reported scRNA-seq data. This underscores the chip’s potential in comprehensively understanding tissue architectures and in early-stage pathological detection.

[0141] At 69% sequencing saturation, our sample yielded over 643 million effective reads from a single section with a median of 8,274 UMIs and 2, 181 genes per 20 x 20 pm area (the average size of an adult mouse liver cell), demonstrating exceptional molecular capture capacity. Additional sequencing is ongoing and we expect to obtain additional unique reads. By integrating the analysis techniques of U-CIE and Cell2Location, we have been able to decode complex gene expression patterns within liver tissues, enhancing our understanding of cellular diversity and structure. We believe the advancements demonstrated here establish a new benchmark for future research in the field. As the field evolves, integrating computational strategies and advanced microscopic methodologies with platforms like this will be paramount to deepening our insights into tissue dynamics and cellular intricacies.

[0142] While the present disclosure has been described with reference to the specific embodiments thereof, it should be understood by those skilled in the art that various changes may be made and equivalents may be substituted without departing from the true spirit and scope of the disclosure. In addition, many modifications may be made to adapt a particular situation, material, composition of matter, process, process step or steps, to the objective, spirit and scope of the present disclosure. All such modifications are intended to be within the scope of the claims appended hereto.

[0143] All references, issued patents and patent applications mentioned or cited within the body of the instant specification are hereby incorporated by reference in their entirety, to thesame extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference, for all purposes.

Claims

WHAT IS CLAIMED:

1. A system for analyzing a sample for the spatial distribution of at least one target molecule comprising:(a) a chip comprising a plurality of features, wherein each feature of the plurality of features comprises a plurality of barcode probes, and wherein each feature of the plurality of features is at a distinct location on the chip associated with two or more coordinates; and(b) a computer processor coupled to the chip and programmed to (i) measure at least one signal while the chip is in contact with the sample; and (ii) determine, based on the at least one signal, the spatial distribution of said at least one target molecule at the plurality of features.

2. A system for analyzing a sample for the spatial distribution of at least one target molecule comprising:(a) a chip comprising a plurality of features, wherein each feature of the plurality of features is at a distinct location on the chip associated with two or more coordinates; and(b) a plurality of barcode probes, wherein each barcode probe is attached to a feature, wherein the plurality of barcode probes are capable of binding to at least one target molecule, and wherein the location of each of the barcode probe is identified by the two or more coordinates.

3. The system of claim 1 or 2, wherein the plurality of barcode probes are nucleic acid probes.

4. The system of any one of claims 1 to 3, wherein the system comprises between approximately 5 million and approximately 10 million features, between approximately 10 million and approximately 15 million features, between approximately 15 million and approximately 20 million features, between approximately 20 million and approximately 25 million features, between approximately 25 million and approximately 30 million features, between approximately 30 million and approximately 35 million features, between approximately 35 million and approximately 40 million features, between approximately 40 million and approximately 45 million features, and between approximately 45 million and approximately 50 million features.

5. The system of claim 4, wherein the system comprises approximately 25 million features.

6. The system of any one of claims 1 to 5, wherein each feature comprises approximately 1 million to approximately 3 million barcode probes, approximately 3 million to approximately 5 million barcode probes, approximately 5 million to approximately 7 million barcode probes, approximately 7 million to approximately 9 million barcode probes, or approximately 9 million to approximately 11 million barcode probes.

7. The system of claim 6, wherein each feature comprises approximately 5 million barcode probes.

8. The system of any one of claims 1 to 7, wherein each of the features is a square of approximately 0.25 pm by 0.25 pm to approximately 0.75 pm by 0.75 pm in size, approximately 0.75 pm by 0.75 pm to approximately 1.25 pm by 1.25 pm in size, approximately 1.25 pm by 1.25 pm to approximately 1.75 pm by 1.75 pm in size, approximately 1.75 pm by 1.75 pm to approximately 2.25 pm by 2.25 pm in size, approximately 2.25 pm by 2.25 pm to approximately 2.75 pm by 2.75 pm in size, approximately 2.75 pm by 2.75 pm to approximately 3.25 pm by 3.25 pm in size, approximately 3.25 pm by 3.25 pm to approximately 4.75 pm by 4.75 pm in size, or approximately 4.75 pm by 4.75 pm to approximately 5.25 pm by 5.25 pm in size.

9. The system of claim 8, wherein each of the features is a square of about 2 pm by about 2 pm.

10. The system of any one of claim 1 to 9, wherein at least approximately 50%, at least approximately 60%, at least approximately 70%, at least approximately 80%, at least approximately 90%, or at least approximately 100% of the surface area of the chip is covered by at least one barcode probe.

11. The system of claim 10, wherein at least approximately 90% of the surface area of the chip is covered by at least one barcode probe.

12. The system of any one of claims 1 to 11, wherein at least two of the plurality of features are adjacent with no interstitial space between adjacent features.

13. A method for detecting the spatial distribution of a plurality of target molecules within a sample, comprising:(a) providing a substrate comprising a plurality of features, wherein each of the plurality of features comprises two or more coordinates, and wherein at least 90% of the surface area of the plurality of features is covered by at least one barcode probe;(b) contacting the surface with the sample, wherein the target molecules are capable of coupling with the barcode probes; and(c) determining spatial distribution of the target molecules on the substrate.

14. A spatial distribution pattern of target molecules from a sample, wherein the spatial distribution pattern comprises two or more coordinates for each of a plurality of features on a chip, wherein at least 90% of the surface area of the chip is covered by at least one barcode probe, and wherein the barcode probe is capable of binding to the target molecules.

15. A method of forming a pattern of oligonucleotides on a substrate having a plurality of functional groups, comprising: forming a vertical region with an X barcode sequence comprising a first nucleotide by at least coupling the first nucleotide to a portion of the plurality of functional groups in a first vertical exposed region of the substrate; and forming a second vertical region having an X barcode sequence comprising the first nucleotide and a second nucleotide by at least coupling the second nucleotide to a portion of the plurality of functional groups in a second vertical exposed region of the substrate that overlaps with the first vertical region.

16. The method of claim 15, further comprising: forming a third vertical region having an X barcode sequence comprising the first nucleotide, the second nucleotide, and a third nucleotide by at last coupling the third nucleotide toa portion of the plurality of functional groups in a third vertical exposed region of the substrate that overlaps with the second vertical region.

17. The method of claim 16, further comprising: forming, from a non-overlapping region between the third vertical exposed region and the second vertical region, a fourth vertical region having an X barcode sequence that comprises the first nucleotide and the second nucleotide but not the third nucleotide.

18. The method of any one of claims 15 to 17, wherein a location within each vertical region is identified by a corresponding X barcode sequence.

19. The method of any one of claims 15 to 18, further comprising: forming a first horizontal region with a Y barcode sequence comprising a third nucleotide by at least coupling the third nucleotide to a portion of the plurality of functional groups in a first horizontal exposed region of the substrate; and forming a second horizontal region with a Y barcode sequence comprising the third nucleotide and a fourth nucleotide by at least coupling the fourth nucleotide to a portion of the plurality of functional groups in a second horizontal exposed region of the substrate that overlaps with the first horizontal region.

20. The method of claim 19, further comprising: forming a third horizontal region having a Y barcode sequence comprising the third nucleotide, the fourth nucleotide, and a fifth nucleotide by at least coupling the fifth nucleotide to a portion of the plurality of functional groups in a third horizontal exposed region of the substrate that overlaps with the second horizontal region.

21. The method of claim 20, further comprising: forming, from a non-overlapping region between the third horizontal exposed region and the second horizontal region, a fourth horizontal region having a Y barcode sequence that includes the third nucleotide and the fourth nucleotide but not the fifth nucleotide.

22. The method of any one of claims 19 to 21 , wherein a location within each horizontal region of the substate is identified by a corresponding Y barcode sequence.

23. The method of claim 22, wherein an intersection between a vertical region and a horizontal region comprises at least one feature whose location is identified by a corresponding X barcode sequence and Y barcode sequence.

24. The method of any one of claims 15 to 23, wherein the coupling comprises:(i) forming a photoresist layer by applying a photoresist composition onto an underlying layer of a substrate comprising the plurality of functional groups, wherein the plurality of functional groups are protected by protective groups;(ii) exposing a dose of light through a patterned mask onto the substrate;(iii) removing the protective groups on a section of the plurality of functional groups within at least one exposed region of the substrate; and(iv) contacting the functional groups within the at least one exposed region of the substrate with a nucleotide reagent; thereby coupling a fraction of the functional groups within the at least one exposed region of the substrate with a nucleotide.

25. The method of any one of claims 15 to 24, wherein adjacent spatially-defined vertical and / or horizontal regions in the substrate have barcode sequences that differ by a single nucleotide26. The method of any one of claims 19 to 25, wherein at least 90% of the surface area of the substrate is covered with at least one barcode probe having a unique spatially-defined X barcode sequence and a unique spatially-defined Y barcode sequence.

Citation Information

Patent Citations

  • Methods, compositions, and systems for detecting exogenous nucleic acids

    US20230416850A1

  • Methods for performing spatial profiling of biological molecules

    WO2018217862A1