Method and system for splitting and merging spatial transcriptomics data for a sample

WO2026182892A1PCT designated stage Publication Date: 2026-09-03ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/013466
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2026-02-02
Publication Date
2026-09-03

Smart Images

  • Figure US2026013466_03092026_PF_FP_ABST
    Figure US2026013466_03092026_PF_FP_ABST
Patent Text Reader

Abstract

To split and merge spatial transcriptomics data for large samples to minimize memory storage and / or processing power, spatial transcriptomics data is obtained which includes sequence reads having a spatial barcode and a target nucleic acid. The sample is divided into sub-samples by assigning subsets of spatial barcodes for the sample to the sub-samples. Then respective sequence reads are demultiplexed into appropriate sub-samples according to the spatial barcodes for the respective sequence reads. Each sub-sample is analyzed. Then the results of the analysis are merged to generate a set of results for the sample. Another method includes obtaining a spatial barcode within a sequence read, performing base mismatch correction by partitioning the spatial barcode into substrings and comparing the substrings to substrings for predetermined or known spatial barcodes for the sample to identify an approximate match, and generating a corrected spatial barcode for the sequence read using the approximate match.
Need to check novelty before this filing date? Find Prior Art

Description

33080 / IP-2937 METHOD AND SYSTEM FOR SPLITTING AND MERGING SPATIAL TRANSCRIPTOMICS DATA FOR A SAMPLE CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 765,006, filed February 28, 2025, entitled “Method and System for Splitting and Merging Spatial Transcriptomics Data for a Sample,” the entire disclosure of which is hereby expressly incorporated by reference herein.FIELD OF THE INVENTION

[0002] The present disclosure generally relates to techniques for processing and analyzing spatial transcriptomics data, and more particularly, to techniques for splitting and merging spatial transcriptomics data for a sample on a substrate, such as demultiplexing respective sequence reads in accordance with spatial barcodes and correcting base mismatches.BACKGROUND

[0003] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in this background section, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure.

[0004] In the field of spatial transcriptomics, the precise localization of transcripts within tissue samples is fundamental for understanding cellular functions and interactions in a spatial context. Transcripts may be obtained from tissue samples placed on a substrate after tissue permeabilization and subsequent capture using target capture sites on the surface. The captured transcripts are then assigned locations to generate a spatial transcriptomics map.

[0005] However, the manipulation and analysis of data derived from spatial transcriptomics pose substantial computational challenges, notably due to the sheer volume of data. Each sample may include millions or even billions of sequence reads which requires significant memory storage and processing power. For example, a substrate may have an active area of 15 mm x 50 mm or 750 mm2which is large enough for a tissue sample with up to 7.5 million cells.33080 / IP-2937SUMMARY

[0006] The field of spatial transcriptomics has seen significant advancements in recent years, enabling researchers to obtain high-resolution spatial maps of gene expression within tissue samples (also referred to herein as “samples”). This technology provides invaluable insights into the cellular composition and organization of tissues, which is crucial for understanding various biological processes and disease mechanisms. However, the analysis of spatial transcriptomics data presents substantial computational challenges, primarily due to the vast amount of data generated from each sample. These challenges include the need for substantial memory storage and processing power, especially when dealing with large tissue samples that can contain millions of cells. The system introduces a novel approach to address these challenges, focusing on sample partitioning, efficient barcode processing, and parallel analysis to improve processing, network usage, and memory usage in computers handling spatial transcriptomics data.

[0007] To address these challenges, the present techniques involve a method for splitting and merging spatial transcriptomics data for a sample on a substrate. This method begins with obtaining a set of spatial transcriptomics data for a sample, which includes generating a tissue mask to define the boundaries of the sample and capturing sequence reads within these boundaries. The sample is then divided into multiple sub-samples, each assigned a different subset of spatial barcodes. This division is based on various criteria, including the memory size of the spatial transcriptomics data for the sample and the processing power available, to ensure that the memory sizes of the sub-samples do not exceed a threshold memory storage capacity.

[0008] One of the key improvements offered by the present techniques is the enhanced processing efficiency. By dividing the sample into sub-samples and analyzing these subsamples simultaneously, the techniques significantly reduce the run-time for analyzing large samples. This parallel processing approach allows for the handling of samples that would otherwise be too large to analyze as a single entity due to memory and processing constraints.

[0009] Another significant improvement is the optimization of memory usage. By dividing the sample into sub-samples based on memory size considerations, the system ensures that the analysis can be conducted within the available memory resources, thereby preventing memory overflow issues that could lead to system crashes or failed analyses. This approach prevents system overloads and crashes that could occur when attempting to process large amounts of data simultaneously. Moreover, this approach enables the analysis of large samples without33080 / IP-2937 requiring an impractical increase in memory capacity, making spatial transcriptomics more accessible and feasible for a broader range of research applications.

[0010] Additionally, by partitioning the sample and analyzing the sub-samples in parallel, the system significantly reduces the run-time for analyzing large samples. This approach not only makes it feasible to process large tissue samples that were previously beyond the capabilities of existing tools but also optimizes the use of computational resources, leading to a more efficient processing workflow.

[0011] Furthermore, the present techniques improve the accuracy of spatial transcriptomics data analysis. The method includes steps for performing base mismatch correction. This ensures that the spatial barcodes are accurately matched to the predetermined spatial barcodes for the substrate, thereby enhancing the reliability of the spatial transcriptomics map generated from the analysis. The base mismatch correction, in particular, is performed in a manner that significantly reduces the run-time and processing power required, by partitioning the spatial barcode into nucleotide substrings and utilizing hash tables for efficient comparison and correction.

[0012] Furthermore, the system introduces an innovative method for improving network usage through efficient barcode processing. The system utilizes hash tables to perform base mismatch correction for spatial barcodes. By partitioning the nucleotide string for each predetermined spatial barcode into multiple substrings and storing each substring in a separate hash table, the system reduces the computational complexity of identifying approximate matches. This method significantly decreases the run-time for barcode processing.

[0013] The system's approach to sample partitioning, barcode processing, and parallel analysis represents a comprehensive solution to the computational challenges associated with spatial transcriptomics. By addressing the need for improved processing efficiency, optimized memory usage, and efficient network usage, the system enables the analysis of large tissue samples that were previously unmanageable with existing tools. This advancement opens up new possibilities for spatial transcriptomics research, allowing for the exploration of larger and more complex tissue samples than ever before. The ability to efficiently process and analyze these samples is expected to contribute to the understanding of tissue organization, disease progression, and the underlying mechanisms of various biological processes.

[0014] In one aspect, a method for splitting and merging spatial transcriptomics data for a sample on a substrate includes: (1) obtaining, by one or more processors, a set of spatial33080 / IP-2937 transcriptomics data for a sample on a substrate, the set of spatial transcriptomics data including a plurality of sequence reads of nucleic acid molecules having a spatial barcode and a target nucleic acid; (2) dividing, by the one or more processors, the sample into a plurality of sub-samples each assigned a different subset of spatial barcodes; (3) demultiplexing, by the one or more processors, respective sequence reads into one of the plurality of sub-samples in accordance with the spatial barcodes for the respective sequence reads; (4) analyzing, by the one or more processors, the plurality of sub-samples; and (5) merging, by the one or more processors, results of the analysis of each of the plurality of sub-samples to generate a set of results for the sample.

[0015] In another aspect, a method for base mismatch correction of spatial barcodes includes: (1) obtaining, by one or more processors, a spatial barcode within a sequence read of a sample on a substrate; (2) performing, by the one or more processors, base mismatch correction of at least one nucleobase for the spatial barcode to match a predetermined spatial barcode in a set of predetermined spatial barcodes for the substrate by: (a) partitioning the spatial barcode into a plurality of nucleotide substrings; (b) comparing the plurality of nucleotide substrings for the spatial barcode to a plurality of sets of nucleotide substrings for the set of predetermined spatial barcodes to identify the predetermined spatial barcode having nucleotide substrings which do not differ from the plurality of nucleotide substrings for the spatial barcode by more than a threshold number of nucleobases; and (c) generating, by the one or more processors, a corrected spatial barcode for the sequence read using the predetermined spatial barcode.

[0016] Advantages will become more apparent to those of ordinary skill in the art from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects.Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The figures described below depict various aspects of the system and methods disclosed herein. It should be understood that each figure depicts an embodiment of a particular aspect of the disclosed system and methods, and that each of the figures is intended to accord with a possible embodiment thereof.33080 / IP-2937

[0018] There are shown in the drawings arrangements which are presently discussed, it being understood, however, that the present embodiments are not limited to the precise arrangements and instrumentalities shown, wherein:

[0019] Fig. 1 depicts a block diagram of an example computing environment for splitting and merging spatial transcriptomics data for a sample on a substrate, according to some aspects.

[0020] Fig. 2 depicts an example substrate having an active area for sequencing nucleic acid molecules from a tissue sample, according to some aspects.

[0021] Fig. 3A depicts example samples on a substrate each divided into multiple subsamples, according to some aspects.

[0022] Fig. 3B depicts a flow diagram of example method for splitting a sample into multiple sub-samples which are stored as separate files, analyzing each sub-sample separately using a genomic analysis tool, and merging the results of the analyses to generate a single set of results for the sample, according to some aspects.

[0023] Figs. 4A and 4B depict example data structures for storing a predetermined set of spatial barcodes for a substrate, according to some aspects.

[0024] Fig. 5A depicts an example data structure for separately storing nucleotide substrings of predetermined spatial barcodes for a substrate and respective indexes for retrieving the nucleotide substrings, according to some aspects.

[0025] Fig. 5B depicts an example data structure for storing predetermined spatial barcodes and indications of the corresponding sub-samples to which they belong, according to some aspects.

[0026] Fig. 6 depicts a comparison of the number of sequence reads having a valid spatial barcode with and without base mismatch correction, according to some aspects.

[0027] Fig. 7 illustrates example spatial transcriptomics maps for the sub-samples and for the merged sample, according to some aspects.

[0028] Fig. 8 is a flow diagram of an example method for splitting and merging spatial transcriptomics data for a sample on a substrate, which may be implemented by a server device, according to some aspects.

[0029] Fig. 9 is a flow diagram of an example method for base mismatch correction of spatial barcodes, which may be implemented by a server device, according to some aspects.33080 / IP-2937

[0030] The Figures depict preferred implementations for purposes of illustration only.Alternative implementations of the systems and methods illustrated herein may be employed without departing from the principles of the invention described herein.DETAILED DESCRIPTION

[0031] Although the following text discloses a detailed description of implementations of methods, apparatuses and / or articles of manufacture, it should be understood that the legal scope of the property right is defined by the words of the claims set forth at the end of this document. Accordingly, the following detailed description is to be construed as examples only and does not describe every possible implementation, as describing every possible implementation would be impractical, if not impossible. Numerous alternative implementations could be implemented, using either current technology or technology developed after the filing date of this patent. It is envisioned that such alternative implementations would still fall within the scope of the claims.

[0032] To analyze very large samples while taking into account processing power and memory constraints, a sample partitioning system divides sequence reads from a sample into multiple sub-samples. The sample partitioning system may determine the number of subsamples in which to divide the sample based on the size of the sample and memory and / or processing power constraints. For example, a 750 mm2sample may require about X gigabytes (GB) of memory and the maximum amount of memory for processing a sub-sample may be Y GB. Therefore, the number of sub-samples may be X / Y. This may also reduce the run-time for analyzing a sample by processing each of the sub-samples in parallel (e.g., on multiple processors).

[0033] Each sequence read may be a nucleic acid molecule including a spatial barcode used to identify the location of the sequence read, a target nucleic acid (e.g., DNA, cDNA, mRNA, or RNA), and / or a molecular identifier (e.g., a unique molecular identifier (UMI)). The sample partitioning system may assign spatial barcodes to each sub-sample. For example, during a first sequencing step (first seq), the sample partitioning system may identify each of the spatial barcodes on the substrate and their respective locations.

[0034] The sample partitioning system may also generate a tissue mask indicating the boundaries of the sample within the substrate and may identify the spatial barcodes within the boundaries of the tissue mask. In any event, the sample partitioning system may assign the spatial barcodes, so they are evenly distributed across the sub-samples. For example, the33080 / IP-2937 sample partitioning system may randomly assign the spatial barcodes to the sub-samples. In other implementations, the sample partitioning system may assign the spatial barcodes to the sub-samples based on their location (e.g., such that spatial barcodes corresponding to a top subsection of the sample correspond to sub-sample 1 , spatial barcodes corresponding to a middle subsection of the sample correspond to sub-sample 2, and spatial barcodes corresponding to a bottom subsection of the sample correspond to sub-sample 3), or may assign spatial barcodes to sub-samples in any suitable manner. In some implementations, the sample partitioning system may assign spatial barcodes outside of the boundaries of the tissue mask to its own sample which may then be divided into sub-samples.

[0035] The sample partitioning system may then demultiplex sequence reads into the subsamples based on their respective spatial barcodes. For example, sequence reads having spatial barcodes within a first set of spatial barcodes may be assigned to sub-sample 1 , sequence reads having spatial barcodes within a second set of spatial barcodes may be assigned to sub-sample 2, etc. The sample partitioning system may store each sub-sample in a separate file (e.g., a Fastq file).

[0036] To demultiplex sequence reads based on their respective spatial barcodes, the sample partitioning system may compare a spatial barcode for a sequence read to a set of predetermined spatial barcodes for the substrate to ensure that the spatial barcode for the sequence read is accurate. The sample partitioning system may also perform base mismatch correction for the spatial barcode if the nucleobases in the spatial barcode do not differ by more than a threshold number of nucleobases from a predetermined spatial barcode (e.g., one).

[0037] However, performing the base mismatch correction by comparing the nucleotide string in the spatial barcode to the nucleotide strings for every predetermined spatial barcode to identify an approximate match would require extensive processing power and would be very time consuming (e.g., on the order of minutes per sequence read for the 750 million possible spatial barcode sequences which would take years to analyze a single sample). Instead, the sample partitioning system partitions the nucleotide string for each predetermined spatial barcode into multiple substrings and stores each substring in a separate hash table. For example, each predetermined spatial barcode may be 30 nucleotides long, and the sample partitioning system may store two 15 nucleotide hash tables for the first half and the second half of the predetermined spatial barcodes, respectively.

[0038] Then the sample partitioning system may divide a spatial barcode for a sequence read into substrings and may compute a hash function for each substring which may be used to33080 / IP-2937 index the hash tables. For example, the hash function may be converting the nucleobases in the substring into a numerical value. The sample partitioning system may compare the computed indexes for the substrings to the indexes in the hash tables to identify a match. If there is a match for at least one of the substrings of the spatial barcode for a sequence read in one of the hash tables, the sample partitioning system can filter out each of the predetermined spatial barcodes which do not include the matching substring. The resulting subset may be a small number of predetermined spatial barcodes compared to the 750 million possible spatial barcode sequences. Then the sample partitioning system may compare the remaining substring(s) of the spatial barcode for the sequence read to the remaining substring(s) in the filtered set to identify an approximate match which does not differ from the spatial barcode for the sequence read by more than a threshold number of nucleobases (e.g., one). The sample partitioning system may correct the spatial barcode for the sequence read using the identified spatial barcode from the predetermined set. This process requires a minimal number of operations compared to a process for performing an approximate match comparison with each of the predetermined spatial barcodes, and therefore the run-time is significantly reduced.Utilizing this base mismatch correction process, the sample partitioning system may demultiplex 33 billion reads in about 2 hours.

[0039] In any event, the sample partitioning system may separately analyze each of the subsamples (e.g., using Dynamic Read Analysis for GENomics (DRAGEN) or any other sequencing analysis tools) to generate a subset of results for each sub-sample. For example, the analysis may include de-duplicating sequence reads of the same molecule having the same spatial barcode, target nucleic acid, and / or molecular identifier in a sub-sample. In this manner, the sub-sample may include a set of unique molecules. The analysis may also include identifying a gene and / or genetic variant corresponding to a target nucleic acid. Moreover, the analysis may include identifying the location of the gene and / or genetic variant according to the spatial barcode. In some implementations, the sample partitioning system can analyze each of the sub-samples in parallel to reduce the run-time for the analysis.

[0040] Additionally, the sample partitioning system may merge the results of the analysis for each of the sub-samples to generate a set of results for the sample. For example, the set of results may include a superset of the sets of unique molecules for the sub-samples. Then the sample partitioning system may present the results of the analysis, for example for display to a user. More specifically, the sample partitioning system may present a spatial transcriptomics map indicating the transcripts and their respective locations within the sample.33080 / IP-2937 EXEMPLARY COMPUTING ENVIRONMENT

[0041] Fig. 1 is a block diagram of an example computing environment 100 for splitting and merging spatial transcriptomics data for a sample on a substrate, as described herein. This computing environment 100 is designed to address the challenges of processing large-scale spatial transcriptomics data by partitioning the data into manageable sub-samples, analyzing these sub-samples, and then merging the results. The example computing environment 100 may include a sequencing device 110, a server device 120, a client device 130, tissue samples 102a-102d on a substrate 140, and a network 150.

[0042] The substrate 140 (e.g., a flow cell) may include a very large set of (e.g. up to hundreds of millions) of barcoded locations, each containing a sequence of capture oligonucleotides constituting a spatial barcode unique to that location. During first seq, the sequencing device 110 may decode the spatial barcodes for creating a map of barcode sequences and coordinate locations. A tissue sample 102a-102d is then placed on the substrate 140 and target nucleic acids (e.g., polyA mRNA molecules) within the tissue 102a-102d diffuse to the features and are captured on the substrate 140 during permeabilization. In some embodiments, the captured target nucleic acid is RNA that is reverse transcribed into cDNA, wherein the spatial barcode is linked with the cDNA sequence.

[0043] During permeabilization, the tissue sample 102a-102d may release mRNA which binds to capture oligonucleotides from a proximal location on the tissue 102a-102d. A reverse transcription reaction may occur while the tissue 102a-102d is still in place, generating a library that incorporates the spatial barcodes and preserves spatial information. For example, each library may include a combination of mRNA, a barcode, an index, a molecular identifier (e.g., a UMI), and / or other known sequences. The barcoded sequences are mapped back to a specific location within the flow cell 140.

[0044] In further implementations, capture oligonucleotides and spatially barcoded oligonucleotides having a spatial barcode may be separated but proximal to one another on the substrate. In such implementations, a mechanism may be used after molecule capture, where the captured molecules may bridge to a nearby spatially barcoded nucleotide, for example as described in U.S. Provisional Patent Application No. 63 / 615,558, which is incorporated by reference herein in its entirety. A reverse transcription reaction may occur resulting in a library having both mRNA and a spatial barcode, which the sequencing device 110 may sequence.

[0045] This is followed by library preparation, which typically involves amplification (e.g., PCR). The sequencing library is then subjected to sequencing (e.g., Next Generation33080 / IP-2937 Sequencing). During analysis of the sequencing data, the spatial barcode is used to map the physical location of the molecule from which the sequence read is derived.

[0046] As used herein, the term "barcode" is intended to mean a series of nucleotides in an oligonucleotide that can be used to identify the oligonucleotide, a spatial address on a surface (i.e., a “spatial barcode”), a characteristic of the oligonucleotide, or a manipulation that has been carried out on the oligonucleotide. The barcode can be a naturally occurring nucleotide sequence or a nucleotide sequence that does not occur naturally in the organism from which the barcoded nucleic acid was obtained. A barcode sequence can be unique to a single nucleic acid species in a population, or a barcode sequence can be shared by several different nucleic acid species in a population.

[0047] In some embodiments, one or more of the plurality of spatially barcoded oligonucleotides comprises a molecule identifier (Ml). In further embodiments, the Ml is a UMI.

[0048] As used herein, the term “molecular identifier” or “Ml” refers to a molecular tag, either random, non-random, or semi-random, that may be attached to a nucleic acid. In various embodiments, an Ml is a UMI. When incorporated into a nucleic acid, an Ml can be used to correct for subsequent amplification bias by directly counting Mis that are sequenced after amplification. An Ml (e.g., a UMI) can be attached to similar nucleic acids, e.g., adapters, making each nucleic acid unique. Mis (e.g., UM Is) may also be used to uniquely tag individual molecules (e.g., individual mRNA molecules) in a sample (e.g., individual mRNA molecules in a tissue sample, cell sample, or sample library). In some embodiments, a UMI is a random nucleotide sequence (e.g., N9).

[0049] UMIs may be applied to or identified in individual DNA molecules. In some implementations, the UMIs may be applied to the DNA molecules by methods that physically link or bond the UMIs to the DNA molecules, e.g., by ligation or transposition through polymerase, endonuclease, transposases, etc. These “applied” UMIs are therefore also referred to as physical UMIs. In some contexts, they may also be referred to as exogenous UMIs. The UMIs identified within source DNA molecules are referred to as virtual UMIs. In some context, virtual UMIs may also be referred to as endogenous UMI.

[0050] UMIs are uniquely associated with a single DNA fragment in a sample including a source polynucleotide and its complementary strand. A physical UMI is a sequence of an oligonucleotide linked to the source polynucleotide, its complementary strand, or a polynucleotide derived from the source polynucleotide. A virtual UMI is a sequence of an33080 / IP-2937 oligonucleotide within the source polynucleotide, its complementary strand, or a polynucleotide derived from the source polynucleotide. Within this scheme, one may also refer to the physical UMI as an extrinsic or exogenous UMI, and the virtual UMI as an intrinsic or endogenous UMI.

[0051] In any event, during a second sequencing step (second seq), the sequencing device 110 may image the flow cell 140 having molecules from the tissue samples 102a-102d with barcoded sequences. As used herein, the term “first seq” generally refers to a first sequencing step where barcode sequences on a substrate 140 are read out together with their spatial coordinates. The term “second seq” generally refers to a second sequencing step where the substrate 140 of the first sequencing step is exposed to a sample 102a-102d whose molecules are captured and read out together with corresponding barcode sequences. "Reading out” the barcode sequences and spatial coordinates may include obtaining and interpreting genetic information (e.g., RNA transcripts) from a biological sample 102a-102d, which includes preserving the spatial context of where each transcript is located within the biological sample / tissue 102a-102d. Steps involved in “reading out” may include tissue preparation and sequencing by sectioning a tissue sample 102a-102d onto a substrate 140 that includes an array of barcoded probes, hybridization and reverse transcription, sequencing using nextgeneration sequencing such as the sequencing device 110, and data analysis and spatial mapping by mapping sequence reads back to locations on the tissue slide using the barcode sequences and interpretation.

[0052] The sequencing device 110 may include a computing device, image sensors, and a sequencing application 112 for sequencing a genomic sample or other nucleic-acid polymer. In some versions, by executing the sequencing application 112 using a processor, the sequencing device 110 may analyze nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on the sequencing device 110.

[0053] More particularly, the sequencing device 110 may receive the flow cell 140 with barcoded nucleotide fragments extracted from the tissue sample 102a-102d, and the sequencing device 110 may determine the nucleobase sequence of such extracted nucleotide fragments. The sequencing device 110 may be the sequencing system described in U.S.Patent Application No. 18 / 340,795, titled “Split-Read Alignment by Intelligently Identifying and Scoring Candidate Split Groups filed on July 23, 2023, which is hereby incorporated by reference in its entirety.33080 / IP-2937

[0054] In some versions, the sequencing device 110 may utilize sequencing-by-synthesis (SBS) to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. By executing the sequencing application 112, the sequencing device 110 may further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and send the BCL file to the server device 120. The sequencing device 110 may communicate the BCL file and / or other data to the sever device 120 via one or more network(s) 150 or directly (e.g., bypassing the one or more network(s) 150).

[0055] In some implementations, prior to permeabilization, the sequencing device 110 may decode the spatial barcodes for creating a map of barcode sequences and coordinate locations. Then after permeabilization, the cDNA along with the barcode may be eluted and sequenced on a second flow cell.

[0056] As used herein, the term “cluster of oligonucleotides” (or simply “cluster(s)” or “DNA cluster(s)”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell 140, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell 140 or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a nonpatterned flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.

[0057] The server device 120 receives the spatial transcriptomics data including sequence reads from the sequencing device 110. As mentioned above, a sequence read may include a spatial barcode, a target nucleic acid (e.g., cDNA, DNA, RNA, or mRNA), a molecular identifier (e.g., a UMI), an index, and / or other known sequences. The server device 120 may include a memory and one or more processors (CPUs). The memory can be a non-transitory memory and can include one or several suitable memory modules, such as random access memory (RAM), read-only memory (ROM), flash memory, other types of persistent memory, etc.

[0058] The memory may store an operating system (OS), which can be any type of suitable mobile or general-purpose operating system. The memory also stores a sample split and merge engine 122 that divides a sample 102a into sub-samples, assigns a set of spatial barcodes to each sub-sample, and demultiplexes sequence reads into the appropriate sub-33080 / IP-2937 samples based on their spatial barcodes. The sample split and merge engine 122 stores each sub-sample as a separate file (e.g., a Fastq file). Then the sample split and merge engine 122 separately analyzes each sub-sample for example, using a genomic analysis tool (e.g., DRAGEN) to de-duplicate duplicate sequence reads and generate a set of results for each subsample. The sample split and merge engine 122 merges the sets of results for each subsample to generate a combined set of results for the sample 102a. The results may include transcripts (e.g., of gene and / or genetic variants) along with their respective locations within the sample 102a.

[0059] The server device 120 may then transmit the combined set of results for the sample 102a to the client device 130. The client device 130 may include a memory and one or more processors (CPUs). The memory can be a non-transitory memory and can include one or several suitable memory modules, such as random access memory (RAM), read-only memory (ROM), flash memory, other types of persistent memory, etc.

[0060] The memory may store an operating system (OS), which can be any type of suitable mobile or general-purpose operating system. The memory also stores a spatial transcriptomics application 132 which may present the results of the genomic analysis for the sample 102a. The results may include transcripts (e.g., of genes and / or genetic variants) corresponding to a set of unique molecules identified for the sample 102a and their respective locations.Additionally, the spatial transcriptomics application 132 may present a spatial transcriptomics map using the combined set of results for the sample 102a via a user interface. For example, the spatial transcriptomics application 132 may present a heatmap indicating the number of transcripts at various locations within the tissue sample 102a.

[0061] In this manner, users can identify which transcript corresponds to which cell, determine cell types based on corresponding transcripts, assess transcript density within cells, identify elevated gene expression, and analyze cell-to-cell interactions. The client device 130 may also perform this analysis and provide an indication of the transcript density within cells, elevated gene expression in particular cells, cell-to-cell interactions, etc. for display to the user.

[0062] While the sample split and merge engine 122 is shown as part of the server device 120, the functionality of the sample split and merge engine 122 may be included in the client device 130, another computing device, or any suitable combination of these. For example, some portions of the sample split and merge engine 122 may be included in a client device 130 or local computing device while other portions of the sample split and merge engine 122 may be included in a server device 120 or remote computing device.33080 / IP-2937

[0063] More specifically, a local computing device communicatively coupled to the sequencing device 110 (e.g., a laptop, a desktop computer, a tablet computer, etc.) may divide the sample into sub-samples and demultiplex sequence reads into the sub-samples. Then the local computing devices may provide the sub-samples to the server device 120. The server device 120 may then analyze the sub-samples and merge the results of the analysis to generate a result of results for the sample.

[0064] Fig. 2 illustrates an example substrate 140 for sequencing nucleic acid molecules derived from a tissue sample 102a. As mentioned above, the substrate 140 includes clusters of oligonucleotides having spatial barcodes which are decoded during first seq. During permeabilization, the tissue sample 102a may release mRNA which binds to the capture oligonucleotides from a proximal location on the tissue sample 102a. The active area 142 of the substrate 140, where the binding and subsequent sequencing takes place, may be about 15 mm by 50 mm. This sizeable area allows for the accommodation of tissue samples as large as 750 mm2, which could contain approximately 7.5 million cells. The ability to process such a large number of cells is significant for comprehensive analysis but presents challenges in terms of the computational resources required, specifically memory capacity and processing time.

[0065] To address these challenges, the server device 120 receives the sequence reads for a sample 102a from the sequencing device 110 and divides the sample 102a into subsamples 302-306, as shown in Fig. 3A. The server device 120 may determine the number of subsamples in which to divide the sample based on the size of the sample and memory and / or processing power constraints, so that the downstream analysis can run successfully.

[0066] For example, a sample may require about 9 terabytes (TB) of memory based on the size of the sample and the maximum amount of memory for processing a sub-sample (and performing downstream analysis) may be 1 TB. Therefore, the number of sub-samples may be calculated as the total memory required to store the sample divided by a maximum threshold amount of memory for performing the genomic analysis or 9. In another example, the number of sub-samples may be calculated as an estimated run-time for analyzing a sample divided by a maximum threshold run-time for performing the genomic analysis. In other implementations, the number of sub-samples may be determined based on any suitable combination of the memory size of the sample, an estimated run-time for analyzing the sample, a maximum threshold amount of memory for performing the genomic analysis, and / or a maximum threshold run-time for performing the genomic analysis.33080 / IP-2937

[0067] In some implementations, the sample split and merge engine 122 may obtain sample metadata information indicating the density of transcripts (e.g., based on nuclei density detection in a microscope image of the tissue sample). The sample split and merge engine 122 may then estimate the total number of molecules in the sample and may determine the processing requirements ( / .e., compute and memory requirements) for the sample based on the number of molecules. Then the sample split and merge engine 122 can determine the number of sub-samples in which to divide the sample based on the processing requirements for the sample.

[0068] In any event, the sample split and merge engine 122 assigns spatial barcodes to each of the sub-samples 302-306. The sample split and merge engine 122 may generate a tissue mask indicating the boundaries of the sample 102a within the substrate 140 and may identify the spatial barcodes within the boundaries of the tissue mask.

[0069] In some implementations, rather than including a few large samples, the substrate 140 may include hundreds of small samples (e.g., a tissue microarray). In these implementations, the sample split and merge engine 122 may divide regions of the substrate 140 (e.g., tens of small sample regions) into sub-regions and may assign spatial barcodes to each of these subregions. The sample split and merge engine 122 may determine the number of sub-regions in which to divide a region based on the memory size of the region, an estimated run-time for analyzing the region, a maximum threshold amount of memory for performing the genomic analysis, and / or a maximum threshold run-time for performing the genomic analysis

[0070] To generate the tissue mask, the sample split and merge engine 122 may obtain a microscope image of the substrate 140 from a microscope. The microscope image may include the sample 102a on the substrate 140 with fiducials or markers at known, physical locations within the substrate 140. The sample split and merge engine 122 may use the fiducials and / or other alignment techniques to stitch fields of view of the microscope image together to generate a complete view of the microscope image of the substrate 140 and align the coordinates of the microscope image and the spatial barcodes to a common coordinate system. Example techniques for aligning the coordinates to a common coordinate system are described in U.S. Provisional Patent Application Nos. 63 / 735,519 entitled “Microscope Image and Spatial Transcript Data Alignment Using Fiducial and Feature-Based Alignment,” and 63 / 735,535 entitled “Spatial Image Registration to Ground Truth,” each of which are incorporated by reference herein. In any event, the sample split and merge engine 122 may generate the tissue mask indicating the boundaries of the sample 102a within the substrate 140 based on the33080 / IP-2937 respective coordinates of the boundaries of the sample 102a in the common coordinate system. The sample split and merge engine 122 may generate the tissue mask with some margin of error around the tissue boundaries to account for noise in the boundary detection process and / or to capture transcripts near the tissue boundaries that likely resulted from short range diffusion.

[0071] Then the sample split and merge engine 122 may assign the spatial barcodes, so they are evenly distributed across the sub-samples 302-306. For example, the sample split and merge engine 122 may randomly assign the spatial barcodes to the sub-samples 302-306 to randomly distribute the spatial barcodes across the sub-samples.

[0072] In other implementations, the sample split and merge engine 122 may assign the spatial barcodes to the sub-samples based on their location, segmenting the sample 102a into a specific number of similarly sized sections. This can be visualized as drawing grid lines across the sample 102a, either from top to bottom or from left to right, depending on the desired orientation. Each grid section then represents a sub-sample 302-306, with the spatial barcodes contained within that section belonging to that sub-sample.

[0073] For example, spatial barcodes corresponding to a top subsection 302 of the sample may correspond to sub-sample 1a, spatial barcodes corresponding to a middle subsection 304 of the sample may correspond to sub-sample 1b, and spatial barcodes corresponding to a bottom subsection 306 of the sample may correspond to sub-sample 1c. In other implementations, the sample split and merge engine 122 may assign spatial barcodes to subsamples in any suitable manner. Also in some implementations, the sample split and merge engine 122 may assign spatial barcodes outside of the boundaries of a tissue mask to its own sample which may then be divided into sub-samples.

[0074] Furthermore, as shown in Fig. 3A, multiple samples 102a-102d on the substrate 140 may be divided into sub-samples. While the process described herein focuses on the division of one sample 102a into sub-samples, this is for ease of explanation only. This process may be repeated for several samples 102a-102d on the substrate 140. The server device 120 may obtain each of the sequence reads for each of the samples on the substrate 140 (e.g., in a base call (BCL) file) which are then demultiplexed into sub-samples.

[0075] Additionally, while the process described herein divides a sample into sub-samples and then merges the results of the analyses for the sub-samples into a set of results for the sample, the substrate 140 may be divided into sub-regions (e.g., a grid of X pm x X pm sub-33080 / IP-2937 regions). Then the server device 120 may analyze the sequence reads in the sub-regions and identify whether and / or which sample a sub-region corresponds to. The server device 120 may then merge the results of the analyses for sub-regions for a particular sample into a set of results for the sample.

[0076] Fig. 3B depicts a flow diagram of an example method 350 for splitting and merging spatial transcriptomics data for a sample 102a. The method 350 may be implemented by the server device 120, and more specifically, the sample split and merge engine 122.

[0077] First, the server device 120 receives spatial transcriptomics data including sequence reads for a sample 102a from the sequencing device 110 obtained during second seq (ref. no.352). Then the server device 120 demultiplexes or splits sequence reads into sub-samples (e.g., sub-sample 1a, sub-sample 1b, and sub-sample 1c) based on their respective spatial barcodes. For example, sequence reads having spatial barcodes within a first set of spatial barcodes may be assigned to sub-sample 1a, sequence reads having spatial barcodes within a second set of spatial barcodes may be assigned to sub-sample 1 b, etc. The server device 120 stores the sub-samples in separate files (ref. no. 354), typically in a format conducive to genomic analysis, such as a Fastq file. For example, the server device 120 may store the sequence reads for the sub-samples in a sub-sample 1a file, a sub-sample 1b file, and a subsample 1c file.

[0078] The server device 120 then proceeds to analyze each of the sub-sample files separately (ref. no. 356) using a genomic analysis tool, such as DRAGEN. This tool performs several functions, including identifying genes and genetic variants present in each sub-sample and de-duplicating sequence reads having the same spatial barcode, target nucleic acid, and / or UMI. De-duplication removes artificial duplicates that may have been introduced during the sequencing process, ensuring that the analysis accurately reflects the number of genes and / or genetic variants in the original sample 102a.

[0079] In some implementations, the server device 120 analyzes each sub-sample file one at a time. In other implementations, the server device 120 analyzes the sub-sample files simultaneously in parallel. For example, the server device 120 may include several processors, where a separate processor analyzes each sub-sample file. In another example, the server device 120 may include multiple server devices where a separate server device analyzes each sub-sample file.33080 / IP-2937

[0080] In any event, the server device 120 generates results for each sub-sample (e.g., subsample 1a results, sub-sample 1b results, and sub-sample 1c results). The results may include transcripts (e.g., of genes and / or genetic variants) corresponding to a set of unique molecules identified for the sub-samples and their respective locations. Then the server device 120 merges the results of these analyses into a single set of results for the entire sample 102a (ref. no. 358). The server device 120 may then store the merged set of results for the sample 102a, typically in a database or a file system, for further analysis or reporting.

[0081] To demultiplex sequence reads based on their respective spatial barcodes, the server device 120 may compare a spatial barcode for a sequence read to a set of predetermined spatial barcodes for the substrate 140 and / or within the boundaries of the tissue mask for the sample 102a to ensure that the spatial barcode for the sequence read is accurate. The server device 102a may also perform base mismatch correction for the spatial barcode if the nucleobases in the spatial barcode do not differ by more than a threshold number of nucleobases from a predetermined spatial barcode (e.g., one).

[0082] The server device 120 may store a set of predetermined or known spatial barcodes for the substrate 140 and / or within the boundaries of the tissue mask for the sample 102a in a data structure. Fig. 4A illustrates an example data structure 400 for storing the set of predetermined spatial barcodes. The data structure 400 may be a hash table indexed by computing the output of a hash function on a predetermined spatial barcode. For example, the hash function may be converting the nucleobases in a predetermined spatial barcode into a numerical value. More specifically, each nucleobase may correspond to a base 5 value, where for example “A” corresponds to the number ‘1 ,’ “C” corresponds to the number ‘2.’ “T” corresponds to the number ‘3,’ and “G” corresponds to the number ‘4, ’ and no nucleotide corresponds to the number ‘0.’ In other implementations, the nucleobases may be represented with less than five different digits. For example, if the spatial barcodes are not random, each nucleotide position may one of three nucleobases instead of four (e.g., cycle 1 could be A / C / G, cycle 2 could be A / G / T, etc.). The hash table design may be optimized based on the known structure of the spatial barcodes.

[0083] In this manner, the server device 120 may determine whether a spatial barcode for a sequence read is within the set of predetermined or known spatial barcodes for a sample 102a by computing the hash function on the nucleotide sequence for the spatial barcode to generate an index value and looking up whether there is a predetermined spatial barcode at the index value in the hash table 400. This allows for fast lookup with only one operation, 0(1 ), and does33080 / IP-2937 not requiring searching every predetermined spatial barcode in the data structure 400, O(n), which may include hundreds of millions of predetermined spatial barcodes.

[0084] If the hash table does not have a predetermined spatial barcode at the index value, the server device may determine that the spatial barcode for the sequence read is not within the set of predetermined or known spatial barcodes for the sample 102a. This may indicate that there is an error in the spatial barcode for the sequence read.

[0085] As shown in Fig. 4A, the set of predetermined spatial barcodes are 30 nucleotides long. However, the set of predetermined spatial barcodes may be any suitable length.Additionally, while the data structure 400 is shown as a hash table in Fig. 4A, the data structure may be any suitable data structure such as an array, a linked list, etc.

[0086] In any event, to correct errors in the spatial barcode, the spatial barcode may be compared to each of the predetermined spatial barcodes in the data structure 400 to identify the closest match, and one or more nucleobases in the spatial barcode may be corrected according to the closest matching predetermined spatial barcode. However, this process may require searching every predetermined spatial barcode in the data structure 400, O(n), which may include hundreds of millions of predetermined spatial barcodes. This can take minutes for each sequence read in the sample, requiring years to analyze the entire sample 102a.

[0087] Instead, the server device 120 may partition the nucleotide sequence or string for each predetermined spatial barcode into multiple predetermined nucleotide substrings. Then the server device stores each nucleotide substring in a separate data structure, as shown in Fig.4B. As shown in Fig. 4B, the server device 120 stores two 15 nucleotide hash tables 450, 460 for the first half (nucleotide positions 1-15) and the second half (nucleotide positions 16-30) of the predetermined spatial barcodes, respectively.

[0088] The ordered substructures 450, 460 may reference each other so that an entire predetermined spatial barcode may be obtained from the ordered substructures 450, 460. More specifically, each hash table 450, 460 may be indexed by computing the output of a hash function on the predetermined nucleotide substring. Then for a particular index, the first hash table 450 may store the first predetermined nucleotide substring corresponding to the index and indexes for retrieving second predetermined nucleotide substrings in the second hash table 460 that combine with the first predetermined nucleotide substring to make up the predetermined nucleotide string for a predetermined spatial barcode. For a particular index, the second hash table 460 may store the second predetermined nucleotide substring corresponding to the index33080 / IP-2937 and indexes for retrieving first predetermined nucleotide substrings in the first hash table 450 that combine with the second predetermined nucleotide substring to make up the predetermined nucleotide string for a predetermined spatial barcode.

[0089] The ordered substructures 450, 460 may indicate the nucleotide positions of the predetermined nucleotide substrings stored in the ordered substructures 450, 460 so the predetermined nucleotide substrings in the ordered substructures 450, 460 can be reconstructed into a predetermined spatial barcode. For example, ordered substructure 450 may be associated with nucleotide positions 1-15 and ordered substructure 460 may be associated with nucleotide positions 16-30 of the predetermined spatial barcodes. In another example, ordered substructure 450 may be associated with nucleotide positions 1 , 3, 5, 7, 9, ...29 and ordered substructure 460 may be associated with nucleotide positions 2, 4, 6, 8, 10, ...30 of the predetermined spatial barcodes.

[0090] While Fig. 4B illustrates two ordered substructures 450, 460, this is merely one example for ease of illustration only. The set of predetermined spatial barcodes may be partitioned into any suitable number of nucleotide substrings stored in a corresponding number of ordered substructures that reference each other. Each additional partition, k, increases the base mismatch tolerance, k-1 , by one. For example, if the nucleotide string is partitioned into two partitions, the base mismatch tolerance is one. If the nucleotide string is partitioned into three partitions, the base mismatch tolerance is two, etc.

[0091] Fig. 5A depicts an example process for base mismatch correction for a spatial barcode 502 for a sequence read using ordered substructures 510, 520 that store respective predetermined nucleotide substrings of a predetermined spatial barcode. The ordered substructures 510, 520 include a first hash table 510 that stores first halves (nucleotides 1-15) of predetermined spatial barcodes and a second hash table 520 that stores second halves (nucleotides 16-30) of predetermined spatial barcodes. As in Fig. 4B, the values in the hash tables 510, 520 may also include indexes for retrieving the other halves of predetermined spatial barcodes (Table 2 Reference / Table 1 Reference).

[0092] To perform the base mismatch correction for the spatial barcode 502, the server device 120 may partition the spatial barcode into two nucleotide substrings 504a (CGGCGICCACCGAGA), 504b (TCTACACTAAGTGCG). Then the server device 120 computes a hash function for the first nucleotide substring 504a to generate a first index (244243221224141) and computes a hash function for the second nucleotide substrings 504b to generate a second index (323121231123424).33080 / IP-2937

[0093] The server device 120 determines whether there is a value at the first index in the first hash table 510 and determines whether there is a value at the second index in the second hash table 520. This only requires two operations, 0(2). If both hash tables 510, 520 return a value at the respective indexes, the server device 120 determines that the spatial barcode 502 matches a predetermined spatial barcode for the sample. Otherwise, if one of the hash tables 510, 520 returns a value and the other does not, the server device 120 may identify each of the predetermined nucleotide substrings in the other hash table linked to the matching nucleotide substring. This is likely to be a small number of predetermined nucleotide substrings (e.g., about 1-3 predetermined nucleotide substrings).

[0094] The server device 120 may perform approximate matching with the nucleotide substring that did not match and the linked predetermined nucleotide substrings in the other hash table to determine whether the nucleotide substring does not differ from one of the linked predetermined nucleotide substrings by more than a threshold number of nucleobases (e.g., one). If the nucleotide substring does not differ from a linked predetermined nucleotide substring by more than the threshold number of nucleobases, the server device 120 may correct the spatial barcode for the sequence read using the predetermined spatial barcode corresponding to the linked predetermined nucleotide substring.

[0095] In the example shown in Fig. 5A, the server device 120 retrieves the value in the first hash table 510 at index 244243221224141 corresponding to the first nucleotide substring 504a and does not return a value. The server device 120 retrieves the value in the second hash table 520 at index 323121231123424 corresponding to the second nucleotide substring 504b and returns the second nucleotide substring (TCTACACTAACTGCG). The second hash table 520 may also store indexes 244241221224141 and 222211133414332 with Index323121231123424 as Table 1 References to the first hash table 510.

[0096] As a result, the server device 120 compares the first nucleotide substring 504a to the predetermined nucleotide substrings at indexes 244241221224141 and 222211133414332 in the first hash table 510 to identify an approximate match. The first nucleotide substring 504a CGGCGTCCAGCGAGA only differs by the predetermined nucleotide substring (CGGCGACCACCGAGA) at index 244241221224141 by one nucleobase and is identified as an approximate match. Accordingly, the server device 120 corrects the spatial barcode 502 for the sequence read using the predetermined spatial barcode (CGGCGACCACCGAGATCTACACTAACTGCG).33080 / IP-2937

[0097] While in this example, the spatial barcode is partitioned into two partitions and compared to two ordered substructures, the spatial barcode may be partitioned into any suitable number of partitions and compared to a corresponding number of ordered substructures. If there is a match for at least one of the substrings of the spatial barcode for a sequence read in one of the ordered substructures, the server device 120 can filter out each of the predetermined spatial barcodes which do not include the matching substring. Then the server device 120 may compare the remaining substring(s) of the spatial barcode for the sequence read to the remaining substring(s) in the filtered set to identify an approximate match which does not differ from the spatial barcode for the sequence read by more than a threshold number of nucleobases (e.g., where the threshold number is one less (k-1) than the number of partitions (k)).

[0098] Each additional partition, k, increases the base mismatch tolerance, k-1, by one. For example, if the spatial barcode is partitioned into two partitions, the base mismatch tolerance is one. If the spatial barcode is partitioned into three partitions, the base mismatch tolerance is two, etc. In this manner, if there are k-1 errors in a spatial barcode partitioned into k partitions, at least one of the partitions will match with a nucleotide substring in the predetermined spatial barcodes, because there cannot be an error in each of the k partitions. Spatial barcodes with more errors than the base mismatch tolerance may not be corrected.

[0099] As a result, the base mismatch correction requires k lookup operations, where k is the number of partitions of the spatial barcode and m substring comparison operations, where m is the number of predetermined nucleotide substrings linked to a predetermined nucleotide substring that matches a nucleotide substring of the spatial barcode. The resulting total number of operations (k + m) is very small (on the order of 1 , 10, or 100 operations) compared to the total number of operations to compare the spatial barcode to each of the n predetermined spatial barcodes in the data structure, where n can be 750 million predetermined spatial barcodes.

[0100] Additionally, while the spatial barcode shown in Figs. 4B and 5 is partitioned in contiguous portions of the spatial barcode (a first half and a second half), the spatial barcode may be partitioned in non-contiguous portions of the spatial barcode. The spatial barcode may be partitioned into any suitable non-overlapping portions. For example, the spatial barcode may be partitioned according to the nucleotide positions of the nucleotide substring for the spatial barcode. More specifically, the spatial barcode may be partitioned into a first nucleotide substring of odd nucleotide positions (nucleotide positions 1, 3, 5, 7, 9, ... 29) and a second33080 / IP-2937 nucleotide substring of even nucleotide positions (nucleotide positions 2, 4, 6, 8, 10, ... 30) within the spatial barcode.

[0101] The use of contiguous vs non-contiguous partitioning could be based on a priori information about error profiles of first seq and second seq spatial barcodes. For instance, if homopolymer-like sequences are prone to sequence specific errors that can result in consecutive bases in error, it may be advantageous to use contiguous partitioning. In this instance, it may be more likely that for example, one half of the sequence matches a predetermined or known spatial barcode, while the error(s) is / are in the other half. This may make it a bit easier to detect low edit distances that could be corrected. The edit distance may be the minimum number of edits needed to transform a spatial barcode into the closest matching predetermined spatial barcode.

[0102] Fig. 5B illustrates an example data structure 550 for storing predetermined spatial barcodes and indications of the corresponding sub-samples to which they belong. The data structure 550 may also be a hash table. For each predetermined spatial barcode, the data structure 550 may include an indication of the corresponding sub-sample for the predetermined spatial barcode, such as a sub-sample identifier (ID). For example, spatial barcode AACCACCCACCATACCCCTTTTTATGCTAA may correspond to sub-sample 1, spatial barcode AACCACCCACCATACAACCGCCGGGTATAG may correspond to sub-sample 0, spatial barcode CTCCCTGATCGGGCCCAACCAAATCGTGTT may correspond to sub-sample 3, etc.

[0103] Fig. 6 illustrates a comparison of the number of sequence reads having a valid spatial barcode with and without base mismatch correction. More specifically, Fig. 6 includes two tables 600, 650, where the first table 600 indicates the number of sequence reads having a valid spatial barcode 602 (a spatial barcode matching one of the predetermined spatial barcodes for the sample) without base mismatch correction. This scenario is indicative of the baseline performance, where the sequence reads are processed in their raw form without any intervention to correct potential errors in the base matching process

[0104] The second table 650 indicates the number of sequence reads having a valid spatial barcode 652 with base mismatch correction for spatial barcodes having a mismatch in one nucleobase. This scenario represents an enhanced processing step where errors in base matching are identified and corrected, thereby potentially increasing the accuracy of matching sequence reads to their correct spatial barcodes.33080 / IP-2937

[0105] This comparison reveals a significant improvement in the number of sequence reads having a valid spatial barcode when base mismatch correction is applied. As shown in the first table 600, 62.6 percent of the sequence reads have a valid spatial barcode 602 without error correction. With base mismatch correction for one nucleobase mismatches, the second table 650 indicates that 71.9 percent of the sequence reads have a valid spatial barcode 652 after error correction. This results in an increase in accurately identifying the spatial barcodes for an additional 9.3 percent of the sequence reads.

[0106] Additionally, as mentioned above, by partitioning the spatial barcode into nucleotide substrings and using ordered substructures, the base mismatch correction process increases the number of sequence reads having a valid spatial barcode without significantly increasing the processing time needed for the analysis. Each spatial barcode for the sequence reads can be analyzed in just a few operations.

[0107] In any event, after base mismatch correction is applied, the sample split and merge engine 122 may demultiplex the sequence reads into sub-samples using the corrected spatial barcodes. In other implementations, the sequence reads are not demultiplexed and instead are provided for genomic analysis. After the sample split and merge engine 122 analyzes the subsamples or the sample, for example using a genomic analysis tool, and / or merges the results of the analysis to generate a set of results for the sample, the server device 120 may provide the results of the analysis to the client device 130. The client device 130 may present the results of the genomic analysis for the sample via a spatial transcriptomics application 132. The results may include transcripts (e.g., of genes and / or genetic variants) corresponding to a set of unique molecules identified for the sample and their respective locations. Additionally, the spatial transcriptomics application 132 may present a spatial transcriptomics map using the combined set of results for the sample via a user interface.

[0108] The pool of sequence reads with spatial barcodes that are not part of the set of matching or error corrected spatial barcodes can be designated as sequence reads with “unmapped spatial barcodes.” Depending on the size of this pool of sequence reads, the server device 120 may apply a random subsample split to the pool for secondary analysis (e.g., mapping, alignment, etc.). Then the server device 120 may subsequently merge and / or deduplicate the subsamples in the pool of sequence reads with unmapped spatial barcodes. The server device 120 may then perform quality control on this pool of sequence reads / UMIs for example, to determine a read distribution per spatial barcode or a UMI distribution per spatial barcode. The server device 120 may then determine whether these spatial barcodes likely33080 / IP-2937 correspond to actual clusters that were not detected in first seq based on the read or UMI distributions. Alternatively, further error correction can be applied using additional information for the sequence reads to utilize more of the transcripts in the sequence reads and estimate their spatial locations. For example, the additional information may include a spatial barcode edit distance to the closest matching predetermined / first seq spatial barcode, a payload match for the sequence read, a priori knowledge about expected genes for a given tissue sample, etc.

[0109] Fig. 7 illustrates example spatial transcriptomics maps for each of the sub-samples 710-740 and for the sample 750. The spatial transcriptomics maps 710-750 may include heat maps of the transcripts at various locations within the sub-samples and the sample according to the spatial barcodes (which may be corrected spatial barcodes). The spatial transcriptomics heat maps 710-750 may use colors to indicate the concentration of transcripts, with darker colors typically representing higher concentrations.

[0110] More specifically, the spatial transcriptomics maps 710-750 may include a first spatial transcriptomics map 710 indicating the transcripts at various locations within sub-sample 1a, a second spatial transcriptomics map 720 indicating the transcripts at various locations within subsample 1b, a third spatial transcriptomics map 730 indicating the transcripts at various locations within sub-sample 1c, and a fourth spatial transcriptomics map 740 indicating the transcripts at various locations within sub-sample 1d. Furthermore, the client device 130 may also present a merged spatial transcriptomics map 750 for the sample indicating the transcripts at various locations within the sample based on the combined set of results for the sample.

[0111] Fig. 8 is a flow diagram of an example method 800 for splitting and merging spatial transcriptomics data for a sample 102a on a substrate 140. The method 800 may be implemented by a server device 120, and more specifically, a sample split and merge engine 122.

[0112] At block 802, the server device 120 obtains a set of spatial transcriptomics data for a sample 102a on a substrate 140. The spatial transcriptomics data includes multiple sequence reads of nucleic acid molecules having a spatial barcode, a target nucleic acid (e.g., DNA, cDNA, mRNA, or RNA), a molecular identifier (e.g., a UMI), an index, and / or other known sequences. For example, the sequencing device 110 captures nucleotide sequences from the tissue sample 102a bound to oligonucleotide barcodes on the substrate 140.

[0113] At block 804, the server device 120 divides the sample 102a into sub-samples, each assigned a different subset of spatial barcodes. For example, the server device 120 may33080 / IP-2937 generate a tissue mask corresponding to the boundaries of the sample 102a and obtain the spatial barcodes within the boundaries of the tissue mask. The server device 120 may then randomly assign spatial barcodes within the tissue mask to one of the sub-samples, ensuring an even distribution across the sub-samples. The number of sub-samples may be selected based on the memory size of the set of spatial transcriptomics data, such that the memory sizes of the sub-samples do not exceed a threshold memory storage capacity.

[0114] At block 806, the server device 120 demultiplexes respective sequence reads into one of the sub-samples in accordance with the spatial barcodes for the respective sequence reads. Prior to demultiplexing, the server device 120 may perform base mismatch correction of at least one nucleobase for a spatial barcode in a sequence read to match the spatial barcode to a predetermined spatial barcode in a set of predetermined spatial barcodes for the substrate 140 and / or the sample 102a. The base mismatch correction process may include partitioning the spatial barcode into multiple nucleotide substrings and comparing these substrings to sets of predetermined nucleotide substrings for the set of predetermined spatial barcodes, as detailed in Figs. 4B and 5. The sequence reads for each of the sub-samples may be stored in a separate file (e.g., a Fastq file).

[0115] At block 808, the server device 120 analyzes the sub-samples, for example using a genomic analysis tool. In some implementations, the server device 120 or multiple server devices may analyze the sub-samples simultaneously. In other implementations, the server device 120 may analyze the sub-samples one at a time.

[0116] For respective sub-samples, the server device 120 de-duplicates sequence reads for the same molecule having the same spatial barcode, target nucleic acid, and / or UMI in the respective sub-sample to generate a set of unique molecules for the respective sub-sample. In some implementations, a unique molecule includes a single instance of a particular spatial barcode, target nucleic acid, and UMI. The particular spatial barcode, target nucleic acid, and UMI may be duplicated during amplification resulting in several sequence reads having the same molecule. Additionally, the server device 120 may identify genes or genetic variants and corresponding locations of these genes or genetic variants on the substrate 140 in the respective sub-sample.

[0117] Upon completion of the analysis, the results of the analysis of each of the subsamples are merged to generate a set of results for the sample (block 810). The set of results for the sample 102a may include a superset of the sets of unique molecules in each sub-33080 / IP-2937 sample. The set of results may also include transcripts (e.g., of gene and / or genetic variants) along with their respective locations within the sample 102a.

[0118] In some implementations, the server device 120 provides the set of results for the sample 102a to a client device 130 for display to a user. For example, the client device 130 may present a spatial transcriptomics map indicating the transcripts and their respective locations within the sample 102a.

[0119] Fig. 9 is a flow diagram of an example method 900 for base mismatch correction of spatial barcodes. The method 900 may be implemented by a server device 120, and more specifically, a sample split and merge engine 122.

[0120] At block 902, the sever device 120 obtains a spatial barcode within a sequence read of a sample 102a on a substrate 140. The server device 120 then performs base mismatch correction of at least one nucleobase for the spatial barcode to match a predetermined spatial barcode in a set of predetermined spatial barcodes for the substrate 140.

[0121] This correction process involves partitioning the spatial barcode into nucleotide substrings (block 904). For example, the server device 120 may partition the spatial barcode into two nucleotide substrings of equal length, three nucleotide substrings of equal length, or any suitable number of nucleotide substrings.

[0122] The server device 120 compares these nucleotide substrings for the spatial barcode to sets of nucleotide substrings for a set of predetermined or known spatial barcodes for the substrate 140 and / or sample 102a (block 906). This comparison aims to identify the predetermined spatial barcode having nucleotide substrings which do not differ from the nucleotide substrings for the spatial barcode by more than a threshold number of nucleobases. Upon identifying a predetermined spatial barcode which is an approximate match to the spatial barcode, the server device 120 generates a corrected spatial barcode for the sequence read using the predetermined spatial barcode (block 908).

[0123] To facilitate this correction process, the server device 120 stores the set of predetermined spatial barcodes for the substrate 140 and / or the sample 120a in a data structure. This data structure is partitioned into multiple ordered substructures, where each substructure includes a respective set of nucleotide substrings from nucleotide strings in the set of predetermined spatial barcodes. The respective sets of nucleotide substrings are indexed according to representations of the respective set of nucleotide substrings in the ordered substructures. For example, the ordered substructures may be hash tables with indexes that33080 / IP-2937 are determined by computing a hash function of the nucleotide substrings. In some implementations, the hash function may be converting the nucleobases in a nucleotide substring into a numerical value which is n digital long, where n is the number of nucleotides in the nucleotide substring (e.g., a base 5 value where each nucleobase corresponds to a different number and no nucleobase also corresponds to a number). The spatial barcode is divided into the same number of nucleotide substrings as the number of ordered substructures.

[0124] The server device 120 then performs lookup operations to determine whether index values for respective nucleotide substrings for the spatial barcode match index values in the respective ordered substructures. If a match is identified for at least one nucleotide substring for the spatial barcode but a match is not identified for at least one remaining nucleotide substring, the server device 120 may identify remaining predetermined nucleotide substrings in the other ordered substructure(s) linked to the matching nucleotide substring in a subset of the predetermined spatial barcodes.

[0125] The server device 120 proceeds to compare the one or more remaining substrings for the spatial barcode to the remaining substrings in the subset of the predetermined spatial barcodes. The server device 120 identifies an approximate match between a remaining substring in the spatial barcode and a remaining substring in the subset for an ordered substructure when the remaining substring in the spatial barcode does not differ from the remaining substring in the ordered substructure by more than a threshold number of nucleobases (e.g., one). Then the server device 120 generates the corrected spatial barcode by combining the matching nucleotide substring with the remaining substring(s) identified in the remaining substructure(s).

[0126] More specifically, the server device 120 may partition the set of predetermined spatial barcodes for the substrate 140 and / or sample 102a into two nucleotide substrings stored in two ordered substructures. The server device 120 may divide a spatial barcode of a sequence read into two nucleotide substrings and perform base mismatch correction for one mismatched nucleobase.

[0127] The spatial barcodes may be 30 nucleotides long and the two ordered substructures may include a first ordered substructure that stores the odd nucleotide positions of the predetermined spatial barcodes (nucleotide positions 1, 3, 5, 7, 9, ... 29) and a second ordered substructure that stores even nucleotide positions of the predetermined spatial barcodes (nucleotide positions 2, 4, 6, 8, 10, ... 30). The server device 120 may also divide the spatial barcode for the sequence read into odd nucleotide positions and even nucleotide positions.33080 / IP-2937

[0128] Then the server device 120 may compute a first index value for the odd nucleotide positions and a second index value for the even nucleotide positions. The server device 120 may perform a first lookup using the first index value in the first ordered substructure and a second lookup using the second index value in the second ordered substructure. If the server device 120 identifies a match in the first ordered substructure but does not identify a match in the second ordered substructure, the server device 120 may identify a subset of nucleotide substrings in the second ordered substructure that are linked to the matching nucleotide substring in the first ordered substructure. Then the server device 120 compares the even nucleotide positions of the spatial barcode to the identified subset to identify a nucleotide substring in the subset which differs from the even nucleotide positions of the spatial barcode by one nucleobase. The server device 120 corrects the even nucleotide positions of the spatial barcode with the identified nucleotide substring from the second ordered substructure.

[0129] In some implementations, the server device 120 demultiplexes the sequence read according to the corrected spatial barcode. Also in some implementations, the server device 120 corrects mismatches in multiple spatial barcodes for the sample 102a and presents a spatial transcriptomics map providing information regarding transcripts at various locations within the sample 102a in accordance with the corrected spatial barcodes.

[0130] Aspects of the techniques described in the present disclosure may include any of the following aspects, either alone or in combination:

[0131] 1. A method for splitting and merging spatial transcriptomics data for a sample on a substrate, the method comprising: obtaining, by one or more processors, a set of spatial transcriptomics data for a sample on a substrate, the set of spatial transcriptomics data including a plurality of sequence reads of nucleic acid molecules having a spatial barcode and a target nucleic acid; dividing, by the one or more processors, the sample into a plurality of subsamples each assigned a different subset of spatial barcodes: demultiplexing, by the one or more processors, respective sequence reads into one of the plurality of sub-samples in accordance with the spatial barcodes for the respective sequence reads; analyzing, by the one or more processors, the plurality of sub-samples; and merging, by the one or more processors, results of the analysis of each of the plurality of sub-samples to generate a set of results for the sample.

[0132] 2. The method of aspect 1 , wherein obtaining the set of spatial transcriptomics data for the sample includes: generating, by the one or more processors, a tissue mask33080 / IP-2937 corresponding to boundaries of the sample; and obtaining, by the one or more processors, the plurality of sequence reads having spatial barcodes within the tissue mask.

[0133] 3. The method of any of aspects 1-2, wherein dividing the sample into the plurality of sub-samples includes: randomly assigning, by the one or more processors, each of the spatial barcodes within the tissue mask to one of the plurality of sub-samples.

[0134] 4. The method of any of aspects 1-3, further comprising: selecting, by the one or more processors, a number of sub-samples based on a memory size of the set of spatial transcriptomics data, such that the memory sizes of the plurality of sub-samples do not exceed a threshold memory storage capacity.

[0135] 5. The method of any of aspects 1 -4, wherein analyzing the plurality of sub-samples includes: for respective sub-samples, de-duplicating, by the one or more processors, sequence reads for a same molecule having a same spatial barcode and target nucleic acid in the respective sub-sample to generate a set of unique molecules for the respective sub-sample, wherein the results of the analysis include a superset of the sets of unique molecules for the plurality of sub-samples.

[0136] 6. The method of any of aspects 1-5, wherein analyzing the plurality of sub-samples includes: for respective sub-samples, identifying, by the one or more processors, a plurality of genes or genetic variants and corresponding locations of the plurality of genes or genetic variants on the substrate in the respective sub-sample.

[0137] 7. The method of any of aspects 1 -6, wherein the plurality of sub-samples are analyzed simultaneously.

[0138] 8. The method of any of aspects 1-7, wherein demultiplexing respective sequence reads in accordance with the spatial barcodes includes: performing, by the one or more processors, base mismatch correction of at least one nucleobase for a spatial barcode in a sequence read to match the spatial barcode to a predetermined spatial barcode in a set of predetermined spatial barcodes for the substrate.

[0139] 9. The method of any of aspects 1-8, further comprising: presenting, by the one or more processors, a spatial transcriptomics map providing information regarding transcripts at a plurality of locations within the sample in accordance with the results of the analysis.

[0140] 10. The method of any of aspects 1-9, wherein the substrate includes an active area of approximately 750 mm2.33080 / IP-2937

[0141] 11. A computing device for splitting and merging spatial transcriptomics data for a sample on a substrate, the computing device comprising: one or more processors; and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon, that when executed by the one or more processors, cause the computing device to: perform the steps of any of aspects 1-10.

[0142] 12. A non-transitory computer-readable memory storing instructions thereon, that when executed by one or more processors, cause the one or more processors to: perform the steps of any of aspects 1-10.

[0143] 13. A method for base mismatch correction of spatial barcodes, the method comprising: obtaining, by one or more processors, a spatial barcode within a sequence read of a sample on a substrate; performing, by the one or more processors, base mismatch correction of at least one nucleobase for the spatial barcode to match a predetermined spatial barcode in a set of predetermined spatial barcodes for the substrate by: partitioning the spatial barcode into a plurality of nucleotide substrings, wherein the base mismatch correction has a base mismatch tolerance which increases with each partition of the spatial barcode; comparing the plurality of nucleotide substrings for the spatial barcode to a plurality of sets of nucleotide substrings for the set of predetermined spatial barcodes to identify the predetermined spatial barcode having nucleotide substrings which do not differ from the plurality of nucleotide substrings for the spatial barcode by more than a threshold number of nucleobases; and generating, by the one or more processors, a corrected spatial barcode for the sequence read using the predetermined spatial barcode.

[0144] 14. The method of aspect 13, further comprising: storing, by one or more processors, the set of predetermined spatial barcodes for the substrate in a data structure; partitioning, by the one or more processors, the data structure into a plurality of ordered substructures each including a respective set of nucleotide substrings from nucleotide strings in the set of predetermined spatial barcodes, wherein the respective sets of nucleotide substrings are indexed according to representations of the respective set of nucleotide substrings in the plurality of ordered substructures; dividing, by the one or more processors, the spatial barcode into a same number of nucleotide substrings as a number of the ordered substructures; performing, by the one or more processors, lookup operations to determine whether index values for respective nucleotide substrings for the spatial barcode match index values in the respective ordered substructures; in response to identifying a match of an index value for a nucleotide substring for the spatial barcode to an index value in a corresponding ordered33080 / IP-2937 substructure: identifying, by the one or more processors, remaining substrings in one or more other ordered substructures for a subset of the predetermined spatial barcodes having the matching index value in the corresponding ordered substructure; comparing, by the one or more processors, one or more remaining substrings for the spatial barcode to the remaining substrings in the subset of the predetermined spatial barcodes; and correcting, by the one more processors, a mismatch in the one or more remaining substrings of the spatial barcode by identifying an approximate match between the one or more remaining substrings and a remaining substring in the subset for the one or more other ordered substructures to generate the corrected spatial barcode.

[0145] 15. The method of any of aspects 13-14, wherein the set of predetermined spatial barcodes are partitioned into two ordered substructures, the spatial barcode of the sequence read is divided into two nucleotide substrings, and the spatial barcode of the sequence read includes one mismatch error.

[0146] 16. The method of any of aspects 13-15, wherein performing lookup operations includes: determining, by the one or more processors, a first index value based on odd positions within the spatial barcode; determining, by the one or more processors, a second index value based on even positions within the spatial barcode; performing, by the one or more processors, a first lookup using the first index value in a first ordered substructure; and performing, by the one or more processors, a second lookup using the second index value in a second ordered substructure.

[0147] 17. The method of any of aspects 13-16, further comprising: in response to identifying a match of the first index value in the first ordered substructure and not identifying a match of the second index value in the second ordered substructure: identifying, by the one or more processors, a subset of nucleotide substrings in the second ordered substructure for the subset of the predetermined spatial barcodes having the matching index value in the first ordered substructure; comparing, by the one or more processors, the even positions of the spatial barcode to the subset of nucleotide substrings in the second ordered substructure to identify a nucleotide substring in the subset which differs from the even positions of the spatial barcode by one nucleobase; and correcting, by the one or more processors, the even positions of the spatial barcode with the identified nucleotide substring.

[0148] 18. The method of any of aspects 13-17, wherein the respective sets of nucleotide substrings are indexed in the plurality of ordered substructures by converting each nucleobase33080 / IP-2937 in a particular nucleotide substring into a numerical value and using resulting numerical values for the particular nucleotide substring as the index value.

[0149] 19. The method of any of aspects 13-18, wherein the index value is a base five value which is n digits long, wherein n is a number of nucleotides in each nucleotide substring.

[0150] 20. The method of any of aspects 13-19, further comprising: demultiplexing, by the one or more processors, the sequence read according to the corrected spatial barcode.

[0151] 21 . The method of any of aspects 13-20, further comprising: correcting, by the one or more processors, mismatches in a plurality of spatial barcodes for the sample; and presenting, by the one more processors, a spatial transcriptomics map providing information regarding transcripts at a plurality of locations within the sample in accordance with the corrected spatial barcodes.

[0152] 22. A computing device for base mismatch correction of spatial barcodes, the computing device comprising: one or more processors; and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon, that when executed by the one or more processors, cause the computing device to: perform the steps of any of aspects 13-21 .

[0153] 23. A non-transitory computer-readable memory storing instructions thereon, that when executed by one or more processors, cause the one or more processors to: perform the steps of any of aspects 13-21.ADDITIONAL CONSIDERATIONS

[0154] Although the disclosure herein sets forth a detailed description of numerous different implementations, it should be understood that the legal scope of the description is defined by the words of the claims set forth at the end of this patent and equivalents. The detailed description is to be construed as exemplary only and does not describe every possible implementation since describing every possible implementation would be impractical.Numerous alternative implementations may be implemented, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.

[0155] The following additional considerations apply to the foregoing discussion. Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are33080 / IP-2937 illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.

[0156] Additionally, certain implementations are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In example implementations, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.

[0157] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example implementations, comprise processor-implemented modules.

[0158] Similarly, the methods or routines described herein may be at least partially processor implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example implementations, the processor or processors may be located in a single location, while in other implementations the processors may be distributed across a number of locations.

[0159] The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example implementations, the one or more processors or processor-implemented modules may be located in a single geographic location (e.g., within a home33080 / IP-2937 environment, an office environment, or a server farm). In other implementations, the one or more processors or processor-implemented modules may be distributed across a number of geographic locations.

[0160] This detailed description is to be construed as exemplary only and does not describe every possible implementation, as describing every possible implementation would be impractical, if not impossible. A person of ordinary skill in the art may implement numerous alternate implementations, using either current technology or technology developed after the filing date of this application.

[0161] Those of ordinary skill in the art will recognize that a wide variety of modifications, alterations, and combinations may be made with respect to the above described implementations without departing from the scope of the invention, and that such modifications, alterations, and combinations are to be viewed as being within the ambit of the inventive concept.

[0162] The patent claims at the end of this patent application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s). The systems and methods described herein are directed to an improvement to computer functionality and improve the functioning of conventional computers.

Claims

33080 / IP-2937CLAIMSWhat is claimed is:

1. A method for splitting and merging spatial transcriptomics data for a sample on a substrate, the method comprising:obtaining, by one or more processors, a set of spatial transcriptomics data for a sample on a substrate, the set of spatial transcriptomics data including a plurality of sequence reads of nucleic acid molecules having a spatial barcode and a target nucleic acid;dividing, by the one or more processors, the sample into a plurality of sub-samples each assigned a different subset of spatial barcodes:demultiplexing, by the one or more processors, respective sequence reads into one of the plurality of sub-samples in accordance with the spatial barcodes for the respective sequence reads;analyzing, by the one or more processors, the plurality of sub-samples; and merging, by the one or more processors, results of the analysis of each of the plurality of sub-samples to generate a set of results for the sample.

2. The method of claim 1 , wherein obtaining the set of spatial transcriptomics data for the sample includes:generating, by the one or more processors, a tissue mask corresponding to boundaries of the sample; andobtaining, by the one or more processors, the plurality of sequence reads having spatial barcodes within the tissue mask.

3. The method of claim 2, wherein dividing the sample into the plurality of subsamples includes:randomly assigning, by the one or more processors, each of the spatial barcodes within the tissue mask to one of the plurality of sub-samples.

4. The method of claim 3, further comprising:selecting, by the one or more processors, a number of sub-samples based on a memory size of the set of spatial transcriptomics data, such that the memory sizes of the plurality of subsamples do not exceed a threshold memory storage capacity.33080 / IP-2937 5. The method of claim 1 , wherein analyzing the plurality of sub-samples includes: for respective sub-samples, de-duplicating, by the one or more processors, sequence reads for a same molecule having a same spatial barcode and target nucleic acid in the respective sub-sample to generate a set of unique molecules for the respective sub-sample, wherein the results of the analysis include a superset of the sets of unique molecules for the plurality of sub-samples.

6. The method of claim 1 , wherein analyzing the plurality of sub-samples includes: for respective sub-samples, identifying, by the one or more processors, a plurality of genes or genetic variants and corresponding locations of the plurality of genes or genetic variants on the substrate in the respective sub-sample.

7. The method of claim 1 , wherein the plurality of sub-samples are analyzed simultaneously.

8. The method of claim 1 , wherein demultiplexing respective sequence reads in accordance with the spatial barcodes includes:performing, by the one or more processors, base mismatch correction of at least one nucleobase for a spatial barcode in a sequence read to match the spatial barcode to a predetermined spatial barcode in a set of predetermined spatial barcodes for the substrate.

9. The method of claim 1 , further comprising:presenting, by the one or more processors, a spatial transcriptomics map providing information regarding transcripts at a plurality of locations within the sample in accordance with the results of the analysis.

10. The method of claim 1 , wherein the substrate includes an active area of approximately 750 mm2.

11. A computing device for splitting and merging spatial transcriptomics data for a sample on a substrate, the computing device comprising:one or more processors; and33080 / IP-2937 a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon, that when executed by the one or more processors, cause the computing device to:perform the steps of any of claims 1-10.

12. A non-transitory computer-readable memory storing instructions thereon, that when executed by one or more processors, cause the one or more processors to:perform the steps of any of claims 1-10.

13. A method for base mismatch correction of spatial barcodes, the method comprising:obtaining, by one or more processors, a spatial barcode within a sequence read of a sample on a substrate;performing, by the one or more processors, base mismatch correction of at least one nucleobase for the spatial barcode to match a predetermined spatial barcode in a set of predetermined spatial barcodes for the substrate by:partitioning the spatial barcode into a plurality of nucleotide substrings, wherein the base mismatch correction has a base mismatch tolerance which increases with each partition of the spatial barcode;comparing the plurality of nucleotide substrings for the spatial barcode to a plurality of sets of nucleotide substrings for the set of predetermined spatial barcodes to identify the predetermined spatial barcode having nucleotide substrings which do not differ from the plurality of nucleotide substrings for the spatial barcode by more than a threshold number of nucleobases; andgenerating, by the one or more processors, a corrected spatial barcode for the sequence read using the predetermined spatial barcode.

14. The method of claim 13, further comprising:storing, by one or more processors, the set of predetermined spatial barcodes for the substrate in a data structure;partitioning, by the one or more processors, the data structure into a plurality of ordered substructures each including a respective set of nucleotide substrings from nucleotide strings in the set of predetermined spatial barcodes, wherein the respective sets of nucleotide substrings33080 / IP-2937 are indexed according to representations of the respective set of nucleotide substrings in the plurality of ordered substructures;dividing, by the one or more processors, the spatial barcode into a same number of nucleotide substrings as a number of the ordered substructures;performing, by the one or more processors, lookup operations to determine whether index values for respective nucleotide substrings for the spatial barcode match index values in the respective ordered substructures;in response to identifying a match of an index value for a nucleotide substring for the spatial barcode to an index value in a corresponding ordered substructure:identifying, by the one or more processors, remaining substrings in one or more other ordered substructures for a subset of the predetermined spatial barcodes having the matching index value in the corresponding ordered substructure;comparing, by the one or more processors, one or more remaining substrings for the spatial barcode to the remaining substrings in the subset of the predetermined spatial barcodes; andcorrecting, by the one more processors, a mismatch in the one or more remaining substrings of the spatial barcode by identifying an approximate match between the one or more remaining substrings and a remaining substring in the subset for the one or more other ordered substructures to generate the corrected spatial barcode.

15. The method of claim 14, wherein the set of predetermined spatial barcodes are partitioned into two ordered substructures, the spatial barcode of the sequence read is divided into two nucleotide substrings, and the spatial barcode of the sequence read includes one mismatch error.

16. The method of claim 15, wherein performing lookup operations includes: determining, by the one or more processors, a first index value based on odd positions within the spatial barcode;determining, by the one or more processors, a second index value based on even positions within the spatial barcode;performing, by the one or more processors, a first lookup using the first index value in a first ordered substructure; and33080 / IP-2937 performing, by the one or more processors, a second lookup using the second index value in a second ordered substructure.

17. The method of claim 16, further comprising:in response to identifying a match of the first index value in the first ordered substructure and not identifying a match of the second index value in the second ordered substructure: identifying, by the one or more processors, a subset of nucleotide substrings in the second ordered substructure for the subset of the predetermined spatial barcodes having the matching index value in the first ordered substructure;comparing, by the one or more processors, the even positions of the spatial barcode to the subset of nucleotide substrings in the second ordered substructure to identify a nucleotide substring in the subset which differs from the even positions of the spatial barcode by one nucleobase; andcorrecting, by the one or more processors, the even positions of the spatial barcode with the identified nucleotide substring.

18. The method of claim 14, wherein the respective sets of nucleotide substrings are indexed in the plurality of ordered substructures by converting each nucleobase in a particular nucleotide substring into a numerical value and using resulting numerical values for the particular nucleotide substring as the index value.

19. The method of claim 18, wherein the index value is a base five value which is n digits long, wherein n is a number of nucleotides in each nucleotide substring.

20. The method of claim 13, further comprising:demultiplexing, by the one or more processors, the sequence read according to the corrected spatial barcode.

21. The method of claim 13, further comprising:correcting, by the one or more processors, mismatches in a plurality of spatial barcodes for the sample; andpresenting, by the one more processors, a spatial transcriptomics map providing information regarding transcripts at a plurality of locations within the sample in accordance with the corrected spatial barcodes.33080 / IP-293722. A computing device for base mismatch correction of spatial barcodes, the computing device comprising:one or more processors; anda non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon, that when executed by the one or more processors, cause the computing device to:perform the steps of any of claims 13-21.

23. A non-transitory computer-readable memory storing instructions thereon, that when executed by one or more processors, cause the one or more processors to:perform the steps of any of claims 13-21.