Vertical integration and analysis of spatial multi-omics data

The method addresses limitations in existing spatial multi-omics data analysis by coregistering and granularity matching spatial readouts, enhancing pathway activity detection and enabling advanced downstream analyses.

JP2026506366APending Publication Date: 2026-02-24ASPECT ANALYTICS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025544725
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-02
Filing Date
2024-02-02
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing methods for analyzing spatial multi-omics data are limited by the need for multi-resolution registration, which can lead to local optima convergence, information loss, noise sensitivity, and overfitting, and are not suitable for integrating data from multiple tissue types or classes.

Method used

A method for analyzing biomolecular associations in vertically integrated spatial multi-omics data, involving coregistration and granularity matching of spatial readouts to create a common spatial reference and match data points across different analytical techniques, allowing for differential expression analysis across multiple tissue types.

Benefits of technology

Enables detection of more changes in pathway activity by incorporating multiple readouts, facilitating downstream analyses such as biomarker discovery, drug target identification, and precision medicine through improved spatial correlation analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026506366000001_ABST
    Figure 2026506366000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method for analyzing associations between biomolecules in vertically integrated spatial multi-omics data. The method integrates association analyses of matched readouts across spatial readout sets obtained within a specific tissue type and compares the results of such association analyses across different tissue types. By including multiple readouts (biomolecular classes), the method provides a systematic framework for analyzing relationships between biomolecular classes, enabling deeper biological understanding compared to methods limited to a single readout.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for the molecular analysis of biological samples. [Background technology]

[0002] Prade et al. (2020) describe a computational multimodal workflow that combines metabolite mass spectrometry with multiplexed immunohistochemistry (IHC). IHC images are co-registered pixel-by-pixel to the coordinates of the mass spectrum. After co-registration, the stained images are scaled to the measurement resolution, and the pixel-by-pixel color values ​​are used to define regions or perform pixel-precise analysis of metabolic correlations and heterogeneity.

[0003] Bergholt et al. (2018) described a correlative heterospectral lipidomics (HSL) method based on coregistration of Raman spectroscopy, desorption electrospray ionization mass spectrometry (DESI-MS), and immunofluorescence imaging. Pixel-by-pixel co-registration of Raman spectroscopy and DESI-MS imaging generated heterospectral maps that correlated the biomolecular structure and composition of myelin.

[0004] The methods disclosed in Prade et al. 2020 and Bergholt et al. 2018 are directed to the analysis of matched readouts in a single tissue type or a single class of readouts.

[0005] International Publication No. WO 2022 / 051546 discloses a method for identifying cross-modal features from two or more spatially resolved datasets. The method includes registering the two or more spatially resolved datasets to generate an aligned feature image containing the two or more spatially aligned datasets, and extracting cross-modal features from the aligned feature image. The image alignment and further analysis is an affine multi-resolution registration. Multi-resolution registration is known to be sensitive to the quality of the initial transformation estimate. If the initial estimate deviates significantly from the optimal transformation, the algorithm may converge to a local optimum rather than a global optimum. Furthermore, the multi-resolution approach increases the likelihood of information loss, noise sensitivity, and overfitting. Summary of the Invention

[0006] The present invention aims to overcome at least one of the above-mentioned drawbacks.

[0007] The present invention and its embodiments provide a means for overcoming one or more of the above-mentioned drawbacks. To this end, the present invention relates to a method for analyzing biomolecular associations in vertically integrated spatial multi-omics data, as set forth in claim 1. First, an initial spatial readout set is typically obtained from one or more biological tissue sections. Multiple sections are often necessary because the desired analytical technique cannot be used on the same section, for example, because the technique may destroy the tissue or interfere with other techniques to be used. In the typical case of using multiple sections, an appropriate approach is needed to create a common spatial reference between sections and determine whether readouts were performed on sufficiently comparable tissue regions. Depending on the application, this may be the same cell / cell type / cell state, the same functional tissue unit, the same gross anatomical structure, etc. Therefore, the method includes a coregistration step. Depending on the analytical technique selected, the spatial scale at which readouts are performed can vary significantly. This applies both to the physical distance between individual data points within a single spatial omics experiment and to the surface area covered by each data point, which can range from square nanometers to hundreds of square micrometers. Differences in the former spatial scale, when applicable, are typically addressed as part of the coregistration step. The method further includes a granularity matching step, which allows for the generation of fairly comparable data points across readouts with widely differing surface areas. These steps provide input for various downstream analyses, such as differential expression analysis of spatial relationships between matched readouts. By incorporating multiple readouts (i.e., classes of biomolecules), the method can detect more changes in pathway activity than other methods limited to a single readout.

[0008] Preferred embodiments of the device are set out in claims 2 to 14. A particularly preferred embodiment is the invention according to claim 8. Claim 8 further extends the method of claim 1 to analyse spatial relationships between matched readouts across multiple spatial readout sets. This allows for the detection of changes in relationships between analytes across tissue types or instances, thereby enabling a variety of downstream analyses (e.g., detection of differential pathways between tissues, or assessment of the response of one or more tissues to a treatment or drug). [Brief explanation of the drawings]

[0009] The following description of specific embodiments of the present invention is for illustrative purposes only and is not intended to limit the present disclosure, its scope, or uses. Corresponding reference numerals indicate the same or corresponding parts and features throughout the drawings.

[0010] [Figure 1] Shows a sub-method for generating matched datasets between spatial readouts. [Figure 2] Schematic showing how biomolecular associations can be analyzed from different spatial readout sets. [Figure 3] The data structure including matched readouts of ST data and MSI data is shown. [Figure 4] This shows the pairwise spatial correlation between matched readouts (ST and MSI), which is an example of the output of an association analysis. [Figure 5] An example of complex granularity matching is shown in Figure 1. A shows segmented cells, which are matched to the circular readout of a parallel assay in Figure 2. C shows the weighted readout contribution of each cell as a result of the match. DETAILED DESCRIPTION OF THE INVENTION

[0011] The present invention relates to a method for analyzing biomolecular associations in vertically integrated spatial multi-omics data. This method enables differential expression analysis of spatial relationships between matched readouts across multiple tissue types. Detecting changes in relationships between analytes across tissue types enables a variety of downstream analyses, such as detecting differentially expressed pathways between tissue types. By including multiple readouts (classes of biomolecules), the method can detect more changes in pathway activity compared to methods limited to a single readout (e.g., a combination of genes, proteins, and lipids is much more informative than proteins alone).

[0012] As used herein, the term "tissue" is intended to mean a biological sample, including but not limited to plant material, cell cultures, organoids, xenografts, animal / human material, etc.

[0013] As used herein, "vertical integration" refers to spatial multi-omics datasets that relate to the same physical sample (e.g., serial sections obtained from a single organ) and are analyzed together. As used herein, spatial multi-omics dataset and spatial readout set are interpreted to mean multiple spatial omics measurements and multiple spatial readouts, respectively.

[0014] Unless otherwise defined, all terms used in disclosing the present invention, including technical and scientific terms, have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. As a further guide, term definitions are included to facilitate a better understanding of the present invention.

[0015] As used herein, the following terms have the following meanings:

[0016] As used herein, "a," "an," and "the" refer to both the singular and the plural unless the context clearly dictates otherwise. By way of example, "a compartment" refers to one or more compartments.

[0017] As used herein, the terms "comprise," "comprising," "comprises," and "comprised of" are synonymous with "include," "including," "includes," or "contain," "containing," and "contains" and are inclusive or open-ended terms that identify the presence of what follows the term (e.g., components) but do not exclude the presence of additional, unrecited components, features, elements, materials, steps, etc.

[0018] Furthermore, the terms "first," "second," "third," and the like in the specification and claims are used to distinguish between similar elements and do not necessarily denote a sequential or chronological order unless otherwise specified. These terms are interchangeable under appropriate circumstances, and it should be understood that embodiments of the invention may function in operations other than those described or illustrated herein.

[0019] When numerical ranges are expressed by endpoints, all integers and fractions subsumed within the range, as well as the stated endpoints themselves, are intended to be included in the range.

[0020] The terms "one or more" or "at least one" (e.g., one or more elements or at least one element of a collection of elements) are self-explanatory, but for clarity, they also include any one of the elements, any two or more—e.g., three or more, four or more, five or more, six or more, seven or more, etc.—and even all of the elements.

[0021] Unless otherwise defined, all terms used in disclosing the present invention, including technical and scientific terms, have the meanings that are commonly understood by those skilled in the art to which the present invention belongs. For further guidance, definitions of terms are included in the description to better understand the teachings of the present invention. The terms or definitions used herein are provided solely to aid in the understanding of the present invention.

[0022] References throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the present invention. Thus, the appearance of the phrase "in one embodiment" or "in an embodiment" in multiple places throughout this specification do not necessarily refer to the same embodiment, but may. Moreover, as will be apparent to those skilled in the art, particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Also, even if some embodiments described herein include some features and not others included in other embodiments, combinations of features from different embodiments are within the scope of the present invention and form different embodiments, as will be understood by those skilled in the art. For example, in the following claims, any of the embodiments described in any of the claims may be used in any combination.

[0023] In a first aspect, the present invention provides a method for analyzing biomolecular associations in vertically integrated spatial multi-omics data, the method comprising the steps of: obtaining a first set of at least two biological sample spatial readouts from at least one biological sample using at least one analytical platform per readout; establishing a common spatial coordinate system between the spatial readouts by coregistration; generating a matched dataset by granular matching; performing an association analysis at least between the spatial readouts; After establishing a common spatial coordinate system, the data point granularity of all spatial measurements is matched, including the substeps of finding the coarsest data point, finding finer data points in other readouts that overlap each of the coarsest data points, and aggregating and associating each group of overlapping finer data points with its overlapping coarse data point. This method advantageously generates co-registered, spot-matched datasets, enabling the generation of matched datasets, in which each data point contains information from all complementary readouts (e.g., mRNA, protein, and lipid for each matched spot) that can reasonably be considered to have originated from a comparable tissue region. The matched dataset serves as the underlying data structure for many types of downstream analyses. In one embodiment, the step of obtaining a first set of at least two spatial readouts of the biological sample is performed using at least two different analytical platforms for each readout.

[0024] Granularity describes the spatial scale at which spatial readout is performed, both in terms of the physical distance between individual data points within a single spatial omics experiment and the surface area covered by each data point, which can range from the order of square nanometers to hundreds of square micrometers.

[0025] In this context, "data set" or "dataset" is understood to mean any multiple spatial readouts. Multiple spatial readouts within a single dataset are defined by a combination of at least four variables: sample, tissue type, tissue state (e.g., healthy or cancerous), and time. Any of these four variables may vary from readout to readout within the same dataset, or may be the same. For example, each readout within a dataset may be collected from a different tissue sample, and the tissue may be of the same type, ranging from healthy to advanced tumor, and collected at different time points.

[0026] In one embodiment, before performing the association analysis and after matching the data point granularity of the spatial readout sets, each matched spatial readout is stored in a data structure that is identical for all matched spatial readouts in the set, which has the advantage of maintaining the correspondence of each data point of each matched readout with data points of other readouts in the same set, which facilitates further processing of the matched readouts.

[0027] Association analysis is performed to estimate one or more measures of co-expression of the multiple readouts based on the matched spatial readouts. For each of the multiple readouts, the association analysis may generate one or more quantitative values, such as the spatial correlation of each gene-lipid pair (single value) for all genes and proteins in the matched dataset, or parameters of a statistical model fitted to the joint distribution of each gene-lipid pair (multiple values).

[0028] In one embodiment, the results of the association analysis are stored as at least one matrix, which allows values, such as point estimates, to be easily stored and more efficiently accessed. This type of data structure is also particularly well suited for data visualization and interpretation by a human operator.

[0029] In one embodiment, the results of an association analysis are stored as a tensor. This type of data structure has the advantage that more detailed information, such as quantiles and confidence intervals, can be easily and efficiently stored. This type of data structure is also particularly suitable for data visualization and interpretation by a human operator. Tensors are particularly useful when the association analysis produces multiple numerical values, such as correlation values ​​plus statistics on the correlation distribution (e.g., those obtained by bootstrap) and rank distribution tests (e.g., Mann-Whitney, Kolmogorov-Smirnov, etc.).

[0030] In one embodiment, all spatial readouts in the same spatial readout set are acquired from a single section of each biological sample, which has the advantage of facilitating the process of co-registering the readouts of the set.

[0031] In one embodiment, all spatial readouts in the same spatial readout set have the same surface area relative to the data points, which has the advantage that it facilitates the process of granularity matching of the readouts after spatially co-registering the sets.

[0032] In one embodiment, the association analysis data for each spatial readout set is further subjected to downstream analysis using the results of multiple previously performed association analyses, where these different association analyses correspond to, for example, different functional areas / healthy vs. tumor / treated vs. untreated, etc.

[0033] In one embodiment, spatial readouts within the same spatial readout set are obtained from at least two sections of each biological sample. In another embodiment, each spatial readout within the same spatial readout set is obtained from a different section of the biological sample. In any embodiment where multiple sections are required for each biological sample to obtain all spatial readouts in a single set, it is desirable that these sections be extracted as close to each other as possible, and most preferably adjacent.

[0034] In one embodiment, the method further comprises acquiring additional sets of spatial readouts and performing co-registration, granularity matching and association analysis for each set. Preferably, the analysis platform used to acquire the spatial readouts to obtain each additional set is the same as that used for all other spatial readout sets. This allows association analyses to be generated for each readout set, and advantageously these association analyses can be compared with each other.

[0035] Preferably, each association analysis for each set of spatial readouts is stored in the same data structure, which has the advantage that any further analysis steps can be automated, such as (but not limited to) performing association analysis between different sets of association analysis groups. This further analysis can be differential expression analysis, clustering, segmentation, etc.

[0036] In one embodiment, each additional set of spatial readouts corresponds to a different region of the sample. Furthermore, in another embodiment, each additional set of spatial readouts corresponds to a different time point. In further or alternative embodiments, each additional set of spatial readouts corresponds to a different sample. The ability to collect additional readout sets from different regions, different time points, and / or different samples greatly expands the biological understanding gained from the spatial readouts. This new understanding provides avenues for, but is not limited to, biomarker and drug target discovery, novel pathway identification, therapeutic improvement and patient stratification, precision medicine, and the like.

[0037] The present invention is further illustrated by the following non-limiting examples which further illustrate the invention, but which are not intended, and should not be construed, to limit the scope of the invention.

[0038] The invention will now be explained in more detail with reference to the following non-limiting examples.

[0039] In order to more clearly illustrate the features of the present invention, the following describes some preferred application examples of the method for analyzing biomolecular associations in vertically integrated spatial multi-omics data based on the present invention, but these are merely examples and do not in any way limit other potential applications.

[0040] FIG. 1 illustrates a sub-method (1) for generating matched datasets between spatial readouts. The illustrated embodiment of sub-method (1) begins by collecting a first spatial readout (S1) of a first spatial readout set and a second spatial readout (S2) of the first spatial readout set. After acquiring each spatial readout (S1, S2), each readout may undergo an additional pre-processing step (S9). The pre-processing step (S9) may include data cleanup, data conversion, readout annotation, etc. The pre-processing step (S9) for the first spatial readout set (S1) may be different from another pre-processing step (S9) for the second spatial readout set (S2). In the illustrated embodiment of sub-method (1), the pre-processing step (S9) for the first spatial readout includes obtaining information from a database (3). This information may include annotated readouts that serve as a comparison standard for the first spatial readout, enabling, for example, automated annotation of the first spatial readout (e.g., identification of MSI peaks by LC-MS) or information extension of the first spatial readout (e.g., augmenting spatial transcriptomics data with scRNA-seq). After the preprocessing step (S9), preprocessed spatial readout files (S10, S15) are generated for each spatial readout, and these preprocessed spatial readout files (S10, S15) are used as input for the coregistration step (S11). If one or more preprocessing steps (S9) are omitted, the spatial readouts for which preprocessing was omitted are themselves used as input for the coregistration step (S11) instead of the preprocessed spatial readout files. In this embodiment, the granularity matching step (S12) is performed after the coregistration step (S11) establishes a common coordinate system between the preprocessed spatial readout files (S10, S15). The granularity matching step (S12) generates a matched first spatial readout (S13) file and a matched second spatial readout (S16) file, which may undergo additional pre-processing (S14) before performing the association analysis (S3).Pre-processing of the matched readouts (S14) may include, for example, removal of batch effects and may be performed based on either the matched spatial readouts (S13, S16) and / or a combination of both (e.g., by using joint embedding / deep learning algorithms). After the association analysis (S3), an association analysis data structure (S4) is generated.

[0041] FIG. 2 illustrates a schematic diagram of a method (2) for analyzing biomolecular associations from different spatial readout sets. The method (2) involves processing each collected spatial readout set by executing a submethod that generates and analyzes matched datasets between spatial readouts. The diagram shows a first spatial readout set (S1, S2) and a second spatial readout set (S5, S6). In this example, S1 and S5 are equivalent spatial readouts in different tissue regions, as are S2 and S6. Each set is processed using the submethod to generate and analyze matched datasets between spatial readouts (1), resulting in association analysis data structures (S4, S17) for each readout set, which are then used for downstream analysis (S7). In the illustrated embodiment, the co-registration step (S11) and the granularity matching step (S12) are performed for all acquired spatial readouts (S1, S2, S5, S6) during execution of the sub-method (1) for each spatial readout set. In this way, the association analysis data structure (S4) for the first readout set and the association analysis data structure (S17) for the second readout set are already co-registered with each other and already match in granularity with each other.

[0042] To further illustrate the present invention, reference is made to the following examples, which are not intended to limit the invention to the embodiments shown in the following examples or figures.

[0043] Example 1 relates to a specific implementation of the submethod disclosed in Figure 1. This example demonstrates the integration of spatial transcriptomics and spatial lipidomics. In this example, mass spectrometry imaging (MSI) is used to perform lipid imaging on the first of two serial sections, separated vertically by 20 microns. This MSI image corresponds to the first spatial readout (S2). Spatial transcriptomics (ST) is used to obtain a second spatial readout (S1) from the second section. In this example, it was advantageous to set a very short vertical distance between the sections to minimize the probability that different cell types appear in corresponding regions of both sections.

[0044] The MSI and ST data from the first readout set (S1, S2) were preprocessed (S9) using a common method for each readout (S1, S2). Preprocessing (S9) for each readout (S1 and S2) included data normalization, filtering, feature extraction, and realignment (for MSI data). Furthermore, the ST data (S1) was deconvolved using information obtained from single-cell RNA sequencing performed on the bulk sample, resulting in preprocessed ST readouts (S10) that were expanded beyond the original ST readouts included in S1.

[0045] Coregistration (S11) of lipid imaging (S2) and spatial transcriptomics data (S1) was performed by coregistrating both modalities to a third hematoxylin-eosin (H&E)-stained microscopy image. The preprocessed lipid imaging (S15) and preprocessed spatial transcriptomics data (S10) were indirectly registered to each other by combining the MSI and microscopy image registration and the ST and microscopy image registration. This approach of coregistration of multiple sections via a single proxy section (S11) is unavoidable for applications involving more than two sections, in which case all spatial readouts are typically coregistrated to a readout near the middle of the vertical stack of sections.

[0046] Pathologists provided annotations of various tissue types on the microscopy data, including healthy tissue and various tumor grades. The pathologist-annotated regions were then co-registered with the microscopy image (S11) and transferred to the spatial readout. Each pathologist-annotated tissue type resulted in one spatial readout set (e.g., tumor grade A, which contains multiple spots / pixels in the MSI and ST data).

[0047] After coregistration (S11) of the MSI and ST data (using H&E microscopy images as a proxy), differences in the scale and size of the spatial readouts are corrected by granularity matching (S12). In this example, the MSI data were acquired with a spatial resolution (pixel size) of 30 microns, while the ST data were acquired on a Visium platform with a spot diameter of 55 microns. The granularity of both spatial readouts was matched (S12) by applying Gaussian smoothing to the MSI pixels close to each ST spot, and a representative aggregate lipid spectrum associated with each ST spot (i.e., matched spatial readouts for MSI and ST (S13 and S16)) was obtained. To illustrate an example of a data structure for storing and representing matched spatial readouts for MSI and ST (S13 and S16), Figure 3 shows a data structure containing matched readouts for ST (S13) and MSI data (S16). This data structure can be thought of as a data frame, with each row representing one matched spot in the spatial readout set and each column containing information about an individual feature (organized by readout, in this example, RNA transcripts and lipids).

[0048] Pairwise associations between matched spatial readouts (S3) were then determined by calculating the spatial correlation of gene and lipid expression for all genes and lipids contained in the matched spatial readouts (S13 and S16). This resulted in a correlation matrix consisting of correlation values ​​for each gene-lipid pair within each set of matched spatial readouts S13 and S16 (tissue type). To illustrate such a correlation matrix and the data it contains, a highly simplified example is shown in Figure 4, which shows the numerical spatial correlations for each pair of matched readouts (ST and MSI) (S13 and S16) as output of the association analysis (S4).

[0049] Finally, differential expression analysis was performed between tissue types—e.g., healthy and tumor tissues—to identify whether gene-lipid correlations significantly changed between tissue types (healthy vs. tumor). To this end, the microscopy data were annotated by a pathologist for various tissue types (including healthy tissues and multiple tumor grades). The pathologist-annotated regions were then transferred to spatial readouts after coregistration with the microscopy images (S11). Each pathologist-annotated tissue type generated a spatial readout set (e.g., tumor grade A, including multiple spots / pixels in the MSI and ST data). This enabled downstream analysis across different spatial readout sets (S7). Using this approach, we identified sets of correlation changes—gene-lipid pairs that were spatially co-expressed in healthy tissues but not in tumor tissues (or vice versa). Such understanding is biologically significant, strongly suggests changes in the underlying biological mechanisms, and provides a strong starting point for further research (e.g., pathway analysis starting from known genes and lipids).

[0050] Example 2 illustrates the application of the method of the present invention, where spatial readout sets are acquired at different time points after drug administration to a comparable model animal. Exemplary readout sets include mass spectrometry imaging (localization of drug metabolites) and spatial transcriptomics (analysis of transcriptome-level changes), which are typically performed on serial sections and have different spatial resolutions, requiring both coregistration and granularity matching (S11, S12). Spatial correlation can be used as an association analysis (S3) to identify gene expression associated with the presence / absence of a specific drug or its metabolites within a single spatial readout set (i.e., a single time point after administration). Downstream analysis (S7) is performed by analyzing multiple spatial readout sets (corresponding to different time points after administration) (S1, S2 and S5, S6) to determine the time-dependent effects of drug metabolites on transcriptomic expression.

[0051] An example application involves a single biopsy containing multiple tissue types. These tissue types are determined, for example, by annotation, where a pathologist assigns multiple tumor grades to specific regions within the tissue. Each annotated tissue type results in a spatial readout set derived from the same original biopsy. In this example, there are multiple tissue types (e.g., multiple cancer grades), resulting in an equal number of spatial readout sets. Exemplary readouts for such applications may include immunohistochemistry, lipid imaging by mass spectrometry, and spatial transcriptomics. Spatial correlation can be used in association analysis to identify co-expressed lipids and genes within each tissue type (spatial readout set). These lists of co-expressed lipids and genes can be compared across tissue types (spatial readout sets) to identify which tissue types are similar in comparable gene-lipid associations (e.g., tumor grades A and B are similar to each other, but different from grades C and D) using, for example, dimensionality reduction via linear or non-linear embedding (e.g., PCA, UMAP, NMF, t-SNE) or clustering methods (e.g., k-means, DBSCAN).

[0052] Many data analysis workflows do not use the original raw readouts (S1, S2, S5, S6), but instead use transformed readouts or other derived understandings derived from the raw readouts by preprocessing (S9) the raw readouts (S1, S2, S5, S6). One example of a derived understanding is using deep learning (e.g., convolutional neural networks) to create vector embeddings that capture tissue morphology starting from high-resolution microscopy images. In subsequent analyses (S10 and beyond), the vector embeddings can be used instead of or in conjunction with the pixel values ​​of the microscopy images.

[0053] Another type of transformed readout (S9) is a technique that extends the original spatial readout with additional information (3). A first example of such an extended readout is the annotation of mass spectrometry imaging peaks based on identifying information obtained from database searches or parallel experiments (e.g., liquid chromatography-mass spectrometry). A second example is the augmentation or deconvolution of spatial transcriptomics data with information obtained from bulk single-cell analyses such as scRNAseq, imputing gene information not measured in the spatial transcriptomics readout based on the cell type and its expression observed in the single-cell data. A third example is the use of anatomical atlases to augment the spatial readout with anatomical labels and additional information embedded in the atlas. This can be achieved by mapping tissue samples to corresponding tissue sections within such atlases and transferring the atlas labels to the tissue samples. This anatomical atlas can also be used as a common reference frame and intermediary for registering different tissue sections. For example, the Allen Mouse Brain Atlas (AMBA) can be used as a reference for aligning sections of the mouse brain measured with different modalities. AMBA includes anatomical labels for different regions of the mouse brain, as well as spatial omics measurements and data from various biomedical imaging modalities. This information can be integrated to extend the spatial readout during preprocessing (S9), where the atlas is used as an external database (3).

[0054] Granular matching (S12) can be performed in a variety of ways. A simple approach is to use spatial nearest neighbor techniques to identify individual data points that correspond well between readouts. A more advanced approach to granular matching (S12) is to meaningfully combine multiple data points from the finest spatial resolution readout to generate a representative aggregate readout that corresponds to the coarsest spatial resolution readout, as shown in Example 1.

[0055] Another method for granularity matching (S12) is to extract statistics of the matching spot areas (such as the mean or median expression level or expression density of the readouts) and use this to match readouts of different granularities.

[0056] A variety of techniques can be used for association analysis (S3). A simple example would be to calculate a point estimate of a suitable spatial association measure, such as a (spatial) correlation coefficient. This results in a tabular data structure (e.g., a matrix) containing point estimates of the statistic between readout pairs. Instead of using a single point estimate, it is also possible to estimate confidence intervals or other parameters for the correlation between readout pairs, resulting in a multidimensional tabular data structure (e.g., a tensor).

[0057] Often, association analysis is performed pairwise for individual features between two readouts (e.g., calculating the spatial correlation between one gene and one lipid when combining lipid imaging and spatial transcriptomics), but in other applications, association analysis (S3) may involve multiple features per readout (e.g., modeling the joint distribution of multiple genes and multiple lipids based on matched spatial readouts) and / or more than two readouts (e.g., spatial proteomics, lipidomics, transcriptomics).

[0058] Association analysis (S3) can also be performed by first performing dimensionality reduction techniques, such as principal component analysis or nonnegative matrix factorization, on each of the different modalities, either collectively across all samples or individually (preprocessing S14). This process extracts subsets of correlated analytes within each modality. These subsets are then correlated across modalities using methods such as canonical correlation analysis.

[0059] Association analysis (S3) can be performed by directly mapping analytes to biochemical pathways. Analytes across different modalities are clustered together using one or more biochemical pathways, such as pathways extracted from databases, pathways created from multiple different databases, or pathways based on other existing knowledge sources. The expression levels of each analyte within a pathway can then be used to directly distinguish biological responses between tissue regions within a single section or between tissue regions across different samples.

[0060] It is important to recognize that the final output of this invention will provide new insights for further downstream analysis. For example, by following the workflow detailed above to determine differential correlations between genes and lipids across tissue types (based on spatial transcriptomics and lipid imaging by mass spectrometry, respectively), a key insight is that it is possible to identify sets of genes and lipids that are co-expressed in healthy tissues but not in cancer tissues. These genes and lipids found to be differentially expressed in spatial correlations across different tissue types can be used as input for various forms of pathway analysis, potentially providing a deeper understanding of the underlying biological mechanisms and revealing biological pathways that may be affected in cancer tissues.

[0061] Another application example of the method of the present invention is return to normalcy analysis. In the pharmaceutical field, return to normalcy analysis refers to administering a drug compound to a diseased individual, and then evaluating whether the detected biochemical signal has shifted from the diseased state to the biochemical signal observed in a healthy state, thereby determining whether the drug has shown an effective effect.

[0062] Starting with a wild-type (WT) model, a disease model, and a disease model treated with a novel drug compound, samples are subjected to multimodal measurements, including a spatial transcriptomics dataset and a protein-centric dataset for each model. Then, spatial registration between two (or more) datasets is performed for each measured individual, followed by a granular matching step. The molecules contained in both datasets are correlated to construct an association analysis data structure in the form of a correlation matrix for each of the three models. The distance between the correlation matrices (or portions thereof) of the WT model, the disease model, and the disease model treated with a novel drug compound is then calculated. This allows us to evaluate whether the distance between the WT model and the drug-treated disease model is reduced compared to the distance between the disease model and the WT model.

[0063] In studies evaluating different drug doses and / or different drugs with the same objective, this approach can be used to assess the effects of different doses and / or different drugs. Association analysis data structures, such as correlation matrices, kernel matrices, tensors, or sets of multiple correlation matrices and / or sets of kernel matrices and / or sets of multiple tensors, may be stored in a database for a particular individual, tissue, or biological state (e.g., healthy, diseased). The resulting set of correlated molecules serves as a reference point for comparison with newly acquired data. This database can be used for several purposes, including quality control (i.e., whether newly measured healthy samples are comparable to previously measured samples), the normality analysis described above, and the detection of drug-induced responses.

[0064] Before calculating the association analysis data structure, the spot locations can be grouped based on cell segmentation or tissue segmentation algorithms. Similar segments can be grouped together based on specific markers (e.g., cell typing). An association analysis data structure is then calculated for each segment or group of segments. The resulting association analysis data structures can be compared to extract cell- or tissue-specific differential association pairs, allowing users to discover co-expression differences between cell types. Similarly, the resulting association analysis data structure can be stored in a database as a reference point for a certain state of a particular cell or tissue type, using a method similar to that described above.

[0065] The resulting association analysis data structure can be used to analyze cell or tissue state transitions (trajectories), i.e., to examine the process of transitioning from one cell or tissue state to another, by mapping small differences between the association analysis data structures of a cell or tissue segment as intermediate states between the two cell or tissue states.

[0066] Association analysis data structures can be computed on a per-spot or per-cell basis. These association analysis data structures can be clustered using clustering algorithms (DBSCAN, Leiden clustering, Louvain clustering) with or without preprocessing using dimensionality reduction (UMAP, t-SNE, etc.). The resulting clusters can be used to group multiple spots in spatial data (e.g., cell types, tissue regions).

[0067] Not all molecules can be easily targeted by drug compounds; therefore, proteins and RNA molecules are often used as drug compounds. Proteins and RNA molecules provide suitable binding sites for drugs to interact with, resulting in, for example, protein inhibition. The set of correlated molecules extracted from the association analysis data structure can be used to find suitable drug targets starting from a specific molecule of interest. For example, if a lipid is found to be differentially expressed between healthy and diseased individuals, proteins correlated with that lipid can be extracted and used as potential drug targets.

[0068] In such a setting, users use this pipeline to calculate association analysis data structures for each organism and / or state. By performing differential analysis on the resulting association analysis data structures, they can extract association pairs that are differentially expressed between different states. This differential analysis can also calculate differences across multiple technical and / or biological replicate experiments to obtain statistically robust results. This differential analysis yields molecules of interest that may be involved in the underlying pathomechanisms of the disease under study. Furthermore, by examining sets of molecules that correlate with the differentially expressed association pairs, they can extract related potential drug targets.

[0069] Furthermore, molecules of interest and their correlates can be mapped to known biological pathways, which can then be used to identify other potential drug targets.

[0070] As a further application, our multi-omics pipeline can be used for biomarker discovery, i.e., the discovery of molecules that indicate specific disease states. Once a biomarker is identified, stains and antibodies can be developed to easily detect it in diagnostic and other settings. However, not all molecules can be easily targeted with stains or antibodies, and high specificity may not be achieved. Proteins and RNA molecules are often easier to detect than lipids.

[0071] Finding such suitable detectable biomarkers can utilize a similar workflow to that described for searching for relevant drug targets.

[0072] By using the correlated molecule set extracted from the association analysis data structure, appropriate biomarkers can be extracted based on specific molecules of interest. For example, if a lipid shows differential expression between healthy and diseased individuals, proteins correlated with that lipid can be extracted as potential biomarkers.

[0073] In such a setting, users use this pipeline to calculate association analysis data structures for each organism and / or state. By performing differential analysis on the resulting association analysis data structures, they can extract association pairs that are differentially expressed between different states. This differential analysis can also calculate differences across multiple technical and / or biological replicate experiments to obtain statistically robust results. This differential analysis yields molecules of interest that may be involved in the underlying pathomechanisms of the disease under study. Furthermore, by examining sets of molecules that correlate with the differentially expressed association pairs, they can extract related potential drug targets.

[0074] Furthermore, molecules of interest and their correlates can be mapped to known biological pathways, which can then be used to identify other potential drug targets.

[0075] When performing granular matching, diverse strategies are required, as one may start with readouts with complex spatial footprints. It is not uncommon for individual readouts to exhibit such complex spatial footprints, such as segmented cells or regions obtained by GeoMx or laser capture microdissection. This highlights the fact that spatial smoothing, while a rudimentary form of spot matching, can generally make the procedure significantly more informative.

[0076] Figure 5 shows an example of complex granularity matching, where the readout of a segmented cell is matched to a circular capture area of ​​another assay (e.g., Visium). Depending on the application, various approaches can be used to perform complex matching. For example, one approach weights the readouts of multiple cells based on the overlap of the capture area with the circular spot. Another approach is to average the readouts of cells whose associated capture area is mostly contained within the circular spot. In this figure, A shows segmented cells, which are matched to the circular readout of a parallel assay (e.g., Visium) in B. C shows the weighted readout contributions of each cell in the matched result.

[0077] Designing an appropriate granularity matching procedure is a crucial step in generating a data structure that integrates (matches) spatial multi-omics data. The illustration in Figure 5 shows how granularity matching is performed from assays with finer spatial footprints to assays with coarser footprints, but it is also possible to devise a granularity matching procedure that works in the reverse direction.

[0078] Another approach to the granular matching process is to start with cell-segmented data, where cell segmentation (i.e., delineating cell boundaries based on spatial readouts and then dividing the image into individual cells based on those boundaries) is performed as a pre-processing / annotation step on one or more spatial readouts, after which association analysis is performed at the cell level on the resulting cell types, cell states, or aggregated analyses at the cell level.

[0079] Cells are the functional units and building blocks for spatial biology inference, and the state of each cell, including its cell type, structure, and molecular composition, plays a pivotal role in its functionality and interactions within larger biological systems. Association analysis across multiple spatial omics readouts allows for improved characterization of cell types and cell states, and even allows for the identification of novel cell types and cell states. The resulting (in this case, cell-based) association analysis data structure can be used to generate cell-level databases of healthy and diseased samples, identify disease states, screen for novel drug targets, and even evaluate normality analysis.

[0080] Understanding the interactions between different cells within a region of interest is crucial for elucidating the biological functions of individual organisms and deepening our understanding of disease pathogenesis. Therefore, analyzing and comparing the cellular composition, cell type density, and cell type abundance within spatially localized regions (commonly referred to as "neighborhood analysis") is often incorporated into spatial omics data analysis based on individual omics modalities. In the presented workflow, neighborhood analysis algorithms can be utilized as a form of association analysis, which is enhanced by combined spatial readouts across different omics modalities. Furthermore, incorporating additional spatial omics readouts can aid cell deconvolution, i.e., improve the determination of the proportions of cell types and cell states present within a measured region of interest. [Explanation of symbols]

[0081] 1. Sub-methods for generating and analyzing matched datasets between spatial readouts 2. Methods for analyzing biomolecular associations from different spatial readout sets 3 Database S1 First Spatial Readout Set S2 First Second Spatial Readout Set S3 Association Analysis S4 Association Analysis Data Structure S5 Second Primary Spatial Readout Set S6 Second spatial readout set S7 Downstream data analysis S8 Downstream Data Analysis Data Structure S9 Pretreatment S10 Preprocessed Spatial Readout S11 Coregistration S12 Granularity Matching S13 Matched Spatial Readout S14 Preprocessing of matched spatial readout S15 Preprocessed spatial readout S16 Matched Spatial Readout S17 Association analysis data structure for the second readout set

[0082] The present invention is not limited to the illustrated embodiments or disclosed examples, but rather the method according to the present invention can be embodied in various forms without departing from the scope of the present invention.

Claims

1. 1. A method for analyzing biomolecular associations in vertically integrated spatial multi-omics data, the method comprising: obtaining a first set of at least two biological sample spatial readouts from at least one biological sample using at least one analytical platform for each readout; establishing a common spatial coordinate system between the spatial readouts by coregistration; generating a matched dataset by granular matching; performing an association analysis between at least the matched spatial readouts; wherein the data point granularity of all spatial measurements is matched after the step of establishing a common spatial coordinate system, the step comprising the substeps of (i) finding the coarsest-grained data point, (ii) finding finer-grained data points in other readouts that overlap with each of the coarsest-grained data points, and (iii) aggregating and associating each group of overlapping finer-grained data points with the coarsest-grained data point that overlaps them.

2. 2. The method of claim 1, wherein, prior to the step of performing the association analysis and after matching the data point granularity of the spatial readout sets, each matched spatial readout is stored in a data structure spatially referenced to each data structure of all other spatial readouts.

3. 3. The method according to claim 1, wherein the results of the association analysis are stored as at least one matrix.

4. The method according to any one of claims 1 to 3, characterized in that the results of the association analysis are stored as tensors.

5. The method according to any one of claims 1 to 4, characterized in that all spatial readouts in the same spatial readout set are obtained from a single section of at least one biological sample.

6. The method according to any one of claims 1 to 4, characterized in that the spatial readouts in one and the same spatial readout set are obtained from at least two sections of at least one biological sample.

7. 7. The method of claim 6, wherein each spatial readout in the same spatial readout set is obtained from a different section of at least one biological sample.

8. 8. The method of claim 1, further comprising the steps of acquiring additional spatial readout sets and performing the steps of co-registration, granular matching and association analysis for each set.

9. 9. The method of claim 8, wherein the association analysis data for each spatial readout set is subjected to further downstream analysis.

10. 9. The method of claim 8, wherein the analysis platform used to obtain the spatial readouts for each of the additional sets is the same as the analysis platform used for all other spatial readout sets.

11. A method according to any one of claims 1 to 10, characterized in that all spatial readouts in the same spatial readout set have the same surface area relative to the data point.

12. A method according to any one of claims 8 to 11, characterized in that each of the additional spatial readout sets corresponds to a different region of the sample.

13. A method according to any one of claims 8 to 12, characterized in that each of the additional spatial readout sets corresponds to a different time point.

14. The method according to any one of claims 8 to 13, characterized in that the additional spatial readout sets correspond to different samples.