Predicting spatial transcriptomics from medical images
Patent Information
- Application Number
- PCT/US2026/016587
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-25
- Filing Date
- 2026-02-25
- Publication Date
- 2026-09-03
Smart Images

Figure US2026016587_03092026_PF_FP_ABST
Abstract
Description
PREDICTING SPATIAL TRANSCRIPTOMICS FROM MEDICAL IMAGESCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U. S. Provisional Patent Application No. 63 / 762,952, filed on February 25, 2025, and entitled AI-DRIVEN 3D SPATIAL TRANSCRIPTOMICS, which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] This disclosure relates to diagnostic systems and is specifically directed to a system for predicting spatial transcriptomics from medical images.BACKGROUND
[0003] Understanding intratumoral morphological and molecular heterogeneity in human tissue is critical for developing personalized treatments and predicting therapeutic responses. Spatially-resolved transcriptomics (ST) provides expression profiles for many genes at high spatial resolution on two-dimensional (2D) tissue sections. By analyzing ST with its associated high-resolution tissue morphology, researchers can holistically characterize intratumoral heterogeneity with multimodal views, investigate how changes in molecular profile influence underlying morphology, and vice versa.
[0004] The molecular and morphological traits captured within a 2D tissue section only represent a small fraction of the tissue volume and the patient.Therefore, increasing attention has recently been directed toward extending molecular characterization from within a single tissue section to many adjacent tissue sections or across a larger volume. Recent three-dimensional (3-D) pathology studies, fueled by substantial advances in high-resolution 3-D tissue imaging modalities, such as micro-computed tomography (micro-CT) or open¬ top light-sheet microscopy, showed that 3-D morphological characterization can lead to better patient prognostication or cancer biomarker discovery. Parallel efforts have been devoted to creating 3-D molecular atlases of tissue, either with in-situ sequencing or by registering serial sections of 2-D ST data meticulously obtained from a single tissue volume. While promising, in-situ approaches remain limited in terms of capture area and depth and require longprocessing times. Serial section-based approaches provide discontinuous coverage along the axial dimension of thick tissues. Such approaches are impractical for scaling to whole-volume transcriptomic profiling in terms of cost and effort, with up to several days of processing for a single clinical sample.SUMMARY
[0005] In one example, a system includes a processor, an output device, and at least one non-transitory computer readable medium, storing executable instructions executable by the processor. The executable instructions provide an imaging interface that receives a three-dimensional image of a region of tissue from an imaging system, a three-dimensional image encoder that extracts a lowdimensional embedding of at least a portion of the three-dimensional image as a set of token features, and a machine learning model, trained on pairs of images and spatial transcriptomic data representing a same region of interest, that predicts spatial transcriptomic data for the region of tissue based on the set of token features.
[0006] In another example, a method is provided for predicting spatial transcriptomic data for a region of tissue from a three-dimensional medical image. The three-dimensional image of a region of tissue is received from an imaging system. A low-dimensional embedding of at least a portion of the three-dimensional image is extracted as a set of token features. Spatial transcriptomic data for the region of tissue is predicted based on the set of token features at a machine learning model. The machine learning model is trained on pairs of images and spatial transcriptomic data representing a same region of interest.
[0007] In a further example, a method is provided for training a machine learning model to predict spatial transcriptomic data. The machine learning model is pretrained on a plurality of images of stained tissue samples with associated spatial transcriptomics data. Three-dimensional images of volume of tissues are obtained along with a set of two-dimensional images of stained tissue samples and associated spatial transcriptomics data from within each volume of tissue. The machine learning model is fine-tuned using the three-dimensional images and the set of two- dimensional images of stained tissue samples and associated spatial transcriptomics data.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 illustrates an example of a system for predicting spatial transcriptomics for a region of interest from a three-dimensional medical images;
[0009] FIG. 2 illustrates another example of a system for predicting spatial transcriptomics for a region of interest from a three-dimensional medical images;
[0010] FIG. 3 illustrates a system for training a system for predicting spatial transcriptomics from medical images;
[0011] FIG. 4 illustrates one example of a method for predicting spatial transcriptomic data for a region of tissue represented by a three-dimensional (3-D) image;
[0012] FIG. 5 illustrates one example of a method for training a system for predicting spatial transcriptomic data from three-dimensional (3-D) images; and
[0013] FIG. 6 is a schematic block diagram illustrating an exemplary system of hardware components capable of implementing examples of the systems and methods disclosed in FIGS. 1-5.DETAILED DESCRIPTION
[0014] A “three-dimensional image,” as described herein, refers to any image having at least two voxels in each of three-dimensions and is intended to refer to images of both three-dimensional tissue sections as well as topographically structured surfaces or “2.5-dimensional” tissue sections.
[0015] The systems and methods provided herein provide a deep learning model that enables 3-D spatial transcriptomics prediction for 3-D tissue images captured with high-resolution, non-destructive 3-D pathology modalities, which are anticipated to become more common as a complementary approach to serial tissue sectioning. Non-destructive imaging preserves tissues for downstream assays, thereby facilitating morphomolecular analyses. Upon modeling the link between 3-D tissue morphology and corresponding spatially resolved gene expression profiles in local 3- D regions (or patches), the system processes each 2-D section of the test volume, or volume of interest, and its neighboring sections together to provide 2-D spatial transcriptomic predictions for all sections. This stack of predicted 2-D spatial transcriptomic images constitutes a 3-D spatial transcriptomic prediction for thick tissue specimens, accommodating any tissue volume size.
[0016] FIG. 1 illustrates one example of a system 100 for predicting spatial transcriptomics from medical images. The system 100 includes a processor 102, an output device 104, and a non-transitory computer readable medium 110 storing executable instructions, executed by the processor 102. It will be appreciated that the executable instructions can be spread across multiple non-transitory computer readable media that are operatively connected via an appropriate data connection, such that the executable instructions can be executed by multiple processors. In particular, the system 100 can be trained across multiple graphics processing units, with a number of graphics processing units and storage at the non-transitory computer readable medium 110 used for a given application be scalable with the application and the amount of available training data.
[0017] The executable instructions stored on the non-transitory computer readable medium 110 include an imager interface 112 that receives a three- dimensional image from an associated imaging system (not shown) and conditions the three-dimensional image for analysis. In one example, the imager interface 112 can apply one or more of normalization, clipping, and filtering to the image. In one example, the three-dimensional image can be acquired at a micro-computed tomography scanner or an open-top light-sheet microscopy imager. A three-dimensional image encoder 114 converts the received three-dimensional image into a lower dimensionality representation of the image as a set of tokens features. In one example, the three-dimensional image encoder 114 can be implemented as an artificial neural network that is trained on a corpus of training images, such as a convolutional neural network or an autoencoder. In one example, the three- dimensional image encoder 114 can comprise a transformer encoder, trained on a corpus of histology images, that operates directly on patches of the image to assign one or more categorical parameters to the image.
[0018] In one example, the three-dimensional image encoder 114 can be implemented using a transformer-based architecture pretrained on a set of images. Instead of treating each 3-D patch as a volume, the three-dimensional image encoder 114 processes the patch as a stack of 2-D patches. In one example, the three-dimensional image encoder 114 extracts a set of 2-D patch token features for every 2-D section of the 3-D patch. A depth-specific learnable embedding is then added to each set of token features. In one example, the same learnable embedding is added to all the token features in the same depth, without additional 2-D positional embeddings. Subsequently, the sets are merged to form a larger set of patch token embeddings. In another example, the three-dimensional image encoder 114 can be implemented using a foundation model similar to that found in Towards a general-purpose foundation model for computational pathology by Chen et al., Nature Medicine (March 2024), which is hereby incorporated by reference. In one example, the three-dimensional image encoder 114 can further include one or more attention poolers that facilitate the encoding of interactions between the token features set and a set of learnable embeddings. In this example, each embedding is then projected to a lower dimension through a linear layer to provide the token features.
[0019] A machine learning model 116 predicts spatial transcriptomic data for the region of tissue based on the set of token features representing the three-dimensional image. The machine learning model 116 is trained on pairs of images and spatial transcriptomic data representing a same region of interest. In one implementation, the machine learning model is implemented as a transformer with a single layer followed by a fully-connected layer that provides an output representing the spatial transcriptomic data from the set of token features. The resulting spatial transcriptomic data for the region of tissue represented by the three-dimensional image can then be provided to the user at the input device.
[0020] FIG. 2 illustrates another example of a system 200 for predicting spatial transcriptomics from medical images. The system 200 includes a processor 202, a three-dimensional imaging system 204, a display 206, and a non-transitory computer readable medium 210 storing executable instructions, executed by the processor 202. In one implementation, the three-dimensional imaging system 204 can be a micro-computed tomography scanner or an open-top light-sheet microscopy imager, and the three-dimensional image can be an image of a sample of tissue or a topographically structured surface. It will be appreciated that the executable instructions can be spread across multiple non-transitory computer readable media that are operatively connected via an appropriate data connection, such that the executable instructions can be executed by multiple processors. In particular, the system 200 can be trained across multiple graphics processing units, with a number of graphics processing units and storage at the non-transitory computer readable medium 210 used for a given application be scalable with the application and the amount of available training data.
[0021] The executable instructions stored on the non-transitory computer readable medium 210 include an imager interface 212 that receives a three- dimensional image from the imaging system (not shown) and conditions the three- dimensional image for analysis. A 3-D image encoder 220 can be implemented with a transformer-based architecture pretrained on a set of images. In the illustrated example, the 3-D image encoder first extracts a set of 2-D patch token features for every 2-D section of the 3-D patch, and a depth-specific learnable embedding is then added to each set of token features. Subsequently, the sets are merged to form a larger set of patch token embeddings. The 3-D image encoder 220 includes a set of attentional poolers 222 and 224. Each attentional pooler 324 and 332 can be implemented as a single-layer transformer, which facilitates the encoding of interactions between the token features set and a set of learnable embeddings and a linear layer that projects each embedding into a lower dimensionality representation. The encoded embeddings are then used for subsequent downstream tasks. A contrastive attentional pooler 222 provides cross-modal alignment with contrastive learning and, for each image patch, uses a single embedding to provide a global representation of the patch. A reconstruction attentional pooler 224 uses multiple embeddings to capture more localized and fine-grained image details for spatial transcriptomic prediction.
[0022] A machine learning model 230 predicts spatial transcriptomic data for the region of tissue based on the set of token features. In one example, the machine learning model 230 is implemented as a transformer with a fully-connected output layer. The machine learning model 230 is trained on pairs of images and spatial transcriptomic data representing a same region of interest. In one example, a subset of the pairs of images and spatial transcriptomics data each represent a two-dimensional region of tissue, with each of the images being acquired from a stained tissue section. Additionally or alternatively, a subset of the pairs of images and spatial transcriptomics data can represent each represent a volume of tissue, with each of the images being a three-dimensional image and the spatial transcription data representing one or more two-dimensional sections within the volume of tissue. For example, the machine learning model is pretrained on a set of pairs of two- dimensional stained images and spatial transcriptomic data representing a same region of interest and fine-tuned on a set of pairs of three-dimensional stained images and spatial transcriptomic data representing a same region of interest.
[0023] The output of the machine learning model 230 can be provided to a user at a display. In one example, the machine learning model output can represent a predicted spatial transcriptomic data for the region of tissue, for example, in the form of an image. Alternatively, the machine learning model 230 can be trained or prompted to respond to a provided 3-D image as a query and return other 3-D images from the training set that have a similar spatial transcriptomic profile to that predicted for the received 3-D image.
[0024] FIG. 3 illustrates a system 300 for training a system for predicting spatial transcriptomics from medical images. The illustrated system 300 aligns spatial transcriptomics to a set of corresponding image modalities, specifically two- dimensional (2-D) histologically stained image patches and three-dimensional (3-D) micro-computed tomography (micro-CT) patches, to allowing the trained system to perform cross-modal retrieval tasks in addition to spatial transcriptomics prediction, providing a flexible framework for diverse tasks. The system 300 includes a processor 302 and a non-transitory computer readable medium 310 storing instructions executable by the processor 302. It will be appreciated that the executable instructions can be spread across multiple non-transitory computer readable media that are operatively connected via an appropriate data connection, such that the executable instructions can be executed by multiple processors. In addition, the processor 302 can be implemented as multiple graphics processing units, with a number of graphics processing units and storage at the non-transitory computer readable medium 310 used for a given application be scalable with the application and the amount of available training data.
[0025] The stored instructions include four primary functional components, a 2-D image encoder 320, a 3-D image encoder 330, a transcriptomics encoder 340, and a machine learning model 350. The 2-D image encoder 320 can be implemented using a transformer model pretrained on histology regions with diverse types and stains, including frozen tissue, formalin-fixed paraffin embedded (FFPE) tissue, and immunohistochemistry, yielding image features robust to different tissue processing protocols across data sources. In one implementation, the image encoder uses a set of 196 (e.g., 14 x 14) patch token embeddings, each of which has a dimensionality of 768. This provides additional flexibility in using image encoder output embeddings for different downstream tasks. To address the use of image and transcriptomics encoders pretrained on diverse data sources, a lightweightmultilayer perceptron (MLP) 322 can be trained to encode a source or batch identification to distill biological variations while removing batch-associated variations during training through a domain adaptation loss. Upon training, the MLP module 322 can be discarded for downstream tasks.
[0026] In the illustrated implementation, the 3-D image encoder 330 is implemented with a transformer-based architecture pretrained on a set of images. In the illustrated example, the 3-D image encoder first extracts a set of 196 2-D patch token features for every 2-D section of the 3-D patch. A depth-specific learnable embedding is then added to each set of token features. The same learnable embedding is added to all the token features in the same depth, without additional 2-D positional embeddings. Subsequently, the sets are merged to form a larger set of patch token embeddings. For example, a 3-D patch with depth of twenty-one would result in 4,116 patch token features. The pretraining can utilize natural images, such as those in the ImageNet training set, as it has been determined that image encoders pretrained on natural images provide better transfer performance for microCT data compared to other radiology-specific image encoders, due to inherent texture and resolution differences between MRI / CT and microCT.
[0027] Each of the 2-D image encoder 320 and the 3-D image encoder 330 includes a set of attentional poolers 324 and 332. Each attentional pooler 324 and 332 can be implemented as a single-layer transformer, that facilitates the encoding of interactions between the token features set and a set of learnable embeddings, each of which has a dimensionality of 768, and a linear layer that projects each embedding into a lower dimensionality representation, for example, a 512- dimensional embedding. The encoded embeddings are then used for subsequent downstream tasks. A contrastive attentional pooler provides cross-modal alignment with contrastive learning and, for each image patch, uses a single embedding to provide a global representation of the patch. A reconstruction attentional pooler uses multiple embeddings (e.g., 32) to capture more localized and fine-grained image details for spatial transcriptomic prediction.
[0028] The transcriptomics encoder 340 encodes spatial transcriptomic data using a foundation model pretrained on transcriptomics data from millions of cells of various cancer types. The transcriptomics encoder 340 is configured to encode transcriptomics data from Visium and spatial transcriptomics spots, which typically contain about ten and twenty cells, respectively. In one example, the foundationmodel used in the transcriptomics encoder 340 can be pretrained on single cell data and then fine-tuned to encode transcriptomics data representing multiple cells. The transcriptomics encoder 340 features three key components: a gene-name encoder, an expression-value encoder, and a transformer encoder. In the illustrated implementation, the gene-name encoder 342 comprises an embedding layer that maps each gene to a fixed-length embedding vector, for example, a vector of dimension 512. The expression-value encoder 344 consists of two fully connected layers with rectified linear unit (ReLU) activation, which transform each gene expression value into a 512-dimensional vector. The output of the gene-name encoder and the expression-value encoder are then combined through element-wise addition, forming the input to the transformer encoder 346. In one implementation, the transformer coder 346 is implemented as a stack of transformer layers, each with eight attention heads. In this example, the token from the last transformer layer is fed into a single fully-connected layer for the transcriptomics embedding, with each transcription spot represented by a feature vector. In one example, each transcription spot is represented by a 512-dimensional vector.
[0029] In the illustrated example, the machine learning model 350 is implemented as a single layer followed by a single fully-connected layer, and trained such that, during operation, the machine learning model accepts an input from the 3-D encoder 330, specifically the output of the set of attentional poolers 334, representing a three- dimensional image and predicts corresponding spatial transcriptomics data for the tissue represented by the image. In the illustrated example, the system is trained over three stages designed to gradually build the capacity of 3-D spatial transcriptomic data prediction for a volume-of-interest (VOI). The first two stages utilize both 2-D and 3-D images of all the volumes except VOI in the same cancer cohort. If the 2-D spatial transcriptomic measurements from VOI are available, a third stage is performed to fine-tune the model. All three stages use loss functions designed to predict transcriptomics profiles from image embeddings while also aligning them with transcriptomics embeddings.
[0030] During the first stage, a pretraining stage, is trained on 2-D morphology and transcriptomics data, including transcriptomics expression data, 2-D morphology images from histology image patches centered at the location of each of a plurality of spatial transcriptomics spots associated with the transcriptomics expression data, and a source identifier for correcting for batch effects. The 2-D morphology imagesare encoded at the 2-D image encoder 320 and the transcriptomics expression data is encoded at the transcriptomics encoder 340. The contrastive and reconstruction attentional poolers 324 are randomly initialized and trained. The last three transformer layers from the 2-D image encoder 320 and the transcriptomic encoder 340 are also fine-tuned to provide task-specific embeddings. Data augmentation can be applied image patches, including horizontal flips, vertical flips, and color jittering.
[0031] This stage of training uses a combination of three loss functions: symmetric cross-modal contrastive learning objective, spatial transcriptomics reconstruction loss, and domain adaptation loss. The embedding spaces of the 2-D image encoder 320 and the transcriptomic encoder 340 are aligned using a symmetric cross-model contrastive learning objective. The loss function for the symmetric cross-modal contrastive learning objective includes a first term representing histology-to-gene loss and a second term representing gene-to-histology loss and aims to minimize the distance between paired embeddings while maximizing the distance between unpaired embeddings. The reconstruction loss function minimizes the error between the predicted gene expression and the ground truth spatial transcriptomic profiles. In one example, the loss function minimizes a mean squared error (MSE) between the smoothed ground truth gene expression and the predicted expression obtained from the histology image embeddings. Finally, the domain adaptation loss addresses potential batch effects by integrating spatial transcriptomics samples from multiple data sources, by training the multilayer perceptron 322 to infer the batch source identifier from the transcriptomic embedding and use a cross-entropy loss. As the aim is to make the model invariant to the batch attribute, the negative of the attribute prediction loss is backpropagated, making the system poor in predicting the data source. The total loss function is a weighted linear combination of these three loss functions. In one example, the contrastive and reconstruction losses are weighted substantially equally, and the weight applied to the domain adaptation loss is about one-tenth that of the other losses.
[0032] A second training stage focuses on further fine-tuning the system to capture the relationship between the morphology present in 3-D tissue imaging data and transcriptomics. Specifically, in this stage, the morphology of 3-D tissue image data is encoded with the 3-D image encoder 330 and this embedding is aligned to the corresponding 2-D histology images and spatial transcriptomic embeddings. The transcriptomics predictor 340 is fine-tuned such that the model can transition frompredicting spatial transcriptomics from 2-D images of stained tissue patches to predicting spatial transcriptomics from 3-D image patches. To preserve the morphology-transcriptomics embedding space from the first stage, the weights of the 2-D image encoder 320 and the transcriptomics encoder 340 are kept frozen. Data with paired 3-D images and spatial transcriptomics data can be more difficult to acquire than the corresponding 2-D training data. To account for the smaller size of the data set compared to the dataset used in the first stage, the 3-D image encoder 330 can also be frozen to prevent overfitting. Instead, the initialized contrastive and reconstruction attentional poolers 332 and 334 are trained.
[0033] The loss function used in the second training stage includes a symmetric cross-modal contrastive learning objective, a direct alignment loss, and a reconstruction loss. The symmetric cross-modal contrastive learning objective rewards alignment of the embedding space of the 3-D image encoder 330 to that formed between the 2-D image encoder 320 and transcriptomic encoder 340 using a dual symmetric cross-modal contrastive learning objective in a manner similar to the contrastive loss in the first stage. A second alignment loss minimizes the Euclidean distance between the 2-D image patch token embeddings and the 3-D image patch token embeddings from the reconstruction attentional pooler. Since the 3-D imaging modality presents different intensity, texture, and resolved structures compared to the 2-D stained images, the alignment loss minimizes the gap between different imaging modalities, allowing the system 300 to leverage the first pretraining stage based on the 2-D images. The reconstruction loss function minimizes the error between the predicted gene expression and the ground truth spatial transcriptomic profiles, for example, as a mean squared error (MSE) between the smoothed ground truth gene expression and the predicted expression. The total loss function is a weighted linear combination of these three loss functions. In one example, all three losses are weighted substantially equally. In another example, where the three- dimensional images are images of topographically structured surfaces, the contrastive and reconstruction losses can be substantially equally weighted, and the weight of the alignment loss can be zero or close to zero.
[0034] In a final stage, the system 300 is fine-tuned with sample-specific data, with all layers that were trainable during previous stages being trained. This includes the last three transformer layers of the 2-D image encoder 320, the last three transformer layers of the transcriptomics encoder 340, the trainable layers of the 3-Dimage encoder 330 from the previous stage, the contrastive and reconstruction attentional poolers 324 and 332 in both the 2-D image encoder 320 and 3-D image encoder 330, and the machine learning model 350. During this stage, the system is trained using the same direct loss and reconstruction loss as in the second training stage. A contrastive loss is defined as the sum of the symmetric contrastive losses from the first two stages, and a final loss function is determined as a weighted linear combination of these losses. Once trained, the 3-D image encoder 330 and the machine learning model are active during operation to either predict spatial transcriptomic data for a region of tissue represented by a 3-D image or to find 3-D images with spatial transcriptomic data similar to a received 3-D image.
[0035] In view of the foregoing structural and functional features described above in FIGS. 1-3, example methods will be better appreciated with reference to FIGS. 4 and 5. While, for purposes of simplicity of explanation, the methods of FIGS. 4 and 5 are shown and described as executing serially, it is to be understood and appreciated that the present invention is not limited by the illustrated order, as some actions could in other examples occur in different orders and / or concurrently from that shown and described herein.
[0036] FIG. 4 illustrates one example of a method 400 for predicting spatial transcriptomic data for a region of tissue represented by a three-dimensional (3-D) image. At 402, a three-dimensional image of a region of tissue is received from an imaging system. For example, the three-dimensional image of the region of tissue can be received as a micro-computed tomography image from a micro-computed tomography scanner. In another example, the three-dimensional image is received from an open-top light-sheet microscopy imager. At 404, a low-dimensional embedding of at least a portion of the three-dimensional image is extracted as a set of token features. In one example, the three-dimensional image is divided into a plurality of three-dimensional patches, and each three-dimensional patch is divided into a plurality of two-dimensional slices, and a low-dimensional embedding is extracted for each of the plurality of two-dimensional slices associated with each image patch.
[0037] At 406, spatial transcriptomic data for the region of tissue is predicted based on the set of token features at a machine learning model. The machine learning model is trained on pairs of images and spatial transcriptomic data representing a same region of interest. In one example, a subset of the pairs ofimages and spatial transcriptomics data each represent a two-dimensional region of tissue, with each of the images being acquired from a stained tissue section.Additionally or alternatively, a subset of the pairs of images and spatial transcriptomics data can each represent a volume of tissue, with each of the images being a three-dimensional image and the spatial transcription data representing one or more two-dimensional section within the volume of tissue. The output of the machine learning model can be provided to a user at an associated output device.
[0038] FIG. 5 illustrates one example of a method 500 for training a system for predicting spatial transcriptomic data from three-dimensional (3-D) images.
[0039] At 502, the machine learning model is pretrained on a plurality of images of stained tissue samples with associated spatial transcriptomics data. At 504, three-dimensional images of volumes of tissues are obtained. For example, the three-dimensional images can be obtained via a micro-CT imager or an open-top light-sheet microscopy imager. At 506, a set of two-dimensional images of stained tissue samples and associated spatial transcriptomics data are obtained from within each volume of tissue. At 508, the machine learning model is fine-tuned using the three-dimensional images and the set of two-dimensional images of stained tissue samples and associated spatial transcriptomics data.
[0040] In one example, the machine learning model is fine-tuned by training the machine learning model with the weights frozen on each of a two-dimensional image encoder associated with the set of two-dimensional images of stained tissue samples and a spatial transcriptomics encoder associated with the spatial transcriptomics data and then trained in concert with each of the two-dimensional image encoder, the spatial transcriptomics encoder associated with the spatial transcriptomics data, and a three-dimensional image encoder associated with the three-dimensional images. In another example, at least some weights are frozen on each of the two-dimensional image encoder, the spatial transcriptomics encoder, and the three-dimensional image encoder during the first training stage, while a set of weights associated with a reconstruction attentional pooler, which encodes interactions between tokens provided from the two-dimensional image encoder and the three-dimensional image encoder with a set of learned embeddings, are adjusted using a fitness function containing a term representing an alignment between a set of two-dimensional token embeddings and a set of three-dimensional token embeddings from the reconstruction attentional pooler. Additionally or alternatively,the machine learning model is fine-tuned using a fitness function incorporating a term representing a difference between a set of predicted spatial transcriptomics data for each three-dimensional image and the spatial transcriptomics data associated with the three-dimensional image.
[0041] FIG. 6 is a schematic block diagram illustrating an exemplary system 600 of hardware components capable of implementing examples of the systems and methods disclosed in FIGS. 1-5. The system 600 can include various systems and subsystems. The system 600 can be a personal computer, a laptop computer, a workstation, a computer system, an appliance, an application-specific integrated circuit (ASIC), a server, a server blade center, a server farm, etc.
[0042] The system 600 can includes a system bus 602, a processing unit 604, a system memory 606, memory devices 608 and 610, a communication interface 612 (e.g., a network interface), a communication link 614, a display 616 (e.g., a video screen), and an input device 618 (e.g., a keyboard and / or a mouse). The system bus 602 can be in communication with the processing unit 604 and the system memory 606. The additional memory devices 608 and 610, such as a hard disk drive, server, stand-alone database, or other non-volatile memory, can also be in communication with the system bus 602. The system bus 602 interconnects the processing unit 604, the memory devices 606-610, the communication interface 612, the display 616, and the input device 618. In some examples, the system bus 602 also interconnects an additional port (not shown), such as a universal serial bus (USB) port.
[0043] The processing unit 604 can be a computing device and can include an application-specific integrated circuit (ASIC). The processing unit 604 executes a set of instructions to implement the operations of examples disclosed herein. The processing unit can include a processing core.
[0044] The additional memory devices 606, 608, and 610 can store data, programs, instructions, database queries in text or compiled form, and any other information that can be needed to operate a computer. The memories 606, 608 and 610 can be implemented as computer-readable media (integrated or removable) such as a memory card, disk drive, compact disk (CD), or server accessible over a network. In certain examples, the memories 606, 608 and 610 can comprise text, images, video, and / or audio, portions of which can be available in formats comprehensible to human beings. Additionally or alternatively, the system 600 canaccess an external data source or query source through the communication interface 612, which can communicate with the system bus 602 and the communication link 614.
[0045] In operation, the system 600 can be used to implement one or more parts of a clinical decision support system in accordance with the present invention.Computer executable logic for implementing the clinical decision support system resides on one or more of the system memory 606, and the memory devices 608, 610 in accordance with certain examples. The processing unit 604 executes one or more computer executable instructions originating from the system memory 606 and the memory devices 608 and 610. The term "computer readable medium" as used herein refers to any medium that participates in providing instructions to the processing unit 604 for execution, and it will be appreciated that a computer readable medium can include multiple computer readable media each operatively connected to the processing unit.
[0046] Specific details are given in the above description to provide a thorough understanding of the embodiments. However, it is understood that the embodiments can be practiced without these specific details. For example, physical components can be shown in block diagrams in order not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques can be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0047] Implementation of the techniques, blocks, steps, and means described above can be done in various ways. For example, these techniques, blocks, steps, and means can be implemented in hardware, software, or a combination thereof. For a hardware implementation, the processing units can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro¬ controllers, microprocessors, other electronic units designed to perform the functions described above, and / or a combination thereof.
[0048] Also, it is noted that the embodiments can be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart can describe the operations as a sequential process, many of the operations can be performed in parallel orconcurrently. In addition, the order of the operations can be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to a return of the function to the calling function or the main function.
[0049] Furthermore, embodiments can be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting language, and / or microcode, the program code or code segments to perform the necessary tasks can be stored in a machine-readable medium such as a storage medium. A code segment or machine-executable instruction can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a script, a class, or any combination of instructions, data structures, and / or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, ticket passing, network transmission, etc.
[0050] For a firmware and / or software implementation, the methodologies can be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. Any machine-readable medium tangibly embodying instructions can be used in implementing the methodologies described herein. For example, software codes can be stored in a memory. Memory can be implemented within the processor or external to the processor. As used herein the term “memory” refers to any type of long term, short term, volatile, nonvolatile, or other storage medium and is not to be limited to any particular type of memory or number of memories, or type of media upon which memory is stored.
[0051] Moreover, as disclosed herein, the term "storage medium" can represent one or more memories for storing data, including read only memory (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage mediums, optical storage mediums, flash memory devices and / or other machine- readable mediums for storing information. The term "machine-readable medium"includes, but is not limited to portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage mediums capable of storing that contain or carry instruction(s) and / or data.
[0052] What have been described above are examples of the invention. It is, of course, not possible to describe every conceivable combination of components or methodologies, but one of ordinary skill in the art will recognize that many further combinations and permutations of the invention are possible. Accordingly, the invention is intended to embrace all such alterations, modifications, and variations that fall within the scope of the appended claims and the application. Additionally, where the disclosure or claims recite "a," "an," "a first," or "another" element, or the equivalent thereof, it should be interpreted to include one or more than one such element, neither requiring nor excluding two or more such elements. As used herein, the term “includes” means includes but not limited to, the term “including” means including but not limited to. The term “based on” means based at least in part on.
Claims
CLAIMSWhat is claimed is:
1. A system comprising:a processor;an output device; andat least one non-transitory computer readable medium, storing executable instructions executable by the at least one processor to provide:an imaging interface that receives a three-dimensional image of a region of tissue from an imaging system;a three-dimensional image encoder that extracts a low-dimensional embedding of at least a portion of the three-dimensional image as a set of token features; anda machine learning model, trained on pairs of images and spatial transcriptomic data representing a same region of interest, that predicts spatial transcriptomic data for the region of tissue based on the set of token features.
2. The system of claim 1, wherein the machine learning model is implemented as a transformer with a fully-connected output layer.
3. The system of claim 1, wherein the imaging interface divides the image into a plurality of image patches, the image encoder extracting a low-dimensional embedding for each of the plurality of image patches.
4. The system of claim 1, further comprising the imaging system, wherein the imaging system is one of a micro-computed tomography scanner or an open-top light-sheet microscopy imager.
5. The system of claim 1, the image encoder dividing the three-dimensional image into a plurality of three-dimensional patches, dividing each three-dimensional patch into a plurality of two-dimensional slices, and extracting a low-dimensional embedding for each of the plurality of image patches.
6. The system of claim 1, wherein a subset of the pairs of images and spatial transcriptomics data representing the same region of interest each represent a two- dimensional region of tissue, each of the images being acquired from a stained tissue section.
7. The system of claim 1, wherein a subset of the pairs of images and spatial transcriptomics data representing the same region of interest each represent a volume of tissue, each of the images being a three-dimensional image and the spatial transcription data representing one or more two-dimensional sections within the volume of tissue.
8. The system of claim 1, wherein the three-dimensional image represents as topographically structured surfaces.
9. The system of claim 1, wherein the machine learning model is pretrained on a set of pairs of two-dimensional stained images and spatial transcriptomic data representing a same region of interest and fine-tuned on a set of pairs of three- dimensional stained images and spatial transcriptomic data representing a same region of interest.
10. A method comprising:receiving a three-dimensional image of a region of tissue from an imaging system;extracting a low-dimensional embedding of at least a portion of the three- dimensional image as a set of token features; andpredicting spatial transcriptomic data for the region of tissue based on the set of token features at a machine learning model, the machine learning model being trained on pairs of images and spatial transcriptomic data representing a same region of interest.
11. The method of claim 10, further comprising dividing the three-dimensional image into a plurality of three-dimensional patches and dividing each three-dimensional patch into a plurality of two-dimensional slices, wherein extracting thelow-dimensional embedding comprises extracting a low-dimensional embedding for each of the plurality of two-dimensional slices associated with each image patch.
12. The method of claim 10, wherein receiving the three-dimensional image of the region of tissue comprises receiving the three-dimensional image from a open-top light-sheet microscopy imager.
13. The method of claim 10, wherein a subset of the pairs of images and spatial transcriptomics data representing the same region of interest each represent a two- dimensional region of tissue, each of the images being acquired from a stained tissue section.
14. The method of claim 10, wherein a subset of the pairs of images and spatial transcriptomics data representing the same region of interest each represent a volume of tissue, each of the images being a three-dimensional image and the spatial transcription data representing one or more two-dimensional section within the volume of tissue.
15. The method of claim 10, wherein receiving the three-dimensional image of the region of tissue comprises receiving a micro-computed tomography image from a micro-computed tomography scanner.
16. A method for training a machine learning model to predict spatial transcriptomic data, the method comprising:pretraining the machine learning model on a plurality of images of stained tissue samples with associated spatial transcriptomics data;obtaining three-dimensional images of volumes of tissues;obtaining a set of two-dimensional images of stained tissue samples and associated spatial transcriptomics data from within each volume of tissue; and fine-tuning the machine learning model using the three-dimensional images and the set of two-dimensional images of stained tissue samples and associated spatial transcriptomics data.
17. The method of claim 16, where fine-tuning the machine learning model using the three-dimensional images and the set of two-dimensional images of stained tissue samples and associated spatial transcriptomics data comprising:training the machine learning model with the weights frozen on each of a two- dimensional image encoder associated with the set of two-dimensional images of stained tissue samples and a spatial transcriptomics encoder associated with the spatial transcriptomics data; andtraining the machine learning model in concert with each of the two-dimensional image encoder, the spatial transcriptomics encoder associated with the spatial transcriptomics data, and a three-dimensional image encoder associated with the three-dimensional images.
18. The method of claim 17, wherein training the machine learning model with the weights frozen on each of the two-dimensional image encoder and the spatial transcriptomics encoder further comprises training the machine learning model with the weights frozen on each of the two-dimensional image encoder, the spatial transcriptomics encoder, and the three-dimensional image encoder.
19. The method of claim 17, wherein training the machine learning model with the weights frozen on each of the two-dimensional image encoder and the spatial transcriptomics encoder further comprises adjusting a set of weights associated with a reconstruction attentional pooler, which encodes interactions between tokens provided from the two-dimensional image encoder and the three-dimensional image encoder with a set of learned embeddings, using a fitness function containing a term representing an alignment between a set of two-dimensional token embeddings and a set of three-dimensional token embeddings from the reconstruction attentional pooler.
20. The method of claim 16, where fine-tuning the machine learning model using the three-dimensional images and the set of two-dimensional images of stained tissue samples and associated spatial transcriptomics data comprises fine-tuning the machine learning model using a fitness function incorporating a term representing a difference between a set of predicted spatial transcriptomics data for each three-dimensional image and the spatial transcriptomics data associated with the three- dimensional image.