Machine learning histological analysis for identification of molecular features

Through a histological computational model that blends multi-instance learning and attention mechanisms, the problem of feature recognition at different scales in histological images is solved, and more accurate analysis of tumor cell molecular characteristics is achieved, supporting more accurate disease diagnosis and treatment prediction.

CN120359549APending Publication Date: 2025-07-22GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085009.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-14
Filing Date
2023-12-13
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify molecular features at different scales in histological images, resulting in inaccurate analysis of phenotypic heterogeneity of tumor cells, affecting disease diagnosis and treatment response prediction.

Method used

Using a hybrid multi-instance learning method, tile features of different sizes are extracted from biological sample images through histological calculation models, and molecular features in the image are determined based on feature stitching and attention mechanisms, and bag-level label recognition is combined with clustering and pooling technology.

Benefits of technology

It improves the accuracy of identification of tumor cell molecular characteristics, supports more accurate disease diagnosis and treatment response prediction, and enhances the analytical ability of biological samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359549A_ABST
    Figure CN120359549A_ABST
Patent Text Reader

Abstract

A method may include determining a first plurality of tiles having a first tile size and a second plurality of tiles having a second tile size within an image of a biological sample. A feature extraction model may be applied to extract features from tiles of different sizes. A tiled set of features may be formed, each of the tiled set of features including a first feature from a first tile of the first plurality of tiles, a second feature from a second tile of the second plurality of tiles, and a third feature from a third tile of the second plurality of tiles. Molecular features present in the biological sample may be determined based on an attention-weighted location embedding of the stitched feature set and a joint representation of features across spatially adjacent tile clusters in the image. Related systems and computer program products are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 387,462, filed on December 14, 2022, entitled "Machine - Learning Histological Analysis for Identification of Molecular Features", which is hereby incorporated by reference in its entirety. Technical Field

[0003] The subject matter described herein generally relates to digital and computational pathology, and more particularly to deep - learning methods for identifying molecular features in histological images. Background Art

[0004] The phenotype of a cell can refer to a unique combination of morphological and functional characteristics produced by various cellular processes, including, for example, gene expression, protein expression, etc. In some cases, the complex interactions between a cell's genome, epigenome, and local environment can give rise to a set of observable characteristics, collectively referred to as the cell's phenotype. Although cell phenotypes, including those of tumor cells, are often attributed to genomic instability, more recently epigenetic and microenvironmental influences have received increasing attention. Such non - genetic factors can further increase the intrinsic diversity and plasticity of tumor cells. At the tumor level, non - genetic factors may lead to greater phenotypic heterogeneity, which enables tumor cells to evade immune responses and resist drug interventions. Summary of the Invention

[0005] Systems, methods, and articles of manufacture (including computer program products) are provided for machine - learning - enabled identification of molecular features in histological images. In one aspect, a system is provided that includes at least one processor and at least one memory. The at least one memory can include program code that, when executed by the at least one processor, provides operations. The operations can include: determining a first plurality of tiles having a first tile size within an image of a biological sample; determining a second plurality of tiles having a second tile size within the image of the biological sample; applying a feature extraction model to extract a first plurality of features from the first plurality of tiles of the first size; applying a feature extraction model to extract a second plurality of features from the second plurality of tiles of the second size; and determining one or more molecular features present in the biological sample depicted in the image based at least on the first plurality of features and the second plurality of features.

[0006] In another aspect, a machine learning-enabled method for identifying molecular features in histological images is provided. The method may include: determining a first plurality of tiles having a first tile size within an image of a biological sample; determining a second plurality of tiles having a second tile size within the image of the biological sample; applying a feature extraction model to extract a first plurality of features from the first plurality of tiles of the first size; applying the feature extraction model to extract a second plurality of features from the second plurality of tiles of the second size; and determining one or more molecular features present in the biological sample depicted in the image based at least on the first plurality of features and the second plurality of features.

[0007] In another aspect, a computer program product for machine learning-enabled identification of molecular features in histological images is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations. The operations may include: determining a first plurality of tiles having a first tile size within an image of a biological sample; determining a second plurality of tiles having a second tile size within the image of the biological sample; applying a feature extraction model to extract a first plurality of features from the first plurality of tiles of the first size; applying the feature extraction model to extract a second plurality of features from the second plurality of tiles of the second size; and determining one or more molecular features present in the biological sample depicted in the image based at least on the first plurality of features and the second plurality of features.

[0008] Specific implementations of the present subject matter may include, but are not limited to, methods consistent with the description provided herein and articles of manufacture including tangible embodied machine-readable media that are operable to cause one or more machines (e.g., computers, etc.) to cause operations implementing one or more of the described features. Similarly, a computer system may be described that may include one or more processors and one or more memories coupled to the one or more processors. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. A computer-implemented method consistent with one or more implementations of the present subject matter may be implemented by one or more data processors present in a single computing system or multiple computing systems. Such multiple computing systems may be connected and may exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.) or via a direct connection between one or more of the multiple computing systems.

[0009] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Referring to the specification, the drawings, and the claims, other features and advantages of the subject matter described herein will become apparent. Although certain features of the presently disclosed subject matter are described for illustrative purposes related to enabling machine learning-based identification of gene expression, protein expression, and gene signature expression in histological images, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are incorporated into and constitute a part of this specification, showing certain aspects of the subject matter disclosed herein, and together with the description, help to explain some of the principles associated with the disclosed embodiments. In the drawings,

[0011] Figure 1 a system diagram depicting an example of a digital pathology system according to some exemplary embodiments is shown;

[0012] Figure 2A a flowchart depicting an example of a process for machine learning-based identification of molecular features in histological images according to some exemplary embodiments is shown;

[0013] Figure 2B a flowchart depicting another example of a process for machine learning-based identification of molecular features in histological images according to some exemplary embodiments is shown;

[0014] Figure 3 a schematic diagram depicting an example of a histological computational model according to some exemplary embodiments is shown;

[0015] Figure 4A a schematic diagram depicting an example of a tile extractor and a feature extractor according to some exemplary embodiments is shown;

[0016] Figure 4B a schematic diagram depicting an example of cross-cluster attention according to some exemplary embodiments is shown;

[0017] Figure 5 an example of a preprocessed histological image according to some exemplary embodiments is shown;

[0018] Figure 6 a graph depicting the structural similarity (SSIM) index as a measure of the consistency between transforming growth factor (TGF)-β inhibitory membrane-associated protein (TIMAP) cell type prediction and tile-level gene expression prediction made by a histological computational model according to some exemplary embodiments is shown;

[0019] Figure 7A Depicts a histological image showing the concordance between tumor cells identified by transforming growth factor (TGF)-β inhibitory membrane-associated protein (TIMAP) cell type prediction and tile-level gene expression prediction made by a histological computational model according to some exemplary embodiments;

[0020] Figure 7B Depicts a histological image showing the concordance between lymphocytes identified by transforming growth factor (TGF)-β inhibitory membrane-associated protein (TIMAP) cell type prediction and tile-level gene expression prediction made by a histological computational model according to some exemplary embodiments;

[0021] Figure 7C Depicts a histological image showing the concordance between fibroblasts identified by transforming growth factor (TGF)-β inhibitory membrane-associated protein (TIMAP) cell type prediction and tile-level gene expression prediction made by a histological computational model according to some exemplary embodiments;

[0022] Figure 8A Depicts a histological image of a tumor region located based on molecular features identified by a histological computational model according to some exemplary embodiments;

[0023] Figure 8B Depicts a histological image of intratumoral heterogeneity captured based on molecular features identified by a histological computational model according to some exemplary embodiments; and

[0024] Figure 8C Depicts the concordance between the cyclin spatial pattern and the cyclin bulk RNA sequence expression pattern identified by a histological computational model according to some exemplary embodiments;

[0025] Figure 9A Depicts various examples of signatures associated with tiles depicting lymphocytes in a histological image according to some exemplary embodiments;

[0026] Figure 9B Depicts various examples of signatures associated with tiles depicting adipose, tumor, and mucin tissue structures in a histological image according to some exemplary embodiments;

[0027] Figure 10A Depicts a histological image showing the co-localization of fatty acid oxidation and proton transport signatures as predictive biomarkers for predicting clinical outcomes according to some exemplary embodiments;

[0028] Figure 10Bdepicts a histological image showing the co - localization of amino acid catabolism and neuronal signatures as predictive biomarkers for predicting clinical outcomes according to some exemplary embodiments; and

[0029] Figure 11 depicts a block diagram of an example of a computing system according to some exemplary embodiments.

[0030] When actually applied, like reference numerals denote like structures, features, or elements. Detailed Description

[0031] In highly heterogeneous diseases such as cancer, in - depth understanding of the molecular features present in diseased tissue and the surrounding microenvironment may be indispensable for accurately predicting clinical endpoints. For example, certain molecular features, such as gene expression, protein expression, and gene signature expression, can serve as biomarkers for diagnosing disease subtypes, prognosticating disease progression, and predicting responses to various treatments. However, conventional histological analysis techniques for identifying molecular features in microscopic images (e.g., hematoxylin and eosin (H&E) - stained whole - slide images, multiplex immunofluorescence (MxIF) - stained whole - slide images, etc.), including deep - learning - based methods, focus on features of a fixed size, while key insights are often obtained across a range of features of different sizes (e.g., from millimeter - scale features such as blood vessels to cell - scale features such as tissue microenvironment).

[0032] In some exemplary embodiments, a histological computational model can apply a mixed multi-instance learning (MIL) method to patches of different sizes in an image of a biological sample (e.g., a whole slide image (WSI), etc.). For example, the histological computational model can extract a first plurality of patches of a first size (e.g., 224×224 pixels) that capture features at a first scale (e.g., the cellular scale) and a second plurality of patches of a second size (e.g., 56×56 pixels) that capture features at a second scale (e.g., the millimeter scale) from an image of a biological sample. Additionally, the histological computational model can concatenate a first plurality of features extracted from the first plurality of patches of the first size with a second plurality of features extracted from the second plurality of patches of the second size. For example, in some cases, the histological computational model can apply pyramid concatenation, where features from a larger patch covering a portion of the image are concatenated with features from two or more smaller patches covering the same (or a similar) portion of the image. Thus, a first feature associated with a first patch of the first size is concatenated with at least a second feature associated with a second patch of the second size and a third feature associated with a third patch of the second size. Additionally, in some cases, a first feature associated with a first patch of the first size can be concatenated with a second feature from the first patch of the first size, a third feature from a second patch of the second size, and a fourth feature from a third patch of the second size.

[0033] In some exemplary embodiments, the histological computational model can determine one or more bag-levels for an image of a biological sample based on a joint representation of key instances from the first plurality of patches of the first size and the second plurality of patches of the second size. For example, the bag-level label of the image can indicate whether the biological sample depicted in the image is associated with a molecular feature such as gene expression, protein expression, or gene signature expression. In such a case, if the biological sample is positive for (or exhibits) the molecular feature, the biological sample can be associated with the molecular feature, and if the biological sample is negative for (or does not exhibit) the molecular feature, the biological sample may not be associated with the molecular feature.

[0034] In some exemplary embodiments, a bag-level label can be determined based at least on a joint representation of key instances included in a first plurality of patches and a second plurality of patches. In some cases, a bag-level label of an image can be determined based on a positional embedding of a first plurality of features extracted from a first plurality of patches of a first size, where the first plurality of features of the first size are concatenated with a second plurality of features extracted from a second plurality of patches of a second size. For example, the positional embedding can include a first position of a first patch of the first size, which embeds a first feature extracted from the first patch; a second position of a second patch of the second size, which embeds a second feature extracted from the second patch; and a third position of a third patch of the second size, which embeds a third feature extracted from the third patch. Thus, a bag-level label of the image can be determined to account for different scale features from patches of different sizes and the spatial distribution of these features within the image.

[0035] In some exemplary embodiments, a histology computational model can include an attention mechanism to identify one or more key instances on individual patches when determining a bag-level label of an image. Thus, in some cases, the histology computational model can include an attention generator network that is trained to determine, for each positional embedding (e.g., of a first feature of a first patch of a first size concatenated with a second feature of a second patch of a second size and a third feature of a third patch of the second size), a corresponding attention weight indicating whether the corresponding instance triggers the bag-level label of the image. For example, in some cases, a bag-level label of an image of a biological sample can be a binary value indicating whether the biological sample is associated with a particular molecular feature. A key instance in such a case can refer to a patch (or a cluster of patches) that triggers the bag-level label of the image by at least causing the bag-level label to assume a first value indicating that the biological sample is associated with (or is positive for) the molecular feature or a second value indicating that the biological sample is not associated with (or is negative for) the molecular feature.

[0036] In some exemplary embodiments, a histology computational model can determine a plurality of bag-level labels for an image of a biological sample, where each of the bag-level labels indicates, for example, whether the biological sample depicted in the image is associated with a molecular feature such as gene expression, gene signature expression, protein expression, etc. For example, the histology computational model can determine a first bag-level label for the image based on attention-weighted instances of a feature set of positional embeddings and mosaics from a first plurality of tiles and / or a second plurality of tiles. In such a case, each instance included in the image of the biological sample can refer to a positional embedding of the mosaicked feature set, including, for example, a first feature associated with a first tile of a first size mosaicked with at least a second feature associated with a second tile of a second size and a third feature associated with a third tile of the second size. Additionally, in some cases, the histology computational model can perform attention-based tile selection and pooling, and then instance regression, to determine a first bag-level label for the image of the biological sample based at least on the attention-weighted instances.

[0037] In some exemplary embodiments, the histology computational model can also determine a second bag-level label for the image based on different tile clusters within the image of the biological sample. For example, in some cases, the histology computational model can perform location-based clustering to identify one or more clusters of spatially adjacent tiles within a first plurality of tiles and a second plurality of tiles in the image. The histology computational model can perform cross-cluster attention map distillation in order to determine a label for each tile cluster that identifies molecular features present in tiles found in other tile clusters. Additionally, the histology computational model can determine a set of cross-cluster attention weights for each tile cluster, the set of cross-cluster attention weights including a first average attention weight of tiles within the tile cluster and a second average attention weight of tiles in other tile clusters. In some cases, the histology computational model can determine a second bag-level label for the image of the biological sample based at least on the set of cross-cluster attention weights associated with each tile cluster. For example, in some cases, the histology computational model can perform attention-based cluster selection and pooling, and then bag-level regression, to determine a second bag-level label for the image as a whole based at least on the attention-weighted tile clusters. In some cases, the histology computational model can determine an overall label for the image based at least on the first bag-level label determined by instance-level regression and the second bag-level label determined by bag-level regression, the overall label indicating, for example, whether the biological sample depicted in the image is associated with a molecular feature such as gene expression, gene signature expression, protein expression, etc.

[0038] Figure 1 A system diagram depicting an example of a digital pathology system 100 according to some exemplary embodiments is shown. Referring Figure 1 , the digital pathology system 100 can include a digital pathology platform 110, an imaging system 120, and a client device 130. AsFigure 1 As shown, the digital pathology platform 110, the imaging system 120, and the client device 130 can be communicatively coupled via a network 140. The network 140 can be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. The imaging system 120 can include one or more imaging devices (including, for example, a microscope, a digital camera, a whole slide scanner, an automated microscope, etc.). The client device 130 can be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc.

[0039] Referring again to Figure 1 , the digital pathology platform 110 can include a histology computational model 115 and an analysis engine 117. In the Figure 1 example shown, the digital pathology platform 110 can apply the histology computational model 115 to an image 125 of a biological sample to identify one or more molecular features present in the biological sample. Examples of molecular features can include gene expression, gene signature expression, and protein expression, as well as gene mutations, copy number alterations (CNA), cell phenotypes, etc. In some cases, the first image 125 can be a stained whole slide image (WSI), including, for example, a hematoxylin and eosin (H&E)-stained whole slide image, a multiplex immunofluorescence (MxIF)-stained whole slide image, an immunohistochemistry (IHC)-stained whole slide image, etc. In some cases, the analysis engine 117 can determine at least one of a disease diagnosis, disease progression, disease burden, treatment, treatment response, and survival prediction of a patient associated with the biological sample based at least in part on one or more molecular features present in the biological sample. Alternatively and / or additionally, the analysis engine 117 can identify one or more biomarkers and disease-modifying target genes based at least in part on one or more molecular features present in the biological sample. In some cases, the analysis engine 117 can also perform bulk RNA sequence prediction and in silico spatial transcriptomics based at least in part on one or more molecular features present in the biological sample to determine the spatial distribution of gene activity occurring within the biological sample.

[0040] Figure 2A FIG. depicts a flowchart of an example of a process 200 for enabling machine learning-based identification of molecular features in histological images according to some exemplary embodiments. Referring to Figure 2A , the process 200 can be performed by the digital pathology platform 110 to determine, for example, a bag-level label that indicates whether a biological sample depicted in the image 125 is associated with a molecular feature, which includes, for example, gene expression, gene signature expression, and protein expression, gene mutations, copy number alterations (CNA), cell phenotypes, etc.

[0041] At 202, the digital pathology platform 110 may determine a first plurality of tiles having a first size within an image of a biological sample. In some exemplary embodiments, the digital pathology platform 110 may extract tiles of different sizes from the image 125 of the biological sample in order to capture features at different scales, such as, for example, millimeter-scale features (such as blood vessels) and cell-scale features (such as tissue microenvironment). For further illustration, Figure 3 FIG. depicts a schematic diagram showing an example of a histology computational model 115 according to some exemplary embodiments. As Figure 3 shown, the histology computational model 115 may include a tile extractor 302. Figure 4A FIG. depicts a schematic diagram showing an example of the tile extractor 302 according to some exemplary embodiments. As Figure 4A shown, the tile extractor 302 may perform tile extraction to extract a first plurality of tiles 410 having a first size (e.g., 224×224 pixels) from the image 125 of the biological sample. In some cases, the first plurality of tiles 410 of the first size may capture features at a first scale, such as, for example, global features or cell-scale features present in the image 125.

[0042] In some cases, before applying the histology computational model 115, the image 125 may undergo various forms of image preprocessing. For example, in some cases, the image 125 may be preprocessed to reduce and / or remove artifacts. Alternatively and / or additionally, in some cases, the image 125 may be preprocessed to remove one or more background portions of the image 125. Figure 5 FIG. depicts an example in which the image 125 is preprocessed to remove artifacts and background. Further, in some cases, when determining the first plurality of tiles 410, the tile extractor 302 may exclude one or more tiles in which less than a threshold portion of the tile (e.g., less than 50% or another threshold portion of the tile) is covered by the biological sample.

[0043] At 204, the digital pathology platform 110 may determine a second plurality of tiles having a second size within an image of a biological sample. Again referring to Figure 3 and 4A, in some exemplary embodiments, a tile extractor 302 (or a different tile extractor) of the histological computational model 115 may extract a second plurality of tiles 420 of a second size from an image 125 of a biological sample. In some cases, the second plurality of tiles 420 of the second size may include a different number of pixels (e.g., 64×64 pixels) compared to the first plurality of tiles 420. Additionally, in some cases, a single tile in the first plurality of tiles 410 may cover the same (or similar) portion of the image 125 as two or more tiles in the second plurality of tiles 420. Thus, in some cases, the second plurality of tiles 420 of the second size may capture features at a second scale, such as local features or millimeter-scale features present in the image 125. Additionally, although Figure 4A each tile in the first plurality of tiles 410 and the second plurality of tiles 420 is shown as being the same size, the first plurality of tiles 410 and / or the second plurality of tiles 420 may also include tiles of different sizes. For example, in some cases, two or more tiles in the second plurality of tiles 420 cover the same (or similar) portion of the image 125 because a single tile in the first plurality of tiles 410 may have the same size or different sizes. Additionally, the same number or different numbers of tiles from the second plurality of tiles 420 of the second size may be associated with each tile in the first plurality of tiles 410 of the first size. For example, while a first number of tiles from the second plurality of tiles 420 of the second size may be associated with a first tile in the first plurality of tiles 410 of the first size, the same first number of tiles or a different second number of tiles from the second plurality of tiles 420 may be associated with a second tile in the first plurality of tiles 410. In some cases, when determining the second plurality of tiles 420 of the second size, the tile extractor 302 may exclude one or more tiles in which less than a threshold portion of the tile (e.g., less than 50% or another threshold portion of the tile) is covered by the biological sample.

[0044] At 206, the digital pathology platform 110 may extract a first plurality of features from the first plurality of tiles of the first size. Referring again to Figure 3 and 4A , the histological computational model 115 may include a feature extractor 304 (e.g., including a machine learning model 400, such as a vision transformer, etc.) that is trained to extract a first plurality of features from the first plurality of tiles 410 of the first size. In some cases, the first plurality of features may be at a first scale (e.g., global scale or cellular scale) corresponding to the first size of the first plurality of tiles 410. In Figure 3In the example shown, the feature extractor 304 is an indication - specific feature extractor trained to identify and extract features associated with a specific disease or a specific disease subclass (such as cancer). However, it should be understood that the feature extractor 304 can also be implemented as a general - purpose feature extractor trained to identify and extract features associated with multiple diseases or multiple disease subclasses.

[0045] At 208, the digital pathology platform 110 can extract a second plurality of features from a second plurality of tiles of a second size. In some cases, the feature extractor 304 (or a different feature extractor) of the histological computing model 115 can also be trained to extract a second plurality of features from a second plurality of tiles 420 of a second size. In some cases, the second plurality of features can be a second scale corresponding to the second size of the second plurality of tiles 420 (e.g., a local scale or a millimeter scale).

[0046] At 210, the digital pathology platform 110 determines one or more molecular features present in the biological sample depicted in the image based at least on the first plurality of features and the second plurality of features. In some exemplary embodiments, the histology computational model 115 may determine one or more bag-level labels of the image 125 based at least on the first plurality of features extracted from the first plurality of tiles 410 of the first size and the second plurality of features extracted from the second plurality of tiles 420 of the second size. In some cases, the bag-level label of the image 125 may be determined based at least on the position embedding 308 of the concatenated feature set 306. For example, in some cases, the concatenated feature set 306 may include a first feature associated with a first tile from the first plurality of tiles 410 of the first size, the first feature being concatenated with at least a second feature associated with a second tile from the second plurality of tiles 420 of the second size and a third feature associated with a third tile from the second plurality of tiles 420 of the second size. Additionally, in some cases, the concatenated feature set 306 may include a first feature associated with a first tile of the first size, a second feature from the first tile of the first size, a third feature from the second tile of the second size, and a fourth feature from the third tile of the second size. Meanwhile, the position embedding 308 of the concatenated feature set 306 may further include a first position of the first tile, a second position of the second tile, and / or a third position of the third tile. In some cases, the first position of the first tile, the second position of the second tile, and the third position of the third tile may each include a set of coordinates of one or more pixels included in the corresponding tile, such as, for example, one or more pixels occupying a corner (e.g., the upper left corner) of the corresponding tile. As will be described in more detail below, in some cases, the bag-level label indicating whether the biological sample depicted in the image 125 is associated with the molecular feature may be determined based at least on the attention-weighted instances, where each of the attention-weighted instances corresponds to the position embedding of the concatenated features from the first plurality of tiles 410 and the second plurality of tiles 420.

[0047] At 212, the digital pathology platform 110 may perform one or more downstream analysis tasks based at least on one or more molecular features present in a biological sample. In some exemplary embodiments, the digital pathology platform 110, such as the analysis engine 117, may perform various downstream analysis tasks based on one or more molecular features, such as gene expression, gene signature expression, and protein expression, as well as gene mutations, copy number alterations (CNAs), cell phenotypes, etc., identified as being present (or absent) in the biological sample depicted in the image 125. For example, in some cases, the analysis engine 117 may determine at least one of a disease diagnosis, disease progression, treatment, treatment response, and survival prediction of a patient associated with the biological sample based at least on one or more molecular features identified as being present (or absent) in the biological sample. In some cases, the analysis engine 117 may also identify one or more biomarkers and disease-modifying target genes based at least on one or more molecular features identified as being present (or absent) in the biological sample. Alternatively and / or additionally, in some cases, the analysis engine 117 may perform bulk RNA sequence prediction and in silico spatial transcriptomics based at least on one or more molecular features identified as being present (or absent) in the biological sample to determine the spatial distribution of gene activity occurring within the biological sample.

[0048] Figure 2B FIG. depicts a flowchart of an example of a process 250 for machine learning-enabled identification of molecular features in a histological image. Referring Figure 2B thereto, the process 250 may be performed by the digital pathology platform 110 to determine a bag-level label based on features extracted from tiles of different sizes in the image 125, the bag-level label indicating whether the biological sample depicted in the image 125 is associated with a molecular feature, such as gene expression, gene signature expression, and protein expression, gene mutations, copy number alterations (CNAs), cell phenotypes, etc. In some cases, the process 250 may implement the operation 210 of the process 200 described in Figure 2A reference

[0049] At 252, the digital pathology platform 110 may concatenate a first plurality of features extracted from a first plurality of tiles of a first size with a second plurality of features extracted from a second plurality of tiles of a second size. For example, as Figure 3 and 4AAs shown, the digital pathology platform 110 can generate a concatenated feature set 306 based at least on a first plurality of features extracted from a first plurality of tiles 410 of a first size and a second plurality of features extracted from a second plurality of tiles 420 of a second size. In some cases, the concatenated feature set 306 can include, for example, a concatenation of a first feature of a first tile from the first plurality of tiles 410 of the first size, a second feature of a second tile from the second plurality of tiles 420 of the second size, and a third feature of a third tile from the second plurality of tiles 420 of the second size.

[0050] At 254, the digital pathology platform 110 can determine a positional embedding for each concatenated feature set that includes a first feature of a first tile from a first plurality of tiles of a first size, a second feature of a second tile from a second plurality of tiles of a second size, and a third feature of a third tile from a second plurality of tiles of a second size. For example, in some cases, the histology computing model 115 can generate a corresponding positional embedding 308 for each concatenated feature set 306. In some cases, the positional embedding 308 of the concatenated feature set 306 can include, for example, a first position of the first tile, a second position of the second tile, and / or a third position of the third tile. Additionally, in some cases, the first position of the first tile, the second position of the second tile, and the third position of the third tile can each include a set of coordinates of one or more pixels included in the corresponding tile. For example, in some cases, the positional embedding 308 of the concatenated feature set 306 can be generated based on pixels that occupy a corner (e.g., the upper left corner) of one or more of the first tile from the first plurality of tiles 410 of the first size, the second tile from the second plurality of tiles 420 of the second size, and the third tile from the second plurality of tiles 420 of the second size.

[0051] At 256, the digital pathology platform 110 can determine a first bag-level label of an image of a biological sample based at least on the attention-weighted positional embedding of the concatenated feature set. Referring again to Figure 3 , in some exemplary embodiments, the histology computing model 115 can include an attention generator network 310 that is configured to determine, for the positional embedding 308 of each concatenated feature set 306, attention weights that indicate the relative importance of individual instances that include the corresponding tile and the features contained therein. In this case, the attention weights assigned to the positional embedding 308 of the concatenated feature set 306 can be values corresponding to the degree of contribution of that particular instance to the bag-level label of the image 125. Thus, important (or key) instances that trigger (or contribute to) the bag-level label can be associated with higher attention weights compared to minor instances that have less impact on the bag-level label. For example, in Figure 3In the example shown, the attention generator network 310 can determine corresponding attention weights a1, a2, …, a for each position embedding in the N position embeddings 308 N .

[0052] Referring again to Figure 3 , the histology computational model 110 can include an attention-based tile selection and pooling network 312, followed by an instance regressor 314 that is trained to determine a first bag-level label based at least on attention-weighted instances (e.g., attention-weighted position embeddings 308 of the concatenated feature set 306), where the first bag-level label indicates whether the biological sample depicted in the image 125 is associated with a molecular feature such as gene expression, gene signature expression, protein expression, gene mutation, copy number alteration (CNA), cell phenotype, etc. For example, in some cases, the first bag-level label can be a binary label having a first value (e.g., “1”) indicating that the biological sample is associated with (or is positive for) the molecular feature or a second value (e.g., “0”) indicating that the biological sample is not associated with (or is negative for) the molecular feature. In some cases, the instance regressor 314 can be implemented using a neural network, a Hopfield network, etc.

[0053] At 258, the digital pathology platform 110 can identify one or more tile clusters based at least on the positions of each tile in the first plurality of tiles and the second plurality of tiles. In some exemplary embodiments, the histology computational model 115 can apply a clustering algorithm 316 to identify one or more clusters of similar tiles within the first plurality of tiles 410 and / or the second plurality of tiles 420. In some cases, the histology computational model 115 can apply a clustering algorithm 316 to perform location-based clustering such that the resulting tile clusters include spatially proximate tiles (e.g., tiles that occupy the same or similar regions of the image 125). In Figure 3 the example shown, the histology computational model 115 can apply a clustering algorithm 316 to determine k-number of tile clusters represented as C1, C2, …, C k within the first plurality of tiles 410 of a first size and the second plurality of tiles 420 of a second size.

[0054] At 260, the digital pathology platform 110 can determine a set of cross-cluster attention weights for each tile cluster, where the set of cross-cluster attention weights includes a first average attention weight of the tiles in the tile cluster and a second average attention weight of the tiles in other tile clusters. Referring again to Figure 3 , the histology computational model 115 can perform cross-cluster attention map (CAM) distillation in order to determine a set of cross-cluster attention weights for each tile cluster, where the set of cross-cluster attention weights includes, for example, a first average attention weight C of the tiles within tile cluster k ka and a second average attention weight nC of a tile that is not in the tile cluster k but in other tile clusters k a. In Figure 3 the example shown, the histology computational model 115 can determine a set of cross-cluster attention weights by performing cross-cluster attention 320 across tile clusters. For further illustration, Figure 4B a schematic diagram depicting an example of cross-cluster attention 320 is shown, where the histology computational model 115 determines a first average attention weight C of tiles within the tile cluster k based at least on the cross-cluster attention map 450 k a and a second average attention weight nC of a tile that is not within the tile cluster k k a. As Figure 3 shown, the first average attention weight C k a can be associated with a first joint representation of the features of tiles in the tile cluster k, while the second average attention weight nC k a can be associated with a second joint representation of the features of tiles not in the tile cluster k.

[0055] At 262, the digital pathology platform 110 can determine a second bag-level label indicating whether the biological sample depicted in the image is associated with a molecular feature based at least on the cross-cluster attention weights of each tile cluster. In some exemplary embodiments, Figure 3 the example of the histology computational model 115 shown can include an attention-based cluster selection and pooling network 322 and a bag regressor 324 that is trained to determine a second bag-level label indicating whether the biological sample depicted in the image 125 is associated with a molecular feature based at least on the attention-weighted tile clusters. For example, as Figure 3 shown, for each tile cluster k, a second bag-level label can be determined based on a first joint representation of the features present in the tiles within the tile cluster k (weighted by the first average attention weight C k a) and a second joint representation of the features present in the tiles outside the tile cluster k (weighted by the second average attention weight nC k a). In some cases, the bag regressor 324 can be implemented using a neural network, a Hopfield network, etc.

[0056] At 264, the digital pathology platform 110 can determine an overall label indicating whether the biological sample depicted in the image is associated with a molecular feature based at least on the first bag-level label and the second bag-level label. For example, as Figure 3As shown, the histological computational model 115 can determine an overall label based at least on a first bag-level label of the image 125 determined by the instance regressor 314 and a second bag-level label of the image 125 determined by the packet regressor 324, where the overall label indicates whether the biological sample depicted in the image 125 is associated with molecular features such as gene expression, gene signature expression, protein expression, gene mutation, copy number alteration (CNA), cell phenotype, etc.

[0057] In some exemplary embodiments, the performance of the histological computational model 115 in determining whether the biological sample depicted in the image 125 is associated with molecular biomarkers (e.g., gene expression, gene signature expression, protein expression, gene mutation, copy number alteration (CNA), cell phenotype, etc.) can be evaluated based on, for example, the consistency between the transforming growth factor (TGF)-β inhibitory membrane-associated protein (TIMAP) cell mask and the tile-level gene expression prediction made by the histological computational model 115. Figure 6 A graph is depicted showing the structural similarity (SSIM) index as a measure of the consistency between the TIMAP cell type prediction and the tile-level gene expression prediction made by the histological computational model according to some exemplary embodiments. Figure 7A A histological image is depicted showing the consistency between the tumor cells identified by the TIMAP cell type prediction and the tile-level gene expression prediction made by the histological computational model 115. Figure 7B A histological image is depicted showing the consistency between the lymphocytes identified by the TIMAP cell type prediction and the tile-level gene expression prediction made by the histological computational model 115. Figure 7C A histological image is depicted showing the consistency between the fibroblasts identified by the TIMAP cell type prediction and the tile-level gene expression prediction made by the histological computational model 115.

[0058] In some cases, the performance of the histological computational model 115 can also be evaluated based on the consistency with expert annotations. For example, Figure 8A A histological image is depicted showing the tumor region located based on the molecular features identified by the histological computational model 115 and the corresponding expert annotations of the same image. Figure 8B A histological image is depicted showing the intra-tumor heterogeneity captured based on the molecular features identified by the histological computational model 115 and the corresponding expert annotations of the same image. In some cases, the predictions made by the histological computational model 115 can also be verified based on the underlying bulk RNA sequence expression patterns. For example, Figure 8C A graph is depicted showing the consistency between the cyclin spatial pattern identified by the histological computational model 115 and the same cyclin pattern identified by bulk RNA sequence expression.

[0059] As described above, in some exemplary embodiments, the digital pathology platform 110 may apply a histology computational model 115 to determine whether a biological sample depicted in an image 125 is associated with one or more molecular features, the one or more molecular features including, for example, gene expression, gene signature expression, protein expression, gene mutations, copy number alterations (CNAs), cell phenotypes, and the like. For example, in some cases, the histology computational model 115 may output a binary label for a particular molecular feature, the binary label having a first value (e.g., "1") indicating that the biological sample is associated with (or is positive for) the molecular feature or a second value (e.g., "0") indicating that the biological sample is not associated with (or is negative for) the molecular feature. Figure 9A depicts various examples of signatures associated with tiles depicting lymphocytes in a histological image (such as image 125), while Figure 9B depicts various examples of signatures associated with tiles depicting adipose, tumor, and mucin tissue structures in a histological image (such as image 125).

[0060] In some exemplary embodiments, the digital pathology platform 110 may perform various downstream analysis tasks based at least on one or more molecular features. For example, in some cases, one or more molecular features identified within a biological sample depicted in an image 125 may be used as biomarkers for determining at least one of a disease diagnosis, disease progression, disease burden, treatment, treatment response, and prediction for a patient associated with the biological sample. For example, Figure 10A depicts a histological image showing co-localization of fatty acid oxidation and proton transport signatures as predictive biomarkers for predicting clinical outcomes according to some exemplary embodiments. Figure 10B depicts a histological image showing co-localization of amino acid catabolism and neuronal signatures as predictive biomarkers for predicting clinical outcomes according to some exemplary embodiments.

[0061] Figure 11 depicts a block diagram of an example of a computing system 1100. Referring to Figures 1 to 11 , the computing system 1100 may be used to implement the digital pathology platform 110, the imaging system 120, the client device 130, and / or any components thereof.

[0062] As Figure 11As shown, the computing system 1100 may include a processor 1110, a memory 1120, a storage device 1130, and an input / output device 1140. The processor 1110, the memory 1120, the storage device 1130, and the input / output device 1140 may be interconnected via a system bus 1150. The processor 1110 is capable of processing instructions for execution within the computing system 1100. Such executed instructions may implement one or more components, such as, for example, a digital pathology platform 110, an imaging system 120, a client device 130, etc. In some exemplary embodiments, the processor 1110 may be a single-threaded processor. Alternatively, the processor 1110 may be a multi-threaded processor. The processor 1110 is capable of processing instructions stored in the memory 1120 and / or the storage device 1130 to display graphical information for a user interface provided via the input / output device 1140.

[0063] The memory 1120 is a computer-readable medium for storing information within the computing system 1100, such as a volatile or non-volatile computer-readable medium. For example, the memory 1120 may store data structures representing a configuration object database. The storage device 1130 is capable of providing persistent storage for the computing system 1100. The storage device 1130 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device or other suitable persistent storage means. The input / output device 1140 provides input / output operations for the computing system 1100. In some exemplary embodiments, the input / output device 1140 includes a keyboard and / or a pointing device. In various specific implementations, the input / output device 1140 includes a display unit for displaying a graphical user interface.

[0064] According to some exemplary embodiments, the input / output device 1140 may provide input / output operations for a network device. For example, the input / output device 1140 may include an Ethernet port or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0065] In some exemplary embodiments, computing system 1100 can be used to execute various interactive computer software applications that can be used to organize, analyze, and / or store data in various formats. Alternatively, computing system 1100 can be used to execute any type of software application. These applications can be used to perform various functions, such as, for example, scheduling functions (e.g., generating, managing, editing spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functions, communication functions, etc. The applications can include various additional functions or can be stand-alone computing products and / or functions. Once activated within an application, the functions can be used to generate a user interface provided via input / output device 1140. The user interface can be generated by computing system 1100 and presented to a user (e.g., on a computer screen monitor, etc.).

[0066] One or more aspects or features of the subject matter described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include a particular implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor (which can be special purpose or general purpose, coupled to receive data and instructions from, and to send data and instructions to, a storage system, at least one input device, and at least one output device). The programmable system or computing system can include clients and servers. Typically, the clients and servers are remotely located from each other and generally interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and the client-server relationship between them.

[0067] These computer programs may also be referred to as programs, software, software applications, applications, components, or code, including machine instructions for a programmable processor, and may be implemented in a high-level procedural and / or object-oriented programming language and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device (such as, for example, a disk, optical disk, memory, and programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the instructions of the machine as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium may non-transitorily store such machine instructions (such as, for example, in non-transitory solid-state memory or a magnetic hard disk drive or any equivalent storage medium). The machine-readable medium may alternatively or additionally store such machine instructions in a transitory manner (such as, for example, in a processor cache or other random access memory associated with one or more physical processor cores).

[0068] To provide for interaction with a user, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), or a light emitting diode (LED) monitor for displaying information to the user) and a keyboard and a pointing device (such as, for example, a mouse or a trackball by which the user may provide input to the computer). Other kinds of devices may also be used to provide for interaction with the user. For example, the feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; the input from the user may be received in any form (including sound, voice, or tactile input). Other possible input devices include a touch screen or other touch-sensitive device, such as a single-point or multi-point resistive or capacitive trackpad, speech recognition hardware and software, an optical scanner, an optical indicator, a digital image capture device, and associated interpretation software, etc.

[0069] Embodiment

[0070] The provided embodiments are as follows:

[0071] 1. A computer-implemented method, comprising:

[0072] Determining a first plurality of tiles having a first tile size within an image of a biological sample;

[0073] Determining a second plurality of tiles having a second tile size within the image of the biological sample;

[0074] Applying a feature extraction model to extract a first plurality of features from the first plurality of tiles of the first size;

[0075] Apply a feature extraction model to extract a second plurality of features from a second plurality of tiles of a second size; and

[0076] Determine one or more molecular features present in the biological sample depicted in the image based at least on the first plurality of features and the second plurality of features.

[0077] 2. The method according to embodiment 1, wherein the feature extraction model includes a first machine learning model that is trained to extract a first plurality of features from a first plurality of tiles having a first tile size, and wherein the feature extraction model further includes a second machine learning model that is trained to extract a second plurality of features from a second plurality of tiles having a second tile size.

[0078] 3. The method according to any one of embodiments 1 to 2, wherein the first tile size of the first plurality of tiles includes a different number of pixels compared to the second tile size of the second plurality of tiles.

[0079] 4. The method according to any one of embodiments 1 to 3, wherein the first plurality of features extracted from the first plurality of tiles having the first tile size include global features present in the image of the biological sample, and wherein the second plurality of features extracted from the second plurality of tiles having the second tile size include local features present in the image of the biological sample.

[0080] 5. The method according to any one of embodiments 1 to 4, wherein the first plurality of features extracted from the first plurality of tiles having the first tile size include cell-scale features, and wherein the second plurality of features extracted from the second plurality of tiles having the second tile size include millimeter-scale features.

[0081] 6. The method according to any one of embodiments 1 to 5, wherein the first tile size of the first plurality of tiles is 56 pixels x 56 pixels.

[0082] 7. The method according to any one of embodiments 1 to 6, wherein the second tile size of the second plurality of tiles is 224 pixels x 224 pixels.

[0083] 8. The method according to any one of embodiments 1 to 7, further comprising:

[0084] Determine a third plurality of tiles having a third size within the image of the biological sample;

[0085] Apply a feature extraction model to extract a third plurality of features from the third plurality of tiles; and determine one or more molecular features in the biological sample based at least on the third plurality of features.

[0086] 9. The method according to any one of embodiments 1 to 8, further comprising:

[0087] concatenating the first plurality of features and the second plurality of features; and

[0088] determining one or more molecular features present in the biological sample based at least on the concatenation of the first plurality of features and the second plurality of features.

[0089] 10. The method according to embodiment 9, wherein the concatenation of the first plurality of features and the second plurality of features comprises concatenating a first feature associated with a first tile of the first plurality of tiles of a first size and a second feature associated with a second tile of the second plurality of tiles of a second size.

[0090] 11. The method according to embodiment 10, wherein the concatenation of the first plurality of features and the second plurality of features further comprises concatenating a second feature associated with the first tile of the first plurality of tiles of the first size and a second feature associated with the second tile of the second plurality of tiles of the second size.

[0091] 12. The method according to embodiment 10, wherein the concatenation of the first plurality of features and the second plurality of features further comprises concatenating the first feature of the first tile, the second feature of the second tile, and a third feature of a third tile of the second plurality of tiles of the second size.

[0092] 13. The method according to embodiment 12, further comprising:

[0093] generating a positional embedding of the concatenated feature set, the concatenated feature set comprising the first feature of the first tile, the second feature of the second tile, and the third feature of the third tile.

[0094] 14. The method according to embodiment 13, wherein the positional embedding comprises a first position of the first tile, a second position of the second tile, and / or a third position of the third tile.

[0095] 15. The method according to embodiment 14, wherein the first position of the first tile, the second position of the second tile, and the third position of the third tile each comprise a set of coordinates of at least one pixel from the corresponding tile.

[0096] 16. The method according to any one of embodiments 1 to 15, further comprising:

[0097] determining a first bag-level label indicative of one or more molecular features present in the biological sample based at least on the positional embeddings of the concatenated feature sets associated with each tile of the first plurality of tiles and two or more corresponding tiles of the second plurality of tiles.

[0098] 17. The method according to embodiment 16, wherein one or more molecular features are further determined based at least on the attention weights associated with each position embedding.

[0099] 18. The method according to embodiment 17, wherein one or more molecular features are determined by applying an attention-based patch selection and pooling network and an instance regressor to a plurality of attention-weighted position embeddings.

[0100] 19. The method according to embodiment 16, further comprising:

[0101] determining a second bag-level label indicative of one or more molecular features present in the biological sample based at least on a joint representation of features associated with one or more patch clusters.

[0102] 20. The method according to embodiment 19, further comprising:

[0103] clustering a first plurality of patches and a second plurality of patches into one or more patch clusters.

[0104] 21. The method according to embodiment 20, wherein the clustering is performed based at least on the position information of each patch in the first plurality of patches and the second plurality of patches.

[0105] 22. The method according to embodiment 20, wherein the first plurality of patches and the second plurality of patches are clustered into a configurable number of clusters.

[0106] 23. The method according to embodiment 20, further comprising:

[0107] for each patch cluster, determining a first average attention weight for the features of the patches in the cluster and a second average attention weight for the features of the patches not in the cluster.

[0108] 24. The method according to embodiment 23, further comprising:

[0109] determining one or more molecular features present in the biological sample based at least on the first average attention weight applied to the first joint representation of the features of the patches in the cluster and the second average attention weight applied to the second joint representation of the features of the patches not in the cluster.

[0110] 25. The method according to embodiment 24, wherein one or more molecular features present in the biological sample are determined by applying at least an attention-based cluster selection and pooling network and a bag-level regressor to: the first average attention weight applied to the first joint representation of the features of the patches in the cluster, and the second average attention weight applied to the second joint representation of the features of the patches not in the cluster.

[0111] 26. The method according to embodiment 19, further comprising:

[0112] Determining an overall label indicative of one or more molecular features present in the biological sample based at least on the first bag-level label and the second bag-level label.

[0113] 27. The method according to any one of embodiments 1 to 26, wherein the feature extraction model is a vision transformer.

[0114] 28. The method according to any one of embodiments 1 to 27, wherein the image is a whole slide image.

[0115] 29. The method according to any one of embodiments 1 to 28, wherein the image is a hematoxylin and eosin (H&E) stained whole slide image.

[0116] 30. The method according to any one of embodiments 1 to 29, wherein the biological sample comprises one or more tissue fragments, free cells, and / or body fluids.

[0117] 31. The method according to any one of embodiments 1 to 30, wherein the biological sample comprises tumor tissue.

[0118] 32. The method according to any one of embodiments 1 to 31, wherein the feature extraction model is trained to extract features associated with a specific disease or a specific disease subtype.

[0119] 33. The method according to any one of embodiments 1 to 32, wherein the feature extraction model is trained to extract features associated with a specific cancer or a specific cancer subtype.

[0120] 34. The method according to any one of embodiments 1 to 33, wherein the one or more molecular features comprise gene expression, gene signature expression, protein expression, gene mutation, copy number alteration (CNA), and / or cell phenotype.

[0121] 35. The method according to any one of embodiments 1 to 34, further comprising:

[0122] Identifying one or more biomarkers and disease modifying target genes based at least on one or more molecular features present in the biological sample.

[0123] 36. The method according to any one of embodiments 1 to 35, further comprising:

[0124] Performing bulk RNA sequence prediction based at least on one or more molecular features present in the biological sample.

[0125] 37. The method according to any one of embodiments 1 to 36, further comprising:

[0126] Performing computational spatial transcriptomics based at least on one or more molecular features present in a biological sample to determine the spatial distribution of gene activities occurring within the biological sample.

[0127] 38. The method according to any one of embodiments 1 to 37, further comprising:

[0128] Determining at least one of disease diagnosis, disease progression, disease burden, treatment, treatment response, and survival prediction for a patient associated with the biological sample based at least on one or more molecular features present in the biological sample.

[0129] 39. A system, comprising:

[0130] At least one data processor; and

[0131] At least one memory storing instructions that, when executed by the at least one data processor, cause operations including the method according to any one of embodiments 1 to 38.

[0132] 40. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations including the method according to any one of embodiments 1 to 38.

[0133] In the above description and claims, phrases such as "at least one" or "one or more" may appear, followed by a list of combinations of elements or features. The term "and / or" may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to mean any element or feature listed individually, or any other recited element or feature in combination with any other recited element or feature. For example, the phrases "at least one of A and B"; "one or more of A and B"; "A and / or B" are each intended to mean "A alone, B alone, or A and B together". Similar interpretations apply to lists including three or more items. For example, the phrases "at least one of A, B, and C"; "one or more of A, B, and C" and "A, B, and / or C" are each intended to mean "A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together". The use of the term "based on" in the above and claims is intended to mean "at least partially based on", such that unrecited features or elements are also permissible.

[0134] Depending on the desired configuration, the subject matter described herein may be embodied in a system, apparatus, method, and / or article. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are only some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, other features and / or variations may be provided in addition to those features and / or variations set forth herein. For example, the above-described specific embodiments may be directed to various combinations and sub-combinations of the disclosed features and / or to combinations and sub-combinations of several further features disclosed above. Additionally, the logical flows depicted in the figures and / or described herein need not be in the particular order shown or in sequential order to achieve the desired result. Other specific embodiments may be within the scope of the following claims.

Claims

1. A computer-implemented method, comprising: Determining a first plurality of tiles having a first tile size within an image of a biological sample; Determining a second plurality of tiles having a second tile size within the image of the biological sample; Applying a feature extraction model to extract a first plurality of features from the first plurality of tiles of the first size; Applying the feature extraction model to extract a second plurality of features from the second plurality of tiles of the second size; And Determining one or more molecular features present in the biological sample depicted in the image based at least on the first plurality of features and the second plurality of features.

2. The method according to claim 1, wherein the feature extraction model comprises a first machine learning model trained to extract the first plurality of features from the first plurality of tiles having the first tile size, and wherein the feature extraction model further comprises a second machine learning model trained to extract the second plurality of features from the second plurality of tiles having the second tile size.

3. The method according to claim 1, wherein the first tile size of the first plurality of tiles comprises a different number of pixels compared to the second tile size of the second plurality of tiles.

4. The method according to claim 1, wherein the first plurality of features extracted from the first plurality of tiles having the first tile size comprise global features present in the image of the biological sample, and wherein the second plurality of features extracted from the second plurality of tiles having the second tile size comprise local features present in the image of the biological sample.

5. The method according to claim 1, wherein the first plurality of features extracted from the first plurality of tiles having the first tile size comprise cell-scale features, and wherein the second plurality of features extracted from the second plurality of tiles having the second tile size comprise millimeter-scale features.

6. The method according to claim 1, wherein the first tile size of the first plurality of tiles is 56 pixels x 56 pixels.

7. The method according to claim 1, wherein the second tile size of the second plurality of tiles is 224 pixels x 224 pixels.

8. The method according to claim 1, further comprising: Determining a third plurality of tiles having a third size within the image of the biological sample; Applying the feature extraction model to extract a third plurality of features from the third plurality of tiles; And Determining the one or more molecular features in the biological sample based at least on the third plurality of features.

9. The method according to claim 1, further comprising: Concatenating the first plurality of features and the second plurality of features; And Determining the one or more molecular features present in the biological sample based at least on the concatenation of the first plurality of features and the second plurality of features.

10. The method according to claim 9, wherein the stitching of the first plurality of features and the second plurality of features includes stitching a first feature associated with a first tile of the first plurality of tiles of the first size and a second feature associated with a second tile of the second plurality of tiles of the second size.

11. The method according to claim 10, wherein the stitching of the first plurality of features and the second plurality of features further includes stitching a second feature associated with the first tile of the first plurality of tiles of the first size and the second feature associated with the second tile of the second plurality of tiles of the second size.

12. The method according to claim 10, wherein the stitching of the first plurality of features and the second plurality of features further includes stitching the first feature of the first tile and the second feature of the second tile with a third feature of a third tile of the second plurality of tiles of the second size.

13. The method according to claim 12, further comprising: generating a positional embedding of the stitched feature set, the stitched feature set including the first feature of the first tile, the second feature of the second tile, and the third feature of the third tile.

14. The method according to claim 13, wherein the positional embedding includes a first position of the first tile, a second position of the second tile, and / or a third position of the third tile.

15. The method according to claim 14, wherein the first position of the first tile, the second position of the second tile, and the third position of the third tile each include a set of coordinates of at least one pixel from the corresponding tile.

16. The method according to claim 1, further comprising: determining a first bag-level label indicative of the one or more molecular features present in the biological sample, based at least on the positional embeddings of the stitched feature sets associated with each tile of the first plurality of tiles and two or more corresponding tiles of the second plurality of tiles.

17. The method according to claim 16, wherein the one or more molecular features are further determined based at least on the attention weights associated with each positional embedding.

18. The method according to claim 17, wherein the one or more molecular features are determined by applying an attention-based tile selection and pooling network and an instance regressor to a plurality of attention-weighted positional embeddings.

19. The method according to claim 16, further comprising: determining a second bag-level label indicative of the one or more molecular features present in the biological sample, based at least on a joint representation of the features associated with one or more tile clusters.

20. The method according to claim 19, further comprising: clustering the first plurality of tiles and the second plurality of tiles into the one or more tile clusters.

21. The method according to claim 20, wherein the clustering is performed based at least on the position information of each tile in the first plurality of tiles and the second plurality of tiles.

22. The method according to claim 20, wherein the first plurality of tiles and the second plurality of tiles are clustered into a configurable number of clusters.

23. The method according to claim 20, further comprising: For each tile cluster, determining a first average attention weight for the features of the tiles in the cluster and a second average attention weight for the features of the tiles not in the cluster.

24. The method according to claim 23, further comprising: Determining the one or more molecular features present in the biological sample based at least on the first average attention weight of the first joint representation applied to the features of the tiles in the cluster and the second average attention weight of the second joint representation applied to the features of the tiles not in the cluster.

25. The method according to claim 24, wherein the one or more molecular features present in the biological sample are determined by applying at least an attention-based cluster selection and pooling network and a bag-level regressor to: the first average attention weight of the first joint representation applied to the features of the tiles in the cluster, and the second average attention weight of the second joint representation applied to the features of the tiles not in the cluster.

26. The method according to claim 19, further comprising: Determining an overall label indicating the one or more molecular features present in the biological sample based at least on the first bag-level label and the second bag-level label.

27. The method according to claim 1, wherein the feature extraction model is a vision transformer.

28. The method according to claim 1, wherein the image is a whole slide image.

29. The method according to claim 1, wherein the image is a hematoxylin and eosin (H&E)-stained whole slide image.

30. The method according to claim 1, wherein the biological sample comprises one or more tissue fragments, free cells, and / or body fluids.

31. The method according to claim 1, wherein the biological sample comprises tumor tissue.

32. The method according to claim 1, wherein the feature extraction model is trained to extract features associated with a specific disease or a specific disease subtype.

33. The method according to claim 1, wherein the feature extraction model is trained to extract features associated with a specific cancer or a specific cancer subtype.

34. The method according to claim 1, wherein the one or more molecular features comprise gene expression, gene signature expression, protein expression, gene mutation, copy number alteration (CNA), and / or cell phenotype.

35. The method according to claim 1, further comprising: Identifying one or more biomarkers and disease-modifying target genes based at least on the one or more molecular features present in the biological sample.

36. The method according to claim 1, further comprising: Performing bulk RNA sequence prediction based at least on the one or more molecular features present in the biological sample.

37. The method according to claim 1, further comprising: Performing in silico spatial transcriptomics based at least on the one or more molecular features present in the biological sample to determine the spatial distribution of gene activity occurring within the biological sample.

38. The method according to claim 1, further comprising: Determining at least one of disease diagnosis, disease progression, disease burden, treatment, treatment response, and survival prediction for a patient associated with the biological sample based at least on the one or more molecular features present in the biological sample.

39. A system, comprising: At least one data processor; And At least one memory storing instructions that, when executed by the at least one data processor, cause operations including the method according to any one of claims 1 to 38.

40. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations including the method according to any one of claims 1 to 38.