Systems and methods for processing electronic images using deep underlying models

Deep underlying models trained in a self-supervised manner address the challenges of large-scale image processing by enabling efficient, pan-cancer, and pan-tissue analysis of high-resolution medical images, enhancing predictive accuracy and data integration for disease prediction and treatment planning.

JP2025539398APending Publication Date: 2025-12-05PAIGE AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025530708
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-11-28
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing systems for large-scale image processing, particularly in computational pathology, face challenges due to annotation requirements, variations in tissue types, scarcity of data, and limited generalization across applications, especially when dealing with high-resolution images, leading to computational intensity and poor training performance.

Method used

The use of deep underlying models trained in a self-supervised manner to generate and modify foundational models for inferring metadata from digital medical images, allowing for pan-cancer and pan-tissue analysis without exhaustive annotation, and enabling feature sharing while preserving differential privacy.

Benefits of technology

Enables efficient analysis of large, variable, and unannotated data across various modalities, improving predictive accuracy and generalization to new domains, and facilitating the integration of diverse medical data types for enhanced disease prediction and treatment planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539398000001_ABST
    Figure 2025539398000001_ABST
Patent Text Reader

Abstract

Systems and methods are disclosed for processing digital medical images to infer metadata from those images. In some aspects, the digital medical images may be processed to infer metadata by receiving a plurality of digital medical images, receiving a prompt, the prompt being a request for a particular type of metadata to be inferred from the plurality of digital medical images, determining at least one feature descriptor from the plurality of digital medical images using a trained foundational model based on the prompt, and providing the at least one feature descriptor for each of the plurality of digital medical images as an output.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Patent Application No. 63 / 385,364, filed November 29, 2022, the contents of which are incorporated herein by reference.

[0002] Various embodiments of the present disclosure relate generally to large-scale image processing. More specifically, certain embodiments of the present disclosure relate to systems and methods for large-scale image processing using deep underlying models for inferring metadata from images. [Background technology]

[0003] In general, analyzing large amounts of data using, for example, machine learning systems can be limited by annotation requirements, variations in tissue types within samples, and the like. Furthermore, to fully capture the diversity of complex fields, the model's parameter complexity must be substantial, necessitating very large datasets for training. Training systems for analyzing large, variable, and / or unannotated data can be computationally intensive, especially when the data includes high-resolution images such as those used in computational pathology. Another challenge is the scarcity of data, even when the data does not require exhaustive annotation. Furthermore, even when using supervised or weakly supervised training methods, the ability to generalize across applications can be limited, the availability of clinical labels or manual annotations can be reduced, and training can generalize poorly to long-tail distributions or rare events. Conventional techniques, including those described above, often do not consider the need to analyze large amounts of unannotated data across various modalities. There is a need for systems and / or methods that operate in a pan-cancer and / or pan-tissue manner.

[0004] The foregoing general description and the following detailed description are exemplary and explanatory only and are not intended to limit the present disclosure. The background discussion provided herein is for the purpose of providing an overall context for the disclosure. Unless otherwise indicated herein, the material described in this section is not prior art, is not admitted to be prior art, or is an indication of prior art, by inclusion in this section to the claims in this application. Summary of the Invention [Means for solving the problem]

[0005] According to certain aspects of the present disclosure, methods and systems for generating and modifying foundation models are disclosed. Each of the aspects disclosed herein may include one or more of the features described in connection with any of the other disclosed aspects.

[0006] According to one example of the present disclosure, a method for processing digital medical images to infer metadata from the images may be described. The example method may include receiving a plurality of digital medical images and receiving a prompt requesting a particular type of metadata to be inferred from the plurality of digital medical images, determining at least one feature descriptor from the plurality of digital medical images based on the prompt using a trained foundational model, and outputting or providing the one or more feature descriptors for each of the plurality of digital medical images.

[0007] According to another example of the present disclosure, a method for processing digital medical images and training a foundation model to infer metadata from those images may be described. An example method may include receiving a plurality of digital medical images, generating a plurality of image tokens from the digital medical images, where the image tokens are fixed-size patches, removing a subset of the plurality of image tokens from each of the digital medical images to generate a remaining plurality of image tokens from each of the digital medical images, encoding the remaining plurality of image tokens from each of the digital medical images using an encoder, adding classification tokens to the encoded image tokens, appending a masked token having a positional encoding to each respective encoded image token, and reconstructing the image tokens using a decoder such that the image tokens match pixel values ​​of the original image.

[0008] According to another example of the present disclosure, a system for processing digital medical images to infer metadata from the images may include at least one memory storing instructions and at least one processor configured to execute the instructions to perform operations. The operations may include receiving a plurality of digital medical images and receiving a prompt requesting a particular type of metadata to be inferred from the plurality of digital medical images, using a trained foundational model to determine one or more feature descriptors from the plurality of digital medical images based on the prompt, and outputting or providing the one or more feature descriptors for each of the plurality of digital medical images.

[0009] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various exemplary technologies and, together with the description, serve to explain the principles of the disclosed technology. [Brief explanation of the drawings]

[0011] [Figure 1A] 1 illustrates a block diagram of an exemplary system for generating and modifying foundation models according to one or more techniques. [Figure 1B] 1 illustrates a block diagram of an example system for generating a foundation model according to one or more techniques. [Figure 1C] 1 illustrates a block diagram of an example system for modifying an underlying model according to one or more techniques. [Figure 2] 1 shows an exemplary schematic diagram for training a masked autoencoder (MAE) model according to one or more techniques. [Figure 3] 1 shows an exemplary schematic diagram for training a distillation MAE model according to one or more techniques. [Figure 4] 1 shows an exemplary schematic diagram for training a hierarchical MAE model according to one or more techniques. [Figure 5] 1 shows an exemplary schematic diagram for training a multimodal model according to one or more techniques. [Figure 6] 1 shows an exemplary schematic diagram for training downstream task implementations according to one or more techniques. [Figure 7] FIG. 1 illustrates an exemplary method for processing digital medical images to infer metadata from those images, according to one or more techniques. [Figure 8] FIG. 1 illustrates an exemplary method for processing digital medical images to infer at least one embedding from those images, according to one or more techniques. [Figure 9] FIG. 1 illustrates an example system or device capable of implementing the techniques presented herein according to one or more techniques. DETAILED DESCRIPTION OF THE INVENTION

[0012] Reference will now be made in detail to the exemplary embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts.

[0013] The systems, devices, and methods disclosed herein are described in detail by way of example and with reference to the drawings. The examples discussed herein are examples only and are provided to aid in the explanation of the systems, devices, and methods described herein. Unless expressly designated as essential, none of the features or components shown in the drawings or described below should be construed as essential to particular embodiments of these systems, devices, or methods.

[0014] Also, with respect to any method described, whether or not the method is described with reference to a flowchart, it should be understood that any explicit or implicit order of steps performed in performing the method does not imply that the steps must be performed in the order presented, but rather may be performed in a different order or in parallel, unless otherwise specified or required by context.

[0015] As used herein, the term "exemplary" is used in the sense of "example" as opposed to "ideal." Furthermore, the words "a" and "an" as used herein do not denote a limitation of quantity, but rather denote the presence of one or more of the referenced item.

[0016] preface The technology disclosed herein may describe systems and related methods for processing large-scale images using foundational models. The foundational models may include large-scale deep neural networks trained in a self-supervised manner and adaptable to downstream tasks. For example, millions of slides across hundreds of tissue types may be analyzed by the foundational models for universal whole-slide representations for pattern discovery applied to cancer detection or segmentation, as well as one or more downstream prognostic clinical or biomarker tasks.

[0017] The system and / or method can expose data in a manner that preserves differential privacy, allowing feature sharing for model development without the need to employ federated learning strategies. Additionally, the system can be leveraged to analyze all biomarker signals across all tissue types to discover how multiple tests can be combined to more effectively predict disease, outcome, and / or treatment response.

[0018] Overall system overview The system described herein includes a foundational model trained using a large collection of pathology slides or other samples covering a wide variety of organs, sampling types (i.e., biopsies, resections, aspirates, etc.), and a long tail of rare pathologies and subtypes. The system may receive a plurality of medical images and a prompt to infer metadata from the plurality of medical images. The one or more medical images may include, but are not limited to, whole slide images (WSIs), such as pathology WSIs, radiology images, confocal microscopy images, etc. The prompt may be a request for a specific type of metadata to be inferred from the image. The type of metadata may include, but is not limited to, supplemental medical images from cases, structured diagnostic reports, unstructured free-text reports, genomic data, proteomic data, treatment data, responses, diagnoses, etc.

[0019] Various additional inputs may be received to alter the output of the trained foundation model. In some techniques, the foundation model may receive one or more query constraints. The query constraints may include judgments or hypotheses from a clinician or expert, e.g., a clinician or expert using the system. If one or more query constraints are received, metadata estimates consistent with the query constraints may be output by the foundation model. In some techniques, the trained foundation model may receive free text, which may enable the trained foundation model to output a structural synoptic diagnostic report based on a plurality of digital medical images and the user-entered free-text query constraints. The input free-text content may include at least one of: (1) specific histological details; (2) clinical context involving patient history and other examination modalities; (3) diagnostic criteria, such as application of WHO standards; (4) stain- and marker-specific information; (5) morphological findings; (6) report formatting instructions; (7) input data quality concerns; (8) comparative analysis; and (9) special instructions. The output synoptic diagnostic report can be iteratively adjusted based on new input text prompts provided by the user.

[0020] In some techniques, the foundation model may be run with queries. The foundation model may use one or more queries to request inference of one modality from another. For example, a query may be used to request a diagnostic report from an image. In some techniques, the foundation model may be run without queries. The foundation model may generate general feature descriptors that can be used to train downstream models on more specialized tasks.

[0021] In some techniques, a content-based search system may be used to determine a collection of related digital medical images or cases based on metadata associated with each of the digital medical images or cases. The content-based search system may receive content-based constraints to modify queries for content searches. The content-based constraints may include instructions to include or exclude particular types of metadata or attributes of metadata.

[0022] FIG. 1A illustrates an exemplary system for processing digital medical images and inferring metadata from those images according to one or more techniques. Shown in FIG. 1A is an electronic network 120 that may be connected, for example, via one or more computers, servers, and / or handheld mobile devices, to a physician server 121, a hospital server 122, a clinical trial server 123, a laboratory server 124, and / or a laboratory information system 125. According to exemplary aspects of the present disclosure, network 120 may be connected to a server system 110, which may include, for example, one or more processing devices 100 configured to implement or execute a foundational model generation system 101, and a storage device 109. Foundational model generation system 101 may be configured to process the digital medical images to generate a foundational model. According to exemplary aspects of the present disclosure, downstream foundational model system 102 may be configured to modify the foundational model (e.g., a foundational model generated by foundational model generation system 101 or another system) for at least one downstream application. It should be understood that although the foundational model generation system 101 and the downstream foundational model system 102 are shown as separate systems in FIG. 1A, in other examples, the foundational model generation system 101 and the downstream foundational model system 102 may be subsystems of a larger system.

[0023] The physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125 can create or otherwise acquire pathology slides, digital medical images, clinical reports, free-text reports, immunohistochemistry (IHC) or immunofluorescence slides, computed tomography (CT) scans, genomic data (e.g., gene expression data, genomic mutations, etc.), proteomic data, clinical data, etc. For example, pathology slides can include a wide variety of organs, sampling types (e.g., biopsy, resection, aspirate, etc.), disease states, and / or disease subtypes. In another example, digital medical images can include digital pathology images including whole slide image(s), cytology specimen(s), histopathology specimen(s), slide(s) of cytology specimens, digital image(s) of histopathology specimen slide(s), or any combination thereof, for one or more patients, which can be created or acquired. Additionally or alternatively, digital medical images may include images of other modality types, including digital multiplexed immunofluorescence images, digital multiplexed immunohistochemistry images, magnetic resonance imaging (MRI), computed tomography (CT), x-ray, nuclear medicine imaging, or ultrasound, that may be created or acquired.

[0024] Expression data can include patient-specific or non-patient-specific tumor sequence data, protein expression levels, and / or non-coding RNA expression levels. Expression data can be used for training purposes by both medical professionals (e.g., pathologists, physicians, etc.) and AI systems alike, to improve the predictive accuracy of underlying models, among other tasks. The increasing availability of expression data indicative of specific conditions or diseases increases the variability in expression data, improving the learning capabilities of both medical professionals and AI systems. However, large amounts of expression data are still unavailable for individual genetic mutations in each tumor type, inevitably limiting the amount of variation that can be learned. For example, treating a patient-specific tumor can be challenging due to genotypic differences compared to another patient with the same phenotype but a different genotype.

[0025] Genomic variants may include mutations in individual genes of a given gene complex or signaling pathway, such as the switch / sucrose non-fermenting (SWI / SNF) complex (e.g., ARID1A, ARID1B, ARID2, PBRM1, SMARCA4, and SMARCB1) or the receptor tyrosine kinase (RTK) / Ras / MAP kinase (MAPK) pathway (e.g., ERBB2, ERBB3, ERBB4, SOS1, HRAS, BRAF, MAP2K1, and MAPK1). Clinical data may include age, medical history, cancer treatment history, family history, previous biopsy or cytology information, tumor sequence information, messenger ribonucleic acid (mRNA) expression levels, gene network graphs (pre- and / or post-treatment), overall survival data, progression-free survival with corresponding censored data, 5-year survival rate, drug treatment outcome data, etc.

[0026] The data described herein may be communicated between server system 110 and physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125 in digital and / or electronic form via network 120.

[0027] The server system 110 may include one or more storage devices 109 for storing data received from at least one of the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. For example, the foundational model generated by the foundational model generation system 101 may be stored in one or more data stores, such as the storage devices 109.

[0028] Server system 110 may include processing device 100 to process data stored on storage device 109. Server system 110 may include one or more machine learning tool(s) or functionality. For example, processing device 100 may execute one or more machine learning systems utilized by foundation model generation system 101 and / or downstream foundation model systems 102 according to one or more techniques. In some examples, output of the machine learning systems may be stored on storage device 109 for use in other systems or processes, as described in more detail below. Alternatively, or in addition, the present disclosure (or portions of the systems and methods of the present disclosure) may be performed on a local processing device (e.g., a laptop).

[0029] According to exemplary aspects of the present disclosure, the foundational model generation system 101 may be configured to generate a foundational model. The foundational model may be generated, for example, without annotation, using a large number (e.g., hundreds, thousands, millions, etc.) of slides across a large number (e.g., tens, hundreds, thousands, etc.) of tissue types. According to exemplary aspects of the present disclosure, the downstream foundational model system 102 may be configured to modify at least one foundational model for downstream tasks, applications, etc. The techniques discussed herein may have features, such as the ability to function across various tissues, anomalies, and tasks even with small sample sizes, to operate in a pan-cancer and / or pan-tissue manner, and to further enable improved generalization to new domains (e.g., other scanners, hospitals, etc.).

[0030] 1B illustrates an exemplary system for generating a foundational model system (e.g., foundational model generation system 101) according to an exemplary embodiment of the present disclosure. Foundational model generation system 101 may include a training foundational model generation platform 131 and / or a target foundational model generation platform 135.

[0031] According to one technique, the training basis model generation platform 131 can generate or receive one or more data of training data for the foundation model that are used to generate and train one or more machine learning models that, when implemented, generate a foundation model. According to one technique, the training basis model generation platform 131 can include multiple software modules, including a training data ingestion module 132 and a cross-organizational training data population module 133. The data and / or machine learning system output by the training basis model generation platform 131 can be stored, for example, in the storage device 109, or can be used by other systems, such as the target foundation model generation platform 135.

[0032] According to one aspect, training data ingestion module 132 can create or receive foundational model training data that can be used to generate at least one foundational model. The foundational model training data can be received from any one or combination of server system 110, physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125. The foundational model training data can be obtained from real sources (e.g., humans, animals, etc.) or synthetic sources (e.g., graphics simulators, graphics rendering engines, 3D models, etc.).

[0033] The foundational model training data may include one or more data corresponding to pathology slides, digital medical images, clinical reports, free text reports, IHC or immunofluorescence slides, CT scans, genomic data, proteomic data, clinical data. In some examples, the subset of foundational model training data may overlap among various data, such as pathology slides, digital medical images, clinical reports, free text reports, IHC or immunofluorescence slides, CT scans, genomic data, proteomic data, clinical data, etc. The foundational model training dataset may be stored on a digital storage device, for example, one of storage devices 109.

[0034] The cross-institutional training module 133 may be configured to generate a trained foundational model based on the foundational model training data. As described in further detail herein, the cross-institutional training module 133 may be configured to train the foundational model using any suitable technique(s), e.g., multimodal, annotation-free, etc. In one technique, the cross-institutional training module 133 may be trained to learn at least one relationship between modalities (e.g., clinical reports, free text reports, IHC or immunofluorescence slides, CT scans, genomic data, proteomic data, etc.) and / or to make inferences across various modalities. The foundational model training data may be received by the cross-institutional training module 133 from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, the laboratory information system 125, and / or the training data ingestion module 132. The cross-organizational training module 133 may output, for example, a trained foundational model that may be stored, for example, in the storage device 109 and / or utilized by the target foundational model generation platform 135.

[0035] According to one technique, the target-based model generation platform 135 can include software modules such as a target data ingestion module 136, a cross-tissue module 137, and an output interface 138. According to one aspect, the target-based model generation platform 135 can receive a request to create image features representative of various tissue morphologies and architectures for both healthy and disease states, for example. As discussed in more detail below (see, e.g., FIG. 1C ), these features can be used to train downstream base models for a variety of tasks, such as cancer detection, segmentation, and biomarker identification. The target-based model generation platform 135 can be configured to execute one or more of the base models trained by the training base model generation platform 131. For example, the target-based model generation platform 135 can further train the base model using embeddings (e.g., embeddings generated by the trained base model) to generate an image analysis model. In some techniques, the request can be received from any one or any combination of the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. In some techniques, a request may be automatically received from the downstream foundational model system 102 in response to the downstream foundational model system 102 receiving a request to generate, train, etc., a foundational model and / or modify a trained foundational model.

[0036] According to one aspect, the target data ingestion module 136 can create or receive target data that can be used as input for one or more trained machine learning systems to modify the foundational models. For example, the target data ingestion module 136 can receive digital medical images that can be used as input for one or more trained foundational models. The target data can be received from any one or combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. The target data can be obtained from real sources (e.g., humans, animals, etc.) or synthetic sources (e.g., graphics simulators, graphics rendering engines, 3D models, etc.). The target data ingestion module 136 can create or receive target data. The target data can include at least one of digital medical images, clinical reports, free text reports, IHC and / or immunofluorescence slides, CT scans, genomic data, proteomic data, other medical data, etc. In some examples, the subset of target data can overlap among various data of images and / or clinical data. The target data may be stored on a digital storage device, such as one of the storage devices 109.

[0037] The cross-tissue module 137 may include any suitable foundational model machine learning system, including, but not limited to, graph neural networks, convolutional neural networks, transformer neural networks, etc. The cross-tissue module 137 may execute various foundational models, such as those generated by the training graph generation platform 131. The cross-tissue module 137 may determine at least one relationship between modalities (e.g., between clinical reports, free-text reports, IHC or immunofluorescence slides, CT scans, genomic data, proteomic data, etc.) and / or make inferences between various modalities. For example, the cross-tissue module 137 may synthesize natural language descriptions or structured diagnostic reports from hematoxylin and eosin (H&E) slides, render synthetic IHC based on H&E, and / or predict complete genomic panels. In another example, the model may generate representations of cells, local regions (e.g., patches), entire slide images, and / or groups of slides, etc.

[0038] The output interface 138 may be used to output (e.g., to a screen, monitor, storage device, web browser, etc.) the trained foundational model (e.g., generated by the training foundational model generation platform 131) and / or the modified foundational model (e.g., modified by the target foundational model generation platform 135). According to some techniques, the output interface 138 may output the trained foundational model to the downstream foundational model system 102 for use as input in subsequent processes described below. The foundational models and other data generated or used by the foundational model generation system 101 may be stored in one or more storage devices 109.

[0039] 1C illustrates an exemplary system for modifying at least one foundation model, e.g., for a downstream task, according to an exemplary embodiment of the present disclosure, such as a downstream foundation model system 102. The downstream foundation model system 102 may include a training downstream platform 141 and / or a target downstream platform 145.

[0040] According to one technique, the training downstream platform 141 can include software modules such as a training data ingestion module 142 and a downstream training module 147. According to one aspect, the training data ingestion module 142 can create or receive training data (e.g., foundational model training data) that can be used to train one or more machine learning systems to modify at least one foundational model for a downstream task. For example, the downstream training module 143 can further train a foundational model on embeddings (e.g., embeddings generated by the trained foundational model) to generate an image analysis model. The training data can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. The training data can be obtained from real sources (e.g., humans, animals, etc.) or synthetic sources (e.g., graphics simulators, graphics rendering engines, 3D models, etc.). The training data ingestion module 142 can create or receive datasets of downstream training data. For example, the downstream training data may include embeddings generated by foundational models (e.g., by the cross-tissue module 137) for cancer and / or biomarker detection, cancer and / or biomarker scoring, multi-model diagnosis, prognosis, treatment planning, content-based search, etc. In some examples, subsets of the downstream training data may overlap among various downstream training data.

[0041] In some examples, the training data may be the direct output of one or more machine learning systems (e.g., the underlying model). In other examples, one or more outputs of the machine learning systems may be used as inputs to further processing that effect changes to the underlying model. The downstream training dataset may be stored in a digital storage device, such as one of storage devices 109.

[0042] The downstream training module 143 can use the downstream training data as input to, for example, generate one or more modified foundational models to be used for downstream purposes. In some examples, a third party can generate one or more trained machine learning systems and provide the trained machine learning system(s) to the server system 110 for storage (e.g., in the storage device 109) and / or execution by the downstream foundational model system 102. The downstream training module 143 can train a Transformer, a graph neural network, or any other suitable type of machine learning system to modify (e.g., further train) the foundational model (e.g., a model obtained from the foundational model generation system 101) for a given downstream application. The training prediction module 143 can store the modified foundational model in a database, e.g., in the storage device 109, along with other foundational models, such as the foundational model and the modified foundational model.

[0043] According to one technique, the target downstream platform 145 can include software modules such as a target data ingestion module 146, downstream modules 147, and an output interface 148. The target data ingestion module 146 can receive one or more target inputs, including, but not limited to, a modified foundation model, embeddings (e.g., modality-specific embeddings) generated by a trained foundation model, etc. For example, the target data can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125.

[0044] The target data ingestion module 146 can provide one or more inputs to the downstream module 147 to generate an output via the modified foundation model. The downstream module 147 can execute the various modified foundation models generated by the training downstream platform 141 to generate at least one output for a downstream task.

[0045] According to one aspect, downstream module 147 can receive a request to execute one or more machine learning systems (e.g., modified foundational models) trained by training downstream platform 141 to predict an output of at least one downstream task. For example, the request can be received from any one or combination of server system 110, physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125. In another example, the request can be automatically generated by downstream foundational model system 102 in response to detection of an output from another system, such as, for example, from foundational model generation system 101. In some implementations, downstream module 147 can be configured to automatically predict an output for at least one downstream task based, for example, on input target data and / or modified foundational models.

[0046] An output interface 148 may be used to output (eg, to a screen, monitor, storage device, web browser, etc.) the predicted output for at least one downstream task.

[0047] Exemplary Techniques The base model (e.g., trained by the training base model generation platform 131) may be trained in a self-supervised manner using at least a plurality of medical images. Exemplary training methods may include masked autoencoder (MAE) training, distilled MAE training, hierarchical MAE training, hierarchical distilled MAE training, contrastive methods, multimodal training, etc.

[0048] Masked Autoencoder (MAE) model The MAE-based model may follow an auto-encoding scheme that reconstructs the original data given partially masked input data. The MAE-based model may have an encoder, such as a Vision Transformer (ViT) encoder, that maps observed data to a latent space, and a decoder, such as a ViT decoder, that reconstructs the original data from the latent space. The MAE-based model may operate based on an asymmetric design that allows the encoder to operate on partial observed data without mask tokens and the decoder to reconstruct the complete data from the latent space and mask tokens.

[0049] MAE Training As shown in Figure 2, an exemplary method for training an MAE model may include receiving a plurality of inputs 205, such as medical images, one or more prompts, etc. As shown in Figure 2, the plurality of medical images may be divided into a plurality of fixed-size patches (image tokens 210). A subset of the image tokens 215 may be intentionally removed from the plurality of medical images, leaving a remaining plurality of image tokens.

[0050] An encoder, such as the ViT encoder 220, may output one or more encoded image tokens 225 based on the remaining image tokens. In some techniques, a classification token 230 may be added to the beginning of the encoded image token 225. The classification token 230 may be a network-specific numeric vector and may be used to summarize the image tile representation. In some techniques, a masked token 235 may be appended with a positional encoding applied to each encoded image token.

[0051] The masked tokens 235 and optional classification tokens 230 may be provided to a ViT decoder 240. The ViT decoder 240 may be used to reconstruct the original image tokens (e.g., image tokens 210) so that they match or substantially match the original image pixel values. The ViT decoder 240 may generate at least one output, such as a reconstructed token 245, a reconstructed tile 250, etc. In some techniques, the network may be optimized by applying an L2 image reconstruction loss only to the removed visual tokens (e.g., masked token(s) 253) (e.g., see step 255 of FIG. 2). In some techniques, the ViT encoder 220 may be trained to match a subset of the image tokens 215 with the reconstructed tokens 245, thereby allowing the network to learn a global structure (see step 260).

[0052] Distillation MAE model The distilled MAE model, like any MAE model, may include an encoder and a decoder (e.g., ViT decoder 321). The encoder of the distilled MAE model further includes a student encoder (e.g., student ViT encoder) and a teacher encoder (e.g., teacher ViT encoder). The distilled MAE model may have any of the features or training steps discussed in connection with any of the other models described herein, unless otherwise specified.

[0053] Distillation MAE Training An exemplary method for training a distilled MAE model is described in FIG. 3. The distilled MAE model can receive multiple inputs 305, such as medical images and one or more prompts. For each of the multiple inputs 305, a hue-saturation-value (HSV)-based foreground detection algorithm (not shown) can extract image tokens 310 of a specific size, e.g., 224 x 224, from the tissue of interest. Glass regions can be discarded. During training, the inputs 305 can be randomly sampled from the entire database using balanced metadata sampling. Referring to FIG. 3, a student encoder, e.g., a student ViT encoder 320, can divide the multiple inputs 305 (e.g., multiple medical images) into multiple fixed-size patches (image tokens 325). A subset of the image tokens 325 can be intentionally removed from the multiple medical images, and the student ViT encoder 320 outputs one or more masked tokens 337.

[0054] Because the classification tokens 330 may not have been explicitly trained, a multilayer perceptron (MLP) similar to self-distillation without labels (DINO) can be applied on top of the classification tokens 330 to project the tile embedding. The student ViT encoder 320 with masked tokens 337 can be trained with distillation loss to predict the classification tokens 330 predicted by the teacher ViT encoder 335 fed with all image tokens 310 without masking. The teacher ViT encoder 335 can be updated using a running average of the student ViT encoder 320 at the end of each training batch. The distillation loss can be combined with the image reconstruction loss using a weighted sum to encourage the network to learn local-scale patterns while retaining the ability to summarize the content of the image tiles. The projection head 340 can convert the classification tokens 330 from a one-dimensional vector to another one-dimensional vector of a different size, which can train the student ViT encoder 320 to match the distribution from the teacher ViT encoder 335 and force the network to learn a global structure (see step 345).

[0055] Hierarchical MAE model The hierarchical MAE model may include a tile-level ViT encoder that feeds either an MAE or a distilled MAE. A first set of tiles may be fed into the tile-level (e.g., pre-trained) ViT encoder to output class tokens. The class tokens may then be fed into either an MAE or a distilled MAE to output reconstructed tokens.

[0056] DINOv2 model The DINOv2-based model may utilize a self-distillation mechanism to implement a novel self-supervised learning architecture. At the heart of the DINOv2-based model are two Vision Transformers (ViTs): a student network and a teacher network. These networks may process augmented views of the same image in parallel, with the student network tasked with predicting the output of the teacher network. The teacher network may be independently updated as an exponential moving average of the student network's weights. This architecture enables DINOv2 to learn by maintaining consistency across different image views, eliminating the need for labeled data in its training process.

[0057] DINOv2 Training DINOv2 training may be similar to MAE training (described in more detail below), which may involve a diverse array of medical images. Each image may undergo a set of augmentations to create multiple and / or varied views. These augmented views may be divided into a series of fixed-size patches and / or may be similar to image tokens in the MAE method. The student network may receive a subset of these patches, and the teacher network may receive either the same or a different subset, which may include patches that may be withheld from the student network to create knowledge gaps.

[0058] For example, a student network with a Vision Transformer (ViT) architecture may process the patches to generate a set of encoded representations. A teacher network operating in a non-gradient fashion may update its weights as an exponentially moving average of the student's parameters over time. This setup can generate continuously evolving goals for the student network, which may encourage the student network to learn richer representations as it attempts to predict the teacher's outputs.

[0059] During training, both networks can utilize a self-attention mechanism inherent in ViT, which may enable the networks to focus on the informative portions of the input data. The predictions of the student network may be compared to the outputs of the teacher network, and the discrepancy between them may be minimized using a distillation loss function. This loss may be calculated based solely on predictions of the encoded representations, which may encourage the student to mimic the behavior of the teacher. As training progresses, the student network may be encouraged to learn both local and global structures in the data, which may facilitate a comprehensive understanding of visual content independent of manual annotations or labels.

[0060] Hierarchical MAE (or Slide-Level Foundation Model) Training An exemplary method for training a hierarchical foundation model is shown in FIG. 4. As shown, a tile-level ViT encoder (e.g., a pre-trained ViT encoder 415) may be trained to output tile-level class tokens 420 based on multiple inputs 405 (e.g., image tiles), e.g., using methods described herein. The tile-level ViT encoder (e.g., a pre-trained ViT encoder 415) may be trained using either MAE training or distillation training, as described herein. Once the tile-level ViT encoder is trained, the predicted tile-level class tokens 420 may be grouped into a grid of tiles, e.g., 16x16, and fed to a grid-level ViT encoder (e.g., a ViT encoder 425), as described above. The grid-level ViT encoder may have a similar architecture to the tile-level ViT encoder to represent the spatial organization of tiles, similar to a hierarchical image pyramid transformer (HIPT). The same distilled MAE combination loss may be applied to the class tokens 420 instead of the image tokens 430.

[0061] The masked tokens 445 and optional classification tokens 440 may be provided to a ViT decoder 450. The ViT decoder 450 may be used to reconstruct the original image tokens (e.g., image tokens 410) to match or substantially match the original image pixel values. The ViT decoder 450 may generate at least one output, e.g., a reconstructed token 455, a reconstructed tile / image, etc. In some techniques, the ViT encoder 425 may be trained to match a subset of the class tokens 420 with the reconstructed tokens 455, and may force the network to learn a global structure (see step 460).

[0062] Grid-level class tokens from the same slide can be aggregated using an aggregation network (not shown) to obtain a representation of the entire slide. The aggregation network can be a similar ViT network or a traditional MLP / long short-term memory (LSTM) / convolutional neural network (CNN) network. The slide-level aggregation network can be pre-trained using at least one of three main methodologies: (1) self-supervised learning, (2) image-text pre-training, and / or (3) supervised training with slide- or group-level ground truth labels.

[0063] Self-supervised learning can involve training an aggregation network to understand and interpret complex patterns within slide images without explicit labeling. By analyzing the inherent structure and features of slides, the model can learn to identify subtle nuances essential for accurate interpretation in digital pathology. This self-learning approach can foster a deeper understanding of slide images and lay a solid foundation for further, more specialized training.

[0064] Following the principles of models such as Contrastive Language-Image Pretraining (CLIP) or Contrastive Captioners are Image-Text Foundation Models (COCA), the aggregation network may undergo an image-text pretraining phase. In this case, the aggregation network learns to associate text descriptions, such as diagnostic reports or molecular test reports, with their corresponding slide images. This training phase may enable the model to recognize visual features and / or understand contextual information presented in text format. Such dual-modality training can improve the model's ability to process and interpret complex medical images in conjunction with associated text data, such as clinical notes or research findings.

[0065] To improve the accuracy of the model and adapt it to the specific requirements of digital pathology, supervised training using slide- or group-level ground truth labels can be employed. Ground truth labels for slides or groups of slides can be used in this stage. These labels can be extracted automatically using advanced large-scale language models (LLMs), which can analyze text data associated with slides to generate accurate and / or contextually relevant labels. The LLMs can be provided with detailed system prompts, including detailed descriptions of the ground truth labels and / or the expected extracted format. Ground truth labels can be obtained by feeding text reports to an LLM agent and / or analyzing texture responses using a heuristic pipeline. This approach can train the model on visual data and / or substantially accurately labeled datasets, thereby improving its predictive accuracy and / or reliability.

[0066] Multimodal Model The foundational model may also be a multimodal model and / or be trained in a multimodal manner, incorporating additional relevant data such as clinical reports, free-text reports, hematoxylin and eosin (H&E) stains, immunohistochemistry (IHC) slides, immunofluorescence slides, CT scans, genomic data, and proteomic data. Models trained in this manner can learn relationships between various modalities, which may enable inferences to be drawn between them. For example, the model cloud can synthesize natural-language descriptions or structured diagnostic reports from H&E slides, render synthetic IHC based on H&E, or predict complete genomic panels. The model can also generate representations of cells, local regions (e.g., patches), WSIs, groups of slides, and so on.

[0067] Multimodal Training An exemplary method for training a multimodal model is described below. The multimodal model may perform a search to find slides with similar features. For example, if trained to map WSIs to proteomes, the multimodal model may generate a proteome for a particular patient given only the WSIs, or it may use the proteome to find WSIs similar to the proteome. The multimodal input may have embeddings for each modality, and the combination of modalities may then be used for downstream tasks such as treatment recommendations.

[0068] A multimodal model may be trained to analyze data based on modality, as shown in Figure 5. In an example where the multimodal model may analyze IHC and H&E images, the multimodal model may be trained to match H&E image 1 (505) with IHC image 1 (510) based on a shared tissue type, H&E image 2 (506) with IHC image 2 (511) based on a shared tissue type, H&E image 3 (507) with IHC image 3 (512) based on a shared tissue type, H&E image 4 (508) with IHC image 4 (513) based on a shared tissue type, etc.

[0069] Each piece of data may be input to a modality-specific encoder. For example, H&E images (e.g., H&E images 505, 506, 507, 508, etc.) may be input to an H&E encoder 520, and IHC images (e.g., IHC images 510, 511, 512, 513, etc.) may be input to an IHC encoder 515. The modality-specific encoders may output encodings, such as an H&E encoding 530 and an IHC encoding 525. The encodings may be used to learn relationships between various modalities. In some techniques, a multimodal model may be trained to maximize the diagonal entries 535 and minimize the off-diagonal entries 540 of the encodings.

[0070] The training of the multimodal model and the modality-specific encoders may be done simultaneously, or the multimodal model may be trained after the modality-specific encoders.

[0071] Implementation of downstream tasks In general, the foundational models can be adapted for a wide range of downstream tasks, significantly improving functionality such as cancer detection, segmentation, biomarker identification, or morphology-based image retrieval. The foundational models can be further optimized or fine-tuned independently or collaboratively for downstream tasks. For example, image analysis models can be trained on embeddings generated by the foundational models, including cancer or biomarker detection, cancer or biomarker scoring, multi-model diagnosis, prognosis, treatment planning, content-based retrieval, etc.

[0072] In one technique, a multimodal model can be configured to generate embeddings for each modality, and the combination of modalities can then be configured to generate downstream tasks, such as treatment recommendations. In another technique, the foundational model can be configured for biomarker-related tasks. For example, given a WSI, downstream task implementations can be configured to detect the presence or score of biomarkers such as human epidermal growth factor receptor 2 (HER2), Kirsten rat sarcoma virus (KRAS), microsatellite instability (MSI), epidermal growth factor receptor (EGFR), etc. In another technique, the foundational model can be configured for clinically relevant tasks. For example, given a WSI, downstream task implementations can be configured to detect the presence, subtype, grade, and site of cancer in the WSI. In another technique, the foundational model can be configured for survival / outcome-related tasks. For example, given a WSI, downstream task implementations can be configured to predict a patient's overall survival, disease-free survival, and / or prognosis-free survival. In another technique, the foundational model can also be configured for treatment outcome-related tasks. For example, given a WSI, an implementation of a downstream task can be configured to predict the immuno-oncological response to a particular treatment. In any of the above examples, given a certain number of WSI image / label pairs, the WSI image / label pairs can be used to train a second model using the outputs of the base model as inputs:

[0073] In some techniques, downstream task implementations may include predicting output targets not included within the metadata type based on one or more feature descriptors for each of the plurality of digital medical images, which may include any combination of learning markers for drug response, building models that replicate existing biomarkers, learning new biomarkers from test data or other gold standard indicators, or predicting additional disease states or diagnoses.

[0074] An implementation of a content-based search system using embeddings from a digital pathology foundation model may include several key steps to optimize performance and storage efficiency. In a first step, embeddings of image tiles, which may be derived from whole slide images (WSIs), may be stored in a vector database. This database may be designed for fast query execution and / or may allow for efficient retrieval of similar image tiles based on their embeddings. To optimize storage, the system may identify and / or store only a few representative tiles from each WSI. This approach may reduce the storage footprint while maintaining the ability to search a comprehensive range of similar tiles.

[0075] When a user selects a region of interest (ROI) on the WSI, the system may process the ROI by extracting embeddings of tiles enclosed within it. Each of these embeddings may serve as a separate query against a vector database. The system may, for example, search the database for slides and / or ROIs with embeddings similar to those of the query tile. To rank the returned slides and ROIs, the system may calculate similarity scores. These scores may measure how closely tiles found in the searched slides match the bag of query tiles in the user-selected ROI. The system may support searching for the most similar tiles and / or the least similar tiles by manipulating similarity statistics.

[0076] This method may allow users to search for the most relevant WSIs and / or regions thereof based on, for example, a user-specific query. The system's flexibility may allow users to select any shape, size, and / or magnification of the ROI, ensuring a highly customizable search experience. The foundational model embedding-based approach may be applicable to various medical image types, e.g., H&E, IHC, and CT images, and may be configured to process data sources ranging from individual WSIs to large, cross-institutional data lakes of WSIs. This versatility may improve the usability of the search system and / or make it an essential tool in digital pathology for tasks such as contextualizing downstream task predictions using relevant text and image references. Additionally, it may provide an interface for domain experts to examine the learned information stored within the foundational model embedding, enabling the exploration of new research activities and the formulation of novel downstream tasks.

[0077] Implementation of downstream tasks can reduce artificial intelligence development time and costs, improve performance, and enable new product development, especially for tasks involving poorly labeled data for training deep neural networks. Additionally, embedding generation can be used to search WSI's data lake, improving search time and use of data stored within the data lake.

[0078] Downstream task implementation training Exemplary methods for training implementations of downstream tasks (e.g., modifying the trained underlying model) are described below.

[0079] 6, input(s) 605, such as image tiles, may be decomposed into image tokens 610 using methods described herein. The image tokens 610 may be input to a ViT encoder 620, which may output encoded image tokens 625 using methods described herein. In some implementations, the ViT encoder 620 may be a fixed encoder or a low-rank (LR) encoder. The encodings may be input to a classifier 630, which may be trained to classify the encodings (e.g., encoded image tokens 625) and output predicted labels 635. The predicted labels may be trained to match ground truth labels 645 (step 640).

[0080] 7 illustrates an exemplary method 700 for processing digital medical images and inferring metadata from those images, according to one or more techniques. The metadata may include any combination of supplemental medical images, structured diagnostic reports, unstructured free text reports, genomic data, proteomic data, treatment data, responses, diagnoses, etc.

[0081] A plurality of digital medical images may be acquired in step 702. For example, the plurality of digital medical images may be received from a physician server 121, a hospital server 122, a clinical trial server 123, a laboratory server 124, and / or a laboratory information system 125, etc.

[0082] Optionally, at least one query constraint or free text, or both, may be received in step 704. In some techniques, the query constraint may include judgments or hypotheses from a clinician or expert.

[0083] A prompt may be received at step 706. In some techniques, the prompt may request a particular type of metadata to be inferred from the plurality of digital medical images.

[0084] At step 708, at least one feature descriptor may be determined from the plurality of digital medical images based on the prompts. In some techniques, the at least one feature descriptor may be determined using a foundation model. The foundation model may be trained using techniques described herein.

[0085] Optionally, at step 710, a collection of related digital images or cases may be identified. This collection of related digital images or cases may be based on metadata associated with each of the digital medical images or cases. In some techniques, the collection of related digital images or cases may be identified using a content-based search system. In some techniques, content-based constraints may be received. Content-based constraints may include instructions to include or exclude certain types of metadata, or attributes of that metadata, in the content search query.

[0086] Optionally, at step 712, at least one output target may be predicted. The output target may be predicted using a downstream task model. The downstream task may be trained using techniques described herein. In some techniques, the output target may not be included within a metadata type based, for example, on one or more feature descriptors of each of the multiple digital medical images. The output target may include any combination of learned markers for drug response, building a model that replicates existing biomarkers, learning new biomarkers from test data or other gold standard indicators, or predicting additional disease states or diagnoses.

[0087] At least one feature descriptor, metadata estimate, and / or structured synoptic diagnostic report may be output for each of the plurality of digital medical images at step 714. In some techniques, the metadata estimate may be consistent with one or more query constraints. In some techniques, the structured synoptic diagnostic report may be based on the plurality of digital medical images and free text.

[0088] 8 illustrates an example method 800 for processing digital medical images and inferring at least one embedding from those images according to one or more techniques. At step 802, a plurality of digital medical images may be acquired. For example, the plurality of digital medical images may be received from a physician server 121, a hospital server 122, a clinical trial server 123, a laboratory server 124, and / or a laboratory information system 125, etc.

[0089] At step 804, multiple embeddings may be obtained. In some techniques, the multiple embeddings may be obtained from a base model. The base model may be trained for downstream tasks using techniques described herein. Multiple embeddings may be obtained for each of multiple modalities.

[0090] Optionally, at least one output target may be predicted in step 806. In some techniques, at least one output may be predicted using the trained foundational model (e.g., the trained downstream foundational model). The at least one output may include detecting the presence and / or score of a biomarker (e.g., HER2, KRAS, MSI, EGFR, etc.), detecting the presence, subtype, grade, site of cancer in the WSI, predicting the overall survival, disease-free survival, and / or prognostic survival of a patient, predicting response to a particular treatment, etc.

[0091] 9 illustrates an exemplary computer-implemented system or device 900 capable of executing the techniques presented herein. The device 900 may include a central processing device (CPU) 920. The CPU 920 may be any type of processor device, including, for example, any type of special-purpose or general-purpose microprocessor device. As will be appreciated by those skilled in the art, the CPU 920 may also be a single processor in a multi-core / multi-processor system operating alone or in a cluster of computing devices operating in a cluster or server farm. The CPU 920 may be connected to a data communications infrastructure 910, such as, for example, a bus, a message queue, a network, or a multi-core message passing scheme.

[0092] The device 900 may also include a main memory 940, such as, for example, random access memory (RAM), and may also include a secondary memory 930. The secondary memory 930, such as, for example, read-only memory (ROM), may be, for example, a hard disk drive or a removable storage drive. Such removable storage drives may include, for example, floppy disk drives, magnetic tape drives, optical disk drives, flash memory, etc. The removable storage drive in this example reads from and / or writes to a removable storage unit in a well-known manner. The removable storage device may include a floppy disk, magnetic tape, optical disk, etc., which are read from and written to by the removable storage drive. As will be appreciated by those skilled in the art, such removable storage units typically include computer-usable storage media having computer software and / or data stored thereon.

[0093] In alternative implementations, secondary memory 930 may include similar means for allowing computer programs or other instructions to be loaded into device 900. Examples of such means may include program cartridges and cartridge interfaces (such as those found in video game devices), removable memory chips (such as EPROMs or PROMs) and associated sockets, and other removable storage units and interfaces that allow software and data to be transferred from removable storage units to device 900.

[0094] Device 900 may also include a communications interface (COM) 960. Communications interface 960 allows software and data to be transferred between device 900 and external devices. Communications interface 960 may include a modem, a network interface (such as an Ethernet card), a communications port, a PCMCIA slot and card, or the like. The software and data transferred via communications interface 960 may be in the form of signals, which may be electronic, electromagnetic, optical, or other signals receivable by communications interface 960. These signals may be provided to communications interface 960 via a communications path in device 900, which may be implemented using, for example, wire or cable, fiber optics, a telephone line, a cellular phone link, an RF link, or other communications channel.

[0095] The hardware elements, operating systems, and programming languages ​​of such equipment are conventional in nature and are assumed to be sufficiently familiar to those skilled in the art. Device 900 may also include input / output ports 950 for connecting input / output devices such as a keyboard, mouse, touchscreen, monitor, display, etc. Of course, various server functions may be implemented in a distributed manner on multiple similar platforms to distribute the processing load. Alternatively, the server may be implemented by appropriate programming of a single computer hardware platform.

[0096] Throughout this disclosure, references to components or modules generally refer to items that can be logically grouped together to perform a function or group of related functions. Like reference numbers are generally intended to refer to the same or similar components. Components and / or modules may be implemented in software, hardware, or a combination of software and / or hardware.

[0097] The tools, modules, and / or functions described above may be executed by one or more processors. The "storage" type of medium may include any or all of the tangible memory of a computer, processor, etc., or associated modules such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory storage for software programming at any time.

[0098] The software may be communicated over the Internet, a cloud service provider, or other telecommunications network. For example, the communication may enable the software to be loaded from one computer or processor to another. As used herein, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution, unless limited to non-transitory, tangible "storage" media.

[0099] The foregoing general description is exemplary and explanatory only and is not intended to limit the present disclosure. Other embodiments of the invention will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure. It is intended that the specification and examples be considered as exemplary only.

Claims

1. 1. A computer-implemented method for processing digital medical images to infer metadata from those images, comprising: receiving a plurality of digital medical images; receiving a prompt, the prompt being a request for a particular type of metadata to be inferred from the plurality of digital medical images; determining at least one feature descriptor from the plurality of digital medical images using a trained foundational model based on the prompt; and providing as output the at least one feature descriptor for each of the plurality of digital medical images.

2. 10. The method of claim 1, wherein the plurality of digital medical images comprises at least one of whole slide images (WSIs), hematoxylin and eosin (H&E) stains, immunohistochemistry (IHC) slides, immunofluorescence slides, or CT scans.

3. 10. The computer-implemented method of claim 1, wherein the metadata includes any combination of supplemental medical images, structured diagnostic reports, unstructured free text reports, genomic data, proteomic data, treatment data, responses, diagnoses, and the like.

4. receiving at least one query constraint, the query constraint including a judgment or hypothesis from a clinician or expert; The computer-implemented method of claim 1 , further comprising: providing as output from the trained foundation model metadata estimates that are consistent with the at least one query constraint.

5. receiving free texts; 10. The computer-implemented method of claim 1, further comprising: providing as output from the trained foundational model a structured synoptic diagnostic report based on the plurality of digital medical images and free text.

6. 10. The computer-implemented method of claim 1, further comprising using a content-based search system to determine a collection of related digital medical images or cases based on the metadata associated with each of the digital medical images or cases.

7. 7. The computer-implemented method of claim 6, further comprising receiving a content-based constraint, the content-based constraint including instructions to include or exclude particular types of metadata, or attributes of that metadata, from a content search query.

8. 2. The computer-implemented method of claim 1, further comprising: using a downstream task model to determine an output target not included within a metadata type based on the at least one feature descriptor for each of the plurality of digital medical images.

9. 9. The computer-implemented method of claim 8, wherein the output targets include any combination of learned markers for drug response, building models that replicate existing biomarkers, learning new biomarkers from test data or other gold standard indicators, or predicting additional disease states or diagnoses.

10. 1. A method for processing digital medical images to train a foundation model to infer metadata from those images, comprising: receiving a plurality of digital medical images; generating a plurality of image tokens from the digital medical image, the image tokens being fixed size patches; removing a subset of the plurality of image tokens from each of the digital medical images to generate a remaining plurality of image tokens from each of the digital medical images; encoding the remaining plurality of image tokens from each of the digital medical images using an encoder; adding classification tokens to the encoded image tokens; adding a masked token having a positional encoding to each respective encoded image token; using a decoder to reconstruct the image tokens so that they match pixel values ​​of the original image.

11. The method of claim 10 , wherein the encoder is a Vision Transformer (ViT) encoder.

12. The method of claim 10 , wherein the classification tokens are network-specific numeric vectors that summarize image tile representations.

13. The method of claim 10, wherein the decoder is a ViT decoder.

14. The method of claim 13 , wherein the ViT decoder can be optimized using an L2 image reconstruction loss applied to the masked tokens.

15. The method of claim 10 , further comprising training the encoder and decoder to match the image tokens and the masked reconstructed tokens.

16. 1. A system for processing digital medical images to infer metadata from those images, comprising: at least one memory storing instructions; and at least one processor configured to execute the instructions to perform operations, the operations comprising: receiving a plurality of digital medical images; receiving a prompt, the prompt being a request for a particular type of metadata to be inferred from the plurality of digital medical images; determining at least one feature descriptor from the plurality of digital medical images using a trained foundational model based on the prompt; and providing as an output the at least one feature descriptor for each of the plurality of digital medical images.

17. 17. The system of claim 16, wherein the plurality of digital medical images comprises at least one of whole slide images (WSI), hematoxylin and eosin (H&E) stains, immunohistochemistry (IHC) slides, immunofluorescence slides, or CT scans.

18. 17. The system of claim 16, wherein the metadata includes any combination of supplemental medical images, structured diagnostic reports, unstructured free text reports, genomic data, proteomic data, treatment data, responses, diagnoses, and the like.

19. 17. The system of claim 16, wherein the operations further include using a downstream task model to determine an output target not included within a metadata type based on the at least one feature descriptor for each of the plurality of digital medical images.

20. 20. The system of claim 19, wherein the output targets include any combination of learned markers for drug response, building models that replicate existing biomarkers, learning new biomarkers from test data or other sound indicators, or predicting additional disease states or diagnoses.