Image processing of biomedical images using machine learning models for rapid screening

An AI-based two-step method for EGFR mutation prediction from WSIs addresses the limitations of current testing by providing rapid, accurate, and cost-effective EGFR mutation assessment, enhancing treatment decisions and tissue preservation.

WO2026060029A1PCT designated stage Publication Date: 2026-03-19MEMORIAL SLOAN KETTERING CANCER CENT +2

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current methods for EGFR mutation testing in lung cancer are costly, time-consuming, and not widely accessible due to high infrastructure requirements, leading to suboptimal treatment decisions for patients.

Method used

A two-step AI-based approach using a vision transformer for patch-level feature extraction followed by multiple instance learning (MIL) for slide-level prediction, leveraging weakly supervised learning and self-supervised contrastive learning to analyze whole slide images (WSIs) for rapid EGFR mutation assessment.

Benefits of technology

The approach achieves accurate and rapid EGFR mutation prediction in 2 minutes per slide, reducing the need for costly and time-consuming next-generation sequencing, and optimizing tissue utilization for comprehensive genetic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045795_19032026_PF_FP_ABST
    Figure US2025045795_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Presented herein are systems and methods for classifying biomedical images for executing operations. A computing system may identify a biomedical image of a slide with a biological sample obtained from a subject at risk of a condition; apply the biomedical image to a machine learning model; generate, based on applying the biomedical image to the ML model, a classification corresponding to the biomarker associated with the condition in biological sample on the slide; and execute an operation with respect to the slide for testing of the biological sample, in accordance with the classification for the biomedical image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Atty. Dkt. No.: 115872-3307 IMAGE PROCESSING OF BIOMEDICAL IMAGES USING MACHINE LEARNING MODELS FOR RAPID SCREENING CROSS-REFERENCE TO RELATED APPLICATIONS The present application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 693,412, filed September 11, 2024, which is incorporated herein by reference in its entirety. BACKGROUND A computing device may store and maintain data on a database. Upon receipt of a query, the computing device may search for data corresponding to the data on the database. SUMMARY Aspects of the present disclosure are directed to systems and methods of classifying biomedical images for executing operations. One or more processors may identify a biomedical image of a slide with a biological sample obtained from a subject at risk of a condition. The one or more processors may apply the biomedical image to a machine learning (ML) model. The ML model may be established using a plurality of example biomedical images, each of the plurality of example biomedical images labeled with a respective indication of one of a presence or an absence of a biomarker associated with the condition. The one or more processors may generate, based on applying the biomedical image to the ML model, a classification corresponding to the biomarker associated with the condition in biological sample on the slide. The one or more processors may execute an operation with respect to the slide for testing of the biological sample, in accordance with the classification for the biomedical image. In some embodiments, the one or more processors may execute the operation to perform rapid testing using at least a portion of the biological sample on the slide, responsive to the classification indicating an uncertainty of the presence. In some embodiments, the one or -1- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 more processors may execute the operation to skip at least one of rapid testing or next-generation sequencing (NGS) testing using the biological sample on the slide, responsive to the classification indicating the absence of the biomarker in the biological sample. In some embodiments, the one or more processors may execute the operation to perform NGS testing using the biological sample on the slide, responsive to the classification indicating the presence of the biomarker in the biological sample. In some embodiments, the one or more processors may identify a test result indicating one of the presence or the absence of the biomarker associated with the condition in the biological sample in slide, based on performing the operation with respect to the slide for the testing, wherein the operation comprises at least one of rapid testing or NGS testing. The one or more processors may provide, for presentation via a user interface, an output based on the test result of the testing of the biological sample of the slide. In some embodiments, the one or more processors may determine that the subject is a candidate for administration of therapy for the condition, responsive to the test result indicating the presence of the biomarker. The one or more processors may provide the output to indicate that the subject is subject is the candidate for the administration of therapy. In some embodiments, the subject may be administered with a therapeutic effective amount of the therapy for the condition. The condition may include cancer, and wherein the therapy comprises chemotherapy or targeted therapy for the cancer. In some embodiments, the one or more processors may determine that the subject is a non-candidate for administration of therapy for the condition, responsive to the test result indicating the absence of the biomarker. The one or more processors may provide the output to indicate that the subject is subject is the non-candidate for the administration of therapy. The subject may be withheld from the administration of the therapy for the condition, subsequent to provision of the output. In some embodiments, the one or more processors may determine, based on applying the biomedical image to the ML model, a value indicating a likelihood of the presence or the absence of the biomarker associated with the condition in the biological sample on the slide. The one or more processors may generate, based on a comparison of the value with a threshold, the classification to indicate one of the presence, the absence, or an uncertainty. -2- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 In some embodiments, the one or more processors may generate a plurality of patches using the biomedical image, each patch of the plurality of patches comprising a respective portion of the biomedical image. The one or more processors may select from the plurality of patches, a subset of patches corresponding to a region of interest (ROI) in the biomedical image. The one or more processors may apply the subset of patches to the ML model. In some embodiments, the one or more processors may apply a subset of patches of the biomedical image to the ML model The ML model may include a patch encoder configured to generate, for each patch of the subset of patches, a respective feature vector of a plurality of feature vectors; and a classifier configured to generate, using the plurality of feature vectors for the subset of patches, the classification corresponding to the biomarker associated with the condition in biological sample. In some embodiments, the ML model may be established by: identifying, from the plurality of example biomedical images, an example biomedical image of a respective slide with a respective biological sample; applying the example biomedical image to the ML model to generate a respective classification indicating one of the presence or the absence of the biomarker associated with the condition; determining a loss metric based on a comparison between the respective classification and the respective indication associated with the example biomedical image; and updating at least one of a plurality of weights of at least a portion of the ML model in accordance with the loss metric. In some embodiments, the one or more processors may receive the biomedical image in accordance with at least one of a plurality of imaging modalities, wherein the plurality of imaging modalities comprises whole slide imaging (WSI) or immunohistochemistry (IHC) imaging. The biological sample may include tissue obtained from an anatomical site associated with the condition. In some embodiments, the biomarker may include at least one of: AKT1, ALK, APC, AR, ARAF, ARID1A, ARID2, ATM, B2M, BCL2, BCOR, BRAF, BRCA1, BRCA2, CARD11, CBFB, CCND1, CDH1, CDK4, CDKN2A, CIC, CREBBP, CTCF, CTNNB1, DICER1, DIS3, DNMT3A, EGFR, EIF1AX, EP300, ERBB2, ERBB3, ERCC2, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXA1, FOXL2, FOXO1, FUBP1, HER2, -3- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 GATA3, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HRAS, IDH1, IDH2, IKZF1, INPPL1, JAK1, KDM6A, KEAP1, KIT, KNSTRN, KRAS, MAP2K1, MAPK1, MAX, MED12, MET, MLH1, MSH2, MSH3, MSH6, MTOR, MYC, MYCN, MYD88, MYOD1, NF1, NFE2L2, NOTCH1, NRAS, NTRK1, NTRK2, NTRK3, NUP93, PAK7, PDGFRA, PIK3CA, PIK3CB, PIK3R1, PIK3R2, PMS2, POLE, PPP2R1A, PPP6C, PRKCI, PTCH1, PTEN, PTPN11, RAC1, RAF1, RB1, RET, RHOA, RIT1, ROS1, RRAS2, RXRA, SETD2, SF3B1, SMAD3, SMAD4, SMARCA4, SMARCB1, SOS1, SPOP, STAT3, STK11, STK19, TCF7L2, TERT, TGFBR1, TGFBR2, TP53, TP63, TSC1, TSC2, U2AF1, VHL, or XPO1. In some embodiments, the condition may include cancer. The cancer may include at least one of carcinoma, sarcoma, hematopoietic cancer, adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, neuroblastoma, non- Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vascular tumor, and metastases thereof. BRIEF DESCRIPTION OF THE DRAWINGS The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which: FIG. 1A is a block diagram illustrating a multi-step process for predicting epidermal growth factor receptor (EGFR) mutation status, in accordance with an illustrative embodiment. -4- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 FIG. 1B is a block diagram illustrating an end-to-end training process for predicting EGFR mutation status, in accordance with an illustrative embodiment. FIG. 2 is a flow diagram illustrating a method for EGFR prediction, in accordance with an illustrative embodiment. FIG. 3A is a block diagram of a workflow for processing whole-slide images (WSI), in accordance with an illustrative embodiment. FIG. 3B is a block diagram comparing workflows for processing WSI, in accordance with an illustrative embodiment. FIG. 3C is a block diagram illustrating WSI processing using artificial intelligence, in accordance with an illustrative embodiment. FIG. 4A is a graph illustrating retrospective internal validation of EGFR AI Genomic Lung Evaluation (EAGLE) performance, in accordance with an illustrative embodiment. FIG. 4B is a graph illustrating retrospective external validation of EAGLE performance, in accordance with an illustrative embodiment. FIG. 4C is a graph illustrating retrospective pretrial cohort validation of EAGLE performance, in accordance with an illustrative embodiment. FIG. 4D is a graph illustrating retrospective prospective silent trial cohort validation of EAGLE performance, in accordance with an illustrative embodiment. FIG. 5 is a block diagram illustrating a silent trial workflow, in accordance with an illustrative embodiment. FIG. 6A illustrates pretrial tuning and silent trial results, in accordance with an illustrative embodiment. -5- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 FIG. 6B illustrates pretrial tuning and silent trial results, in accordance with an illustrative embodiment. FIG. 7 illustrates a plurality of tissue samples analyzed using EAGLE, in accordance with an illustrative embodiment FIGs 8A–8G illustrate a validation and test ROC curves comparison. Each test cohort is plotted separately. 95% confidence interval of the ROC curve was computed via bootstrapping with 1,000 samples. FIGs. 9A and 9B illustrate the analysis of model performance on metastatic samples. (A) Internal validation ROC curves stratified by sample type: samples from primary site of disease vs metastatic sites. (B) AUC performance on metastatic samples stratified by metastasis location. Each barplot represents a single AUC for each location. The red vertical line represents the overall performance on samples from metastatic sites. FIGs. 10A and 10B illustrate the analysis of model performance stratified by tissue area (in squared millimeters). (A) The Distribution of tissue area per sample for the internal MSKCC validation cohort. (B)The distribution of tissue area was divided by deciles. For each bucket the AUC is plotted against the median area of the bucket. The x axis errors bar represents the range of the bucket, while the y axis error bar is the 95% confidence interval estimated via bootstrapping. The analysis was performed for primary and metastatic samples independently. FIG. 11 illustrates the comparison of model outputs for different EGFR mutation variants. Fig. 12 illustrates the validation AUC performance stratified by EGFR mutation variant. All variants achieved AUC scores that were not significantly different from the overall AUC score, highlighting the robustness of EAGLE across variants. -6- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 FIG. 13 illustrates the comparison of model outputs for paired slides scanned with different scanners. For each comparison, the Pearson correlation coefficient is shown. FIGs. 14A–14H illustrate the model performance results on the overall TCGA cohort and stratified by artifact type. For each artifact type, the result is obtained by removing slides containing the artifact. The shaded ROC region represents the 95% confidence interval calculated via bootstrapping with 1000 iterations. FIGs. 15A and 15B illustrates (A) Summary showing the median time from molecular accession to completion of EAGLE, rapid molecular testing, comprehensive genomic sequencing for all samples (N=197) from the silent trial. Error bars show the interquartile range. (B) Same as a) but starting from accession of surgical pathology specimen. Also includes time to scanning and time to molecular accession. The red vertical line demonstrates where time zero on plot a) is present on plot b). FIG. 16 depicts a block diagram of a system for identifying biomarkers in subject tissue samples, in accordance with an illustrative embodiment. FIG. 17 depicts a block diagram of a process to train a machine learning (ML) architecture to identify biomarkers in subject tissue samples, in accordance with an illustrative embodiment. FIG.18 depicts a block diagram of a process to identify biomarkers in subject tissue samples; in accordance with an illustrative embodiment. FIG. 19 depicts a block diagram of a process to provide an output based on the identification of biomarkers in subject tissue samples, in accordance with an illustrative embodiment. FIGs.20A and 20B depicts a flowchart of a method of training the ML architecture to identify biomarkers in subject tissue samples, in accordance with an illustrative embodiment. -7- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 FIG. 21 is a block diagram of a computing environment according to an example implementation of the present disclosure. DETAILED DESCRIPTION Following below are more detailed descriptions of various concepts related to, and embodiments of, systems and methods for classifying biomedical images to identify biomarkers in subject tissue samples. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in numerous ways, as the disclosed concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes. A. Scalable Artificial Intelligence Framework for Rapid EGFR Mutation Screening Presented herein is system and methods to predict Epidermal Growth Factor Receptor (EGFR) mutation status in non-small cell lung cancer (NSCLC) patients from H&E stained histopathology whole slide images (WSIs). The first uses a two-step process: training a vision transformer on lung histology classification, then using it as a frozen feature extractor for a multiple instance learning (MIL) aggregator. The second implements end-to-end training of a pre- trained foundation model encoder and an MIL aggregator using distributed training. An in-real- time pipeline is presented for rapid clinical EGFR screening. Experiments on a large patient cohort demonstrate effectiveness, with the best model achieving 0.83 AUC and 2-minute inference time per slide, offering a potential rapid, cost-effective alternative to conventional molecular testing in a live clinical setting. Lung cancer is the deadliest cancer in the United States (US). EGFR testing is the cornerstone of early treatment protocols and determine first line of therapy with EGFR tyrosine kinase inhibitors (TKIs). This targeted approach can lead to improved progression-free survival, enhanced quality of life, and potentially reduced overall disease burden in patients with EGFR- mutant NSCLC. For example, 22% of advanced NSCLC patients in the US had access to NGS testing between 2010-2018, with even lower rates in other developed countries. Limited adoption -8- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 may result from high costs, infrastructure gaps, tissue sampling issues, and time constraints. Where NGS is unavailable, less sensitive rapid tests may be used. Given these constraints, complementary techniques may be used to efficiently screen NSCLC patients, and / or any other type of patients, for EGFR status prediction. Artificial intelligence (AI)-based approaches utilizing digital whole slide images (WSIs) of routine hematoxylin and eosin (H&E) stained histopathology slides may be used. AI-based EGFR screening utilizes only digitized H&E slides, making it potentially cost-effective and accessible even in resource-limited settings. This could serve as a valuable adjunct to genomic testing allowing for lower cost, faster turnaround time (TAT), and better tissue utilization. Based on these considerations, this systems and methods described herein provide: (a) development and validation of AI-based approaches to predict EGFR mutation status directly from WSIs of NSCLC patients, (b) evaluation of the various EGFR prediction methods, and (c) an example of the clinical implementation of an EGFR status prediction system, designed to assess WSIs of NSCLC patients in real-time and at scale. Context Deep learning may be used for automated analysis of histopathology WSIs. In various embodiments, WSIs may be gigapixel images. As such, WSIs may contain, for example, around 100,000 pixels per dimension at 20× magnification (0.5 microns per pixel). Thus, pixel- level annotations may be time-consuming and expensive. To overcome this, weakly supervised learning techniques that rely on slide-level labels instead of pixel-level annotations may be used. Weakly Supervised Learning Due to the large size of WSIs, slides may be divided into a number of smaller “patches” (e.g., segments, portions, etc.). For example, a WSI may be divided into tens of thousands of smaller patches. The patches may then be analyzed using multiple instance learning (MIL). MIL formulates WSI classification as a bag-of-instances problem, where each WSI is a bag containing many patches (i.e., instances). A WSI may be considered positive if at least one -9- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 patch is positive. An attention-based MIL approach that learns to identify the most informative regions in a WSI may be used for cancer detection, such as breast and colon cancer detection. A deep learning system may be trained on WSIs (e.g., over 44,000 WSIs). The deep learning system may be trained using only slide-level labels, thereby achieving performance comparable to pathologists (e.g., humans) in cancer diagnosis. Various foundation models that may be developed through self-supervised learning may be used as feature extractors, which may deliver improved performance in MIL- based methods for various downstream pathology tasks. Other approaches may include, for example, vision transformer models pretrained on datasets of histopathology slides using self- supervised learning techniques. For example, a vision transformer model may be trained on, for example, 1.5 million H&E slides, and may be used in pan-cancer detection. As another example, a vision transformer model may be pretrained on 100 million image patches from 100,000 whole slide images across 20 tissue types. This model may improve upon previous models on various pathology tasks As yet another example, a vision transformer model may be trained on 1.3 billion image tiles from 171,000 slides, and be used in biomarker prediction and pan-cancer detection. Models may use various types of algorithms for pretraining, such as, for example, a DINO v2 algorithm. EGFR Status Prediction In various embodiments, a MIL may be used to predict EGFR mutations from H&E stained WSIs. A convolutional neural network (CNN) may be used on TCGA data for EGFR prediction in NSCLC patients. For example, a CNN may be trained on patches from manually curated regions of interest. Patch-level predictions may be aggregated to make slide- level classifications. However, reliance on manual curation and a dataset with high tumor purity may limit clinical applicability. In some embodiments, self-supervised contrastive learning may be used to pretrain a feature extractor, followed by attention-based MIL to aggregate patch features for slide-level prediction without requiring manual annotation. Similarly, another approach may utilize CNN to extract patch features. Patch features may then be aggregated -10- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 using, for example, an attention mechanism, to produce slide-level predictions. However, these models may be developed and validated on surgical resection specimens, with limited evaluation on biopsy samples that are more common in clinical practice. Additionally, current systems may not integrate these models in clinical practice for real-time EGFR status assessment. Methods Presented herein the following two approaches to predict EGFR mutation status directly from H&E stained WSIs: Approach 1: This approach may include training two separate models. A patch-level encoder may be trained. The patch-level encoder may extract a feature vector for each patch of a given WSI. The encoder's weights may be frozen after training, and the patch-encoder may serve as a feature extractor for the next stage. A slide-level aggregator may also be trained. The slide-level aggregator takes, as input, the feature vectors produced by the patch-level encoder for all patches in a WSI. The slide-level aggregator then aggregates these features to produce a slide-level prediction of EGFR mutation status. FIG. 1A shows an illustration of the Approach 1 described above, according to an example embodiment. During patch-level encoder training, a tumor-containing region of interest (ROI) is extracted for each WSI in the training set. The ROI may be extracted using, for example, Otsu’s thresholding. Morphological dilation may then be performed for background and small (e.g., less than 64 pixels) object (e.g., tissue area) removal. Non-overlapping patches of a particular size (e.g., of size 224 × 224 pixels) are then extracted from the identified ROI for every WSI. The patches are subsequently used as input to a vision transformer (ViT) that learns to assign each input patch to one of various histological types of cancer. For example, as shown in FIG 1A, each input patch may be assigned to one of the histological types of lung cancer (e.g., acinar, solid, lepidic, papillary, or micropapillary). Though reference is made throughout to various systems and methods used to classify or identify types of lung cancers, it should be understood that the systems and methods described herein may be used to classify or identify any -11- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 type of cancer. The ground truth histological labels for each WSI may be obtained by extracting genomics results from comprehensive genomics testing, and curating the EGFR mutations that are biologically relevant and targets of mutation specific therapy. The weights of the ViT are frozen after training and the ViT is used as an encoder to extract feature vectors (e.g., a CLS token) for input patches of a particular size (e.g., a size of 224 × 224 pixels) in the next stage. Specifically, as shown during multiple instance learning (MIL), patch extraction is performed for ROI of an input WSI. Feature extraction is performed on each patch using the pre-trained encoder. The extracted features are then concatenated. The concatenated features are used to compute attention scores. The computed attention scores are input into a weighted feature vector, and a classification is made. The classification may be expressed as a binary (e.g., “1” indicates EGFR and “0” indicates no EGFR). To summarize Approach 1, an encoder (e.g., a vision transformer) is trained on a surrogate task of histology classification (e.g., lung histology classification). Subsequently, the pre-trained encoder, with frozen weights, is used as a feature extractor for patches, followed by a slide-level aggregator (e.g., gated attention) to predict EGFR mutation status. The slide-level aggregator obtains feature vectors for every (non-overlapping) patch obtained from an ROI in a given WSI. The slide-level aggregator applies MIL pooling on plurality of patch-level feature vectors to compute a slide-level feature vector. The slide-level feature vector is subsequently used for EGFR status prediction. In various embodiments, gated multi-head attention (GMA) may be used for slide-level aggregation. In various embodiments, the weights of slide-level aggregator are only updated in the second stage of Approach 1 while the pre-trained encoder weights are frozen as illustrated in FIG. 1A. Approach 2: In Approach 2, the encoder weights may not be longer frozen and may be updated along with the MIL weights, as illustrated in FIG. 1B. Specifically, FIG. 1B shows an end-to-end training process for predicting EGFR mutation status, in accordance with an illustrative embodiment. Approach 2 updates both the encoder and MIL weights simultaneously. -12- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 This approach may be computationally expensive, as multiple WSIs are processed during training, which involves ROI identification, patch extraction, a forward pass through the encoder followed by MIL mechanism, a loss function computation, and backpropagation to update the weights of both the encoder and the MIL module. Thus, to implement Approach 2, a distributed training method presented in and summarized in Algorithm A1 was adopted. The algorithm utilizes two or more GPUs during training. One GPU may be dedicated to an aggregator process and N number of GPUs may serve as encoder processes. For each slide, the algorithm distributes batches of image patches across the N number of encoder GPUs. These GPUs perform parallel forward passes through the encoder to generate patch-level features. The resulting features are then gathered on the aggregator GPU, where they are concatenated and processed by the aggregator to produce a slide-level feature vector. The slide level feature vector is used to predict the EGFR status, and a loss is computed. The backward pass begins on the aggregator GPU, computing gradients up to the concatenated features. These gradients are then split and scattered back to the encoder GPUs. Each encoder GPU computes a pseudo-loss using its portion of the gradients, and backpropagates through the encoder. The encoder and aggregator models are then updated using their respective gradients. This approach may allow for efficient processing of large WSIs by leveraging multiple GPUs, thereby improving training speed and enabling the handling of high- resolution image data. Further, Approach 2 does not require a surrogate task of histology classification to train an encoder. In Real Time EGFR Assessment: FIG. 2 illustrates an in-real-time (IRT) pipeline to identify and process WSIs (e.g., using Approach 1 or 2 described above) for EGFR prediction in a live clinical setting, according to an example embodiment. As an overview, the process of FIG. 2 involves scanning and storing WSIs, retrieving slides of cancer patients that require EGFR testing, processing slides for EGFR prediction, and assessing results. The last step may include applying for regulatory approvals after completing required evaluations. -13- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 The purpose of the IRT pipeline is that the processes included in the pipeline may be generalized such that any new WSI will perform as it would in a live clinical scenario (e.g., as opposed to methods that utilize retrospective sample for which models can be built that overfit datasets). In various embodiments, approximately 90–100 NSCLC cases per month can be processed for which EGFR testing is clinically indicated. The IRT pipeline identifies slides that are scanned for molecular testing, as well as the slides scanned from the same surgical pathology block for which molecular testing was ordered. One or more (e.g., two) watchers may run every hour to identify 1) which slides have been scanned, and 2) which cancer cases are sent for molecular analysis. When a slide is determined to match the molecular case, the slide is transferred from archive, and the AI model is run (e.g., immediately upon transfer, within a period of time after transfer, etc.). This may allow for a real-time or substantially real-time EGFR prediction. If two or more WSIs are scanned, the mean probability score may be used. When a large deviation exists in score from one slide to another, the case is flagged to assess (e.g., by human) if one of the slides is from a different block than molecular testing. Referring now to FIG. 3A, a block diagram of a workflow for processing whole- slide images (WSI) is shown, according to an example embodiment. Specifically, the block diagram of FIG. 3A shows WSI processing using approaches that do not leverage artificial intelligence. When a biopsy or resection WSI is obtained, the image is first analyzed by a clinician for identification or presence of a biomarker, such as a protein (e.g., ALK or PDL1 IHC). Subsequently, if further analysis is to be performed, a portion of the tissue on the WSI slide is removed and taken for rapid testing (e.g., for rapid EGRF, KRAS, ERBB2 analysis, etc.). Further, additional portion(s) of the tissue on the WSL slide can be removed for comprehensive genetic analysis (e.g., using next-generation sequencing (NGS)). However, due to a lack of tissues after removal of one or more portions of the sample, a certain portion or number (e.g., 25%) of samples may fail to produce NGS. Referring now to FIG. 3B, a block diagram comparing a workflow of other approaches and an AI-enhanced workflow is shown, according to an example embodiment. Under the AI-enhanced workflow, the IRT pipeline identifies slides that are scanned for -14- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 molecular testing, as well as the slides scanned from the same surgical pathology block for which molecular testing was ordered. One or more (e.g., two) watchers run every hour to identify (1) which slides have been scanned and (2) which cancer cases are sent for molecular analysis. Responsive to a determination of a match between a slide and the molecular case, the slide may be transferred from archive and passed through the artificial intelligence (AI) model. The slide may be passed through the AI model immediately, within a predetermined time period, etc. This allows for a real-time EGFR prediction. Relative to the approach without AI, the AI-enhanced workflow can reduce the number of slides that are assigned to undergo rapid testing, which, as described above, can involve removal of a portion of the tissue from the WSI slide. In this manner, more WSI slides may have the original tissue on them, thereby allowing for a higher NGS success rate. Referring now to FIG. 3C, a block diagram illustrating WSI processing using artificial intelligence is shown, in accordance with an illustrative embodiment. Specifically, model training and validation may occur in or as part of a retrospective study. Performance tuning and silent trials may occur as part of a prospective study. Further, regulatory approval and clinical deployment may occur during a deployment phase of implementing the described WSI processing methods. In a standard clinical workflow for patients with lung adenocarcinoma (LUAD) and / or any other type of cancer, rapid tests for EGFR and other biomarkers are performed, reducing the tissue available for NGS and leading to up to one-quarter of the cases being unsuitable for NGS. By contrast, the clinical application of the proposed EGFR biomarker will allow a drastic reduction in the number of cases unsuitable for NGS. As soon as slides are digitized, the computational biomarker can be calculated and may be available to the pathologist before they review the case and sign it out. Based on the model’s outputs, the rapid test may be avoided, increasing the tissue available for NGS. Experiments and Results Data Description -15- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 The approaches discussed herein on a real-world clinical data were applied to predict EGFR mutation status including digitized slides of LUAD patient - scanned 20× magnification with Aperio AT2 digital slide scanner from Leica Biosystems, paired with ground truth EGFR mutational status obtained from the IMPACT sequencing panel. Additionally, the performance of the methods described herein was evaluated on an external dataset from the TCGA LUAD cohort. Table A1 below presents the distribution of WSIs across different stages for two distinct approaches. While Approach 2 utilizes all slides in the training set, the training set is divided into two subsets: one for encoder pre-training and another for MIL model training for Approach 1. Both approaches share identical model selection (validation) and independent testing sets, ensuring a fair comparison. All slides in the training, model selection, and independent test sets were scanned at 20× magnification. For the IRT set, both 20× and 40× lung cancer slides (requiring EGFR testing) scanned in the past 90 days are used. For 40× slides in the IRT set, 448 × 448 sized patches are extracted and downsampled to 224 × 224 for a forward pass through the trained models of both approaches. Results The learning parameters for the two approaches described above are presented in Table A2 in Appendix A. EGFR status prediction performance of all experiments is reported in terms of the area under curve (AUC), sensitivity, and specificity metrics for binary classification (EGFR positive or negative with 0.5 threshold). The best model for each approach is selected on the basis of the best AUC value on the model selection (validation) set (see Table A1). Table A3 in Appendix A presents the results of EGFR status prediction with the Approach 1 using different encoder architectures with GMA on the validation set. Table A3 shows that ViT-B achieved the highest AUC among all encoders. While ViT-L shows marginally better performance, the improvement may not justify its increased parameter count. Consequently, ViT-B is used as the encoder with GMA based MIL for all subsequent experiments with Approach 1. Prov-GigaPath is used as the encoder model for fine-tuning with -16- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 GMA based MIL in Approach 2 as it has shown better performance for several downstream tasks. Results in Table 1 reveal that while Approach 1 marginally outperformed on the model selection set, Approach 2 demonstrated superior generalization, consistently achieving higher AUC values on both the independent test set and IRT cases across 20× and 40× magnifications. This enhanced performance of Approach 2 can be attributed to two key factors: (1) the use of the Prov-GigaPath foundation model, pre-trained on a large corpus of pathology images, providing more robust initial weights, and (2) the end-to-end learning paradigm that allows fine-tuning of encoder weights based on EGFR prediction errors in the training set. These advantages enable Approach 2 to capture intricate pathological features more effectively and adapt its feature extraction process specifically to the EGFR prediction task. Consequently, Approach 2 shows greater potential for EGFR mutation prediction in clinical settings, despite being more expensive to train than Approach 1. Table 1: Performance metrics for the approaches across different datasets at 0.5 threshold for binary classification. Bold text indicates the best model (in terms of AUC). Dataset Approach 1 Approach 2 (n = # of WSIs) AUC Sensitivity Specificity AUC Sensitivity Specificity Model Selection 0.96 0.83 0.92 0.90 0.88 0.93 (n = 260) Independent 0.89 0.72 0.96 0.90 0.84 0.95 Testing (n = 6300) TCGA-LUAD (n 0.78 0.60 0.81 0.86 0.78 0.74 = 519) IRT (n = 1000) 0.80 0.65 0.80 0.83 0.62 0.83 For IRT, the inference time for both approaches may be about 2 minutes per slide, offering a significant improvement in turnaround time compared to conventional methods. The most common IdyllaTMrapid (PCR-based) EGFR testing has a 2-day turnaround time while suffering from issues with technical sensitivity, potentially missing important mutations. In -17- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 contrast, the more comprehensive IMPACT-based genomic assessment has an 18-days turnaround time and is more costly. The method offers a significantly shorter turnaround time, processing all slides scanned in a single day at the hospital within 30-45 minutes (median turnaround time is 20 minutes). As such, the systems and methods described herein illustrate the feasibility and clinical value of AI-driven EGFR mutation prediction from histopathology images in NSCLC patients. The described approaches offer rapid turnaround (2 minutes per slide), cost- effectiveness, and performance comparable to rapid PCR-based tests. The in-real-time (IRT) pipeline shows promise for expediting mutation detection, potentially leading to faster treatment decisions and optimized tissue sample utilization. As models are refined based on IRT performance, further improvements are anticipated. This work represents a significant step towards integrating AI-powered diagnostics into NSCLC management, advancing precision oncology. Appendix A Table A1: Number of WSIs in the training, model selection (validation), and testing sets EncoderMIL Model Model Independent Pre-training Training SelectionTestingTCGA IRTApproach 1 3475 1558 260 6300 514 1000 Approach 2 5033 Table A2: Training parameters for the approaches Approach 1 Parameter Approach 2 Encoder Pre-training MIL Training Number of GPUs 1 1 4 GPU NVIDIA A100 -18- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Number of epochs(^^^^^^^^) 80 50 150Learning rate (^^) 1 × 10ିହ 1 × 10ିସOptimizer Adam Weight decay (^^) 0.01 1 × 10ିସ 1 × 0.001Loss function Cross-EntropyWeighted BinaryWeighted Binary Cross-EntropyCross-Entropy Learning rateStepLR (^^^௧^^schedulerCosine with restarts=15, ^^ = 0.1 Cosine with restartsClass weighting Not used ^^^^^^௧^௩^ = 0.7 ^^^^^^௧^௩^ = 0.7Table A3: Performance metrics of various encoders with MIL using GMA on the validation set (260 slides, Table A1): Encoder ArchitectureAUC Sensitivity SpecificityResNet50 0.82 0.76 0.80 ViT-S 0.84 0.79 0.81 ViT-B 0.96 0.83 0.92 ViT-L 0.97 0.85 0.91 Algorithm A1 Distributed Encoder-Aggregator Training 1: Input: N + 1 GPUs, WSI Dataset D, Encoder E, Aggregator A 2: Output: Trained Encoder E and Aggregator A 3: Initialize: 4: Rank 0 GPU as aggregator process 5: Ranks 1 to N GPUs as encoder processes 6: Wrap E with DistributedDataParallel 7: for each epoch do 8: for each slide in D do 9: Distribute N batches of image patches to N encoder GPUs using distributed sampler ▷ Forward Pass 10: for rank r in 1 to N in parallel do 11: fr← E(batchr) ▷ Generate patch-level features 12: end for 13: F ← Gather(f1, ... , fN) to rank 0 ▷ Break computing graph -19- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 14: On rank 0: 15: Fconcatenated← Concatenate(F) 16: slide_feature ← A(Fconcatenated) ▷ Slide-level aggregator 17: output ← Project(slide_feature) ▷ EGFR status prediction 18: loss l ← ComputeLoss(output, ground_truth) ▷ Backward Pass 19: On rank 0: 20: Compute gradients from l back to Fconcatenated21: Split gradients into N chunks: g1, . . . , gN 22: Scatter(g1, . . . , gN) to ranks 1 to N 23: for rank r in 1 to N in parallel do24: pseudo_loss le ← ^^ × ∑|ி|^ୀ^ (^^^⌊^^⌋ × ^^^⌊^^⌋)25: E 26: end for ▷ Update 27: Update E on ranks 1 to N (handled by DistributedDataParallel) 28: Update A on rank 0 29: end for 30: end for 31: return E, A B. Deployment of a Fine-Tuned Pathology Foundation Model for Cancer Biomarker Detection An in-real-time (IRT) pipeline may be used to identify and process whole slide images (WSIs) for various predictions (e.g., Epidermal Growth Factor Receptor (EGFR) prediction) in a live clinical setting. The IRT pipeline may generalize such that any new WSI may perform as it would in a live clinical scenario. For example, approximately 90-110 non- small cell lung cancer (NSCLC) cases per month can be processed for which EGFR testing is clinically indicated. The IRT pipeline identifies slides that are scanned for molecular testing as well as the slides scanned from the same surgical pathology block for which molecular testing was ordered. Watchers run every hour to identify (1) which slides have been scanned and (2) which cancer cases are sent for molecular analysis. When a slide that matches the molecular -20- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 case, the slide is transferred from archive and passed through the artificial intelligence (AI) model immediately. This allows for a real-time EGFR prediction. If two or more WSIs are scanned, the mean probability score is used. When there is a large deviation in score from one slide to the next, the case is flagged to assess if one of the slides is from a different block than molecular testing. The in-real-time (IRT) pipeline may expedite EGFR mutation detection in live clinical settings, potentially leading to faster treatment decisions and improved patient care. By avoiding the need for expensive and potentially inaccurate rapid tests, this approach can optimize tissue sample utilization for comprehensive DNA testing. This approach may integrate AI- powered diagnostics into the clinical workflow for management (e.g., management of NSCLC), thereby causing more efficient and personalized treatment strategies in precision oncology. Artificial intelligence models using digital histopathology slides stained with hematoxylin and eosin may offer promising, tissue-preserving diagnostic tools for patients with cancer. The use of AI models offers many advantages. Assessing EGFR mutations in lung adenocarcinoma and / or other types of cancer demands rapid, accurate and cost-effective tests that preserve tissue for genomic sequencing. PCR-based assays may provide rapid results, but with reduced accuracy compared with next-generation sequencing, and require additional tissue. Computational biomarkers leveraging modern foundation models may address these limitations. The systems and methods described herein assembled and utilized a large international clinical dataset of digital lung adenocarcinoma slides (N = 8,461) to develop a computational EGFR biomarker. The fine-tuned foundation model fine-tunes improves task-specific performance with out-of-center generalization and clinical-grade accuracy on primary and metastatic specimens (e.g., with a mean area under the curve: internal 0.847, external 0.870). In one example, an area under the curve of 0.890 was achieved in a silent trial of the biomarker on primary samples. The artificial-intelligence-assisted workflow reduced the number of rapid molecular tests needed by up to 43% while maintaining the current clinical standard performance. The retrospective and prospective analyses shown in FIG. 3C demonstrate the real-world clinical utility of a computational pathology biomarker. -21- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Lung adenocarcinoma (LUAD) is the most prevalent form of lung cancer and has been found to have multiple somatic mutations in kinase genes, among which EGFR is the most prevalent, that are treated with tyrosine kinase inhibitor (TKI) therapy. Because tumors with EGFR mutations are treated with EGFR-specific TKIs, accurate EGFR testing is necessary for patients to receive the correct first-line therapy. Clinical testing guidelines provide guidance for the standard of care for molecular testing in lung cancer, and EGFR testing is a requirement for patients with advanced-stage LUAD (e.g., stage IB or higher). Despite these recommendations, EGFR testing is not performed on 24–28% of lung cancer cases in the USA. The reason for the discrepancy between clearly published guidelines and actual clinical practice may be related to technical hurdles in obtaining and processing samples for testing. Genomic sequencing, including targeted EGFR assays, is even less common in many regions of the world. Given the high prevalence of EGFR mutation in LUAD, the lack of EGFR testing results in tens of thousands of patients worldwide receiving suboptimal therapy for EGFR-mutated LUAD every year. Though reference is made throughout to LUAD, it should be understood that the systems and methods described herein may be applied to any type or types of cancer. Even in well-resourced centers that have adopted standards for universal EGFR testing for lung cancer, many samples fail to be assessed for EGFR status owing to insufficient material being available from diagnostic biopsies. Lung biopsies may not be performed regularly, due to the challenge of safely acquiring lung tissue. The number of tissue-based tests necessary for a proper diagnosis and to obtain comprehensive biomarker testing is large and increasing. As discussed herein, a standard lung cancer diagnostic biopsy workup may include standard hematoxylin and eosin (H&E) sections, PDL1 immunohistochemistry (IHC), diagnostic IHC (for example, TTF-1, p40 and so on), ALK fusion IHC, rapid EGFR testing and comprehensive genomic sequencing. Turnaround times (TATs) may be a challenge for the treatment of LUAD, especially for comprehensive sequencing with next-generation sequencing (NGS). NGS may have a TAT of approximately 2–3 weeks from the date of the biopsy. Following guidelines, first-line therapy may not be given until an EGFR mutation status is known because (1) EGFR mutant tumors benefit from first-line TKI rather than chemotherapy / -22- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 immunotherapy and (2) EGFR results further inform the likelihood of response to immune-based therapies. Rapid TAT testing has been developed and implemented to overcome this fundamental limitation of NGS. Owing to the targeted nature of rapid tests, rapid TAT testing may fail to detect less common EGFR variants, including many EGFR exon 20 insertions, uncommon exon 19 deletions and dinucleotide mutations resulting in common missense mutations (for example, p.L858R). This results in a technical sensitivity of 85–90% and a negative predictive value (NPV) of 90–95% on clinical experience, indicating that 5–10% of samples that screen negative for EGFR mutations actually harbor a targetable mutation and would receive the incorrect first-line therapy. Computational methods to detect EGFR mutations that can be deployed with little cost, rapid TAT and automated implementation while preserving tissue for comprehensive genomic sequencing may improve the clinical workflow for lung cancer diagnostic biopsies. Such a method may increase the detection of EGFR mutant lung cancer and reduce the number of suboptimal treatment regimens. Detecting mutations directly from H&E slides, if adequate performance characteristics are achieved, offers an opportunity to overcome many of the limitations of current clinical sequencing protocols for EGFR. A computational EGFR mutation assessment would use, as substrate, only the digitized pathology slides from the diagnostic H&E biopsy. A result could be reported with little cost and no physical processing. Such technology can also produce results immediately (e.g., in real time or substantially real time), which allows the results to inform all other downstream decisions. Prior studies have shown that molecular biomarkers for somatic mutations can be predicted directly from routine H&E slides. In the context of EGFR prediction in LUAD, various existing models have required manual delineation of tumor boundaries before analysis, indicating that such models may lack clinical relevance. Further, highly effective EGFR detection models have been built with modern weakly supervised techniques and convolutional neural network-based feature extraction encoders that can detect EGFR mutations with high accuracy on internal datasets. However, despite the evidence that EGFR mutational status can be -23- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 predicted with reasonable accuracy from pathology slides, and despite the potential positive impact to patient care, these models may have difficulty being clinically implemented. As described herein, systems and methods for EGFR AI Genomic Lung Evaluation (EAGLE) may be used as an H&E-based computational biomarker. EAGLE may predict the EGFR mutational status from diagnostic biopsies of patients with LUAD or any other type of cancer, enhancing the standard molecular workflow, as shown in FIG. 3B. Compared with the traditional workflow, the AI-assisted screening precludes rapid testing in a substantial amount of cases while maintaining overall high screening performance. In various embodiments, and as will be discussed herein, samples that are screened positive may still require NGS-based testing. The system (e.g., EAGLE system) may be and evaluated on digitized slides, illustrating the broad technical and biological variability expected from real clinical deployment. The system may be trained by fine-tuning a pathology foundation model on a number of slides (e.g., 5,174). To assess the robustness of the model in and prove generalization across institutions and scanners, validation may be performed on a number of internal slides (e.g., 1,742 internal slides) and on external test cohorts. The model may be deployed in real time to simulate performance in a real-world setting and provide an indication of whether the application of EAGLE can effectively reduce the need for rapid testing of EGFR without sacrificing the performance characteristics of the mutation screening process. Results Rapid Testing performance and rapid test benchmarking To assess the clinical performance of the EGFR rapid test assay, real-time PCR molecular-based testing (e.g., rapid testing) may be performed on patients with LUAD undergoing the described diagnostic workflow. The rapid test results may be compared with -24- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 NGS testing. In one example, rapid test results may have, for example, a sensitivity of 0.918, a specificity of 0.993, a positive predictive value (PPV) of 0.988 and a NPV of 0.954 in a time period analyzed. Fine-tuning foundation model performance The performance of the trained model may be assessed using the internal validation set of slides. Referring generally to FIGS. 4A-4D, EAGLE performance on the internal and external cohorts described above is shown, according to an example embodiment. Receiver operating characteristic (ROC) curves and respective AUCs are shown. The ROC confidence interval (e.g., ROC 85% confidence interval), shown in FIGS. 4A-4D as the shaded area, may be calculated via bootstrapping with a number (e.g., 1000) iterations. Specifically, FIG. 4A illustrates retrospective internal validation of EAGLE performance, in accordance with an illustrative embodiment. As shown in FIG. 4A and Table 5, the model achieved an AUC of 0.847. Model performance may be more accurate in primary samples (e.g., AUC 0.90) than in metastatic specimens (e.g., AUC 0.75) as shown in FIG. 9A. The analysis of metastasis location showed a pattern of differential performance as shown in FIG. 9B. These results are in support of the clinical application of EAGLE on primary samples. FIG. 4B illustrates retrospective external validation of EAGLE performance, according to an example embodiment. Analysis of metastasis location may indicate or show a pattern of differential performance, as shown in FIG. 4B. These results are in support of the clinical application of EAGLE on primary samples. FIG. 4C illustrates retrospective pretrial cohort validation of EAGLE performance, according to an example embodiment. FIG. 4D is a graph illustrating retrospective prospective silent trial cohort validation of EAGLE performance, according to an example embodiment. The performance of the model may be evaluated with respect to an amount of tissue present in the sample. Tissue amount may be used as a proxy for tumor amount. Tissue surface area may be calculated on the basis of the tiles used for model inference. The distribution -25- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 of tissue area over an internal validation set (FIG. 10A) may be divided into a number of “buckets” (e.g., ten buckets) by deciles, for example. As metastatic samples contain on average less tissue, analysis may be performed independently for primary and metastatic samples. The performance over tissue size is shown in FIG. 10B and Supplementary Tables 6 and 7. As the area of the tissue being analyzed increases, performance may increase. To ensure that the model does not systematically underperform for specific EGFR mutation variants, the model’s probability scores may be compared across mutation variants. Probability distributions across variants may be observed and compared with an overall distribution. In various embodiments, the variant distributions and overall distribution may be non-significantly different (FIG. 11 and Supplementary Table 8). This may indicate that the model may be able to detect all of the clinically relevant EGFR mutations. The validation AUC restricted to each EGFR mutation variant (FIG.12) may also be assessed. In one example, the validation AUC may indicate that the exon 19 T790M mutation has the highest AUC score, while the group of variants classified as ‘other’ have the lowest AUC score. In some embodiments, all variants may achieve AUC scores that are not significantly different from the overall AUC score. This may indicate the robustness of EAGLE across variants. External validation is consistent with internal results The performance of the model may be assessed on a variety of external cohorts. In various embodiments, performance of the model using external cohorts may align with internal validation. For example, the performance of the model on the external cohorts may achieve an overall AUC of 0.870 on 1,484 slides, as shown in FIG. 4B, thereby highlighting the generalization capacity of the model. Various external cohorts may show AUCs of, for example, 0.870 (N = 294), 0.877 (N = 241) and 0.884 (N = 259) for slides scanned with various scanners, respectively. In one example, the model may achieve an AUC of 0.772 (N = 95). In other examples, the model may achieve an AUC of 0.808 (N = 76) or an AUC of 0.860 (N = 519). Supplementary Figs. 8A–G and Supplementary Table 5 summarize these results. -26- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Slides may be scanned using different scanners or types of scanners. The pairwise linear correlations between the model scores for a slide scanned on two different scanners may be compared. For example, Pearson coefficients of 0.828, 0.832, and 0.935 may be obtained, indicating the robustness of the model across a variety of scanners (FIG. 13 and Supplementary Table 9). In various embodiments, slides may have one or more artifacts present, which may, in some embodiments, affect performance. For each type of artifact, performance in terms of AUC may be calculated when removing slides containing that artifact. The model may have stable AUC performance across various stratifications, indicating robustness against all artifact types (FIGs. 14A–H). Further, by removing the slides that contain the most severe artifacts that may obfuscate the tissue morphology, AUC may increase (e.g., from 0.860 on an overall cohort to 0.918 with slides removed). Silent trial supports the clinical use of EAGLE In various embodiments, the use of a silent trial may validate the clinical use of EAGLE for the prediction of EGFR mutations in LUAD samples extracted from a primary tumor site, based on internal and external validation results. A silent trial may include one or more (e.g., two) stages. First, a pretrial cohort may be used to simulate an outcome of model deployment under various assistive strategies and select appropriate model score thresholds. Second, using the threshold set, the in-real-time (IRT) silent trial may be performed, capturing time and results from EAGLE, rapid test and MSK-IMPACT for analysis. FIG. 5 shows a flow diagram of a silent trial, according to an example embodiment. Specifically, in the clinical workflow portion of FIG. 5, relevant components of the standard clinical workflow are shown along a timeline. ΔT indicates the time from molecular accession to the availability of a result. The silent trial components occurring in parallel with the clinical workflow are indicated in the “in-real-time silent trial” portion of FIG 5. -27- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 FIG. 5 illustrates the IRT pipeline to identify and process WSIs of primary samples of LUAD specimens for EGFR prediction in a live, real-world, clinical setting in the context of a silent trial. The IRT pipeline automatically identifies slides that are scanned for molecular testing, as well as the slides scanned from the same surgical pathology block for which molecular testing was ordered. One or more watcher applications may be run automatically on an, for example, hourly cadence to identify (1) which slides have been scanned and (2) which cancer cases are sent for molecular analysis. When a slide matches a molecular case of interest, the slide may be transferred from the digital pathology system to the GPU compute infrastructure, and inference on the AI model may be run immediately. This setup allows automated, real-time EGFR prediction. If two or more WSIs are scanned, the first scanned slide is used. During the silent trial, the results from EAGLE, rapid testing and MSK-IMPACT may be collected. In addition, timestamps from key events were recorded, specifically when the rapid test is accessioned, which triggers the execution of EAGLE, when the result from EAGLE is produced, when the rapid test result is generated, and when the MSK-IMPACT test result is ready. Based on this information, the performance of the assisted screening pipeline consisting of EAGLE and the rapid test may be analyzed against existing workflows consisting of the rapid test alone. A pretrial cohort (e.g., N = 765) may be analyzed with EAGLE, obtaining, for example, an overall AUC of 0.853, (e.g., in line with an expected performance). In concordance with the internal validation, model performance may be observed to be higher in primary samples (N = 374, AUC 0.896, as shown in FIG. 4C) than in metastatic specimens (e.g., AUC 0.760). Specific locations of metastases may have particularly poor performances, such as the lymph nodes (AUC 0.74) and bone (AUC 0.71). Other metastatic samples may perform closer to a level seen on primary samples, such as the liver (AUC 0.83) and brain (AUC 0.79). These results further support the deployment of EAGLE for primary samples. In an existing workflow, polymerase chain reaction (PCR)-based rapid tests may be run on all LUAD samples. Under an artificial intelligence (AI)-assisted workflow, some of the samples may be spared from rapid testing based on the output of the EAGLE model. In a high- -28- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 NPV domain (high-score regime), samples scored negative by EAGLE may be true negatives and may not undergo rapid testing. Conversely, in a high-PPV domain (e.g., low-score regime), samples scored positive by EAGLE may be likely to be true positives, and may also not undergo rapid testing. Given these two tunable parameters, the AI-assisted screening workflow may be defined as follows: (1) if the sample’s EAGLE score is below the NPV threshold, the sample is negative, no rapid test may be needed or performed; (2) if the EAGLE score is above the PPV threshold, the sample is positive, no rapid test may be needed or performed; (3) if the score is in between, the rapid test is performed for confirmation. To simulate deployment, various performance metrics associated with the AI- assisted screening may be calculated when modulating the NPV and PPV thresholds. FIGS. 6A and 6B show the pretrial tuning and silent trial results. In the AI-assisted EGFR screening, samples with EAGLE scores below the NPV threshold or above the PPV threshold can be spared from the rapid test. FIG. 6A shows a heatmap of the reduction of rapid tests with isolines corresponding to the historical PCR rapid test performance when modulating the NPV and PPV thresholds. Each point in the heatmap of FIG. 6A may be expressed as (average, average). The 95% confidence intervals (CI) may be estimated via bootstrapping with 1,000 iterations. FIG. 6A further shows a zoomed-in area from a focusing on the top right corner where NPV and PPV are maximized. Threshold points are chosen within the rapid test noninferiority region with increasing levels of rapid test reduction. FIG. 6B shows pretrial deployment along the line established in FIG. 6A with the selected thresholds (thresh.) as vertical lines. The historical rapid test performance is shown with solid horizontal lines, and the dashed horizontal lines represent the 95% confidence intervals estimated via bootstrapping with 1,000 iterations. From top to bottom, NPV, PPV and rapid test reduction associated with AI-assisted workflow are presented. Shaded areas represent the 95% confidence interval estimated via bootstrapping with 1,000 iterations. FIG. 6B also shows a similar analysis for the silent trial cohort. The thresholds may be chosen on the pretrial cohort and not on the silent trial cohort. -29- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 As shown in FIG. 6A, the reduction of rapid tests is shown as heatmap and with isolines the historical performance of rapid PCR testing for NPV and PPV. The top right corner (thresholds 0, 1), enlarged in FIG. 6A, constitutes the original rapid test performance without AI assistance (e.g., rapid test decrease equal to 0), whereas the bottom left corner (thresholds 0.5, 0.5) may be equivalent to completely replacing the rapid test by EAGLE. Areas of high NPV and PPV intersect in the top right corner. Identifying points of the two-dimensional space where the assisted workflow is noninferior to the rapid test alone may be identified. As an example, three points corresponding to three sets of thresholds with increasing levels of rapid test reduction may be identified, all within the noninferiority area. These points may lie on a line that balances the joint optimization of NPV, PPV and rapid test reduction. FIG. 6B shows the NPV, PPV and rapid test reduction associated with the AI- assisted workflow in the pretrial cohort along the path identified in Fig. 6A. The selected thresholds are also shown as vertical lines. With the most conservative threshold (0.004, 0.999), the assisted workflow may be expected to yield, for example, 0.954 NPV, 0.991 PPV and 25% rapid test reduction. With the least conservative threshold (0.038, 0.995), the assisted workflow may be expected to yield, for example, 0.952 NPV, 0.981 PPV and 43% rapid test reduction. Table 1 | Performance of the AI-assisted EGFR screening for the pretrial and silent trial cohorts Cohort Threshold AI-assisted NPV AI-assisted PPV Test reduction NPV PPV Average 95% CI Average 95% CI Average 95% CI Pretrial 0.004 0.999 0.954 0.929–0.976 0.991 0.929–0.976 0.250 0.210–0.293 Pretrial 0.023 0.997 0.952 0.925–0.974 0.981 0.925–0.974 0.389 0.341–0.437 Pretrial 0.038 0.995 0.952 0.924–0.975 0.981 0.924–0.975 0.433 0.383–0.481 Silent trial 0.004 0.999 0.971 0.941–0.993 1.000 0.941–0.993 0.178 0.127–0.229 Silent trial 0.023 0.997 0.970 0.941–0.993 0.984 0.941–0.993 0.371 0.310–0.442 Silent trial 0.038 0.995 0.963 0.930–0.992 0.984 0.930–0.992 0.431 0.360–0.503 -30- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 On the samples analyzed (N = 197), EAGLE obtained an AUC of 0.890 (Fig. 4D) using the final MSK-IMPACT results as the ground truth, in line with the expected performance. Deployment of the AI-assisted workflow may be evaluated at the three operating points defined for the pretrial cohort, described above. In a most stringent setting, the assisted workflow achieved an NPV, PPV and rapid test reduction of 0.971, 1.0 and 18%, respectively. In the least stringent setting, these metrics were equal to 0.963, 0.984 and 43%. The complete set of results are listed in Table 1 above. Figure 6B shows the continuous NPV, PPV and rapid test reduction of the AI-assisted screening in the silent trial. Overall, these results demonstrate the noninferiority of the AI-assisted workflow when deployed IRT in the clinical setting. The TATs of each test may also be analyzed. As an example, it may be determined that EAGLE results are available with a median TAT from the time of molecular accession of 0.74 h (44 min), while PCR rapid testing has a median TAT of 48.78 h, and MSK- IMPACT has a median TAT of 435.26 h (FIG. 15A). FIG. 15B demonstrates that TAT relative to the surgical accession could be greatly improved if, for example, an order from the clinician for AI was provided. EAGLE model introspection The silent trial enables understanding of testing protocol performance in a real- world setting, including possible sources of false positive and false negative results. This ability for introspection is enhanced by generating image overlays to highlight the areas that are most attended to by the model (e.g., ‘high attention’ areas). FIG. 7 specifically illustrates model introspection using the attention scores from the aggregation function. As inference is deterministic, images may be generated in one shot. Repeat inference may generate identical image. As shown in FIG. 7, each figure is an example from the silent trial. Case a indicates a true positive, case b indicates a true negative, case c indicates a false positive, and case d indicates a false negative. In each image, the top left may be the thumbnail of the H&E WSI. The top middle includes an overlay of the thumbnail, with the -31- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 full spectrum of the attention mask from a score of −4 to 4. The attention is the level to which the model is attending to the region on the image for making the decision of positive or negative (that is, does not indicate whether the model is interpreting the area as positive or negative, but only weighting). The bottom left has an overlay of the regions of the slide that have an attention score >3 (that is, high attention). The label at the bottom of panel is the quantity of pixels with high attention. The bottom middle has an inverted mask so that the non-high-attention regions are obscured. The red box in the panel indicates the region of the WSI that has the highest density of high-attention pixels. To the right is a high-resolution image of the portion of the slide highlighted by the red box in the prior panel. FIG. 7 shows examples of the attention maps discussed above. Analysis of the attention maps alongside the respective genomic profiles of each case, can provide insight into false positives and false negatives. For example, when evaluating cases predicted by EAGLE to have a high probability of being EGFR positive but that are negative by MSK-IMPACT, the cases may either have (1) a biologically related mutation (for example, ERBB2 exon 20 insertions) or (2) certain kinase activating events (for example, MET exon 14 skipping mutations and ROS1 fusions). Case c in FIG. 7, for example, was predicted by the model to be positive for EGFR but has a ERBB2 exon 20 insertion. In the specific example, false negative cases by the stated thresholds may be too few to draw clear conclusions from the silent trial. To assess potential sources of false negatives, EGFR-positive cases with EAGLE-predicted probability <0.5 may then analyzed. These cases may include (1) cytology specimens lacking fragments with preserved tumor architecture, (2) biopsies with mostly blood and minimal fragments of tumor tissue, and / or (3) tumors with unusual morphologies for EGFR mutant LUAD (for example, tumors with high tumor-infiltrating lymphocytes and spindle-cell morphologies). Case d of FIG. 7 has an EAGLE score of 0.43, but is positive for an EGFR mutation. The sample is largely blood and has very little tissue and almost no tumor, yet the attention mechanism correctly highlights the rare areas of tumor. Furthermore, the molecular results report the variant fraction as <5%. Thus, this sample is also borderline for molecular assessment owing to the -32- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 relatively low tumor content. In various embodiments, manual interpretations by pathologists, with tools like the masks presented, may lower the error rate. Discussion The systems and methods described herein therefore demonstrate the real-world clinical-level performance of a computational pathology biomarker in a real-time setting. While many AI models have shown reasonable performance in cross-validation experiments on retrospective datasets, the systems and methods described herein provide a link between model development and clinical deployment. The silent trial in which the model is applied prospectively to slides scanned in a live clinical setting may represent a shift from controlled experimental conditions to clinical utility. Further, in an IRT scenario, the model may be applied to slides that do not exist when the model is finalized, thereby ensuring a realistic test of the model’s ability to generalize to new, unseen data, which may be necessary step toward regulatory approval and use in clinical practice. Various existing solutions may focus on achieving high-performance metrics using retrospective, held-out test datasets. These solutions may rely on curated datasets with strict inclusion criteria and closely matched data distributions between training and test sets. Although retrospective studies facilitate performance optimization for a specific dataset, they do not adequately simulate how a model would perform in real-world clinical deployment. Furthermore, existing solutions may not demonstrate that these models can generalize to slides prepared in various laboratories, where variations in sample preparation, scanner hardware and staining protocols occur. The systems and methods described herein demonstrate cross-location and multi-scanner performance for a challenging computational biomarker. In addition, the IRT method evaluates EAGLE under conditions of a real-world clinical settings, using routinely processed slides from a diverse patient population that reflects circumstances when utilized clinically. The silent trial further establishes a critical benchmark for assessing the true clinical readiness of AI models in computational pathology. -33- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 The silent trial results may support using EAGLE as a screening test complementary to existing rapid tissue-based tests, thereby enhancing clinical genomics workflows. AI-assisted screening can reduce the number of costly, tissue-consuming rapid tests. Unlike rapid PCR-based assays, EAGLE may detect all clinically relevant EGFR mutations. EAGLE may be used as a screening test, not a replacement for NGS sequencing. That is, EAGLE may rule out EGFR mutations and identify likely positive cases. However, because EAGLE may not distinguish between mutations requiring different TKIs, sequencing confirmation may be used before TKI therapy initiation. In the example silent trial discussed herein, EAGLE maintained high accuracy on real-time clinical samples, achieving an AUC of 0.890, consistent with internal and external validation, highlighting the model’s robustness and applicability to clinical implementation. In various embodiments (e.g., as shown by simulated clinical use), EAGLE predictions may be available within a median of, for example, 44 min. EAGLE predictions may therefore be available substantially faster than rapid tissue-based tests (e.g., 48 h) and comprehensive genomic sequencing (e.g., 2–3 weeks). This rapid turnaround allows clinicians to make informed decisions sooner, potentially initiating therapies earlier and improving patient outcomes. The use of EAGLE may also conserve biopsy material for comprehensive genomic testing, resulting in fewer test failures owing to insufficient material and the need for repeat biopsies. Automated deployment delivers results directly to pathologists, thereby enhancing interpretation within clinical contexts. In various embodiments, incorporating pathologists into the workflow could further improve model performance. The performance of a rapid molecular test, which requires a molecular pathologist’s interpretation, may also be enhanced as equivocal results are deferred until confirmation or refutation of the mutation is provided by NGS. A pathologist-in-the-loop workflow may similarly benefit clinical AI models. In this regard, it has been shown that analysis of model attention maps alongside histology and mutation profiles identified trends associated with false positives and negatives. Samples with minimal tumor architecture, such as cytology samples, may generally exhibit lower -34- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 scores and attention. For these samples, different triage strategies (for example, additional quality check and separate models) may be necessary to deliver similar performance as larger biopsies. The false positives may indicate overlapping tumor morphologies in certain biologically similar mutations that activate kinase domains of protooncogenes. Advantageously, the EAGLE computational methods may provide rapid inference, with median inference time of, for example, 68 s on a consumer-grade graphics processing unit (GPU). In existing solutions, model inference may be triggered by molecular accession, although slides are typically scanned several days earlier. Thus, the use of an ‘AI service’ clinical workflow in which a clinician could order the AI result along with the biopsy, the EAGLE result could be produced and reported even sooner. EAGLE may also offer greater efficiency gains in hospitals relying solely on NGS. In this setting, EAGLE may allow earlier initiation of chemotherapy or immunotherapy for tumors screened negative. The systems and methods described herein may further utilize a foundation model for feature extraction. This may enable the development of a highly generalizable and robust EGFR prediction model. Foundation models, which may be trained on large, diverse datasets, offer several advantages over traditional deep learning models, including improved transfer learning capabilities and feature representations that are adaptable to a wide range of tasks. In various embodiments, a foundational model, such as a vision transformer, may be fine-tuned to achieve clinical-level performance for the specific task of EGFR mutation detection in LUAD. The fine-tuning of a foundational model may allow EAGLE to generalize across institutions, patient populations and scanning equipment, which is a challenge that has hindered the clinical translation of computational pathology models. This flexibility may allow for real- world deployment, where variability in slide preparation and digitization may exist. The fine- tuning strategy described herein further enhances the foundation model’s performance, yielding a significant improvement in AUC compared with previous studies using traditional convolutional -35- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 neural networks. For instance, the model achieves a mean AUC of 0.861 across multiple external cohorts, surpassing the AUC of existing solutions on a smaller, retrospective dataset. The application of EAGLE in a real-time clinical setting demonstrates the potential of computational pathology models to transform traditional diagnostic workflows. With the use of the EAGLE model, computational pathology may be used in precision oncology, offering scalable, low-cost, automated and highly accurate solutions for a wide range of clinical tasks. Unlike traditional diagnostic tools, AI-based computational biomarkers are inherently digital, enabling remote deployment and democratization of advanced diagnostic capabilities to underserved regions worldwide. By reducing or eliminating the need for tissue-based testing and offering rapid, reproducible results, AI models like EAGLE can help to bridge gaps in care and ensure that all patients, regardless of geographic or economic barriers, have access to state-of- the-art diagnostic tools. As such, the systems and methods described herein provide a real-time, clinically validated computational pathology model for EGFR mutation detection in LUAD. By utilizing the strengths of foundation models and validating the model in a prospective IRT setting, a benchmark for clinical-level performance in computational pathology can be established. The EAGLE system may improve diagnostic efficiency, reduce tissue consumption and accelerate the adoption of AI in routine clinical practice. In various embodiments, the model may be expanded to include additional biomarkers and evaluating its impact on therapeutic outcomes in a prospective clinical trial. Methods Whole-slide images and sequencing datasets This research study was approved by the respective institutional review boards at the Icahn School of Medicine at Mount Sinai (protocol 19-00951) and MSKCC (protocol 18- 013). Informed consent was waived as per the institutional review board protocols. Participants -36- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 were not compensated. Sex and / or gender was not considered in the study design as cohorts were generated as random samples of the patient population. Slide and sample preparation is described herein. For example, a dataset may be divided into, for example, a training dataset, a validation dataset, a dataset for calibrating a clinical threshold, and a dataset including slides processed in real time for a silent trial. For the retrospective cohorts (e.g., training and validation), the last section taken from each formalin- fixed paraffin-embedded (FFPE) tissue block (e.g., after all unstained slides used for molecular sequencing have been cut) may be stained with H&E. All samples, including cytology cell block samples, are part of the retrospective cohort. The calibrating and IRT datasets may be evaluated using the diagnostic slides, including from cytology cell blocks, the first section from the block, before either the rapid test and / or genomic sequencing. Slides may be digitized with a mix of, for example, Aperio AT2 (e.g., at 20× magnification) and GT450 (e.g., at 40× magnification) digital slide scanners. All slides that are part of a standard clinical workflow may be utilized. Thus, EAGLE may utilize or model the full extent of biological and technical variability of the clinical setting. Cases for prospective sequencing may be selected in real time by identifying LUAD samples for which rapid EGFR genomic sequencing is ordered prospectively. The ground truth for EGFR mutations may be established using, for example, the MSK-IMPACT targeted genomic sequencing assay, performed on the same tissue block from which the digital slide is created. MSK-IMPACT is a hybridization capture-based NGS assay used to detect clinically relevant somatic mutations, copy number alterations and gene fusions across cancers. This assay may screen for variants in up to 505 unique cancer-related genes, including EGFR, in all tumor types. All sequencing may occur in a Clinical Laboratory Improvement Amendments (CLIA)-certified laboratory, and each variant may be reviewed by a board-certified molecular pathologist. For LUAD samples, the standard clinical workflow may include rapid EGFR mutation testing via a PCR testing platform. In this process, tissue is scraped directly from unstained slides and placed into a cartridge. The cartridge, once mounted onto the analyzer, -37- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 facilitates both DNA extraction and all PCR reactions in a fully automated, self-contained manner. Output data may then be uploaded to the platform’s website, where the platform provides analysis and interpretation of common EGFR mutations. A board-certified pathologist may interpret the results, which may be subsequently reported in the clinical setting. In various embodiments, an H&E-stained section of samples adjacent to those used for sequencing may be reviewed for tumor cellularity and serve as a source of imaging. For example, slides may be obtained and scanned using one or more scanners or types of scanners. Some samples and / or cohorts may include LUAD specimens consecutively ascertained and subjected to molecular profiling. Genomic DNA and RNA may be extracted from unstained FFPE tissue sections and then sent for profiling. Clinically relevant variants (single-nucleotide variants, indels, gain and fusions) may extracted from reports. In some embodiments still, consecutive FFPE material of LUAD tumors (e.g., N = 95) may be obtained through surgical resections (e.g., sublobar wedge resections or lobectomy). The slides may be scanned with digital slide scanner (e.g., a 40× mode scanner with a resolution of 0.23 µm per pixel). In various embodiments, a targeted, multi- biomarker assay that enables detection of hotspots, single-nucleotide polymorphisms, indels, copy number variations and gene fusions from DNA and RNA in a single workflow, may be used and analyzed. The analysis may cover variants across, for example, 52 major genes with frequent alterations in non-small-cell lung cancer (NSCLC). In some embodiments, FFPE tissue sections of patients with LUAD may be digitized using one or more slide scanners. For the molecular analysis, any type of assay may be used, such as a pan-cancer assay. A pan-cancer may assay allow targeted-capture sequencing of, for example, 523 cancer-related genes at the DNA level, and translocation detection of, for example, 50 driver fusion genes at the RNA level. Sequencing may be performed and data may be processed and analyzed, followed by a pipeline using a second variant caller (Mutect2) and ANNOVAR for annotation of alterations. For DNA analysis, single-nucleotide variants, insertions and deletions, copy number variations, total mutation burden and microsatellite -38- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 instability may be calculated. For RNA analysis, putative gene fusion of around, for example, 50 fusion driver genes and RNA splice variants from EGFR, AR or MET (for example, MET exon 14 skipping) may be explored. In various embodiments, a cohort of primary resection specimens may be used. For example, samples may have comprehensive genomic profiling of, for example, 230 resected LUADs, by whole-exome sequencing (WES) and RNA sequencing. The corresponding diagnostic digital slides may be downloaded and reviewed by expert pathologist. The pathologist may annotate the presence of different types of artifact on the slides, such as: low-quality stain, blue and red saturation, blur, widespread tissue necrosis, freeze artifacts, severe artifacts that completely obscured the morphology of the sample, etc. EGFR mutations may be were clinically characterized using a database. In various embodiments, mutations outside of the EGFR kinase domain (exons 18–24) may not be oncogenic and therefore may be excluded from analysis. Oncogenic EGFR mutations may be grouped into the common subtypes: (1) exon 19 deletions, (2) L858R, (3) exon 20 insertions, (4) T790M, and (5) other kinase domain mutations. EAGLE—EGFR prediction model The model may include: (1) a multi-parameter (e.g., 1.1-billion-parameter) vision transformer (ViT-g) that encodes high-resolution (e.g., 20× magnification, 0.5 µm per pixel) 224-pixel patches into a feature vector (e.g., 1536 feature vector); (2) a gated multiple instance learning (MIL) attention (GMA) aggregator that integrates all encoded patches from a slide into a global slide-level feature representation; and (3) a linear classifier that outputs the probability of an EGFR mutation based on the input slide data. During training, the encoder may be initialized with a pathology foundation model. The full model may then be trained end to end using a parallelization strategy. In the parallelization strategy, the encoding is parallelized across numerous processes to divide the GPU memory burden across several GPUs thereby allowing joint optimization of the encoder, aggregator and classifier. A separate GPU may receive the encoded images, aggregate them with -39- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 GMA, and produce the classification loss. During backpropagation, the gradients may be directed to each process and synchronized. For each slide, a number of tissue patches may be sampled at each training step and divided across a number of GPUs for encoding. For example, 6624 patches may be sampled and divided across 23 GPUs to achieve 96 patches per GPU. The patch encoding may be performed in 16-bit float precision to enable the use of larger image batches. The model was trained on a number of GPUs for a certain number of epochs in a certain time period. At inference time, the trained model can be run on a single GPU. For the silent trial, EAGLE may be deployed with full floating-point precision using a single GPU. The time required to process a slide may be a median time of 68 s, making it suitable for real-time application in the clinical workflow. On lower-capacity hardware, the deployment of EAGLE is still possible by trading off memory consumption with inference speed. Supplemental Materials Cohort Descriptions Supp. Table 1 Clinical Characteristics of MSKCC Cohorts Used for Training, Validating, Calibrating, and Real-Time Analysis. Values for Sex, Smoking History, Race, and Stage are obtained from internal CbioPortal instance. Smoking status is obtained by natural language processing (NLP) of clinical notes from the patients charts. Stage is highest recorded stage in the patient’s clinical history at the time data is obtained, not the stage at the time of diagnosis. Category Training Dataset Validation Calibrating In Real Time Dataset Threshold Dataset Dataset Sex Female 3129 (64.29%) 1032 (62.93%) 471 (61.65%) 210 (66.67%) Male 1738 (35.71%) 608 (37.07%) 293 (38.35%) 105 (33.33%) Smoking Former / Current 3291 (67.62%) 1093 (66.65%) 487 (63.74%) 215 (68.25%) History Smoker (NLP) Never 1466 (30.12%) 501 (30.55%) 259 (33.9%) 93 (29.52%) -40- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Unknown 110 (2.26%) 46 (2.8%) 18 (2.36%) 7 (2.22%) Race Asian-Far 556 (11.42%) 164 (10.0%) 89 (11.65%) 41 (13.02%) East / Indian Subcontinent Black or African 259 (5.32%) 88 (5.37%) 51 (6.68%) 16 (5.08%) American Native American- 6 (0.12%) 1 (0.06%) 0 (0.0%) 0 (0.0%) Am Ind / Alaska Native Hawaiian 3 (0.06%) 1 (0.06%) 1 (0.13%) 0 (0.0%) or Pacific Islander White 3772 (77.5%) 1294 (78.9%) 553 (72.38%) 248 (78.73%) Unknown or 250 (5.14%) 84 (5.12%) 68 (8.9%) 10 (3.17%) Other Stage Stage 1-3 2621 (53.85%) 892 (54.39%) 485 (63.48%) 132 (41.9%) (Highest Stage 4 2002 (41.13%) 661 (40.3%) 243 (31.81%) 43 (13.65%) Recorded) Unknown 244 (5.01%) 87 (5.3%) 36 (4.71%) 140 (44.44%) Age at years, mean [95% 65.67 [65.37, 65.22 [64.70, 68.57 [67.81, 68.63 [67.53, Diagnosis Confidence 65.97] 65.74] 69.33] 69.73] Interval] Source Surgical 4070 (83.62%) 1404 (85.61%) 620 (81.15%) 258 (81.90%) Material Cytology 797 (16.38%) 236 (14.39%) 144 (18.85%) 57 (18.09%) Sample Primary 3104 (63.79%) 1038 (63.29%) 480 (62.83%) 225 (71.43%) Type Metastatic 1762 (36.21%) 602 (36.71%) 284 (37.17%) 90 (28.57%) Supp. Table 2 MSHS Cohort Description. Slides EGFR+% Patients -41- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Overall 294 35.0 287 Sex NA 56 16.1 55 Female 133 44.4 130 Male 105 33.3 102 Race NA 66 16.7 65 White 106 35.8 105 Asian 33 54.5 32 Black 40 35.0 39 Other 49 44.9 46 Smoking NA 43 30.2 Never 80 65.0 Pas 136 25.5 Current 35 5.7 Age of Diagnosis NA 56 16.1 30-40 4 75.0 40-50 8 37.5 50-60 26 19.2 60-70 59 44.1 70-80 83 42.2 80-90 54 37.0 90-100 4 50.0 Stage NA 173 36.4 1 61 41.0 -42- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 2 8 50.0 3 16 0.0 4 36 30.6 Supp. Table 3 SUH Cohort Description. Slides / Patients Overall 95 EGFR Status EGFR mut 58 EGFR wt 37 Sample Type Primary 51 Metastatic 44 Sex Female 53 Male 49 Age at Diagnosis 20-30 0 30-40 3 40-50 4 50-60 4 60-70 20 70-80 48 80-90 8 90-100 0 NA 8 Supp. Table 4 TUM Cohort Description. -43- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Slides / Patients Overall 76 / 41 EGFR Status EGFR mut 23 EGFR wt 18 Sample Type Primary 27 Metastatic 14 Sex Female 17 Male 24 Scanner AT2 38 GT450Dx 38 Validation and Test Results Supp. Table 5 Evaluation of the performance on the internal validation cohort and external test sets. 95% confidence interval (CI) calculated via bootstrapping with 1,000 iterations. Cohort N AUC 95% CI Internal Validation 1742 0.847 0.828-0.866 External Test 1484 0.870 0.851-0.889 MSHS Philips 294 0.870 0.827-0.907 MSHS Aperio 241 0.877 0.832-0.918 MSHS Pramana 259 0.883 0.843-0.922 SUH 95 0.772 0.672-0.860 TUM 76 0.808 0.708-0.896 -44- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 TCGA 519 0.860 0.814-0.902 Sample Type Analysis Tissue Area Analysis Supp. Table 6 Analysis of model performance stratified by tissue area (in squared millimeters) for the primary samples in the internal validation cohort. The number of slides and the number of positive slides within each area range is also provided. AUC Median Area Slides EGFR+Area Range (0.115, 3.751] 0.812 2.270 111 39(3.751, 7.215] 0.929 5.639 110 30(7.125, 11.854] 0.873 9.145 110 38(11.854, 25.514] 0.895 15.624 110 39(25.514, 121.84] 0.808 60.443 110 34(121.84, 196.289] 0.890 163.116 110 39(196.289, 237.42] 0.915 218.723 110 30(237.42, 288.236] 0.957 266.102 110 26(288.236, 344.182] 0.935 312.709 110 30(344.182, 522.558] 0.910 374.840 110 25Supp. Table 7 Analysis of model performance stratified by tissue area (in squared millimeters) for the metastatic samples in the internal validation cohort. The number of slides and the number of positive slides within each area range is also provided. AUC Median Area Slides EGFR+Area Range -45- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 (0.391, 4.002] 0.703 2.195 65 14(4.002, 7.439] 0.693 5.701 64 21(7.439,11.164] 0.733 9.659 64 25(11.164, 14.865] 0.738 12.983 64 20(14.865, 22.303] 0.749 17.812 64 23(22.303, 40.354] 0.755 31.479 64 26(40.354, 66.395] 0.787 52.177 64 25(66.395, 127.773] 0.824 95.341 64 17(127.773, 210.601] 0.652 165.405 64 21(210.601, 531.0] 0.852 280.001 64 14EGFR Variant Analysis Supp. Table 8 Statistical significance of score distribution comparisons across mutation variants. Statistical significance was estimated using the 2-sample 2-sided Kolmogorov– Smirnov test. The p-values were corrected using the Bonferroni method. Sample 1 Sample 2 p-value Any Ex 19 Del 9.754 Any Ex 19 L858R 6.181 Any Ex 19 T790M 0.033 Any Ex 20 Ins 1.991 Any Other 5.276 Ex 19 Del Ex 19 L858R 1.026 -46- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Ex 19 Del Ex 19 T790M 0.014 Ex 19 Del Other 0.450 Ex 19 Del Ex 19 T790M 9.962 Ex 19 L858R Ex 20 Ins 0.449 Ex 19 L858R Other 5.916 Ex 19 L858R Other 0.967 Ex 19 T790M Ex 20 Ins 6.473 Ex 19 T790M Other 0.029 Ex 20 Ins Other 0.818 MSHS Scanner Analysis Supp. Table 9 Comparison of model outputs for slides scanned with different scanner vendors (N=224). Linear relationship was measured using the Pearson correlation coefficient alongside the calculated p-value for testing non-correlation. Scanner 1 Scanner 2 r p-value Philips Aperio 0.828 1.3e-57 Philips Pramana 0.832 1.2e-58 Aperio Pramana 0.935 5.2e-102 TCGA Artifact Analysis -47- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Supp. Table 10 Summary of slide artifacts for the TCGA dataset curated by a thoracic pathologist. Artifact N Red Saturation 437 Blur 338 Low Quality Stain 325 Freeze Artifacts 265 Severe Artifacts 230 Tissue Necrosis 208 Blue Saturation 63 C. Systems and Methods for Classifying Biomedical Images for Executing Operations FIG. 16 depicts a block diagram of a system 100 for identifying biomarkers in subject tissue samples. In a brief overview, the system 100 can include at least one data processing system 105, at least one imaging device 110, at least one administrative device 115, at least one testing device 120, and at least one database 155, among others, communicatively coupled via at least one network 125. The data processing system 105 can include at least one data retriever 130, at least one model trainer 135, at least one model applier 140, at least one output evaluator 145, and at least one machine learning (ML) architecture 150, among others. The ML architecture 150 may include at least one patch encoder 165, at least one aggregator 170, and at least one classifier 175, among others. Each of the components of the system 100 can be implemented using the computing system as described in Section D. The system 100 may be used to carry out the functionalities described herein in Sections A and B. In further detail, the data processing system 105 can be any computing device comprising one or more processors coupled with memory and software capable of performing the various processes and tasks described herein. The data processing system 105 can be housed within -48- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 a computing system (e.g., laptop, PC, smart device) or within a server group (e.g., a data center, a branch office, or a server site), and include instructions to manage the identifying of images, generating a classification, and storing an association. The data processing system 105 can be in communication with the imaging device 110, administrative device 115, the testing device 120, and the database 155, among others. The data processing system 105 can include or execute any number of modules, processes, components, or subcomponents to performing the various processes and tasks described herein. On the data processing system 105, the data retriever 130 within the data processing system 105 can receive, retrieve, or otherwise identify a dataset from the imaging device 110 and the administrative device 115. The model trainer 135 can train, establish or otherwise initialize the ML architecture 150. The model applier 140 can apply, execute, or otherwise run the ML architecture 150. The output evaluator 145 can generate, determine, or otherwise identify an output based on a classification. The ML architecture 150 can be any type of ML algorithm or model to analyze tissue samples to identify or determine the presence of one or more biomarkers. The ML architecture 150 can be maintained on the data processing system 105. The ML architecture 150 can be, for example, a deep learning artificial neural network (ANN) (e.g., an encoder-decoder model with a convolution neural network architecture, a transformer architecture, a diffusion model, or an encoder-classifier model), among others. In some embodiments, the ML architecture 150 can also include, for example, a clustering algorithm (e.g., K-means clustering), a support vector machine (SVM), a Naïve Bayesian classifier, a decision tree, or a random forest classifier, among others. In general, the ML architecture 150 can have a tissue sample from a subject (e.g., a human) on a slide (e.g., a glass slide, etc.) as an input and a determination of the presence or absence of a biomarker (e.g., indicative of cancer, etc.) as an output. The ML architecture 150 may have been initialized, trained, and established using training data in accordance with learning techniques (e.g., supervised or semi-supervised). The training data can include or identify images or non-image data set and a label for the subject. In some embodiments, the ML architecture 150 may include the models described herein in conjunction with FIGs.1A and 1B. -49- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 The ML architecture 150 can include at least one patch encoder 165, at least one aggregator 170, and at least one classifier 175, among others. The patch encoder 165 can execute dimensionality reduction, feature extraction, or sequence modeling to generate, determine, or otherwise create a set of embeddings. Specifically, the ML architecture 150 may generate a plurality of patches using a biomedical image, each patch including a portion of the biomedical image. The ML architecture 150 may select, from the plurality of patches, a subset of patches corresponding to a region of interest (ROI) in the biomedical image. The patch encoder 165 may therefore be configured to generate, for each patch of the subset of patches, a respective feature vector of a plurality of feature vectors. The aggregator 170 can combine the feature vectors from the patch encoder 165. The classifier 175 can generate, determine, or otherwise identify a classification for the value indicating the presence of absence of the biomarker. In some embodiments, at least a portion of the ML architecture 150 (e.g., the patch encoder 165) may be pre-trained. In some embodiments, the ML architecture 150 may lack the aggregator 170. In some embodiments, the ML architecture 150 may be an instance of the model architecture shown in FIG. 3A. The imaging device 110 can be any device capable of acquiring images of tissue samples of subjects. The image can be captured via various techniques such as fluorescence microscopy, phase-contrast microscopy, bright-field microscopy, confocal microscopy, scanning electron microscopy, transmission electron microscopy, quantitative phase imaging, and automated digital microscopy, among others. It should be understood that other imaging modalities besides those listed above may be supported by the data processing system 105. The imaging device 110 can be in communication with the data processing system 105 and the administrative device 115 to provide acquired images. The administrative device 115 can be any device comprising one or more processors coupled with memory and software and capable of providing an output projection image. The administrative device 115 can be associated with an entity (e.g., clinician, physician, doctor) examining the subject or biomedical images from the subject. The administrative device 115 can be in communication with the data processing system 105 and the imaging device 110 to -50- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 exchange data. The administrative device 115 can display images acquired from the imaging device 110 on a display. The testing device 120 may be any device to perform genetic testing using biological samples. The testing device 120 may perform at least one of next-generation sequence (NGS) or rapid testing. NGS testing may include any sequencing method that determines the nucleotide sequence of individual nucleic acid molecules (e.g., in single molecule sequencing) or clonally expanded proxies for individual nucleic acid molecules in a high throughput parallel fashion (e.g., greater than 103, 104, 105 or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of the nucleic acid species in the library can be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment. When performing NGS, the testing device 120 may carry output genetic sequencing on a gene segment (e.g., deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sample) in a sample taken from a subject and generate sequencing data using the genetic sequencing. The genetic sequencing carried out may be a high throughput, massively parallel sequencing technique (sometimes herein referred to as next-generation sequencing), such as whole genome sequencing (WGS), pyrosequencing, Reversible dye-terminator sequencing, SOLiD sequencing, Ion semiconductor sequencing, and Helioscope single molecule sequencing, among others. NGS testing may have a turnaround time (TAT) between 2 to 3 weeks. In some embodiments, using the genetic sequencing data, the testing device 120 may generate genomic dataset profiling the subject. In generating the genomic dataset, the testing device 120 may execute read alignment, variant calling, and gene expression quantification, among others. In some embodiments, the testing device 120 may be part of a genomic profiling platform. The testing device 120 may use the gene sequencing to generate a sequencing dataset. The sequencing dataset may lack identifiers for at least a portion of genes (e.g., due to differences in sequencing or acquisition protocols across the different genomic profiling platforms). The -51- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 sequencing dataset may be maintained using one or more files according to a format (e.g., FASTQ, BAM, SAM, BCL, or VCF formats). The testing device 120 may also be configured to perform rapid testing. Rapid testing may include a diagnostic test performed to detect or identify the presence of pathogens and / or biomarkers in a biological sample (e.g., blood, saliva, various tissues, etc.). Rapid tests may be performed quickly (e.g., at or below a threshold amount of time). Rapid testing may have a turnaround time of between 10 minutes to 72 hours. The turnaround time of rapid testing may be shorter than NGS testing. A rapid test may facilitate timely diagnosis and treatment. Rapid tests may be performed using techniques such as lateral flow assays, antigen tests, nucleic acid tests, and the like. Although rapid testing is described primarily as polymerase chain reaction (PCR)- based testing, other rapid testing that rely on tissue samples with similar turnaround time (e.g., 10 minutes to 72 hours) may be used, such as loop-mediated isothermal amplification (LAMP)-based testing, recombinase polymerase amplification (RPA) based testing, or nucleic acid sequence- based amplification (NASBA) based testing, among others. The database 155 may store and maintain various resources and data associated with the data processing system 105, the imaging device 110, the administrative device 115, and the testing device 120, among others. The database 155 may include a database management system (DBMS) to arrange and organize the data maintained thereon. The database 155 may be in communication with the data processing system 105, the imaging device 110, the administrative device 115, and the testing device 120, via the network 125. While running various operations, the data processing system 105, the testing device 120, and the administrative device 115 may access the database 155 to retrieve identified data therefrom. The data processing system 105, the imaging device 110, the administrative device 115, and the testing device 120 may also write data onto the database 155 from running such operations. FIG.17 depicts a block diagram of a process 200 to train a machine learning (ML) architecture to determine a value indicating one of a presence or an absence of a biomarker associated with a condition (e.g., cancer) for a subject. Under the process 200, the data retriever -52- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 130 can retrieve, receive, obtain, or otherwise extract training data 205 from the database 155. The training data 205 can be used to initialize, train, and establish the ML architecture 150. The training data 205 can be for a subject 210 and be associated with a biological sample 215 retrieved from the subject 210 and placed on a slide 211 (e.g., a glass slide, a plastic slide, etc.). The training data 205 can include a plurality of examples to train the ML architecture 150. Each example of the training data 205 can include one or more of at least one biomedical image 218 and at least one label 235 for the subject 210, among others. The subject 210 (sometimes herein referred to as a sample subject 210 when associated with the training data 205) can be a human or animal subject, among others. The subject 210 can have, can be at risk of or diagnosed with a condition, such as cancer. The cancer can be at least one of carcinoma, sarcoma, hematopoietic cancer, adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, neuroblastoma, non- Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vascular tumor, or metastases thereof, among others. In some embodiments, the subject 210 may lack cancer or any condition. At least one biological sample 215 may be isolated, extracted, or otherwise obtained from an anatomical site associated with the condition. The anatomical site may include, for example, lung, breast, prostate, soft tissue, bone, bone marrow, blood, adrenal glands, bladder, bones, brain, breast, cervix, colon, colon, rectum, uterus, nasal cavity, pharynx, larynx, endometrium, esophagus, stomach, intestines, oral cavity, pharynx, larynx, lymph nodes, small intestine, colon, kidneys, larynx, blood, bone marrow, liver, lymph nodes, lymph nodes, spleen, lungs, skin, pleura, peritoneum, bone marrow, nasopharynx, adrenal glands, nerve tissue, lymph nodes, spleen, oral cavity, ovaries, pancreas, penis, pharynx, prostate gland, rectum, testes, skin, -53- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 stomach, ovaries, testes, testes, thyroid gland, uterus, vagina, blood vessels, or other organs to which the cancer metastasized to. A physician can extract, obtain, or otherwise retrieve the biological sample 215 from the subject 210. The biological sample 215 may be extracted from an anatomical site associated with the condition (e.g., cancer). For example, when testing for a biomarker indicative of lung cancer, tissue may be taken from the lung. Once extracted, the physician can place or secure the biological sample 215 on the slide 211, and then use or control the imaging device 110 to acquire images of the biological sample 215 on the slide 211. The biomedical image 218 may be of the slide 211 with the biological sample 215. The biomedical image 218 may be obtained or acquired (e.g., using the imaging device 110) in accordance with at least one imaging modality. The imaging modality may include, for example, whole slide imaging (WSI) or immunohistochemistry (IHC) imaging. For WSI imaging, the biological sample 215 may be stained with hematoxylin and eosin (H&E), periodic Acid-Schiff (PAS), trichrome, Giemsa stain, silver stain, toluidine blue, oil red O, or Alcian blue, among others. For IHC imaging, the biological sample 215 may be stained with a stain to increase contrast of cells associated with target biomarkers, such as chromogenic stains (e.g., DAB, AP red, AEC), fluorescent stains (e.g., DAPI), or counterstains (e.g., H&E), among others. The biomedical image 218 may be acquired in accordance with any number of imaging techniques, such as fluorescence microscopy, phase-contrast microscopy, bright-field microscopy, confocal microscopy, scanning electron microscopy, transmission electron microscopy, quantitative phase imaging, and automated digital microscopy, among others. In some embodiments, the biomedical image 218 may include a set of patches 220A–N (hereinafter generally referred to as patches 220). Each patch 220 of the biomedical image may be a portion of the image of the biological sample 215. An entire image of the biological sample 215 may be segmented into non-overlapping patches. The set of patches 220 can derived from at least one biological sample 215 from the subject 210. For example, each WSI or IHC image may be divided or segmented into a plurality of segments or patches. Each patch 220 may correspond to at least one cell in the biological sample 215. In various embodiment, a subset of the patches 220 may include a region of interest (ROI) of the biomedical image. The ROI may be -54- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 a portion of the image that includes cells or other structures to be further analyzed that may indicate or be used to identify a biomarker indicating a condition of the subject 210. For example, the ROI may include one or more portions of potentially cancerous cells. The label 235 can correspond to the classification of the subject 210 based on a value from the aggregator 170. The label 235 can identify or indicate the presence (or occurrence or detection) or the absence (or lack of occurrence or detection) of a biomarker within the biological sample 215. The biomarker being identified may be a gene associated with one or more of the types of cancers listed above. That is, the presence of one or more genes may indicate a likelihood that the subject 310 may have or get cancer of the type associated with the gene(s) / biomarker. The biomarker may include at least one of: AKT Serine / Threonine Kinase 1 (AKT1), Anaplastic Lymphoma Kinase (ALK), Adenomatous Polyposis Coli (APC), Androgen Receptor (AR), A-Raf Proto-Oncogene, Serine / Threonine Kinase (ARAF), AT-Rich Interaction Domain 1A (ARID1A), AT-Rich Interaction Domain 2 (ARID2), Ataxia Telangiectasia Mutated (ATM), Beta-2-Microglobulin (B2M), B-Cell CLL / Lymphoma 2 (BCL2), BCL6 Corepressor (BCOR), B-Raf Proto-Oncogene, Serine / Threonine Kinase (BRAF), Breast Cancer Type 1 Susceptibility Protein (BRCA1), Breast And Ovarian Cancer Susceptibility Protein 2 (BRCA2), Caspase Recruitment Domain Family Member 11 (CARD11), Core-Binding Factor Subunit Beta (CBFB), Cyclin D1 (CCND1), Cadherin 1 (CDH1), Cyclin Dependent Kinase 4 (CDK4), Cyclin Dependent Kinase Inhibitor 2A (CDKN2A), Capicua Transcriptional Repressor (CIC), CREB Binding Protein (CREBBP), CCCTC-Binding Factor (CTCF), Catenin Beta 1 (CTNNB1), Dicer 1, Ribonuclease III (DICER1), DIS3 Homolog, Exosome Endoribonuclease And 3'-5' Exoribonuclease (DIS3), DNA Methyltransferase 3 Alpha (DNMT3A), Epidermal Growth Factor Receptor (EGFR), Eukaryotic Translation Initiation Factor 1A X-Linked (EIF1AX), E1A Binding Protein P300 (EP300), Erb-B2 Receptor Tyrosine Kinase 2 (ERBB2), Erb-B2 Receptor Tyrosine Kinase 3 (ERBB3), ERCC Excision Repair 2, TFIIH Core Complex Helicase Subunit (ERCC2), Estrogen Receptor 1 (ESR1), Enhancer Of Zeste 2 Polycomb Repressive Complex 2 Subunit (EZH2), F-Box And WD Repeat Domain Containing 7 (FBXW7), Fibroblast Growth Factor Receptor 1 (FGFR1), Fibroblast Growth Factor Receptor 2 (FGFR2), Fibroblast Growth Factor -55- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Receptor 3 (FGFR3), Fibroblast Growth Factor Receptor 4 (FGFR4), Fms Related Receptor Tyrosine Kinase 3 (FLT3), Forkhead Box A1 (FOXA1), Forkhead Box L2 (FOXL2), Forkhead Box O1 (FOXO1), Far Upstream Element Binding Protein 1 (FUBP1), GATA Binding Protein 3 (GATA3), G Protein Subunit Alpha 11 (GNA11), G Protein Subunit Alpha Q (GNAQ), GNAS Complex Locus (GNAS), H3 Histone, Family 3A (H3F3A), Histone Cluster 1, H3b (HIST1H3B), HRas Proto-Oncogene, GTPase (HRAS), Isocitrate Dehydrogenase (NADP(+)) 1 (IDH1), Isocitrate Dehydrogenase (NADP(+)) 2 (IDH2), IKAROS Family Zinc Finger 1 (IKZF1), Inositol Polyphosphate Phosphatase Like 1 (INPPL1), Janus Kinase 1 (JAK1), Lysine Demethylase 6A (KDM6A), Kelch Like ECH Associated Protein 1 (KEAP1), KIT Proto-Oncogene, Receptor Tyrosine Kinase (KIT), Kinetochore Localized Astrin (SPAG5) Binding Protein (KNSTRN), KRAS Proto-Oncogene, GTPase (KRAS), Mitogen-Activated Protein Kinase Kinase 1 (MAP2K1), Mitogen-Activated Protein Kinase 1 (MAPK1), MYC Associated Factor X (MAX), Mediator Complex Subunit 12 (MED12), MET Proto-Oncogene, Receptor Tyrosine Kinase (MET), MutL Homolog 1 (MLH1), MutS Homolog 2 (MSH2), MutS Homolog 3 (MSH3), MutS Homolog 6 (MSH6), Mechanistic Target Of Rapamycin Kinase (MTOR), MYC Proto-Oncogene, BHLH Transcription Factor (MYC), MYCN Proto-Oncogene, BHLH Transcription Factor (MYCN), Myeloid Differentiation Primary Response 88 (MYD88), Myogenic Differentiation 1 (MYOD1), Neurofibromin 1 (NF1), NFE2 Like BZIP Transcription Factor 2 (NFE2L2), Notch Receptor 1 (NOTCH1), NRAS Proto-Oncogene, GTPase (NRAS), Neurotrophic Receptor Tyrosine Kinase 1 (NTRK1), Neurotrophic Receptor Tyrosine Kinase 2 (NTRK2), Neurotrophic Receptor Tyrosine Kinase 3 (NTRK3), Nucleoporin 93 (NUP93), P21 (RAC1) Activated Kinase 7 (PAK7), Platelet Derived Growth Factor Receptor Alpha (PDGFRA), Phosphatidylinositol-4,5- Bisphosphate 3-Kinase Catalytic Subunit Alpha (PIK3CA), Phosphatidylinositol-4,5- Bisphosphate 3-Kinase Catalytic Subunit Beta (PIK3CB), Phosphoinositide-3-Kinase Regulatory Subunit 1 (PIK3R1), Phosphoinositide-3-Kinase Regulatory Subunit 2 (PIK3R2), PMS1 Homolog 2, Mismatch Repair System Component (PMS2), DNA Polymerase Epsilon, Catalytic Subunit (POLE), Protein Phosphatase 2 Scaffold Subunit Aalpha (PPP2R1A), Protein Phosphatase 6 Catalytic Subunit (PPP6C), Protein Kinase C Iota (PRKCI), Patched 1 (PTCH1), Phosphatase And Tensin Homolog (PTEN), Protein Tyrosine Phosphatase Non-Receptor Type 11 (PTPN11), Rac -56- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Family Small GTPase 1 (RAC1), Raf-1 Proto-Oncogene, Serine / Threonine Kinase (RAF1), RB Transcriptional Corepressor 1 (RB1), Ret Proto-Oncogene (RET), Ras Homolog Family Member A (RHOA), Ras Like Without CAAX 1 (RIT1), ROS Proto-Oncogene 1, Receptor Tyrosine Kinase (ROS1), RAS Related 2 (RRAS2), Retinoid X Receptor Alpha (RXRA), SET Domain Containing 2, Histone Lysine Methyltransferase (SETD2), Splicing Factor 3b Subunit 1 (SF3B1), SMAD Family Member 3 (SMAD3), SMAD Family Member 4 (SMAD4), SWI / SNF Related BAF Chromatin Remodeling Complex Subunit ATPase 4 (SMARCA4), SWI / SNF Related BAF Chromatin Remodeling Complex Subunit B1 (SMARCB1), SOS Ras / Rac Guanine Nucleotide Exchange Factor 1 (SOS1), Speckle Type BTB / POZ Protein (SPOP), Signal Transducer And Activator Of Transcription 3 (STAT3), Serine / Threonine Kinase 11 (STK11), Serine / Threonine Kinase 19 (STK19), Transcription Factor 7 Like 2 (TCF7L2), Telomerase Reverse Transcriptase (TERT), Transforming Growth Factor Beta Receptor 1 (TGFBR1), Transforming Growth Factor Beta Receptor 2 (TGFBR2), Tumor Protein P53 (TP53), Tumor Protein P63 (TP63), TSC Complex Subunit 1 (TSC1), TSC Complex Subunit 2 (TSC2), U2 Small Nuclear RNA Auxiliary Factor 1 (U2AF1), Von Hippel-Lindau Tumor Suppressor (VHL), or Exportin 1 (XPO1), among others. The biomarkers identified in the label 235 may be associated with various types of cancers. The biomarkers can include sequence variants, gene fusions, and copy number alterations (e.g., amplifications, deletions, and codeletions). For instance, for carcinoma, the biomarkers may include EGFR, KRAS, TP53, NTRK1 / 2 / 3 fusions, ERBB2 (HER2) amplification, among others. For sarcoma, the biomarkers may include FOXO1, KIT, CIC, SS18-SSX (synovial sarcoma), FUS-DDIT3 (myxoid liposarcoma), EWSR1-partner fusions (e.g., EWSR1-FLI1), CIC-DUX4, BCOR-CCNB3, NAB2-STAT6 (solitary fibrous tumor), and MDM2 and / or CDK4 amplifications (dedifferentiated liposarcoma), among others. For hematopoietic cancer, the biomarkers may include BCL2, MYC, IKZF1, FLT3, IGH translocations (e.g., IGH-BCL2), and deletions such as del(17p) or del(13q), among others. For adrenal cancer, the biomarkers may include TP53, ARID1A, among others. For bladder cancer, the biomarkers may include FGFR3, HRAS, ERCC2, KDM6A, and occasional ERBB2 (HER2) amplification, among others. For blood cancer, the -57- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 biomarkers may include MYC, BCL2, FLT3, among others. For bone cancer, the biomarkers may include RB1, TP53, and EWSR1-partner fusions (e.g., EWSR1-FLI1 in Ewing sarcoma), among others. For brain cancer, the biomarkers may include IDH1, IDH2, TERT, H3F3A, 1p / 19q codeletion, EGFR amplification, CDKN2A deletion, and gene fusions such as FGFR3-TACC3 or KIAA1549-BRAF, among others. For breast cancer, the biomarkers may include BRCA1, BRCA2, ESR1, HER2 (ERBB2)—including ERBB2 amplification—PIK3CA, and CCND1 amplification, among others. For cervical cancer, the biomarkers may include FGFR3, TP53, among others. For colon cancer, the biomarkers may include APC, KRAS, TP53, PIK3CA, with rare NTRK fusions or ERBB2 amplification, among others. For colorectal cancer, the biomarkers may include APC, KRAS, TP53, PIK3CA, SMAD4, with occasional NTRK fusions or ERBB2 amplification, among others. For corpus uterine cancer, the biomarkers may include PTEN, among others. For ENT cancer, the biomarkers may include TP53, NOTCH1, with frequent CCND1 or EGFR amplifications in subsets, among others. For endometrial cancer, the biomarkers may include PTEN, ARID1A, PIK3CA, and ERBB2 (HER2) amplification in serous histologies, among others. For esophageal cancer, the biomarkers may include TP53, HER2 (ERBB2) (including ERBB2 amplification), among others. For gastrointestinal cancer, the biomarkers may include CDH1, PIK3CA, and fusions such as FGFR2 in cholangiocarcinoma, among others. For head and neck cancer, the biomarkers may include TP53, NOTCH1, EGFR and / or CCND1 amplification, and CRTC1-MAML2 fusions in mucoepidermoid carcinoma, among others. For Hodgkin’s disease, the biomarkers may include BCL6 and 9p24.1 amplification (CD274 / PDCD1LG2), among others. For intestinal cancer, the biomarkers may include APC, KRAS, among others. For kidney cancer, the biomarkers may include VHL, MET, PBRM1, chromosome 3p loss, and TFE3 / TFEB fusions (MiT family translocation renal cell carcinoma), among others. For larynx cancer, the biomarkers may include TP53, among others. For leukemia, the biomarkers may include BCR-ABL, FLT3, TP53, PML-RARA (t(15;17)), ETV6-RUNX1, KMT2A (MLL) rearrangements, CBFB-MYH11, and copy-number lesions such as del(5q) or del(7q), among others. For liver cancer, the biomarkers may include CTNNB1, TP53, ARID2, DNAJB1-PRKACA fusion (fibrolamellar carcinoma), and FGF19 amplification, among others. For lymph node cancer, the biomarkers may include BCL2, BCL6, and MYC rearrangements, -58- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 among others. For lymphoma, the biomarkers may include BCL2, MYD88, EZH2, IGH-BCL2, MYC, and concurrent rearrangements (e.g., “double-hit” MYC / BCL2), among others. For lung cancer, the biomarkers may include EGFR, ALK, KRAS, TP53, ROS1, RET, NTRK1 / 2 / 3 fusions, MET exon 14 skipping and / or MET amplification, and ERBB2 (HER2) amplification, among others. For melanoma, the biomarkers may include BRAF, NRAS, KIT, TP53, and CDKN2A deletion, among others. For mesothelioma, the biomarkers may include WT1, NF2, with frequent CDKN2A deletion and BAP1 alterations, among others. For myeloma, the biomarkers may include DIS3, B2M, IGH translocations (e.g., t(11;14) CCND1), 1q gain, and del(17p), among others. For nasopharynx cancer, the biomarkers may include EBV, among others. For neuroblastoma, the biomarkers may include ALK, MYCN amplification, 1p deletion, 11q deletion, and 17q gain, among others. For non-Hodgkin’s lymphoma, the biomarkers may include BCL2, MYC, EZH2, IGH-BCL2, and combined MYC / BCL2 / BCL6 rearrangements, among others. For oral cancer, the biomarkers may include TP53, among others. For ovarian cancer, the biomarkers may include BRCA1, BRCA2, FOXL2, and CCNE1 amplification, among others. For pancreatic cancer, the biomarkers may include KRAS, TP53, CDKN2A, among others. For penile cancer, the biomarkers may include HPV, among others. For pharynx cancer, the biomarkers may include TP53, NOTCH1, among others. For prostate cancer, the biomarkers may include AR, PTEN, TMPRSS2-ERG fusion, and AR amplification, among others. For rectal cancer, the biomarkers may include KRAS, TP53, PIK3CA, with occasional NTRK fusions or ERBB2 amplification, among others. For seminoma, the biomarkers may include OCT3 / 4, PLAP, and isochromosome 12p [i(12p)] / 12p amplification, among others. For skin cancer, the biomarkers may include BRAF, NRAS, and CDKN2A deletion, among others. For stomach cancer, the biomarkers may include HER2 (ERBB2)—including ERBB2 amplification—CDH1, FGFR2 amplification, and CLDN18-ARHGAP fusions, among others. For teratoma, the biomarkers may include AFP, hCG, and i(12p) / 12p amplification, among others. For testicular cancer, the biomarkers may include OCT3 / 4, PLAP, and i(12p) / 12p amplification, among others. For thyroid cancer, the biomarkers may include BRAF, RET (including RET / PTC fusions), KRAS, and NTRK1 / 3 fusions, among others. For uterine cancer, the biomarkers may include PTEN, PIK3CA, among others. For vaginal cancer, the biomarkers may include HPV, among others. For vascular tumor, the biomarkers may -59- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 include VEGF, MYC amplification (e.g., radiation-associated angiosarcoma), and fusions such as WWTR1-CAMTA1 or YAP1-TFE3 (epithelioid hemangioendothelioma), among others. Other biomarkers may also include features associated with prediction of survival outcomes and responses to specific targeted therapies, immunotherapies, or chemotherapeutic regimens. From the training data 205, the data retriever 130 can retrieve or identify the biomedical image 218. The data retriever 130 can produce or generate the set of patches 220 from the biomedical image 218. The data retriever 130 can partition or divide the biomedical image 218 into the set of patches 220. Each patch 220 can correspond to a respective portion of the biomedical image 218. The set of patches 220 can be partially overlapping or non-overlapping with one another. In some embodiments, the data retriever 130 can identify, or select a subset of patches 220’A–N (hereinafter generally referred to as patches 220’) from the patches 220. The patches 220’ may correspond to a region of interest (ROI) of the biomedical image 218. The selection can be based on the visual characteristics of the biomedical image 218 or by preprocessing the biomedical image 218. In some embodiments, the patches 220 can include a label (e.g., a segmentation mask) to identify the patches 220 as images having or being a region of interest (ROI) of the WSI. For example, the label can be a binary value assigned to the patches 220. The binary value can be a 0 or a 1, such that a 0 indicates an image that does not include an ROI and a 1 indicates an image that includes a ROI. The data retriever 130 can preprocess the patches 220 to remove noise, distortions, artifacts, and stains, among other blemishes associated with the image 218 to perform or execute contrast enhancement between the cells of the biological sample 215 and the background of the image 218. In some embodiments, the data retriever 130 can execute a segmentation or contouring model to identify at least one portion, region, or object of the biomedical image 218 to separate ROIs from the background and non-similar cells. The data retriever 130 can use thresholding algorithms (e.g., Otsu threshold or watershed algorithm), object detection, or edge detection, among others to execute the segmentation. In some embodiments, the data retriever 130 can extract, identify, or otherwise determine features (e.g., visual characteristics) from each cell -60- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 within the biomedical image 218 to distinguish the ROIs from non-ROIs. The features can include shape, size, color, intensity, or granularity, among other distinguishing features. The model trainer 135 can provide, feed, or otherwise apply the patches 220’ (or the biomedical image 218) to the ML architecture 150 for each example of the training data 205. The ML architecture 150 can include a plurality of weights to determine values indicating one of a presence or an absence of a biomarker. The model trainer 135 can fine-tune, update, or otherwise modify the plurality of weights associated with the ML architecture 150. The plurality of weights may be updated in accordance with a determines loss metric, as will be described herein. In some embodiments, at least a portion of the ML model is updated (e.g., in some embodiments, one or more portions of the ML model may be pre-trained). By modifying the plurality of weights, the model trainer 135 can improve the aggregator 170 to determine more accurate values indicating the presence or absence of a biomarker associated with the condition. The plurality of parameters can establish the architecture, design, or configuration of the ML architecture 150 to determine values indicating the presence or absence of a biomarker and a classification based on the value. Within the ML architecture 150, the plurality of weights can be arranged across the patch encoder 165, the aggregator 170, and the classifier 175. The plurality of parameters can include a learning rate, number of epochs, a batch size, a model architecture (e.g., deep learning convolutional neural network, ensemble network, or clustering algorithm, or any combination thereof), and regularization parameters. In some implementations, the plurality of parameters can include a plurality of model weights. The plurality of model weights can be variables learned from the training data 205 (e.g., one or more examples) that establish, dictate, or otherwise define a link between the input (e.g., patches 220’ containing a ROI) and the output (e.g., a classification). The model weights can continuously optimize during training to minimize a loss function of the ML architecture 150. The plurality of model parameters can include target weights or biases (e.g., influence the output of the ML architecture 150 based on accurate targets), or loss function (e.g., compound loss), among others. This is not limited to the patch encoder 165, but the weights of the ML architecture 150 can be arranged and adjusted for the aggregator 170 and the classifier 175 in a similar manner. -61- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 In applying the set of patches 220’, the model trainer 135 can process the set of patches 220’ using the plurality of weights of the ML architecture 150. The patch encoder 165 can extract, generate, or otherwise generate a plurality of feature vector 225A–N (referred to as feature vector 225 herein) using the set of patches 220’. For each patch 220’, the patch encoder 165 can calculate, determine, or otherwise generate at least one respective feature vector 225. Each feature vector 225 may be a lower dimensional representation of the latent features in the patch 220’, such as morphological features in the biological sample 215 or visual characteristics of the patch 220’ associated with the identification of the biomarker, among others. The patch encoder 165 can feed forward or provide the set of feature vectors 225 to the aggregator 170. In some embodiments, the patch encoder 165 can provide the set of feature vectors 225 to the classifier 175 (e.g., when the ML architecture 150 lacks the aggregator 170). The aggregator 170 can extract, generate, or otherwise generate a plurality of feature vectors 225’A–N (referred to as feature vector 225’ herein) using the plurality of feature vectors 225. Each feature vector 225 may have been generated independently from one another by the patch encoder 165. The set of feature vectors 225’ may be a combination or derivation of the plurality of feature vectors 225’. To generate, the aggregator 170 can join or combine the plurality of feature vectors 225. In some embodiments, the aggregator 170 can perform alignment operations to combine one or more of the feature vectors 225 based on feature similarity. In some embodiments, the aggregator 170 can join or concatenate the plurality of feature vectors 225 to generate the plurality of feature vectors 225’. The aggregator 170 can feed forward or provide the set of feature vectors 225 to the classifier 175. The classifier 175 can calculate, identify, or otherwise determine a value 240 based on the set of feature vectors 225’. The value 240 can indicate a likelihood of the presence or absence of the biomarker in the biological sample 215. To determine the value 240, the classifier 175 can evaluate a combination of the feature vector 225’ (or the set of feature vectors 225) by neural attention, pooling, bagging, boosting, stacking, majority voting, or weighted averaging (e.g., via attention scores), among other methods to determine the value 240. In some instances, the classifier 175 can establish a plurality of predictions of values 240 for the subject. By averaging -62- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 or combining the predictions, the classifier 175 can determine the value 240. In some embodiments, the classifier 175 can execute, for example, linear transformation, non-linear activation, Softmax functions, and sigmoid functions, among other functions, to combine the set of feature vectors 225’ to generate the value 240. Using the value 240, the classifier 175 can identify, determine, or otherwise generate the classification 245 for the biomedical image 218, the slide 211, the sample 215, or the subject 210. The classification 245 can corresponding to the biomarker associated with the condition in biological sample 215 on the slide 211. The classification 245 can indicate an absence, presence, or uncertainty of the biomarker being present (or absent) in the biological sample 215. To generate the classification 245, the classifier 175 can compare the value 240 with a set of ranges, including a first range of values (e.g., 0–10% probability) for classifying as absent, a second range of values for classifying as present (e.g., 75–100% probability), and a third range of values to classify as uncertain (e.g., 10–75% probability), among others. If the value satisfies (e.g. within) the first range, the classifier 175 may generate the classification 245 to indicate the absence of the biomarker in the biological sample 215. If the value satisfies (e.g. within) the second range, the classifier 175 may generate the classification 245 to indicate the presence of the biomarker in the biological sample 215. If the value satisfies (e.g. within) the third range, the classifier 175 may generate the classification 245 to indicate the uncertainty of the present (or absence) of the biomarker in the biological sample 215. In some embodiments, the classifier 175 can compare the value 240 with a threshold to determine the classification 245. The threshold may be used to generate a binary classification and may define a value at which to classify as present or absent. When the value 240 satisfies (e.g., greater than or equal to) the threshold, the classifier 175 can generate the classification 245 indicating presence of a biomarker. Conversely, when the value 240 does not satisfy (e.g., less than) the threshold, the classifier 175 can generate the classification 245 indicating absence of a biomarker. The model trainer 135 can calculate, generate, or otherwise determine at least on loss metric 250 to use to update the ML architecture 150. The loss metric 250 can indicate a discrepancy between the classification 245 and the label 235. In some embodiments, the loss -63- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 metric 250 may correspond to a loss function to quantify the difference between the classification 245 and the label 235 within the training data 205 to optimize ML architecture 150 during training. The loss function can be at least one of number of loss functions, such as a norm loss (e.g., L1 or L2), mean squared error (MSE), mean average error (MAE), a quadratic loss, a cross-entropy loss, and a Huber loss, or Wasserstein loss, among others. To determine the loss metric 250, the model trainer 135 can compare the classification 245 with the label 235 of the training data 205. Based on the comparison, the model trainer 135 can generate the loss metric 250. In general, when the classification 245 differs from the label 235, the higher the loss metric 250 may be. Conversely, when the classification 245 is the same as the label 235, the lower the loss metric 250 may be. The model trainer 135 can modify, change, or otherwise update at least one of the set of weights in the ML architecture 150 (e.g., the patch encoder 165, the aggregator 170, and the classifier 175) in accordance with the loss metric 250. The updating of the weights may be in accordance with a back propagation and optimization function (sometimes referred to herein as an objective function) with one or more parameters (e.g., learning rate, momentum, weight decay, and number of iterations). The optimization function may define one or more parameters at which the weights of the ML architecture 150 are to be updated. The optimization function may be in accordance with stochastic gradient descent, and may include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and adaptive gradient algorithm (AdaGrad), among others. The model trainer 135 can iteratively train the ML architecture 150 until convergence. In some embodiments, the patch encoder 165 may have been pretrained to generate proper feature vector 225, and the model trainer 135 may use the loss metric 250 to update the weights in the aggregator 170 or the classifier 175, or both. In some embodiments, the model trainer 135 may use the loss metric 250 to update the ML architecture 150 from end-to-end (e.g., including the patch encoder 165, the aggregator 170, and the classifier 175). In some embodiments, the model trainer 135 can adjust the plurality of weights so the ML architecture 150 can generate more accurate classifications 245. In some embodiments, the model trainer 135 can adjust the -64- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 plurality of weights so the ML architecture 150 can generate more accurate values 240. For example, the model trainer 135 can use the loss metric 250 to compute gradients for the direction and magnitude of parameter adjustments to minimize the loss metric 250. In another instance, the model trainer 135 can use optimization algorithms such as, stochastic gradient descent or RMS prop to update model parameters iteratively. In some implementations, the direction and step size of the parameter update according to the gradients and the optimization algorithm. Upon convergence, the model trainer 135 can store and maintain the set of weights for the set of layers of the ML architecture 150 for use in inference. FIG.18 depicts a block diagram of a process 300 to determine values indicating the presence or absence of a biomarker associated with a condition for a subject based on a tissue sample of the subject. Under the process 300, the imaging device 110 can produce or generate at least one biomedical image 318 of a slide 311 with a biological sample 315 obtained from a subject 310. The subject 310 can be a human or animal subject, among others. The subject 310 can have, can be at risk of or diagnosed with a condition, such as cancer. The cancer can be at least one of carcinoma, sarcoma, hematopoietic cancer, adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, neuroblastoma, non- Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vascular tumor, or metastases thereof, among others. In some embodiments, the subject 310 may lack cancer or any condition. At least one biological sample 315 may be isolated, extracted, or otherwise obtained from an anatomical site associated with the condition. The anatomical site may include, for example, lung, breast, prostate, soft tissue, bone, bone marrow, blood, adrenal glands, bladder, bones, brain, breast, cervix, colon, colon, rectum, uterus, nasal cavity, pharynx, larynx, -65- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 endometrium, esophagus, stomach, intestines, oral cavity, pharynx, larynx, lymph nodes, small intestine, colon, kidneys, larynx, blood, bone marrow, liver, lymph nodes, lymph nodes, spleen, lungs, skin, pleura, peritoneum, bone marrow, nasopharynx, adrenal glands, nerve tissue, lymph nodes, spleen, oral cavity, ovaries, pancreas, penis, pharynx, prostate gland, rectum, testes, skin, stomach, ovaries, testes, testes, thyroid gland, uterus, vagina, blood vessels, or other organs to which the cancer metastasized to. A physician can extract, obtain, or otherwise retrieve the biological sample 315 from the subject 310. The biological sample 315 may be extracted from an anatomical site associated with the condition (e.g., cancer). For example, when testing for a biomarker indicative of lung cancer, tissue may be taken from the lung. Once extracted, the physician can place or secure the biological sample 315 on the slide 311, and then use or control the imaging device 110 to acquire images of the biological sample 315 on the slide 311. The biomedical image 318 may be of the slide 311 with the biological sample 315. The biomedical image 318 may be obtained or acquired (e.g., using the imaging device 110) in accordance with at least one imaging modality. The imaging modality may include, for example, whole slide imaging (WSI) or immunohistochemistry (IHC) imaging. For WSI imaging, the biological sample 315 may be stained with hematoxylin and eosin (H&E), periodic Acid-Schiff (PAS), trichrome, Giemsa stain, silver stain, toluidine blue, oil red O, or Alcian blue, among others. For IHC imaging, the biological sample 315 may be stained with a stain to increase contrast of cells associated with target biomarkers, such as chromogenic stains (e.g., DAB, AP red, AEC), fluorescent stains (e.g., DAPI), or counterstains (e.g., H&E), among others. The biomedical image 318 may be acquired in accordance with any number of imaging techniques, such as fluorescence microscopy, phase-contrast microscopy, bright-field microscopy, confocal microscopy, scanning electron microscopy, transmission electron microscopy, quantitative phase imaging, and automated digital microscopy, among others. In some embodiments, the biomedical image 318 may include a set of patches 320A–N (hereinafter generally referred to as patches 320). Each patch 320 of the biomedical image may be a portion of the image of the biological sample 315. An entire image of the biological sample 315 may be segmented into non-overlapping patches. The set of patches 320 can derived -66- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 from at least one biological sample 315 from the subject 310. For example, each WSI or IHC image may be divided or segmented into a plurality of segments or patches. Each patch 320 may correspond to at least one cell in the biological sample 315. In various embodiment, a subset of the patches 320 may include a region of interest (ROI) of the biomedical image. The ROI may be a portion of the image that includes cells or other structures to be further analyzed that may indicate or be used to identify a biomarker indicating a condition of the subject 310 (e.g., EGFR for lung cancer). For example, the ROI may include one or more portions of potentially cancerous cells. The data retriever 130 can retrieve or identify the biomedical image 318. The data retriever 130 can produce or generate the set of patches 320 from the biomedical image 318. The data retriever 130 can partition or divide the biomedical image 318 into the set of patches 320. Each patch 320 can correspond to a respective portion of the biomedical image 318. The set of patches 320 can be partially overlapping or non-overlapping with one another. In some embodiments, the data retriever 130 can identify, or select a subset of patches 320’A–N (hereinafter generally referred to as patches 320’) from the patches 320. The patches 320’ may correspond to a region of interest (ROI) of the biomedical image 318. The selection can be based on the visual characteristics of the biomedical image 318 or by preprocessing the biomedical image 318. The data retriever 130 can preprocess the patches 320 to remove noise, distortions, artifacts, and stains, among other blemishes associated with the image 318 to perform or execute contrast enhancement between the cells of the biological sample 315 and the background of the image 318. In some embodiments, the data retriever 130 can execute a segmentation or contouring model to identify at least one portion, region, or object of the biomedical image 318 to separate ROIs from the background and non-similar cells. The data retriever 130 can use thresholding algorithms (e.g., Otsu threshold or watershed algorithm), object detection, or edge detection, among others to execute the segmentation. In some embodiments, the data retriever 130 can extract, identify, or otherwise determine features (e.g., visual characteristics) from each cell within the biomedical image 318 to distinguish the ROIs from non-ROIs. The features can include shape, size, color, intensity, or granularity, among other distinguishing features. -67- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 The model applier 140 can provide, feed, or otherwise apply the set of patches 320’ (or the biomedical image 318) the ML architecture 150. In applying the set of patches 320’, the model trainer 135 can process the set of patches 320’ using the plurality of weights of the ML architecture 150. The patch encoder 165 can extract, generate, or otherwise generate a plurality of feature vector 325A–N (referred to as feature vector 325 herein) using the set of patches 320’. For each patch 320’, the patch encoder 165 can calculate, determine, or otherwise generate at least one respective feature vector 325. Each feature vector 325 may be a lower dimensional representation of the latent features in the patch 320’, such as morphological features in the biological sample 315 or visual characteristics of the patch 320’ associated with the identification of the biomarker, among others. The patch encoder 165 can feed forward or provide the set of feature vectors 325 to the aggregator 170. In some embodiments, the patch encoder 165 can provide the set of feature vectors 325 to the classifier 175 (e.g., when the ML architecture 150 lacks the aggregator 170). The aggregator 170 can extract, generate, or otherwise generate a plurality of feature vectors 325’A–N (referred to as feature vector 325’ herein) using the plurality of feature vectors 325. Each feature vector 325 may have been generated independently from one another by the patch encoder 165. The set of feature vectors 325’ may be a combination or derivation of the plurality of feature vectors 325’. To generate, the aggregator 170 can join or combine the plurality of feature vectors 325. In some embodiments, the aggregator 170 can perform alignment operations to combine one or more of the feature vectors 325 based on feature similarity. In some embodiments, the aggregator 170 can join or concatenate the plurality of feature vectors 325 to generate the plurality of feature vectors 325’. The aggregator 170 can feed forward or provide the set of feature vectors 325 to the classifier 175. The classifier 175 can calculate, identify, or otherwise determine a value 340 based on the set of feature vectors 325’. The value 340 can indicate a likelihood of the presence or absence of a biomarker in the biological sample 315. To determine the value 340, the classifier 175 can evaluate a combination of the feature vector 325’ (or the set of feature vectors 325) by neural attention, pooling, bagging, boosting, stacking, majority voting, or weighted averaging (e.g., via attention scores), among other methods to determine the value 340. In some instances, -68- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 the classifier 175 can establish a plurality of predictions of values 340 for the subject. By averaging or combining the predictions, the classifier 175 can determine the value 340. In some embodiments, the classifier 175 can execute, for example, linear transformation, non-linear activation, Softmax functions, and sigmoid functions, among other functions, to combine the set of feature vectors 325’ to generate the value 340. Using the value 340, the classifier 175 can generate, determine, or otherwise identify the classification 345 for the biomedical image 318, the slide 311, the sample 315, or the subject 310. The classification 345 can corresponding to the biomarker associated with the condition in biological sample 315 on the slide 311. The classification 345 can indicate an absence, presence, or uncertainty of the biomarker being present (or absent) in the biological sample 315. To determine, the classifier 175 can compare the value 340 with a set of ranges, including a first range of values (e.g., 0–10% probability) for classifying as absent, a second range of values for classifying as present (e.g., 75–100% probability), and a third range of values to classify as uncertain (e.g., 10–75% probability), among others. If the value satisfies (e.g. within) the first range, the classifier 175 may generate the classification 345 to indicate the absence of the biomarker in the biological sample 315. If the value satisfies (e.g. within) the second range, the classifier 175 may generate the classification 345 to indicate the presence of the biomarker in the biological sample 315. If the value satisfies (e.g. within) the third range, the classifier 175 may generate the classification 345 to indicate the uncertainty of the present (or absence) of the biomarker in the biological sample 315. In some embodiments, the classifier 175 can compare the value 340 with a threshold to determine the classification 345. The threshold may be used to generate a binary classification and may define a value at which to classify as present or absent. When the value 340 satisfies (e.g., greater than or equal to) the threshold, the classifier 175 can generate the classification 345 indicating presence of a biomarker. Conversely, when the value 340 does not satisfy (e.g., less than) the threshold, the classifier 175 can generate the classification 345 indicating absence of a biomarker. FIG. 19 depicts a block diagram of a process 400 to provide instructions to testing devices based on the classification and output identifying the subject as a candidate or non- -69- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 candidate for administration of medication. Under the process 400, the output evaluator 145 can store, house, or otherwise maintain an association between the subject 310 and the classification 345. The association can be a link, a map, or a connection between the subject 310, the classification 345, and the value 340. To store the association, the output evaluator 145 can generate one or more data structures. The one or more data structures can include an array, a linked list, a stack, a tree, and a hash table, among others. For example, the data structure can be a hash table where the subject 310 is the key to the hash table and the classification 345 is the value of the hash table. The hash table can include a plurality of keys (e.g., for each subject 310) mapped to a plurality of values (e.g., classification 345 of the subject 310). In another example, the data structure can be a plurality of linked lists. A first linked list can correspond to a first subject 310. Each node in the first linked list corresponds to the value 340 of the first subject 310. Concurrently, a second linked list can correspond to a second subject 310. Each node in the second linked list corresponds to the value 340 of the second subject 310. In some implementations, the output evaluator 145 can store an association among the biomedical image 318, patches 320, patches 320’ including an ROI, the classification 345, and the value 340, among others, on the database 155. The output evaluator 145 may produce or generate at least one instruction 405 based on the classification 345 (or the value 340). When the classification 345 or the value 340 indicating an uncertainty of the presence of the biomarker in the biological sample 315, the output evaluator 145 may generate the instruction 405 for the testing device 120 to execute an operation to perform rapid testing (e.g., PCR-based rapid testing) using at least a portion (e.g., portion 408A) of the sample 315. When the classification 345 or the value 340 indicating a presence of the biomarker, the output evaluator 145 may generate the instruction 405 for the testing device 120 to execute an operation to perform next-generation sequencing (NGS) testing using at least a portion (e.g., portions 408A and 408B) of the biological sample 315. When the classification 345 or the value 340 indicates an absence of the biomarker, the output evaluator 145 may generate the instruction 405 to indicate to the testing device 120 to execute an operation to skip NGS testing and / or rapid testing using any portion of the biological sample 315. As shown in FIG. 19, the portion 408A of the biological sample 315 used for performing rapid testing may be a smaller than the portions -70- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 408A and 408B of the biological sample 315 used for performing NGS testing. With the generation, the output evaluator 145 may provide, send, or otherwise transmit the instruction 405 to the testing device 120 to cause the testing device 120 to perform the indicated operation. Upon receipt of the instruction 405, the testing device 120 may carry out or perform the indicated operation. When the instruction 405 specifies performance of rapid testing, the testing device 120 may perform rapid testing on the portion of the biological sample 315. For instance, upon display of the instruction 405, the clinician examining the biological sample 315 may remove the portion 408A of the biological sample 315 from the slide 311, and may carry out rapid testing using the testing device 120. By taking a smaller portion, more of the biological sample 315 may remain on the slide 311, and may be used for further testing (e.g., NGS testing). When the instruction 405 specifies performance of NGS testing, the testing device 120 may perform NGS testing on the portion of the biological sample 315. For example, upon display of the instruction 405, the clinician examining the biological sample 315 may remove the portion 408A and 408B (e.g., as one piece) of the biological sample 315 from the slide 311, and may carry out NGS testing using the testing device 120. When the instruction 405 specifies skipping of NGS or rapid testing, the testing device 120 may refrain from performing any testing. Upon performing one of NGS testing or rapid testing, the testing device 120 may output or generate at least one test result 410. The test result 410 may indicate one of the presence or absence of the biomarker in the biological sample 315 on the slide 311. When the NGS or rapid testing indicates the presence of the biomarker, the test result 410 may indicate a confirmation of the presence of the biomarker in the biological sample 315. Conversely, when the NGS or rapid testing indicates the absence of the biomarker, the test result 410 may indicate a refutation of the presence of the biomarker in the biological sample 315. The testing device 120 may return, send, or otherwise transmit the test result 410 to the output evaluator 145. The output evaluator 145 may identify or determine whether the subject 310 is a candidate for administration of therapy 420 for the condition based on the test result 410 received from the testing device 120. The therapy 420 can include at least one of a chemotherapy or targeted -71- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 therapy for treatment of the condition (e.g., the type of cancer) in the subject 310. When the test result 410 indicates or confirms the presence of the biomarker, the output evaluator 145 may determine that the subject 310 is the candidate for the administration of therapy 420. Conversely, when the test result 410 indicates the absence of the biomarker or refutes the initial determination of the presence (e.g., as indicated by the classification 345), the output evaluator 145 may determine that the subject 310 is the non-candidate for the administration of therapy 420. In some embodiments, the output evaluator 145 can identify or determine whether the subject 310 is a candidate for therapy 420 based on the classification 345. When the classification 345 corresponds to the presence of the biomarker, the output evaluator 145 may determine that the subject 310 is the candidate for the administration of therapy 420. Conversely, when the classification 245 indicates the absence of the biomarker, the output evaluator 145 may determine that the subject 310 is a non-candidate for the administration of therapy 420. In some embodiments, the output evaluator 145 may determine or identify the condition based on the detected biomarker. The determination may be based on a mapping between the condition (e.g., cancer) and one or more biomarkers. The biomarkers can include sequence variants, gene fusions, and copy number alterations (e.g., amplifications, deletions, and codeletions). For instance, for carcinoma, the biomarkers may include EGFR, KRAS, TP53, NTRK1 / 2 / 3 fusions, ERBB2 (HER2) amplification, among others. For sarcoma, the biomarkers may include FOXO1, KIT, CIC, SS18-SSX (synovial sarcoma), FUS-DDIT3 (myxoid liposarcoma), EWSR1-partner fusions (e.g., EWSR1-FLI1), CIC-DUX4, BCOR-CCNB3, NAB2-STAT6 (solitary fibrous tumor), and MDM2 and / or CDK4 amplifications (dedifferentiated liposarcoma), among others. For hematopoietic cancer, the biomarkers may include BCL2, MYC, IKZF1, FLT3, IGH translocations (e.g., IGH-BCL2), and deletions such as del(17p) or del(13q), among others. For adrenal cancer, the biomarkers may include TP53, ARID1A, among others. For bladder cancer, the biomarkers may include FGFR3, HRAS, ERCC2, KDM6A, and occasional ERBB2 (HER2) amplification, among others. For blood cancer, the biomarkers may include MYC, BCL2, FLT3, among others. For bone cancer, the biomarkers may include RB1, TP53, and EWSR1-partner fusions (e.g., EWSR1-FLI1 in Ewing sarcoma), among others. For brain cancer, -72- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 the biomarkers may include IDH1, IDH2, TERT, H3F3A, 1p / 19q codeletion, EGFR amplification, CDKN2A deletion, and gene fusions such as FGFR3-TACC3 or KIAA1549-BRAF, among others. For breast cancer, the biomarkers may include BRCA1, BRCA2, ESR1, HER2 (ERBB2)— including ERBB2 amplification—PIK3CA, and CCND1 amplification, among others. For cervical cancer, the biomarkers may include FGFR3, TP53, among others. For colon cancer, the biomarkers may include APC, KRAS, TP53, PIK3CA, with rare NTRK fusions or ERBB2 amplification, among others. For colorectal cancer, the biomarkers may include APC, KRAS, TP53, PIK3CA, SMAD4, with occasional NTRK fusions or ERBB2 amplification, among others. For corpus uterine cancer, the biomarkers may include PTEN, among others. For ENT cancer, the biomarkers may include TP53, NOTCH1, with frequent CCND1 or EGFR amplifications in subsets, among others. For endometrial cancer, the biomarkers may include PTEN, ARID1A, PIK3CA, and ERBB2 (HER2) amplification in serous histologies, among others. For esophageal cancer, the biomarkers may include TP53, HER2 (ERBB2) (including ERBB2 amplification), among others. For gastrointestinal cancer, the biomarkers may include CDH1, PIK3CA, and fusions such as FGFR2 in cholangiocarcinoma, among others. For head and neck cancer, the biomarkers may include TP53, NOTCH1, EGFR and / or CCND1 amplification, and CRTC1-MAML2 fusions in mucoepidermoid carcinoma, among others. For Hodgkin’s disease, the biomarkers may include BCL6 and 9p24.1 amplification (CD274 / PDCD1LG2), among others. For intestinal cancer, the biomarkers may include APC, KRAS, among others. For kidney cancer, the biomarkers may include VHL, MET, PBRM1, chromosome 3p loss, and TFE3 / TFEB fusions (MiT family translocation renal cell carcinoma), among others. For larynx cancer, the biomarkers may include TP53, among others. For leukemia, the biomarkers may include BCR-ABL, FLT3, TP53, PML-RARA (t(15;17)), ETV6-RUNX1, KMT2A (MLL) rearrangements, CBFB-MYH11, and copy-number lesions such as del(5q) or del(7q), among others. For liver cancer, the biomarkers may include CTNNB1, TP53, ARID2, DNAJB1-PRKACA fusion (fibrolamellar carcinoma), and FGF19 amplification, among others. For lymph node cancer, the biomarkers may include BCL2, BCL6, and MYC rearrangements, among others. For lymphoma, the biomarkers may include BCL2, MYD88, EZH2, IGH-BCL2, MYC, and concurrent rearrangements (e.g., “double-hit” MYC / BCL2), among others. For lung cancer, the biomarkers may include EGFR, ALK, KRAS, -73- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 TP53, ROS1, RET, NTRK1 / 2 / 3 fusions, MET exon 14 skipping and / or MET amplification, and ERBB2 (HER2) amplification, among others. For melanoma, the biomarkers may include BRAF, NRAS, KIT, TP53, and CDKN2A deletion, among others. For mesothelioma, the biomarkers may include WT1, NF2, with frequent CDKN2A deletion and BAP1 alterations, among others. For myeloma, the biomarkers may include DIS3, B2M, IGH translocations (e.g., t(11;14) CCND1), 1q gain, and del(17p), among others. For nasopharynx cancer, the biomarkers may include EBV, among others. For neuroblastoma, the biomarkers may include ALK, MYCN amplification, 1p deletion, 11q deletion, and 17q gain, among others. For non-Hodgkin’s lymphoma, the biomarkers may include BCL2, MYC, EZH2, IGH-BCL2, and combined MYC / BCL2 / BCL6 rearrangements, among others. For oral cancer, the biomarkers may include TP53, among others. For ovarian cancer, the biomarkers may include BRCA1, BRCA2, FOXL2, and CCNE1 amplification, among others. For pancreatic cancer, the biomarkers may include KRAS, TP53, CDKN2A, among others. For penile cancer, the biomarkers may include HPV, among others. For pharynx cancer, the biomarkers may include TP53, NOTCH1, among others. For prostate cancer, the biomarkers may include AR, PTEN, TMPRSS2-ERG fusion, and AR amplification, among others. For rectal cancer, the biomarkers may include KRAS, TP53, PIK3CA, with occasional NTRK fusions or ERBB2 amplification, among others. For seminoma, the biomarkers may include OCT3 / 4, PLAP, and isochromosome 12p [i(12p)] / 12p amplification, among others. For skin cancer, the biomarkers may include BRAF, NRAS, and CDKN2A deletion, among others. For stomach cancer, the biomarkers may include HER2 (ERBB2)—including ERBB2 amplification—CDH1, FGFR2 amplification, and CLDN18-ARHGAP fusions, among others. For teratoma, the biomarkers may include AFP, hCG, and i(12p) / 12p amplification, among others. For testicular cancer, the biomarkers may include OCT3 / 4, PLAP, and i(12p) / 12p amplification, among others. For thyroid cancer, the biomarkers may include BRAF, RET (including RET / PTC fusions), KRAS, and NTRK1 / 3 fusions, among others. For uterine cancer, the biomarkers may include PTEN, PIK3CA, among others. For vaginal cancer, the biomarkers may include HPV, among others. For vascular tumor, the biomarkers may include VEGF, MYC amplification (e.g., radiation-associated angiosarcoma), and fusions such as WWTR1-CAMTA1 or YAP1-TFE3 (epithelioid hemangioendothelioma), among others. Other biomarkers may also include features associated -74- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 with prediction of survival outcomes and responses to specific targeted therapies, immunotherapies, or chemotherapeutic regimens. The output evaluator 145 can create, produce, or otherwise generate the output 415 based on at least one of the value 340 or the classification 345. The output 415 can include, for example, a presentation based on the test result 410 of the testing of the biological sample 315. The output 415 can include, for example, the classification 345 for presentation to the physician, subject 310, or lab technician, among other clinicians, or the value 340 for presentation to the same. In some embodiments, the output evaluator 145 may generate the output 415 based on the identification of the subject as one of a candidate or non-candidate for the administration of the therapy 420. In some embodiments, the output evaluator 145 may select or identify at least one therapy 420 based on the cancer corresponding the biomarker detected to be present. For example, for lung cancer, the therapies may include chemotherapy (e.g., isplatin and carboplatin) and targeted therapy (e.g., crizotinib for ALK / ROS1-positive cancers or dacomitinib for EGFR- mutated cancers), among others. When the subject 310 is identified as the candidate, the output evaluator 145 can generate the output 415 to indicate that the subject is to be administered with the therapy 420. In some embodiments, the output 415 can be a notification to administer medication within the time interval or a notification for the subject 310 to request medical attention. When the subject 310 is identified as a non-candidate, the output evaluator 145 can generate the output 415 identifying the subject as a non-candidate for administration of a therapy for the condition associated with the biomarker (e.g., cancer). For example, responsive to the test result 410 indicating absence of the biomarker in the subject 310, the output evaluator 14 generates the output 415 indicating that the subject 310 is a non-candidate for administration of therapy. The output 415 may be a notification or indication to one or more parties (e.g., a physician, a technician, a clinician, the subject 310, etc.) that the subject 310 is a non-candidate. The subject 310 may then be withheld from the administration of the therapy subsequent to providing the output 415. -75- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 The output evaluator 145 can send, transmit, or otherwise or provide the output 415 for presentation at the user interface 425 of the administrative device 115. The output 415 can be provided as a system call (e.g., to cause the administrative device 115 to render a message), a notification, an alert, a text message, and electronic mail (e-mail), among other messages in a human-readable format. In some cases, the output 415 can include instructions to render an application for presentation on the user interface 425. The application can include the value 340, the classification 345, and the biomedical image 318, among other information associated with the subject 310. The message or notification within the output 415 can indicate that the subject 310 is a candidate or a non-candidate for the therapy 420. Upon receipt, the administration device 115 can render, display, or present the output 415 on the user interface 425. The administrative device 115 can present the user interface 425 based on the instructions within the output 415. The instruction within the output 415 can cause the user interface 425 to display the notification or message associated with the output 415. The information presented via the user interface 425 may include, for example, at least one of the biomedical image 318, the patches 320’ containing an ROI, the value 340, the classification 345, or indication of whether the subject 310 is a candidate or non-candidate for administration of a specified therapy 420, among others. When the output 415 identified the subject as a candidate for the administration of the therapy 420, the subject 310 can be provided or administered (e.g., by a clinician examining the subject 310 or by the subject 310) with a therapeutically effective amount of the therapy 420. When the output 415 identified the subject as the non-candidate for the administration of the therapy 420, the administration of the therapy 420 may be refrained or withheld from the subject 310. In some instances, the subject 310 may be directed (e.g., by the clinician) to perform other preventative measures, such as exercise or administration of medication for other conditions in the subject 310. The systems and methods described herein can be used to control or regulate the number of slides from which to take portions of biological samples for further testing and analysis in a rapid and speedy manner. The ML architecture can process patches from a given biomedical image of a slide containing a biological sample to infer the presence, absence, or uncertainty of a -76- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 target biomarker in the biological sample. As the ML architecture may be trained to initially process each patch individual and amalgamate the resultant features to derive the classification, the accuracy and precision in determining the presence of the biomarker may be improved, relative to other automated or semi-automated techniques. From a computer resource perspective, the use of the integrated ML architecture to first make a preliminary determination of the presence of absence of a biomarker improves or reduces computing processing times. Furthermore, this process may reduce the number of slides assigned to undergo rapid testing, preserving more tissue for next-generation sequencing (NGS) and increasing the NGS success rate. From a clinical perspective, by enabling timely and accurate analysis, the systems and methods herein may lead to better patient outcomes. For instance, early detection of EGFR mutations in non-small cell lung cancer (NSCLC) patients allows for targeted therapy with EGFR tyrosine kinase inhibitors (TKIs), improving progression-free survival and quality of life. The systems and methods herein can be deployed in a clinical setting. For instance, a patient at risk of a particular cancer (e.g., lung cancer) may visit a clinic, and a physician may collect a tissue sample from the patient’s organ (e.g., lung) and place the sample on a slide. The slide can be scanned using a whole slide imaging (WSI) device, producing a high-resolution digital image of the tissue sample. The digital image can be then divided into patches, with a subset of patches being provided to a ML architecture. The ML architecture may extract feature vectors from each patch, and use the feature vectors to derive a value indicating likelihood of presence or absence of the biomarker in the biological sample. Using the value, the ML architecture may also generate a classification indicating the presence, absence, or uncertainty of biomarkers associated with cancer in the tissue sample. The classification may then be used to perform additional testing. If the classification indicates the presence of biomarkers, next-generation sequencing (NGS) testing may be on the tissue sample for further analysis. If the classification indicates the absence of biomarkers, the NGS and rapid testing may be skipped, saving time and resources on the part of -77- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 the data processing system as well as staff at the clinic. If the classification indicates uncertainty, rapid testing may be performed on the tissue sample to confirm the presence or absence of biomarkers. The results of the testing may then be presented to the physician via a user interface. If the patient is identified as having biomarkers associated with the condition, the data processing system can function as a clinical decision support tool to recommend appropriate therapy, such as targeted therapy or chemotherapy. The physician examining the patient may decide whether to deliver the therapy to the patient. Based on the test results, the patient may receive timely and appropriate treatment, thereby potentially improving clinical outcomes FIGs. 20A and 20B depicts a flowchart of a method 500 of training the ML architecture to classify biomedical images. The method 500 can be implemented or performed by any of the components in the system 100 or system 700 as detailed herein. In brief overview, a computing system can identify a biomedical image (505). The computing system can apply a machine learning (ML) (510). The computing system can generate a classification for a biomarker (515). The computing system can determine certainty of presence of the biomarker based on the classification (520). If the biomarker is determined to be present, the computing system can provide an instruction to execute NGS testing (525). If the biomarker is uncertain, the computing system can provide an instruction to execute rapid testing (530). If the biomarker is determined to be absent, the computing system can provide an instruction to skip NGS or rapid testing (535). When a testing is run, the computing system can identify test results (540). The computing system can identify a presence or absence of the biomarker (455). The computing system can determine whether the biomarker is confirmed to be present (550). When confirmed to be absent, the computing system can determine that a subject as a non-candidate for administration of a therapy (555). When confirmed to be present, the computing system can determine that a subject as a non- candidate for administration of a therapy (560). The computing system can provide an output based on the determination (545). D: Computing and Network Environment -78- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 Various operations described herein can be implemented on computer systems. FIG.21 shows a simplified block diagram of a representative server system 700, client computing system 714, and network 726 usable to implement certain embodiments of the present disclosure. In various embodiments, server system 700 or similar systems can implement services or servers described herein or portions thereof. Client computing system 714 or similar systems can implement clients described herein. The system 700 described herein can be similar to the server system 700. Server system 700 can have a modular design that incorporates a number of modules 702 (e.g., blades in a blade server embodiment); while two modules 702 are shown, any number can be provided. Each module 702 can include processing unit(s) 704 and local storage 706. Processing unit(s) 704 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 704 can include a general- purpose primary processor as well as one or more special-purpose co-processors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 704 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 704 can execute instructions stored in local storage 706. Any type of processors in any combination can be included in processing unit(s) 704. Local storage 706 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 706 can be fixed, removable or upgradeable as desired. Local storage 706 can be physically or logically divided into various subunits, such as a system memory, a read-only memory (ROM), and a permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 704 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 704. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module -79- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 702 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections. In some embodiments, local storage 706 can store one or more software programs to be executed by processing unit(s) 704, such as an operating system and / or programs implementing various server functions, such as functions of the system 100 of FIG.16 or any other system described herein, or any other server(s) associated with system 100 or any other system described herein. “Software” refers generally to sequences of instructions that, when executed by processing unit(s) 704, cause server system 700 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 704. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 706 (or non-local storage described below), processing unit(s) 704 can retrieve program instructions to execute and data to process in order to execute various operations described above. In some server systems 700, multiple modules 702 can be interconnected via a bus or other interconnect 708, forming a local area network that supports communication between modules 702 and other components of server system 700. Interconnect 708 can be implemented using various technologies, including server racks, hubs, routers, etc. A wide area network (WAN) interface 710 can provide data communication capability between the local area network (interconnect 708) and the network 726, such as the Internet. Technologies can be used, including wired (e.g., Ethernet, IEEE 702.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 702.24 standards). -80- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 In some embodiments, local storage 706 is intended to provide working memory for processing unit(s) 704, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 708. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 712 that can be connected to interconnect 708. Mass storage subsystem 712 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 712. In some embodiments, additional data storage resources may be accessible via WAN interface 710 (potentially with increased latency). Server system 700 can operate in response to requests received via WAN interface 710. For example, one of the modules 702 can implement a supervisory function and assign discrete tasks to other modules 702 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 710. Such operation can generally be automated. Further, in some embodiments, WAN interface 710 can connect multiple server systems 700 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation. Server system 700 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG. 16 as client computing system 714. Client computing system 714 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on. For example, client computing system 714 can communicate via WAN interface 710. Client computing system 714 can include computer components such as processing unit(s) -81- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 716, storage device 718, network interface 720, user input device 722, and user output device 724. Client computing system 714 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like. Processing unit(s) 716 and storage device 718 can be similar to processing unit(s) 704 and local storage 706 described above. Suitable devices can be selected based on the demands to be placed on client computing system 714; for example, client computing system 714 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Client computing system 714 can be provisioned with program code executable by processing unit(s) 716 to enable various interactions with server system 700. Network interface 720 can provide a connection to the network 726, such as a wide area network (e.g., the Internet) to which WAN interface 710 of server system 700 is also connected. In various embodiments, network interface 720 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, LTE, etc.). User input device 722 can include any device (or devices) via which a user can provide signals to client computing system 714; client computing system 714 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 722 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on. User output device 724 can include any device via which client computing system 714 can provide information to a user. For example, user output device 724 can include a display to present images generated by or delivered to client computing system 714. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light- emitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog- -82- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that functions as both input and output device. In some embodiments, other user output devices 724 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on. Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer-readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer-readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operations indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 704 and 716 can provide various functionality for server system 700 and client computing system 714, including any of the functionality described herein as being performed by a server or client, or other functionality. It will be appreciated that server system 700 and client computing system 714 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 700 and client computing system 714 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments -83- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 of the present disclosure can be realized in a variety of apparatus, including electronic devices implemented using any combination of circuitry and software. -84- 4905-0844-9366.1

Claims

1. Atty. Dkt. No.: 115872-3307 WHAT IS CLAIMED IS:

1. A method of classifying biomedical images for executing operations, comprising: identifying, by one or more processors, a biomedical image of a slide with a biological sample obtained from a subject at risk of a condition; applying, by the one or more processors, the biomedical image to a machine learning (ML) model, wherein the ML model is established using a plurality of example biomedical images, each of the plurality of example biomedical images labeled with a respective indication of one of a presence or an absence of a biomarker associated with the condition; generating, by the one or more processors, based on applying the biomedical image to the ML model, a classification corresponding to the biomarker associated with the condition in biological sample on the slide; and executing, by the one or more processors, an operation with respect to the slide for testing of the biological sample, in accordance with the classification for the biomedical image.

2. The method of claim 1, wherein executing the operation further comprises executing the operation to perform rapid testing using at least a portion of the biological sample on the slide, responsive to the classification indicating an uncertainty of the presence.

3. The method of claim 1, wherein executing the operation further comprises executing the operation to skip at least one of rapid testing or next-generation sequencing (NGS) testing using the biological sample on the slide, responsive to the classification indicating the absence of the biomarker in the biological sample.

4. The method of claim 1, wherein executing the operation further comprises executing the operation to perform NGS testing using the biological sample on the slide, responsive to the classification indicating the presence of the biomarker in the biological sample.

5. The method of claim 1, further comprising: -85- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 identifying, by the one or more processors, a test result indicating one of the presence or the absence of the biomarker associated with the condition in the biological sample in slide, based on performing the operation with respect to the slide for the testing, wherein the operation comprises at least one of rapid testing or NGS testing; and providing, by the one or more processors, for presentation via a user interface, an output based on the test result of the testing of the biological sample of the slide.

6. The method of claim 5, further comprising determining, by the one or more processors, that the subject is a candidate for administration of therapy for the condition, responsive to the test result indicating the presence of the biomarker; wherein providing the output further comprises providing the output to indicate that the subject is subject is the candidate for the administration of therapy.

7. The method of claim 6, wherein the subject is administered with a therapeutic effective amount of the therapy for the condition, wherein the condition comprises cancer, and wherein the therapy comprises chemotherapy or targeted therapy for the cancer.

8. The method of claim 5, further comprising determining, by the one or more processors, that the subject is a non-candidate for administration of therapy for the condition, responsive to the test result indicating the absence of the biomarker, wherein providing the output further comprises providing the output to indicate that the subject is subject is the non-candidate for the administration of therapy, and wherein the subject is withheld from the administration of the therapy for the condition, subsequent to provision of the output.

9. The method of claim 1, further comprising determining, by the one or more processors, based on applying the biomedical image to the ML model, a value indicating a likelihood of the presence or the absence of the biomarker associated with the condition in the biological sample on the slide, and -86- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 wherein generating the classification further comprises generating, based on a comparison of the value with a threshold, the classification to indicate one of the presence, the absence, or an uncertainty.

10. The method of claim 1, further comprising: generating, by the one or more processors, a plurality of patches using the biomedical image, each patch of the plurality of patches comprising a respective portion of the biomedical image; and selecting, by the one or more processors, from the plurality of patches, a subset of patches corresponding to a region of interest (ROI) in the biomedical image; and wherein applying the biomedical image to the ML model further comprises applying the subset of patches to the ML model.

11. The method of claim 1, wherein applying the biomedical image to the ML model further comprises applying a subset of patches of the biomedical image to the ML model, the ML model comprising: a patch encoder configured to generate, for each patch of the subset of patches, a respective feature vector of a plurality of feature vectors; and a classifier configured to generate, using the plurality of feature vectors for the subset of patches, the classification corresponding to the biomarker associated with the condition in biological sample.

12. The method of claim 1, wherein the ML model is established by: identifying, from the plurality of example biomedical images, an example biomedical image of a respective slide with a respective biological sample; applying the example biomedical image to the ML model to generate a respective classification indicating one of the presence or the absence of the biomarker associated with the condition; -87- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 determining a loss metric based on a comparison between the respective classification and the respective indication associated with the example biomedical image; and updating at least one of a plurality of weights of at least a portion of the ML model in accordance with the loss metric.

13. The method of claim 1, wherein receiving the biomedical image further comprise receiving the biomedical image in accordance with at least one of a plurality of imaging modalities, wherein the plurality of imaging modalities comprises whole slide imaging (WSI) or immunohistochemistry (IHC) imaging, wherein the biological sample comprises tissue obtained from an anatomical site associated with the condition.

14. The method of claim 1, wherein the biomarker comprises at least one of: AKT1, ALK, APC, AR, ARAF, ARID1A, ARID2, ATM, B2M, BCL2, BCOR, BRAF, BRCA1, BRCA2, CARD11, CBFB, CCND1, CDH1, CDK4, CDKN2A, CIC, CREBBP, CTCF, CTNNB1, DICER1, DIS3, DNMT3A, EGFR, EIF1AX, EP300, ERBB2, ERBB3, ERCC2, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXA1, FOXL2, FOXO1, FUBP1, HER2, GATA3, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HRAS, IDH1, IDH2, IKZF1, INPPL1, JAK1, KDM6A, KEAP1, KIT, KNSTRN, KRAS, MAP2K1, MAPK1, MAX, MED12, MET, MLH1, MSH2, MSH3, MSH6, MTOR, MYC, MYCN, MYD88, MYOD1, NF1, NFE2L2, NOTCH1, NRAS, NTRK1, NTRK2, NTRK3, NUP93, PAK7, PDGFRA, PIK3CA, PIK3CB, PIK3R1, PIK3R2, PMS2, POLE, PPP2R1A, PPP6C, PRKCI, PTCH1, PTEN, PTPN11, RAC1, RAF1, RB1, RET, RHOA, RIT1, ROS1, RRAS2, RXRA, SETD2, SF3B1, SMAD3, SMAD4, SMARCA4, SMARCB1, SOS1, SPOP, STAT3, STK11, STK19, TCF7L2, TERT, TGFBR1, TGFBR2, TP53, TP63, TSC1, TSC2, U2AF1, VHL, or XPO1.

15. The method of claim 1, wherein the condition comprises cancer, wherein the cancer comprises at least one of carcinoma, sarcoma, hematopoietic cancer, adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon -88- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, neuroblastoma, non- Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vascular tumor, and metastases thereof.

16. A system for classifying biomedical images for executing operations, comprising: one or more processors coupled with memory, configured to: identify a biomedical image of a slide with a biological sample obtained from a subject at risk of a condition; apply the biomedical image to a machine learning (ML) model, wherein the ML model is established using a plurality of example biomedical images, each of the plurality of example biomedical images labeled with a respective indication of one of a presence or an absence of a biomarker associated with the condition; generate, based on applying the biomedical image to the ML model, a classification corresponding to the biomarker associated with the condition in biological sample on the slide; and execute an operation with respect to the slide for testing of the biological sample, in accordance with the classification for the biomedical image.

17. The system of claim 16, wherein the one or more processors are further configured to execute the operation to perform rapid testing using at least a portion of the biological sample on the slide, responsive to the classification indicating an uncertainty of the presence.

18. The system of claim 16, wherein the one or more processors are further configured to execute the operation to skip at least one of rapid testing or next-generation sequencing (NGS) testing -89- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 using the biological sample on the slide, responsive to the classification indicating the absence of the biomarker in the biological sample.

19. The system of claim 16, wherein the one or more processors are further configured to execute the operation to perform NGS testing using the biological sample on the slide, responsive to the classification indicating the presence of the biomarker in the biological sample.

20. The system of claim 16, wherein the one or more processors are configured to: identify a test result indicating one of the presence or the absence of the biomarker associated with the condition in the biological sample in slide, based on performing the operation with respect to the slide for the testing, wherein the operation comprises at least one of rapid testing or NGS testing; and provide, for presentation via a user interface, an output based on the test result of the testing of the biological sample of the slide.

21. The system of claim 20, wherein the one or more processors are configured to: determine that the subject is a candidate for administration of therapy for the condition, responsive to the test result indicating the presence of the biomarker; and provide the output to indicate that the subject is subject is the candidate for the administration of therapy.

22. The system of claim 21, wherein the subject is administered with a therapeutic effective amount of the therapy for the condition, wherein the condition comprises cancer, and wherein the therapy comprises chemotherapy or targeted therapy for the cancer.

23. The system of claim 20, wherein the one or more processors are configured to: determine that the subject is a non-candidate for administration of therapy for the condition, responsive to the test result indicating the absence of the biomarker, -90- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 provide the output to indicate that the subject is subject is the non-candidate for the administration of therapy, and wherein the subject is withheld from the administration of the therapy for the condition, subsequent to provision of the output.

24. The system of claim 16, wherein the one or more processors are configured to: determine, based on applying the biomedical image to the ML model, a value indicating a likelihood of the presence or the absence of the biomarker associated with the condition in the biological sample on the slide, and generate, based on a comparison of the value with a threshold, the classification to indicate one of the presence, the absence, or an uncertainty.

25. The system of claim 16, wherein the one or more processors are configured to: generate a plurality of patches using the biomedical image, each patch of the plurality of patches comprising a respective portion of the biomedical image; and select from the plurality of patches, a subset of patches corresponding to a region of interest (ROI) in the biomedical image; and apply the subset of patches to the ML model.

26. The system of claim 16, wherein the one or more processors are configured to apply a subset of patches of the biomedical image to the ML model, the ML model comprising: a patch encoder configured to generate, for each patch of the subset of patches, a respective feature vector of a plurality of feature vectors; and a classifier configured to generate, using the plurality of feature vectors for the subset of patches, the classification corresponding to the biomarker associated with the condition in biological sample.

27. The system of claim 16, wherein the ML model is established by: -91- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 identifying, from the plurality of example biomedical images, an example biomedical image of a respective slide with a respective biological sample; applying the example biomedical image to the ML model to generate a respective classification indicating one of the presence or the absence of the biomarker associated with the condition; determining a loss metric based on a comparison between the respective classification and the respective indication associated with the example biomedical image; and updating at least one of a plurality of weights of at least a portion of the ML model in accordance with the loss metric.

28. The system of claim 16, wherein the one or more processors are configured to receive the biomedical image in accordance with at least one of a plurality of imaging modalities, wherein the plurality of imaging modalities comprises whole slide imaging (WSI) or immunohistochemistry (IHC) imaging, wherein the biological sample comprises tissue obtained from an anatomical site associated with the condition.

29. The system of claim 16, wherein the biomarker comprises at least one of: AKT1, ALK, APC, AR, ARAF, ARID1A, ARID2, ATM, B2M, BCL2, BCOR, BRAF, BRCA1, BRCA2, CARD11, CBFB, CCND1, CDH1, CDK4, CDKN2A, CIC, CREBBP, CTCF, CTNNB1, DICER1, DIS3, DNMT3A, EGFR, EIF1AX, EP300, ERBB2, ERBB3, ERCC2, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXA1, FOXL2, FOXO1, FUBP1, HER2, GATA3, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HRAS, IDH1, IDH2, IKZF1, INPPL1, JAK1, KDM6A, KEAP1, KIT, KNSTRN, KRAS, MAP2K1, MAPK1, MAX, MED12, MET, MLH1, MSH2, MSH3, MSH6, MTOR, MYC, MYCN, MYD88, MYOD1, NF1, NFE2L2, NOTCH1, NRAS, NTRK1, NTRK2, NTRK3, NUP93, PAK7, PDGFRA, PIK3CA, PIK3CB, PIK3R1, PIK3R2, PMS2, POLE, PPP2R1A, PPP6C, PRKCI, PTCH1, PTEN, PTPN11, RAC1, RAF1, RB1, RET, RHOA, RIT1, ROS1, RRAS2, RXRA, SETD2, SF3B1, SMAD3, SMAD4, SMARCA4, -92- 4905-0844-9366.1 Atty. Dkt. No.: 115872-3307 SMARCB1, SOS1, SPOP, STAT3, STK11, STK19, TCF7L2, TERT, TGFBR1, TGFBR2, TP53, TP63, TSC1, TSC2, U2AF1, VHL, or XPO1.

30. The system of claim 16, wherein the condition comprises cancer, wherein the cancer comprises at least one of carcinoma, sarcoma, hematopoietic cancer, adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, neuroblastoma, non- Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vascular tumor, and metastases thereof. -93- 4905-0844-9366.1

Citation Information

Patent Citations

  • Artificial intelligence assisted precision medicine enhancements to standardized laboratory diagnostic testing

    US20210118559A1

  • Biological image transformation using machine-learning models

    US20220076067A1

  • Multiple instance learner for tissue image classification

    US20220237788A1

  • Non-tumor segmentation to support tumor detection and analysis

    US20220351379A1

  • Systems and methods for generating histology image training datasets for machine learning models

    US20230343074A1

Cited By

  • Exosome multi-parameter joint quantitative analysis method and device based on fluorescence intensity of coded microspheres

    CN122175984A