Biomarker detection engine(s) for predicting treatment outcomes for patients diagnosed with myeloid malignancies

The biomarker detection engine addresses the challenge of human-dependent diagnosis in myeloid malignancies by using deep learning to analyze bone marrow images, improving treatment prediction and personalization.

WO2026097010A1PCT designated stage Publication Date: 2026-05-07THE REGENTS OF THE UNIVERSITY OF COLORADO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE REGENTS OF THE UNIVERSITY OF COLORADO
Filing Date
2025-11-03
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current diagnostic approaches for myeloid malignancies rely heavily on human interpretation of bone marrow biopsy images, lacking computational systems to analyze morphological features and predict treatment responses, leading to delayed and suboptimal treatment selection.

Method used

A biomarker detection engine using deep learning models to analyze digital bone marrow images, automatically extracting morphological features and predicting treatment outcomes, enabling timely and personalized treatment plans.

Benefits of technology

Enhances computational efficiency and reduces latency in treatment decisions by objectively identifying predictive biomarkers, minimizing subjectivity and variability, and facilitating earlier and more precise treatment choices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025053762_07052026_PF_FP_ABST
    Figure US2025053762_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods herein provide a biomarker detection engine and its related functions. In an aspect, a biomarker detection engine receives a smear image of cells from a biopsy sample collected from a patient diagnosed with a myeloid malignancy. Responsive to receipt, the biomarker detection engine may preprocess the smear image to generate processed image data, for example by removing areas lacking cellular presence. Using the processed image data, the biomarker detection engine may generate, via one or more deep learning models, a prediction of a treatment outcome for a treatment therapy. The prediction may indicate whether the patient will be responsive to the treatment therapy. Based on the prediction, the biomarker detection engine may generate a recommendation for a treatment plan for the patient.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 303.0138WOBIOMARKER DETECTION ENGINE(S) FOR PREDICTING TREATMENTOUTCOMES FOR PATIENTS DIAGNOSED WITH MYELOID MALIGNANCIESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 715,721, titled BIOMARKER DETECTION ENGINE(S) FOR PREDICTING TREATMENT OUTCOMES FOR PATIENTS DIAGNOSED WITH MYELOID MALIGNANCIES, filed on November 4, 2024, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD

[0002] Aspects of the disclosure are related to the field of computer software applications and services and, in particular, to biomarker detection engines for detecting biomarkers within sample smears containing cells affected by a myeloid malignancy and predicting a treatment outcome for a respective patient in response to a treatment therapy based on the respective patient’s morphological variation.BACKGROUND

[0003] Myeloid malignancies are a group of cancers that arise from the bone marrow, where blood cells are produced. These malignancies include conditions such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndromes (MDS), and myeloproliferative neoplasms (MPN). In recent years, the prevalence of myeloid malignancies has been increasing, driven in part by an aging population and improvements in diagnostic techniques. These cancers primarily affect the production and function of myeloid cells — cells that differentiate into important components of the immune system, such as red blood cells, platelets, and certain types of white blood cells. As more individuals live longer, age-related genetic mutations in bone marrow cells are believed to contribute to the rising incidence of these diseases.

[0004] There are a variety of treatment options available for patients with myeloid malignancies, but the appropriate choice largely depends on the specific type and stage of the disease, as well as the patient’s overall health. Treatment strategies may include a variety of treatment therapies, including chemotherapy, targeted therapies, immunotherapy, bone marrow transplants, and in some cases, watchful waiting for slower-progressing conditions like certain myelodysplastic syndromes. For aggressive forms such as AML, early and intensive treatmentAttorney Docket No. 303.0138WO may be necessary, while chronic conditions like chronic myeloid leukemia (CML) often respond well to targeted therapies such as tyrosine kinase inhibitors (TKIs). However, determining the most effective treatment approach can be challenging due to the complexity of these diseases. Diagnostic uncertainties, overlapping symptoms, and the often-evolving nature of myeloid malignancies can delay the selection of a primary treatment. In some cases, this delay results in initiating treatment too late to adequately control or eliminate the malignancy, limiting the effectiveness of therapeutic options and negatively impacting patient outcomes. In addition, response rates for particular treatment therapies are not 100%, and the ability to predict responses is very limited.

[0005] Accordingly, there is a need for a biomarker detection engine, and its related functions, for predicting treatment outcomes for a patient and recommending a treatment plan tailored to the patient’s specific biomarkers present in the affected cells.SUMMARY

[0006] Systems and methods for providing a biomarker detection engine and its related functions are provided herein. As will be expanded on below, the biomarker detection engine may receive an image from a bone marrow biopsy and / or aspirate smear collected from a patient. The image may include the affected cells. Responsive to receiving the image, the biomarker detection engine may preprocess the image to generate image data. Preprocessing the image may include performing one or more image enhancement operations such as normalization, contrast adjustment, color-space conversion, and geometric transformations to standardize the input data and improve analysis consistency. In some cases, preprocessing may also include identifying segments of the image lacking cellular presence and removing those segments. In such an embodiment, the biomarker detection engine may segment the image into multiple regions, each region containing a grid formed by a collection of tiles. The biomarker detection engine may process each region to identify tiles lacking cellular presence and remove those tiles from the image, thereby generating the image data.

[0007] Once the image data is generated, the biomarker detection engine may generate, using one or more deep learning models, a prediction of a treatment outcome for a treatment therapy based on the image data. In some cases, the one or more deep learning models may include a convolutional neural network that is trained to predict whether a patient, based on morphological variations present in the image data, will respond positively to a selected treatment therapy. Based on the prediction, the biomarker detection engine may generate a recommended treatment plan for the patient. The recommendation may be provided to aAttorney Docket No. 303.0138WO provider, via a respective client device, and together, the provider and patient can determine a treatment plan best tailored to the patient based on this information and other relevant clinical data.

[0008] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Technical Disclosure. It may be understood that this Overview is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Many aspects of the disclosure may be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure. Moreover, in the drawings, like reference numerals designate corresponding parts throughout the several views. While several embodiments arc described in connection with these drawings, the disclosure is not limited to the embodiments disclosed herein. On the contrary, the intent is to cover all alternatives, modifications, and equivalents.

[0010] Figure 1 illustrates an operational environment for providing a biomarker detection engine, according to an embodiment herein.

[0011] Figure 2 illustrates an example biomarker detection engine, according to an embodiment provided herein.

[0012] Figure 3 illustrates a process for providing a biomarker detection engine and its related functions, according to an embodiment herein.

[0013] Figure 4 illustrates an example smear image created from a bone marrow biopsy, according to an embodiment herein.

[0014] Figure 5 illustrates an example image of a region containing a grid of tiles, according to an embodiment herein.

[0015] Figure 6 illustrates an example image of the region from Figure 5 with empty tiles removed, according to an embodiment herein.

[0016] Figure 7 shows an example client device suitable for providing a biomarker detection engine and related functions, according to an embodiment herein.DETAILED DESCRIPTION

[0017] Myeloid malignancies, a group of cancers originating in the bone marrow, include conditions such as AML, CML, MDS, and MPN. The prevalence of these cancers has beenAttorney Docket No. 303.0138WO increasing, driven by an aging population and advances in diagnostic techniques. These malignancies disrupt the production and function of myeloid cells, which are critical for immune system function. With more individuals living longer, age-related genetic mutations in bone marrow cells are thought to contribute to this rise in incidence. Treatment options vary depending on the specific type and stage of the malignancy, and may include chemotherapy, targeted therapies, immunotherapy, and bone marrow transplants. However, determining the most appropriate treatment therapy can be complex. The evolving nature of these diseases, along with diagnostic uncertainties, can lead to delays in selecting a primary treatment therapy. In some cases, treatment therapies may be initiated too late to adequately control or cure the malignancy, underscoring the importance of early diagnosis and careful evaluation in improving patient outcomes.

[0018] Recent advancements in the treatment of myeloid malignancies have increasingly focused on identifying specific molecular alterations or mutations within the affected cells, allowing for more personalized and targeted therapeutic approaches. These breakthroughs have led to the development of treatment therapies that are tailored to the genetic makeup of the cancer, improving outcomes for some patients. Tor example, targeted therapies, such as FLT3 inhibitors for acute myeloid leukemia, are designed to address specific mutations driving the disease. However, analyzing the mutational data for a patient can be a time-intensive process, often requiring complex genomic sequencing and interpretation. This can delay the initiation of treatment, especially in aggressive cases where time is critical. Moreover, the majority of patients do not have a genomically-targeted therapy available. The need to balance the benefits of personalized medicine with the urgency of prompt treatment highlights a challenge in modem cancer care, where waiting for detailed genetic analysis can sometimes come at the cost of lost treatment opportunities, potentially affecting patient survival and prognosis.

[0019] Current approaches to treatment therapy selection for myeloid malignancies rely heavily on human interpretation of diagnostic data and are limited by the absence of any software-based or computationally automated means of predicting a patient’s response to a specific therapy. Physicians and pathologists typically review a combination of laboratory results, cytogenetic panels, and bone marrow biopsy images to inform treatment planning. While molecular testing can identify certain driver mutations, it does not provide a comprehensive predictive framework for therapeutic response across diverse patient populations. Moreover, no existing diagnostic or clinical decision-support systems are configured to computationally evaluate raw image data from bone marrow samples to infer treatment outcomes. Available digital pathology platforms perform basic image visualizationAttorney Docket No. 303.0138WO or annotation but do not execute algorithmic feature extraction, pattern recognition, or predictive modeling to correlate morphological characteristics with therapeutic efficacy. As a result, the determination of which treatment therapy is most likely to succeed remains primarily a qualitative and experience-dependent judgment made by human experts.

[0020] This reliance on human observation imposes inherent technical and practical constraints. Bone marrow biopsy images contain millions of individual cells exhibiting subtle variations in nuclear morphology, cytoplasmic texture, and cellular organization — patterns that are statistically correlated with underlying genomic alterations and treatment outcomes but imperceptible to the human eye. Even experienced hematopathologists cannot manually discern or quantify such high-dimensional morphological signatures, nor can they systematically integrate image-based features with genomic or clinical data at the speed required for time-sensitive treatment decisions. The absence of computational systems capable of performing this type of multimodal data analysis creates a technological gap: there is presently no automated mechanism that can process digital biopsy images to detect latent biomarkers, compute treatment-response likelihoods, or generate objective, reproducible recommendations for therapeutic selection. This deficiency reflects a fundamental limitation in existing diagnostic technology, which is constrained by human perceptual bandwidth and the lack of specialized software architectures for predictive image analysis.

[0021] To address at least the above shortcomings of conventional approaches to selecting a treatment therapy for a patient having a myeloid malignancy diagnosis, a biomarker detection engine is provided. The biomarker detection engine is implemented as a computer- implemented system configured to analyze one or more images, such as a digital bone marrow biopsy image, containing affected cells and to predict the respective patient’s probability of achieving a positive outcome from one or more treatment therapies. In operation, the biomarker detection engine may receive one or more digital images derived from bone marrow samples, including but not limited to bone marrow aspirate smears, core biopsy sections, or touch preparation slides. The biomarker detection engine may execute automated image preprocessing and segmentation to isolate cellular regions of interest, and identify one or more biomarkers or morphologic alterations present within those regions. Based on the presence of various biomarkers, the biomarker detection engine may generate a recommendation for a treatment therapy that is tailored to the patient’s disease. As such, the biomarker detection engine may enable providers to select a treatment therapy that provides the highest probability of success for the patient, thereby improving patient survival and prognosis. The biomarkerAttorney Docket No. 303.0138WO detection engine is further configured to output these results substantially in real time following image submission, thereby enabling clinicians to make timely treatment decisions.

[0022] From a technological standpoint, the biomarker detection engine represents an improvement over conventional diagnostic computing systems. Unlike standard image storage or visualization tools, the disclosed biomarker detection engine integrates specialized artificial intelligence (Al) and machine-learning modules that execute data transformations not achievable through manual analysis or conventional software. For example, the biomarker detection engine may employ deep convolutional or transformer-based neural architectures trained on annotated biopsy image datasets to automatically extract and classify subtle morphological features correlated with treatment response. These operations allow the biomarker detection engine to detect latent, non-obvious image biomarkers that are imperceptible to human observers, while simultaneously increasing processing throughput and reducing the latency associated with manual review and laboratory sequencing workflows. The architecture thereby enhances the computational efficiency and functional capability of digital pathology systems, providing a specific and non-generic improvement in medical image analysis technology.

[0023] In practical clinical use, the biomarker detection engine enables a more objective, reproducible, and rapid approach to treatment selection. By automatically identifying image- derived biomarkers predictive of therapeutic efficacy, the system minimizes subjectivity in interpretation, reduces reliance on post hoc genetic testing, and supports early initiation of the most promising treatment regimen. The real-time inference capability allows providers to obtain actionable insight directly from routinely collected biopsy images, eliminating days or weeks of delay associated with traditional molecular assays. Additionally, because the biomarker detection engine standardizes feature extraction and predictive modeling, it ensures consistent results across institutions and operators, reducing diagnostic variability. Collectively, the biomarker detection engine facilitates earlier and more precise treatment decisions, improve patient management efficiency, and enhance overall clinical workflow performance without requiring modification of existing imaging or laboratory infrastructure.

[0024] Turning now to Figure 1, Figure 1 illustrates an operational environment 100 for providing a biomarker detection engine, according to an embodiment herein. In particular, the environment 100 illustrates the biomarker detection engine 114 leveraged by a provider to determine a treatment plan 120 for a patient 102. The patient 102 may be diagnosed with a myeloid malignancy, such as AML and as such may be seeking treatment with guidance from the provider. For ease of explanation, the following discussion focuses on the specific myeloidAttorney Docket No. 303.0138WO malignancy of AML and a treatment therapy of a combination of venetoclax and azacitidine, however it should be appreciated that the following discussion is equally applicable to other forms of myeloid malignancies and their respective treatment therapies.

[0025] As those skilled in the art readily appreciate, AML is a type of cancer that originates in the bone marrow and affects the myeloid cells responsible for producing red blood cells, white blood cells, and platelets. In AML, immature myeloid cells, called blasts, proliferate uncontrollably and fail to mature into functional blood cells, leading to a variety of symptoms such as fatigue, frequent infections, and easy bruising or bleeding. Treatment for AML often requires aggressive therapies, particularly for older or high-risk patients. A combination of venetoclax, a BCL-2 inhibitor, and azacitidine, a hypomethylating agent, has become a therapeutic standard for individuals diagnosed with AML, especially those who are ineligible for intensive chemotherapy. Venetoclax works by promoting the death of cancerous cells by inhibiting proteins that allow them to survive, while azacitidine interferes with the DNA of cancer cells, reducing their ability to grow and divide. Together, this combination therapy has shown to improve survival rates and produce durable remissions, offering a promising treatment therapy for many patients with AML. This specific treatment therapy for AML that includes a combination of venetoclax and azacitidine is referred to herein as a VEN / AZA treatment.

[0026] While most patients with AML respond positively to the combination of venetoclax (VEN) and azacitidine (AZA), a substantial proportion are refractory to this therapy, meaning their disease does not respond or relapses soon after treatment. These refractory patients face significantly poorer outcomes, as the available treatment options become more limited, and the aggressive nature of AML often leads to rapid disease progression. For these individuals, alternative therapies, such as salvage chemotherapy or experimental treatments in clinical trials, are sometimes considered, but these approaches generally offer lower success rates. Refractory AML patients managed with the VEN / AZA treatment typically have worse overall survival and reduced quality of life compared to those who achieve remission with this regimen.

[0027] While recent advancements in the study of myeloid malignancies have suggested an association between molecular alterations in specific genes, such as NPM1, IDH1 / 2, NRAS, KRAS, and TP53, and a patient’s potential responsiveness to certain treatment, capturing mutational data is both time and cost intensive. That is, recent studies indicate that genetic mutations can provide valuable insight into how the disease might progress and which therapies may be more effective. For instance, patients with NPM1 mutations may respond better toAttorney Docket No. 303.0138WO certain targeted therapies, while alterations in TP53 are often associated with poorer outcomes and resistance to standard treatments. However, identifying these mutations requires genomic sequencing, which can be both time-consuming and costly. The time needed to acquire this mutational data can delay treatment decisions, particularly for aggressive cancers like AML, where prompt intervention is crucial. Moreover, relying solely on mutational data to guide a treatment plan often overlooks other important factors, such as the differentiation stage of the leukemic cells, which can have a significant impact on treatment outcomes.

[0028] To aid the patient 102 with selecting a treatment therapy that is tailored to the patient’s 102 specific AML, the provider, via a client device 106 may use the biomarker detection engine 114. The biomarker detection engine 114 may generate a prediction 116 indicating the patient’s 102 likelihood to respond positively to a particular treatment therapy, here the VEN / AZA treatment. Based on this prediction 116, the provider, along with input from the patient 102, can create or tailor a treatment plan 120 for the patient’s 102 diagnosis. When dealing with AML, time is critical, and selecting a treatment plan 120 with the highest likelihood of success from the outset is crucial. Delaying the initiation of effective therapy can rapidly worsen the patient’s 102 condition, reducing the chances of a favorable outcome. Trying a therapy that the patient 102 is unlikely to positively respond to not only degrades the patient's 102 health but also allows the disease to progress further, making it more difficult to control and diminishing the effectiveness of future treatments. This can severely limit the patient’s 102 options and worsen their overall prognosis. Accordingly, by leveraging the biomarker detection engine 114 from the onset of treatment, the provider can select a treatment therapy and tailor the treatment plan 120 to the patient 102, thereby increasing the chances of a favorable outcome for the patient 102.

[0029] To generate the treatment plan 120 for the patient 102, including identifying a treatment therapy likely to be successful, the patient 102 may first provide a sample 104. The sample 104 may be a biological specimen containing cells affected by the respective disease, such as AML, and may be collected from the bone marrow of the patient 102. The nature of the sample 104 may vary depending on the clinical circumstances and the type of image 108 to be generated. In some embodiments, the sample 104 may be a bone marrow aspirate sample, in which a liquid suspension of marrow cells is drawn from the marrow cavity and subsequently used to create a smear or cytological preparation. In other embodiments, the sample 104 may be a bone marrow core biopsy specimen, which yields a solid tissue section containing both hematopoietic and stromal elements. In cases where an aspirate cannot be obtained, such as during a “dry tap,” the provider may generate a touch preparation by gently pressing the freshlyAttorney Docket No. 303.0138WO obtained core biopsy specimen against a glass slide to transfer cellular material suitable for microscopic evaluation.

[0030] From the collected sample 104, the provider may prepare one or more digital images 108 suitable for computational analysis. The images 108 may include, for example, aspirate smear images (hereinafter “smear images”), histologic images of core biopsy sections, or touch preparation images derived from core samples. Each image type provides distinct cytologic and morphologic information — aspirate smears typically allow detailed visualization of individual cell morphology, while core biopsies and touch preparations preserve architectural context and spatial cell relationships. In some embodiments, the images 108 may be digitized using a slide scanner or microscope-mounted imaging system at high optical magnification to capture the relevant cellular and subcellular features.

[0031] For purposes of clarity and consistency in the following description, the following discussion makes reference to a smear image. However, it should be appreciated that the techniques and systems described herein arc equally applicable to other bone marrow image types, including biopsy section images and touch preparation images, unless explicitly stated otherwise.

[0032] Afterwards, the slide may be stained with one or more specialized dyes or staining protocols configured to enhance visualization of cellular and subcellular structures within the sample 104. In some embodiments, the slide may be stained using a Wright-Giemsa stain, which differentially colors the nuclei, cytoplasm, and granules of hematopoietic cells to facilitate assessment of cell lineage and maturity. In other embodiments, a Romanowsky-type stain, such as May-Grunwald-Giemsa or Leishman stain, may be employed to accentuate fine nuclear chromatin patterns and cytoplasmic details. Additional or alternative staining methods, such as Hematoxylin and Eosin (H&E), Prussian Blue, or immunohistochemical stains targeting lineage-specific markers (e.g., CD34, MPG, or GDI 17), may also be applied depending on the diagnostic protocol and the desired contrast of specific cellular features.

[0033] These staining techniques improve image quality and analytical precision by increasing the contrast between cellular components, enabling both human observers and automated image-processing systems to better distinguish nuclear morphology, cytoplasmic granularity, and cellular organization. Enhanced staining contrast allows the biomarker detection engine to more accurately segment individual cells, extract quantitative morphological features, and identify subtle patterns that may correlate with treatment response. From a technological standpoint, the use of standardized and high-contrast staining protocols also reduces variability across samples and imaging systems, thereby improving the reliabilityAttorney Docket No. 303.0138WO and reproducibility of downstream computational analysis. As will be described in greater detail below, consistent staining enhances the biomarker detection engine’s 104 ability to generalize across patient datasets and ensures that its trained models perform robustly in clinical and research environments.

[0034] Once prepared, the provider may provide the images 108, via the client device 106, to an application service 110. As illustrated, the client device 106 may be in operational communication with the application service 110 for one or more functions or features. Broadly speaking, the application service 1 10 provides software application services to end points, such as the client device 106, examples of which include medical software for preparing treatment plans or recording medical events for patients, such as the patient 102. In the illustrated example, the application service 110 may provide a diagnostic tool for generating treatment plans for patients, such as the treatment plan 120 for the patient 102. As such, the client device 106 may load and execute software applications locally that interface with services and resources provided by the application service 110, such as the biomarkcr detection engine 114. The applications may be natively installed and executed applications, web-based applications that execute in the context of a local browser application, mobile applications, streaming applications, or any other suitable type of application. Example services and resources provided by the application service 110 include front-end servers, application servers, content storage services, authorization and authentication services, and the like.

[0035] To interact with the application service 110, the client device 106 may communicate with the application service 110 via one or more internets and intranets, the Internet, wired and wireless networks, local area networks (LANs), wide area networks (WANs), or any other type of network or combination thereof. Examples of the client device 106 may include personal computers, tablet computers, mobile phones, gaming consoles, wearable devices, Internet of Things (loT) devices, and any other suitable devices, of which computing apparatus 791 in Figure 7 is also broadly representative.

[0036] In the illustrated example, the application service 110 operates in a cloud-based environment. As such, the application service 110 employs one or more server computers 112 co-located with respect to each other or distributed across one or more data centers to deliver its functionalities and services. Example servers include web servers, application servers, virtual or physical servers, or any combination or variation thereof, of which computing apparatus 791 in Figure 7 is broadly representative.

[0037] In some embodiments, the biomarker detection engine 114 may be executed remotely by the application service 110 or a third party, while in other embodiments theAttorney Docket No. 303.0138WO biomarker detection engine 114 may be installed and executed locally on the client device 106. In still other embodiments, one or more functions of the biomarker detection engine 114, as described herein, may be installed and executed locally on the client device 106, while the remaining functions are integrated and executed remotely via the application service 110 or a third party.

[0038] As illustrated, the application service 110 may include an integration with the biomarker detection engine 114 to generate a prediction of the treatment outcome of the patient 102 when treated according to a specific treatment therapy, as described herein. A treatment therapy may be or include any therapeutic intervention, regimen, or combination of interventions designed to treat, manage, or cure a myeloid malignancy. Example treatment therapies may include, but are not limited to, pharmaceutical agents, biological therapies, radiation treatments, surgical procedures, or combinations thereof administered according to established medical protocols.

[0039] As noted above, to generate a prediction of a treatment outcome for the patient 102, the provider, via the client device 106 may submit the images 108 to the biomarker detection engine 1 14, such as via the application service 110. Responsive to receiving the images 108, the biomarker detection engine 1 14 may generate the prediction 116 and provide the prediction 116 and / or the treatment plan 120 to the client device 106. The prediction 116 may indicate a likelihood that the patient 102 is to respond positively to a selected treatment therapy, such as the VEN / AZA treatment in the case that the patient 102 is diagnosed with AML. The generation of the prediction 1 16 by the biomarker detection engine 1 14 is described in greater detail below with respect to Figures 2-6.

[0040] In some embodiments, the biomarker detection engine 114 may provide a treatment plan 120 along with the prediction 116. The treatment plan 120 (or the prediction 116) may be displayed to the provider via a user interface 118 of an application executing on the client device 106. The application may correspond to the application service 110 or with an application associated with the biomarker detection engine 114. As will be described in greater detail below, the treatment plan 120 may include the prediction 116, and in some cases, a recommendation of a treatment therapy based on the prediction 116. Through the user interface 118, the provider can interact with the treatment plan 120 to develop a management strategy that optimally addresses the patient's current prognosis.

[0041] Referring now to Figure 2, an example biomarker detection engine 214 is provided, according to an embodiment herein. For ease of illustration, Figure 2 is described with respect to Figure 3, which provides a process 300 for providing a biomarker detectionAttorney Docket No. 303.0138WO engine and its related functions, such as the biomarker detection engine 214, according to an embodiment herein. Although Figure 3 is described in relation to Figure 2, it should be appreciated that the process 300 of Figure 3 is equally applicable to the remaining Figures and components therein. Figure 2 is also described with respect to Figures 4-6, each of which is referenced in turn below.

[0042] To generate a prediction of a patient’s response to a treatment therapy, such as the patient’s 102 response to a VEN / AZA treatment, the biomarker detection engine 214 may receive an image 208 from a client device 206 (305). The client device 206 may be the same or similar to the client device 106, and be associated with a provider. The smear image 208 may be a smear image of a biopsy sample collected from a patient. As such, the smear image 208 may include cells affected by the respective malignancy.

[0043] With brief reference to Figure 4, an example perspective 400 of a smear image that may be submitted to the biomarker detection engine 214 is provided, according to an embodiment herein. The perspective 400 is a smear image of a bone marrow biopsy sample collected from a patient having AML. As shown, the perspective 400 contains various bone marrow cells 448 having various morphological variations and adipocyte cells 452. Example cell 450 illustrates an affected cell (e.g., leukemia cell) that may be used by the biomarker detection engine 214 to determine a responsiveness of a respective patient to a selected treatment plan.

[0044] Returning now to Figure 2, responsive to receiving the smear image 208, the biomarker detection engine 214 may preprocess the smear image 208 to generate a preprocessed image 224. In particular, the biomarker detection engine 214 may include an image preprocessor 222 that may perform one or more preprocessing operations on the smear image 208. The preprocessing operations may standardize the input data, enhance image quality, and optimize the smear image 208 for subsequent analysis by the deep learning model 238. For example, the image preprocessor 222 may digitize and / or magnify the smear image 208, such as at a resolution equivalent to 63 x optical magnification, to capture fine cellular structures and morphological characteristics relevant to biomarker identification. The image preprocessor 222 may also perform normalization operations to standardize pixel intensity values across imaging conditions and equipment types, thereby minimizing variability among samples obtained from different sources or time points.

[0045] To further refine the image data, the image preprocessor 222 may execute a sequence of enhancement and transformation procedures. Contrast adjustment operations may increase differentiation between cellular components and background regions, improving theAttorney Docket No. 303.0138WO visibility of nuclear morphology, cytoplasmic boundaries, and other subcellular details. Colorspace conversion operations may transform the image from an RGB color space to an alternative representation such as HSV or LAB, enabling more effective feature extraction by the deep learning model 238. Geometric transformation operations, including rotation, scaling, translation, and shearing, may augment the dataset and improve model robustness to variations in cell orientation or spatial positioning. Additionally, noise-reduction filters (e.g., Gaussian or median filtering) may be applied to remove imaging artifacts while preserving diagnostically relevant detail. Histogram equalization and edge enhancement operations may further improve contrast and sharpen cellular boundaries, enhancing the delineation of morphological features critical for downstream biomarker detection. Once these preprocessing operations are completed, the image preprocessor 222 may output the processed image 224 that has been standardized, denoised, and contrast-optimized.

[0046] In some embodiments, the biomarker detection engine 214 may generate a smear image data 228 from the processed image 224 (310). In particular, the biomarkcr detection engine 214 may include a smear image data generator 226 that generates the smear image data 228. In some embodiments, the smear image data 228 may be generated directly from the smear image 208, prior to performing one or more of the preprocessing operations, while in others, the smear image data 228 may be generated subsequent to the preprocessing operations. The smear image data generator 226 may be configured to identify and retain regions (also referred to herein as segments) of the smear image 208 that contain high-quality cellular structures suitable for biomarker analysis while removing regions that lack diagnostic value.

[0047] To generate the smear image data 228, the smear image data generator 226 may section the smear image 208 (or the processed image 224) into multiple regions (315). Each region may include a grid formed from a collection of tiles, thereby segmenting the smear image 208 into equi-sized segments. This tiling approach enables systematic analysis of the entire smear image while facilitating efficient processing and quality assessment of individual image segments. The smear image data generator 226 may then process each region to identify tiles lacking cellular presence (320) and remove those tiles from the smear image (325), thereby generating refined smear image data 228 that contains only tiles with meaningful cellular content for subsequent analysis by the deep learning model 238. In some cases, the remaining tiles may be filtered according to quality of cellular presence (330) to identify tiles containing quality cellular presence for inclusion in the smear image data 228. Each of these steps are described in turn below.Attorney Docket No. 303.0138WO

[0048] Referring now to Figure 5, an example image 500 of a region 554 containing a grid of tiles 556 is illustrated, according to an embodiment herein. The region 554 may be generated from a smear image, such as the smear image 208, by partitioning the smear image 208 into multiple regions. As illustrated, each region 554 may be further segmented into multiple tiles 556 arranged according to a grid pattern. Each tile 556 represents an equi-sized segment of the original smear image 208, thereby allowing for uniform processing and analysis. As shown, the presence of cellular components may vary across the tiles 556 within the region 554. That is, certain tiles may contain diagnostically relevant cellular structures and morphological features, while other tiles may consist primarily of background regions with minimal or no cellular presence. For example, some tiles 556 may be densely populated with bone marrow cells, affected cells, and adipocytes, whereas other tiles, such as tiles 558A-E, may lack observable cellular content. The tiles 558A-E lacking sufficient cellular presence may be referred to herein as empty tiles. In some embodiments, the absence of cellular presence may be defined quantitatively, such as when the number of pixels corresponding to cellular material within a given tile falls below a predetermined threshold, thereby indicating that the tile contains primarily non-cellular background regions rather than diagnostically useful content.

[0049] The smear image data generator 226 may process each tile 556 to assess the degree of cellular presence and image quality, enabling the system to selectively retain tiles containing diagnostically valuable information while excluding empty or low-information tiles. In some embodiments, the smear image data generator 226 may analyze the smear image 208 (or the processed image 224) to detect and tag tiles 556 that lack sufficient cellular presence, designating them as empty tiles 558A-E. Based on these tags, the smear image data generator 226 may remove the empty tiles 558A-E from the dataset to prevent inclusion of irrelevant or low-value data in subsequent computational analysis.

[0050] In addition to excluding empty tiles, the smear image data generator 226 may identify tiles exhibiting quality cellular presence. As used herein, quality cellular presence refers to areas of the smear image that contain cells with adequate morphological clarity, structural integrity, and staining definition suitable for automated feature extraction and biomarker detection. Quality may be determined based on one or more parameters, including cell density, visible nuclear and cytoplasmic boundaries, absence of imaging artifacts, appropriate color contrast, and preservation of cellular morphology. Tiles 556 exhibiting these characteristics are differentiated from those that contain partial or degraded cells, excessive background noise, or other image distortions that hinder reliable computational analysis. AsAttorney Docket No. 303.0138WO described below, the smear image data generator 226 may filter the tiles 556 according to these criteria, retaining only those with sufficient quality cellular presence for inclusion in the smear image data 228, while discarding tiles 558A-E that fail to meet established thresholds.

[0051] Referring now to Figure 6, an example image 600 of the region 554 of Figure 5 is provided, illustrating the removal of the empty tiles 558A-E. As shown, the remaining tiles 660A-DD include varying degrees of cellular presence. Some tiles, such as tiles 660I-K, 660O-P, 660U-V, and 660BB, contain dense clusters of intact cellular components, whereas others, such as tiles 660D-E, 660G-H, and 660Y, contain only partial or fragmented cells. Because these latter tiles offer limited information about overall cellular morphology, the smear image data generator 226 may classify them as exhibiting minimal quality cellular presence and exclude them from the final smear image data 228. This filtering process ensures that only tiles containing diagnostically informative cellular content are used in subsequent computational analysis, improving both model accuracy and data efficiency.

[0052] In some embodiments, the smear image data generator 226 may incorporate or interface with a machine-learning model trained to automatically identify tiles with quality cellular presence. The machine-learning model may be trained using a labeled subset of the training dataset 230, where historical images 232 have been preprocessed and segmented into tiles in accordance with the methods described above. During training, each tile may be annotated to indicate the level of cellular content quality, allowing the model to leam distinguishing visual features — such as cellular density, boundary sharpness, texture uniformity, and staining consistency — that differentiate high-quality tiles from low-quality or empty ones. The machine-learning model may employ, for example, a convolutional neural network (CNN), a vision transformer (ViT), or a hybrid architecture optimized for spatial feature recognition in histopathologic images. Once trained, the model may automatically evaluate tiles within incoming smear images and classify each tile according to its cellular presence quality. By automating this filtering process, the smear image data generator 226 enhances throughput, reduces variability between operators, and ensures consistent, objective selection of diagnostically relevant tiles across different samples and imaging platforms. This automation improves the technical efficiency of the overall biomarker detection engine 214 by reducing computational load and increasing the fidelity of the data used for downstream feature extraction and inference.

[0053] Returning now to Figure 2, responsive to generating the smear image data 228, the biomarker detection engine 214 may generate a prediction 216 representing a treatment outcome for a targeted therapy based on the smear image data 228 (335). To generate theAttorney Docket No. 303.0138WO prediction 216, the smear image data 228 may be provided as an input to one or more deep learning models 238 (340), which may responsively compute a probability value indicating the patient’s likelihood of responding positively to the treatment therapy (345). The probability 216 may be expressed as a continuous value between 0 and 1, representing a normalized likelihood score, or as a binary classification output in which the patient is categorized as Responsive or Non-Responsive to the selected treatment.

[0054] Responsivity to a particular treatment therapy may be quantitatively defined as the probability that the patient will achieve Complete Remission (CR) or Complete Remission with incomplete recovery of blood count (CRi), as established under standardized hematologic response criteria. In some embodiments, the biomarker detection engine 214 may compute a continuous probability value 216 representing the patient’s likelihood of achieving CR or CRi following the treatment therapy. When the computed probability 216 exceeds a predetermined threshold (for example, 0.70 or another empirically derived cutoff), the patient may be classified as Responsive to the treatment therapy. Conversely, when the computed probability 216 falls below the threshold, the patient may be classified as Non-Responsive, indicating a low expected likelihood of remission or a higher probability of disease persistence or progression.

[0055] In other embodiments, the probability 216 may be represented as a binary classification 242 output directly generated by the deep learning model 238. For instance, the deep learning model 238 may output a value of “1” to indicate that the patient is Responsive — that is, likely to achieve CR or CRi — and a value of “0” to indicate that the patient is Non- Responsive. In some configurations, the binary classification 242 may be produced through a sigmoid or softmax activation layer that converts the model’s 238 raw prediction score into a discrete output class. This quantitative framework provides an objective and reproducible metric for evaluating treatment efficacy, enabling consistent interpretation of model outputs across patient populations and therapeutic regimens.

[0056] The deep learning model 238 may operate as a multi-layer neural network (often called deep neural networks) that automatically learns patterns and features from datasets, such as the training dataset 230 described in greater detail below. Unlike traditional machine learning models, which often require manual feature engineering, the deep learning model 238 processes unstructured data such as the smear image 208 (or smear image data 228), to learn high-level abstractions directly from the raw input. The deep learning model 238, provided herein, is trained to detect variations in biomarkers present in the affected cells captured in the smear images 208 (or smear image data 228) and correlate those biomarker variations toAttorney Docket No. 303.0138WO treatment outcomes. As noted above, the deep learning model 238 is able to detect variations in biomarkers that are imperceptible to the human eye, and therefore able to identify subtle morphological or molecular differences within the cellular structures that correspond to particular treatment responses or resistance profiles.

[0057] In some embodiments, the deep learning model 238 implements a CNN architecture, which includes multiple convolutional, pooling, and fully connected layers configured to automatically extract and hierarchically represent visual features from the input images. The convolutional layers may capture low-level spatial patterns such as edges, textures, and shapes, while deeper layers learn higher-level abstractions corresponding to complex biomarker structures and morphological variations. The CNN may include hyperparameters such as learning rate, batch size, number of convolutional filters, kernel size, dropout rate, and regularization strength, which the trainer 236 optimizes during training to improve performance, convergence stability, and model generalization. In certain embodiments, the deep learning model 238 may include or be configured as a Wide Residual Network (Widc- ResNet) architecture, such as a Wide-ResNet 5-2 model containing three to six groups of residual layers with widening factors ranging from two to four, where the widening factor indicates the number of filters in each layer relative to a standard ResNet architecture. Increasing the widening factor allows the deep learning model 238 to learn a richer set of visual representations, thereby enhancing its ability to capture subtle and complex variations in biomarker morphology across the smear images.

[0058] To train the deep learning model 238, the biomarker detection engine 214 may include a trainer 236. The trainer 236 trains the deep learning model 238 using a training dataset 230. The training dataset 230 includes a plurality of historical images 232 comprising cell smear images from patients who have previously received a particular treatment therapy, such as the VEN / AZA treatment regimen. For each of the historical images 232, the training dataset 230 further includes an associated treatment outcome 234. Continuing with the example of AML patients receiving VEN / AZA treatment, the treatment outcomes 234 may include CR, CRi, and non-Responsive (non-R), where CR and CRi collectively represent “Responsive” outcomes.

[0059] Prior to feeding the historical images 232 into the training pipeline, the historical images 232 may each be preprocessed. For example, the historical images 232 may be preprocessed by the image preprocessor 222. In some embodiments, the image preprocessor 222 may perform preprocessing operations such as normalization, contrast adjustment, or color-space conversion to standardize the input data and improve training consistency. TheAttorney Docket No. 303.0138WO image preprocessor 222 may also augment the historical images 232 by employing one or more geometric transformations to increase dataset diversity and improve model robustness. That is, the image preprocessor 222 may modify each of the historical images’ 232 spatial properties to expand the effective dataset and strengthen the subsequent analysis performed by the biomarker detection engine 214, and in particular, the deep learning model 238. For example, the image preprocessor 222 may perform operations such as rotating, scaling, flipping, translating, or shearing the historical images 232 to aid in training the deep learning model 238 to recognize visual patterns across a range of orientations and spatial configurations. By introducing controlled variability, the image preprocessor 222 enables the deep learning model 238 to develop invariance to geometric and spatial transformations, thereby enhancing its ability to generalize to unseen data during inference.

[0060] In an example embodiment, the trainer 236 may train the deep learning model 238 using a High-Performance Computing (HPC) cluster, which refers to a distributed computing environment composed of multiple interconnected servers or compute nodes that operate in parallel to perform large-scale computational tasks with high throughput and efficiency. The HPC cluster may include multiple graphics processing units (GPUs) or tensor processing units (TPUs) configured to accelerate matrix operations and gradient computations commonly used in deep learning. The trainer 236 utilizes the HPC cluster to predict whether a patient is likely to respond positively to a specific treatment based on biomarker variations identified within a submitted image.

[0061] To begin, the training dataset 230 is prepared by labeling each historical image 232 with its corresponding treatment outcome 234 of either “Responsive” or “Non-responsive.” The trainer 236 then executes a distributed training process across the HPC cluster, wherein each GPU or compute node processes a subset of the historical images 232 in parallel, synchronizing gradients and model weights across nodes to ensure consistent learning. In one embodiment, the trainer 236 deploys the deep learning model 238 on an HPC cluster comprising eight Nvidia A100 GPUs and performs a hyperparameter search across approximately 100 independent training runs to identify optimal model configurations that maximize training performance, convergence stability, and predictive accuracy.

[0062] During training, each node processes batches of the historical images 232, passing them through the convolutional and residual layers of the deep learning model 238 to produce classification predictions. The trainer 236 compares the predicted outcomes to the actual treatment outcomes 234 to compute a loss value. Using backpropagation, the trainer 236 updates the internal weights and biases of the deep learning model 238 through gradient descentAttorney Docket No. 303.0138WO optimization. The trainer 236 may employ a stochastic gradient descent (SGD) or Adam optimizer to refine the deep learning model’ s 238 parameters iteratively across multiple epochs. The trainer 236 synchronizes gradients across distributed nodes to ensure consistent model updates. Over successive iterations, the deep learning model 238 progressively improves its ability to identify discriminative biomarkers and predict patient treatment responses.

[0063] The trainer 236 may leverage the HPC cluster to achieve efficient data transfer, gradient aggregation, and parallel execution, enabling faster convergence compared to singlenode training. Throughout the training process, the trainer 236 monitors performance metrics such as training accuracy, validation accuracy, and loss. The trainer 236 periodically saves checkpoints of the deep learning model 238 to preserve the best-performing weight configurations. Once trained, the deep learning model 238 applies the learned parameters to predict whether a given patient will likely be responsive or non-responsive to a specified treatment therapy, such as VEN / AZA therapy, based on features extracted from a submitted image 208.

[0064] In some embodiments, the performance of the deep learning model 238 may be measured or validated using a five-fold stratified cross-validation procedure with an approximately 80:20 train-test split. In this approach, the training dataset 230 is partitioned into five subsets (folds) that maintain the same proportional distribution of responsive and non- responsive outcomes as the full dataset, thereby preserving statistical balance between classes. For each fold, the trainer 236 trains the deep learning model 238 on four of the partitions and validates the trained model on the remaining partition, repeating this process five times so that each subset serves once as a validation set. The trainer 236 may aggregate the validation results across folds to compute average performance metrics and assess the deep learning model’s 238 generalization capability.

[0065] The stratified cross-validation process also mitigates overfitting by ensuring that model evaluation occurs on multiple independent data splits. This validation procedure may be particularly advantageous when training on relatively small datasets or limited patient cohorts. For example, when the deep learning model 238 is trained to determine whether an AML patient is likely to be responsive or non-responsive to VEN / AZA treatment therapy and the dataset 230 includes fewer than 200 patients, the five-fold stratified cross-validation ensures that each sample contributes to both training and validation, providing a more reliable estimate of model accuracy. The trainer 236 may further perform hyperparameter optimization across the folds to identify the configuration that yields the highest mean validation performance and stable convergence behavior.Attorney Docket No. 303.0138WO

[0066] Using the five-fold stratified cross-validation process described above, the trainer 236 may evaluate the performance of the deep learning model 238 using metrics such as the Area Under the Receiver Operating Characteristic (AUROC) curve. The AUROC quantifies the overall ability of the deep learning model 238 to distinguish between responsive and non- responsive outcomes across the validation sets. The Receiver Operating Characteristic (ROC) curve plots the true positive rate (sensitivity) against the false positive rate (1 - specificity) at various decision thresholds, thereby illustrating how well the deep learning model 238 discriminates between positive and negative classes over the entire range of possible thresholds. The area under this curve (AUC) provides a single scalar value representing overall classification performance. An AUC value closer to 1.0 indicates strong model performance with excellent class separation, whereas an AUC of 0.5 suggests that the deep learning model 238 performs no better than random guessing. AUC values below 0.5 indicate poor discriminative capability, where the deep learning model 238 misclassifies more than chance would predict. In many embodiments, an AUROC above approximately 0.8 is considered desirable to ensure that the deep learning model 238 is accurately classifying the validation data and generalizing effectively to unseen samples.

[0067] During an example training run, the deep learning model 238 achieved, for a fivefold stratified validation dataset, AUC values of approximately 0.8240, 0.7744, 0.8590, 0.8624, and 0.8624 for Folds 0 through 4, respectively, corresponding to a mean AUC of approximately 0.84. These results demonstrate consistent discriminative performance across folds, confirming that the deep learning model 238 generalizes well and reliably distinguishes treatment responders from non-responders. These results provided a performance baseline that informed subsequent hyperparameter optimization (I IPO) runs designed to further refine model generalization and predictive stability.

[0068] In another example training run, the deep learning model 238, when evaluated on a five-fold stratified dataset, achieved an AUROC of approximately 0.802 on held-out validation datasets following multiple hyperparameter optimization (HPO) runs. Each HPO run involved systematic variation of model hyperparameters, such as learning rate, batch size, number of convolutional filters, and regularization parameters, to identify optimal configurations that maximized predictive accuracy and training stability. The reported AUROC value represents the mean discriminative performance obtained after convergence of the optimized configuration across multiple independent training cycles. These results further demonstrate that the deep learning model 238 maintains strong generalization capability andAttorney Docket No. 303.0138WO consistent predictive accuracy on unseen, held-out data, confirming the reproducibility and robustness of its learned representations.

[0069] In a further example training run, the deep learning model 238 achieved an area under the receiver operating characteristic (AUC) of approximately 0.970635 for the training dataset corresponding to Fold 1 of the five-fold stratified dataset. This high AUC value indicates that the deep learning model 238 effectively learned discriminative representations from the training data, accurately distinguishing between responsive and non-responsive outcomes during the training phase. The trainer 236 monitored training and validation losses throughout this process to ensure convergence stability and to mitigate overfitting, thereby maintaining generalization capability. The observed training AUC of approximately 0.970635 for Fold 1 demonstrates that the deep learning model 238 is capable of achieving high classification fidelity on the training data while preserving strong validation performance across folds.

[0070] In addition to AUROC analysis, the trainer 236 may further validate the deep learning model 238 using a confusion matrix to provide a detailed breakdown of classification outcomes for each validation fold. The confusion matrix compares the predicted treatment outcomes with the actual treatment outcomes 234 and categorizes results into four groups: true positives (TP), where the deep learning model 238 correctly predicts the responsive class; true negatives (TN), where it correctly predicts the non-responsive class; false positives (FP), where it incorrectly predicts responsiveness; and false negatives (FN), where it incorrectly predicts non-responsiveness. By analyzing these categories, the trainer 236 can identify specific areas of strength and misclassification, thereby obtaining a more granular understanding of model behavior than is provided by a single scalar metric such as accuracy. The confusion matrix also provides the basis for calculating additional evaluation metrics, including sensitivity (recall), specificity, precision, and Fl -score. These derived measures further inform the ROC and AUROC analyses, as sensitivity and specificity directly determine the ROC curve’s shape. During five-fold stratified cross-validation, the trainer 236 may generate a confusion matrix for each fold to monitor model stability and consistency across independent data partitions.

[0071] In an illustrative example, the deep learning model 238 may be a CNN-based Wide-ResNet 5-2 architecture trained by the trainer 236 using the HPC cluster to predict treatment outcomes for patients diagnosed with AMU receiving VEN / AZA treatment therapy. In this example, the trainer 236 validated the deep learning model 238 using the five-fold stratified cross-validation process and achieved a mean AUROC of approximately 0.84 (range: 0.77-0.86) in distinguishing between responders and non-responders. The confusion matricesAttorney Docket No. 303.0138WO across the five folds showed accuracies ranging from 0.84 to 0.90, Fl-scores from 0.77 to 0.81, sensitivities from 0.93 to 0.96, specificities from 0.49 to 0.67, positive predictive values from 0.88 to 0.92, and negative predictive values from 0.62 to 0.79. These results are consistent with those observed across additional training and held-out validation runs, collectively demonstrating that the deep learning model 238 exhibits strong generalization, stable classification accuracy, and robust predictive capability when trained and tuned by the trainer 236 using the training dataset 230 of historical images 232 and outcomes 234. The overall model performance across embodiments may be further characterized by the performance metrics described below.

[0072] In various embodiments, the deep learning model 238, when trained and validated as described herein, demonstrates performance metrics indicative of high classification accuracy, robust generalization, and stable convergence. The deep learning model 238 may achieve an average AUROC curve of approximately 0.84 across a five-fold stratified validation dataset, with individual fold AUG values ranging from approximately 0.77 to 0.86. In some embodiments, the deep learning model 238 may achieve an AUROC of approximately 0.802 on held-out validation datasets following multiple hyperparameter optimization runs, and an AUG of approximately 0.9706 on training data for certain folds, such as Fold 1. The trainer 236 may further compute additional performance metrics, including classification accuracy ranging from approximately 0.84 to 0.90, Fl-scores from approximately 0.77 to 0.81, sensitivities from approximately 0.93 to 0.96, specificities from approximately 0.49 to 0.67, positive predictive values from approximately 0.88 to 0.92, and negative predictive values from approximately 0.62 to 0.79. Collectively, these performance metrics demonstrate that the deep learning model 238 provides reliable and repeatable classification of treatment response outcomes while maintaining consistency across multiple validation folds and held-out datasets.

[0073] In some embodiments, the deep learning model 238 may be trained according to a specific cytological staining technique used to prepare the smear image 208. Because different staining techniques highlight distinct cellular and subcellular structures, the visual and colorimetric characteristics of the training dataset 230 may vary depending on the staining method employed. For example, the deep learning model 238 may be trained using Romanowsky-type stained bone marrow smear images, such as those prepared with Wright- Giemsa or May-Grunwald-Giemsa stains. These staining techniques produce characteristic polychromatic coloration of cellular components — differentially tinting nuclei, cytoplasm, and cytoplasmic granules — which enhances the deep learning model’s 238 ability to learn andAttorney Docket No. 303.0138WO distinguish morphological patterns associated with hematopoietic cell lineages, maturation stages, and disease-specific cytologic features.

[0074] In other embodiments, the deep learning model 238 may be trained on Hematoxylin and Eosin (H&E)-stained images or immunohistochemical (IHC) stains that mark lineage-specific antigens such as CD34, myeloperoxidase (MPO), or CD117. Training the deep learning model 238 with data derived from a consistent staining technique ensures spectral and structural uniformity across samples, reducing color-space variability and improving generalization performance. In yet other embodiments, the biomarker detection engine 214 may include multiple deep learning models 238, each trained on image datasets 230 corresponding to different staining methods, allowing the system to automatically identify or adapt to the staining protocol applied to a given input image. This configuration enables robust performance across laboratory settings and supports accurate biomarker detection regardless of the cytological preparation method used.

[0075] Once the deep learning model 238 is trained, the model 238 can rapidly and accurately predict a patient’s outcome (e.g., responsive vs. non-responsive) to a particular treatment therapy based on smear images of the patient’s affected cells (e.g., leukemic cells). From its training, the deep learning model 238 may detect biomarker variations present within the affected cells, variations that are often undetectable by the human eye. In some embodiments, the deep learning model 238 may predict how likely a patient is to respond positively to a treatment therapy (e.g., be responsive) agnostic to typical response predictors such as gene mutations or other biological patient data.

[0076] In some embodiments, the deep learning model 238 may be trained to generate predictions across multiple treatment therapies for a given myeloid malignancy. For malignancies with several approved or investigational treatment options, the deep learning model 238 may learn to associate distinct morphological and biomarker patterns present in the smear image data 228 with responsiveness to each available therapy. During training, the deep learning model 238 may be provided with a labeled dataset 230 in which each training instance corresponds to a digital smear image annotated with clinical outcome data specifying whether the patient achieved CR or CRi after receiving a particular treatment therapy. Through exposure to this multi-label dataset 230, the deep learning model 238 learns to infer treatmentspecific response probabilities that quantitatively indicate the likelihood of a favorable therapeutic outcome based on the patient’ s cellular morphology and biomarker profile.

[0077] When the smear image data 228 is submitted to the trained deep learning model 238, the model may generate multiple prediction values 216A-C — each corresponding to aAttorney Docket No. 303.0138WO respective treatment therapy available for the diagnosed malignancy. For example, in the case of AML, the deep learning model 238 may output a first prediction 216A representing the probability of responsiveness to an anthracycline- and cytarabine-based chemotherapy regimen (e.g., “7+3”), a second prediction 216B representing the probability of responsiveness to a targeted therapy such as an FLT3 inhibitor, and a third prediction 216C representing the probability of responsiveness to a hypomethylating agent (e.g., azacitidine or decitabine) alone or in combination with a BCL-2 inhibitor such as venetoclax. Each prediction 216A-216C may be expressed as a normalized probability value between 0 and 1, or as a binary classification indicating whether the patient is likely or unlikely to achieve CR or CRi for the corresponding treatment option. In this embodiment, a single multi-output deep learning model 238 is capable of providing a comprehensive, treatment-specific set of response predictions from a single input of the smear image data 228.

[0078] In other embodiments, rather than employing a single deep learning model trained to output multiple treatment predictions 216A-C, the biomarkcr detection engine 214 may include multiple deep learning models 238A-C, each individually trained to predict responsiveness for a specific treatment therapy. For example, a first deep learning model 238A may be trained using image data and outcome labels corresponding to patients treated with anthracycline -based chemotherapy, a second deep learning model 238B may be trained using data associated with FLT3 inhibitor therapy, and a third deep learning model 238C may be trained using data from patients treated with hypomethylating agents such as azacitidine or decitabine. Each model may independently learn morphological and biomarker features that are most predictive of therapeutic response for its respective treatment type. During operation, the smear image data 228 may be submitted in parallel or sequentially to each of the deep learning models 238A-238C, and each model may generate a separate prediction 216A-216C corresponding to the likelihood of responsiveness for its associated therapy.

[0079] This multi-model configuration allows each deep learning model 238A-C to specialize in the unique cellular and morphological indicators relevant to a specific therapeutic mechanism of action. Such specialization can improve predictive accuracy by enabling each model to focus on feature distributions that are most discriminative for its corresponding treatment class, without being influenced by confounding patterns associated with unrelated therapies. In some implementations, the biomarker detection engine 214 may aggregate the predictions from the multiple models 238A-C to produce a comparative output, such as a ranked treatment list or a probability distribution across all available therapies. This architecture further enables modular updating — individual models 238A-C can be retrained orAttorney Docket No. 303.0138WO replaced as new treatment options become available or as additional outcome data are collected, without requiring retraining of the entire system. Accordingly, the multi-model embodiment provides flexibility, scalability, and improved interpretability in assessing patientspecific therapeutic responsiveness.

[0080] As noted above, in some embodiments, the prediction 216 output from the deep learning model 238 may include a binary classification 242 that indicates whether the patient is predicted to be Responsive or Non-Responsive to a particular treatment therapy. In some cases, the prediction 216 may further include a confidence score 244 corresponding to the deep learning model’s 238 level of certainty in its classification. The confidence score 244 may represent a normalized probability value, such as between 0 and 1, or a percentage value indicating the statistical confidence associated with the binary classification 242. For example, a classification of Responsive with a confidence score 244 of 0.85 may indicate that, based on the features detected in the smear image data 228, the deep learning model 238 estimates an 85% probability that the patient will achieve CR or CRi following the selected treatment therapy. Conversely, a classification of Non-Responsive with a confidence score of 0.20 may indicate a low likelihood of therapeutic success.

[0081] When multiple predictions 216A-N are generated — such as when the deep learning model 238 outputs treatment-specific predictions for several available therapies, or when multiple deep learning models 238A-238N are each trained for a different treatment therapy — the biomarker detection engine 214 may generate a corresponding binary classification 242 and confidence score 244 for each treatment option. For example, for a particular patient diagnosed with AML, the biomarker detection engine 214 may produce separate binary classifications 242A-C and confidence scores 244A-C for chemotherapy, targeted inhibitor, and hypomethylating agent therapies. Each classification 242A-C may indicate whether the patient is predicted to respond to that therapy, while each confidence score 244 A-C quantifies the relative strength or reliability of that prediction. The combination of binary classifications 242A-C and associated confidence scores 244A-C enables the biomarker detection engine 214 to provide a comprehensive, interpretable, and quantitative representation of therapeutic responsiveness across multiple treatment therapies. These outputs may be further utilized by a treatment recommendation module or displayed in ranked order according to confidence, allowing clinicians to assess which therapies present the highest likelihood of achieving a favorable outcome for the individual patient.

[0082] In some embodiments, the biomarker detection engine 214 may generate a recommendation 246 for a treatment plan for the patient based on the prediction 216 (350).Attorney Docket No. 303.0138WOFollowing the above example, if the deep learning model 238 generates the prediction 216 that the AML patient is likely to be “Responsive” to the VEN / AZA treatment, then a treatment recommendation generator 240 of the biomarker detection engine 214 may generate a recommendation 246 that the patient proceed with the VEN / AZA treatment. In contrast, if the prediction 216 indicates that the AML patient is likely to be “Non-responsive” to theVEN / AZA treatment, then the treatment recommendation generator 240 may generate the recommendation 246 that the patient select an alternative treatment. The recommendation 246 may include the classification 242 and the confidence score 244 to provide context for the recommended treatment plan. In some cases, the treatment recommendation generator 240 may include a listing of one or more other treatment therapies for the specific malignancy, such as intensive chemotherapy, hypomethylating agents, or targeted therapies for AML.

[0083] Table 1 provided below illustrates various treatment therapies that may be options to treat different myeloid malignancies.Table 1Attorney Docket No. 303.0138WO

[0084] As illustrated, the recommendation 246 may be transmitted by the biomarker detection engine 214 to the client device 206, for example via the application service 110. The biomarker detection engine 214 may be configured for real-time or near- real-time processing, enabling the generation and transmission of the recommendation 246 immediately following submission of the smear image data 228. The recommendation 246 may be presented to a provider through the treatment plan 120 displayed on the user interface 118. In some embodiments, the recommendation 246 may include one or more predicted treatment options ranked according to the corresponding probabilities 216, binary classifications 242, or confidence scores 244 generated by the deep learning models 238. Because the biomarker detection engine 214 performs automated image preprocessing, feature extraction, and inference in real time, the provider may receive treatment predictions within seconds or minutes of image submission. This capability allows for prompt clinical decision-making in time-sensitive cases, such as aggressive myeloid malignancies requiring rapid therapeutic intervention. By incorporating individualized biomarkcr-bascd predictions derived from the patient’s smear image data 228, the system enables both the provider and the patient to make informed, data-driven decisions regarding therapeutic selection, tailored to the patient’s unique morphological and biomarker profile.

[0085] Referring to Figure 7, Figure 7 illustrates a computing apparatus 791 that may be used for providing a biomarker detection engine and related functions, as described herein. For example, the client device 106 may be or include the computing apparatus 791. As illustrated, the computing apparatus 791 includes a processing system 792 that includes a microprocessor and other circuitry that retrieves and executes software 795 from storage system 793. The processing system 792 may be implemented within a single processing device but may also be distributed across multiple processing devices or sub-systems that cooperate in executing program instructions. Examples of the processing system 792 include general purpose central processing units, graphical processing units, application specific processors, and logic devices, as well as any other type of processing device, combinations, or variations thereof.

[0086] The storage system 793 may comprise any computer-readable storage media or medium readable by processing system 792 and capable of storing software 795. The storage system 793 may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules, or other data. Examples of storage media include random access memory, read only memory, magnetic disks, optical disks, flash memory, virtual memory and non-virtual memory, magnetic cassettes, magnetic tape, magneticAttorney Docket No. 303.0138WO disk storage or other magnetic storage devices, or any other suitable storage media. In no case is the computer readable storage media a propagated signal.

[0087] In addition to computer readable storage media, in some implementations the storage system 793 may also include computer readable communication media over which at least some of the software 795 may be communicated internally or externally. The storage system 793 may be implemented as a single storage device but may also be implemented across multiple storage devices or sub-systems co-located or distributed relative to each other. The storage system 793 may comprise additional elements, such as a controller capable of communicating with the processing system 792 or possibly other systems.

[0088] The software 795 (including biomarker detection engine process 796) may be implemented in program instructions and among other functions may, when executed by the processing system 792, direct the processing system 792 to operate as described with respect to the various operational scenarios, sequences, and processes illustrated herein. For example, the software 795 may include program instructions for implementing a biomarkcr detection engine and related functions, such as the process 300, as described herein.

[0089] 1'he term “engine” as used herein include a “component”, “module”, “system,” and the like is intended to refer to a computer-related entity, either software-executing general- purpose processor, hardware, firmware and a combination thereof. For example, an engine may be, but is not limited to being, a process running on a hardware processor, a hardware-based processor, an object, an executable, a thread of execution, a program, and / or a computer.

[0090] The program instructions of the software 795 may include various components or modules that cooperate or otherwise interact to carry out the various processes and operational scenarios described herein. The various components or modules may be embodied in compiled or interpreted instructions, or in some other variation or combination of instructions. The various components or modules may be executed in a synchronous or asynchronous manner, serially or in parallel, in a single threaded environment or multi-threaded, or in accordance with any other suitable execution paradigm, variation, or combination thereof. The software 795 may include additional processes, programs, or components, such as operating system software, virtualization software, or other application software. The software 795 may also comprise firmware or some other form of machine-readable processing instmctions executable by the processing system 792.

[0091] In general, the software 795 may, when loaded into the processing system 792 and executed, transform a suitable apparatus, system, or device (of which computing apparatus 791 is representative) overall from a general-purpose computing system into a special-purposeAttorney Docket No. 303.0138WO computing system customized to generate features, functionality, and user experiences provided by the biomarker detection engine. Indeed, encoding the software 795 on the storage system 793 may transform the physical structure of the storage system 793. The specific transformation of the physical structure may depend on various factors in different implementations of this description. Examples of such factors may include, but are not limited to, the technology used to implement the storage media of the storage system 793 and whether the computer-storage media are characterized as primary or secondary storage, as well as other factors.

[0092] For example, if the computer readable storage media are implemented as semiconductor-based memory, the software 795 may transform the physical state of the semiconductor memory when the program instructions are encoded therein, such as by transforming the state of transistors, capacitors, or other discrete circuit elements constituting the semiconductor memory. A similar transformation may occur with respect to magnetic or optical media. Other transformations of physical media arc possible without departing from the scope of the present description, with the foregoing examples provided only to facilitate the present discussion.

[0093] Communication interface system 797 may include communication connections and devices that allow for communication with other computing systems (not shown) over communication networks (not shown). Examples of connections and devices that together allow for inter-system communication may include network interface cards, antennas, power amplifiers, RF circuitry, transceivers, and other communication circuitry. The connections and devices may communicate over communication media to exchange communications with other computing systems or networks of systems, such as metal, glass, air, or any other suitable communication media. The aforementioned media, connections, and devices are well known and need not be discussed at length here.

[0094] User interface system 799 may include various components and devices that enable interaction between the user and the computing system. Examples of these components and devices may include display screens, touchscreens, keyboards, mice, trackpads, styluses, voice recognition microphones, and other input / output devices. The user interface system 799 facilitates user commands and feedback through graphical user interfaces (GUIs), commandline interfaces (CLIs), or other interaction models. These interfaces may display information, receive user inputs, and provide visual, auditory, or tactile responses. The components and devices within the user interface system 799 are designed to ensure seamless and intuitive userAttorney Docket No. 303.0138WO interaction, leveraging well-established technologies and practices that need not be elaborated upon here.

[0095] Communication between the computing apparatus 791 and other computing systems (not shown), may occur over a communication network or networks and in accordance with various communication protocols, combinations of protocols, or variations thereof. Examples include intranets, internets, the Internet, local area networks, wide area networks, wireless networks, wired networks, virtual networks, software defined networks, data center buses and backplanes, or any other type of network, combination of network, or variation thereof. The aforementioned communication networks and protocols are well known and need not be discussed at length here.

[0096] While some examples of methods and systems herein are described in terms of software executing on various machines, the methods and systems may also be implemented as specifically-configured hardware, such as field-programmable gate array (FPGA) specifically to execute the various methods according to this disclosure. For example, examples can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in a combination thereof. In one example, a device may include a processor or processors. The processor comprises a computer-readable medium, such as a random access memory (RAM) coupled to the processor. The processor executes computer-executable program instructions stored in memory, such as executing one or more computer programs. Such processors may comprise a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), field programmable gate arrays (FPGAs), and state machines. Such processors may further comprise programmable electronic devices such as PLCs, programmable interrupt controllers (PICs), programmable logic devices (PLDs), programmable read-only memories (PROMs), electronically programmable read-only memories (EPROMs or EEPROMs), or other similar devices.

[0097] Such processors may comprise, or may be in communication with, media, for example one or more non-transitory computer-readable media, which may store processorexecutable instructions that, when executed by the processor, can cause the processor to perform methods according to this disclosure as earned out, or assisted, by a processor. Examples of may include, but are not limited to, an electronic, optical, magnetic, or other storage device capable of providing a processor, such as the processor in a web server, with processor-executable instructions. Other examples of non-transitory computer- readable media include, but are not limited to, a floppy disk, CD-ROM, magnetic disk, memory chip, ROM, RAM, ASIC, configured processor, all optical media, all magnetic tape or other magneticAttorney Docket No. 303.0138WO media, or any other medium from which a computer processor can read. The processor, and the processing, described may be in one or more structures, and may be dispersed through one or more structures. The processor may comprise code to carry out methods (or parts of methods) according to this disclosure.

[0098] Examples are described herein in the context of systems and methods for providing an biomarker detection engine and related functions. Those of ordinary skill in the art will realize that the foregoing description is illustrative only and is not intended to be in any way limiting. Reference is made in detail to implementations of examples as illustrated in the accompanying drawings. The same reference indicators will be used throughout the drawings and the following description to refer to the same or like items.

[0099] Additionally, the foregoing description of some examples has been presented only for the purpose of illustration and description and is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the spirit and scope of the disclosure. In the interest of clarity, not all of the routine features of the examples described herein are shown and described. It will, of course, be appreciated that in the development of any such actual implementation, numerous implementation-specific decisions must be made in order to achieve the developer’s specific goals, such as compliance with application- and business-related constraints, and that these specific goals will vary from one implementation to another and from one developer to another.

[0100] Reference herein to an example or implementation means that a particular feature, structure, operation, or other characteristic described in connection with the example may be included in at least one implementation of the disclosure. The disclosure is not restricted to the particular examples or implementations described as such. The appearance of the phrases “in one example,” “in an example,” “in one implementation,” or “in an implementation,” or variations of the same in various places in the specification does not necessarily refer to the same example or implementation. Any particular feature, structure, operation, or other characteristic described in this specification in relation to one example or implementation may be combined with other features, structures, operations, or other characteristics described in respect of any other example or implementation.

[0101] Use herein of the word “or” is intended to cover inclusive and exclusive OR conditions. In other words, A or B or C includes any or all of the following alternative combinations as appropriate for a particular usage: A alone; B alone; C alone; A and B only; A and C only; B and C only; and A and B and C.Attorney Docket No. 303.0138WOEXAMPLES

[0102] These illustrative examples are mentioned not to limit or define the scope of this disclosure, but rather to provide examples to aid understanding thereof. Illustrative examples are discussed above in the Detailed Description, which provides further description. Advantages offered by various examples may be further understood by examining this specification.

[0103] As used below, any reference to a series of examples is to be understood as a reference to each of those examples disjunctively (e.g., “Examples 1-4” is to be understood as “Examples 1, 2, 3, or 4”).

[0104] Example 1 is a biomarker detection engine for predicting a treatment outcome for a patient diagnosed with a myeloid malignancy in response to a treatment therapy, the biomarker detection engine comprising: a computer- readable storage media comprising processor-executable instructions stored thereon; and one or more processors coupled to the computer-readable storage media and configured to execute the processor-executable instructions that, when executed by the one or more processors, direct the biomarker detection engine to at least: receive a smear image of cells from a biopsy sample collected from the patient; preprocess the smear image to generate a smear image data; generate, using one or more deep learning models, a prediction of the treatment outcome for the treatment therapy based on the smear image data; and generate a recommendation for a treatment plan for the patient based on the prediction of the treatment outcome.

[0105] Example 2 is the biomarker detection engine of any previous or subsequent Example, wherein the processor-executable instructions to preprocess the smear image to generate the smear image data, when executed by the one or more processors, further direct the biomarker detection engine to: section the smear image into a plurality of regions, each region comprising a grid formed of a plurality of tiles; process a first region of the plurality of regions to identify at least one tile of the plurality of tiles within the first region comprising a lack of cellular presence; tag the at least one tile for removal; and generate the smear image data by removing a plurality of tagged tiles, wherein the plurality of tagged tiles comprises the at least one tile.

[0106] Example 3 is the biomarker detection engine of any previous or subsequent Example, wherein the processor-executable instructions to preprocess the smear image to generate the smear image data, when executed by the one or more processors, further direct the biomarker detection engine to: section the smear image into a plurality of regions, each regionAttorney Docket No. 303.0138WO comprising a grid formed of a plurality of tiles; identify a plurality of tiles within each region comprising cellular presence; identify a subset of tiles from the plurality of tiles comprising quality cellular presence; and generate the smear image data by removing remaining tiles of the plurality of tiles from the smear image.

[0107] Example 4 is the biomarker detection engine of any previous or subsequent Example, wherein the processor-executable instructions to generate, using the one or more deep learning models, the prediction of the treatment outcome for the patient, when executed by the one or more processors, further direct the biomarker detection engine to: submit the smear image data as an input into the one or more deep learning models, wherein the one or more deep learning models generate, responsive to the input, an output comprising: a binary classification indicating whether the patient will respond to the target therapy; and a confidence score associated with the binary classification.

[0108] Example 5 is the biomarker detection engine of any previous or subsequent Example, wherein: the myeloid malignancy is acute myeloid leukemia (AML); the treatment therapy comprises a combination therapy comprising venetoclax and azacitidine; the prediction of treatment outcome is a prediction of whether the patient will achieve complete remission (CR) or complete remission with incomplete recovery of blood count (CRi) in response to the combination therapy; and the recommendation for the treatment plan comprises recommending the combination therapy when the prediction indicates the patient is likely to achieve CR or CRi, or recommending an alternative treatment when the prediction indicates the patient is unlikely to achieve CR or CRi.

[0109] Example 6 is the biomarker detection engine of any previous or subsequent Example, wherein the one or more deep learning models comprise a convolutional neural network having hyperparameters tuned through a hyperparameter optimization process configured to improve biomarker detection accuracy in smear images.

[0110] Example 7 is the biomarker detection engine of any previous or subsequent Example, wherein the convolutional neural network exhibits an average area under a receiver operating characteristic (AUROC) curve greater than 0.8 when validated using a five-fold stratified cross-validation process.

[0111] Example 8 is a method for predicting treatment outcomes for a patient diagnosed with a myeloid malignancy in response to a treatment therapy, the method comprising: receiving, by a biomarker detection engine, a smear image of cells from a biopsy sample collected from the patient; generating, by the biomarker detection engine, a smear image data from the smear image; generating, using one or more deep learning models, aAttorney Docket No. 303.0138WO prediction of a treatment outcome for the treatment therapy based on the smear image data; and generating, by the biomarker detection engine, a recommendation for a treatment plan for the patient based on the prediction of the treatment outcome.

[0112] Example 9 is the method of any previous or subsequent Example, wherein the method further comprises: training the one or more deep learning models using a training dataset comprising a plurality of historical smear images and corresponding treatment outcomes; and tuning model hyperparameters through a hyperparameter optimization process executed by a trainer to improve predictive performance; and wherein the one or more deep learning models comprise a convolutional neural network trained using a five-fold stratified cross-validation process to detect biomarker variations within the historical smear images and correlate the variations with treatment outcomes.

[0113] Example 10 is the method of any previous or subsequent Example, wherein: the smear image is stained according to a cytological staining technique configured to enhance visualization of cellular morphology; and the one or more deep learning models are trained on images of smear samples prepared using the cytological staining technique.

[0114] Example 11 is the method of any previous or subsequent Example, wherein the cytological staining technique comprises a Romanowski-type stain configured to enhance visualization of cellular nuclear and cytoplasmic features.

[0115] Example 12 is the method of any previous or subsequent Example, wherein generating, by the biomarker detection engine, the smear image data from the smear image comprises: processing, using a machine-learning model, the smear image to identify regions of the smear image comprising quality cellular presence; and generating the smear image data by retaining the identified regions and removing regions lacking quality cellular presence.

[0116] Example 13 is the method of any previous or subsequent Example, wherein: the treatment therapy comprises two or more treatment therapies; the one or more deep learning models comprises two or more deep learning models, each deep learning model trained to generate a prediction for a treatment outcome of a respective treatment therapy of the two or more treatment therapies; generating, using the one or more deep learning models, the prediction of the treatment outcome for the treatment therapy comprises: generating, by a first deep learning model of the two or more deep learning models, a first prediction of the treatment outcome for a first treatment therapy of the two or more treatment therapies based on the smear image data; and generating, by a second deep learning model of the two or more deep learning models, a second prediction of the treatment outcome for a second treatment therapy of the two or more treatment therapies based on the smear image data; and generating the recommendationAttorney Docket No. 303.0138WO for the treatment plan comprises evaluating the first prediction and the second prediction to identify a treatment option with a highest likelihood of positive outcome for the patient.

[0117] Example 14 is the method of any previous or subsequent Example, wherein: the myeloid malignancy is acute myeloid leukemia (AML); the treatment therapy comprises a combination therapy comprising venetoclax and azacitidine; and the prediction of treatment outcome is a prediction of whether the patient will achieve complete remission (CR) or complete remission with incomplete recovery of blood count (CRi) in response to the combination therapy.

[0118] Example 15 is a computer- readable storage media comprising processorexecutable instructions configured to cause one or more processors to: receive a smear image of cells from a biopsy sample collected from a patient diagnosed with a myeloid malignancy; generate, using a machine-learning model, a smear image data from the smear image; generate, using a deep learning model, a prediction of a treatment outcome for a treatment therapy of the myeloid malignancy based on the smear image data.

[0119] Example 16 is the computer- readable storage media of any previous or subsequent Example, wherein the processor-executable instructions to generate, using the machine-learning model, the smear image data, when executed by the one or more processors, further direct the one or more processors to: section the smear image into a plurality of regions, each region comprising a grid formed of a plurality of tiles; identify, by the machine-learning model, a plurality of empty tiles within the plurality of regions lacking cellular presence; remove the plurality of empty tiles to generate a plurality of remaining tiles; filter, by the machine-learning model, the plurality of remaining tiles based on quality of cellular presence within a respective remaining tile; and generate the smear image data by selecting a subset of remaining tiles comprising quality cellular presence.

[0120] Example 17 is the computer- readable storage media of any previous or subsequent Example, wherein the one or more processor-executable instructions to generate, using the deep learning model, the prediction of the treatment outcome, when executed by the one or more processors, further direct the one or more processors to: generate, by the deep learning model, a binary classification indicating whether the patient will achieve complete remission (CR) or complete remission with incomplete recovery of blood count (CRi) in response to the treatment therapy; and generate, by the deep learning model, a confidence score associated with the binary classification.

[0121] Example 18 is the computer- readable storage media of any previous or subsequent Example, wherein: the treatment therapy comprises two or more treatmentAttorney Docket No. 303.0138WO therapies; and the one or more processor-executable instructions to generate the prediction of the treatment outcomes for the treatment therapy, when executed by the one or more processors, further direct the one or more processors to: generate, using the deep learning model, a first prediction of the treatment outcome for a first treatment therapy of the two or more treatment therapies based on the smear image data; and generate, using the deep learning model, a second prediction of the treatment outcome for a second treatment therapy of the two or more treatment therapies based on the smear image data.

[0122] Example 19 is the computer-readable storage media of any previous or subsequent Example, wherein the deep learning model comprises a convolutional neural network that exhibits an area under a receiver operating characteristic (AUROC) curve greater than 0.8 when validated using cross-validation.

[0123] Example 20 is the computer-readable storage media of any previous or subsequent Example, wherein: the myeloid malignancy is acute myeloid leukemia (AML); and the treatment therapy comprises a combination therapy comprising vcnctoclax and azacitidine.

Claims

Attorney Docket No. 303.0138WOCLAIMSWhat is claimed is:

1. A biomarker detection engine for predicting a treatment outcome for a patient diagnosed with a myeloid malignancy in response to a treatment therapy, the biomarker detection engine comprising: a computer-readable storage media comprising processor-executable instructions stored thereon; and one or more processors coupled to the computer-readable storage media and configured to execute the processor-executable instructions that, when executed by the one or more processors, direct the biomarker detection engine to at least: receive a smear image of cells from a biopsy sample collected from the patient; preprocess the smear image to generate a smear image data; generate, using one or more deep learning models, a prediction of the treatment outcome for the treatment therapy based on the smear image data; and generate a recommendation for a treatment plan for the patient based on the prediction of the treatment outcome.

2. The biomarker detection engine of claim 1, wherein the processor-executable instructions to preprocess the smear image to generate the smear image data, when executed by the one or more processors, further direct the biomarker detection engine to: section the smear image into a plurality of regions, each region comprising a grid formed of a plurality of tiles; process a first region of the plurality of regions to identify at least one tile of the plurality of tiles within the first region comprising a lack of cellular presence; tag the at least one tile for removal; and generate the smear image data by removing a plurality of tagged tiles, wherein the plurality of tagged tiles comprises the at least one tile.

3. The biomarker detection engine of claim 1, wherein the processor-executable instructions to preprocess the smear image to generate the smear image data, when executed by the one or more processors, further direct the biomarker detection engine to: section the smear image into a plurality of regions, each region comprising a grid formed of a plurality of tiles: identify a plurality of tiles within each region comprising cellular presence;Attorney Docket No. 303.0138WO identify a subset of tiles from the plurality of tiles comprising quality cellular presence; and generate the smear image data by removing remaining tiles of the plurality of tiles from the smear image.

4. The biomarker detection engine of claim 1, wherein the processor-executable instructions to generate, using the one or more deep learning models, the prediction of the treatment outcome for the patient, when executed by the one or more processors, further direct the biomarker detection engine to: submit the smear image data as an input into the one or more deep learning models, wherein the one or more deep learning models generate, responsive to the input, an output comprising: a binary classification indicating whether the patient will respond to the target therapy; and a confidence score associated with the binary classification.

5. The biomarker detection engine of claim 1, wherein: the myeloid malignancy is acute myeloid leukemia (AML); the treatment therapy comprises a combination therapy comprising venetoclax and azacitidine: the prediction of treatment outcome is a prediction of whether the patient will achieve complete remission (CR) or complete remission with incomplete recovery of blood count (CRi) in response to the combination therapy; and the recommendation for the treatment plan comprises recommending the combination therapy when the prediction indicates the patient is likely to achieve CR or CRi, or recommending an alternative treatment when the prediction indicates the patient is unlikely to achieve CR or CRi.

6. The biomarker detection engine of claim 1, wherein the one or more deep learning models comprise a convolutional neural network having hyperparameters tuned through a hyperparameter optimization process configured to improve biomarker detection accuracy in smear images.Attorney Docket No. 303.0138WO7. The biomarker detection engine of claim 6, wherein the convolutional neural network exhibits an average area under a receiver operating characteristic (AUROC) curve greater than 0.8 when validated using a five-fold stratified cross-validation process.

8. A method for predicting treatment outcomes for a patient diagnosed with a myeloid malignancy in response to a treatment therapy, the method comprising: receiving, by a biomarker detection engine, a smear image of cells from a biopsy sample collected from the patient; generating, by the biomarker detection engine, a smear image data from the smear image; generating, using one or more deep learning models, a prediction of a treatment outcome for the treatment therapy based on the smear image data; and generating, by the biomarker detection engine, a recommendation for a treatment plan for the patient based on the prediction of the treatment outcome.

9. 1’he method of claim 8, wherein the method further comprises: training the one or more deep learning models using a training dataset comprising a plurality of historical smear images and corresponding treatment outcomes; and tuning model hyperparameters through a hyperparameter optimization process executed by a trainer to improve predictive performance; and wherein the one or more deep learning models comprise a convolutional neural network trained using a five-fold stratified cross-validation process to detect biomarker variations within the historical smear images and correlate the variations with treatment outcomes.

10. The method of claim 8, wherein: the smear image is stained according to a cytological staining technique configured to enhance visualization of cellular morphology; and the one or more deep learning models are trained on images of smear samples prepared using the cytological staining technique.Attorney Docket No. 303.0138WO11. The method of claim 10, wherein the cytological staining technique comprises a Romanowski-type stain configured to enhance visualization of cellular nuclear and cytoplasmic features.

12. The method of claim 8, wherein generating, by the biomarker detection engine, the smear image data from the smear image comprises: processing, using a machine-learning model, the smear image to identify regions of the smear image comprising quality cellular presence; and generating the smear image data by retaining the identified regions and removing regions lacking quality cellular presence.

13. The method of claim 8, wherein: the treatment therapy comprises two or more treatment therapies; the one or more deep learning models comprises two or more deep learning models, each deep learning model trained to generate a prediction for a treatment outcome of a respective treatment therapy of the two or more treatment therapies; generating, using the one or more deep learning models, the prediction of the treatment outcome for the treatment therapy comprises: generating, by a first deep learning model of the two or more deep learning models, a first prediction of the treatment outcome for a first treatment therapy of the two or more treatment therapies based on the smear image data; and generating, by a second deep learning model of the two or more deep learning models, a second prediction of the treatment outcome for a second treatment therapy of the two or more treatment therapies based on the smear image data; and generating the recommendation for the treatment plan comprises evaluating the first prediction and the second prediction to identify a treatment option with a highest likelihood of positive outcome for the patient.

14. The method of claim 8, wherein: the myeloid malignancy is acute myeloid leukemia (AML); the treatment therapy comprises a combination therapy comprising venetoclax and azacitidine; andAttorney Docket No. 303.0138WO the prediction of treatment outcome is a prediction of whether the patient will achieve complete remission (CR) or complete remission with incomplete recovery of blood count (CRi) in response to the combination therapy.

15. A computer- readable storage media comprising processor-executable instructions configured to cause one or more processors to: receive a smear image of cells from a biopsy sample collected from a patient diagnosed with a myeloid malignancy; generate, using a machine-learning model, a smear image data from the smear image; generate, using a deep learning model, a prediction of a treatment outcome for a treatment therapy of the myeloid malignancy based on the smear image data.

16. The computer-readable storage media of claim 15, wherein the processorexecutable instructions to generate, using the machine-learning model, the smear image data, when executed by the one or more processors, further direct the one or more processors to: section the smear image into a plurality of regions, each region comprising a grid formed of a plurality of tiles; identify, by the machine-learning model, a plurality of empty tiles within the plurality of regions lacking cellular presence; remove the plurality of empty tiles to generate a plurality of remaining tiles; filter, by the machine-learning model, the plurality of remaining tiles based on quality of cellular presence within a respective remaining tile; and generate the smear image data by selecting a subset of remaining tiles comprising quality cellular presence.

17. The computer-readable storage media of claim 15, wherein the one or more processor-executable instructions to generate, using the deep learning model, the prediction of the treatment outcome, when executed by the one or more processors, further direct the one or more processors to: generate, by the deep learning model, a binary classification indicating whether the patient will achieve complete remission (CR) or complete remission with incomplete recovery of blood count (CRi) in response to the treatment therapy; and generate, by the deep learning model, a confidence score associated with the binary classification.Attorney Docket No. 303.0138WO18. The computer-readable storage media of claim 15, wherein: the treatment therapy comprises two or more treatment therapies; and the one or more processor-executable instructions to generate the prediction of the treatment outcomes for the treatment therapy, when executed by the one or more processors, further direct the one or more processors to: generate, using the deep learning model, a first prediction of the treatment outcome for a first treatment therapy of the two or more treatment therapies based on the smear image data; and generate, using the deep learning model, a second prediction of the treatment outcome for a second treatment therapy of the two or more treatment therapies based on the smear image data.

19. The computer-readable storage media of claim 15 , wherein the deep learning model comprises a convolutional neural network that exhibits an area under a receiver operating characteristic (AUROC) curve greater than 0.8 when validated using cross-validation.

20. The computer- readable storage media of claim 15, wherein: the myeloid malignancy is acute myeloid leukemia (AML); and the treatment therapy comprises a combination therapy comprising venetoclax and azacitidine.

Citation Information

Patent Citations

  • Machine learning for predicting cancer genotype and treatment response using digital histopathology images

    WO2023042184A1

  • Methods relating to treatment of acute myeloid leukemia

    WO2023081721A1