System and method for virtual pathology

WO2026178060A1PCT designated stage Publication Date: 2026-08-27VENTANA MEDICAL SYSTEMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/015572
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-20
Filing Date
2026-02-17
Publication Date
2026-08-27

Smart Images

  • Figure US2026015572_27082026_PF_FP_ABST
    Figure US2026015572_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A method of determining a target response to a user query by a virtual pathology system based on machine learning includes identifying, by a virtual pathology encoder of the virtual pathology system, visual data corresponding to the user query; identifying, by the virtual pathology encoder, textual data corresponding to the user query; identifying, by the virtual pathology encoder, classification and detection data corresponding to the visual data; encoding, by the virtual pathology encoder, the visual data to generate first embeddings; encoding, by the virtual pathology encoder, the textual data and the classification and detection data to generate second embeddings; and providing, by the virtual pathology encoder, the first and second embeddings to a large language model (LLM) to generate the target response corresponding to the user query.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR VIRTUAL PATHOLOGYCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 760,867, filed on February 20, 2025, the entire content of which is incorporated herein by reference.FIELD

[0002] Aspects of some embodiments of the present disclosure relate to a system and method for computational pathology.BACKGROUND

[0003] Before digitization, examining tissue required a microscope, giving pathologists exclusive access to tissue images and the related information. Today, however, slides are being scanned, allowing patients and clinicians to view pathology images through electronic medical records. While this advancement increases accessibility, it also raises questions for patients who may not understand what they are seeing, highlighting the need for timely pathology consultations.Globally, digital pathology consultation is a rapidly expanding industry, with millions of slides being shared virtually between labs and servers. Given the shortage of trained pathologists and the tedium involved in answering basic queries to nonpathologists, this growth underscores the urgent demand for real-time, clear language explanations of digital pathology images for purposes such as screening, quality assurance, image analysis, and diagnostic consultations.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background and therefore the information discussed in this Background section does not necessarily constitute prior art.SUMMARY

[0005] Aspects of some embodiments of the present disclosure are directed to a virtual pathology system capable of improving the interpretation of digital pathology images. This system is capable of receiving a user query and analyzing it to identify both textual and visual data relevant to the query. By generating embeddings from these data types, the system employs a large language model to produce a target response that offers a clear and concise explanation of the pathology images in response to the user query. This innovative approach not only enhances the understanding of slide images but also streamlines the process of obtainingpathology consultations, thereby increasing the overall efficiency and effectiveness of digital pathology services.

[0006] According to some embodiments of the present disclosure, there is provided a method of determining a target response to a user query by a virtual pathology system based on machine learning, the method including: identifying, by a virtual pathology encoder of the virtual pathology system, visual data corresponding to the user query; identifying, by the virtual pathology encoder, textual data corresponding to the user query; identifying, by the virtual pathology encoder, classification and detection data corresponding to the visual data; encoding, by the virtual pathology encoder, the visual data to generate first embeddings; encoding, by the virtual pathology encoder, the textual data and the classification and detection data to generate second embeddings; and providing, by the virtual pathology encoder, the first and second embeddings to a large language model (LLM) to generate the target response corresponding to the user query.

[0007] In some embodiments, the method further includes: receiving, by the virtual pathology encoder, the user query from a user device; and transmitting, by the LLM, the target response to the user device.

[0008] In some embodiments, the visual data includes a plurality of image patches of a whole slide image corresponding to the user query.

[0009] In some embodiments, the whole slide image includes a histology hematoxylin and eosin (H&E) image.

[0010] In some embodiments, the textual data includes natural language data corresponding to one or more pathologist reports and meta data corresponding to a whole slide image, and wherein the whole slide image corresponds to the user query.

[0011] In some embodiments, the identifying the classification and detection data includes retrieving the classification and detection data from an external source, the classification and detection data corresponding to the user query.

[0012] In some embodiments, the identifying the classification and detection data includes: determining, by one or more models, a biomarker associated with a whole slide image, wherein the whole slide image corresponds to the user query, and wherein the classification and detection data includes the biomarker.

[0013] In some embodiments, the biomarker includes at least one of an estrogen receptor (ER), a progesterone receptor (PR), a human epidermal growth factor receptor 2 (HER2), a programmed cell death ligand 1 (PD-L1) marker, a c-MET marker, or the like.

[0014] In some embodiments, the identifying the classification and detection data includes: classifying, by one or more classifiers, at least one of a tissue type or atissue category associated with a whole slide image, wherein the whole slide image corresponds to the user query, and wherein the classification and detection data includes the at least one of the tissue type or the tissue category.

[0015] In some embodiments, the tissue type includes at least one of breast cancer tissue, skin, lung tissue, colorectal cancer (CRC), ovarian tissue, and wherein the tissue category includes at least one of necrosis, tumor, stroma, immune cells, collagen, red blood cells, non-tumor regions, or the like.

[0016] According to some embodiments of the present disclosure, there is provided a system for determining a target response to a user query by a virtual pathology system based on machine learning, the virtual pathology system including: a processor; and a memory storing instructions that, when executed on the processor, cause the processor to perform: identifying visual data corresponding to the user query; identifying textual data corresponding to the user query; identifying classification and detection data corresponding to the visual data; encoding the visual data to generate first embeddings; encoding the textual data and the classification and detection data to generate second embeddings; and generating the target response corresponding to the user query.

[0017] In some embodiments, the processor of the virtual pathology system is further caused to perform: receiving the user query; and transmitting the target response.

[0018] In some embodiments, the visual data includes a plurality of image patches of a whole slide image corresponding to the user query.

[0019] In some embodiments, the whole slide image includes a histology hematoxylin and eosin (H&E) image.

[0020] In some embodiments, the textual data includes natural language data corresponding to one or more pathologist reports and meta data corresponding to a whole slide image, and wherein the whole slide image corresponds to the user query.

[0021] In some embodiments, the identifying classification and detection data includes: generating the classification and detection data based on a whole slide image corresponding to the user query.

[0022] In some embodiments, the generating classification and detection data includes: classifying, by one or more classifiers, at least one of a tissue type or a tissue category associated with a whole slide image, wherein the whole slide image corresponds to the user query, and wherein the classification and detection data includes the at least one of the tissue type or the tissue category.

[0023] In some embodiments, the tissue type includes at least one of breast cancer tissue, skin, lung tissue, colorectal cancer (CRC), ovarian tissue, and whereinthe tissue category includes at least one of necrosis, tumor, stroma, immune cells, collagen, red blood cells, non-tumor regions, or the like.

[0024] In some embodiments, the generating classification and detection data includes: determining, by one or more models, a biomarker associated with a whole slide image, wherein the whole slide image corresponds to the user query, and wherein the classification and detection data includes the biomarker.

[0025] In some embodiments, the biomarker includes at least one of an estrogen receptor (ER), a progesterone receptor (PR), a human epidermal growth factor receptor 2 (HER2), a programmed cell death ligand 1 (PD-L1) marker, a c-MET marker, or the like.BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Non-limiting and non-exhaustive embodiments according to the present disclosure are described with reference to the following figures, wherein like reference numerals refer to like parts throughout the various views unless otherwise specified.

[0027] FIG. 1 is a block diagram illustrating a virtual pathology system, according to some embodiments of the present disclosure.

[0028] FIG. 2 is a block diagram illustrating the internal structure of the virtual pathology encoder, according to some embodiments of the present disclosure.

[0029] FIG. 3 is a block diagram illustrating the internal structure of the image data extractor of the virtual pathology encoder, according to some embodiments of the present disclosure.

[0030] FIG. 4 illustrates a user interface (III) of the virtual pathology system, according to some embodiments of the present disclosure.

[0031] FIG. 5 is a flow diagram illustrating a process of generating a target response based on a user query by the virtual pathology system, according to some embodiments of the present disclosure.DETAILED DESCRIPTION

[0032] Hereinafter, aspects of some example embodiments will be described in more detail with reference to the accompanying drawings, in which like reference numbers refer to like elements throughout. The present invention, however, may be embodied in various different forms, and should not be construed as being limited to only the illustrated embodiments herein. Rather, these embodiments are provided as examples so that this disclosure will be thorough and complete, and will fully convey the aspects and features of the present invention to those skilled in the art. Accordingly, processes, elements, and techniques that are not necessary to thosehaving ordinary skill in the art for a complete understanding of the aspects and features of the present invention may not be described. Unless otherwise noted, like reference numerals denote like elements throughout the attached drawings and the written description, and thus, descriptions thereof will not be repeated. In the drawings, the relative sizes of elements, layers, and regions may be exaggerated for clarity.

[0033] In the field of pathology, traditional methods of examining tissue samples using microscopes have limited access to pathology information to pathologists, restricting patients and clinicians who lack specialized equipment and expertise. With digitization, pathology slides are now accessible through electronic medical records, broadening access but introducing challenges for patients who may not understand the images. This highlights the need for timely and comprehensible pathology consultations. Despite advancements, current digital pathology solutions often fail to provide real-time, clear language explanations, focusing instead on image sharing and storage. This gap can delay diagnosis and treatment, as patients and clinicians struggle to interpret images without expert guidance, underscoring the inadequacies of existing systems in meeting the global demand for effective communication of pathology findings.

[0034] The present system addresses these challenges by introducing a virtual pathology system designed to enhance the interpretation of digital pathology images. This system is configured to receive a user query and analyze the query to identify both textual and visual data corresponding to the query. By generating embeddings from these data types, the system utilizes a large language model to produce a target response that provides a clear and concise explanation of the pathology images. This approach not only facilitates a better understanding of slide images but also streamlines the process of obtaining pathology consultations, thereby improving the overall efficiency and effectiveness of digital pathology services.

[0035] In some examples, the vision-to-language capability of the virtual pathology system may be utilized to assist pathologists in their analysis of a tissue sample, may aid in quality assurance and pre-screening, and may also enable interpretation of pathology images for non-experts, such as patients and nonpathologist clinicians.

[0036] FIG. 1 is a block diagram illustrating a virtual pathology system 1 , according to some embodiments of the present disclosure.

[0037] In some embodiments, the virtual pathology system 1 is configured to receive a user query 10 that is inputted by a user, such as a clinician or a patient, via a user device 300, to process and analyze the query 10 based on available visualand textual data corresponding to the query 10, and to generate a target response for transmission to the user device 300.

[0038] The user device 300 may be configured to receive the user query 10 from a user, for example, via a display (e.g., touch sensitive display), a keyboard, etc. The user device 300 may further be configured to receive the user query 10 from the user, for example, via a microphone operatively connected to and / or integrated with the user device 300. In some examples, the user device 300 may include a mobile device such as a smart phone or tablet, a personal computer, or the like.

[0039] According to some embodiments, the virtual pathology system 1 includes a virtual pathology encoder 100 and a large language model (LLM) 200. The virtual pathology encoder 100 analyzes the user query 10 to extract relevant data. In some embodiments, the virtual pathology encoder 100 generates first embeddings (e.g., image embeddings) 20 and second embeddings (e.g., textual embeddings) 30 from the user query 10. The first embeddings 20 may be derived from visual data (e.g., a whole slide image of a patient’s tissue sample) associated with the query 10, while the second embeddings 30 may be derived from textual data and classification and detection data associated with the query 10.

[0040] The LLM 200 then processes the first embeddings 20 and second embeddings 30 to generate a target response 40. In some examples, the user query 10 may be a request for an explanation of a pathology image related to the user query 10, and the target response 40 may provide a natural-language, clear and concise explanation of the pathology image.

[0041] Once the target response 40 is generated, the virtual pathology system 1 transmits the response back to the user device 300. The user device 300 may display the target response 40 to the user, facilitating further interaction. For example, the virtual pathology system 1 may process further (follow-up) user queries in the same way, thus engaging the user in a conversation.

[0042] In some embodiments, the target response 40 may be sent to a server (e.g., a remote server or a cloud server) 350 for further processing or storage, ensuring that the information is accessible for future reference or analysis.

[0043] FIG. 2 is a block diagram illustrating the internal structure of the virtual pathology encoder 100, according to some embodiments of the present disclosure.

[0044] According to some embodiments, the virtual pathology encoder 100 includes a query analyzer 110, a first encoder (e.g., an image encoder) 120, and a second encoder (e.g., a text encoder) 130.

[0045] In some embodiments, the query analyzer 110 is configured to analyze the user query 10 and to identify visual data 12 as well as textual data 14 corresponding to the user query 10. In some examples, the query analyzer 110 may utilize one ormore machine learning models (MLs; e.g., deep learning models) in parsing and analyzing the user query 10; however, embodiments of the present disclosure are not limited thereto.

[0046] In some embodiments, the visual data 12 identified by the query analyzer 110 may include a WSI associated with the user query 10. The WSI may be a histology hematoxylin and eosin (H&E) image of a tissue sample; however, embodiments of the present disclosure are not limited thereto. In some examples, the query analyzer 110 may use an identifier included within the user query 10 and associated with the WSI image to retrieve the WSI from an external source (e.g., a data base) 450, which is communicatively coupled to the virtual pathology encoder 100. In other examples, the WSI may be provided by the user as part of the user query 10 or the interaction with the user (which includes the user query 10).However, embodiments of the present disclosure are not limited thereto, and the query analyzer 110 may analyze the user query 10 via the one or more machine learning models to identify visual data 12. In some embodiments, rather than identify / retrieve an entire WSI, the query analyzer 110 may identify / retrieve a plurality of image patches / tiles that are extracted from the WSI.

[0047] In some embodiments, the first encoder 120 is configured to receive the visual data 12 (e.g., the WSI or at least some of its constituent image patches / tiles), and to encode the visual data 12 to generate the first embeddings 20 for consumption by the LLM 200. For example, the first encoder 120 may encode the visual information within the image patches / tiles and their locations relative to the WSI. The encoding process performed by the first encoder 120 includes processing and transforming the visual data 12 into a format that the LLM 200 can understand.

[0048] According to some embodiments, the textual data 14 includes natural language data from one of one or more pathologist reports / annotations and / or meta data corresponding to a whole slide image (WSI) of a tissue sample, which is associated with the user or is provided as part of the interaction with the user (which the user query 10 is a part of). The meta data may include digital information corresponding to the WSI, such as the author, date and / or time of capture, access records, amongst other information corresponding to the WSI. In some embodiments, the query analyzer 110 may retrieve the one of one or more pathologist reports / annotations from the external source (e.g., the database) 450 based on a comparison between the reference pathologist reports / annotations stored thereon and the user query 10. The meta data may be embedded within the WSI or be stored at the external source 450 and associated with the WSI.

[0049] In some embodiments, the textual data 14 also includes classification and detection data 15 corresponding to (e.g., derived from) the WSI. In some examples,the classification and detection data 15 may include text explanations corresponding to probabilities for tissue types (e.g., breast cancer tissue, skin, lung), tissue categories (e.g., necrosis, tumor), or specific markers (e.g., an estrogen receptor (ER), a progesterone receptor (PR), etc.). The query analyzer 110 may retrieve the classifications and detection data 15, which may be derived by external MLs (e.g., deep learning models), from the external source (e.g., database) 450. However, embodiments of the present disclosure are not limited thereto.

[0050] According to some embodiments, the virtual pathology encoder 100 also includes an image data extractor 140 that utilizes one or more deep learning models to extract the classification and detection data 15 from the WSI or a plurality of image tiles / patches corresponding to the WSI. The image data extractor 140 may then convert these probability predictions into corresponding text explanations for processing by the second encoder 130.

[0051] Thus, by integrating and fusing the visual and textual data, the LLM 200 is able to generate text-based image explanations. This architecture allows users to input images, ask questions, and receive answers from the virtual pathology system 1.

[0052] FIG. 3 is a block diagram illustrating the internal structure of the image data extractor 150 of the virtual pathology encoder 100, according to some embodiments of the present disclosure.

[0053] According to some embodiments, the image data extractor 150 includes a biomarker detector 152 and a classifier 154.

[0054] In some embodiments, the biomarker detector 152 is configured to analyze the features (e.g., the tile-level and / or cell-level features) of a given WSI and to generate a corresponding prediction (e.g., a biomarker prediction) 15a regarding the presence or absence of one or more biomarkers. In some examples, the biomarker detector 152 utilizes machine learning-based models to identify the one or more biomarkers based on morphology from the WSI. In some examples, the biomarker detector 152 includes a feature analyzer for analyzing and extracting features (e.g., cell-level and / or tile-level features) from a WSI and generating corresponding embeddings data, and a biomarker predictor for generating the biomarker prediction 15a, which may be a slide-level prediction, based on the embeddings data.

[0055] The biomarkers detected by the biomarker detector 152 may include at least one of an estrogen receptor (ER), a progesterone receptor (PR), a human epidermal growth factor receptor 2 (HER2), an MYC-driven high-grade B-cell lymphoma (HGBL), a mutation of an individual gene (e.g., loss of function single nucleotide variation in the TP53 gene), a gene mutation signature (e.g., MCDsignature based on co-occurrence of MYD88 and CD79B mutations), the expression level of an individual gene or protein (e.g., MYC), a gene expression profile or signature (e.g., cell-of-origin signature), the infiltration of immune cells in the microenvironment (e.g., lymphocytes), or the like.

[0056] The biomarker prediction 15a that is output by the biomarker detector 152 may be a binary output (e.g., ‘0’ or T, or '+’ or ‘-‘) indicating the presence or absence of a biomarker for which the biomarker detector 152 is trained. In some examples, the biomarker prediction 15a may be a confidence level or probability that the biomarker is present in the tissue sample associated with the WSI. However, these are merely examples, and embodiments of the present disclosure are not limited thereto.

[0057] In some embodiments, the classifier 155 is configured to classify at least one of the tissue type 15b and / or the tissue category 15c corresponding to the visual data 12. The classifier 155 may include one or more machine learning (ML) models configured to identify the particular tissue type and / or tissue category of the tissue sample represented by the WSI. Each of the ML models may be capable of tissue classification by inputting an image, assigning importance (e.g., via learnable weights and biases) to various aspects / objects in the image, and differentiating one from the other. The neural network may include a convolutional neural network (ConvNet / CNN), a recurrent neural network (RNN) with convolution operation, and / or the like.

[0058] In some embodiments, the tissue type 15b includes at least one of breast cancer tissue, skin, lung tissue, colorectal cancer (CRC), ovarian tissue, or the like. In some embodiments, the tissue category 15c may include at least one of necrosis, tumor, stroma, immune cells, collagen, red blood cells, non-tumor regions, or the like. However, embodiments of the present disclosure are not limited thereto; for example, the classifier 154 may be trained to estimate tumor grade and to identify tumor types or sub-types.

[0059] In some embodiments, the biomarker detector 152 and the classifier 154 operate on a plurality of image tiles / patches extracted from the WSI that may or may not be included in the visual data 12 applied to the image data extractor 150. In embodiments in which the visual data does not include the image tiles / patches, the biomarker detector 152 includes a WSI processor that extracts the tiles / patches from a WSI and normalizes them to ensure uniformity across all tiles / patches, and then passes on these normalized tiles / patches to the biomarker detector 152 and the classifier 154 for further analysis and processing.

[0060] FIG. 4 illustrates a user interface (Ul) 400 of the virtual pathology system 1 , according to some embodiments of the present disclosure.

[0061] According to some embodiments, the III 400 is configured to receive the user query 10 from the user. For example, the III 400 may be configured to receive a textual input 10a and a visual input 10b associated with the user query 10. The textual input 10a may correspond to the particular question being asked by the user, and the visual input 10b may correspond to the particular whole slide image corresponding to the user query (e.g., the whole slide image inputted by the user via the user device).

[0062] In some embodiments, the III 400 is configured to receive the visual input 10b from the user via the upload box 11 a. For example, the user may select the upload box 11a to upload the visual input 10b, such as a WSI, to the user device 300. The user device 300 may be configured to associate the visual input 10b with the textual input 10a inputted by the user to determine the user query 10, which is received by the virtual pathology encoder 100 to generate the first embeddings 20 and the second embeddings 30.

[0063] As shown in FIG. 4, through the III 400, a user is able to have a naturallanguage conversation with the virtual pathology system 1 , which is made up of a sequence of user queries 10 and target responses 40. While not specifically shown, the III may also present to the user a set of pre-engineered prompts that are customized for specific diagnosis use cases. Additionally, the III may allow the user to access any linked lab information system, which may assist pathologists in generating a complete lab report.

[0064] The user interface exemplified by FIG. 4 may be implemented as an application running on the user device 300 or as an web application running on a remote server. In some examples, the web application may include three components: a controller, a model worker, and a web client. The controller may implement the REST API for the web client to request LLM queries. The model worker may perform the inference work of the virtual pathology system 1 that is dispatched by the controller. The web client may implement the web III. Users may upload the digital pathology images and input the text questions on the web Ul. The web client then sends them to the controller's REST API. The model worker generates the answers to users’ questions and returns the answers to the controller, which is then returned to the web client and displayed in the Ul.

[0065] FIG. 5 illustrates a process 500 of determining a target response to a user query by the virtual pathology system, according to some embodiments of the present disclosure.

[0066] In some embodiments, the virtual pathology encoder 100 of the virtual pathology system 1 identifies visual data 12 corresponding to the user query 10 (S502). The virtual pathology encoder 100 may receive the user query 10 from auser device 300. The visual data 12 may include a plurality of image patches of a WSI corresponding to the user query 10. The WSI may include a histology hematoxylin and eosin (H&E) image.

[0067] The virtual pathology encoder 100 also identifies textual data 14 corresponding to the user query 10 (S504). The textual data 14 may include natural language data corresponding to one or more pathologist reports and meta data corresponding to the WSI.

[0068] Additionally, the virtual pathology encoder 100 may identify classification and detection data 15 corresponding to the visual data 12 (S506). In some examples, the virtual pathology encoder 100 may retrieve the classification and detection data 15 from an external source (e.g., a database) 450. The identifying the classification and detection data 15 may include: determining, by one or more models, a biomarker associated with a whole slide image and / or at least one of a tissue type or a tissue category associated with the WSI. Thus, the classification and detection data may include the biomarker and / or at least one of the tissue type or the tissue category.

[0069] The virtual pathology encoder 100 encodes the visual data 12 to generate first embeddings 20 (S508), and encodes the textual data 14 and the classification and detection data 15 to generate second embeddings 30 (S510).

[0070] The virtual pathology encoder 100 provides the first and second embeddings 20 and 30 to the large language model (LLM) 200 to generate the target response 40 corresponding to the user query 10 (S512). The target response may then be sent to the user device 300.

[0071] According to various embodiments of the present disclosure, the virtual pathology system 1 is implemented using one or more processing circuits or electronic circuits configured to perform various operations as described above. Types of electronic circuits may include a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence (Al) accelerator (e.g., a vector processor, which may include vector arithmetic logic units configured efficiently perform operations common to neural networks, such dot products and softmax), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processor (DSP), or the like. For example, in some circumstances, aspects of embodiments of the present disclosure are implemented in program instructions that are stored in a non-volatile computer readable memory where, when executed by the electronic circuit (e.g., a CPU, a GPU, an Al accelerator, or combinations thereof), perform the operations described. The operations performed by the virtual pathology system 1 may be performed by a single electronic circuit (e.g., a single CPU, a single GPU, or the like) or may beallocated between multiple electronic circuits (e.g., multiple GPUs or a CPU in conjunction with a GPU). The multiple electronic circuits may be local to one another (e.g., located on a same die, located within a same package, or located within a same embedded device or computer system) and / or may be remote from one other (e.g., in communication over a network such as a local personal area network such as Bluetooth®, over a local area network such as a local wired and / or wireless network, and / or over wide area network such as the internet, such a case where some operations are performed locally and other operations are performed on a server hosted by a cloud computing service). One or more electronic circuits operating to implement the virtual pathology system 1 may be referred to herein as a computer or a computer system, which may include memory storing instructions that, when executed by the one or more electronic circuits, implement the systems and methods described herein.

[0072] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms “a” and “an” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” "includes," and "including," when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items. Expressions such as “at least one of,” when preceding a list of elements, modify the entire list of elements and do not modify the individual elements of the list.

[0073] It will be understood that, although the terms “first,” “second,” “third,” etc., may be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms are used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Thus, a first element, component, region, layer or section described below could be termed a second element, component, region, layer or section, without departing from the spirit and scope of the present invention.

[0074] As used herein, the term "substantially," "about," and similar terms are used as terms of approximation and not as terms of degree, and are intended to account for the inherent deviations in measured or calculated values that would be recognized by those of ordinary skill in the art. Further, the use of “may” when describing embodiments of the present invention refers to “one or moreembodiments of the present invention.” As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively.

[0075] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or the present specification, and should not be interpreted in an idealized or overly formal sense, unless expressly so defined herein.

[0076] Although aspects of some example embodiments of the system and method of generating a target response based on a user query by a virtual pathology system have been described and illustrated herein, various modifications and variations may be implemented, as would be understood by a person having ordinary skill in the art, without departing from the spirit and scope of embodiments according to the present disclosure. Accordingly, it is to be understood that a virtual pathology system and method according to the principles of the present disclosure may be embodiments other than as specifically described herein. The disclosure is also defined in the following claims, and equivalents thereof.

Claims

WHAT IS CLAIMED IS:

1. A method of determining a target response to a user query by a virtual pathology system based on machine learning, the method comprising:identifying, by a virtual pathology encoder of the virtual pathology system, visual data corresponding to the user query;identifying, by the virtual pathology encoder, textual data corresponding to the user query;identifying, by the virtual pathology encoder, classification and detection data corresponding to the visual data;encoding, by the virtual pathology encoder, the visual data to generate first embeddings;encoding, by the virtual pathology encoder, the textual data and the classification and detection data to generate second embeddings; and providing, by the virtual pathology encoder, the first and second embeddings to a large language model (LLM) to generate the target response corresponding to the user query.

2. The method of claim 1 , further comprising:receiving, by the virtual pathology encoder, the user query from a user device; andtransmitting, by the LLM, the target response to the user device.

3. The method of claim 1 , wherein the visual data comprises a plurality of image patches of a whole slide image corresponding to the user query.

4. The method of claim 3, wherein the whole slide image comprises a histology hematoxylin and eosin (H&E) image.

5. The method of claim 1 , wherein the textual data comprises natural language data corresponding to one or more pathologist reports and meta data corresponding to a whole slide image, andwherein the whole slide image corresponds to the user query.

6. The method of claim 1 , wherein the identifying the classification and detection data comprises retrieving the classification and detection data from an external source, the classification and detection data corresponding to the user query.

7. The method of claim 1 , wherein the identifying the classification and detection data comprises:determining, by one or more models, a biomarker associated with a whole slide image,wherein the whole slide image corresponds to the user query, and wherein the classification and detection data comprises the biomarker.

8. The method of claim 7, wherein the biomarker comprises at least one of an estrogen receptor (ER), a progesterone receptor (PR), a human epidermal growth factor receptor 2 (HER2), a programmed cell death ligand 1 (PD-L1) marker, or a c-MET marker.

9. The method of claim 1 , wherein the identifying the classification and detection data comprises:classifying, by one or more classifiers, at least one of a tissue type or a tissue category associated with a whole slide image,wherein the whole slide image corresponds to the user query, and wherein the classification and detection data comprises the at least one of the tissue type or the tissue category.

10. The method of claim 9, wherein the tissue type comprises at least one of breast cancer tissue, skin, lung tissue, colorectal cancer (CRC), or ovarian tissue, andwherein the tissue category comprises at least one of necrosis, tumor, stroma, immune cells, collagen, red blood cells, or non-tumor regions.

11. A virtual pathology system for determining a target response to a user query based on machine learning, the virtual pathology system comprising:a processor; anda memory storing instructions that, when executed on the processor, cause the processor to perform:identifying visual data corresponding to the user query;identifying textual data corresponding to the user query;identifying classification and detection data corresponding to the visual data; encoding the visual data to generate first embeddings;encoding the textual data and the classification and detection data to generate second embeddings; andgenerating the target response corresponding to the user query.

12. The virtual pathology system of claim 11 , wherein the instructions further cause the processor to perform:receiving the user query; andtransmitting the target response.

13. The virtual pathology system of claim 11 , wherein the visual data comprises a plurality of image patches of a whole slide image corresponding to the user query.

14. The virtual pathology system of claim 13, wherein the whole slide image comprises a histology hematoxylin and eosin (H&E) image.

15. The virtual pathology system of claim 11 , wherein the textual data comprises natural language data corresponding to one or more pathologist reports and meta data corresponding to a whole slide image, andwherein the whole slide image corresponds to the user query.

16. The virtual pathology system of claim 11 , wherein the identifying classification and detection data comprises:generating the classification and detection data based on a whole slide image corresponding to the user query.

17. The virtual pathology system of claim 16, wherein the generating classification and detection data comprises:classifying at least one of a tissue type or a tissue category associated with a whole slide image,wherein the whole slide image corresponds to the user query, and wherein the classification and detection data comprises the at least one of the tissue type or the tissue category.

18. The virtual pathology system of claim 17, wherein the tissue type comprises at least one of breast cancer tissue, skin, lung tissue, colorectal cancer (CRC), ovarian tissue, andwherein the tissue category comprises at least one of necrosis, tumor, stroma, immune cells, collagen, red blood cells, or non-tumor regions.

19. The virtual pathology system of claim 16, wherein the generating classification and detection data comprises:determining a biomarker associated with a whole slide image,wherein the whole slide image corresponds to the user query, and wherein the classification and detection data comprises the biomarker.

20. The virtual pathology system of claim 19, wherein the biomarker comprises at least one of an estrogen receptor (ER), a progesterone receptor (PR), a human epidermal growth factor receptor 2 (HER2), a programmed cell death ligand 1 (PD-L1) marker, or a c-MET marker.