System and method for processing electronic images to identify diagnostic tests

A machine learning-based system processes digital images to identify and prioritize diagnostic tests, addressing issues of physician unfamiliarity, availability, and cost, ensuring effective treatment methods are utilized.

JP2026083299APending Publication Date: 2026-05-19PAIGE AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PAIGE AI INC
Filing Date
2026-03-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Diagnostic tests for identifying treatment methods for diseased tissues are often not performed due to physician unfamiliarity, lack of availability, insufficient samples, low pre-test expectations, or high costs, leading to ineffective treatments.

Method used

A system and method using machine learning to process digital images of pathological specimens to identify and prioritize applicable diagnostic tests based on patient-specific information, availability, and cost considerations.

Benefits of technology

Facilitates the identification and prioritization of beneficial diagnostic tests, ensuring appropriate tests are performed, thereby improving treatment efficacy and reducing unnecessary or costly tests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083299000001
    Figure 2026083299000001
  • Figure 2026083299000002
    Figure 2026083299000002
  • Figure 2026083299000003
    Figure 2026083299000003
Patent Text Reader

Abstract

To provide a system and method for processing electronic images to identify suitable diagnostic tests. [Solution] A system and method for processing digital images to identify diagnostic tests is disclosed, the method comprising: receiving one or more digital images associated with a pathological specimen; determining a plurality of diagnostic tests; applying a machine learning system to one or more digital images to identify any of the prerequisites for each of the plurality of diagnostic tests to be applicable, the machine learning system being trained by processing a plurality of training images; using a machine learning model to identify applicable diagnostic tests from the plurality of diagnostic tests based on one or more digital images and prerequisites; and outputting the applicable diagnostic tests to a digital storage device and / or display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims priority to U.S. Provisional Application No. 63 / 104,923, filed October 23, 2020, the entire disclosure of which is hereby incorporated by reference in its entirety.

[0002] Various embodiments of the present disclosure generally relate to image processing methods. More specifically, certain embodiments of the present disclosure relate to systems and methods for processing electronic images to prioritize and / or identify diagnostic tests.

Background Art

[0003] Diagnostic tests for identifying treatment methods and treatment processes for diseased tissues have been continuously developed and are now available for clinical use. Diagnostic tests can benefit patients by excluding ineffective treatments and / or by identifying the treatment methods that are most likely to provide significant benefits in treating the patient's disease by detecting the presence or absence of biomarkers (e.g., the practice known as "precision medicine"). However, various factors such as physicians being unfamiliar with the tests, the tests not being available within the facility, a lack of viable samples for successfully performing the recommended tests, low pre-test expectations that a particular test may yield a positive result for this patient, or the high cost of the treatment identified by the test may result in important diagnostic tests not being performed on patients. The technology presented herein can address this clinical need by identifying and prioritizing which tests may be beneficial to a patient and making that information available to the patient and the physician.

[0004] The background information provided herein is intended to provide a general context for the disclosure. Unless otherwise indicated herein, the matters described in this section are not prior art to the claims of this application, nor are they deemed to be prior art or suggestive of prior art by their inclusion in this section. [Overview of the project] [Means for solving the problem]

[0005] According to certain aspects of this disclosure, a system and method for processing electronic images to recommend diagnostic tests based on tissue samples are disclosed.

[0006] A method for processing digital images to identify diagnostic tests, comprising: receiving one or more digital images associated with a pathological specimen; determining a plurality of diagnostic tests; applying a machine learning system to one or more digital images to identify any of the prerequisites for each of the plurality of diagnostic tests to be applicable, wherein the machine learning system is trained by processing a plurality of training images; using a machine learning model to identify applicable diagnostic tests among the plurality of diagnostic tests based on one or more digital images and prerequisites; and outputting the applicable diagnostic tests to a digital storage device and / or display.

[0007] A system for processing digital images to identify diagnostic tests, the method comprising: receiving one or more digital images associated with a pathological specimen; determining a plurality of diagnostic tests; applying a machine learning system to one or more digital images to identify any of the prerequisites for each of the plurality of diagnostic tests to be applicable, wherein the machine learning system is trained by processing a plurality of training images; using the machine learning system to identify applicable diagnostic tests among the plurality of diagnostic tests based on one or more digital images and prerequisites; and outputting the applicable diagnostic tests to a digital storage device and / or display.

[0008] A non-temporary computer-readable medium that stores instructions causing a processor to perform a method for processing digital images to identify diagnostic tests, wherein the method includes: receiving one or more digital images associated with a pathological specimen; determining a plurality of diagnostic tests; applying a machine learning system to one or more digital images to identify any of the prerequisites for each of the plurality of diagnostic tests to be applicable, the machine learning system being trained by processing a plurality of training images; using the machine learning system to identify applicable diagnostic tests among the plurality of diagnostic tests based on one or more digital images and prerequisites; and outputting the applicable diagnostic tests to a digital storage device and / or display.

[0009] It should be understood that both the above general description and the following detailed description are merely illustrative and do not limit the embodiments of the disclosure as claimed. This specification also provides, for example, the following: (Item 1) A computer implementation method for processing digital images to identify diagnostic tests, Receiving one or more digital images associated with a pathological specimen, Interpreting multiple diagnostic tests, The process of applying a machine learning system to one or more digital images to identify any prerequisites for each of the plurality of diagnostic tests to be applicable, wherein the machine learning system is trained by processing a plurality of training images, and the process of identification is as follows: Using the machine learning model, identify an applicable diagnostic test from among the multiple diagnostic tests based on one or more digital images and the preconditions. The computer implementation method, comprising outputting the applicable diagnostic test to a digital storage device and / or display. (Item 2) The computer implementation method described in item 1 further includes determining additional patient information relating to the patient and / or disease, additional diagnostic test information, and / or additional test preference information relating to the pathological specimen. (Item 3) The computer implementation method according to item 1, further comprising identifying the applicable diagnostic tests, and predicting the negative predictive value (NPV) for each of the plurality of diagnostic tests. (Item 4) Processing multiple training images is Receiving a plurality of digital images associated with at least one previous pathological specimen, wherein the digital images are paired with diagnostic test information relating to the results or values ​​of one or more past diagnostic tests performed on the previous pathological specimen, or tests to establish the applicability of the diagnostic tests to the previous pathological specimen, To predict the applicability of the aforementioned past diagnostic tests, the machine learning system is trained using the aforementioned plurality of digital images and the aforementioned diagnostic test information, wherein the machine learning system includes a multi-binary label machine learning system. Determining at least one threshold of one or more binary outputs of the multi-binary label machine learning system, A computer implementation method according to item 3, comprising outputting a set of parameters from the multi-binary label machine learning system to a digital storage device, wherein the set of parameters includes the at least one threshold. (Item 5) The process involves filtering one or more of the aforementioned digital images to identify the tissue region to be analyzed, The computer implementation method according to item 1, further comprising removing one or more regions from the one or more digital images that are not identified as the target tissue region. (Item 6) The machine learning system is provided with a scoring threshold, Based on the scoring threshold, determine one or more applicable diagnostic tests that exceed the scoring threshold. The computer implementation method according to item 1, further comprising outputting one or more applicable diagnostic tests with scores exceeding the scoring threshold. (Item 7) Based on the applicable diagnostic tests, determine one or more treatment options that may be suitable for the patient, The computer implementation method according to item 1, further comprising outputting the one or more treatment methods to a display. (Item 8) The computer implementation method according to item 1, further comprising displaying the applicable diagnostic tests to the user. (Item 9) A system for processing digital images to identify diagnostic tests, At least one memory to store instructions, At least one processor, Receiving one or more digital images associated with a pathological specimen, Interpreting multiple diagnostic tests, The process of applying a machine learning system to one or more digital images to identify any prerequisites for each of the plurality of diagnostic tests to be applicable, wherein the machine learning system is trained by processing a plurality of training images, and the process of identification is as follows: Using the machine learning system, identify an applicable diagnostic test from among the plurality of diagnostic tests based on one or more digital images and the preconditions, The system includes the at least one processor configured to execute the instructions that perform an operation including outputting the applicable diagnostic test to a digital storage device and / or a display. (Item 10) The system described in item 9 further includes determining additional patient information relating to the patient and / or disease, additional diagnostic test information, and / or additional test preference information relating to the pathological specimen. (Item 11) The system according to item 10, further comprising identifying the applicable diagnostic tests, and predicting the negative predictive value (NPV) for each of the plurality of diagnostic tests. (Item 12) Processing multiple training images is Receiving a plurality of digital images associated with at least one previous pathological specimen, wherein the digital images are paired with diagnostic test information relating to the results or values ​​of one or more past diagnostic tests performed on the previous pathological specimen, or tests to establish the applicability of the diagnostic tests to the previous pathological specimen, To predict the applicability of the aforementioned past diagnostic tests, the machine learning system is trained using the aforementioned plurality of digital images and the aforementioned diagnostic test information, wherein the machine learning system includes a multi-binary label machine learning system. Determining at least one threshold of one or more binary outputs of the multi-binary label machine learning system, Outputting a set of parameters from the multi-binary label machine learning system to a digital storage device, the set of parameters including the at least one threshold value, and the outputting, the system according to item 9 including the above. (Item 13) Filtering the one or more digital images to identify an organizational area to be analyzed. Further comprising removing from the one or more digital images one or more areas not identified as the target organizational area, the system according to item 9 including the above. (Item 14) Providing a scoring threshold value to the machine learning system. Based on the scoring threshold value, determining one or more applicable diagnostic tests with scores exceeding the scoring threshold value. Further comprising outputting the one or more applicable diagnostic tests with scores exceeding the scoring threshold value, the system according to item 9 including the above. (Item 15) Based on the applicable diagnostic tests, determining one or more treatment methods that may be suitable for the patient. Further comprising outputting the one or more treatment methods to a display, the system according to item 9 including the above. (Item 16) Further comprising displaying the applicable diagnostic tests to a user, the system according to item 9 including the above. (Item 17) A non-temporary computer-readable medium storing instructions that, when executed by a processor, cause the processor to execute a method for processing digital images to identify diagnostic tests, the method including: Receiving one or more digital images associated with a pathological specimen. Determining a plurality of diagnostic tests. The process of applying a machine learning system to one or more digital images to identify any prerequisites for each of the plurality of diagnostic tests to be applicable, wherein the machine learning system is trained by processing a plurality of training images, and the process of identification is as follows: Using the machine learning model, identify an applicable diagnostic test from among the multiple diagnostic tests based on one or more digital images and the preconditions. The non-temporary computer-readable medium includes outputting the applicable diagnostic test to a digital storage device and / or display. (Item 18) Non-temporary computer-readable media as described in item 17, further including determining additional patient information relating to the patient and / or disease, additional diagnostic test information, and / or additional test preference information relating to the pathological specimen. (Item 19) Identifying the applicable diagnostic tests further includes predicting the negative predictive value (NPV) for each of the multiple diagnostic tests, as described in item 18, for a non-temporary computer-readable medium. (Item 20) Processing multiple training images is Receiving a plurality of digital images associated with at least one previous pathological specimen, wherein the digital images are paired with diagnostic test information relating to the results or values ​​of one or more past diagnostic tests performed on the previous pathological specimen, or tests to establish the applicability of the diagnostic tests to the previous pathological specimen, To predict the applicability of the aforementioned past diagnostic tests, the machine learning system is trained using the aforementioned plurality of digital images and the aforementioned diagnostic test information, wherein the machine learning system includes a multi-binary label machine learning system. Determining at least one threshold of one or more binary outputs of the multi-binary label machine learning system, The non-temporary computer-readable medium according to item 17, comprising outputting a set of parameters from the multi-binary label machine learning system to a digital storage device, wherein the set of parameters includes the at least one threshold.

[0010] The attached drawings are incorporated into and constitute part of this specification, illustrating various exemplary embodiments, and together with this description, serve to illustrate the principles of the disclosed embodiments. [Brief explanation of the drawing]

[0011] [Figure 1A] An exemplary block diagram of a system and network for identifying diagnostic tests applicable to pathological specimens, according to exemplary embodiments of the present disclosure, is shown.

[0012] [Figure 1B] An exemplary block diagram of the treatment analysis platform 100, as illustrated in this disclosure, is shown.

[0013] [Figure 2A] This flowchart shows an exemplary method for identifying diagnostic tests to be applied to a pathological specimen, according to exemplary embodiments of the present disclosure.

[0014] [Figure 2B] This flowchart shows an exemplary method for training a machine learning system to identify relevant diagnostic tests, according to exemplary embodiments of the present disclosure.

[0015] [Figure 2C] This flowchart shows an exemplary method for training a machine learning system according to exemplary embodiments of the present disclosure.

[0016] [Figure 2D]This flowchart shows an exemplary method, according to an exemplary embodiment of the present disclosure, for identifying applicable tests on a pathological specimen using a trained system.

[0017] [Figure 3] This is an exemplary workflow for determining the applicability of an inspection, according to an exemplary embodiment of the present disclosure.

[0018] [Figure 4] This specification shows an exemplary system capable of performing the techniques presented herein. [Modes for carrying out the invention]

[0019] The following will be a detailed description of exemplary embodiments of the present disclosure, examples of which are shown in the accompanying drawings. Where possible, the same reference numerals will be used throughout the drawings to indicate the same or similar parts.

[0020] The systems, devices, and methods disclosed herein are described in detail by example and with reference to the drawings. The examples described herein are merely illustrative and are presented to facilitate the description of the apparatus, devices, systems, and methods described herein. Functions or components shown in the drawings or described below should not be considered essential to any particular embodiment of these devices, systems, or methods unless specifically designated as essential.

[0021] Furthermore, with respect to any of the methods described, regardless of whether the method is described in relation to a flowchart, unless otherwise specified or required by the context, any explicit or implicit ordering of the steps performed in the execution of the method does not imply that those steps must be performed in the order presented, but rather that they may be performed in a different order or in parallel.

[0022] As used herein, the term “exemplary” means “example” and not “ideal.” Furthermore, the terms “a” and “an” as used herein do not imply a limit on quantity, but rather that one or more of the items mentioned exist.

[0023] Computational assays using machine learning may be able to directly determine the results of diagnostic tests, or they may be used to exclude or prioritize tests that are not likely to be of value, and / or to help prioritize among available tests. One or more embodiments of this disclosure implement this functionality by ranking tests that are not excluded based on supplementary information such as availability and cost.

[0024] While existing computer-aided assays focus on identifying the presence or absence of disease / biomarkers, the methods presented herein may also include identifying diagnostic tests that could be therapeutically useful, as well as tests that are less likely to be beneficial to clinicians.

[0025] Figure 1A shows an exemplary block diagram of a system and network for identifying diagnostic tests applicable to pathological specimens, according to an exemplary embodiment of the present disclosure.

[0026] Specifically, Figure 1A shows an electronic network 120 that may be connected to servers in hospitals, laboratories, and / or clinics. For example, a physician server 121, a hospital server 122, a clinical trial server 123, a laboratory server 124, and / or a laboratory information system 125 may each be connected to the electronic network 120, such as the Internet, via one or more computers, servers, and / or handheld mobile devices. According to exemplary embodiments of the present application, the electronic network 120 may also be connected to a server system 110, which may include a processing device configured to implement a therapeutic analysis platform 100, which includes a slide analysis tool 101 for using machine learning to determine the characteristics of a specimen or image characteristic information related to a digital pathology image(s), and to determine whether a disease or infectious agent is present, according to exemplary embodiments of the present disclosure. The slide analysis tool 101 may also predict appropriate diagnostic tests for the pathology specimen.

[0027] The physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125 may create or acquire digital images of one or more patient cytological specimens, histopathological specimens, cytological specimen slides, histopathological specimen slides, or any combination thereof. The physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125 may also acquire any combination of patient-specific information, such as age, medical history, cancer treatment history, family history, and past biopsy or cytological information. The physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125 may transmit the digitized slide images and / or patient-specific information to the server system 110 over the electronic network 120. The server system 110 may include one or more storage devices 109 for storing images and data received from at least one of the physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125. The server system 110 may also include processing devices for processing the images and data stored in the storage devices 109. The server system 110 may further include one or more machine learning tools or functions. For example, according to one embodiment, the processing device may include machine learning tools for the therapeutic analysis platform 100. Alternatively or additionally, the present disclosure (or parts of the systems and methods of the present disclosure) may be run on a local processing device (e.g., a laptop).

[0028] The physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory system 125 refer to systems used by pathologists to review slide images. In a hospital setting, tissue type information may be stored in the laboratory information system 125.

[0029] Figure 1B shows an exemplary block diagram of a therapeutic analysis platform 100 for determining image characteristic information related to specimen characteristics or digital pathology images (or multiple images) using machine learning. The therapeutic analysis platform 100 may include a slide analysis tool 101, a data acquisition tool 102, a slide acquisition tool 103, a slide scanner 104, a slide manager 105, storage 106, a laboratory information system 107, and a display application tool 108.

[0030] As described below, the slide analysis tool 101 refers to a process and system for determining diagnostic information relating to digital pathology images(s). According to exemplary embodiments, machine learning may be used to classify the images. The slide analysis tool 101 may also receive additional information relating to the pathology specimen, as described in the embodiments below.

[0031] According to exemplary embodiments, the data acquisition tool 102 can facilitate the transfer of digital pathology images to various tools, modules, components, and devices used for classifying and processing digital pathology images.

[0032] According to an exemplary embodiment, the slide acquisition tool 103 may scan pathological images and convert them into a digital format. The slides may also be scanned with a slide scanner 104, and the slide manager 105 may process the images on the slides into digitized pathological images and store the digitized images in storage 106.

[0033] According to exemplary embodiments, the display application tool 108 may provide the user with specimen or image characteristic information relating to digital pathology images(s). This information may be provided through various output interfaces (e.g., screens, monitors, storage devices, and / or web browsers).

[0034] The slide analysis tool 101 and one or more of its components may transmit and / or receive digitized slide images and / or patient information via the network 120 to the server system 110, physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125. Furthermore, the server system 110 may include a storage device for storing images and data received from at least one of the slide analysis tool 101, data acquisition tool 102, slide acquisition tool 103, slide scanner 104, slide manager 105, and display application tool 108. The server system 110 may also include a processing device for processing images and data stored in the storage device. The server system 110 may further include one or more machine learning tools or functions, for example, due to the processing device. Alternatively or additionally, the present disclosure (or parts of the systems and methods of the present disclosure) may be run on a local processing device (e.g., a laptop).

[0035] Any of the above devices, tools, and modules may be placed on a device that can connect to an electronic network, such as the Internet or a cloud service provider, via one or more computers, servers, and / or handheld mobile devices.

[0036] Figure 2A illustrates a method for identifying a series of diagnostic tests on a pathological specimen according to an exemplary embodiment of the present disclosure. For example, exemplary method 200 (e.g., steps 202-210) may be performed automatically by the slide analysis tool 101 or at the request of a user.

[0037] According to one embodiment, an exemplary method 200 for identifying a series of diagnostic tests to be applied to a pathological specimen may include one or more of the following steps: In step 202, the method may include receiving one or more digital images related to the pathological specimen (e.g., histology, cytology, etc.) into a digital storage device (e.g., a hard drive, network drive, cloud storage, RAM, etc.).

[0038] Optionally, this method may include receiving additional patient and / or disease-related information associated with the pathological specimen. This additional information may include, but is not limited to, patient demographics, previous medical history, additional clinicopathological and / or biochemical test results, radiographic images, past pathological specimen images, tumor size, cancer grade, cancer stage, and information about the specimen (e.g., location of the specimen sample, position within a block, etc.), which may be stored in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0039] Optionally, this method may include receiving additional testing information. This additional testing information may include, but is not limited to, the availability of the test at local (nearby) healthcare facilities, testing supplies, current clinical guidelines for the test, current regulatory compliance for the test, the average time to obtain results for one or more tests (test speed and turnaround time), current test prices, and available clinical trials, and may be stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0040] Optionally, this method may also include receiving additional test preference information. This additional preference information may include information stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.) such as whether the test is covered by insurance (government healthcare, patient insurance, etc.), the out-of-pocket cost after insurance considerations, which tests the doctor prioritizes (laboratory, hospital), and which tests the patient prefers (e.g., due to religious practices, the patient's age, underlying medical conditions, side effects, etc.).

[0041] In step 204, this method may include determining multiple diagnostic tests.

[0042] In step 206, this method may include applying a machine learning system to one or more digital images to identify any of the prerequisites for each of several diagnostic tests to be applicable, the machine learning system being trained by processing multiple training images. Diagnostic tests include molecular histology (genomic sequencing, immunohistochemistry (IHC), fluorescence in situ hybridization (FISH), colorimetric in situ hybridization (CISH), in situ hybridization (ISH), genetic testing, special stains, algorithmic (computational, artificial intelligence, machine learning) tests, radiography, additional biopsies (specimens), clinical tests (including biochemical and / or chemical pathology tests such as blood, urine, sputum), etc., output to digital storage devices (e.g., hard drives, electronic medical records, laboratory information systems, network drives, etc.) and / or user displays (e.g., monitors, documents, printed materials, etc.).

[0043] In step 208, the method may include using a machine learning model to identify applicable diagnostic tests from among several diagnostic tests based on one or more digital images and preconditions. The scoring of the diagnostic tests may represent several expressions of desirability. Examples may include the benefit of the test for the expected patient, cost-effectiveness, efficiency of the test results for the benefit, and ranking of preferred tests for the benefit and / or availability of the therapeutic agent or approach in the dosing and dosing schedule of the proposed treatment.

[0044] In step 210, this method may include outputting a ranked set of diagnostic tests to a digital storage device and / or display.

[0045] Optionally, this method involves entering a scoring threshold and may output one or more tests that scored above the threshold, or only those tests (tests with a score of zero above the threshold will not be included).

[0046] Optionally, the method may include outputting one or more possible treatments, medications, or medication schedules as a treatment strategy for the patient, or clinical tests available to the patient based on study inclusion and exclusion criteria, and geographical proximity, based on input information and / or additional suggested tests.

[0047] Optionally, this method may include displaying a ranked set of diagnostic tests to the user (e.g., a contracting clinician, testing laboratory, diagnostic company, treatment company, and / or patient). The test results may also be displayed using a customized interface, output documents (such as PDF), printouts, etc.

[0048] One or more exemplary embodiments may include one or more of the following three components: Training a machine learning system to identify the applicability of a test. Identify applicable tests using a trained system. Ranking of applicable tests based on supplementary information Training a machine learning system to identify the applicability of a test.

[0049] Figure 2B is a flowchart illustrating an exemplary method for training a machine learning system to identify the applicability of an inspection using the techniques presented herein. For example, exemplary methods 220 and 240 (e.g., steps 222-224 and steps 242-252) may be performed automatically by the slide analysis tool 101 or at the request of a user.

[0050] According to one embodiment, an exemplary method 220 for training a machine learning system to identify the applicability of a test may include one or more of the following steps: In step 222, the method may include identifying at least prerequisites for a diagnostic test to be applicable. For example, some breast cancer recurrence tests (such as Oncotype DX) may require that a breast cancer patient be estrogen receptor (ER) positive for the test to be applicable. If computational analysis identifies that a patient is unlikely to be ER positive, then using Oncotype DX on that patient is ruled out.

[0051] Step 224 may involve using a machine learning system to predict the negative predictive value for one or more diagnostic tests. For example, because genomic testing can be costly and time-consuming, determining that a patient does not have a mutation associated with the administration of a particular drug may indicate that performing genomic testing would not add value. If the system cannot rule out the presence of the mutation, then genomic testing to examine for the presence of that mutation may be a useful test. Another example is when immunohistochemical and / or genomic testing may be necessary in a population (e.g., assessment of NTRK fusion genes or microsatellite instability in patients with metastatic cancer) but the prevalence of the biomarker in the population is low. If the system cannot rule out the presence of an immunohistochemical and / or genomic feature, then immunohistochemical and / or genomic testing may be performed.

[0052] Method 240 is a flowchart for training a machine learning system according to an exemplary embodiment. For example, exemplary Method 240 (e.g., steps 242-252) may be performed automatically by the slide analysis tool 101 or on request from a user. In step 242, this method may include receiving one or more digital images from a patient relating to a pathological specimen (e.g., histology, cytology, etc.), the one or more digital images being paired with information about the results and / or value of one or more diagnostic tests performed or tests to comply with the applicability of the diagnostic tests, which are placed in a digital storage device (e.g., a hard drive, network drive, cloud storage, RAM, etc.).

[0053] In step 244, the method may include receiving additional patient and / or disease information associated with one or more digital images. This additional information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block) received on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0054] In step 246, the method may include filtering one or more digital images to identify the tissue region to be analyzed and removing non-prominent areas from one or more digital images, where non-prominent areas are, for example, background and / or not identified as the tissue region of interest. The region(s) of interest may be identified at least in part on additional information relating to the patient and / or disease. The determination of the region of interest / prominent area may be performed using the techniques discussed in U.S. Patent Application Publication No. 17 / 313617, which is incorporated herein by reference. Filtering one or more images may be performed by identifying prominent areas (e.g., invasive tumor and / or invasive tumor stroma) by manual annotation or by using a region detector.

[0055] In step 248, this method may include training a multi-binary machine learning system to predict whether one or more diagnostic tests are performed and whether one or more diagnostic tests are applicable. If no tests are performed, it is treated as missing data for the patient and is not used to update the parameters of the machine learning system. Where available, additional patient data (medical history, existing outcomes, etc.) may be input into the machine learning system to provide additional information (for example, this can be done in a neural network-based way by converting this information into vectors and then adjusting the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels of each patient sample, including but not limited to: a. Multilayer perceptron (MLP) b. Convolutional Neural Networks (CNNs) c. Graph Neural Networks d. Support Vector Machine (SVM) e. Random Forest

[0056] In step 250, the method may include setting at least one threshold for one or more binary outputs of the machine learning system. For outputs corresponding to prerequisites for a diagnostic test, at least one threshold may be set to optimize the detection of that prerequisite (e.g., the presence of a biomarker that makes the diagnostic test applicable). For outputs corresponding to individual tests, thresholds may be set to optimize the NPV to exclude the applicability of that diagnostic test.

[0057] In step 252, the method may include outputting a set of parameters from a multi-binary level machine learning system to a digital storage device (e.g., a hard drive, network drive, cloud storage, RAM, etc.). The set of parameters may include at least one threshold and other data to tune the machine learning system. Identify applicable tests using a trained system.

[0058] Figure 2C is a flowchart for using a trained machine learning system on a patient according to an exemplary method disclosed herein. After the machine learning system is trained to determine applicable diagnostic tests, the user can apply the system to the patient. For example, exemplary method 260 (e.g., steps 262-270) may be performed automatically by the slide analysis tool 101 or on request from the user. In step 262, the method may include receiving one or more digital images related to a pathological specimen (e.g., histology, cytology, IHC, etc.) into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0059] In step 264, the method may include receiving additional information relating to a patient and / or disease associated with one or more digital images. This additional information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block) which are placed in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0060] In step 266, the method may include filtering one or more images to identify the tissue region of interest and removing inapplicable regions from one or more images. Filtering may be performed by identifying prominent regions (e.g., invasive tumor and / or invasive tumor stroma) by manual annotation or by using a region detector.

[0061] In step 268, this method may include predicting the applicability of one or more diagnostic tests by applying a trained machine learning system to one or more digital images.

[0062] In step 270, the method may include outputting the predicted applicability of one or more diagnostic tests to a digital storage device (e.g., a hard drive, network drive, cloud storage, RAM, etc.).

[0063] Ranking of applicable tests based on supplementary information Figure 2D is a flowchart illustrating an exemplary method for ranking diagnostic tests applicable to a pathological specimen using the techniques presented herein. After identifying applicable tests, an optional step is to rank the applicable tests based on patient and clinician priorities, test availability, test cost, test speed, etc. For example, exemplary method 280 (e.g., steps 282-290) may be performed automatically by the slide analysis tool 101 or at the request of the user. In step 282, this method may include applying a trained machine learning system to identify a list of one or more diagnostic tests applicable to a pathological specimen, thereby generating an N-dimensional binary vector "y", in which case one or more elements correspond to the applicability of individual tests.

[0064] Step 284 may involve receiving additional testing and priority information regarding the pathological specimen. Additional testing information may include, but is not limited to, the availability of tests at local (nearby) healthcare facilities, testing supplies, current clinical guidelines for the tests, current regulatory compliance for the tests, average time to obtain results for one or more tests (test speed), and current test prices, and may be stored in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). Additional priority information may include, but is not limited to, information on which tests are covered by insurance (government healthcare, patient insurance, etc.), out-of-pocket expenses after considering insurance, tests preferred by the physician (laboratory, hospital), and tests preferred by the patient (e.g., due to religious practices, patient age, underlying medical conditions, side effects, etc.), and may be stored in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0065] In step 286, this method may involve scoring one or more tests to generate an N-dimensional vector "s" of the scores. There are many non-restrictive ways in which this can be done. a. Use only Applicability and Availability: i. Set s = y. For some or all tests that are expected to be applicable, if a test is unavailable, set the corresponding element of s for that test to 0. b. Use of applicability, availability, and speed: i. Set s=y. For any or all tests that are expected to be applicable, if a test is unavailable, set the corresponding element of s for that test to 0; otherwise, set the corresponding element of s to be inversely proportional to speed, so that faster tests receive higher scores. c. Use of applicability, availability, speed, and patient co-payment: i. Set s=y. For any or all tests that are expected to be applicable, if a test is unavailable, set the corresponding element s for that test to 0. Otherwise, based on user preference, set the corresponding element of s to a weighted sum, where the first term of the sum is inversely proportional to speed, so faster tests will have higher scores, and the second term of the sum is inversely proportional to the patient's test cost minus insurance coverage. d. Use applicability, availability, speed, patient co-payment, and patient priority: i. Set s=y. For any or all tests predicted to be applicable, if a test is unavailable or the patient cannot take the test (e.g., due to religious customs, age, discomfort, etc.), set the corresponding element s for that test to 0. Otherwise, based on user preference, set the corresponding element of s to a weighted sum, where the first term of the sum is inversely proportional to speed, so faster tests will have higher scores, and the second term of the sum is inversely proportional to the patient's test cost minus insurance coverage.

[0066] In step 288, this method may involve sorting an N-dimensional vector s such that tests with higher scores are preferred, which may involve sorting the tests in the vector by their test scores.

[0067] Optionally, this method may involve inputting a scoring threshold and outputting one or more tests, or possibly only those tests, that scored above the threshold (tests with a score of zero above the threshold will not be included).

[0068] Optionally, this method may also include outputting one or more treatments that may be suitable for the patient based on the input information in steps 282-288 and / or additional suggested tests.

[0069] In step 290, the method may include displaying the test results to a user (e.g., a contracted clinician, a laboratory, a diagnostic company, a drug company, and / or a patient) using a customized interface, output documents (e.g., PDF), printed materials, etc.

[0070] Figure 3 is an exemplary workflow 300 for determining the applicability of a test using the technique presented herein. Figure 3 is a depiction of a system run on image data from a patient (before ranking) to determine the applicability of N different diagnostic tests, the system outputting 1 if the test is applicable and 0 if it is not.

[0071] In step 302, the workflow may include inputting digital images of the pathology specimen. In step 304, the pathology specimen and any additional available patient data may be input into the machine learning system.

[0072] In step 306, the workflow may include a multi-label output that determines the applicability of each diagnostic test.

[0073] Exemplary embodiment: Directing a patient to undergo genomic testing, IHC testing, or ISH / FISH testing even if the patient is unlikely to have a specific mutation or antigen before testing. Genomic testing can be costly, may not be available at all centers, may incur additional costs, and can be quite time-consuming. The techniques presented herein may be used to determine when genomic testing is likely to be of diagnostic value, thereby avoiding unnecessary testing. One or more exemplary embodiments may be used to determine when IHC, ISH / FISH testing is applicable.

[0074] Training machine learning systems to identify the applicability of genome, IHC, or ISH / FISH testing. The steps involved in training a machine learning system may include the following: 1. Receive one or more digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images may be combined with information regarding the results of genomic testing (e.g., presence or absence of oncogenic mutations / fusions for a list of genes), IHC testing, and / or ISH / FISH testing. 2. Optionally, receive additional patient information for each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block) and may be placed in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove inconspicuous regions from one or more images. 4. Train a multi-binary label machine learning system to predict the presence of one or more oncogenic gene mutations / fusions. Where available, additional patient data (medical history, existing outcomes, etc.) may be input into the machine learning system to provide additional information (for example, this may be done in a neural network-based way by converting this information into vectors and then adjusting the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels of samples from one or more patients, including but not limited to: i. Multilayer perceptron (MLP) ii. Convolutional Neural Networks (CNNs) iii. Graph Neural Networks iv. Support Vector Machines (SVMs) v. Random Forest 5. Thresholds may be set for one or more binary outputs of the system to optimize it so that mutations / fusions of each oncogene are not present. 6. Output the parameters of the trained system to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). A trained system is used to identify whether genomic, IHC, or ISH / FISH testing may be necessary. 1. Receive digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.) and may be stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove inconspicuous regions from one or more images. 4. If available, incorporate additional patient information by running a trained machine learning system on digital images from the patient to generate an N-dimensional vector of multi-labeled outputs corresponding to the definitive absence of mutations / fusions in each oncogene. 5. Output the prediction to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 6. Optionally, users will be notified of any excluded oncogenes and recommended to undergo genomic testing if necessary.

[0075] Exemplary embodiment: Directing multi-parameter gene expression testing for breast cancer, such as MammaPrint, OncotypeDX, EndoPredict, PAM50 (Prosigna), or Breast Cancer Index. The use of multi-parameter gene expression tests is increasing as a guide for determining treatment options for breast cancer. These tests identify patients at high risk of breast cancer recurrence. Tests used include MammaPrint, a 70-gene assay, and Oncotype DX, a 20-gene assay, which help guide treatment decisions regarding whether chemotherapy may benefit patients with invasive breast cancer. A prerequisite for the Oncotype DX test may be that the patient is ER-positive, so ER-negative patients may need to be excluded. Other tests to determine whether a patient may require chemotherapy include EndoPredict (a 12-gene risk score), PAM50 (a 50-gene assay), and the Breast Cancer Index.

[0076] Training a machine learning system to identify the applicability of multi-parameter gene expression testing to breast cancer patients. The steps involved in training a machine learning system may include the following: 1. Receive multiple digital images of invasive primary breast tumors from pathological specimens (e.g., histology) from the patient to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). For each patient, information on whether the patient is ER-positive or negative may be combined with one or more images, and if positive, the Oncotype DX score may also be included. 2. Optionally, receive additional patient information for each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block) and may be placed in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove inconspicuous regions from one or more images. 4. Train a multi-binary label machine learning system to predict whether a patient is ER-positive or ER-negative, and to predict the Oncotype DX score for ER-positive patients, otherwise treating the Oncotype DX score as missing (e.g., not used to update parameters if missing). For other tests, train a multi-label machine learning system to predict a patient's cancer recurrence risk score. Where available, additional patient data (medical history, existing outcomes, etc.) may be input into the machine learning system to provide additional information (e.g., this may be done in a neural network-based way by using vectorization of this information and then adjusting image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels of samples from each patient, including but not limited to: i. Multilayer perceptron (MLP) ii. Convolutional Neural Networks (CNNs) iii. Graph Neural Networks iv. Support Vector Machines (SVMs) v. Random Forest Thresholds can be set for one or more binary outputs of the system, indicating that Oncotype DX is not applicable if the system determines the patient is ER-negative, and that if the patient is determined to have a very low test score, it is likely that performing a multi-parameter breast cancer gene expression test would lead to a prediction of a low recurrence risk. The parameters of the trained system are output to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0077] Use of trained systems After training the system to determine its applicability for multi-parameter breast cancer gene expression testing, the steps for using the trained system on patients may include the following: 1. Receive digital images of invasive primary breast tumors from the patient's pathological specimens (e.g., histology) on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.) and may be stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove any unsuitable regions from one or more images. 4. Run a trained machine learning system on digital images from the patient, incorporating additional patient information where available. If the system predicts the patient is ER-negative, indicate that Oncotype DX is not recommended. If the system predicts the patient is likely to have a low score on a multi-parameter breast cancer gene expression test, indicate this to the user and recommend not using this test. 5. Output the prediction to a digital storage device (hard drive, network drive, cloud storage, RAM, etc.).

[0078] Exemplary Embodiment: Instructions for multi-parameter gene expression tests for prostate cancer, such as Oncotype DX Genome Prostate Score (GPS) or Prolaris. The OncotypeDX GPS (17-gene assay) and Prolaris (46-gene assay) tests help assess the potential invasiveness of prostate cancer and guide treatment decisions. A higher GPS score or Prolaris risk score indicates a higher likelihood of invasive cancer, which may require immediate treatment such as surgery or radiation therapy.

[0079] The steps involved in training a machine learning system may include the following: 1. Receive multiple digital images of prostate tumors from pathological specimens (e.g., histology) from the patient to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images may be combined with gene expression testing for prostate cancer. 2. Optionally, receive additional patient information for each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.), which are stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions. Remove inconspicuous regions from one or more images. 4. Train a multi-binary label machine learning system to predict OncoType DX GPS score / Prolaris score. If available, additional patient data (medical history, existing outcomes, etc.) may be input into the machine learning system to provide additional information (for example, this may be done in a neural network-based way by converting this information into vectors and then adjusting the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels of samples from each patient, including but not limited to: i. Multilayer perceptron (MLP) ii. Convolutional Neural Networks (CNNs) iii. Graph Neural Networks iv. Support Vector Machines (SVMs) v. Random Forest 5. A threshold may be set for one or more binary outputs of the system, indicating that if a patient is judged to have a very low test score, performing a multi-parameter prostate cancer gene expression test is likely to lead to a less invasive prediction of prostate cancer. 6. Output the parameters of the trained system to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0080] Use of trained systems After the system has been trained to determine the applicability of Oncotype DX, the steps for using the trained system on a patient may include the following: 1. Receive digital images of invasive primary breast tumors from the patient's pathological specimens (e.g., histology) on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.) and may be stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions. Remove any unsuitable regions from one or more images. 4. Run a trained machine learning system on digital images from patients, incorporating additional patient information where available. If the system predicts that a patient is likely to have a low Oncotype DX GPS score or Prolaris score, inform the user and recommend not using Oncotype DX GPS or Prolaris. 5. Output the prediction to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0081] Exemplary embodiments: Direct single / multiple immunohistochemistry (IHC) and fluorescence in situ hybridization (FISH) tests for HER2, mismatch repair (MMR) repair proteins, PD-L1, etc. In the treatment of certain types of cancer at specific clinical stages, additional IHC and / or FISH analysis may be essential for treatment decision-making, although the frequency of markers is low. This is exemplified by the need for tumor site-independent testing in all or multiple metastatic cancer patients for the presence of NTRK1, NTRK2, and NTRK3 fusion genes, as well as microsatellite instability for the use of specific treatment regimens (i.e., TRK inhibitors and immune checkpoint inhibitors, respectively). Similarly, testing for the presence of ALK, RET, and ROS1 rearrangements in non-small cell lung cancer patients may be necessary for the treatment of these patients in the metastatic context. Training a machine learning system to identify the applicability of single / multiple immunohistochemistry (IHC) and fluorescence in situ hybridization (FISH) tests.

[0082] The steps involved in training a machine learning system may include the following: 1. Receive multiple digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images may be combined with information regarding the results of IHC / FISH testing or related genomic testing. 2. Optionally, receive additional patient information for each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.), which are stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor, invasive tumor stroma). Remove inconspicuous regions from one or more images. 4. Train a multi-binary label machine learning system to predict the presence of IHC / FISH markers. If available, additional patient data (medical history, existing outcomes, etc.) may be input into the machine learning system to provide additional information (for example, this may be done in a neural network-based way by converting this information into vectors and then adjusting the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels of each patient sample, including but not limited to: i. Multilayer perceptron (MLP) ii. Convolutional Neural Networks (CNNs) iii. Graph Neural Networks iv. Support Vector Machines (SVMs) v. Random Forest 5. Thresholds may be set for one or more binary outputs of the system to optimize the system so that certain IHC / FISH markers are not present. 6. Output the parameters of the trained system to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0083] Use of trained systems After the system has been trained to determine the applicability of single / multiple immunohistochemistry (IHC) tests, the steps for using the trained system on a patient may include the following: 1. Receive digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.) and may be stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove inconspicuous regions from one or more images. 4. If available, run a trained machine learning system on one or more digital images from the patient to generate an N-dimensional vector of multi-labeled outputs corresponding to the definitive absence of specific IHC / FISH markers, and incorporate additional patient information. 5. Output the prediction to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 6. Optionally, notify the user of any excluded IHC / FISH markers and recommend performing IHC / FISH of that type if necessary.

[0084] Exemplary Embodiment: Instructions for a multi-gene sequence analysis panel such as Foundation One CDx or MSK IMPACT Multi-gene panel analysis of tumor and / or tumor-normal pairs has been shown to benefit cancer patients, and multiple studies have shown that in up to >10% of patients with metastatic cancer, multi-gene sequencing assays could allow them to receive more appropriate treatment and / or be enrolled in clinical trials based solely on the results of these molecular tests. However, for the vast majority of patients, the information provided by these assays has limited or no current utility. Furthermore, these assays are relatively expensive, time-consuming, and available at only a limited number of institutions.

[0085] Training a machine learning system to identify the applicability of multiple gene sequencing panels The steps involved in training a machine learning system may include the following: 1. Receive multiple digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images may be combined with information regarding the results of multiple gene sequencing assays. 2. Optionally, receive additional patient information for each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block) and may be placed in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove inconspicuous regions from one or more images. 4. Train a multi-binary label machine learning system to predict the results of a multi-gene sequencing assay. Where available, additional patient data (medical history, existing results, etc.) may be input into the machine learning system to provide additional information (for example, this may be done in a neural network-based way by converting this information into vectors and then adjusting the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels of each patient sample, including but not limited to: i. Multilayer perceptron (MLP) ii. Convolutional Neural Networks (CNNs) iii. Graph Neural Networks iv. Support Vector Machines (SVMs) v. Random Forest 5. A threshold may be set for one or more binary outputs of the system and optimized so that there are no clinically relevant findings resulting from the multi-gene sequencing assay. 6. Output the parameters of the trained system to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0086] Use of trained systems After the system has been trained to determine the applicability of multiple gene sequencing panels, the steps for using the trained system on a patient may include the following: 1. Receive digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.) and may be stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove inconspicuous regions from one or more images.

[0087] If available, incorporate additional patient information by running a trained machine learning system on digital images from patients to generate an N-dimensional vector of multi-labeled outputs that address critical missing clinically relevant results from multi-gene sequencing assays.

[0088] Output the prediction to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0089] The system optionally notifies the user of which genes and genomic changes were excluded and recommends performing a multi-gene sequencing assay if necessary.

[0090] Exemplary embodiment: Instructions for assays to prioritize immuno-oncology (IO) therapy. Immunotherapy is reshaping the treatment landscape for patients with various types of cancer. Tumor-specific biomarkers for treatment decision-making (e.g., PD-L1 evaluation in non-small cell lung cancer and metastatic triple-negative breast cancer), as well as site-agnostic biomarkers (e.g., microsatellite instability (MSI) or mismatch repair deficiency (dMMR) and tumor mutational burden (TMB)), are sometimes needed. However, their evaluation often involves expensive, lengthy turnaround times and requires multimodal assays (e.g., IHC, PCR, and / or multi-gene sequencing assays) that necessitate subsequent integration.

[0091] Furthermore, new panels are being developed to better understand the composition of the tumor microenvironment and the immune characteristics of patients. The PanCancer IO 360 gene expression panel is a 770-target multi-gene expression panel developed to characterize expression patterns from tumors, the immune system, and the stroma. It includes a tumor inflammation signature (TIS) containing 18 functional genes known to be associated with responses to PD-1 / PD-L1 inhibitor pathway blockade. The PanCancerIO360 panel and TIS may be helpful for physicians in determining the treatment of IO therapy.

[0092] Training a machine learning system to identify assays that can help prioritize immuno-oncology therapies. The steps involved in training a machine learning system may include the following: 4. Receive multiple digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images may be combined with information on specific biomarkers related to the response to immunotherapy (e.g., PD-L1 expression, microsatellite instability high / deficit mismatch repair (MSI / dMMR), tumor mutational burden (TMB), PanCancer IO360 panel, TIS). 5. Optionally, receive additional patient information for each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.), which are stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 6. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove inconspicuous regions from one or more images. 7. Train a multi-binary label machine learning system to predict the presence of one or more specific biomarkers for immunotherapy response (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancerIO360 panel, TIS). Where available, additional patient data (medical history, existing outcomes, etc.) may be input into the machine learning system to provide additional information (e.g., this may be done in a neural network-based manner by converting this information into vectors and then adjusting image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels of each patient sample, including but not limited to: i. Multilayer perceptron (MLP) ii. Convolutional Neural Networks (CNNs) iii. Graph Neural Networks iv. Support Vector Machines (SVMs) v. Random Forest 8. Thresholds may be set for one or more binary outputs of the system to optimize the reliable absence of mutations / fusions of one or more specific biomarkers for immunotherapy response (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancerIO360 panel, TIS). 9. Output the parameters of the trained system to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0093] Use of trained systems to identify assays that help prioritize immuno-oncology therapies. Steps involving the use of a trained machine learning system may include the following: 1. Receive digital images of pathological specimens (e.g., histology, cytology, etc.) from the patient into a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, previous medical history, additional test results, radiographic images, past pathological specimen images, and information about specimens (e.g., location of specimen sample, position within a block, etc.), which are stored on a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 3. Optionally, filter one or more images to identify the target tissue regions that need to be used. This may be done by manually annotating or by using a region detector to identify prominent regions (e.g., invasive tumor and / or invasive tumor stroma). Remove unrelevant, inconspicuous regions from one or more images. 4. If a trained machine learning system can be run on digital images from patients to generate an N-dimensional vector of multi-labeled output corresponding to the definitive absence of mutations / fusions of one or more specific biomarkers of immunotherapy response (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancerIO360 panel, TIS, etc.), incorporate additional patient information. 5. Output the prediction to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.). 6. Optionally, notify the user if specific biomarkers for immunotherapy response (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancer IO 360 panel, TIS) have been ruled out, and recommend that IHC and / or genomic testing be performed if necessary.

[0094] As shown in Figure 4, device 400 may include a central processing device (CPU) 420. The CPU 420 can be any type of processor device, including, for example, any type of dedicated or general-purpose microprocessor device. As will be recognized by those skilled in the art, the CPU 420 can be a single processor in a multicore / multiprocessor system, operating alone or in a cluster of computing devices, or in a system operating in a cluster or within a server farm. The CPU 420 may be connected to a communication infrastructure 410, such as a bus, message queue, network, or multicore message passing scheme.

[0095] Device 400 may also include main memory 440, for example, random access memory (RAM), and may also include secondary memory 430. The secondary memory 430, for example, read-only memory (ROM), may be, for example, a hard disk drive or a removable storage drive. Such removable storage drives may include, for example, floppy disk drives, magnetic tape drives, optical disk drives, flash memory, or the like. In this example, the removable storage drive is read from and / or written to by a removable storage unit in a well-known manner. Removable storage may include floppy disks, magnetic tapes, optical disks, etc., which are read from and written to by the removable storage drive. As will be recognized by those skilled in the art, a removable storage unit generally includes a computer-usable storage medium therein which computer software and / or data are stored.

[0096] In an alternative embodiment, the secondary memory 430 may include similar means that enable computer programs of other instructions to be loaded into device 400. Examples of such means may include program cartridges and cartridge interfaces (such as those found in video game machines), removable memory chips (such as EPROMs or PROMs) and associated sockets, and other removable storage units and interfaces that enable the transfer of software and data from the removable storage units to device 400.

[0097] Device 400 may also include a communication interface ("COM") 460. The communication interface 460 enables software and data to be transferred between Device 400 and external devices. The communication interface 460 may include a modem, a network interface (such as an Ethernet® card), a COM port, a PCMCIA slot and card, or similar. The software and data transferred via the communication interface 460 may be in the form of signals, which may be electrical, electromagnetic, optical, or other signals that can be received by the communication interface 460. These signals that may be provided to the communication interface 460 via the communication path of Device 400 may be implemented using, for example, wires or cables, optical fibers, telephone lines, cell phone links, RF links, or other communication channels.

[0098] The hardware elements, operating systems, and programming languages ​​of such devices are essentially conventional and should be well-known to those skilled in the art. Device 400 may also include input and output ports 450 for connecting to input and output devices such as keyboards, mice, touchscreens, monitors, and displays. Of course, to distribute the processing load, various server functions may be implemented in a distributed manner across many similar platforms. Alternatively, the server may be implemented by appropriate programming on a single computer hardware platform.

[0099] Throughout this disclosure, references to components or modules generally refer to items that can be logically grouped together to perform a function or a group of related functions. Similar reference numbers are generally intended to refer to identical or similar components. Components and / or modules may be implemented in software, hardware, or a combination of software and / or hardware.

[0100] The tools, modules, and / or functions described above may be executed by one or more processors. "Storage media" of this kind may include any or all of the following: tangible memory such as computers and processors, or related modules such as various semiconductor memories, tape drives, and disk drives, which may provide non-temporary storage for software programming at any time.

[0101] Software can be transmitted via the Internet, cloud service providers, or other telecommunications networks. For example, communication may allow software to be loaded from one computer or processor to another. As used herein, unless limited to non-temporary, tangible “storage” media, terms such as computer or machine “readable media” refer to any medium involved in giving instructions to a processor for execution.

[0102] The general descriptions set forth herein are illustrative and descriptive only and do not limit the present disclosure. Other embodiments may become apparent to those skilled in the art from consideration herein and the practice of the invention disclosed herein. Specifications and examples are intended to be considered merely illustrative.

Claims

1. A computer implementation method for processing subject data to identify diagnostic tests, Receiving multiple subject data, including one or more digital pathology images, The machine learning system is applied to the aforementioned plurality of subject data to identify any of the prerequisites for each of the plurality of diagnostic tests, and to determine the applicability of the diagnostic test based on the prerequisites, wherein the machine learning system is trained by processing training subject data paired with diagnostic test information that includes at least one of the prerequisite results and the test applicability results. Using the machine learning system, identify applicable diagnostic tests from among the multiple diagnostic tests based on the multiple subject data and the preconditions, A computer implementation method comprising outputting the applicable diagnostic test to a digital storage device and / or display.

2. The computer implementation method according to claim 1, wherein the plurality of subject data further includes one or more of additional subject information relating to the subject and / or disease and additional diagnostic test information.

3. Filtering one or more of the aforementioned digital pathology images to identify the tissue region to be analyzed, The computer implementation method according to claim 1, further comprising removing one or more regions from the one or more digital pathological images that are not identified as the target tissue region.

4. The machine learning system is provided with a scoring threshold, Based on the scoring threshold, determine one or more applicable diagnostic tests that exceed the scoring threshold. The computer implementation method according to claim 1, further comprising outputting one or more applicable diagnostic tests with scores exceeding the scoring threshold.

5. Based on the applicable diagnostic tests, determine one or more treatments that may be suitable for the subject, The computer implementation method according to claim 1, further comprising outputting the one or more treatment methods to a display.

6. Determining a ranked set of diagnostic tests for a subject based on the applicable diagnostic tests, The computer implementation method according to claim 1, further comprising outputting the aforementioned ranked series of diagnostic tests to a display.

7. The computer implementation method according to claim 1, further comprising displaying the applicable diagnostic tests to the user.

8. A system for processing subject data and identifying diagnostic tests, At least one memory to store instructions, A system comprising at least one processor, wherein the at least one processor executes the instruction, Receiving multiple subject data, including one or more digital pathology images, The machine learning system is applied to the aforementioned plurality of subject data to identify any of the prerequisites for each of the plurality of diagnostic tests, and to determine the applicability of the diagnostic test based on the prerequisites, wherein the machine learning system is trained by processing training subject data paired with diagnostic test information that includes at least one of the prerequisite results and the test applicability results. Using the machine learning system, identify applicable diagnostic tests from among the multiple diagnostic tests based on the multiple subject data and the preconditions, A system configured to perform operations including outputting the applicable diagnostic tests to a digital storage device and / or display.

9. The system according to claim 8, wherein the plurality of subject data further comprises one or more of additional subject information relating to the subject and / or disease and additional diagnostic test information.

10. The operation described above is: Filtering one or more of the aforementioned digital pathology images to identify the tissue region to be analyzed, The system according to claim 8, further comprising removing one or more regions from the one or more digital pathological images that are not identified as the target tissue region.

11. The operation described above is: The machine learning system is provided with a scoring threshold, Based on the scoring threshold, determine one or more applicable diagnostic tests that exceed the scoring threshold. The system according to claim 8, further comprising outputting one or more applicable diagnostic tests with scores exceeding the scoring threshold.

12. The operation described above is: Based on the applicable diagnostic tests, determine one or more treatments that may be suitable for the subject, The system according to claim 8, further comprising outputting one or more of the above-mentioned treatment methods to a display.

13. The operation described above is: Based on the applicable diagnostic tests, determine a ranked set of diagnostic tests for the subject, The system according to claim 8, further comprising outputting the aforementioned ranked series of diagnostic tests to a display.

14. The operation described above is: The system according to claim 8, further comprising displaying the applicable diagnostic tests to the user.

15. A non-temporary computer-readable medium that stores instructions causing the processor to perform a method of processing subject data to identify a diagnostic test, wherein the method is Receiving multiple subject data, including one or more digital pathology images, The machine learning system is applied to the aforementioned plurality of subject data to identify any of the prerequisites for each of the plurality of diagnostic tests, and to determine the applicability of the diagnostic test based on the prerequisites, wherein the machine learning system is trained by processing training subject data paired with diagnostic test information that includes at least one of the prerequisite results and the test applicability results. Using the machine learning system, identify applicable diagnostic tests from among the multiple diagnostic tests based on the multiple subject data and the preconditions, A non-temporary computer-readable medium, including outputting the applicable diagnostic tests to a digital storage device and / or display.

16. The non-temporary computer-readable medium according to claim 15, wherein the plurality of subject data further comprises one or more of additional subject information relating to the subject and / or disease and additional diagnostic test information.

17. The method described above is: Filtering one or more of the aforementioned digital pathology images to identify the tissue region to be analyzed, The non-temporary computer-readable medium according to claim 15, further comprising removing one or more regions from the one or more digital pathological images that are not identified as the target tissue region.

18. The method described above is: The machine learning system is provided with a scoring threshold, Based on the scoring threshold, determine one or more applicable diagnostic tests that exceed the scoring threshold. A non-temporary computer-readable medium according to claim 15, further comprising outputting one or more applicable diagnostic tests with scores exceeding the scoring threshold.

19. The method described above is: Based on the applicable diagnostic tests, determine one or more treatments that may be suitable for the subject, The non-temporary computer-readable medium according to claim 15, further comprising outputting one or more of the above-mentioned treatment methods to a display.

20. The method described above is: Based on the applicable diagnostic tests, determine a ranked set of diagnostic tests for the subject, A non-temporary computer-readable medium according to claim 15, further comprising outputting the aforementioned ranked series of diagnostic tests to a display.