Systems and methods for processing electronic images to identify diagnostic tests

CN116635906BActive Publication Date: 2026-09-08PAIGE AI INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180086755.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-23
Filing Date
2021-10-19
Publication Date
2026-09-08
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

然而,由于各种因素,包括医生对测试不熟悉、设施内测试不可用、缺乏成功执行推荐测试的可行标本、测试前对特定测试可能对该患者产生积极效果的预期低或测试所识别出的治疗的成本高,所以可能无法对患者进行重要的诊断测试

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116635906B_ABST
    Figure CN116635906B_ABST
Patent Text Reader

Abstract

Systems and methods for processing digital images to identify diagnostic tests are disclosed, the methods comprising: receiving one or more digital images associated with a pathology specimen; determining a plurality of diagnostic tests; applying a machine learning system to the one or more digital images to identify any preconditions that each of the plurality of diagnostic tests will be applicable, the machine learning system having been trained by processing a plurality of training images; identifying, using the machine learning system, an applicable diagnostic test of the plurality of diagnostic tests based on the one or more digital images and the preconditions; and outputting the applicable diagnostic test to a digital storage device and / or a display.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 104,923, filed October 23, 2020, the entire disclosure of which is incorporated herein by reference. Technical Field

[0003] Various embodiments of this disclosure generally relate to image processing methods. More specifically, specific embodiments of this disclosure relate to systems and methods for processing electronic images to prioritize and / or identify diagnostic tests. Background Technology

[0004] Diagnostic testing methods for identifying therapies and treatments for diseased tissue continue to be developed and made available for clinical practice. Diagnostic tests have the potential to benefit patients by excluding ineffective treatments and / or by identifying therapies most likely to provide significant benefit to a patient's disease through the detection of the absence and / or presence of biomarkers (e.g., practices known as "precision medicine"). However, important diagnostic tests may not be performed on patients due to various factors, including physician unfamiliarity with the tests, unavailability of in-facility tests, lack of feasible specimens for successful execution of recommended tests, low pre-test expectations of a particular test's potential positive effect on the patient, or high cost of treatments identified by the test. The techniques presented in this article address this clinical need by identifying which tests may benefit the patient, prioritizing these tests, and providing this information to both patients and physicians.

[0005] The background description provided herein is intended to provide a general overview of the context of this disclosure. Unless otherwise stated herein, the materials described in this section are not prior art to the claims of this application and are not, by virtue of their inclusion in this section, acknowledged as prior art or suggestions of prior art. Summary of the Invention

[0006] According to certain aspects of this disclosure, systems and methods for processing electronic images to recommend diagnostic tests based on tissue specimens are disclosed.

[0007] A method for processing digital images to identify diagnostic tests, the method comprising: receiving one or more digital images associated with a pathological specimen; determining a plurality of diagnostic tests; applying a machine learning system to the one or more digital images to identify any prerequisites applicable to each of the plurality of diagnostic tests, the machine learning system having been trained by processing a plurality of training images; using the machine learning system to identify applicable diagnostic tests among the plurality of diagnostic tests based on the one or more digital images and the prerequisites; and outputting the applicable diagnostic tests to a digital storage device and / or a display.

[0008] A system for processing digital images to identify diagnostic tests, the method comprising: receiving one or more digital images associated with a pathological specimen; determining a plurality of diagnostic tests; applying a machine learning system to the one or more digital images to identify any prerequisites applicable to each of the plurality of diagnostic tests, the machine learning system having been trained by processing a plurality of training images; using the machine learning system to identify applicable diagnostic tests among the plurality of diagnostic tests based on the one or more digital images and the prerequisites; and outputting the applicable diagnostic tests to a digital storage device and / or a display.

[0009] A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method for processing digital images to identify diagnostic tests, the method comprising: receiving one or more digital images associated with a pathological specimen; determining a plurality of diagnostic tests; applying a machine learning system to the one or more digital images to identify any prerequisites applicable to each of the plurality of diagnostic tests, the machine learning system having been trained by processing a plurality of training images; using the machine learning system to identify applicable diagnostic tests among the plurality of diagnostic tests based on the one or more digital images and the prerequisites; and outputting the applicable diagnostic tests to a digital storage device and / or a display.

[0010] It should be understood that the foregoing general description and the following detailed description are exemplary and illustrative only, and not intended to limit the disclosed embodiments claimed. Attached Figure Description

[0011] Various exemplary embodiments are illustrated in conjunction with the accompanying drawings, which are incorporated in and form part of this specification, and together with the specification serve to explain the principles of the disclosed embodiments.

[0012] Figure 1A An exemplary block diagram of a system and network for identifying diagnostic tests applicable to pathological specimens, according to an exemplary embodiment of the present disclosure, is shown.

[0013] Figure 1B An exemplary block diagram of a treatment analysis platform 100 according to an exemplary embodiment of the present disclosure is shown.

[0014] Figure 2A This is a flowchart illustrating an exemplary method for identifying diagnostic tests applied to pathological specimens according to an exemplary embodiment of the present disclosure.

[0015] Figure 2B This is a flowchart illustrating an exemplary method for training a machine learning system to identify relevant diagnostic tests according to an exemplary embodiment of the present disclosure.

[0016] Figure 2CThis is a flowchart illustrating an exemplary method for training a machine learning system according to an exemplary embodiment of the present disclosure.

[0017] Figure 2D This is a flowchart illustrating an exemplary method for using a trained system to identify applicable tests for pathological specimens according to an exemplary embodiment of the present disclosure.

[0018] Figure 3 This is an exemplary workflow for determining test suitability according to an exemplary embodiment of this disclosure.

[0019] Figure 4 An example system that can implement the techniques proposed in this paper is described. Detailed Implementation

[0020] Exemplary embodiments of this disclosure will now be referenced in detail, examples of which are shown in the accompanying drawings. Where possible, the same reference numerals will be used in all the drawings to denote the same or similar parts.

[0021] The systems, apparatuses, and methods disclosed herein are described in detail by way of example and with reference to the accompanying drawings. The examples discussed herein are merely illustrative and are provided to aid in the explanation of the devices, apparatuses, systems, and methods described herein. Unless specifically designated as mandatory, the features or components shown in the drawings or discussed below should not be considered mandatory for any particular implementation of any of these apparatuses, systems, or methods.

[0022] Furthermore, for any method described, whether or not it is described in conjunction with a flowchart, it should be understood that, unless the context otherwise specifies or requires, any explicit or implicit order of steps performed in the execution of the method does not mean that these steps must be performed in the presented order, but may be performed in different orders or in parallel.

[0023] As used herein, the term "exemplary" is used in the sense of "example" rather than "ideal." Furthermore, the terms "an" and "a" in this document do not indicate a limitation of quantity, but rather the presence of one or more of the mentioned items.

[0024] In some cases, computational analytics using machine learning can directly determine the results of diagnostic tests, while in others, they can be used to exclude tests that are unlikely to be valuable, prioritize those tests, and / or help prioritize among available tests. One or more embodiments of this disclosure implement this functionality, as well as the prioritization of unexcluded tests based on ancillary information such as their availability and cost.

[0025] While existing computational analyses focus on identifying the presence or absence of disease / biomarkers, the techniques proposed in this paper may include identifying diagnostic tests that can better inform treatment, as well as tests that are unlikely to inform clinicians.

[0026] Figure 1A An exemplary block diagram of a system and network for identifying diagnostic tests applicable to pathological specimens, according to an exemplary embodiment of the present disclosure, is shown.

[0027] Specifically, Figure 1A An electronic network 120 is shown that can be connected to servers in locations such as hospitals, laboratories, and / or doctors' offices. For example, doctor servers 121, hospital servers 122, clinical trial servers 123, research laboratory servers 124, and / or laboratory information systems 125 can each be connected to the electronic network 120, such as the Internet, via one or more computers, servers, and / or handheld mobile devices. According to an exemplary embodiment of this application, the electronic network 120 can also be connected to a server system 110, which, according to an exemplary embodiment of this disclosure, may include processing means configured to implement a therapeutic analysis platform 100, which includes a slide analysis tool 101 for determining specimen characteristics or image characteristic information related to digital pathology images and using machine learning to determine the presence of disease or infectious pathogens. The slide analysis tool 101 can also predict appropriate diagnostic tests for pathology specimens.

[0028] Physician server 121, hospital server 122, clinical trial server 123, research laboratory server 124, and / or laboratory information system 125 may create or otherwise obtain images of cytological specimens, histopathological specimens, slides of cytological specimens, digitized images of slides of histopathological specimens, or any combination thereof from one or more patients. Physician server 121, hospital server 122, clinical trial server 123, research laboratory server 124, and / or laboratory information system 125 may also obtain any combination of patient-specific information, such as age, medical history, cancer treatment history, family history, past biopsies, or cytological information. Physician server 121, hospital server 122, clinical trial server 123, research laboratory server 124, and / or laboratory information system 125 may transmit digitized slide images and / or patient-specific information to server system 110 via electronic network 120. Server system 110 may include one or more storage devices 109 for storing images and data received from at least one of a physician server 121, a hospital server 122, a clinical trial server 123, a research laboratory server 124, and / or a laboratory information system 125. Server system 110 may also include processing means for processing the images and data stored in storage devices 109. Server system 110 may also include one or more machine learning tools or functions. For example, according to one embodiment, the processing means may include machine learning tools for a treatment analysis platform 100. Alternatively or additionally, this disclosure (or part of the systems and methods of this disclosure) may be executed on a local processing device (e.g., a laptop computer).

[0029] Physician server 121, hospital server 122, clinical trial server 123, research laboratory server 124, and / or laboratory system 125 refer to the systems used by pathologists to examine slide images. In a hospital setting, tissue type information may be stored in laboratory information system 125.

[0030] Figure 1B An exemplary block diagram of a therapeutic analysis platform 100 that uses machine learning to determine specimen or image characteristic information related to digital pathology images is shown. The therapeutic analysis platform 100 may include a slide analysis tool 101, a data acquisition tool 102, a slide acquisition tool 103, a slide scanner 104, a slide manager 105, a memory 106, a laboratory information system 107, and a viewing application tool 108.

[0031] As described below, slide analysis tool 101 refers to a process and system for determining diagnostic information associated with digital pathology images. According to an exemplary embodiment, machine learning can be used to classify the images. Slide analysis tool 101 can also receive additional information associated with the pathology specimen, as described in the embodiments below.

[0032] According to an exemplary embodiment, data ingestion tool 102 can facilitate the transmission of digital pathology images to various tools, modules, components, and devices for classifying and processing digital pathology images.

[0033] According to an exemplary embodiment, the slide acquisition tool 103 can scan pathological images and convert them into digital form. Slides can be scanned using a slide scanner 104, and a slide manager 105 can process the images on the slides into digitized pathological images and store the digitized images in a memory 106.

[0034] According to an exemplary implementation, the viewing application tool 108 can provide the user with specimen or image characteristic information related to digital pathology images. This information can be provided through various output interfaces (e.g., screen, monitor, storage device, and / or web browser, etc.).

[0035] The slide analysis tool 101 and one or more of its components can transmit and / or receive digitized slide images and / or patient information via network 120 to server system 110, physician server 121, hospital server 122, clinical trial server 123, research laboratory server 124, and / or laboratory information system 125. Furthermore, server system 110 may include storage devices for storing images and data received from at least one of the slide analysis tool 101, data acquisition tool 102, slide acquisition tool 103, slide scanner 104, slide manager 105, and viewing application tool 108. Server system 110 may also include processing means for processing the images and data stored in the storage device. Server system 110 may also include one or more machine learning tools or functions, for example, due to the processing means. Alternatively or additionally, this disclosure (or part of the systems and methods of this disclosure) may be executed on a local processing device (e.g., a laptop computer).

[0036] Any of the aforementioned devices, tools, and modules may be located on a device that can be connected to an electronic network such as the Internet or a cloud service provider via one or more computers, servers, and / or handheld mobile devices.

[0037] Figure 2AA method for identifying a set of diagnostic tests for a pathological specimen is illustrated according to an exemplary embodiment of the present disclosure. For example, exemplary method 200 (e.g., steps 202 to 210) may be performed automatically by slide analysis tool 101 or in response to a request from a user.

[0038] According to one embodiment, an exemplary method 200 for identifying a set of diagnostic tests to be applied to a pathological specimen may include one or more of the following steps. In step 202, the method may include receiving one or more digital images associated with the pathological specimen (e.g., histological, cytological, etc.) into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0039] Optionally, the method may include receiving additional information about the patient and / or the disease associated with the pathological specimen. This additional information may include, but is not limited to, patient demographics, past medical history, additional clinicopathological and / or biochemical test results, radiological imaging, historical pathological specimen images, tumor size, cancer grade, cancer stage, and information about the specimen (e.g., specimen location, position within a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0040] Optionally, the method may include receiving additional testing information. This additional testing information may include, but is not limited to, the availability of tests at a local (nearby) medical facility, test supply, current clinical guidelines for testing, current regulatory directives for testing, average time (test speed and turnaround time) to obtain results for one or more tests, current test pricing, available clinical trials, etc., stored in digital storage devices (e.g., hard drives, network drives, cloud storage, RAM, etc.).

[0041] Optionally, the method may also include receiving additional test preference information. This additional preference information may include information stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.) regarding whether the test is covered by insurance (government healthcare, patient insurance, etc.), out-of-pocket costs after considering insurance, tests preferred by doctors (laboratories, hospitals), and tests preferred by patients (e.g., due to religious practices, patient age, potential medical conditions, side effects, etc.).

[0042] In step 204, the method may include determining multiple diagnostic tests.

[0043] In step 206, the method may include applying a machine learning system to one or more digital images to identify any prerequisites that would apply to each of a plurality of diagnostic tests, the machine learning system having been trained by processing a plurality of training images. Diagnostic tests may include, but are not limited to, molecular tissue tests (genome sequencing, immunohistochemistry (IHC), fluorescence in situ hybridization (FISH), chromogenic in situ hybridization (CISH), in situ hybridization (ISH), genetic tests, special staining, algorithmic (computational, artificial intelligence, machine learning) tests, radiological tests, additional biopsies (specimens), laboratory tests (including biochemical and / or chemical pathology tests, such as blood, urine, sputum, etc.), and output to digital storage devices (e.g., hard drives, electronic medical records, laboratory information systems, network drives, etc.) and / or user displays (e.g., monitors, documents, printed copies, etc.).

[0044] In step 208, the method may include using a machine learning model to identify applicable diagnostic tests among a plurality of diagnostic tests based on one or more digital images and prerequisites. Scoring the diagnostic tests can indicate several representations of desirability. Examples may include a ranking of preferred tests based on the potential benefit to the patient, cost-effectiveness, the benefit of the test results relative to the benefit, and / or the availability of the therapeutic agent or method using the recommended treatment dosage and dosing regimen.

[0045] In step 210, the method may include outputting a sorted set of diagnostic tests to a digital storage device and / or a display.

[0046] Optionally, the method may include inputting a scoring threshold and outputting one or more tests, or outputting only those tests whose scores are higher than the threshold (if zero tests score higher than the threshold, then no tests are included).

[0047] Optionally, the method may include, based on study inclusion and exclusion criteria and geographical proximity, input information and / or additionally suggested tests, an output that may be considered a treatment strategy for the patient or one or more therapies, dosages or dosing regimens available in the patient’s clinical trials.

[0048] Optionally, the method may include displaying a sorted set of diagnostic tests to a user (e.g., referring clinician, testing laboratory, diagnostic company, treatment company, and / or patient). Test results may be displayed using a customized interface, output documents (e.g., PDF), printouts, etc.

[0049] One or more exemplary implementations may include one or more of the following three components:

[0050] Training a machine learning system to identify test suitability

[0051] Use a trained system to identify applicable tests

[0052] Ranking of applicable tests based on auxiliary information

[0053] Training a machine learning system to identify test suitability

[0054] Figure 2B This is a flowchart illustrating an exemplary method for training a machine learning system to identify test suitability according to the techniques proposed herein. For example, exemplary methods 220 and 240 (e.g., steps 222 to 224 and steps 242 to 252) may be performed automatically by slide analysis tool 101 or in response to a request from a user.

[0055] According to one embodiment, an exemplary method 220 for training a machine learning system to identify test applicability may include one or more of the following steps. In step 222, the method may include identifying at least the prerequisites for which a diagnostic test would be applicable. For example, some breast cancer recurrence tests (e.g., Oncotype DX) may require that a breast cancer patient may be estrogen receptor (ER) positive for the test to be applicable; if computational analysis identifies that the patient may not be ER positive, then the use of Oncotype DX on the patient is excluded.

[0056] In step 224, the method may include using a machine learning system to predict negative predictive values ​​for one or more diagnostic tests. For example, because genomic testing can be expensive and time-consuming, determining that a patient does not have a mutation associated with a specific drug treatment that could indicate the need for genomic testing does not provide added value. If the system cannot rule out the presence of a mutation, then a genomic test targeting the presence of that mutation may be a valid test to perform. Another example is when immunohistochemical and / or genomic testing may be required on a population basis (e.g., the application of NTRK fusion gene or microsatellite instability assessment in patients with metastatic cancer), but the prevalence of that biomarker in the population is low. If the system cannot rule out the presence of an immunohistochemical and / or genomic signature, then immunohistochemical and / or genomic testing may be performed.

[0057] Method 240 is a flowchart of a machine learning system for training according to an exemplary embodiment. For example, exemplary method 240 (e.g., steps 242 to 252) may be performed automatically by slide analysis tool 101 or in response to a request from a user. In step 242, the method may include receiving one or more digital images from a patient associated with a pathological specimen (e.g., histology, cytology, etc.) into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.), wherein the one or more digital images are paired with information regarding the results and / or values ​​of one or more diagnostic tests performed or the suitability of tests included in the diagnostic tests.

[0058] In step 244, the method may include receiving additional information about the patient and / or the disease associated with one or more digital images. This additional information may include, but is not limited to, receiving patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position within a block, etc.) from a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0059] In step 246, the method may include filtering one or more digital images to identify regions of interest for analysis, and removing non-salient regions from the one or more digital images, such as background and / or anything not identified as a region of interest. Regions of interest may be identified based at least in part on additional information relating to the patient and / or disease. The determination of regions of interest / salient regions may be performed using the techniques discussed in U.S. Application No. 17 / 313617, which is incorporated herein by reference. Filtering one or more images may be accomplished by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma).

[0060] In step 248, the method may include training a multi-binary machine learning system to predict one or more diagnostic tests and whether one or more diagnostic tests are performed, and determining the applicability of one or more diagnostic tests. If no test is performed, it is considered missing data for the patient and is not used to update the parameters of the machine learning system. If available, additional patient data (medical history, current outcomes, etc.) may be fed into the machine learning system to provide additional information (e.g., this may be accomplished using a neural network-based approach by transforming that information into a vector and then conditioning the image processing using conditional batch normalization). Many machine learning systems may be trained to do this by applying them to image pixels from samples from each patient, including but not limited to:

[0061] a. Multilayer Perceptron (MLP)

[0062] b. Convolutional Neural Network (CNN)

[0063] c. Graphical Neural Networks

[0064] d. Support Vector Machine (SVM)

[0065] e. Random Forest

[0066] In step 250, the method may include setting at least one threshold for one or more binary outputs of the machine learning system. For outputs corresponding to prerequisites for a diagnostic test, at least one threshold may be set to optimize the detection of that prerequisite (e.g., the presence of a biomarker that makes the diagnostic test applicable). For outputs corresponding to individual tests, thresholds may be set to optimize NPV, thereby excluding the applicability of that diagnostic test.

[0067] In step 252, the method may include outputting a set of parameters from a multi-binary-level machine learning system to a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.). This set of parameters may include at least one threshold, as well as other data for tuning the machine learning system.

[0068] Use a trained system to identify applicable tests

[0069] Figure 2C This is a flowchart illustrating the use of a trained machine learning system for a patient according to the exemplary methods disclosed herein. After the machine learning system has been trained to determine applicable diagnostic tests, the user can apply the system to the patient. For example, exemplary method 260 (e.g., steps 262-270) may be performed automatically by slide analysis tool 101 or in response to a request from the user. In step 262, the method may include receiving one or more digital images associated with a pathological specimen (e.g., histological, cytological, IHC, etc.) into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0070] In step 264, the method may include receiving additional information about the patient and / or the disease associated with one or more digital images. This additional information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position within a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0071] In step 266, the method may include filtering one or more images to identify regions of interest and removing unsuitable regions from the one or more images. Filtering may be performed by manually annotating or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma).

[0072] In step 268, the method may include predicting the suitability of one or more diagnostic tests by applying a trained machine learning system to one or more digital images.

[0073] In step 270, the method may include outputting the predictive suitability of one or more diagnostic tests to a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0074] Ranking of applicable tests based on auxiliary information

[0075] Figure 2D This is a flowchart illustrating an exemplary method for ranking applicable diagnostic tests for pathological specimens according to the techniques proposed herein. After identifying applicable tests, optional steps include ranking the applicable tests based on patient and clinician preferences, test availability, test cost, test speed, etc. For example, exemplary method 280 (e.g., steps 282-290) may be performed automatically by slide analysis tool 101 or in response to a user request. In step 282, the method may include applying a trained machine learning system to identify a list of one or more applicable diagnostic tests for the pathological specimen, which produces an N-dimensional binary vector “y”, where one or more elements correspond to the applicability of a single test.

[0076] In step 284, the method may include receiving additional testing and preference information regarding the pathological specimen. The additional testing information may include, but is not limited to, information stored in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.) regarding the availability of tests at a local (nearby) healthcare facility, test supply, current clinical guidelines for testing, current regulatory directives for testing, average time to obtain results for one or more tests (test speed), current test pricing, etc. The additional preference information may include information stored in a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.) regarding which tests are covered by insurance (government healthcare, patient insurance, etc.), out-of-pocket costs after considering insurance, tests preferred by the physician (laboratory, hospital), and tests preferred by the patient (e.g., due to religious practices, patient age, potential medical conditions, side effects, etc.).

[0077] In step 286, the method may include scoring one or more tests to generate an N-dimensional vector “s” of scores. This can be done in many non-limiting ways:

[0078] a. Only use applicability and availability:

[0079] i. Set s = y. For any or all tests for which the prediction is applicable, if the test is unavailable, set the corresponding element s of that test to 0.

[0080] b. Applicability, availability, and speed:

[0081] i. Set s = y. For any or all tests for which the prediction is applicable, if the test is unavailable, set the corresponding element s for that test to 0; otherwise, set the corresponding element s to be inversely proportional to speed, such that faster tests will have higher scores.

[0082] c. Applicability, availability, speed, and patient out-of-pocket costs:

[0083] i. Set s = y. For any or all tests predicted to be applicable, if the test is unavailable, set the corresponding element s of that test to 0; otherwise, set the corresponding element s to a weighted sum based on user preferences, where the first term in the sum is inversely proportional to speed, such that faster tests will have higher scores, and the second term in the sum is inversely proportional to the patient's test cost minus the portion covered by insurance.

[0084] d. Suitability, availability, speed, patient out-of-pocket costs, and patient preferences:

[0085] i. Set s = y. For any or all tests predicted to be applicable, if the test is unavailable or if the test is one that the patient cannot use (e.g., due to religious practices, age, discomfort, etc.), set the corresponding element s of that test to 0; otherwise, set the corresponding element s to a weighted sum based on user preferences, where the first term in the sum is inversely proportional to speed, such that a faster test will have a higher score, and the second term in the sum is inversely proportional to the patient's test cost minus the portion covered by insurance.

[0086] In step 288, the method may include sorting the N-dimensional vector s such that tests with higher scores are preferred, which may involve sorting the tests within the vector by test scores.

[0087] Optionally, the method may include inputting a scoring threshold and outputting one or more tests, or may only output those tests that score above the threshold (if zero tests score above the threshold, then no tests are included).

[0088] Optionally, the method may also include outputting one or more therapies that may be suitable for the patient based on the input information in steps 282 to 288 and / or additional suggested tests.

[0089] In step 290, the method may include displaying test results to users (e.g., referring clinicians, testing laboratories, diagnostic companies, treatment companies, and / or patients) using a customized interface, output documents (e.g., PDF), printouts, etc.

[0090] Figure 3 This is an exemplary workflow 300 for determining the suitability of testing based on the techniques proposed herein. Figure 3 This is a description of a system that runs on image data from a patient to determine the applicability of N different diagnostic tests (before sorting), where the system outputs 1 if the test is applicable and 0 if the test is not applicable.

[0091] In step 302, the workflow may include a digital image of the input pathology specimen. In step 304, the pathology specimen and any available additional patient data may be input into the machine learning system.

[0092] In step 306, the workflow may include multi-label outputs that determine the suitability of each diagnostic test.

[0093] Exemplary implementation: Sequencing genomic, IHC, or ISH / FISH tests is performed even if the patient has a low probability of having a certain mutation or antigen before testing.

[0094] Genomic testing can be expensive, may not be available at all centers, may incur additional costs, and can be time-consuming. The techniques provided in this article can be used to determine when a genomic test may provide diagnostic value, thereby avoiding unnecessary testing. One or more exemplary implementations can be used to determine when IHC, ISH / FISH tests are appropriate.

[0095] Training a machine learning system to identify the suitability of genomic, IHC, or ISH / FISH tests.

[0096] The steps for training a machine learning system may include:

[0097] 1. Receive one or more digital images from the patient's pathological specimens (e.g., histological, cytological, etc.) into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images can be paired with information about the results of genomic testing (e.g., presence / absence of oncogenic mutations / fusions in a range of genes), IHC testing, and / or ISH / FISH testing.

[0098] 2. Optionally, receive additional patient information about each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0099] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove non-salient regions from one or more images.

[0100] 4. Train a multi-binary-label machine learning system to predict the presence of one or more oncogene mutations / fusions. If available, additional patient data (medical history, current outcomes, etc.) can be fed into the machine learning system to provide additional information (e.g., this can be accomplished using neural network-based methods by transforming that information into a vector and then conditioning the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels from samples of one or more patients, including but not limited to:

[0101] i. Multilayer Perceptron (MLP)

[0102] ii. Convolutional Neural Networks (CNN)

[0103] iii. Graphical Neural Networks

[0104] iv. Support Vector Machine (SVM)

[0105] v. Random Forest

[0106] 5. Thresholds can be set for one or more binary outputs of the system to optimize for determining the absence of mutations / fusions for each oncogene.

[0107] 6. Output the parameters of the trained system to a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0108] Use a trained system to identify whether genomic, IHC, or ISH / FISH testing is required.

[0109] 1. Receive digital images (e.g., histological, cytological, etc.) of pathological specimens from patients into digital storage devices (e.g., hard disk drives, network drives, cloud storage, RAM, etc.).

[0110] 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0111] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove non-salient regions from one or more images.

[0112] 4. Run a trained machine learning system on digital images from the patient, and incorporate the additional patient information, whereby it can be used to generate an N-dimensional vector of multi-label outputs corresponding to the absence of mutations / fusions of each oncogene.

[0113] 5. Send the predicted output to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0114] 6. Optionally, inform users which oncogenes have been excluded and suggest whether genomic testing should be performed.

[0115] Exemplary implementation: Ranking of multi-parameter gene expression assays for breast cancer, such as MammaPrint, OncotypeDX, EndoPredict, PAM50 (Prosigna), or breast cancer index.

[0116] The use of multi-parameter gene expression tests to guide breast cancer treatment decisions is increasing. These tests can identify patients at higher risk of breast cancer recurrence. Some of the tests used are MammaPrint, a 70-gene assay, and Oncotype DX, a 20-gene assay. These tests help guide treatment decisions if chemotherapy may be beneficial for patients with invasive breast cancer. A prerequisite for the Oncotype DX test may be that the patient is ER-positive, so ER-negative patients may need to be excluded. Other tests used to determine whether a patient may need chemotherapy include EndoPredict (a 12-gene risk score), PAM50 (a 50-gene assay), and the Breast Cancer Index.

[0117] Training a machine learning system to identify the applicability of a multi-parameter gene expression test for breast cancer patients.

[0118] The steps for training a machine learning system may include:

[0119] 1. Multiple digital images of invasive primary breast tumors from a patient's pathological specimen (e.g., histology) are received into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images can be paired with information about whether the patient is ER positive or negative and whether positivity also includes an Oncotype DX score.

[0120] 2. Optionally, receive additional patient information about each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0121] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove non-salient regions from one or more images.

[0122] 4. Train a multi-label machine learning system to predict whether a patient is ER-positive or ER-negative, and train it to predict the Oncotype DX score for ER-positive patients; otherwise, treat the Oncotype DX score as a missing value (e.g., if missing, it will not be used to update parameters). For other tests, train the multi-label machine learning system to predict the patient's cancer recurrence risk score. If available, additional patient data (medical history, current outcomes, etc.) can be fed into the machine learning system to provide additional information (e.g., this can be done using neural network-based methods by transforming that information into a vector and then conditioning the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels from samples from each patient, including but not limited to:

[0123] i. Multilayer Perceptron (MLP)

[0124] ii. Convolutional Neural Networks (CNN)

[0125] iii. Graphical Neural Networks

[0126] IV. Support Vector Machine (SVM)

[0127] v. Random Forest

[0128] Thresholds can be set for one or more binary outputs of the system such that if the system determines that the patient is ER negative, Oncotype DX is indicated as inapplicable, and if the system determines that the patient has a very low test score, multiparameter breast cancer gene expression testing is indicated as potentially leading to a predicted low risk of recurrence.

[0129] The parameters of the trained system are output to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0130] Using a trained system

[0131] After the system has been trained to determine the suitability of a multi-parameter breast cancer gene expression test, the steps for using the trained system on patients may include:

[0132] 1. Digital images (e.g., histological) of invasive primary breast tumors from a patient's pathological specimen are received into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0133] 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0134] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove unsuitable regions from one or more images.

[0135] 4. Run a trained machine learning system on digital images from the patient, incorporating any additional patient information available. If the system predicts the patient is ER-negative, indicate that Oncotype DX is not recommended. If the system predicts the patient may have a low score on a multiparameter breast cancer gene expression test, inform the user and recommend against using the test.

[0136] 5. Send the predicted output to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0137] Exemplary implementation: Ranking of multi-parameter gene expression tests for prostate cancer, such as Oncotype DX Genomic Prostate Score (GPS) or Prolaris.

[0138] The OncotypeDX GPS (17-gene assay) and Prolaris (46-gene assay) tests assess the likelihood of invasive prostate cancer and help guide treatment decisions. A higher GPS score or Prolaris risk score indicates a greater likelihood of invasive cancer, and that immediate treatment, such as surgery or radiation therapy, may be necessary.

[0139] The steps for training a machine learning system may include:

[0140] 1. Multiple digital images of prostate tumors from a patient's pathological specimen (e.g., histology) are received into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images can be paired with a gene expression test for prostate cancer.

[0141] 2. Optionally, receive additional patient information about each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0142] 3. Optionally, filter one or more images to identify regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions. Remove non-salient regions from one or more images.

[0143] 4. Train a multi-binary label machine learning system to predict OncoTypeDX GPS scores / Prolaris scores. If available, additional patient data (medical history, current outcomes, etc.) can be fed into the machine learning system to provide additional information (e.g., this can be done using neural network-based methods by transforming that information into a vector and then conditioning the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels from samples from each patient, including but not limited to:

[0144] i. Multilayer Perceptron (MLP)

[0145] ii. Convolutional Neural Networks (CNN)

[0146] iii. Graphical Neural Networks

[0147] iV Support Vector Machine (SVM)

[0148] V. Random Forest

[0149] 5. A threshold can be set for one or more binary outputs of the system, such that if a patient is determined to have a very low test score, instructing a multi-parameter prostate cancer gene expression test may result in a prediction of low-invasive prostate cancer.

[0150] 6. Output the parameters of the trained system to a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0151] Using a trained system

[0152] After the system has been trained to determine the suitability of Oncotype DX, the steps for using the trained system with patients may include:

[0153] 1. Digital images (e.g., histological) of invasive primary breast tumors from a patient's pathological specimen are received into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0154] 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0155] 3. Optionally, filter one or more images to identify regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions. Remove unsuitable regions from one or more images.

[0156] 4. Run a trained machine learning system on digital images from the patient, incorporating additional patient information if available. If the system predicts the patient may have a low Oncotype DX GPS or Prolaris score, inform the user and suggest not using Oncotype DX GPS or Prolaris.

[0157] 5. Send the predicted output to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0158] Exemplary implementation: Sequencing of single / multiple immunohistochemistry (IHC) and fluorescence in situ hybridization (FISH) tests for proteins such as HER2, mismatch repair (MMR) repair proteins, and PD-L1.

[0159] For a given clinical stage and type of cancer, additional IHC and / or FISH analyses may be essential for treatment decisions, even though the biomarkers are less frequently used. This can be illustrated by the need for unknown testing of tumor sites in all or multiple metastatic cancer patients for the presence of NTRK1, NTRK2, and NTRK3 fusion genes and microsatellite instability, in order to use specific treatment regimens (i.e., TRK inhibitors and immune checkpoint inhibitors, respectively). Similarly, it may be necessary to test for ALK, RET, and ROS1 rearrangements in non-small cell lung cancer patients to treat these patients in a metastatic setting.

[0160] Training a machine learning system to identify the applicability of single / multiple immunohistochemistry (IHC) and fluorescence in situ hybridization (FISH) tests.

[0161] The steps for training a machine learning system may include:

[0162] 1. Multiple digital images from pathological specimens (e.g., histological, cytological, etc.) from the patient are received into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images can be paired with information about the results of an IHC / FISH test or a related genomic test.

[0163] 2. Optionally, receive additional patient information about each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0164] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors, invasive tumor stroma). Remove non-salient regions from one or more images.

[0165] 4. Train a multi-binary-label machine learning system to predict the presence of IHC / FISH biomarkers. If available, additional patient data (medical history, current outcomes, etc.) can be fed into the machine learning system to provide additional information (e.g., this can be done using neural network-based methods by transforming that information into a vector and then conditioning the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels from samples from each patient, including but not limited to:

[0166] i. Multilayer Perceptron (MLP)

[0167] ii. Convolutional Neural Networks (CNN)

[0168] iii. Graphical Neural Networks

[0169] IV. Support Vector Machine (SVM)

[0170] v. Random Forest

[0171] 5. Thresholds can be set for one or more binary outputs of the system to optimize for the determination of the absence of a given IHC / FISH flag.

[0172] 6. Output the parameters of the trained system to a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0173] Using a trained system

[0174] After the system has been trained to determine the suitability of a single / multiple immunohistochemical (IHC) test, the steps for using the trained system on a patient may include:

[0175] 1. Receive digital images (e.g., histological, cytological, etc.) of pathological specimens from patients into digital storage devices (e.g., hard disk drives, network drives, cloud storage, RAM, etc.).

[0176] 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0177] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove non-salient regions from one or more images.

[0178] 4. Run a trained machine learning system on one or more digital images from a patient, and incorporate additional patient information if it is available to generate an N-dimensional vector of multi-label outputs corresponding to the absence of a given IHC / FISH marker.

[0179] 5. Send the predicted output to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0180] 6. Optionally, inform the user which IHC / FISH markers have been excluded and suggest whether this type of IHC / FISH should be performed.

[0181] Exemplary implementation: Sequencing multi-gene sequencing panels, such as Foundation One CDx or MSK IMPACT.

[0182] Multigenomic assays of tumor and / or tumor-normal pairs have shown benefit for cancer patients, with studies indicating that in up to >10% of patients with metastatic cancer, multigenomic sequencing assays may lead to more appropriate therapies and / or inclusion in clinical trials based solely on the results of these molecular tests. However, for the vast majority of patients, the information provided by these assays is limited or currently not useful. Furthermore, these assays are relatively expensive, have long turnaround times, and are only available in a limited number of institutions.

[0183] Training a machine learning system to identify the suitability of multi-gene sequencing kits

[0184] The steps for training a machine learning system may include:

[0185] 1. Multiple digital images from pathological specimens (e.g., histological, cytological, etc.) from a patient are received into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images can be paired with information about the results of multi-gene sequencing assays.

[0186] 2. Optionally, receive additional patient information about each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0187] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove non-salient regions from one or more images.

[0188] 4. Train a multi-binary-label machine learning system to predict the results of multi-gene sequencing assays. If available, additional patient data (medical history, current results, etc.) can be fed into the machine learning system to provide additional information (e.g., this can be done using neural network-based methods by transforming that information into a vector and then conditioning the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels from samples from each patient, including but not limited to:

[0189] i. Multilayer Perceptron (MLP)

[0190] ii. Convolutional Neural Networks (CNN)

[0191] iii. Graphical Neural Networks

[0192] iV Support Vector Machine (SVM)

[0193] V. Random Forest

[0194] 5. Thresholds can be set for one or more binary outputs of the system to optimize for the determination of the absence of clinically relevant findings derived from multi-gene sequencing assays.

[0195] 6. Output the parameters of the trained system to a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0196] Using a trained system

[0197] After the system has been trained to determine the suitability of a multi-gene sequencing kit, the steps for using the trained system with patients may include:

[0198] 1. Receive digital images (e.g., histological, cytological, etc.) of pathological specimens from patients into digital storage devices (e.g., hard disk drives, network drives, cloud storage, RAM, etc.).

[0199] 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0200] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove non-salient regions from one or more images.

[0201] A trained machine learning system is run on digital images from patients, and this additional patient information is incorporated into the output of an N-dimensional vector of multi-labeled outputs that correspond to clinically relevant results derived from multigene sequencing assays that are not currently available.

[0202] The predicted output is sent to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0203] Optionally, inform users which genetic and genomic alterations have been excluded and suggest whether multigene sequencing should be performed.

[0204] Exemplary implementation: Ranking assays to prioritize immuno-oncology (IO) therapies.

[0205] Immunotherapy is reshaping the treatment landscape for patients with various cancer types. Tumor-specific (e.g., PD-L1 assessment in non-small cell lung cancer and metastatic triple-negative breast cancer) and cancer-site-agnostic (e.g., microsatellite instability (MSI) or mismatch repair deficiency (dMMR) and tumor mutational burden (TMB)) biomarkers may now be needed for treatment decisions. However, their assessment typically involves multiple forms of assays (e.g., IHC, PCR, and / or multi-gene sequencing assays), which are expensive, time-consuming, and require subsequent integration.

[0206] Furthermore, new kits are being developed to better understand the composition of the tumor microenvironment and the immune profile of patients. The PanCancer IO 360 gene expression kit is a multi-gene expression kit developed to characterize expression patterns from the tumor, immune system, and stroma. It contains a tumor inflammation index (TIS), which includes 18 functional genes known to be associated with PD-1 / PD-L1 inhibitor pathway blockade responses. The PanCancer IO 360 kit and the TIS have the potential to assist physicians in making treatment decisions regarding IO therapies.

[0207] Training machine learning systems to identify metrics to help prioritize cancer immunotherapies

[0208] The steps for training a machine learning system may include:

[0209] 4. Multiple digital images from the patient's pathological specimens (e.g., histology, cytology, etc.) are received into a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.). For each patient, one or more images can be paired with information on specific biomarkers of immunotherapy response (e.g., PD-L1 expression, high / deficient microsatellite instability / mismatch repair (MSI / dMMR), tumor mutational burden (TMB), PanCancer IO360 kit, TIS).

[0210] 5. Optionally, receive additional patient information about each patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0211] 6. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove non-salient regions from one or more images.

[0212] 7. Train a multi-binary-label machine learning system to predict the presence of one or more specific biomarkers for immunotherapy response (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancer IO360 kit, TIS). If available, additional patient data (medical history, current outcomes, etc.) can be fed into the machine learning system to provide additional information (e.g., this can be accomplished using neural network-based methods by transforming that information into a vector and then conditioning the image processing using conditional batch normalization). Many machine learning systems can be trained to do this by applying them to image pixels from samples from each patient, including but not limited to:

[0213] i. Multilayer Perceptron (MLP)

[0214] ii. Convolutional Neural Networks (CNN)

[0215] iii. Graphical Neural Networks

[0216] iv. Support Vector Machine (SVM)

[0217] v. Random Forest

[0218] 8. Thresholds can be set for one or more binary outputs of the system to optimize for the determination of the absence of mutations / fusions of one or more specific biomarkers for immunotherapy response (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancer IO360 kit, TIS).

[0219] 9. Output the parameters of the trained system to a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0220] Use trained systems to identify assays to help prioritize cancer immunotherapies.

[0221] The steps involved in using a trained machine learning system may include:

[0222] 1. Receive digital images (e.g., histological, cytological, etc.) of pathological specimens from patients into digital storage devices (e.g., hard disk drives, network drives, cloud storage, RAM, etc.).

[0223] 2. Optionally, receive additional patient information about the patient and / or disease. This additional patient information may include, but is not limited to, patient demographics, past medical history, additional test results, radiological imaging, historical pathological specimen images, and information about the specimen (e.g., specimen location, position in a block, etc.) stored in a digital storage device (e.g., hard disk drive, network drive, cloud storage, RAM, etc.).

[0224] 3. Optionally, filter one or more images to identify tissue regions of interest that should be used. This can be done by manual annotation or by using a region detector to identify salient regions (e.g., invasive tumors and / or invasive tumor stroma). Remove unsuitable salient regions from one or more images.

[0225] 4. Run a trained machine learning system on digital images from the patient, and incorporate the additional patient information into the N-dimensional vector of multi-label outputs that are determined to be absent from one or more specific biomarkers for immunotherapy response (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancer IO 360 kit, TIS).

[0226] 5. Send the predicted output to a digital storage device (e.g., hard drive, network drive, cloud storage, RAM, etc.).

[0227] 6. Optionally, inform the user which specific biomarkers for immunotherapy response have been excluded (e.g., PD-L1 expression, MSI / dMMR, TMB, PanCancer IO 360 kit, TIS) and suggest whether IHC and / or genomic testing should be performed.

[0228] like Figure 4 As shown, device 400 may include a central processing unit (CPU) 420. CPU 420 can be any type of processor device, including, for example, any type of dedicated or general-purpose microprocessor device. Those skilled in the art will understand that CPU 420 can also be a single processor in a multi-core / multi-processor system, such a system operating independently, or a single processor in a cluster of computing devices operating in a cluster or server group. CPU 420 can be connected to data communication infrastructure 410, such as a bus, message queue, network, or multi-core messaging scheme.

[0229] Device 400 may further include main memory 440, such as random access memory (fRAM), and may also include auxiliary memory 430. Auxiliary memory 430, such as read-only memory (ROM), may be, for example, a hard disk drive or a removable storage drive. Such removable storage drives may include, for example, floppy disk drives, magnetic tape drives, optical disk drives, flash memory, etc. The removable storage drive in this example reads from and / or writes to the removable storage unit in a known manner. Removable storage devices may include floppy disks, magnetic tapes, optical disks, etc., which are read from and written to by the removable storage drive. Those skilled in the art will understand that such removable storage units typically include computer-usable storage media in which computer software and / or data are stored.

[0230] In an alternative implementation, auxiliary memory 430 may include similar means for allowing computer programs with additional instructions to be loaded into device 400. Examples of such means may include a program box and box interface (such as those seen in video game devices), a removable storage chip (such as EPROM or PROM) and associated sockets, as well as other removable storage units and interfaces that allow software and data to be transferred from the removable storage unit to device 400.

[0231] Device 400 may also include a communication interface (“COM”) 460. Communication interface 460 allows the transfer of software and data between device 400 and external devices. Communication interface 460 may include a modem, network interface (such as an Ethernet card), communication port, PCMCIA slot, and card, etc. The software and data transferred via communication interface 460 may be in the form of signals, which may be electronic, electromagnetic, optical, or other signals that can be received by communication interface 460. These signals may be provided to communication interface 460 via a communication path of device 400, which may be implemented using, for example, wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, or other communication channels.

[0232] The hardware components, operating system, and programming language of such devices are essentially conventional and are assumed to be sufficiently familiar to those skilled in the art. Device 400 may also include input and output ports 450 connected to input and output devices such as a keyboard, mouse, touchscreen, monitor, display, etc. Of course, various server functions can be implemented in a distributed manner on multiple similar platforms to distribute the load. Alternatively, the server can be implemented through appropriate programming of a single computer hardware platform.

[0233] Throughout this disclosure, references to components or modules generally refer to items that can be logically grouped together to perform a function or a related group of functions. The same reference numerals are generally intended to denote the same or similar components. Components and / or modules may be implemented in software, hardware, or a combination of software and / or hardware.

[0234] The aforementioned tools, modules, and / or functions may be executed by one or more processors. "Storage" media may include any or all tangible memory of a computer, processor, etc., or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time.

[0235] Software can communicate via the Internet, cloud service providers, or other telecommunications networks. For example, communication enables the loading of software from one computer or processor into another. As used herein, unless limited to non-transitory, tangible "storage" media, the term "readable medium" for a computer or machine refers to any medium that participates in providing instructions to a processor for execution.

[0236] The foregoing general description is exemplary and illustrative only, and is not intended to limit the scope of this disclosure. Other embodiments will be apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This specification and examples are intended to be considered exemplary only.

Claims

1. A computer-implemented method for processing digital images to identify diagnostic tests, the method comprising: Receive one or more digital images associated with a pathological specimen; Identify multiple diagnostic tests; Applying a machine learning system to the one or more digital images to identify any prerequisites for each of the plurality of diagnostic tests that would be applicable, the machine learning system having been trained by processing a plurality of training images, wherein processing the plurality of training images includes: receiving a plurality of digital images associated with at least one previous pathological specimen, the digital images being paired with diagnostic test information, the diagnostic test information being results or values ​​of one or more past diagnostic tests performed on the previous pathological specimen, or tests establishing the suitability of a diagnostic test for the previous pathological specimen; and training the machine learning system using the plurality of digital images and the diagnostic test information, the machine learning system including a multi-binary label machine learning system to predict the suitability of the past diagnostic tests, determining at least one threshold for one or more binary outputs of the multi-binary label machine learning system; and outputting a set of parameters from the multi-binary label machine learning system to a digital storage device, the set of parameters including the at least one threshold; The machine learning system is used to identify applicable diagnostic tests among the plurality of diagnostic tests based on the one or more digital images and the prerequisites; and The applicable diagnostic tests are output to a digital storage device and / or a display.

2. The computer-implemented method as described in claim 1, further comprising: Determine additional patient information, additional diagnostic test information, and / or additional test preference information associated with the pathological specimens regarding the patient and / or disease.

3. The computer-implemented method of claim 1, wherein identifying the applicable diagnostic test further includes predicting a negative predictive value (NPV) for each of the plurality of diagnostic tests.

4. The computer-implemented method as described in claim 1, further comprising: Filter the one or more digital images to identify regions of interest for analysis; as well as Remove one or more regions from the one or more digital images that were not identified as the region of interest.

5. The computer-implemented method as described in claim 1, further comprising: Provide a scoring threshold to the machine learning system; Based on the scoring threshold, one or more applicable diagnostic tests are determined that have a score higher than the scoring threshold; as well as Output one or more applicable diagnostic tests whose scores are higher than the stated score threshold.

6. The computer-implemented method as described in claim 1, further comprising: Based on the applicable diagnostic tests, one or more therapies that may be suitable for the patient are determined; as well as The one or more therapies are output to a display.

7. The computer-implemented method of claim 1, further comprising displaying the applicable diagnostic test to the user.

8. A system for processing digital images to identify diagnostic tests, the system comprising: At least one memory, the memory storing instructions; and At least one processor, configured to execute the instructions to perform an operation, the operation including: Receive one or more digital images associated with a pathological specimen; Identify multiple diagnostic tests; Applying a machine learning system to the one or more digital images to identify any prerequisites for each of the plurality of diagnostic tests that would be applicable, the machine learning system having been trained by processing a plurality of training images, wherein processing the plurality of training images includes: receiving a plurality of digital images associated with at least one previous pathological specimen, the digital images being paired with diagnostic test information, the diagnostic test information being results or values ​​of one or more past diagnostic tests performed on the previous pathological specimen, or tests establishing the suitability of a diagnostic test for the previous pathological specimen; and training the machine learning system using the plurality of digital images and the diagnostic test information, the machine learning system including a multi-binary label machine learning system to predict the suitability of the past diagnostic tests, determining at least one threshold for one or more binary outputs of the multi-binary label machine learning system; and outputting a set of parameters from the multi-binary label machine learning system to a digital storage device, the set of parameters including the at least one threshold; The machine learning system is used to identify applicable diagnostic tests among the plurality of diagnostic tests based on the one or more digital images and the prerequisites; and The applicable diagnostic tests are output to a digital storage device and / or a display.

9. The system of claim 8, further comprising: Determine additional patient information, additional diagnostic test information, and / or additional test preference information associated with the pathological specimens regarding the patient and / or disease.

10. The system of claim 9, wherein identifying the applicable diagnostic test further includes predicting a negative predictive value (NPV) for each of the plurality of diagnostic tests.

11. The system of claim 8, further comprising: Filter the one or more digital images to identify tissue regions of interest for analysis; as well as Remove one or more regions from the one or more digital images that were not identified as the region of interest.

12. The system of claim 8, further comprising: Provide a scoring threshold to the machine learning system; Based on the scoring threshold, one or more applicable diagnostic tests are determined that have a score higher than the scoring threshold; as well as Output one or more applicable diagnostic tests whose scores are higher than the stated score threshold.

13. The system of claim 8, further comprising: Based on the applicable diagnostic tests, one or more therapies that may be suitable for the patient are determined; as well as The one or more therapies are output to a display.

14. The system of claim 8, further comprising displaying the applicable diagnostic tests to the user.

15. A non-transitory computer-readable medium storing instructions, said instructions, when executed by a processor, causing the processor to perform a method for processing a digital image to identify diagnostic tests, said method comprising: Receive one or more digital images associated with a pathological specimen; Identify multiple diagnostic tests; Applying a machine learning system to the one or more digital images to identify any prerequisites for each of the plurality of diagnostic tests that would be applicable, the machine learning system having been trained by processing a plurality of training images, wherein processing the plurality of training images includes: receiving a plurality of digital images associated with at least one previous pathological specimen, the digital images being paired with diagnostic test information, the diagnostic test information being results or values ​​of one or more past diagnostic tests performed on the previous pathological specimen, or tests establishing the suitability of a diagnostic test for the previous pathological specimen; and training the machine learning system using the plurality of digital images and the diagnostic test information, the machine learning system including a multi-binary label machine learning system to predict the suitability of the past diagnostic tests, determining at least one threshold for one or more binary outputs of the multi-binary label machine learning system; and outputting a set of parameters from the multi-binary label machine learning system to a digital storage device, the set of parameters including the at least one threshold; The machine learning system is used to identify applicable diagnostic tests among the plurality of diagnostic tests based on the one or more digital images and the prerequisites; and The applicable diagnostic tests are output to a digital storage device and / or a display.

16. The non-transitory computer-readable medium of claim 15, further comprising: Determine additional patient information, additional diagnostic test information, and / or additional test preference information associated with the pathological specimens regarding the patient and / or disease.

17. The non-transitory computer-readable medium of claim 16, wherein identifying the applicable diagnostic test further includes predicting a negative predictive value (NPV) for each of the plurality of diagnostic tests.

Citation Information

Patent Citations

  • Systems and methods to process electronic images to determine salient information in digital pathology

    US20210350166A1

  • Optimization of clinical decision making

    US20180315505A1