Information processing system, information processing method, and program

By developing an information processing system that utilizes feature prediction models and gene mutation prediction models, the treatment delay problem caused by delayed gene mutation test results is solved, and a rapid and timely treatment plan is achieved.

JP7672642B2Active Publication Date: 2025-05-08GENOMEDIA INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022556767
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-10-14
Publication Date
2025-05-08
Estimated Expiration
2040-10-14

AI Technical Summary

Technical Problem

Gene mutation test results for diseases such as cancer usually take 1-2 months, resulting in delays in drug treatment, which is a serious problem especially for patients with advanced cancer.

Method used

An information processing system was developed, which quickly predicts the existence of gene mutations by obtaining the patient's histopathological images, segments the images into multiple regional images, and uses feature prediction models and gene mutation prediction models to quickly predict the existence of gene mutations, thereby formulating treatment plans in a timely manner.

Benefits of technology

The system can predict the existence of gene mutations in a short period of time, helping doctors to formulate treatment plans in a timely manner and avoid the progression of the disease caused by delayed treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672642000001
    Figure 0007672642000001
  • Figure 0007672642000002
    Figure 0007672642000002
  • Figure 0007672642000003
    Figure 0007672642000003
Patent Text Reader

Abstract

The present invention comprises: an acquisition unit which acquires a pathologic tissue image of a patient of a target disease; a division unit which divides the pathologic tissue image of the patient into a plurality of area images; a feature prediction unit which inputs the area images to each of a plurality of respective feature prediction models constructed for each of the types of histopathological features and acquires pieces of prediction information about the presence or absence of the respective histopathological features; a sorting unit which sorts a plurality of area images in which acquired combinations of the presence or absence of the histopathological features match preset combinations of the presence or absence of the histopathological features at the time of sorting; a gene mutation prediction unit which inputs the area images selected by sorting each of gene mutation prediction models constructed for each of the respective types of gene mutations and the combinations of the presence or absence of the histopathological features, and acquires pieces of prediction information about the presence or absence of the gene mutations; and a prediction result output unit which outputs a prediction result of the presence or absence of at least one gene mutation with respect to the patient by using the pieces of prediction information about the presence or absence of the gene mutations for each of the acquired respective area images.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] For cancers such as lung cancer, colon cancer, stomach cancer, breast cancer, GIST (gastrointestinal stromal tumor), and skin cancer (e.g., malignant melanoma), cancer genetic testing is performed if a doctor deems it necessary to examine one or several genes and make a diagnosis, or to select medications for treatment based on the test results (see Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] https: / / ganjoho.jp / public / dia_tre / treatment / genomic_medicine / gentest02.html Summary of the Invention [Problem to be solved by the invention]

[0004] Because it takes one to two months for the results of cancer gene testing to become available, there is a risk that the administration of medicines appropriate to the patient's genetic abnormality will be delayed, and the condition may progress in the meantime, making it too late. This is a particularly serious problem for cancer patients with stage 4 or other cancers, who have little time to spare.

[0005] Similar problems exist not only for cancer, but also for diseases for which drugs can be administered in response to genetic abnormalities.

[0006] One aspect of the present invention has been made in consideration of the above-mentioned problems, and aims to provide an information processing system, an information processing method, and a program that make it possible to suppress delays in administration of pharmaceuticals in response to genetic abnormalities of a target disease.

[0007] Another aspect of the present invention has been made in consideration of the above-mentioned problems, and aims to provide an information processing system, an information processing method, and a program that make it possible to suppress delays in administration of pharmaceuticals in accordance with genetic abnormalities in colorectal cancer patients.

[0008] Another aspect of the present invention has been made in consideration of the above-mentioned problems, and aims to provide an information processing system that makes it possible to suppress delays in administering pharmaceuticals to cancer patients in accordance with their genetic abnormalities. [Means for solving the problem]

[0009] An information processing system according to a first aspect of the present invention comprises an acquisition unit that acquires a pathological tissue image of a patient's tissue with a target disease; a division unit that divides the pathological tissue image of the patient into a plurality of region images; a feature prediction unit that inputs the region images to a plurality of feature prediction models constructed for each type of pathological histological feature and acquires prediction information on the presence or absence of the histological feature, respectively; a selection unit that selects a plurality of region images in which the acquired combinations of the presence / absence of the histological features match a predetermined combination of the presence / absence of the histological features at the time of selection; a gene mutation prediction unit that inputs the region images selected by the selection to each of the combinations of the presence / absence of the histological features and the gene mutation prediction models constructed for each type of gene mutation and acquires prediction information on the presence / absence of the gene mutation, respectively; and a prediction result output unit that uses the prediction information on the presence / absence of the gene mutation for each of the acquired region images to output a prediction result of the presence / absence of at least one gene mutation for the patient.

[0010] According to this configuration, by inputting a pathological tissue image of a patient with a target disease, the prediction result of the presence or absence of a gene mutation can be obtained immediately. This allows a clinician to refer to the prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of the target disease.

[0011] An information processing system according to a second aspect of the present invention is the information processing system according to the first aspect, wherein the acquisition unit further acquires the site of the primary lesion of the target disease, and the selection unit selects the multiple regional images using the acquired site of the primary lesion of the target disease together with a combination of the presence or absence of the acquired histopathological features.

[0012] According to this configuration, by selecting regional images using the site of the primary lesion of the target disease as well, the probability that the regional images input into the gene mutation prediction model can be limited to regional images related to the target disease is improved, thereby improving the accuracy of predicting the presence or absence of gene mutations.

[0013] An information processing system according to a third aspect of the present invention is the information processing system according to the first or second aspect, wherein the feature prediction model is a model trained by machine learning using learning data that takes as input an area image obtained by dividing a pathological tissue image and uses as output pathological histological features assigned to the area image, and the gene mutation prediction model is a model trained by machine learning using learning data that takes as input an area image selected using a combination of the presence or absence of pathological histological features and uses as output information on the presence or absence of a specific gene mutation.

[0014] According to this configuration, since the feature prediction model and the gene mutation prediction model are models obtained through machine learning, it is possible to improve the accuracy of predicting the presence or absence of a gene mutation.

[0015] An information processing system according to a fourth aspect of the present invention is an information processing system for predicting genes in which mutations have occurred in colon cancer tissue of a colon cancer patient, the information processing system including: an acquisition unit that acquires a colon cancer pathological tissue image of the patient; a division unit that divides the colon cancer pathological tissue image of the patient into a plurality of region images; a feature prediction unit that inputs the region images to a plurality of feature prediction models constructed for each type of pathological histological feature, and acquires prediction information on the presence or absence of the histopathological feature, respectively; a selection unit that selects a plurality of region images in which a combination of the presence or absence of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection that is preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI (microsatellite instability); a gene mutation prediction unit that inputs the selected region images to each of the combinations of the presence or absence of the pathological histological features and the gene mutation prediction models constructed for each type of gene mutation, and acquires prediction information on the presence or absence of the gene mutation, respectively; and a prediction result output unit that outputs a prediction result of the presence or absence of at least one gene mutation among V600E, ERBB2, RAS, TP53, or MSI.

[0016] According to this configuration, by inputting the colon cancer pathology tissue image of a colon cancer patient, the prediction result of the presence or absence of gene mutation can be obtained immediately. This allows a clinician to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of colon cancer.

[0017] An information processing system according to a fifth aspect of the present invention is the information processing system according to the fourth aspect, wherein the acquisition unit further acquires a site of a primary lesion of colorectal cancer, and the selection unit selects the plurality of regional images using the acquired site of the primary lesion of colorectal cancer together with the determined at least one histopathological feature.

[0018] According to this configuration, by selecting regional images using the site of the primary lesion of the target disease as well, the probability that the regional images input into the gene mutation prediction model can be limited to regional images related to the target disease is improved, thereby improving the accuracy of predicting the presence or absence of gene mutations.

[0019] An information processing system according to a sixth aspect of the present invention is the information processing system according to the fourth or fifth aspect, wherein the feature prediction model is a model trained by machine learning using learning data that takes as input an area image obtained by dividing a pathological tissue image and uses as output pathological histological features assigned to the area image, and the gene mutation prediction model is a model trained by machine learning using learning data that takes as input an area image selected using a combination of the presence or absence of pathological histological features and uses as output information on the presence or absence of a specific gene mutation.

[0020] According to this configuration, since the feature prediction model and the gene mutation prediction model are models obtained through machine learning, it is possible to improve the accuracy of predicting the presence or absence of a gene mutation.

[0021] An information processing method according to a seventh aspect of the present invention includes an acquisition step of acquiring a pathological tissue image of a patient with a target disease; a division step of dividing the pathological tissue image of the patient into a plurality of region images; a feature prediction step of inputting the region image into each of a plurality of feature prediction models constructed for each type of pathological histological feature, and acquiring prediction information on the presence or absence of the histopathological feature, respectively; a selection step of selecting a plurality of region images in which the acquired combinations of the presence / absence of the histopathological features match a combination of the presence / absence of the histopathological features at the time of selection that is preset; a gene mutation prediction step of inputting the region image selected by the selection into each of the combinations of the presence / absence of the histopathological features and the gene mutation prediction models constructed for each type of gene mutation, and acquiring prediction information on the presence / absence of the gene mutation, respectively; and a prediction result output step of outputting a prediction result of the presence / absence of at least one gene mutation for the patient using the prediction information on the presence / absence of the gene mutation for each acquired region image.

[0022] According to this configuration, by inputting a pathological tissue image of a patient with a target disease, the prediction result of the presence or absence of a gene mutation can be obtained immediately. This allows a clinician to refer to the prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of the target disease.

[0023] An information processing method according to an eighth aspect of the present invention is an information processing method for predicting genes in which mutations have occurred in colon cancer tissue of a colon cancer patient, the information processing method including: an acquisition step of acquiring a colon cancer pathological tissue image of the patient; a division step of dividing the colon cancer pathological tissue image of the patient into a plurality of region images; a feature prediction step of inputting the region images into a plurality of feature prediction models constructed for each type of pathological histological feature, and acquiring prediction information on the presence or absence of the histopathological feature, respectively; a selection step of selecting a plurality of region images in which the acquired combinations of the presence or absence of the histopathological features match a combination of the presence or absence of the histopathological features at the time of selection that is preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; a gene mutation prediction step of inputting the selected region images into each of the combinations of the presence or absence of the histopathological features and the gene mutation prediction models constructed for each type of gene mutation, and acquiring prediction information on the presence or absence of the gene mutation, respectively; and a prediction result output step for outputting a prediction result of the presence or absence of at least one gene mutation among V600E, ERBB2, RAS, TP53, or MSI.

[0024] According to this configuration, by inputting the colon cancer pathology tissue image of a colon cancer patient, the prediction result of the presence or absence of gene mutation can be obtained immediately. This allows a clinician to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of colon cancer.

[0025] A program according to a ninth aspect of the present invention is a program for causing a computer to execute the following steps: an acquisition step of acquiring a pathological tissue image of a patient with a target disease; a division step of dividing the pathological tissue image of the patient into a plurality of region images; a feature prediction step of inputting a region image into each of a plurality of feature prediction models constructed for each type of pathological histological feature, and acquiring prediction information on the presence or absence of the histological feature, respectively; a selection step of selecting a plurality of region images in which the acquired combinations of the presence / absence of the histological features match a predetermined combination of the presence / absence of the histological features at the time of selection; a gene mutation prediction step of inputting the region image selected by the selection into each of the combinations of the presence / absence of the histological features and the gene mutation prediction models constructed for each type of gene mutation, and acquiring prediction information on the presence / absence of the gene mutation, respectively; and a prediction result output step of outputting a prediction result of the presence / absence of at least one gene mutation for the patient, using the prediction information on the presence / absence of the gene mutation for each acquired region image.

[0026] According to this configuration, by inputting a pathological tissue image of a patient with a target disease, the prediction result of the presence or absence of a gene mutation can be obtained immediately. This allows a clinician to refer to the prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of the target disease.

[0027] A program according to a tenth aspect of the present invention is a program for predicting genes in which mutations have occurred in colon cancer tissue of a colon cancer patient, the program including a computer having the following steps: acquiring a colon cancer pathological tissue image of the patient; dividing the colon cancer pathological tissue image of the patient into a plurality of region images; a feature prediction step of inputting the region images into a plurality of feature prediction models constructed for each type of pathological histological feature, and acquiring prediction information on the presence or absence of the histopathological feature, respectively; a selection step of selecting a plurality of region images in which the acquired combinations of the presence or absence of the histopathological features match a combination of the presence or absence of the histopathological features at the time of selection that is preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; a gene mutation prediction step of inputting the selected region images into each of the combinations of the presence or absence of the histopathological features and the gene mutation prediction models constructed for each type of gene mutation, and acquiring prediction information on the presence or absence of the gene mutation, respectively; A prediction result output procedure for outputting a prediction result of the presence or absence of at least one gene mutation among V600E, ERBB2, RAS, TP53, or MSI, and a program for executing the procedure.

[0028] According to this configuration, by inputting the colon cancer pathology tissue image of a colon cancer patient, the prediction result of the presence or absence of gene mutation can be obtained immediately. This allows a clinician to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of colon cancer.

[0029] An information processing system according to an eleventh aspect of the present invention is an information processing system for estimating the presence or absence of a BRAF gene mutation in a tumor, comprising: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image having a tumor cell ratio of more than 50%, and / or an image including a papillary structure, and / or an image including a non-serrated papillary structure, and / or an image including a non-serrated papillary structure and in which mucus is present, and / or an image including a cribriform structure, and / or an image including a cribriform structure and in which mucus is absent, and / or an image including a cribriform structure and in which mucus is present; and a means for estimating the presence or absence of a BRAF gene mutation using the selected images.

[0030] According to this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of a BRAF gene mutation can be obtained immediately. This allows clinicians to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a BRAF gene mutation for the patient, thereby reducing delays in the administration of drugs corresponding to the BRAF gene abnormality of cancer.

[0031] An information processing system according to a twelfth aspect of the present invention is an information processing system for estimating the presence or absence of a BRAF V600E gene mutation in a tumor, comprising: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image including a track pattern, and / or an image including small solid nests, and / or an image in which small solid nests are composed of normal cells, and / or an image in which large solid nests are composed of normal cells, and / or an image including an oblong nucleus, and / or an image in which mucus is present, and / or an image in which a non-serrated papillary structure is not present and / or an image in which a cribriform structure is not present and / or an image in which a cribriform structure is not present and / or an image in which a mucus is present; and a means for estimating the presence or absence of a BRAF V600E gene mutation using the selected images.

[0032] According to this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of the BRAF V600E gene mutation can be obtained immediately. This allows a clinician to refer to the prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of the BRAF V600E gene mutation for the patient, thereby reducing delays in the administration of drugs corresponding to the BRAF V600E gene abnormality of cancer.

[0033] An information processing system according to a thirteenth aspect of the present invention is an information processing system for estimating the presence or absence of an ERBB2 gene mutation in a tumor, comprising: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image containing a cord-like structure and the presence of mucus, and / or an image containing a cord-like structure and mucus leakage; and a means for estimating the presence or absence of an ERBB2 gene mutation by analyzing the selected images.

[0034] According to this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of ERBB2 gene mutation can be obtained immediately. This allows a clinician to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of ERBB2 gene mutation for the patient, thereby reducing delays in administration of drugs corresponding to the ERBB2 gene abnormality of cancer.

[0035] An information processing system according to a fourteenth aspect of the present invention is an information processing system for estimating the presence or absence of a TP53 gene mutation in a tumor, comprising: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image containing signet ring cells and having mucus leakage, and / or an image containing goblet cells, and / or an image containing cord-like structures and no mucus, and / or an image containing mucus leakage; and a means for estimating the presence or absence of a TP53 gene mutation using the selected images.

[0036] According to this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of a TP53 gene mutation can be obtained immediately. This allows clinicians to refer to the prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a TP53 gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the TP53 gene abnormality of cancer.

[0037] An information processing system according to a fifteenth aspect of the present invention is an information processing system for estimating the presence or absence of MSI gene abnormality in a tumor, comprising: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image including a track pattern, and / or an image including a cord-like structure and no mucus, and / or an image with a tumor cell ratio of more than 50%, and / or an image including solid nests, and / or an image including small solid nests and normal cells, and / or an image including an oblong nucleus, and / or an image including a papillary structure, and / or an image including goblet cells, and / or an image including a non-serrated papillary structure, and / or an image including an oval nucleus, and / or an image including a cribriform structure and no mucus, and / or an image including a cribriform structure and mucus present, and / or an image including a cribriform structure and mucus leakage; and a means for estimating the presence or absence of MSI gene abnormality using the selected images.

[0038] According to this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of MSI gene abnormality can be obtained immediately. This allows a clinician to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of MSI gene abnormality for the patient, thereby reducing delays in the administration of drugs corresponding to the MSI gene abnormality of cancer.

[0039] An information processing system according to a sixteenth aspect of the present invention is an information processing system for estimating the presence or absence of a RAS gene mutation in a tumor, comprising: a means for dividing a pathological tissue image of the tumor into one or more images; and selecting from the divided images an image containing signet ring cells and mucus leakage, and / or an image containing tubular structures and no mucus, and / or an image containing cord-like structures and no mucus, and / or an image containing small solid nests, and / or an image containing large solid nests, and / or an image in which large solid nests are composed of normal cells, and / or are images containing papillary structures, and / or images in which mucus is absent, and / or images in which mucus is present, and / or images containing goblet cells, and / or images containing non-serrated papillary structures, and / or images containing non-serrated papillary structures and absence of mucus, and / or images containing non-serrated papillary structures and presence of mucus, and / or images containing cribriform structures, and / or images containing cribriform structures and absence of mucus, and / or images containing cribriform structures and presence of mucus, and / or images containing cribriform structures and mucus leakage, and / or images containing budding Images, and / or images with a tumor cell ratio of more than 50%, and / or images containing small solid nests, and / or images in which small solid nests are composed of normal cells, and / or images containing large solid nests, and / or images in which large solid nests are composed of normal cells, and / or images containing oblong nuclei, and / or images in which mucus is present, and / or images in which mucus leakage is present, and / or images in which non-serrated papillary structures and no mucus is present, and / or images in which cribriform structures and / or images in which cribriform structures and no mucus is present, and / or and selecting images including cribriform structures and the presence of mucus, and / or images including tubular structures and the presence of mucus, and / or images including tubular structures and mucus leakage, and / or images including striated patterns, and / or images including high grade cellular atypia, and / or images including cord-like structures and the presence of mucus, and / or images including cord-like structures and mucus leakage, and / or images including solid nests, and / or images including non-serrated papillary structures, and / or images including non-serrated papillary structures and the presence of mucus, and / or images including round nuclei;and a means for estimating the presence or absence of a RAS gene mutation using the selected image.

[0040] With this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of RAS gene mutation can be obtained immediately. This allows clinicians to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of RAS gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the RAS gene abnormality of cancer. Effect of the Invention

[0041] According to one aspect of the present invention, by inputting a pathological tissue image of a patient with a target disease, a prediction result of the presence or absence of a gene mutation can be obtained immediately. This allows a clinician to refer to the prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby reducing delays in administration of drugs corresponding to the genetic abnormality of the target disease. According to another aspect of the present invention, by inputting a pathological tissue image of colon cancer in a colon cancer patient, the prediction result of the presence or absence of a gene mutation can be obtained immediately. This allows a clinician to refer to the prediction result and prescribe a medicine corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby preventing delays in administration of medicines corresponding to the genetic abnormality of colon cancer. [Brief description of the drawings]

[0042] [Figure 1] 1 is a schematic configuration diagram of an information processing system according to an embodiment of the present invention. [Diagram 2] FIG. 1 is a schematic configuration diagram of a computer system according to an embodiment of the present invention. [Diagram 3] FIG. 1 is a schematic diagram showing the learning process of a gene mutation prediction model. [Figure 4A] FIG. 1 is a schematic diagram for explaining the process of annotating histopathological features to a histopathological image. [Figure 4B] FIG. 1 is a schematic diagram showing an example of a process for annotating histopathological features to a tissue image. [Diagram 5] FIG. 1 is a schematic diagram illustrating learning and testing when constructing a feature prediction model for a certain histopathological feature. [Figure 6A] FIG. 13 is a schematic diagram for explaining the assignment of predicted values ​​of histopathological features to a region image. [Figure 6B] FIG. 13 is a schematic diagram illustrating an example of assigning predicted values ​​of histopathological features to a region image. [Figure 7] FIG. 1 is a schematic diagram for explaining learning and testing when constructing a gene mutation prediction model for a certain gene mutation. [Figure 8] 1 is an example of testing a gene mutation prediction model for colorectal cancer. [Figure 9] In cases where the target gene mutation is BRAF+, this is a set of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection. [Figure 10] In cases where the target gene mutation is BRAF V600E+, this is a set of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection. [Figure 11] In cases where the target gene mutation is ERBB2+, this is a set of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection. [Figure 12] In cases where the target gene mutation is TP53+, this is a set of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection. [Figure 13] In cases where the genetic abnormality is MSI, this is a set of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection. [Figure 14] In cases where the target gene mutation is RAS+, this is a set of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection. [Figure 15] This is a continuation of Figure 14. [Figure 16] 2 is an example of a table stored in the storage 23. [Figure 17] This is a continuation of Figure 16. [Figure 18]FIG. 1 is a schematic diagram showing the process of estimating the presence or absence of each gene mutation. [Figure 19] FIG. 13 is a schematic diagram for explaining a method for assembling prediction results for each patient from prediction results for a regional image. [Figure 20] 10A to 10C are schematic diagrams for explaining an example of a screen transition on a terminal. [Figure 21] 21 is a flowchart showing an example of the process flow of the screen transition in FIG. 20. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0043] Hereinafter, each embodiment will be described with reference to the drawings. However, more detailed description than necessary may be omitted. For example, detailed description of already well-known matters or duplicate description of substantially the same configuration may be omitted. This is to avoid the following description becoming unnecessarily redundant and to facilitate understanding by those skilled in the art.

[0044] The target disease for predicting the presence or absence of a genetic mutation according to this embodiment can be any disease involving genetic mutation in tissues, including cancer (or any disease for which a drug can be administered according to a genetic abnormality), but the target disease according to this embodiment will be described below as colon cancer, as an example. Note that in this embodiment, genetic abnormalities are described as being included in genetic mutations.

[0045] Fig. 1 is a schematic configuration diagram of an information processing system according to this embodiment. As shown in Fig. 1, the information processing system S includes terminals 1-1 to 1-N (N is a natural number) and a computer system 2 connected via a communication network NW.

[0046] The terminals 1-1 to 1-N are used by different users, and may be, for example, mobile phones such as multi-function mobile phones (so-called smartphones), tablets, notebook computers, or desktop computers. The terminals 1-1 to 1-N may display information transmitted from the computer system 2 via, for example, a WEB browser, or may display information transmitted from the computer system 2 on the screen of an application installed on the terminals 1-1 to 1-N. Here, as an example, the following description will be given assuming that the terminals 1-1 to 1-N display information transmitted from the computer system 2 via, for example, a WEB browser.

[0047] The computer system 2 is capable of communicating with the terminals 1-1 to 1-N, and these communications may be wired or wireless. That is, the computer system 2 is connected to the terminals 1-1 to 1-N so as to be able to exchange information with them. The computer system 2 is used, for example, by an administrator who manages the information processing system S according to this embodiment. The computer system 2 may be a single computer or multiple computers.

[0048] 2 is a schematic diagram of a computer system according to this embodiment. As shown in FIG. 2, the computer system 2 includes an input interface 21, a communication module 22, a storage 23, a memory 24, and a processor 25. The input interface 21 receives an input from an administrator of the computer system 2, and outputs to the processor 25 an input signal corresponding to the received input. The communication module 22 is connected to the communication network NW and communicates with the terminals 1-1 to 1-N. This communication may be wired or wireless, but the following description will assume that it is wired.

[0049] The storage 23 stores programs to be read and executed by the processor 25, feature prediction models after machine learning, mutation prediction models after machine learning, and various data. The memory 24 temporarily holds data and programs and is a volatile memory, such as a random access memory (RAM). The processor 25 loads a program from the storage 23 into the memory 24 and executes a series of instructions included in the program, thereby functioning as an acquisition unit 251 and an output unit 250. The output unit 250 has, for example, a division unit 252, a feature prediction unit 253, a selection unit 254, a gene mutation prediction unit 255, and a prediction result output unit 256. Each process will be described later.

[0050] Next, the learning process of the gene mutation prediction model will be described with reference to Fig. 3. Fig. 3 is a schematic diagram showing the learning process of the gene mutation prediction model. (Step S1) The division unit 252 divides an image (e.g., a slide image) of a pathological tissue (here, colon cancer tissue as an example) of a patient with a target disease (here, colon cancer as an example) into one or more region images (e.g., tile images in a tiled shape). Here, the genetic mutation of the diseased tissue (here, colon cancer tissue as an example) of the patient is known. Here, division into tile images in a tiled shape is also called tile division.

[0051] (Step S2) Next, using, for example, an annotation system, the pathologist inputs histopathological features (e.g., oblong nuclei, etc.) for each of the multiple area images. At that time, a terminal device (not shown) used by the pathologist accepts the histopathological features. The acquisition unit 251 acquires the histopathological features assigned to each of the multiple area images.

[0052] (Step S3) Next, the processor 25 learns the relationship between the region image and the histopathological feature for each histopathological feature, and outputs a feature prediction model. Specifically, for example, the processor 25 learns by machine learning (e.g., deep learning) using learning data in which the region image obtained by dividing the pathological tissue image is input and the histopathological feature assigned to the region image is output, and outputs a feature prediction model. As a result, feature prediction models are output as many as the number L (L is a natural number) of histopathological features, and these feature prediction models FM-1, ..., FM-L are stored in the storage 23 by the processor 25. In this way, the feature prediction model is a model that has been machine-learned using learning data in which the region image obtained by dividing the pathological tissue image is input and the histopathological feature assigned to the region image is output.

[0053] (Step S4) Next, processor 25 predicts the presence or absence of each histopathological feature for divided images (tile images as an example here) that have no annotation for the histopathological feature by using feature prediction models FM-1, ..., FM-L. This allows the presence or absence of each histopathological feature to be predicted for divided images that have no annotation for the histopathological feature.

[0054] (Step S5) Next, the selection unit 254 selects region images (here, tile images as an example) from the multiple region images using the combination of the presence and absence of the predicted histopathological features. Specifically, the selection unit 254 uses the histopathological features predicted in step S4 to extract region images (here, tile images as an example) corresponding to a specific combination of the presence and absence of the histopathological features.

[0055] (Step S6) Next, the processor 25 learns the relationship between a group of tile images corresponding to a specific combination of the presence or absence of histopathological features and gene mutations, and outputs a gene mutation prediction model. Specifically, for example, the processor 25 learns by machine learning (for example, deep learning) using learning data that receives as an input an area image (area image extracted in step S5) corresponding to a specific combination of the presence or absence of histopathological features and outputs information on the presence or absence of a specific gene mutation, and outputs a gene mutation prediction model. As a result, gene mutation prediction models are output as many as the number M of gene mutations (M is a natural number), and these gene mutation prediction models GM-1, ..., GM-M are stored in the storage 23 by the processor 25. In this way, the gene mutation prediction model is a model that has been machine-learned using learning data that receives as an input an area image selected using a specific combination of the presence or absence of histopathological features and uses as an output the information on the specific gene mutation.

[0056] Next, the process of annotating histopathological features to a pathological tissue image will be described with reference to Fig. 4A and Fig. 4B. Fig. 4A is a schematic diagram for explaining the process of annotating histopathological features to a pathological tissue image. As shown in Fig. 4A, a pathological tissue image is divided into a plurality of region images. Of the generated region images, region images outside cellular tissue (hereinafter also referred to as white images) and region images in which objects other than cellular tissue (e.g., magic marker ink) occupy more than a standard amount are excluded. This process is called region determination. Histopathological features are manually annotated by a person (e.g., a pathologist) to the region images that have not been excluded.

[0057] FIG. 4B is a schematic diagram showing an example of a process of annotating a pathological tissue image with histopathological features. As shown in FIG. 4B, a pathological tissue image is divided into a plurality of region images. From each region image, region images (also called marker images) in which white images and marker ink occupy more than the standard are excluded. For each region image that is not excluded, the presence or absence of each of the histopathological features is given. In FIG. 4B, if a histopathological feature is present, it is indicated by a circle, and if a histopathological feature is not present, it is indicated by an x, and for each region image, the presence or absence of each of the L histopathological features is given.

[0058] Next, learning and testing when constructing a feature prediction model for a certain histopathological feature will be described with reference to FIG. 5. FIG. 5 is a schematic diagram for explaining learning and testing when constructing a feature prediction model for a certain histopathological feature. As shown in FIG. 5, learning data is a set of a region image and the presence or absence of a certain histopathological feature, and the region image is given as the input of the machine learning model, and the presence or absence of the certain feature is given as the output of the machine learning model, and machine learning is performed. In this way, the relationship between the region image and the presence or absence of a certain histopathological feature is learned, and a feature prediction model that predicts the certain histopathological feature is generated.

[0059] As an example, during training, 5-fold cross validation is performed on 80% of the total training data. In other words, training is performed using 64% of the total training data, and validation is performed using 16% of the total training data. Specifically, validation is performed as follows: First, 80% of the total training data is divided into 5. For ease of explanation, let the subsets of data resulting from this division be s1, s2, s3, s4, and s5. For example, after dividing 80% of the learning data into 5 parts, we first use the subsets s1, s2, s3, and s4 as the training subsets to learn the model. Next, we use s5 as the validation subset to validate the model. The evaluation index (e.g. accuracy or F1 score) obtained at this time is e1. Next, learning is performed using s2, s3, s4, and s5, and evaluation is performed using s1. In the same manner, learning and evaluation are repeated while switching between subsets. By repeating learning and evaluation for all combinations, five evaluation indices are obtained. Then, the model with the best performance among these five evaluation indices is determined as the feature prediction model.

[0060] To verify the accuracy of the obtained feature prediction model, a blind test is conducted on 20% of the entire training data for the model with the best performance in 5-fold cross-validation. This verifies the performance of the determined feature prediction model, and if the performance meets the standards, this feature prediction model is adopted.

[0061] Next, the assignment of predicted values ​​of histopathological features to region images, which is executed in step S4 in Fig. 3, will be described with reference to Fig. 6A and Fig. 6B. Fig. 6A is a schematic diagram for explaining the assignment of predicted values ​​of histopathological features to region images. Fig. 6A shows a feature prediction model FM-1 for histopathological feature 1, a feature prediction model FM-2 for histopathological feature 2, a feature prediction model FM-3 for histopathological feature 3, ..., a feature prediction model FM-L for histopathological feature L. The predicted value of the histopathological feature is determined for each region image based on the prediction results by the feature prediction models FM-1, ..., FM-L, and is the degree to which the region image has the histopathological feature. The predicted value of the histopathological feature may be the prediction result itself (e.g., a numerical value between 0 and 1), or may be a value (e.g., a value of 0 or 1) according to the comparison result obtained by comparing the prediction result with a threshold value. Here, as an example, the predicted value of the histopathological feature is described as being a numerical value between 0 and 1.

[0062] FIG. 6B is a schematic diagram for explaining an example of assigning predicted values ​​of histopathological features to region images. As shown by predicted feature values ​​p11, p12, p13, ..., p1L, ..., p71, p72, p73, ..., p7L in FIG. 6B, predicted feature values ​​of histopathological features 1 to L are assigned to each region image from which white images and magic marker images are excluded. In this case, in step S5 of FIG. 3, the processor 25 may determine that the histopathological feature is present when the predicted feature value exceeds a threshold value, and that the histopathological feature is not present when the predicted feature value is equal to or less than the threshold value. The threshold value may be set to a different value for each histopathological feature, or may be set to the same value, but here, as an example, a description will be given assuming that a different value is set for each histopathological feature.

[0063] Next, the learning and testing performed in constructing a gene mutation prediction model for a certain gene mutation in step S6 of FIG. 3 will be described with reference to FIG. 7. FIG. 7 is a schematic diagram for explaining the learning and testing performed in constructing a gene mutation prediction model for a certain gene mutation. As shown in FIG. 7, the learning data is a set of the region image after selection and the presence or absence of a certain gene mutation, and the region image after selection is given as the input of the machine learning model, and the presence or absence of the certain gene mutation is given as the output of the machine learning model, and machine learning is performed. In this way, the relationship between the region image and the presence or absence of a certain gene mutation is learned, and a gene mutation prediction model that predicts the certain gene mutation is generated.

[0064] As an example, during training, 5-fold cross validation is performed on 80% of the total training data. In other words, training is performed using 64% of the total training data, and validation is performed using 16% of the total training data. Specifically, validation is performed as follows: First, 80% of the total training data is divided into 5. For ease of explanation, let the subsets of data resulting from this division be s1, s2, s3, s4, and s5. For example, after dividing 80% of the learning data into 5 parts, we first use the subsets s1, s2, s3, and s4 as the training subsets to learn the model. Next, we use s5 as the validation subset to validate the model. The evaluation index (e.g. accuracy or F1 score) obtained at this time is e1. Next, learning is performed using s2, s3, s4, and s5, and evaluation is performed using s1. In the same manner, learning and evaluation are repeated while switching between subsets. By repeating learning and evaluation for all combinations, five evaluation indices are obtained. Then, the model with the best performance among these five evaluation indices is determined as the gene mutation prediction model.

[0065] To verify the accuracy of the obtained gene mutation prediction model, a blind test is conducted on 20% of the total training data for the model with the best performance in 5-fold cross-validation. This verifies the performance of the determined gene mutation prediction model, and if the performance meets the standards, this gene mutation prediction model is adopted.

[0066] Using Figures 8 to 17, we will explain how to set the ``combination of the presence or absence of pathological histological features at the time of selection'' when selecting regional images during the process of predicting genetic mutations in colorectal cancer patients with unknown genetic mutations. FIG. 8 shows an example of testing a gene mutation prediction model for colorectal cancer. By performing testing as shown in FIG. 8, a "combination of the presence or absence of histopathological features at the time of selection" is obtained. Specifically, a feature prediction model and a gene mutation prediction model are trained using 80% of the data from stage 2. The trained feature prediction model and the trained gene mutation prediction model are applied to stage 1 data, stage 2.5 data, and TCGA data. The TCGA data here is limited to colon adenocarcinoma (COAD) and rectum adenocarcinoma (READ) among colorectal cancers.

[0067] <Combinations of characteristics that can predict gene mutations (or genetic abnormalities)> As one example, for each type of gene mutation (or gene abnormality) of interest, a combination of a first characteristic (site of the primary lesion) and a second characteristic (histopathological characteristic) that meets all of the following conditions (1) to (3) was selected from all experimental data. (1) In the two-stage test set, either case assembly method 1 or case assembly method 2 shows an AUC (area under the curve) of 0.8 or more (i.e., highly accurate prediction is possible with the cohort used for learning). Here, AUC is the area (integral) under the ROC curve with the first axis being the false positive rate and the second axis being the true positive rate. This area can range from 0 to 1. (2) In any test set other than the second period, either case assembly method 1 or case assembly method 2 exhibits an AUC of 0.8 or higher (i.e., highly accurate predictions are possible even for cohorts not used for learning). (3) The number of tiles required is the smallest among those with the same combination of image type and features.

[0068] Here, case assembly method 1 or case assembly method 2 is as follows. (1) Case Construction Method 1 A method in which, if the average value of the predicted value of the gene mutation model for each region image included in the region image group after selection (for example, the region image group in FIG. 19) is equal to or greater than a threshold th, it is predicted that case 1 has a gene mutation, and if not, it is predicted that there is no gene mutation. (2) Case Construction Method 2 A method for predicting that a target patient has a gene mutation when the ratio of the number of tiles in which the predicted value of a gene mutation model for each region image included in a region image group after selection (e.g., the region image group of FIG. 19 ) is equal to or greater than a first threshold value th1 to the total number of region images in the region image group after selection (e.g., the region image group of FIG. 19 ) is equal to or greater than a second threshold value th2, and predicting that a target patient does not have a gene mutation when this is not the case.

[0069] Figure 9 shows combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the target gene mutation is BRAF+. Here, BRAF+ refers to all BRAF mutations, including BRAF V600E.

[0070] FIG. 10 shows a set of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the target gene mutation is BRAF V600E+. FIG. 11 shows a set of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the target gene mutation is ERBB2+. FIG. 12 shows a set of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the target gene mutation is TP53+. FIG. 13 shows a set of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the state of genetic abnormality is MSI. Here, MSI (microsatellite instability) refers to a certain genetic abnormality state, not a gene. FIG. 14 shows a set of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the target gene mutation is RAS+. FIG. 15 is a continuation of FIG. 14. Here, RAS+ means KRAS+ or NRAS+.

[0071] 9 to 15 show the site of the primary focus, combinations of the presence or absence of histopathological features at the time of selection, the AUC of case composition method 1 and the AUC of case composition method 2 for the 2-period test data, the AUC of case composition method 1 and the AUC of case composition method 2 for the 1-period data, the AUC of case composition method 1 and the AUC of case composition method 2 for the 2.5-period data, and the AUC of case composition method 1 and the AUC of case composition method 2 for the TCGA data. "Small solid nests and normal cells" in the column "Combinations of the presence or absence of histopathological features at the time of selection" in Figs. 10, 13, 14, 15, and Figs. 16 and 17 described below means that small solid nests are composed of normal cells (i.e., normal cells form small solid nests). Similarly, "Large solid nests and normal cells" in the column "Combination of the presence or absence of histopathological characteristics at the time of selection" in Figures 10, 14, 15, and Figures 16 and 17 described below means that the large solid nests are composed of normal cells (i.e., normal cells form the large solid nests).

[0072] In Figures 9 to 17, "XX structure (or cells) and mucus △△" means that there is a feature of XX structure (or cells) and mucus △△ at the same time. Specifically, "sieve structure and absence of mucus" in Figures 9, 10, 13, 14, 15, 16, and 17 means "sieve structure and absence of mucus at the same time." Similarly, "sieve structure and presence of mucus" in Figures 9, 10, 13, 14, 15, 16, and 17 means "sieve structure and presence of mucus at the same time." Similarly, "non-serrated papillary structure and absence of mucus" in Figures 10, 14, 15, 16, and 17 means "non-serrated papillary structure and absence of mucus at the same time." Similarly, "cord-like structures and mucus are present" in Figures 11, 15, 16, and 17 means "cord-like structures and mucus are present at the same time." Similarly, "cord-like structures and mucus leakage" in Figures 11, 15, 16, and 17 means "cord-like structures and mucus leakage at the same time." Similarly, "signet ring cells and mucus leakage" in Figures 12 and 16 means "signet ring cells and mucus leakage at the same time." Similarly, "cord-like structures and no mucus" in Figures 12, 13, and 16 means "cord-like structures and no mucus at the same time." Similarly, "sieve-like structures and mucus leakage" in Figures 13, 14, 15, 16, and 17 means "sieve-like structures and mucus leakage at the same time." Similarly, "signet ring cells and mucus leakage" in Figures 14 and 17 means "signet ring cells and mucus leakage at the same time." Similarly, "tubular structure and absence of mucus" in Figures 14 and 17 means "tubular structure and absence of mucus at the same time." Similarly, "non-serrated papillary structure and presence of mucus" in Figures 14, 15, and 17 means "non-serrated papillary structure and presence of mucus at the same time." Similarly, "tubular structure and presence of mucus" in Figures 15 and 17 means "tubular structure and presence of mucus at the same time." Similarly, "tubular structure and mucus leakage" in Figures 15 and 17 means "tubular structure and mucus leakage at the same time."

[0073] FIG. 16 is an example of a table stored in the storage 23. FIG. 17 is a continuation of FIG. 16. In the BRAF table T1 of FIG. 16, records of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection are stored when the target gene mutation is BRAF+. In addition, in the BRAF V600E table T2 of FIG. 16, records of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection are stored when the target gene mutation is BRAF V600E.

[0074] In addition, the ERBB2 table T3 in Fig. 16 stores records of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the target gene mutation is ERBB2. In addition, the TP53 table T4 in Fig. 16 stores records of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the target gene mutation is TP53. In addition, the MSI table T5 in Fig. 16 stores records of combinations of the site of the selected primary tumor and the presence or absence of histopathological features at the time of selection when the genetic abnormality state is MSI.

[0075] Furthermore, in the RAS table T6 in FIG. 17, records of combinations of the site of the selected primary lesion and the presence or absence of histopathological characteristics at the time of selection are stored when the target gene mutation is RAS.

[0076] Next, the process of estimating the presence or absence of each gene mutation will be described with reference to Figure 18. Figure 18 is a schematic diagram showing the process of estimating the presence or absence of each gene mutation. (Step S11) First, the acquiring unit 251 acquires a pathological tissue image of a patient for whom genetic analysis has not been performed and whose genetic mutation is unknown.

[0077] (Step S12) Next, division unit 252 divides the pathological tissue image of the patient into a plurality of region images.

[0078] (Step S13) Next, the feature prediction unit 253 predicts the presence or absence of a histopathological feature of the target for each region image using the feature prediction model constructed for each type of histopathological feature. More specifically, for example, the feature prediction unit 253 inputs the region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and obtains prediction information on the presence or absence of the histopathological feature.

[0079] (Step S14) Next, the feature prediction unit 253 selects the plurality of region images using a combination of the presence and absence of the acquired histopathological features. More specifically, for example, the feature prediction unit 253 extracts, from the plurality of region images, a region image in which a combination of the presence and absence of the acquired histopathological features matches a specific combination of the presence and absence of the histopathological features set for each genetic abnormality. Here, for example, a specific combination of the presence or absence of histopathological features set for BRAF gene abnormality is, for example, one of the "combinations of the presence or absence of histopathological features at the time of selection" stored in the BRAF table T1 of FIG. 16. Similarly, for example, a specific combination of the presence or absence of histopathological features set for BRAF V600E gene abnormality is, for example, one of the "combinations of the presence or absence of histopathological features at the time of selection" stored in the BRAF V600E table T2 of FIG. 16. Similarly, for example, a specific combination of the presence or absence of histopathological features set for ERBB2 gene abnormality is, for example, one of the "combinations of the presence or absence of histopathological features at the time of selection" stored in the ERBB2 table T3 of FIG. Similarly, for example, a specific combination of the presence or absence of histopathological features set for TP53 gene abnormality may be one of the "combinations of the presence or absence of histopathological features at the time of selection" stored in the TP53 table T4 of FIG. 16, but there are multiple combinations. Similarly, a specific combination of the presence or absence of histopathological features set for, for example, an MSI gene abnormality is, for example, one of the "combinations of the presence or absence of histopathological features at the time of selection" stored in the MSI table T5 of Fig. 16. Similarly, a specific combination of the presence or absence of histopathological features set for, for example, an RAS gene abnormality is, for example, one of the "combinations of the presence or absence of histopathological features at the time of selection" stored in the RAS table T6 of Fig. 16.

[0080] (Step S15) Next, the gene mutation prediction unit 255 predicts the presence or absence of a target gene mutation for each region image using a gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and each type of gene mutation. More specifically, for example, the gene mutation prediction unit 255 inputs the region image selected by selection for each of the gene mutation prediction models constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and obtains prediction information on the presence or absence of a gene mutation.

[0081] (Step S16) Next, the prediction result output unit 256 predicts the presence or absence of each gene mutation in the patient using the prediction result for each region image.

[0082] (Step S17) Next, the prediction result output unit 256 outputs, for example, a list of prediction results for the presence or absence of each gene mutation. Note that the prediction results for the presence or absence of gene mutations do not have to be prediction results for the presence or absence of all gene mutations, and at least one or more will suffice. In this way, the prediction result output unit 256 uses prediction information for the presence or absence of gene mutation for each acquired region image to output a prediction result for the presence or absence of at least one gene mutation for the patient.

[0083] In the above example, a region image that matches the "combination of the presence or absence of histopathological features at the time of selection" is selected for the target gene mutation, the region image after selection is input into the gene mutation prediction model to obtain a gene mutation prediction result for each region image, and the presence or absence of the target gene mutation in the target patient is predicted using the prediction result for each region image, but the present invention is not limited to this. This series of steps may be performed for each "combination of the presence or absence of histopathological features at the time of selection". In this case, it may be possible to output all together information on which combination of "combination of the presence or absence of histopathological features at the time of selection" is predicted to have the target gene mutation. Furthermore, the above processing may be performed across multiple gene mutations, so that the above output may be performed across multiple gene mutations.

[0084] The acquiring unit 251 may further acquire the site of the primary focus of the target disease. In this case, the selecting unit 254 may select the plurality of region images using the acquired site of the primary focus of the target disease together with the acquired combination of the presence or absence of the histopathological features. For example, in the case of colon cancer, in order to predict the presence or absence of a BRAF gene mutation, the selection unit 254 may refer to the BRAF table T1 (see FIG. 16) stored in the storage 23, and select a region image that matches, for example, the "site of the primary focus" and the "combination of the presence or absence of histopathological features at the time of selection" for one record. For example, in the case of the first record (record in the first row), the selection unit 254 may select a region image in which the site of the primary focus is the "left part of the large intestine" and the combination of the presence or absence of histopathological features at the time of selection corresponds to "tumor cell ratio exceeds 50%". In this way, by selecting the region image using the site of the primary focus of the target disease as well, the probability that the region image input to the gene mutation prediction model can be limited to the region image related to the target disease is improved, and therefore the prediction accuracy of the presence or absence of gene mutation can be improved.

[0085] In this case, for example, for the target gene mutation, a region image that matches the combination of "site of primary focus" and "combination of the presence or absence of histopathological features at the time of selection" is selected, and the selected region image is input into the gene mutation prediction model to obtain a gene mutation prediction result for each region image, and the presence or absence of the target gene mutation of the target patient is predicted using the prediction result for each region image. This series of steps may be performed for each combination of "site of primary focus" and "combination of the presence or absence of histopathological features at the time of selection". In this case, it may be output together with which combination of "site of primary focus" and "combination of the presence or absence of histopathological features at the time of selection" the target gene mutation was predicted to be present. Furthermore, the above process may be performed across multiple gene mutations, so that the above output may be performed across multiple gene mutations.

[0086] Next, a method for assembling prediction results for each patient from prediction results for region images will be described with reference to Fig. 19. Fig. 19 is a schematic diagram for explaining a method for assembling prediction results for each patient from prediction results for region images. Fig. 19 shows a pair of image regions after selection derived from a target patient and a predicted value for each region image by a gene mutation prediction model. For example, when the number of tiles included in the selected region image group (for example, the region image group in FIG. 19 ) is K or more, the prediction result output unit 256 predicts the presence or absence of a genetic mutation at the case level by the above-mentioned case assembly method 1 or case assembly method 2. On the other hand, when the number of tiles is less than K, the prediction result output unit 256 does not perform a genetic mutation prediction at the case level (as a result, it is treated as if there is no genetic mutation).

[0087] Next, an example of screen transitions in the terminal 1 will be described with reference to FIG. 20. FIG. 20 is a schematic diagram for explaining an example of screen transitions in a terminal. As shown in FIG. 20, a text box TB1 for inputting a path of a pathological tissue image file to select a pathological tissue image of a patient and a reference button B1 for referring to the pathological tissue image file are provided on the screen G1 of the terminal 1. Also, a radio button B2 for selecting the position of the site of the primary focus of colon cancer is provided. When the send button B3 is pressed with the pathological tissue image of the patient selected and the position of the site of the primary focus of colon cancer selected, a transition to a screen G2 occurs. On the screen G2 after the transition, a list of prediction results of the presence or absence of gene mutation is displayed.

[0088] Next, the process flow of the screen transition in Fig. 20 will be described with reference to Fig. 21. Fig. 21 is a flowchart showing an example of the process flow of the screen transition in Fig. 20. (Step S110) First, the terminal 1 receives a pathological tissue image of a patient and the location of the primary focus of colon cancer.

[0089] (Step S120) Next, the terminal 1 transmits to the computer system 2 the pathological tissue image of the patient and the site of the primary focus of colon cancer.

[0090] (Step S130) Next, the computer system 2 receives the pathological tissue image of the patient and the site of the primary focus of colon cancer, and outputs information for displaying the prediction result of the presence or absence of each gene mutation using the pathological tissue image of the patient and the site of the primary focus of colon cancer. The details of this process have been described in FIG. 18, so the description thereof will be omitted.

[0091] (Step S140) Next, the computer system 2 transmits to the terminal 1 information for displaying the prediction results of the presence or absence of each gene mutation.

[0092] (Step S150) Next, the terminal 1 receives information for displaying the predicted result of the presence or absence of each gene mutation, and uses this information to display the predicted result of the presence or absence of each gene mutation. This ends the processing of this flowchart.

[0093] As described above, the information processing system S according to this embodiment includes an acquisition unit 251 that acquires a pathological tissue image of a patient with a target disease, a division unit 252 that divides the pathological tissue image of the patient into a plurality of region images, a feature prediction unit 253 that inputs a region image for each of a plurality of feature prediction models constructed for each type of pathological histological feature and acquires prediction information on the presence or absence of the histopathological feature, respectively, a selection unit 254 that selects a plurality of region images in which the acquired combinations of the presence or absence of the histopathological features match a combination of the presence or absence of the histopathological features at the time of selection that is set in advance, a gene mutation prediction unit 255 that inputs the region image selected by the selection for each combination of the presence or absence of the histopathological features and the gene mutation prediction model constructed for each type of gene mutation, and acquires prediction information on the presence or absence of the gene mutation, respectively, and a prediction result output unit 256 that outputs a prediction result of the presence or absence of at least one gene mutation for the patient using the prediction information on the presence or absence of the gene mutation for each acquired region image.

[0094] According to this configuration, by inputting a pathological tissue image of a patient with a target disease, the prediction result of the presence or absence of a gene mutation can be obtained immediately. This allows a clinician to refer to the prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of the target disease.

[0095] Moreover, the information processing system according to the present embodiment is, as an example, an information processing system for predicting genes in which mutations have occurred in colon cancer tissues of a colon cancer patient. The information processing system S includes an acquisition unit 251 that acquires a colon cancer pathology tissue image of the patient. The information processing system S further includes a division unit 252 that divides the colon cancer pathology tissue image of the patient into a plurality of region images. The information processing system S further includes a feature prediction unit 253 that inputs a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and acquires prediction information on the presence or absence of the histopathological feature. The information processing system S further includes a selection unit 254 that selects a plurality of region images in which the combination of the presence or absence of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection that is preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI. The information processing system S further includes a gene mutation prediction unit 255 that inputs the selected regional images for each gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and acquires prediction information on the presence or absence of gene mutation, and a prediction result output unit 256 that uses the acquired prediction information on the presence or absence of gene mutation for each regional image to output a prediction result on the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, and MSI for the patient.

[0096] According to this configuration, by inputting the colon cancer pathology tissue image of a colon cancer patient, the prediction result of the presence or absence of gene mutation can be obtained immediately. This allows a clinician to refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the genetic abnormality of colon cancer.

[0097] Moreover, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of a BRAF gene mutation in a tumor, and includes: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image in which the tumor cell ratio exceeds 50%, and / or an image including a papillary structure, and / or an image including a non-serrated papillary structure, and / or an image including a non-serrated papillary structure and in which mucus is present, and / or an image including a cribriform structure, and / or an image including a cribriform structure and in which mucus is absent, and / or an image including a cribriform structure and in which mucus is present; and a means for estimating the presence or absence of a BRAF gene mutation using the selected images.

[0098] With this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of a BRAF gene mutation can be obtained immediately. This allows clinicians to refer to this prediction result and prescribe a drug that corresponds to the prediction result of the presence or absence of a BRAF gene mutation for the patient, thereby reducing delays in administering drugs according to the BRAF gene abnormality of cancer.

[0099] Moreover, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of a BRAF V600E gene mutation in a tumor, and includes: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image including a track pattern, and / or an image including small solid nests, and / or an image in which small solid nests are composed of normal cells, and / or an image in which large solid nests are composed of normal cells, and / or an image including an oblong nucleus, and / or an image in which mucus is present, and / or an image in which a non-serrated papillary structure is not present and / or an image in which a cribriform structure is not present and / or an image in which a cribriform structure is not present and / or an image in which a mucus is present; and a means for estimating the presence or absence of a BRAF V600E gene mutation using the selected images.

[0100] With this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of the BRAF V600E gene mutation can be obtained immediately. This allows clinicians to refer to this prediction result and prescribe a drug that corresponds to the prediction result of the presence or absence of the BRAF V600E gene mutation for the patient, thereby reducing delays in administering drugs according to the BRAF V600E gene abnormality of cancer.

[0101] Furthermore, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of an ERBB2 gene mutation in a tumor, and includes a means for dividing a pathological tissue image of the tumor into one or more images, a means for selecting from the divided images an image that includes a cord-like structure and has mucus present, and / or an image that includes a cord-like structure and has mucus leakage, and a means for estimating the presence or absence of an ERBB2 gene mutation by analyzing the selected images.

[0102] With this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of ERBB2 gene mutation can be obtained immediately. This allows clinicians to refer to this prediction result and prescribe a drug that corresponds to the prediction result of the presence or absence of ERBB2 gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the ERBB2 gene abnormality of cancer.

[0103] Moreover, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of a TP53 gene mutation in a tumor, and includes: a means for dividing a pathological tissue image of the tumor into one or more images; a means for selecting, from the divided images, an image containing signet ring cells and having mucus leakage, and / or an image containing goblet cells, and / or an image containing cord-like structures and no mucus, and / or an image containing mucus leakage; and a means for estimating the presence or absence of a TP53 gene mutation using the selected images.

[0104] With this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of a TP53 gene mutation can be obtained immediately. This allows clinicians to refer to this prediction result and prescribe a drug that corresponds to the prediction result of the presence or absence of a TP53 gene mutation for the patient, thereby reducing delays in administering drugs corresponding to the TP53 gene abnormality of cancer.

[0105] Furthermore, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of MSI gene abnormalities in a tumor, and includes a means for dividing a pathological tissue image of the tumor into one or more images, a means for selecting from the divided images an image including a track pattern, and / or an image including a cord-like structure and no mucus, and / or an image with a tumor cell ratio of more than 50%, and / or an image including solid nests, and / or an image in which small solid nests are composed of normal cells, and / or an image including an oblong nucleus, and / or an image including a papillary structure, and / or an image including goblet cells, and / or an image including a non-serrated papillary structure, and / or an image including an oval nucleus, and / or an image including a cribriform structure and no mucus, and / or an image including a cribriform structure and mucus present, and / or an image including a cribriform structure and mucus leakage, and a means for estimating the presence or absence of MSI gene abnormalities using the selected images.

[0106] With this configuration, by inputting a pathological tissue image, it is possible to immediately obtain the prediction result of the presence or absence of MSI gene abnormalities. This allows clinicians to refer to this prediction result and prescribe a drug for the patient that corresponds to the prediction result of the presence or absence of MSI gene abnormalities, thereby reducing delays in the administration of drugs corresponding to the MSI gene abnormalities of cancer.

[0107] Furthermore, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of a RAS gene mutation in a tumor, and includes a means for dividing a pathological tissue image of the tumor into one or more images, and selecting from the divided images an image containing signet ring cells and mucus leakage, and / or an image containing tubular structures and no mucus, and / or an image containing cord-like structures and no mucus, and / or an image containing small solid nests, and / or an image containing large solid nests, and / or an image in which large solid nests are composed of normal cells, and / or an image containing papillary structures. images containing sarcoid structures, and / or images in which mucus is absent, and / or images in which mucus is present, and / or images containing goblet cells, and / or images containing non-serrated papillary structures, and / or images containing non-serrated papillary structures and the absence of mucus, and / or images containing non-serrated papillary structures and the presence of mucus, and / or images containing cribriform structures, and / or images containing cribriform structures and the absence of mucus, and / or images containing cribriform structures and the presence of mucus, and / or images containing cribriform structures and mucus leakage, and / or images in which budding is present, and / or an image with a tumor cell ratio of more than 50%, and / or an image containing small solid nests, and / or an image in which small solid nests are composed of normal cells, and / or an image containing large solid nests, and / or an image containing oblong nuclei, and / or an image containing mucus, and / or an image containing mucus leakage, and / or an image containing non-serrated papillary structures and no mucus, and / or an image containing cribriform structures, and / or an image containing cribriform structures and no mucus, and / or an image containing cribriform structures and the presence of mucus, and / or an image containing tubular structures and The device comprises a means for selecting an image in which mucus is present, and / or an image including tubular structures and mucus leakage, and / or an image including a track pattern, and / or an image including high-grade cellular atypia, and / or an image including cord-like structures and mucus is present, and / or an image including cord-like structures and mucus leakage, and / or an image including solid nests, and / or an image including non-serrated papillary structures, and / or an image including non-serrated papillary structures and mucus is present, and / or an image including round-shaped nuclei, and a means for estimating the presence or absence of a RAS gene mutation using the selected images.

[0108] With this configuration, by inputting a pathological tissue image, the prediction result of the presence or absence of RAS gene mutation can be obtained immediately. This allows clinicians to refer to this prediction result and prescribe medicines for the patient that correspond to the prediction result of the presence or absence of RAS gene mutation, thereby reducing delays in administering medicines corresponding to the RAS gene abnormality of cancer.

[0109] At least a part of the computer system 2 described in the above embodiment may be configured with hardware or software. When configured with software, a program that realizes at least a part of the functions of the computer system 2 may be stored in a recording medium such as a flexible disk or a CD-ROM, and may be read and executed by a computer. The recording medium is not limited to removable ones such as a magnetic disk or an optical disk, but may be a fixed recording medium such as a hard disk device or a memory.

[0110] Also, a program that realizes at least a part of the functions of the computer system 2 may be distributed via a communication line (including wireless communication) such as the Internet. Furthermore, the program may be encrypted, modulated, or compressed and distributed via a wired line or wireless line such as the Internet, or stored on a recording medium.

[0111] Furthermore, the computer system 2 may be operated by one or more information devices. When multiple information devices are used, one of the devices may be a computer, and the computer may execute a predetermined program to realize the functions as at least one of the means of the computer system 2.

[0112] In the method invention, all the steps may be realized by automatic control using a computer. Alternatively, each step may be performed by a computer while the progress between steps is controlled manually. Furthermore, at least some of the steps may be performed manually.

[0113] As described above, the present invention is not limited to the above-described embodiment as it is, and in the implementation stage, the components can be modified and embodied without departing from the gist of the invention. In addition, various inventions can be formed by appropriately combining the multiple components disclosed in the above-described embodiment. For example, some components may be deleted from all the components shown in the embodiment. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0114] 1 Device 2. Computer Systems 21 Input Interface 22 Communication Module 23. Storage 24 Memory 25 processors 250 Output section 251 Acquisition Department 252 Division 253 Feature Prediction Unit 254 Sorting Department 255 Gene Mutation Prediction Department 256 Prediction result output section

Claims

1. An acquisition unit for acquiring a pathological tissue image of a patient having a target disease; A division unit that divides the pathological tissue image of the patient into a plurality of region images; a feature prediction unit that inputs a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and obtains prediction information on the presence or absence of the histopathological feature; a selection unit that selects a plurality of region images in which a combination of the presence or absence of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection that is set in advance; a gene mutation prediction unit that inputs the region image selected by the selection for each gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and obtains prediction information on the presence or absence of gene mutation; a prediction result output unit that outputs a prediction result of the presence or absence of at least one gene mutation for the patient using the prediction information of the presence or absence of a gene mutation for each of the acquired region images; An information processing system comprising:

2. The acquisition unit further acquires a site of a primary focus of the target disease, The selection unit selects the plurality of regional images using a combination of the presence or absence of the acquired histopathological features and the acquired site of a primary focus of the target disease. The information processing system according to claim 1 .

3. the feature prediction model is a model machine-learned using learning data in which a region image obtained by dividing a pathological tissue image is input and histopathological features assigned to the region image are output, The gene mutation prediction model is a model machine-learned using learning data that uses an area image selected using a combination of the presence or absence of a pathological histological feature as an input and information on the presence or absence of a specific gene mutation as an output.

3. The information processing system according to claim 1 or 2.

4. An information processing system for predicting genes undergoing mutations in colorectal cancer tissue of a colorectal cancer patient, comprising: an acquisition unit for acquiring a colon cancer tissue pathology image of the patient; A division unit that divides the colon cancer pathology tissue image of the patient into a plurality of region images; a feature prediction unit that inputs a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and obtains prediction information on the presence or absence of the histopathological feature; A selection unit that selects a plurality of region images in which the combination of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection that is preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; a gene mutation prediction unit that inputs the selected region image to each of gene mutation prediction models constructed for each combination of the presence or absence of histopathological features and each type of gene mutation to obtain prediction information on the presence or absence of gene mutation; a prediction result output unit that outputs a prediction result of the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, and MSI for the patient using the prediction information of the presence or absence of gene mutation for each of the acquired regional images; An information processing system comprising:

5. The acquiring unit further acquires a site of a primary focus of colorectal cancer, The selection unit selects the plurality of region images using the determined at least one histopathological feature and the acquired site of the primary focus of colorectal cancer.

5. The information processing system according to claim 4.

6. the feature prediction model is a model machine-learned using learning data in which a region image obtained by dividing a pathological tissue image is used as an input and histopathological features assigned to the region image are used as an output, The gene mutation prediction model is a model that has been machine-trained using learning data that uses an area image selected using a combination of the presence or absence of a pathological histological feature as an input and information on the presence or absence of a specific gene mutation as an output.

6. The information processing system according to claim 4 or 5.

7. An acquisition procedure for acquiring a pathological tissue image of a patient with a target disease; a segmentation step of segmenting the pathological tissue image of the patient into a plurality of region images; a feature prediction step of inputting a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and obtaining prediction information on the presence or absence of the histopathological feature; a selection step of selecting a plurality of region images in which a combination of the presence or absence of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection set in advance; a gene mutation prediction procedure for inputting region images selected by the selection process for each gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and obtaining prediction information on the presence or absence of gene mutation; a prediction result output step of outputting a prediction result of the presence or absence of at least one gene mutation for the patient using the prediction information of the presence or absence of a gene mutation for each of the acquired regional images; An information processing method comprising the steps of:

8. An information processing method for predicting genes undergoing mutations in colorectal cancer tissue of a colorectal cancer patient, comprising: An acquisition step of acquiring a colon cancer tissue pathology image of the patient; A segmentation step of segmenting the colon cancer tissue pathology image of the patient into a plurality of region images; a feature prediction step of inputting a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and obtaining prediction information on the presence or absence of the histopathological feature; a selection step of selecting a plurality of region images in which the combination of the presence or absence of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection that is preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; a gene mutation prediction step of inputting the selected region image into each of gene mutation prediction models constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and obtaining prediction information on the presence or absence of gene mutation; a prediction result output step of outputting a prediction result of the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, and MSI for the patient, using the prediction information of the presence or absence of a gene mutation for each of the acquired regional images; An information processing method comprising the steps of:

9. On the computer, An acquisition procedure for acquiring a pathological tissue image of a patient with a target disease; a segmentation step of segmenting the pathological tissue image of the patient into a plurality of region images; a feature prediction step of inputting a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and obtaining prediction information on the presence or absence of the histopathological feature; a selection step of selecting a plurality of region images in which a combination of the presence or absence of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection set in advance; a gene mutation prediction procedure for inputting region images selected by the selection process for each gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and obtaining prediction information on the presence or absence of gene mutation; a prediction result output step of outputting a prediction result of the presence or absence of at least one gene mutation for the patient using the prediction information of the presence or absence of a gene mutation for each of the acquired regional images; A program for executing the above.

10. A program for predicting genes that are mutated in colorectal cancer tissue of a colorectal cancer patient, comprising: An acquisition step of acquiring a colon cancer tissue pathology image of the patient; A segmentation step of segmenting the colon cancer tissue pathology image of the patient into a plurality of region images; a feature prediction step of inputting a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature, and obtaining prediction information on the presence or absence of the histopathological feature; a selection step of selecting a plurality of region images in which the combination of the presence or absence of the acquired histopathological features matches a combination of the presence or absence of the histopathological features at the time of selection that is preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; a gene mutation prediction step of inputting the selected region image into each of gene mutation prediction models constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and obtaining prediction information on the presence or absence of gene mutation; a prediction result output step of outputting a prediction result of the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, and MSI for the patient, using the prediction information of the presence or absence of a gene mutation for each of the acquired regional images; A program for executing the above.

Citation Information

Patent Citations

  • Systems & Methods for Computational Pathology using Points-of-interest

    US20180232883A1

  • Image analysis device, image analysis method, image analysis system, image analysis program, and recording medium

    WO2017010397A1

  • System, program, and method for determining hypermutated tumor

    WO2019159821A1