Information processing system, information processing method, and program

The information processing system quickly predicts gene mutations in pathological tissue images using machine-learned models, addressing the delay in drug administration by enabling immediate treatment decisions.

JP2025108567AInactive Publication Date: 2025-07-23GENOMEDIA INC +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025065843
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The delay in administering drugs based on gene abnormalities, particularly in cancer patients, due to the time required for cancer gene tests, which can progress the disease condition and make treatment ineffective.

Method used

An information processing system that rapidly predicts the presence or absence of gene mutations by analyzing pathological tissue images using machine-learned feature and gene mutation prediction models, enabling immediate drug prescription.

Benefits of technology

Enables immediate prediction of gene mutations, allowing clinicians to promptly prescribe drugs, thereby reducing the delay in administering treatment based on genetic abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108567000001_ABST
    Figure 2025108567000001_ABST
Patent Text Reader

Abstract

To provide an information processing system, an information processing method, and a program for preventing a delay in administration of drugs according to gene abnormality in a target disease.SOLUTION: A system has: an acquisition unit that acquires a pathological tissue image of a patient with a target disease; a dividing unit that divides the pathological tissue image of the patient into a plurality of area images; a feature prediction unit that inputs the area images to a plurality of feature prediction models constituted for every type of histopathological features to acquire prediction information on the presence or absence of the histopathological features; a selection unit that selects an area image in which the acquired presence or absence of the histopathological features matches histopathological features during selection set in advance; a gene mutation prediction unit that inputs the area image selected through the selection to a gene mutation prediction model constituted for the presence or absence of the histopathological features and every type of the gene mutation, to acquire prediction information on the presence or absence of gene mutation; and a prediction result output unit that outputs a result of prediction of the presence or absence of at least one gene mutation for the patient by using the prediction information on the presence or absence of gene mutation for each of the acquired area images.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system, an information processing method, and a program.

Background Art

[0002] In cancers such as lung cancer, colorectal cancer, gastric cancer, breast cancer, GIST (gastrointestinal stromal tumor), and skin cancer (e.g., malignant melanoma), when a doctor determines it is necessary, a cancer gene test is performed to examine one or several genes for diagnosis or to select a drug for treatment based on the test results (see Non-Patent Document 1).

Prior Art Documents

Patent Documents

[0003] Non-Patent Document 1: https: / / ganjoho.jp / public / dia_tre / treatment / genomic_medicine / gentest02.html

Summary of the Invention

Problems to be Solved by the Invention

[0004] Since the results of cancer gene tests take 1 to 2 months, the administration of drugs according to the patient's gene abnormalities is delayed, and during that time, the disease condition may progress and it may become too late for treatment. In particular, for cancer patients in stage 4 etc., there is little time margin, which is a serious problem.

[0005] Also, not only for cancer, but also for diseases for which drugs can be administered according to gene abnormalities, there are similar problems.

[0006] One aspect of the present invention has been made in view of the above problems, and an object thereof is to provide an information processing system, an information processing method, and a program that can suppress the delay in the administration of pharmaceuticals according to gene abnormalities of a target disease.

[0007] Another aspect of the present invention has been made in view of the above problems, and an object thereof is to provide an information processing system, an information processing method, and a program capable of suppressing a delay in administration of a drug according to a gene abnormality of a patient with colorectal cancer.

[0008] Another aspect of the present invention has been made in view of the above problems, and an object thereof is to provide an information processing system capable of suppressing a delay in administration of a drug according to a gene abnormality of a cancer patient.

Means for Solving the Problems

[0009] The information processing system according to the first aspect of the present invention includes an acquisition unit that acquires a pathological tissue image of a tissue of a patient with a target disease, a division unit that divides the pathological tissue image of the patient into a plurality of region images, and for each of a plurality of feature prediction models constructed for each type of histological feature, an input region image is used to obtain prediction information on the presence or absence of histological features respectively, a selection unit that selects a plurality of region images whose combination of the presence or absence of the obtained histological features matches a preset combination of the presence or absence of histological features at the time of screening, and for each of the gene mutation prediction models constructed for each combination of the presence or absence of histological features and each type of gene mutation, an input region image selected by the selection is used to obtain prediction information on the presence or absence of gene mutations respectively, and a prediction result output unit that outputs a prediction result on the presence or absence of at least one gene mutation for the patient using the prediction information on the presence or absence of gene mutations for each of the obtained region images.

[0010] According to this configuration, if a pathological tissue image of a patient with a target disease is input, a prediction result on the presence or absence of a gene mutation can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result on the presence or absence of a gene mutation for the patient, so that a delay in administration of a drug according to a gene abnormality of a target disease can be suppressed.

[0011] The information processing system according to the second aspect of the present invention is the information processing system according to the first aspect, wherein the acquisition unit further acquires the site of the primary focus of the target disease, and the selection unit uses the acquired combination of the presence or absence of the pathological histological features and the acquired site of the primary focus of the target disease to select the plurality of region images.

[0012] According to this configuration, by selecting the region image using the site of the primary focus of the target disease as well, the probability that the region image input to the gene mutation prediction model can be limited to the region image related to the target disease is improved, so that the prediction accuracy of the presence or absence of gene mutation can be improved.

[0013] The information processing system according to the third aspect of the present invention is the information processing system according to the first or second aspect, wherein the feature prediction model is a model machine-learned using learning data that takes as input a region image obtained by dividing a pathological tissue image and uses as output the pathological histological features assigned to the region image, and the gene mutation prediction model is a model machine-learned using learning data that takes as input a region image selected using a combination of the presence or absence of pathological histological features and uses as output information on the presence or absence of a specific gene mutation.

[0014] According to this configuration, since the feature prediction model and the gene mutation prediction model are models after machine learning, the prediction accuracy of the presence or absence of gene mutation can be improved.

[0015] An information processing system according to a fourth aspect of the present invention is an information processing system for predicting genes in which mutations have occurred in colorectal cancer tissue of a patient with colorectal cancer, the system including: an acquisition unit that acquires a colorectal cancer pathological tissue image of the patient; a division unit that divides the colorectal cancer pathological tissue image of the patient into a plurality of region images; a feature prediction unit that inputs the region images to each of a plurality of feature prediction models constructed for each type of histopathological feature, and respectively acquires prediction information on the presence or absence of the histopathological feature; a selection unit that selects a plurality of region images whose combination of the acquired presence or absence of the histopathological feature matches a preset combination of the presence or absence of the histopathological feature at the time of screening for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI (microsatellite instability); a gene mutation prediction unit that inputs the selected region images to each of gene mutation prediction models constructed for each combination of the presence or absence of the histopathological feature and each type of gene mutation, and respectively acquires prediction information on the presence or absence of the gene mutation; and a prediction result output unit that outputs a prediction result on the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI for the patient, using the prediction information on the presence or absence of the gene mutation for each of the acquired region images.

[0016] According to this configuration, if a colorectal cancer pathological tissue image of a patient with colorectal cancer is input, a prediction result on the presence or absence of a gene mutation can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result on the presence or absence of the gene mutation for the patient, thereby suppressing a delay in the administration of a drug corresponding to the genetic abnormality of colorectal cancer.

[0017] An information processing system according to a fifth aspect of the present invention is the information processing system according to the fourth aspect, wherein the acquisition unit further acquires the site of the primary focus of colorectal cancer, and the selection unit selects the plurality of region images using the acquired site of the primary focus of colorectal cancer together with the determined at least one histopathological feature.

[0018] According to this configuration, by selecting region images using the site of the primary focus of the target disease, the probability that the region images input to the gene mutation prediction model can be limited to region images related to the target disease is improved, so that the prediction accuracy of the presence or absence of gene mutations can be improved.

[0019] The information processing system according to the sixth aspect of the present invention is the information processing system according to the fourth or fifth aspect, wherein the feature prediction model is a model machine-learned using learning data that takes as input a region image obtained by dividing a pathological tissue image and uses as output the pathological histological features assigned to the region image, and the gene mutation prediction model is a model machine-learned using learning data that takes as input a region image selected using a combination of the presence or absence of pathological histological features and uses as output information on the presence or absence of a specific gene mutation.

[0020] According to this configuration, since the feature prediction model and the gene mutation prediction model are models after machine learning, the prediction accuracy of the presence or absence of gene mutations can be improved.

[0021] The information processing method according to the seventh aspect of the present invention includes an acquisition procedure for acquiring a pathological tissue image of a patient with a target disease, a division procedure for dividing the pathological tissue image of the patient into a plurality of region images, a feature prediction procedure for inputting the region images to each of a plurality of feature prediction models constructed for each type of pathological histological feature and respectively obtaining prediction information on the presence or absence of the pathological histological features, a selection procedure for selecting a plurality of region images whose combination of the presence or absence of the acquired pathological histological features matches a preset combination of the presence or absence of pathological histological features at the time of selection, a gene mutation prediction procedure for inputting the region images selected by the selection to each of the gene mutation prediction models constructed for each combination of the presence or absence of pathological histological features and each type of gene mutation and respectively obtaining prediction information on the presence or absence of the gene mutation, and a prediction result output procedure for outputting a prediction result on the presence or absence of at least one gene mutation for the patient using the prediction information on the presence or absence of the gene mutation for each of the acquired region images.

[0022] According to this configuration, if a pathological tissue image of a patient with a target disease is input, a prediction result of the presence or absence of gene mutations can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of gene mutations for the patient, thereby suppressing a delay in the administration of a drug according to the gene abnormality of the target disease.

[0023] An information processing method according to an eighth aspect of the present invention is an information processing method for predicting a gene in which a mutation has occurred in a colorectal cancer tissue of a patient with colorectal cancer, the method including: an acquisition step of acquiring a colorectal cancer pathological tissue image of the patient; a division step of dividing the colorectal cancer pathological tissue image of the patient into a plurality of region images; a feature prediction step of inputting a region image to each of a plurality of feature prediction models constructed for each type of pathological histological feature and respectively obtaining prediction information on the presence or absence of the pathological histological feature; a selection step of selecting a plurality of region images whose combination of the presence or absence of the acquired pathological histological feature matches a combination of the presence or absence of the pathological histological feature at the time of screening preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; a gene mutation prediction step of inputting the selected region image to each of gene mutation prediction models constructed for each combination of the presence or absence of the pathological histological feature and each type of gene mutation and respectively obtaining prediction information on the presence or absence of the gene mutation; and a prediction result output step of outputting a prediction result on the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI for the patient by using the prediction information on the presence or absence of the gene mutation for each of the acquired region images.

[0024] According to this configuration, if a colorectal cancer pathological tissue image of a patient with colorectal cancer is input, a prediction result of the presence or absence of gene mutations can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of gene mutations for the patient, thereby suppressing a delay in the administration of a drug according to the gene abnormality of colorectal cancer.

[0025] The program according to the ninth aspect of the present invention causes a computer to execute an acquisition procedure for acquiring a pathological tissue image of a patient with a target disease, a division procedure for dividing the pathological tissue image of the patient into a plurality of region images, and for each of a plurality of feature prediction models constructed for each type of histological feature, an input of a region image to obtain prediction information on the presence or absence of histological features, respectively; a selection procedure for selecting a plurality of region images whose combination of the acquired presence or absence of histological features matches a preset combination of the presence or absence of histological features at the time of screening; a gene mutation prediction procedure for inputting the region images selected by the selection for each of the gene mutation prediction models constructed for each combination of the presence or absence of histological features and each type of gene mutation to obtain prediction information on the presence or absence of gene mutations, respectively; and a prediction result output procedure for outputting a prediction result on the presence or absence of at least one gene mutation for the patient using the prediction information on the presence or absence of gene mutations for each of the acquired region images.

[0026] According to this configuration, if a pathological tissue image of a patient with a target disease is input, a prediction result on the presence or absence of gene mutations can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result on the presence or absence of gene mutations for the patient, thereby suppressing a delay in the administration of a drug corresponding to the gene abnormality of the target disease.

[0027] The program according to the tenth aspect of the present invention is a program for predicting genes in which mutations have occurred in colorectal cancer tissue of a patient with colorectal cancer, and causes a computer to perform an acquisition procedure for acquiring a colorectal cancer pathological tissue image of the patient, a division procedure for dividing the colorectal cancer pathological tissue image of the patient into a plurality of region images, a feature prediction procedure for inputting a region image to each of a plurality of feature prediction models constructed for each type of histopathological feature and acquiring prediction information on the presence or absence of the histopathological feature, a selection procedure for selecting a plurality of region images whose combination of the acquired presence or absence of the histopathological feature matches a preset combination of the presence or absence of the histopathological feature at the time of selection for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI, a gene mutation prediction procedure for inputting the selected region image to each of the gene mutation prediction models constructed for each combination of the presence or absence of the histopathological feature and each type of gene mutation and acquiring prediction information on the presence or absence of the gene mutation, and a prediction result output procedure for outputting a prediction result on the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI for the patient using the prediction information on the presence or absence of the gene mutation for each of the acquired region images.

[0028] According to this configuration, if a colorectal cancer pathological tissue image of a patient with colorectal cancer is input, a prediction result on the presence or absence of a gene mutation can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result on the presence or absence of the gene mutation for the patient, so that it is possible to suppress a delay in the administration of a drug according to the gene abnormality of colorectal cancer.

[0029] An information processing system according to the eleventh aspect of the present invention is an information processing system for estimating the presence or absence of a BRAF gene mutation in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or more images, and from among the divided images, an image having a tumor cell ratio exceeding 50%, and / or an image including a papillary structure, and / or an image including a non-serrated papillary structure, and / or an image including a non-serrated papillary structure and having mucus present, and / or an image including a cribriform structure, and / or an image including a cribriform structure and having no mucus present, and / or an image including a cribriform structure and having mucus present, and means for estimating the presence or absence of a BRAF gene mutation using the selected image.

[0030] According to this configuration, if a pathological tissue image is input, a prediction result of the presence or absence of a BRAF gene mutation can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a pharmaceutical corresponding to the prediction result of the presence or absence of a BRAF gene mutation for the patient, so that a delay in the administration of a pharmaceutical corresponding to the BRAF gene abnormality of cancer can be suppressed.

[0031] An information processing system according to the twelfth aspect of the present invention is an information processing system for estimating the presence or absence of a BRAF V600E gene mutation in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or more images, and from among the divided images, an image including a streaky pattern, and / or an image including a small solid cyst, and / or an image in which the small solid cyst is composed of normal-type cells, and / or an image in which the large solid cyst is composed of normal-type cells, and / or an image including an oval nucleus, and / or an image having mucus present, and / or an image including a non-serrated papillary structure and having no mucus present, and / or an image including a cribriform structure and having no mucus present, and / or an image including a cribriform structure and having mucus present, and means for estimating the presence or absence of a BRAF V600E gene mutation using the selected image.

[0032] According to this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of BRAF V600E gene mutation can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of BRAF V600E gene mutation for the patient, so that the delay in the administration of a drug corresponding to the BRAF V600E gene abnormality of cancer can be suppressed.

[0033] The information processing system according to the 13th aspect of the present invention is an information processing system for estimating the presence or absence of an ERBB2 gene mutation in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or a plurality of images, means for selecting, from among the divided images, an image including a cord-like structure and having mucus, and / or an image including a cord-like structure and having mucus leakage, and means for estimating the presence or absence of an ERBB2 gene mutation using the selected image as an analysis target.

[0034] According to this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of ERBB2 gene mutation can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of ERBB2 gene mutation for the patient, so that the delay in the administration of a drug corresponding to the ERBB2 gene abnormality of cancer can be suppressed.

[0035] The information processing system according to the 14th aspect of the present invention is an information processing system for estimating the presence or absence of a TP53 gene mutation in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or a plurality of images, means for selecting, from among the divided images, an image including signet ring cells and having mucus leakage, and / or an image including goblet cells, and / or an image including a cord-like structure and having no mucus, and / or an image having mucus leakage, and means for estimating the presence or absence of a TP53 gene mutation using the selected image.

[0036] According to this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of TP53 gene mutation can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of TP53 gene mutation for the patient, so that the delay in the administration of a drug corresponding to the TP53 gene abnormality of cancer can be suppressed.

[0037] The information processing system according to the 15th aspect of the present invention is an information processing system for estimating the presence or absence of MSI gene abnormality in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or more images, and from among the divided images, an image including a streak pattern, and / or an image including a cord-like structure and having no mucus, and / or an image having a tumor cell ratio exceeding 50%, and / or an image including a solid nest, and / or an image including a small solid nest and normal-type cells, and / or an image including an oval nucleus, and / or an image including a papillary structure, and / or an image including goblet cells, and / or an image including a non-serrated papillary structure, and / or an image including an ovaloid nucleus, and / or an image including a cribriform structure and having no mucus, and / or an image including a cribriform structure and having mucus, and / or an image including a cribriform structure and having mucus leakage, and means for estimating the presence or absence of MSI gene abnormality using the selected image.

[0038] According to this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of MSI gene abnormality can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of MSI gene abnormality for the patient, so that the delay in the administration of a drug corresponding to the MSI gene abnormality of cancer can be suppressed.

[0039] An information processing system according to the 16th aspect of the present invention is an information processing system for estimating the presence or absence of a RAS gene mutation in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or more images, and from among the divided images, an image containing signet ring cells and having mucus leakage, and / or an image containing a tubular structure and having no mucus, and / or an image containing a cord-like structure and having no mucus, and / or an image containing small solid nests, and / or an image containing large solid nests, and / or an image in which large solid nests are composed of normal-type cells, and / or an image containing a papillary structure, and / or an image having no mucus, and / or an image having mucus, and / or an image containing goblet cells, and / or an image containing a non-serrated papillary structure, and / or an image containing a non-serrated papillary structure and having no mucus, and / or an image containing a non-serrated papillary structure and having mucus, and / or an image containing a cribriform structure, and / or an image containing a cribriform structure and having no mucus, and / or an image containing a cribriform structure and having mucus, and / or an image containing a cribriform structure and having mucus leakage, and / or an image having clustering, and / or an image having a tumor cell ratio exceeding 50%, and / or an image containing small solid nests, and / or an image in which small solid nests are composed of normal-type cells, and / or an image containing large solid nests, and / or an image in which large solid nests are composed of normal-type cells, and / or an image containing oval nuclei, and / or an image having mucus, and / or an image having mucus leakage, and / or an image containing a non-serrated papillary structure and having no mucus, and / or an image containing a cribriform structure, and / or an image containing a cribriform structure and having no mucus, and / or an image containing a cribriform structure and having mucus, and / or an image containing a tubular structure and having mucus, and / or an image containing a tubular structure and having mucus leakage, and / or an image containing a track pattern, and / or an image containing high-grade cellular atypia, and / or an image containing a cord-like structure and having mucus, and / or an image containing a cord-like structure and having mucus leakage, and / or an image containing solid nests, and / or an image containing a non-serrated papillary structure, and / or an image containing a non-serrated papillary structure and having mucus, and / or an image containing roundish nuclei, and means for selecting such images.Means for estimating the presence or absence of a RAS gene mutation using the selected image, and the like.

[0040] According to this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of a RAS gene mutation can be obtained immediately. As a result, the clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a RAS gene mutation for the patient, thereby suppressing the delay in the administration of a drug corresponding to the RAS gene abnormality of cancer.

Effect of the Invention

[0041] According to one aspect of the present invention, if a pathological tissue image of a patient with a target disease is input, the prediction result of the presence or absence of a gene mutation can be obtained immediately. As a result, the clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby suppressing the delay in the administration of a drug corresponding to the gene abnormality of the target disease. According to another aspect of the present invention, if a pathological tissue image of a colorectal cancer patient is input, the prediction result of the presence or absence of a gene mutation can be obtained immediately. As a result, the clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby suppressing the delay in the administration of a drug corresponding to the gene abnormality of colorectal cancer.

Brief Description of the Drawings

[0042]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 5

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Embodiments for Carrying Out the Invention

[0043] Hereinafter, each embodiment will be described with reference to the drawings. However, a more detailed description than necessary may be omitted. For example, a detailed description of well-known matters or a redundant description of substantially the same configuration may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate the understanding of those skilled in the art.

[0044] The target disease for predicting the presence or absence of gene mutations according to the present embodiment is applicable to diseases accompanied by gene mutations in tissues (or diseases for which drugs can be administered according to gene abnormalities), including cancer. However, as an example of the target disease according to the present embodiment, colorectal cancer will be described below. In the present embodiment, gene abnormalities are described as being included in gene mutations.

[0045] FIG. 1 is a schematic configuration diagram of the information processing system according to the present embodiment. As shown in FIG. 1, the information processing system S includes terminals 1-1 to 1-N (N is a natural number) and a computer system 2 connected via a communication circuit network NW.

[0046] The terminals 1-1 to 1-N are used by different users, and are, for example, mobile phones such as multifunctional mobile phones (so-called smartphones), tablets, notebook computers, or desktop computers. The terminals 1-1 to 1-N may display information transmitted from the computer system 2 via, for example, a web browser, or may display information transmitted from the computer system 2 on the screens of applications installed in the terminals 1-1 to 1-N. Here, as an example, the terminals 1-1 to 1-N will be described below as displaying information transmitted from the computer system 2 via, for example, a web browser.

[0047] The computer system 2 can communicate with the terminals 1-1 to 1-N, and these communications may be wired or wireless. That is, the computer system 2 is connected to the terminals 1-1 to 1-N so as to be able to exchange information. The computer system 2 is used, for example, by an administrator who manages the information processing system S according to the present embodiment. The computer system 2 may be a single computer or a plurality of computers.

[0048] FIG. 2 is a schematic configuration diagram of the computer system according to the present embodiment. As shown in FIG. 2, the computer system 2 includes an input interface 21, a communication module 22, a storage 23, a memory 24, and a processor 25. The input interface 21 receives an input from the administrator of the computer system 2 and outputs an input signal corresponding to the received input to the processor 25. The communication module 22 is connected to the communication circuit network NW and communicates with the terminals 1-1 to 1-N. This communication may be wired or wireless, but will be described as being wired.

[0049] The storage 23 stores programs for the processor 25 to read and execute, a feature prediction model after machine learning, a mutation prediction model after machine learning, and various data. The memory 24 temporarily holds data and programs. The memory 24 is a volatile memory, for example, a RAM (Random Access Memory). The processor 25 loads a program from the storage 23 into the memory 24 and functions as an acquisition unit 251 and an output unit 250 by executing a series of instructions included in the program. The output unit 250 has, for example, a division unit 252, a feature prediction unit 253, a selection unit 254, a gene mutation prediction unit 255, and a prediction result output unit 256. Each process will be described later.

[0050] Subsequently, the learning process of the gene mutation prediction model will be described with reference to FIG. 3. FIG. 3 is a schematic diagram showing the learning process of the gene mutation prediction model. (Step S1) The division unit 252 divides an image of a pathological tissue (here, as an example, a colorectal cancer tissue) of a patient with a target disease (here, as an example, colorectal cancer) (for example, a slide image) into one or more region images (for example, tile-shaped tile images). Here, the gene mutation of the diseased tissue (here, as an example, colorectal cancer tissue) of the patient is known. Here, the case of dividing into tile-shaped tile images is also referred to as tile division.

[0051] (Step S2) Subsequently, for example, using an annotation system, a pathologist inputs histological features (for example, oval nuclei, etc.) for each of the plurality of region images. At this time, a terminal device (not shown) used by the pathologist receives the histological features. The acquisition unit 251 acquires the histological features assigned to each of the plurality of region images.

[0052] (Step S3) Subsequently, the processor 25 learns the relationship between the region image and the histopathological feature for each histopathological feature, and outputs a feature prediction model. Specifically, for example, the processor 25 uses the training data with the region image obtained by dividing the histopathological image as the input and the histopathological feature assigned to the region image as the output, and learns by machine learning (e.g., Deep Learning) to output a feature prediction model. As a result, L (L is a natural number) feature prediction models are output for the number of histopathological features, and these feature prediction models FM-1, …, FM-L are stored in the storage 23 by the processor 25. In this way, the feature prediction model is a model learned by machine learning using the training data with the region image obtained by dividing the histopathological image as the input and the histopathological feature assigned to the region image as the output.

[0053] (Step S4) Next, the processor 25 also predicts the presence or absence of each histopathological feature for the divided image (here, as an example, the tile image) without annotation for the histopathological feature, using the feature prediction models FM-1, …, FM-L. As a result, the presence or absence of each histopathological feature is predicted for the divided image without annotation for the histopathological feature.

[0054] (Step S5) Next, the selection unit 254 selects region images (here, as an example, tile images) from the plurality of region images using the combination of the predicted presence or absence of the histopathological feature. Specifically, the selection unit 254 extracts the region image (here, as an example, the tile image) corresponding to a specific combination of the presence or absence of the histopathological feature using the histopathological feature predicted in Step S4.

[0055] (Step S6) Next, the processor 25 learns the relationship between the gene mutations and the group of tile images corresponding to a specific combination of the presence or absence of histopathological features, and outputs a gene mutation prediction model. Specifically, for example, the processor 25 uses the region image (the region image extracted in step S5) corresponding to a specific combination of the presence or absence of histopathological features as the input, and uses the learning data with the information on the presence or absence of a specific gene mutation as the output to perform machine learning (e.g., Deep Learning) to output a gene mutation prediction model. As a result, gene mutation prediction models are output for the number M (M is a natural number) of gene mutations, and these gene mutation prediction models GM-1, …, GM-M are stored in the storage 23 by the processor 25. In this way, the gene mutation prediction model is a model that is machine-learned using the learning data with the region image selected using a specific combination of the presence or absence of histopathological features as the input and the information on the specific gene mutation as the output.

[0056] Subsequently, with reference to FIGS. 4A and 4B, the process of annotating histopathological features to a pathological tissue image will be described. FIG. 4A is a schematic diagram for explaining the process of annotating histopathological features to a pathological tissue image. As shown in FIG. 4A, the pathological tissue image is divided into a plurality of region images. Among the generated region images, the region images outside the cell tissue (hereinafter also referred to as white images) and the region images in which substances other than the cell tissue (e.g., magic ink) occupy more than a reference are excluded. This process is called region determination. For the region images that are not excluded, histopathological features are annotated manually (e.g., by a pathologist).

[0057] FIG. 4B is a schematic diagram showing an example of a process of annotating histological features to a pathological tissue image. As shown in FIG. 4B, the pathological tissue image is divided into a plurality of region images. From each of the region images, a white image and a region image (also referred to as a magic image) in which magic ink occupies more than a reference are excluded. For each of the region images that are not excluded, the presence or absence of each histological feature is assigned. In FIG. 4B, when there is a histological feature, it is indicated by ○, and when there is no histological feature, it is indicated by ×, and for each region image, whether each of the L histological features is present or not is assigned.

[0058] Subsequently, the learning and testing in constructing a feature prediction model for a certain histological feature will be described with reference to FIG. 5. FIG. 5 is a schematic diagram explaining the learning and testing in constructing a feature prediction model for a certain histological feature. As shown in FIG. 5, the learning data is a set of a region image and the presence or absence of a certain histological feature. The region image is given as the input of the machine learning model, and the presence or absence of the certain feature is given as the output of the machine learning model, and machine learning is executed. Thereby, the relationship between the region image and the presence or absence of a certain histological feature is learned, and a feature prediction model for predicting the certain histological feature is generated.

[0059] As an example, during learning, 5-fold Cross Validation is performed on 80% of the entire training data. That is, training is carried out on 64% of the entire training data and validation is performed on 16% of the training data. Specifically, the validation is carried out as follows. First, 80% of the entire training data is divided into 5 parts. For the sake of easy explanation, the data subsets formed by the 5-fold division are denoted as s1, s2, s3, s4, and s5 respectively. For example, after dividing 80% of the training data into 5 parts, first, the subsets s1, s2, s3, and s4 are used as training subsets to train the model. Subsequently, s5 is used as the validation subset to validate the model. Let the evaluation index (such as accuracy or F1 value, etc.) obtained at this time be e1. Next, training is carried out with s2, s3, s4, and s5, and evaluation is performed with s1. Similarly, learning and evaluation are repeated while swapping the subsets. By repeating learning and evaluation for all combinations, 5 evaluation indices are obtained. And among these 5 evaluation indices, the model with the best performance is determined as the feature prediction model.

[0060] For the accuracy verification of the obtained feature prediction model, a blind test is carried out on the model with the best performance in the 5-fold cross-validation using 20% of the entire training data. Thereby, the performance of the determined feature prediction model is verified. If the performance meets the criteria, this feature prediction model is adopted.

[0061] Subsequently, with reference to FIGS. 6A and 6B, the assignment of predicted values of histopathological features to the region images executed in step S4 of FIG. 3 will be described. FIG. 6A is a schematic diagram for explaining the assignment of predicted values of histopathological features to the region images. In FIG. 6A, the feature prediction model FM-1 for histopathological feature 1, the feature prediction model FM-2 for histopathological feature 2, the feature prediction model FM-3 for histopathological feature 3,..., and the feature prediction model FM-L for histopathological feature L are shown. The predicted value of the histopathological feature is defined as the degree of having that histopathological feature, which is determined based on the prediction results of the feature prediction models FM-1, …, FM-L for each region image. The predicted value of the histopathological feature may be the prediction result itself (e.g., a numerical value between 0 and 1), or it may be a value according to the comparison result of comparing the prediction result with a threshold value (e.g., a value of 0 or 1). Here, as an example, it is assumed that the predicted value of the histopathological feature is a numerical value between 0 and 1 for explanation.

[0062] FIG. 6B is a schematic diagram for explaining an example of assigning the predicted value of the histopathological feature to the region image. As shown by the feature predicted values p11, p12, p13, …, p1L, …, p71, p72, p73, …, p7L in FIG. 6B, for each region image excluding the white image and the magic image, the predicted values of the histopathological features 1 to L are respectively assigned. In this case, in step S5 of FIG. 3, the processor 25 may determine that when the feature predicted value exceeds the threshold value, it has that histopathological feature, and when the feature predicted value is below the threshold value, it does not have that histopathological feature. The threshold value may be set to different values for each histopathological feature, or the same value may be set. Here, as an example, it is explained that different values are set for each histopathological feature.

[0063] Subsequently, with reference to FIG. 7, the learning and testing in constructing the gene mutation prediction model for a certain gene mutation, which are executed in step S6 of FIG. 3, are explained. FIG. 7 is a schematic diagram for explaining the learning and testing in constructing the gene mutation prediction model for a certain gene mutation. As shown in FIG. 7, the learning data is a set of the selected region image and the presence or absence of a certain gene mutation. The selected region image is given as the input of the machine learning model, and the presence or absence of the certain gene mutation is given as the output of the machine learning model, and machine learning is executed. Thereby, the relationship between the region image and the presence or absence of a certain gene mutation is learned, and a gene mutation prediction model for predicting the certain gene mutation is generated.

[0064] As an example, during learning, 5-fold Cross Validation is performed on 80% of the total training data. That is, training is carried out on 64% of the total training data and validation is performed on 16% of the total training data. Specifically, the validation is carried out as follows. First, 80% of the total training data is divided into 5 parts. For the sake of easy explanation, the data subsets formed by the 5-fold division are respectively denoted as s1, s2, s3, s4, and s5. For example, after dividing 80% of the training data into 5 parts, first, the subsets s1, s2, s3, and s4 are used as training subsets to proceed with the learning of the model. Subsequently, s5 is used as the validation subset to perform the validation of the model. Let the evaluation index (such as accuracy or F1 value, etc.) obtained at this time be e1. Next, learning is carried out with s2, s3, s4, and s5 and evaluation is performed with s1. Similarly, learning and evaluation are repeated while swapping the subsets. By repeating learning and evaluation for all combinations, 5 evaluation indices are obtained. Then, the model with the best performance among these 5 evaluation indices is determined as the gene mutation prediction model.

[0065] For the accuracy verification of the obtained gene mutation prediction model, a blind test is carried out on the model with the best performance in the 5-fold cross-validation using 20% of the total training data. Thereby, the performance of the determined gene mutation prediction model is verified. If the performance meets the criteria, this gene mutation prediction model is adopted.

[0066] Using FIGS. 8 to 17, the setting method of the "combination of the presence or absence of histological features at the time of selection" when selecting regional images during the process of predicting the gene mutations of colorectal cancer patients with unknown gene mutations will be described. FIG. 8 is an example of the implementation of the test of the gene mutation prediction model for colorectal cancer. By performing the test as shown in FIG. 8, a "combination of the presence or absence of pathological histological features at the time of screening" is obtained. Specifically, the feature prediction model and the gene mutation prediction model are learned from 80% of the data in stage 2. The learned feature prediction model and the learned gene mutation prediction model are applied to the data in stage 1, the data in stage 2.5, and the data of TCGA. Here, the data of TCGA is data limited only to colon adenocarcinoma (COAD) and rectum adenocarcinoma (READ) among colorectal cancers.

[0067] <Regarding the combination of features capable of predicting gene mutations (or gene abnormalities)> As an example, for each type of target gene mutation (or gene abnormality), from all the experimental data, a combination of the first feature (site of the primary tumor) and the second feature (pathological histological feature) that meets all the following conditions (1) to (3) is selected. (1) In the stage 2 test set, the AUC (Area Under the Curve) shows 0.8 or more by either case - building method 1 or case - building method 2 (that is, accurate prediction is possible with high precision in the cohort used for learning). Here, AUC is the area (integral) under the ROC curve with the false - positive rate on the first axis and the true - positive rate on the second axis. The range of this area can take values between 0 and 1. (2) In any test set other than stage 2, the AUC shows 0.8 or more by either case - building method 1 or case - building method 2 (that is, accurate prediction is possible with high precision even in a cohort not used for learning). (3) Among those with the same combination of image type and features, the number of required tiles is the smallest.

[0068] Here, case - building method 1 or case - building method 2 is as follows. (1) Case - building method 1 When the average value of the predicted values of the gene mutation model for each region image included in the selected region image group (for example, the region image group in FIG. 19) is equal to or greater than the threshold value th, it is predicted that case 1 has a gene mutation, and if not, it is predicted that there is no gene mutation. (2) Case grouping method 2 When the ratio of the number of tiles for which the predicted value of the gene mutation model for each region image included in the selected region image group (for example, the region image group in FIG. 19) is equal to or greater than the first threshold value th1 to the total number of region images in the selected region image group (for example, the region image group in FIG. 19) is equal to or greater than the second threshold value th2, it is predicted that the subject patient has a gene mutation, and if not, it is predicted that there is no gene mutation.

[0069] FIG. 9 shows a set of combinations of the selected primary tumor sites and the presence or absence of pathological histological features at the time of selection when the target gene mutation is BRAF+. Here, BRAF+ means all BRAF mutations including BRAF V600E.

[0070] FIG. 10 shows a set of combinations of the selected primary tumor sites and the presence or absence of pathological histological features at the time of selection when the target gene mutation is BRAF V600E+. FIG. 11 shows a set of combinations of the selected primary tumor sites and the presence or absence of pathological histological features at the time of selection when the target gene mutation is ERBB2+. FIG. 12 shows a set of combinations of the selected primary tumor sites and the presence or absence of pathological histological features at the time of selection when the target gene mutation is TP53+. FIG. 13 shows a set of combinations of the selected primary tumor sites and the presence or absence of pathological histological features at the time of selection when the gene abnormality state is MSI. Here, MSI (microsatellite instability) refers to a state of a certain gene abnormality, not a gene. FIG. 14 shows a set of combinations of the selected primary tumor sites and the presence or absence of pathological histological features at the time of selection when the target gene mutation is RAS+. FIG. 15 is a continuation of FIG. 14. Here, RAS+ means KRAS+ or NRAS+. site and a set of combinations of the presence or absence of pathological histological features at the time of selection.

[0071] In FIGS. 9 to 15, the site of the primary focus, the combination of the presence or absence of pathological histological features at the time of screening, the AUC of case grouping method 1 and the AUC of case grouping method 2 in the phase 2 test data, the AUC of case grouping method 1 and the AUC of case grouping method 2 in the phase 1 data, the AUC of case grouping method 1 and the AUC of case grouping method 2 in the phase 2.5 data, and the AUC of case grouping method 1 and the AUC of case grouping method 2 in the TCGA data are shown. In the columns of "Combination of presence or absence of pathological histological features at the time of screening" in FIGS. 10, 13, 14, 15, FIGS. 16 and 17 described later, "Small solid nests and normal-type cells" means that the small solid nests are composed of normal-type cells (that is, normal-type cells form small solid nests). Similarly, in the columns of "Combination of presence or absence of pathological histological features at the time of screening" in FIGS. 10, 14, 15, FIGS. 16 and 17 described later, "Large solid nests and normal-type cells" means that the large solid nests are composed of normal-type cells (that is, normal-type cells form large solid nests).

[0072] The "○○ structure (or cell), and mucus △△" in FIGS. 9 to 17 means that the ○○ structure (or cell) has the characteristic of mucus △△ at the same time. Specifically, the "cribriform structure, and no mucus" in FIGS. 9, 10, 13, 14, 15, 16, and 17 means "cribriform structure, and no mucus at the same time". Similarly, the "cribriform structure, and mucus present" in FIGS. 9, 10, 13, 14, 15, 16, and 17 means "cribriform structure, and mucus present at the same time". Similarly, the "non-serrated papillary structure, and no mucus" in FIGS. 10, 14, 15, 16, and 17 means "non-serrated papillary structure, and no mucus at the same time". Similarly, the "cord-like structure, and mucus present" in FIGS. 11, 15, 16, and 17 means "cord-like structure, and mucus present at the same time". Similarly, the "cord-like structure, and mucus leakage" in FIGS. 11, 15, 16, and 17 means "cord-like structure, and mucus leakage at the same time". Similarly, the "signet ring cell, and mucus leakage" in FIGS. 12 and 16 means "signet ring cell, and mucus leakage at the same time". The "cord-like structure, and no mucus" in FIGS. 12, 13, and 16 means "cord-like structure, and no mucus at the same time". Similarly, the "cribriform structure, and mucus leakage" in FIGS. 13, 14, 15, 16, and 17 means "cribriform structure, and mucus leakage at the same time". Similarly, the "signet ring cell, and mucus leakage" in FIGS. 14 and 17 means "signet ring cell, and mucus leakage at the same time". Similarly, the "tubular structure, and no mucus" in FIGS. 14 and 17 means "tubular structure, and no mucus at the same time". Similarly, the "non-serrated papillary structure, and mucus present" in FIGS. 14, 15, and 17 means "non-serrated papillary structure, and mucus present at the same time". Similarly, the "tubular structure, and mucus present" in FIGS. 15 and 17 means "tubular structure, and mucus present at the same time". Similarly, the "tubular structure, and mucus leakage" in FIGS. 15 and 17 means "tubular structure, and mucus leakage at the same time".

[0073] FIG. 16 shows an example of a table stored in the storage 23. FIG. 17 is a continuation of FIG. 16. In the BRAF table T1 of FIG. 16, records of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection are accumulated when the target gene mutation is BRAF+. Also, in the BRAF V600E table T2 of FIG. 16, records of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection are accumulated when the target gene mutation is BRAF V600E.

[0074] Also, in the ERBB2 table T3 of FIG. 16, records of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection are accumulated when the target gene mutation is ERBB2. Also, in the TP53 table T4 of FIG. 16, records of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection are accumulated when the target gene mutation is TP53. Also, in the MSI table T5 of FIG. 16, records of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection are accumulated when the gene abnormality state is MSI.

[0075] Also, in the RAS table T6 of FIG. 17, records of combinations of the selected primary tumor site and the presence or absence of histopathological features at the time of selection are accumulated when the target gene mutation is RAS.

[0076] Subsequently, the estimation process of the presence or absence of each gene mutation will be described with reference to FIG. 18. FIG. 18 is a schematic diagram showing the estimation process of the presence or absence of each gene mutation. (Step S11) First, the acquisition unit 251 acquires a pathological tissue image of a patient for whom gene analysis has not been performed and the gene mutation is unknown.

[0077] (Step S12) Next, the division unit 252 divides the pathological tissue image of the patient into a plurality of region images.

[0078] (Step S13) Next, the feature prediction unit 253 predicts the presence or absence of the target histological feature for each region image using the feature prediction models constructed for each type of histological feature. More specifically, for example, the feature prediction unit 253 inputs the region image to each of the plurality of feature prediction models constructed for each type of histological feature, and acquires prediction information on the presence or absence of the histological feature respectively.

[0079] (Step S14) Next, the feature prediction unit 253 selects the plurality of region images using the combination of the acquired presence or absence of the histological feature. More specifically, for example, the feature prediction unit 253 extracts, from the plurality of region images, region images in which the combination of the acquired presence or absence of the histological feature matches a specific combination of the presence or absence of the histological feature set for each gene abnormality. Here, for example, a specific combination of the presence or absence of histopathological features set for BRAF gene abnormalities is one of the "combinations of the presence or absence of histopathological features at the time of screening" stored in the BRAF table T1 in FIG. 16. Similarly, for example, a specific combination of the presence or absence of histopathological features set for BRAF V600E gene abnormalities is one of the "combinations of the presence or absence of histopathological features at the time of screening" stored in the BRAF V600E table T2 in FIG. 16. Similarly, for example, a specific combination of the presence or absence of histopathological features set for ERBB2 gene abnormalities is one of the "combinations of the presence or absence of histopathological features at the time of screening" stored in the ERBB2 table T3 in FIG. 16. Similarly, for example, a specific combination of the presence or absence of histopathological features set for TP53 gene abnormalities may be one or a plurality of the "combinations of the presence or absence of histopathological features at the time of screening" stored in the TP53 table T4 in FIG. 16. Similarly, for example, a specific combination of the presence or absence of histopathological features set for MSI gene abnormalities is one of the "combinations of the presence or absence of histopathological features at the time of screening" stored in the MSI table T5 in FIG. 16. Similarly, for example, a specific combination of the presence or absence of histopathological features set for RAS gene abnormalities is one of the "combinations of the presence or absence of histopathological features at the time of screening" stored in the RAS table T6 in FIG. 16.

[0080] (Step S15) Next, the gene mutation prediction unit 255 uses the gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and each type of gene mutation to predict the presence or absence of the target gene mutation for each region image. More specifically, for example, the gene mutation prediction unit 255 inputs the region image selected by screening into each of the gene mutation prediction models constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, and respectively obtains the prediction information on the presence or absence of the gene mutation.

[0081] (Step S16) Next, the prediction result output unit 256 uses the prediction result for each region image to predict the presence or absence of each gene mutation of the patient.

[0082] (Step S17) Next, for example, the prediction result output unit 256 outputs a list of prediction results regarding the presence or absence of each gene mutation. Note that the prediction result of the presence or absence of a gene mutation does not have to be the prediction results of the presence or absence of all gene mutations, and at least one or more are sufficient. In this way, the prediction result output unit 256 uses the prediction information of the presence or absence of gene mutations for each acquired region image to output the prediction result of the presence or absence of at least one gene mutation for the patient.

[0083] In the above example, for the target gene mutation, region images that match the "combination of the presence or absence of histological features at the time of screening" are selected, the selected region images are input into the gene mutation prediction model, the gene mutation prediction results for each region image are obtained, and using the prediction results for each of these region images, the presence or absence of the target gene mutation in the target patient is predicted. However, it is not limited to this. This series of steps may be performed for each "combination of the presence or absence of histological features at the time of screening". In this case, it may be collectively output which "combination of the presence or absence of histological features at the time of screening" resulted in a prediction of the presence of the target gene mutation. Furthermore, by performing the above processing for a plurality of gene mutations, the above output may be performed across a plurality of gene mutations.

[0084] Note that the acquisition unit 251 may further acquire the location of the primary focus of the target disease. In this case, the selection unit 254 may use the acquired combination of the presence or absence of histological features and the acquired location of the primary focus of the target disease to select the plurality of region images. For example, in the case of colorectal cancer, in order to predict the presence or absence of BRAF gene mutations, the screening unit 254 may refer to, for example, the BRAF table T1 (see FIG. 16) stored in the storage 23, and for one record, select, for example, a region image that matches the combination of "site of the primary tumor" and "presence or absence of pathological histological features at the time of screening". For example, in the case of the first record (the record on the first line), the screening unit 254 may select a region image in which the site of the primary tumor is "left side of the large intestine" and the combination of the presence or absence of pathological histological features at the time of screening corresponds to "tumor cell ratio exceeding 50%". By thus selecting region images using also the site of the primary tumor of the target disease, the probability that the region images input to the gene mutation prediction model can be limited to region images related to the target disease is improved, so that the prediction accuracy of the presence or absence of gene mutations can be improved.

[0085] In this case, for example, for the target gene mutation, a region image that matches the set of "site of the primary tumor" and "combination of the presence or absence of pathological histological features at the time of screening" is selected, the selected region image is input to the gene mutation prediction model, gene mutation prediction results for each region image are obtained, and using the prediction results for each of these region images, the presence or absence of the target gene mutation in the target patient is predicted. This series of steps may be performed for each set of "site of the primary tumor" and "combination of the presence or absence of pathological histological features at the time of screening". In this case, it may be collectively output which set of "site of the primary tumor" and "combination of the presence or absence of pathological histological features at the time of screening" was predicted to have the target gene mutation. Further, by performing the above processing for a plurality of gene mutations, the above output may be performed for a plurality of gene mutations.

[0086] Subsequently, a method of assembling the prediction results for each patient from the prediction results for the region images will be described with reference to FIG. 19. FIG. 19 is a schematic diagram for explaining a method of assembling the prediction results for each patient from the prediction results for the region images. FIG. 19 shows a set of the screened image regions derived from the target patient and the prediction values for each region image by the gene mutation prediction model. The prediction result output unit 256 predicts the presence or absence of gene mutations at the case level by, for example, the case assembly method 1 or the case assembly method 2 when the number of tiles included in the region image group after screening (for example, the region image group in FIG. 19) is K or more. On the other hand, for example, when the number of tiles is less than K, the prediction result output unit 256 does not perform the gene mutation prediction itself at the case level (and treats it as having no gene mutations as a result).

[0087] Subsequently, an example of screen transition on the terminal 1 will be described with reference to FIG. 20. FIG. 20 is a schematic diagram for explaining an example of screen transition on the terminal. As shown in FIG. 20, on the screen G1 of the terminal 1, a text box TB1 for inputting the path of the pathological tissue image file to select the patient's pathological tissue image and a reference button B1 for referring to the pathological tissue image file are provided. In addition, a radio button B2 for selecting the position of the primary focus of colorectal cancer is provided. When the transmission button B3 is pressed in a state where the patient's pathological tissue image is selected and the position of the primary focus of colorectal cancer is selected, the screen transitions to G2. On the transitioned screen G2, a list of prediction results of the presence or absence of gene mutations is displayed.

[0088] Subsequently, the flow of the screen transition process in FIG. 20 will be described with reference to FIG. 21. FIG. 21 is a flowchart showing an example of the flow of the screen transition process in FIG. 20. (Step S110) First, the terminal 1 receives the patient's pathological tissue image and the position of the primary focus of colorectal cancer.

[0089] (Step S120) Next, the terminal 1 transmits the patient's pathological tissue image and the position of the primary focus of colorectal cancer to the computer system 2.

[0090] (Step S130) Next, the computer system 2 receives the patient's pathological tissue image and the position of the primary focus of colorectal cancer, and outputs information for displaying the prediction results of the presence or absence of each gene mutation using the patient's pathological tissue image and the position of the primary focus of colorectal cancer. The details of this process have been described in FIG. 18, so the description thereof will be omitted.

[0091] (Step S140) Next, the computer system 2 transmits information for displaying the prediction results of the presence or absence of each gene mutation to the terminal 1.

[0092] (Step S150) Next, the terminal 1 receives the information for displaying the prediction results of the presence or absence of each gene mutation, and uses this information to display the prediction results regarding the presence or absence of each gene mutation. Thus, the processing of this flowchart ends.

[0093] As described above, the information processing system S according to the present embodiment includes an acquisition unit 251 that acquires a pathological tissue image of a patient with a target disease, a division unit 252 that divides the pathological tissue image of the patient into a plurality of region images, a feature prediction unit 253 that inputs the region images to each of a plurality of feature prediction models constructed for each type of histological feature and acquires prediction information on the presence or absence of histological features respectively, a selection unit 254 that selects a plurality of region images whose combinations of the presence or absence of the acquired histological features match the combinations of the presence or absence of histological features at the time of screening set in advance, a gene mutation prediction unit 255 that inputs the region images selected by the selection to each of the gene mutation prediction models constructed for each combination of the presence or absence of histological features and each type of gene mutation and acquires prediction information on the presence or absence of gene mutations respectively, and a prediction result output unit 256 that outputs a prediction result on the presence or absence of at least one gene mutation for the patient using the prediction information on the presence or absence of gene mutations for each of the acquired region images.

[0094] According to this configuration, if a pathological tissue image of a patient with a target disease is input, the prediction result of the presence or absence of a gene mutation can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a gene mutation for the patient, thereby suppressing the delay in the administration of a drug corresponding to the gene abnormality of the target disease.

[0095] Also, as an example, the information processing system according to this embodiment is an information processing system for predicting genes in which mutations have occurred in the colorectal cancer tissue of a patient with colorectal cancer. This information processing system S includes an acquisition unit 251 that acquires the colorectal cancer pathological tissue image of the patient. Further, the information processing system S includes a division unit 252 that divides the colorectal cancer pathological tissue image of the patient into a plurality of region images. Further, the information processing system S includes a feature prediction unit 253 that inputs the region images to each of a plurality of feature prediction models constructed for each type of histopathological feature, and respectively acquires prediction information on the presence or absence of the histopathological feature. Further, the information processing system S includes a selection unit 254 that selects a plurality of region images whose combination of the presence or absence of the acquired histopathological features matches the combination of the presence or absence of the histopathological features at the time of selection preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI. Further, the information processing system S includes a gene mutation prediction unit 255 that inputs the selected region images to each of the gene mutation prediction models constructed for each combination of the presence or absence of the histopathological features and each type of gene mutation, and respectively acquires prediction information on the presence or absence of the gene mutation, and a prediction result output unit 256 that uses the prediction information on the presence or absence of the gene mutation for each of the acquired region images to output a prediction result on the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI for the patient.

[0096] According to this configuration, if a colorectal cancer pathological tissue image of a patient with colorectal cancer is input, a prediction result on the presence or absence of a gene mutation can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result on the presence or absence of a gene mutation for the patient, so that it is possible to suppress a delay in the administration of a drug according to the gene abnormality of colorectal cancer.

[0097] Moreover, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of a BRAF gene mutation in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or more images, and, from among the divided images, means for selecting an image with a tumor cell ratio exceeding 50%, and / or an image including a papillary structure, and / or an image including a non-serrated papillary structure, and / or an image including a non-serrated papillary structure and having mucus present, and / or an image including a cribriform structure, and / or an image including a cribriform structure and having no mucus present, and / or an image including a cribriform structure and having mucus present, and means for estimating the presence or absence of a BRAF gene mutation using the selected images.

[0098] With this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of a BRAF gene mutation can be obtained immediately. Thereby, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of a BRAF gene mutation for the patient, thus suppressing the delay in the administration of a drug corresponding to the BRAF gene abnormality of the cancer.

[0099] Moreover, the information processing system according to this embodiment is an information processing system for estimating the presence or absence of a BRAF V600E gene mutation in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or more images, and, from among the divided images, means for selecting an image including a trabecular pattern, and / or an image including a small solid cyst, and / or an image in which the small solid cyst is composed of normal-type cells, and / or an image in which the large solid cyst is composed of normal-type cells, and / or an image including an oval nucleus, and / or an image having mucus present, and / or an image including a non-serrated papillary structure and having no mucus present, and / or an image including a cribriform structure and having no mucus present, and / or an image including a cribriform structure and having mucus present, and means for estimating the presence or absence of a BRAF V600E gene mutation using the selected images.

[0100] With this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of the BRAF V600E gene mutation can be obtained immediately. As a result, the clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of the BRAF V600E gene mutation for the patient, thereby suppressing the delay in the administration of the drug corresponding to the BRAF V600E gene abnormality of the cancer.

[0101] Moreover, the information processing system according to the present embodiment is an information processing system for estimating the presence or absence of an ERBB2 gene mutation in a tumor, comprising means for dividing the pathological tissue image of the tumor into one or a plurality of images, means for selecting, from among the divided images, an image including a cord-like structure and having mucus, and / or an image including a cord-like structure and having mucus leakage, and means for estimating the presence or absence of an ERBB2 gene mutation using the selected image as an analysis target.

[0102] With this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of the ERBB2 gene mutation can be obtained immediately. As a result, the clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of the ERBB2 gene mutation for the patient, thereby suppressing the delay in the administration of the drug corresponding to the ERBB2 gene abnormality of the cancer.

[0103] Moreover, the information processing system according to the present embodiment is an information processing system for estimating the presence or absence of a TP53 gene mutation in a tumor, comprising means for dividing the pathological tissue image of the tumor into one or a plurality of images, means for selecting, from among the divided images, an image including signet-ring cells and having mucus leakage, and / or an image including goblet cells, and / or an image including a cord-like structure and having no mucus, and / or an image having mucus leakage, and means for estimating the presence or absence of a TP53 gene mutation using the selected image.

[0104] With this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of TP53 gene mutation can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of TP53 gene mutation for the patient, thereby suppressing the delay in the administration of drugs corresponding to the TP53 gene abnormality of cancer.

[0105] Moreover, the information processing system according to the present embodiment is an information processing system for estimating the presence or absence of MSI gene abnormality in a tumor, comprising means for dividing a pathological tissue image of the tumor into one or more images, and from among the divided images, an image including a streak pattern, and / or an image including a cord-like structure and having no mucus, and / or an image having a tumor cell ratio exceeding 50%, and / or an image including a solid nest, and / or an image in which small solid nests are composed of normal-type cells, and / or an image including an oval nucleus, and / or an image including a papillary structure, and / or an image including goblet cells, and / or an image including a non-serrated papillary structure, and / or an image including an oval nucleus, and / or an image including a cribriform structure and having no mucus, and / or an image including a cribriform structure and having mucus, and / or an image including a cribriform structure and having mucus leakage, and means for estimating the presence or absence of MSI gene abnormality using the selected image.

[0106] With this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of MSI gene abnormality can be obtained immediately. As a result, a clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of MSI gene abnormality for the patient, thereby suppressing the delay in the administration of drugs corresponding to the MSI gene abnormality of cancer.

[0107] The information processing system according to the present embodiment is an information processing system for estimating the presence or absence of RAS gene mutations in a tumor, and includes means for dividing a pathological tissue image of the tumor into one or more images, and from among the divided images, an image including signet ring cells and having mucus leakage, and / or an image including a tubular structure and having no mucus, and / or an image including a cord-like structure and having no mucus, and / or an image including small solid nests, and / or an image including large solid nests, and / or an image in which large solid nests are composed of normal-type cells, and / or an image including a papillary structure, and / or an image having no mucus, and / or an image having mucus, and / or an image including goblet cells, and / or an image including a non-serrated papillary structure, and / or an image including a non-serrated papillary structure and having no mucus, and / or an image including a non-serrated papillary structure and having mucus, and / or an image including a cribriform structure, and / or an image including a cribriform structure and having no mucus, and / or an image including a cribriform structure and having mucus, and / or an image including a cribriform structure and having mucus leakage, and / or an image having budding, and / or an image in which the tumor cell ratio exceeds 50%, and / or an image including small solid nests, and / or an image in which small solid nests are composed of normal-type cells, and / or an image including large solid nests, and / or an image including oval nuclei, and / or an image having mucus, and / or an image having mucus leakage, and / or an image including a non-serrated papillary structure and having no mucus, and / or an image including a cribriform structure, and / or an image including a cribriform structure and having no mucus, and / or an image including a cribriform structure and having mucus, and / or an image including a tubular structure and having mucus, and / or an image including a tubular structure and having mucus leakage, and / or an image including a track pattern, and / or an image including high-grade cellular atypia, and / or an image including a cord-like structure and having mucus, and / or an image including a cord-like structure and having mucus leakage, and / or an image including solid nests, and / or an image including a non-serrated papillary structure, and / or an image including a non-serrated papillary structure and having mucus, and / or an image including roundish nuclei, and means for estimating the presence or absence of RAS gene mutations using the selected images.

[0108] With this configuration, if a pathological tissue image is input, the prediction result of the presence or absence of the RAS gene mutation can be obtained immediately. As a result, the clinician can refer to this prediction result and prescribe a drug corresponding to the prediction result of the presence or absence of the RAS gene mutation for the patient, thereby suppressing the delay in the administration of drugs corresponding to the RAS gene abnormality of cancer.

[0109] Note that at least a part of the computer system 2 described in the above-described embodiment may be configured by hardware or may be configured by software. When configured by software, a program for realizing at least a part of the functions of the computer system 2 may be stored in a recording medium such as a flexible disk or a CD-ROM, and read and executed by a computer. The recording medium is not limited to removable ones such as magnetic disks and optical disks, and may be a fixed recording medium such as a hard disk device or a memory.

[0110] Also, a program for realizing at least a part of the functions of the computer system 2 may be distributed via a communication line (including wireless communication) such as the Internet. Furthermore, the program may be distributed in an encrypted, modulated, or compressed state via a wired or wireless line such as the Internet or stored in a recording medium.

[0111] Furthermore, the computer system 2 may be made to function by one or a plurality of information devices. When using a plurality of information devices, one of them may be used as a computer, and the functions may be realized as at least one means of the computer system 2 by the computer executing a predetermined program.

[0112] In the invention of a method, all steps may be realized by automatic control by a computer. Alternatively, while each step is carried out by a computer, the progress control between steps may be carried out manually. Further, at least a part of all steps may be carried out manually.

[0113] As described above, the present invention is not limited to the above-described embodiments as they are, and at the implementation stage, the components can be modified and embodied without departing from the gist thereof. Also, various inventions can be formed by appropriately combining a plurality of components disclosed in the above embodiments. For example, some components may be deleted from all the components shown in the embodiments. Further, components from different embodiments may be appropriately combined.

Explanation of Reference Numerals

[0114] 1 Terminal 2 Computer system 21 Input interface 22 Communication module 23 Storage 24 Memory 25 Processor 250 Output unit 251 Acquisition unit 252 Division unit 253 Feature prediction unit 254 Selection unit 255 Gene mutation prediction unit 256 Prediction result output unit

Claims

1. For each of a plurality of feature prediction models constructed for each type of histological features, a feature prediction unit that inputs a plurality of region images segmented from a histological image of a patient with a target disease and respectively obtains prediction information on the presence or absence of histological features; A selection unit that selects a plurality of region images whose combination of the presence or absence of the obtained histological features matches a preset combination of the presence or absence of histological features at the time of selection; For each of a plurality of gene mutation prediction models constructed for each type of gene mutation, a gene mutation prediction unit that inputs the group of region images selected by selection and obtains prediction information on the presence or absence of gene mutation for each region image from each of the gene mutation prediction models; A prediction result output unit that outputs a prediction result on the presence or absence of a gene mutation for at least one type of gene mutation for the patient by processing the prediction information on the presence or absence of the gene mutation for each region image obtained according to a predetermined rule for each of the plurality of prediction information output from the same gene mutation prediction model; An information processing system comprising the above.

2. The acquisition unit further acquires the site of the primary focus of the target disease, The selection unit selects the plurality of region images using the combination of the presence or absence of the obtained histological features and the site of the primary focus of the obtained target disease. The information processing system according to claim 1.

3. The feature prediction model is a model trained using learning data that takes a region image obtained by dividing a histological image as an input and uses the histological features assigned to the region image as an output. The gene mutation prediction model is a model trained using learning data that takes a region image selected using a combination of the presence or absence of histological features as an input and uses information on the presence or absence of a specific gene mutation as an output. The information processing system according to claim 1 or 2.

4. An information processing system for predicting genes in which mutations have occurred in the colorectal cancer tissue of a patient with colorectal cancer, For each of a plurality of feature prediction models constructed for each type of histological features, a feature prediction unit that inputs a plurality of region images segmented from a histological image of a patient with colorectal cancer and respectively obtains prediction information on the presence or absence of histological features; A selection unit that selects a plurality of region images whose combination of presence or absence of the obtained histopathological features matches a combination of presence or absence of histopathological features at the time of pre-set selection for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; A gene mutation prediction unit that inputs the selected group of region images for each of the gene mutation prediction models constructed for each type of gene mutation, and obtains prediction information on the presence or absence of gene mutations for each region image from each of the gene mutation prediction models; A prediction result output unit that outputs a prediction result on the presence or absence of at least one gene mutation among BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI for the patient by processing the prediction information on the presence or absence of gene mutations for each of the obtained region images according to a predetermined rule for each of the plurality of prediction information output from the same gene mutation prediction model; An information processing system comprising the above.

5. The acquisition unit further acquires the site of the primary focus of colorectal cancer; The selection unit selects the plurality of region images using the site of the primary focus of the obtained colorectal cancer together with the determined at least one histopathological feature. The information processing system according to claim 4.

6. The feature prediction model is a model trained using learning data that takes as input a region image obtained by dividing a histopathological image and uses as output the histopathological features assigned to the region image. The gene mutation prediction model is a model trained using learning data that takes as input a region image selected using a combination of presence or absence of histopathological features and uses as output information on the presence or absence of a specific gene mutation. The information processing system according to claim 4 or 5.

7. A feature prediction procedure that inputs a plurality of region images obtained by dividing a histopathological image of a patient with a target disease for each of a plurality of feature prediction models constructed for each type of histopathological feature, and respectively obtains prediction information on the presence or absence of histopathological features; A selection procedure that selects a plurality of region images whose combination of presence or absence of the obtained histopathological features matches a combination of presence or absence of histopathological features at the time of pre-set selection. For each gene mutation prediction model constructed for each type of gene mutation, input the group of region images selected by screening, and obtain the prediction information on the presence or absence of gene mutations for each region image from each gene mutation prediction model. A gene mutation prediction procedure, and By processing the prediction information on the presence or absence of gene mutations for each of the obtained region images according to a predetermined rule for each of a plurality of pieces of prediction information output from the same gene mutation prediction model, for the patient, output a prediction result on the presence or absence of gene mutations for at least one type of gene mutation. A prediction result output procedure, and An information processing method having the above.

8. An information processing method for predicting genes in which mutations have occurred in the colorectal cancer tissue of a patient with colorectal cancer, For each of a plurality of feature prediction models constructed for each type of histopathological feature, input a plurality of region images segmented from the histopathological image of a patient with colorectal cancer, and respectively obtain prediction information on the presence or absence of histopathological features. A feature prediction procedure, and A screening procedure for selecting a plurality of region images whose combination of the presence or absence of the obtained histopathological features matches the combination of the presence or absence of histopathological features at the time of screening preset for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI, For each of the gene mutation prediction models constructed for each combination of the presence or absence of histopathological features and each type of gene mutation, input the selected group of region images, and obtain the prediction information on the presence or absence of gene mutations for each region image from each gene mutation prediction model. A gene mutation prediction procedure, and By processing the prediction information on the presence or absence of gene mutations for each of the obtained region images according to a predetermined rule for each of a plurality of pieces of prediction information output from the same gene mutation prediction model, for the patient, output a prediction result on the presence or absence of gene mutations for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI. A prediction result output procedure, and An information processing method having the above.

9. On a computer, For each of a plurality of feature prediction models constructed for each type of histopathological feature, input a plurality of region images segmented from the histopathological image of a patient with the target disease, and respectively obtain prediction information on the presence or absence of histopathological features. A feature prediction procedure, and A screening procedure for screening a plurality of region images in which the combination of the presence or absence of the obtained histopathological features matches the combination of the presence or absence of histopathological features at the time of screening set in advance; For each gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and the type of gene mutation, input the group of region images selected by screening, and obtain prediction information on the presence or absence of gene mutation for each region image from each gene mutation prediction model. A gene mutation prediction procedure; A prediction result output procedure for outputting a prediction result on the presence or absence of a gene mutation for at least one type of gene mutation for the patient by processing the prediction information on the presence or absence of a gene mutation for each obtained region image for each of a plurality of prediction information output from the same gene mutation prediction model according to a predetermined rule; A program for causing the above to be executed.

10. A program for predicting genes in which mutations have occurred in colorectal cancer tissue of a patient with colorectal cancer, the computer For each of a plurality of feature prediction models constructed for each type of histopathological feature, input a plurality of region images divided from a histopathological image of a patient with the target disease, and obtain prediction information on the presence or absence of histopathological features respectively. A feature prediction procedure; A screening procedure for screening a plurality of region images in which the combination of the presence or absence of the obtained histopathological features matches the combination of the presence or absence of histopathological features at the time of screening set in advance for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI; For each gene mutation prediction model constructed for each combination of the presence or absence of histopathological features and the type of gene mutation, input the selected group of region images, and obtain prediction information on the presence or absence of gene mutation for each region image from each gene mutation prediction model. A gene mutation prediction procedure; A prediction result output procedure for outputting a prediction result on the presence or absence of a gene mutation for at least one of BRAF, BRAF V600E, ERBB2, RAS, TP53, or MSI for the patient by processing the prediction information on the presence or absence of a gene mutation for each obtained region image for each of a plurality of prediction information output from the same gene mutation prediction model according to a predetermined rule; A program for causing the above to be executed.

Citation Information

Patent Citations

  • Disease data prediction method and device, readable storage medium and electronic equipment

    CN110660055A

  • Auxiliary system and method for predicting gene mutation in lung cancer pathological image

    CN111369534A

  • Pathological diagnosis support device, pathological diagnosis support method, control program for pathological diagnosis support, and recording medium with the program recorded thereon

    JP2011215680A

  • Classification and mutation prediction from histopathology images using deep learning

    US20200184643A1

  • Prediction of effect of EGFR inhibitor by detecting BRAF mutation

    WO2016104794A1