Discrimination System, Program, and Method for High-Frequency Mutation-Type Cancer

The discriminant system addresses the challenges of detecting high-frequency mutant cancers by using a machine learning-based approach with image data, achieving rapid and accurate identification and improving treatment selection.

JP7699361B2Active Publication Date: 2025-06-27NIIGATA UNIVERSITY +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2023015100
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-02-15
Filing Date
2023-02-03
Publication Date
2025-06-27
Estimated Expiration
2039-02-07

AI Technical Summary

Technical Problem

Conventional methods for detecting high-frequency mutant cancers are labor-intensive, time-consuming, and not all high-frequency mutant cancers can be detected using existing techniques.

Method used

A discriminant system comprising an input unit, a holding unit, a machine learning execution unit, and a discrimination unit, which uses image data of pathological sections to generate a discrimination model for accurately identifying high-frequency mutant cancers.

Benefits of technology

The system enables quick and accurate discrimination of high-frequency mutant cancers, overcoming the limitations of conventional methods and facilitating the selection of effective treatment drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699361000001
    Figure 0007699361000001
  • Figure 0007699361000002
    Figure 0007699361000002
  • Figure 0007699361000003
    Figure 0007699361000003
Patent Text Reader

Abstract

The present invention provides a method, program, and method for discriminating highly mutated cancers with higher accuracy than conventional methods. [Solution] A system for discriminating hypermutated cancer is provided, comprising an input unit, a storage unit, a machine learning execution unit, and a discrimination unit, wherein the input unit is configured to be able to input a plurality of first image data, a plurality of second image data, and a plurality of third image data, the storage unit is configured to be able to store the first image data and the second image data, the machine learning execution unit is configured to be able to generate a discrimination model that discriminates whether or not the cancer is hypermutated using the first image data and the second image data stored by the storage unit as training data, and the discrimination unit is configured to input third image data to the discrimination model and be able to discriminate whether or not the third image data is hypermutated cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a discriminant system, program, and method for hypermutated cancer.

Background Art

[0002] It has been found that cancer can be classified by the pattern of gene mutations by widely examining gene mutations in cancer. One of the patterns of mutations characteristic of such cancers is hypermutation (Hypermutation or Hypermutated). Hypermutated cancers are distinguished by a higher somatic mutation rate compared to other types. It is known that there are cancers showing the characteristics of hypermutation in gastric cancer, breast cancer, colorectal cancer, glioblastoma, uterine cancer, etc. Hypermutated cancers often simultaneously have the property of microsatellite instability indicating a defect or incompleteness of the mismatch repair mechanism during DNA replication. This is thought to be due to mutations in the genes of mismatch repair enzymes MLH1, MLH3, MSH2, MSH3, MSH6, PMS2, or suppression of the expression of the MLH1 gene by methylation. It is also known that mutations in DNA polymerase ε (POLE), a DNA replication enzyme, cause somatic mutations at a particularly high frequency and result in hypermutation (Non-Patent Documents 1, 2).

[0003] On the one hand, the cancer immune escape mechanism has been elucidated, and new cancer immunotherapies targeting this mechanism have been clinically applied. Among them, the PD-1 (Programmed cell Death-1) / PD-L1 (PD-1 Ligand1) pathway, also known as the immune checkpoint pathway, is particularly notable. By blocking the immunosuppressive co-signaling PD-1 / PD-L1 pathway, the immunosuppression of T cells is released, T cells are activated, and the suppression of tumors expressing cancer-specific antigens occurs. In addition, CTLA-4 is also expressed on activated T cells, and when the CD28 ligand of antigen-presenting cells binds, the activation of T cells is suppressed. Therefore, by blocking this pathway, it is also possible to release the immunosuppression of T cells and cause tumor suppression. Anticancer drugs applying such principles have been put into practical use (e.g., nivolumab, ipilimumab).

[0004] Furthermore, there are also multiple other such immunosuppressive mechanisms, and it is expected that anti-tumor drugs blocking these immunosuppressive mechanisms will be developed and put into practical use in the future. High-frequency mutant cancers have many cancer-specific antigens that are targets of the immune mechanism, so it has been shown that the effect of therapies blocking the immunosuppressive signaling pathway is high, and a method that can easily discriminate that a cancer is a high-frequency mutant type is desired (Non-Patent Document 3).

[0005] Conventionally, to examine high-frequency mutant cancers, a method of performing comprehensive genetic analysis and counting the number of mutations is known, but there is a problem that it requires a lot of labor and time for the examination. In addition, a method of examining the deficiency or incompleteness of the mismatch repair mechanism, which is one of the causes of high-frequency mutations in cancers, by immunohistochemistry of related genes or microsatellite instability test is also known, but this method has a problem that not all high-frequency mutant cancers can be detected.

[0006] On the other hand, a pathological diagnosis support program as disclosed in Patent Document 1 is known.

Prior Art Documents

Non-Patent Documents

[0007] [Non-Patent Document 1] Nat Rev Cancer.2014 December;14(12):786‐800 [Non-Patent Document 2] J Pathol 2013;230:148‐153 [Non-Patent Document 3] Science 03 Apr 2015 Vol. 348,Issue 6230,pp.124-128 [Patent Document]

[0008] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-346911 [Summary of the Invention] [Problems to be Solved by the Invention]

[0009] In Patent Document 1, it is stated that it is possible to determine the presence or absence of a tumor and whether it is benign or malignant, but nothing is mentioned about a method for discriminating high-frequency mutant cancers.

[0010] The present invention has been made in view of such circumstances, and provides a method, program, and method for discriminating high-frequency mutant cancers with higher accuracy than conventional ones. [Means for Solving the Problems]

[0011] According to the present invention, there is provided a discriminant system for high-frequency mutant cancer, comprising an input unit, a holding unit, a machine learning execution unit, and a discrimination unit. The input unit is configured to be able to input a plurality of first image data, a plurality of second image data, and a plurality of third image data. The first image data is image data representing a pathological section of a stained high-frequency mutant cancer. The second image data is image data representing a pathological section of a cancer that is not a high-frequency mutant cancer and has the same staining as the pathological section of the cancer that is the basis of the first image data. The third image data is image data representing a pathological section of a cancer for which it is newly determined whether it is a high-frequency mutant cancer and has the same staining as the pathological section of the cancer that is the basis of the first image data. The holding unit is configured to be able to hold the first image data and the second image data. The machine learning execution unit is configured to be able to generate a discrimination model for discriminating whether it is a high-frequency mutant cancer using the first image data and the second image data held by the holding unit as teacher data. The discrimination unit is configured to be able to input the third image data into the discrimination model and discriminate whether the third image data is a high-frequency mutant cancer.

[0012] According to the present invention, a discrimination model for discriminating whether it is a high-frequency mutant cancer is generated using the first image data and the second image data as teacher data. Here, the first image data is image data representing a pathological section of a stained high-frequency mutant cancer. The second image data is image data representing a pathological section of a cancer that is not a high-frequency mutant cancer and has the same staining as the pathological section of the cancer that is the basis of the first image data. Then, the third image data is input into the discrimination model, and it is configured to be able to discriminate whether the third image data is a high-frequency mutant cancer. Here, the third image data is image data representing a pathological section of a cancer for which it is newly determined whether it is a high-frequency mutant cancer and has the same staining as the pathological section of the cancer that is the basis of the first image data. As a result, it is possible to quickly and highly accurately determine whether it is a high-frequency mutant cancer, which has been difficult to determine conventionally without performing gene analysis using a next-generation sequencer or the like, and it becomes possible to easily select a drug effective for treatment.

[0013] Hereinafter, various embodiments of the present invention will be exemplified. The embodiments shown below can be combined with each other. Preferably, the staining method of the pathological section is hematoxylin-eosin staining. Preferably, the input unit is configured to be further capable of inputting non-cancer image data, the non-cancer image data is image data that is not a pathological section of cancer, the holding unit is configured to be further capable of holding the non-cancer image data, and the machine learning execution unit uses the non-cancer image data held by the holding unit as teacher data and is further configured to be capable of generating a discrimination model for discriminating whether the image data of the pathological section of cancer is or not, and the discrimination unit is further configured to be capable of discriminating whether the third image data is cancer image data or not. Preferably, an image processing unit is provided, and the image processing unit is configured to be capable of executing a Z-value conversion process for converting each of the RGB colors for each pixel of at least one of the first image data, the second image data, and the non-cancer image data into a Z value in the CIE color system based on the color distribution of the first image data and the second image data or the entire non-cancer image data. Preferably, the image processing unit is configured to be capable of executing a division process for dividing at least one of the first image data, the second image data, and the non-cancer image data input to the input unit. Preferably, the division process is configured to be capable of executing a division process for dividing the image data of the same pathological section for at least one of the first image and the second image data. Preferably, the image processing unit executes the division process so that some regions overlap in the divided image. Preferably, the image processing unit is further configured to be capable of executing a division process for dividing the third image data input to the input unit. Preferably, the discrimination unit discriminates whether the third image data is image data of a pathological section of cancer, and for the image data discriminated to be a pathological section of cancer, further discriminates whether it is a high-frequency mutant cancer or not. Preferably, the determination unit determines whether the cancer is a high-frequency mutant cancer based on the ratio of the image data determined to be the high-frequency mutant cancer in the image data determined to be the image data of the pathological section of the cancer. According to another aspect, a program is provided that causes a computer to function as an input unit, a storage unit, a machine learning execution unit, and a determination unit. The input unit is configured to be able to input a plurality of first image data and a plurality of second image data. The first image data is image data representing a stained pathological section of a high-frequency mutant cancer, and the second image data is image data of a pathological section of a cancer that is not a high-frequency mutant cancer and has the same staining as the pathological section of the cancer that is the basis of the first image data. The storage unit is configured to be able to store the first image data and the second image data. The machine learning execution unit is configured to be able to generate a determination model that determines whether it is a high-frequency mutant cancer, using the first image data and the second image data stored by the storage unit as teacher data. According to another aspect, a method for determining a high-frequency mutant cancer is provided, which is executed using the system according to any one of the above. According to another aspect, a method for determining a high-frequency mutant cancer is provided, which is executed using the program according to any one of the above. Preferably, it includes a step of determining the effectiveness of an immune checkpoint inhibitor.

Brief Description of Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Various characteristic matters shown in the following embodiments can be combined with each other.

[0016] <1. First Embodiment> 1.1. Discrimination of Whether it is a High-Frequency Mutation Type Cancer Hereinafter, with reference to FIGS. 1 to 4, the system 10 according to an embodiment of the present invention will be described.

[0017] (1.1.1. System 10) As shown in FIG. 1, the system 10 includes an input unit 1, an image processing unit 2, a holding unit 3, a machine learning execution unit 4, and a discrimination unit 5.

[0018] The input unit 1 is configured to be able to input a plurality of first image data, a plurality of second image data, and a plurality of third image data. Here, the first image data is image data representing a pathological section of a stained high-frequency mutant cancer. The second image data is image data representing a pathological section of a cancer that is not a high-frequency mutant cancer and is stained the same as the pathological section of the cancer that is the basis of the first image data. Further, the third image data is image data representing a pathological section of a cancer for which it is newly determined whether or not it is a high-frequency mutant cancer and is stained the same as the pathological section of the cancer that is the basis of the first image data. Here, in the present embodiment, the RGB values of these image data can take values from 0 to 255.

[0019] In the present embodiment, 17 pathological tissue stained specimens of each of the colorectal cancer samples determined to be of the Hypermutation type (high-frequency mutant type) and the Non-Hypermutation type (not high-frequency mutant type) from the analysis of the cancer genome DNA sequences were obtained. Here, these 17 cases are 17 cases that could be determined to be Hypermutation as a result of cancer genome sequencing of 201 Japanese colorectal cancer patients (Reference: Nagahashi et al GenomeMed 2017). Then, the pathological tissue stained specimens of colorectal cancer stained with hematoxylin and eosin were used as the first image data and the second image data using digital pathology technology. Here, in the present embodiment, the first image data and the second image data were stored as digital pathology image data conforming to the MIRAX format. Here, the above conditions are not limited to this, and it may be configured to obtain a predetermined number of cancer samples other than colorectal cancer.

[0020] Thus, in the present embodiment, since the hematoxylin and eosin stained image data with many clinical cases is adopted as the first image data and the second image data, it is possible to realize a highly versatile discrimination system.

[0021] However, other staining methods may be adopted according to the conditions. Further, the storage format of the image data is not limited to this.

[0022] The image processing unit 2 is configured to be capable of executing a splitting process for splitting a plurality of first image data, second image data, and third image data input to the input unit 1. In the present embodiment, the image processing unit 2 has a function of splitting the first image data, the second image data, and the third image data into predetermined tiles. As an example, the first image data, the second image data, and the third image data are split by the image processing unit 2 into images having a size of 300 pixels × 300 pixels. Note that such a split size is not particularly limited, but it is preferably a size that can identify whether the image data is a cancer tissue site. And in the present embodiment, by the splitting process, each of the first image data and the second image data is split into 1,000 or more. Further, in the present embodiment, the image processing unit 2 is configured to be capable of executing a splitting process for splitting the image data of the same pathological section for at least one of the first image and the second image data. Note that the split size and the number of splits are not limited thereto, and any conditions can be adopted.

[0023] In this way, by splitting the image data input to the input unit 1, the number of data of the teacher data used for subsequent machine learning can be increased, and the accuracy of the machine learning can be improved.

[0024] Further, in the present embodiment, the image processing unit 2 is further configured to be capable of executing a conversion process of converting each of the RGB colors for each pixel of the divided first image data and second image data into a Z value in the CIE color system based on the color distribution of the entire first image data and second image data. Specifically, the Z value follows a normal distribution centered on 0, and since the RGB values of the image data are values from 0 to 255, it is desirable to keep the Z - valued values of each RGB color within a range of twice the standard deviation (σ). For this reason, the image processing unit 2 has a function of correcting values of 2σ or more to 2σ and values of - 2σ or less to - 2σ. Also, the image processing unit 2 has a function of normalizing all the values to values from 0 to 1 by adding 2 to these values to convert all the values to values of 0 or more and then dividing by 4. Further, the image processing unit 2 has a function of converting such values to values of normal color representation by multiplying by 255. In addition, the image processing unit 2 also performs a process of truncating the decimal part so that such values become integer values. Note that the normalization method is not limited to this.

[0025] Here, if it is defined as "x = int(((min(max(xz, - 2), 2)+2) / 4)×255)", then "xz = the Z - valued RGB value" holds.

[0026] In this way, by converting each of the RGB colors of the first image data and the second image data into a Z value, it is possible to reduce the variation in color tone (color shading) in the staining process, and suppress the influence of the degree of staining on subsequent machine learning. As a result, it is possible to improve the accuracy of machine learning.

[0027] The holding unit 3 is configured to be able to hold the first image data and the second image data. Here, the holding unit 3 is composed of an arbitrary memory, latch, HDD, SSD, or the like.

[0028] The machine learning execution unit 4 is configured to be able to generate a discrimination model for discriminating whether it is a high - frequency mutant cancer or not, using the first image data and the second image data held by the holding unit 3 as teacher data. Details of the discrimination model will be described later with reference to FIG. 3.

[0029] The machine learning algorithm of the machine learning execution unit 4 is not particularly limited. For example, a neural network or deep learning can be used. Also, for example, a CNN (Convolutional Neural Network) for image recognition called "Inception-v3" developed by Google can be used. And such a CNN can be executed using the "Keras" framework. Regarding machine learning itself, in order to prevent overfitting, the accuracy of the model being trained is calculated using the images in the validation set every 1 epoch, and the training can be terminated at the epoch when the variation in the accuracy index stabilizes, and it can be implemented using the "Early Stopping" method. In this embodiment, in the learning with Z-score normalization, machine learning is repeatedly executed for 14 epochs.

[0030] The discrimination unit 5 is configured to be able to discriminate whether the third image data is high-frequency mutant cancer by inputting the third image data into the discrimination model.

[0031] (1.1.2. Flowchart) Next, with reference to FIG. 2, the flow for generating a discrimination model for discriminating whether it is high-frequency mutant cancer according to an embodiment of the present invention will be described.

[0032] First, in S1, the first image data and the second image data are input into the input unit 1.

[0033] Next, in S2, a splitting process is executed by the image processing unit 2 to split the first image data and the second image data. In the present embodiment, each of the first image data and the second image data is split into 1,000 or more tiles. Note that such a number of splits can be set as appropriate. For example, it may be 1,000 to 3,000, preferably 1,000 to 2,000, and more preferably 1,000 to 1,500. Specifically, for example, it may be 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,100, 2,200, 2,300, 2,400, 2,500, 2,600, 2,700, 2,800, 2,900, 3,000, or within the range between any two of the numerical values exemplified herein.

[0034] At the same time, Z - value conversion is executed by the image processing unit 2 to perform Z - value conversion on the split first image data and second image data.

[0035] Next, in S3, for the Z - value - converted first image data and second image data, it is determined whether each image data is a high - frequency mutant cancer tissue site (Hypermutaion type) or a non - high - frequency mutant cancer tissue site (Non - hypermutation type), and a label is attached to each image data. For example, it may be determined by a pathologist specializing in cancer, or the configuration may be such that image data already labeled is acquired from a server. Also, among the image data before splitting, marks are made at locations corresponding to the Hypermutaion type or the Non - hypermutation type, and when the split tile image data corresponds to the marked locations, the split data may be configured to be labeled.

[0036] Next, in S4, from the first image data and the second image data of 17 cases input to the input unit 1, 13 cases of image data to be used for machine learning by the machine learning execution unit 4 are selected. Such selection may be made randomly or may be selected by a pathologist specializing in cancer. Then, the first image data and the second image data with labels are held in the holding unit 3. Such first image data and second image data serve as the "correct set" in machine learning.

[0037] Next, in S5, the machine learning execution unit 4 performs machine learning using the first image data and the second image data held by the holding unit 3 as teacher data to generate a discrimination model for discriminating whether it is a high-frequency mutant cancer or not. Specifically, using the first image data and the second image data of 13 cases labeled in S4, machine learning for discriminating whether such image data is high-frequency mutant cancer or not is performed.

[0038] Next, in S6, it is determined whether the determination accuracy of the discrimination model is equal to or higher than a predetermined accuracy. If the determination accuracy of the discrimination model is not equal to or higher than the predetermined accuracy (NO), the process returns to S4 again, 13 cases of image data with different combinations are selected from the first image data and the second image data of 17 cases, and the process in S5 is executed. On the other hand, if the determination accuracy of the discrimination model is equal to or higher than the predetermined accuracy (YES), such a determination model is adopted and the process proceeds to S7.

[0039] Finally, in S7, the discrimination unit 5 outputs the discrimination model determined in S6 and stores it in the holding unit 3 or a storage unit (not shown).

[0040] (1.1.3. Discrimination of Whether it is High-Frequency Mutant Cancer) Next, with reference to FIGS. 3 and 4, the flow of the third image data when discriminating whether the third image data is high-frequency mutant cancer using the discrimination model will be described.

[0041] As shown in FIG. 3, in the present embodiment, the third image data input to the input unit 1 is output to the image processing unit 2, and the third image data on which the above-described image processing (segmentation processing and Z-value conversion processing) has been executed is output to the determination unit 5. Then, the determination unit 5 uses the determination model output in S7 of FIG. 2 to determine whether the third image data is a high-frequency mutant cancer or not.

[0042] In this way, by performing the segmentation process on the third image data as well, the size of the image data to be discriminated matches the sizes of the first and second image data, and the discrimination accuracy in the discrimination unit 5 can be improved.

[0043] The flowchart at this time is as follows.

[0044] As shown in FIG. 4, first, in S11, the third image data is input to the input unit 1.

[0045] Next, in S12, image processing (segmentation processing and Z-value conversion processing) is executed by the image processing unit 2.

[0046] Next, in S13, the determination unit 5 determines whether the third image data is a high-frequency mutant cancer or not using the above-described determination model.

[0047] Finally, in S14, the determination result by the determination unit 5 is output. The output mode of such a determination result is not particularly limited, and it can be "high-frequency mutant cancer", "not high-frequency mutant cancer", "the probability of being high-frequency mutant cancer is X%", etc.

[0048] (1.1.4. Discrimination by the Discrimination Model) Next, with reference to FIGS. 5 and 6, the discrimination using the determination model in S13 of FIG. 4 will be described. In the present embodiment, the machine learning algorithm is not particularly limited, and a neural network or deep learning can be used. Hereinafter, for simplicity of explanation, an example using a neural network will be described.

[0049] As shown in FIG. 5, the neural network (hereinafter referred to as NN in the drawings) is composed of a plurality of layers (first layer L1 to third layer L3) and a plurality of calculation nodes N (N11 to N31). Here, Nij represents the j-th calculation node N in the i-th layer. In the present embodiment, a neural network is constructed with i = 3 and j = 5. Note that the values of i and j are not limited to this, and for example, they can be integers between i = 1 to 100 and j = 1 to 100 or integers of 100 or more.

[0050] In addition, a predetermined weight w is set for each calculation node N. As shown in FIG. 4, for example, when focusing on the calculation node N23 in the second layer, a weight w is set between the calculation node N23 and all the calculation nodes N11 to N15 in the first layer, which is the previous layer. The weight w is set to a value between, for example, -1 and 1.

[0051] The machine learning execution unit 4 inputs various parameters into the neural network. In the present embodiment, as parameters input into the neural network, the Z value of the third image data, the distribution of the Z value of the third image data, the difference between the Z value of the third image data and the Z value of the first image data, the difference between the Z value of the third image data and the Z value of the second image data, and the difference between the distribution of the Z value of the third image data and the distribution of the Z values of the first image data and the second image data are used. Here, the Z values of the first to third image data are the Z values in pixel units. Also, the distribution of the Z values of the first to third image data is the distribution of the Z values within the image data (300 pixels × 300 pixels). Further, the difference between the distribution of the Z value of the third image data and the distribution of the Z values of the first image data and the second image data is the difference between the distribution of the Z value of the third image data and the distribution of the Z values corresponding to each pixel of the first image data and the second image data, or the sum of the differences in the Z values corresponding to each pixel within the image data.

[0052] Here, as described above, each parameter is normalized to a value between 0 and 1 when input to the neural network. For example, when the input parameter is 0, 0 is input as the input signal. Also, when the input parameter is 1, 1 is input as the input signal.

[0053] Then, the discrimination unit 5 inputs the input signal defined by various parameters to the first layer L1. Such an input signal is output from the calculation nodes N11 to N15 of the first layer to the calculation nodes N21 to N25 of the second layer L2, respectively. At this time, a value obtained by multiplying the values output from the calculation nodes N11 to N15 by the weight w set for each calculation node N is input to the calculation nodes N21 to N25. The calculation nodes N21 to N25 add up the input values, and input a value obtained by adding the bias b shown in FIG. 6 to such a value to the activation function f(). Then, the output value of the activation function f() (the output value from the virtual calculation node N'23 in the example of FIG. 4) is propagated to the next node, which is the calculation node N31. At this time, a value obtained by multiplying the weight w set between the calculation nodes N21 to N25 and the calculation node N31 by the above output value is input to the calculation node N31. The calculation node N31 adds up the input values and outputs the total value as an output signal. At this time, the calculation node N31 may add up the input values, input a value obtained by adding a bias to the total value to the activation function f(), and output the output value thereof as an output signal. Here, in the present embodiment, the value of the output signal is adjusted to be a value between 0 and 1. Then, the machine learning execution unit 4 outputs a value corresponding to the value of the output signal as a probability of determining whether it is a high-frequency mutant cancer.

[0054] As described above, the system 10 of the present embodiment uses the first image data and the second image data as teacher data, and executes machine learning by the machine learning execution unit 4 to generate a discrimination model (neural network and weight w) for determining whether it is a high-frequency mutant cancer. Then, the discrimination unit 5 uses such a discrimination model to determine whether the third image data is a high-frequency mutant cancer.

[0055] 1.2. Generation of Discrimination Model Next, with reference to FIG. 7, the generation of the discrimination model in S5 to S6 of FIG. 2 will be described.

[0056] As shown in FIG. 7, for each calculation node N that constitutes a neural network having the same configuration as the neural network shown in FIG. 5, the machine learning execution unit 4 sets a weight w ranging from, for example, -1 to 1. At this time, in order to reduce the influence of the weight w, it is preferable that the absolute value of the weight w set first is small. Then, five types of parameter sets are input to the neural network. In the present embodiment, as parameters input to the neural network, the Z value of the first image data, the Z value of the second image data, the distribution of the Z value of the first image data, the distribution of the Z value of the second image data, and the difference between the Z values of the first image data and the second image data are used. Here, the Z value of the first image data and the Z value of the second image data are Z values in terms of pixels. Also, the distribution of the Z value of the first image data and the distribution of the Z value of the second image data are distributions of Z values within the image data (300 pixels × 300 pixels). Further, the difference between the Z values of the first image data and the second image data is the difference between the Z values for each corresponding pixel of the first image data and the second image data or the sum of the differences between the Z values for each corresponding pixel within the image data.

[0057] Then, the output signal from the neural network is compared with the teacher data (discrimination by a specialist). When the difference between the output signal and the teacher data (hereinafter referred to as an error) is equal to or greater than a predetermined threshold, the weight w is changed, and the five types of parameter sets are input to the neural network again. At this time, the change of the weight w is executed by a known error propagation method or the like. By repeatedly executing such calculations (machine learning), the error between the output signal from the neural network and the given teacher data is minimized. At this time, the number of learning times of the machine learning is not particularly limited, and for example, it can be set to 1000 to 20000 times. Also, even if the error between the actual output signal and the given teacher data is not minimized, the machine learning may be terminated when such an error becomes equal to or less than a predetermined threshold or at an arbitrary timing of the developer.

[0058] When the machine learning by the machine learning execution unit 4 is completed, the machine learning execution unit 4 sets the weights of each calculation node N at this time in the neural network. That is, in the present embodiment, the weight w is stored in a storage unit such as a memory provided on the neural network. Then, the weight w set by the machine learning execution unit 4 is transmitted to a storage unit (not shown) provided in the system 10 and becomes the weight w of each calculation node N of the neural network in FIG. 5. In the present embodiment, the weight w is stored in a storage unit such as a memory provided on the neural network in FIG. 5. Here, by making the configuration of the neural network in FIG. 7 the same as the configuration of the neural network in FIG. 5, it becomes possible to directly use the weight w set by the machine learning execution unit 4.

[0059] <2. Second Embodiment> The second embodiment of the present invention will be described with reference to FIGS. 8 to 12. Note that the description of the same configurations and functions as those in the first embodiment will not be repeated.

[0060] As shown in FIG. 8, in the system 20 according to the second embodiment, the input unit 21 is configured to be able to further input non-cancer image data in addition to the first image data and the second image data. Here, the non-cancer image data means image data other than the pathological sections of cancer. The image processing unit 22 performs a segmentation process on the input image data. Details of the segmentation process will be described later.

[0061] The holding unit 23 is configured to be able to further hold the divided non-cancer image data in addition to the divided first image data and second image data. The machine learning execution unit 24 uses the first image data, second image data, and non-cancer image data held by the holding unit 3 as teacher data, and is configured to be able to generate a discrimination model (hereinafter referred to as the first discrimination model) for discriminating whether an image is a cancer image or not, and a discrimination model (hereinafter referred to as the second discrimination model) for discriminating whether a cancer image is a high-frequency mutant cancer or not. The discrimination unit 25 is configured to input the third image data into the first and second discrimination models, and to be able to discriminate whether the third image data is cancer image data or not, and whether it is high-frequency mutant cancer image data or not.

[0062] FIG. 9 shows image data P as an example input to the input unit 21. The image data P has a tissue region T and a blank region BL (for example, a region of a preparation). The tissue region T includes a cancer region C1 that is not a high-frequency mutant cancer, a region C2 of high-frequency mutant cancer, and a non-cancer tissue region NC.

[0063] The image processing unit 22 performs a division process on the image data P input to the input unit 21. In the example shown in FIG. 9, the tissue region T is divided into 100 parts with 10 in the vertical direction and 10 in the horizontal direction. That is, 100 tiles D 00 ~D 99 are set.

[0064] In this example, the tile corresponding to the region C2 of high-frequency mutant cancer (for example, tile D54) corresponds to the first image data, and the tile corresponding to the cancer region C1 that is not a high-frequency mutant cancer (for example, tile D34) corresponds to the second image data. Also, tiles corresponding only to the non-cancer tissue region NC (for example, tile D15), tiles corresponding only to the blank region BL (for example, tile D49), and tiles including the non-cancer tissue region NC and the blank region BL (for example, tile D04) all correspond to non-cancer image data.

[0065] As described above, in this embodiment, various images such as tiles corresponding to non-cancerous tissue regions NC as non-cancerous image data, tiles corresponding only to blank regions BL, and tiles including non-cancerous tissue regions NC and blank regions BL are input for machine learning. By increasing the diversity of non-cancerous images in this way, the accuracy of determining whether the test target data is a cancer image is improved.

[0066] Also, in this embodiment, further division processing (hereinafter referred to as the second division processing) can be performed on the image data after the above division processing (hereinafter referred to as the first division processing). In FIG. 10, the tile Dnm divided by the first division processing is further divided into five tiles. Here, in the second division processing, the division processing is executed such that some regions overlap in the divided tiles. That is, some images overlap between tile Dnm1 and tile Dnm2 after the second division processing. Also, some images overlap between tile Dnm2 and tile Dnm3.

[0067] By executing the division processing such that some regions overlap in the divided image in this way, it becomes possible to increase the number of images and improve the learning efficiency in subsequent machine learning.

[0068] FIG. 11 is a processing flow of the discrimination process of the third image data in this embodiment. As shown in FIG. 11, in this embodiment, the discrimination unit 25 discriminates whether the third image data is a cancer image and whether it is a high-frequency mutant cancer.

[0069] Specifically, in step S231 within step S23, the discrimination unit 25 discriminates whether the third image data is a cancer image. If it is not a cancer image (No in step S231), in step S233, it is discriminated that the third image data is a non-cancerous image.

[0070] On the other hand, if it is a cancer image (Yes in step S231), the discrimination unit 25 discriminates in step S232 whether the third image data is an image of a high-frequency mutant cancer. If it is not a high-frequency mutant cancer (No in step S232), in step S235, it is discriminated that the third image data is not an image of a high-frequency mutant cancer. On the other hand, if it is a high-frequency mutant cancer (Yes in step S232), in step S234, it is discriminated that the third image data is an image of a high-frequency mutant cancer.

[0071] In this way, in the present embodiment, discrimination is performed on whether the third image data is a cancer image and whether it is a high-frequency mutant cancer. Therefore, it is not necessary for a pathologist or the like to pre-diagnose whether it is cancer image data, and the working efficiency in the discrimination process can be improved.

[0072] Here, the discrimination unit 25 may discriminate whether the cancer is a high-frequency mutant cancer based on the ratio of the image data discriminated as a high-frequency mutant cancer in the image data discriminated as cancer image data.

[0073] In the example shown in FIG. 12, in the third image data P2, an image E2 discriminated as a high-frequency mutant cancer exists in an image E1 discriminated as cancer image data. At this time, when the ratio determined by (the number of tiles of E2) / (the number of tiles of E1) is larger than a predetermined threshold, the discrimination unit 25 discriminates that the region indicated by the image E1 is a high-frequency mutant cancer.

[0074] By doing so, it is possible to remove false positives that are locally discriminated as high-frequency mutant cancers as noise, and it is possible to improve the accuracy of discrimination.

[0075] As described above, in the second embodiment, the input unit 21 is configured to be further capable of inputting non-cancer image data, the machine learning execution unit 24 is configured to be further capable of generating a discrimination model that uses the non-cancer image data as teacher data to discriminate whether the image data is that of a pathological section of cancer, and the discrimination unit 25 is configured to be further capable of discriminating whether the third image data is cancer image data. With such a configuration, for the third image data, it is no longer necessary for a pathologist or the like to diagnose whether it is cancer, and the working efficiency of the discrimination process is improved.

[0076] <3. Other Embodiments> As described above, various embodiments have been described, but the present invention can also be implemented in the following aspects.

[0077] A computer is caused to function as an input unit, a holding unit, a machine learning execution unit, and an analysis unit, the input unit is configured to be capable of inputting a plurality of first image data and a plurality of second image data, the first image data is image data representing a stained pathological section of a high-frequency mutant cancer, the second image data is a pathological section that is not a high-frequency mutant cancer and is image data representing a pathological section that has the same staining as the cancer pathological section that is the basis of the first image data, the holding unit is configured to be capable of holding the first image data and the second image data, the machine learning execution unit is configured to be capable of generating a discrimination model that uses the first image data and the second image data held by the holding unit as teacher data to discriminate whether it is a high-frequency mutant cancer, Program.

[0078] A method for discriminating high-frequency mutant cancers, which is executed using the system according to any one of the above. Here, the high-frequency mutant cancers include any type of cancer, such as solid cancers like brain tumors, head and neck cancers, breast cancers, lung cancers, esophageal cancers, stomach cancers, duodenal cancers, appendiceal cancers, colorectal cancers, rectal cancers, liver cancers, pancreatic cancers, gallbladder cancers, bile duct cancers, anal cancers, kidney cancers, ureteral cancers, bladder cancers, prostate cancers, penile cancers, testicular cancers, uterine cancers, ovarian cancers, vulvar cancers, vaginal cancers, skin cancers, etc., but are not limited thereto. For the purpose of the present invention, the high-frequency mutant cancers are preferably colorectal cancers, lung cancers, stomach cancers, melanoma (malignant melanoma), head and neck cancers, and esophageal cancers.

[0079] A method for discriminating high-frequency mutant cancers, which is executed using the above program.

[0080] The discrimination method according to any one of the above, including the step of determining the effectiveness of an immune checkpoint inhibitor. Such a discrimination method may further include the step of indicating that a patient determined to have a high-frequency mutant cancer has a high effectiveness of administration of an immune checkpoint inhibitor. High-frequency mutant cancers have been shown to be highly effective in therapies that block immunosuppressive signaling pathways because they have many cancer-specific antigens that are targets of the immune system. Such a discrimination method is advantageous because it can easily discriminate that a cancer is a high-frequency mutant type. The "immune checkpoint" referred to herein is known in the art (Naidoo et al. British Journal of Cancer (2014) 111, 2214-2219), and CTLA4, PD1, and its ligand PDL-1, etc. are known. In addition, TIM-3, KIR, LAG-3, VISTA, BTLA are included. Inhibitors of immune checkpoints inhibit their normal immune functions. For example, they inhibit by negatively regulating the expression of immune checkpoint molecules or by binding to the molecules and blocking normal receptor / ligand interactions. Since immune checkpoints act to brake the immune system response to antigens, their inhibitors reduce this immunosuppressive effect and enhance the immune response. Inhibitors of immune checkpoints are known in the art, and preferred ones are anti-CTLA-4 antibodies (e.g., ipilimumab, tremelimumab), anti-PD-1 antibodies (e.g., nivolumab, pembrolizumab, pidilizumab, and RG7446 (Roche)), and anti-PDL-1 antibodies (e.g., BMS-936559 (Bristol-Myers Squibb), MPDL3280A (Genentech), MSB0010718C (EMD-Serono) and MEDI4736 (AstraZeneca)), etc., which are anti-immune checkpoint antibodies.

[0081] In addition, the holding unit 3 can be configured in a cloud computing mode provided in an information processing device such as an external PC or server. In this case, the external information processing device transmits the data required for each calculation to the system 10.

[0082] Moreover, it can also be provided as a computer-readable non-transitory recording medium storing the above-described program. Further, it can also be provided as an ASIC (application specific integrated circuit), FPGA (field-programmable gate array), or DRP (Dynamic ReConfigurable Processor) that implements the functions of the above-described program.

Explanation of Reference Numerals

[0083] 1, 21: Input unit 2, 22: Image processing unit 3, 23: Holding unit 4, 24: Machine learning execution unit 5, 25: Discrimination unit 10, 20: System

Claims

1. An input unit, a holding unit, a machine learning execution unit, a discrimination unit, and an image processing unit are provided, the input unit is configured to be able to input a plurality of first image data, a plurality of second image data, and a plurality of third image data, the first image data is image data representing a stained pathological section of cancer, the second image data is a pathological section that is not cancer, and is image data representing a pathological section that is stained the same as the pathological section that is the basis of the first image data, the third image data is a pathological section for newly discriminating whether it is cancer or not, and is image data representing a pathological section that is stained the same as the pathological section that is the basis of the first image data, the holding unit is configured to be able to hold the first image data, the second image data, and the third image data, the image processing unit is configured to be able to execute a conversion process of converting each color of RGB for each pixel in at least one of the first image data, the second image data, and the third image data into a Z value in the CIE color system based on the overall color distribution of the first image data, the second image data, or the third image data, the machine learning execution unit is configured to be able to generate a discrimination model that discriminates whether the third image data is image data of a pathological section of cancer or not, using the first image data and the second image data that have been subjected to the conversion process and held by the holding unit as teacher data, the discrimination unit is configured to input the third image data that may have been subjected to the conversion process into the discrimination model, and to be able to discriminate whether the third image data is image data of a pathological section of cancer or not, A cancer discrimination system.

2. The staining method of the pathological section is hematoxylin and eosin staining, The system according to Claim 1.

3. The image processing unit, is configured to be able to execute a division process of dividing at least one of the first image data, the second image data, and the third image data input to the input unit, The system according to Claim 1 or Claim 2.

4. The image processing unit, executes the division process so that some regions overlap in the divided image. The system according to Claim 3.

5. A computer, functions as an input unit, a holding unit, a machine learning execution unit, a discrimination unit, and an image processing unit, the input unit is configured to be able to input a plurality of first image data, a plurality of second image data, and a plurality of third image data, The first image data is image data representing a stained pathological section of cancer, The second image data is image data representing a pathological section that is not cancer and is stained the same as the pathological section that is the basis of the first image data, The third image data is image data of a pathological section for newly determining whether it is cancer or not, and is image data representing a pathological section stained the same as the pathological section that is the basis of the first image data, The holding unit is configured to be able to hold the first image data, the second image data, and the third image data, The image processing unit is configured to be able to execute a conversion process of converting each RGB color for each pixel of at least one of the first image data, the second image data, and the third image data into a Z value in the CIE color system based on the overall color distribution of the first image data, the second image data, or the third image data, The machine learning execution unit is configured to be able to generate a discrimination model that uses the first image data and the second image data that have been subjected to the conversion process and are held by the holding unit as teacher data to discriminate whether the third image data is image data of a pathological section of cancer, The discrimination unit is configured to input the third image data, which may have been subjected to the conversion process, into the discrimination model and be able to discriminate whether the third image data is image data of a pathological section of cancer, Program.

6. A method for discriminating cancer, which is executed using the system according to any one of Claims 1 to 4,

7. A method for discriminating cancer, which is executed using the program according to Claim 5,

8. Including a step of determining the effectiveness of an immune checkpoint inhibitor, The discrimination method according to Claim 6 or Claim 7.

Citation Information

Patent Citations

  • Method for controlling CNG engine based on fuel properties

    JP2004346911A

  • Computer-aided diagnosis apparatus and method

    JP2006320387A

  • Automated quantification of asymmetry

    JP2014508340A

  • Biological tissue image reconfiguration method and device, and image display device using the biological tissue image

    JP2015052581A

  • Image processing device, image processing method, and program

    JP2017167624A