Discrimination system, program, and method for hypermutated cancer

By constructing a machine learning model based on image data and using pathological slide image data stained with hematoxylin and eosin to generate a discriminant model, the problems of slow speed and low accuracy in the existing technology for identifying supermutant cancer are solved, and rapid and high-precision identification of supermutant cancer is achieved.

CN115684569BActive Publication Date: 2026-02-13NIIGATA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211268488.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-02-15
Filing Date
2019-02-07
Publication Date
2026-02-13
Estimated Expiration
2039-02-07

AI Technical Summary

Technical Problem

Existing technologies are insufficient for rapidly and accurately identifying hypermutant cancers, and whole-genome analysis and immunostaining methods cannot detect all hypermutant cancers.

Method used

By constructing a discrimination system for hypermutant cancer, a discrimination model is generated based on the pathological slide image data stained with hematoxylin and eosin, using a machine learning model of image data to quickly and accurately identify hypermutant cancer.

Benefits of technology

It enables rapid and high-precision identification of hypermutant cancers, simplifies the detection process, improves detection efficiency and accuracy, and can guide effective treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115684569B_ABST
    Figure CN115684569B_ABST
Patent Text Reader

Abstract

The present application provides a hypermutated cancer discrimination system, program and method. A hypermutated cancer discrimination method, program and method with higher accuracy than before are provided. According to the present application, a hypermutated cancer discrimination system is provided, characterized by comprising an input unit configured to be able to input a plurality of first image data, a plurality of second image data and a plurality of third image data, a holding unit configured to be able to hold the first image data and the second image data, a machine learning execution unit configured to be able to generate a discrimination model for discriminating whether or not a cancer is hypermutated by using the first image data and the second image data held by the holding unit as teaching data, and a discrimination unit configured to be able to input the third image data to the discrimination model and discriminate whether or not the third image data is hypermutated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Invention Patent Application No. 201980013657.2, filed on February 7, 2019, entitled "System, Program, and Method for Identifying Hypermutated Cancer."

TECHNICAL FIELD

[0002] The present application relates to a system, program, and method for identifying hypermutated cancer.

BACKGROUND

[0003] Through extensive research on genetic mutations of cancer, it has been gradually known that cancer can be classified according to the pattern of genetic mutations. As one of such characteristic mutations of cancer, there is hypermutation. Cancer of hypermutation is distinguished by having a higher somatic mutation rate than other forms. It is known that gastric cancer, breast cancer, colorectal cancer, glioblastoma, uterine cancer, and the like exhibit characteristics of hypermutation. Cancer of hypermutation has, in many cases, the property of microsatellite instability that exhibits a defect or deficiency in the mismatch repair mechanism during DNA replication. It is believed that this is because the genes of MLH1, MLH3, MSH2, MSH3, MSH6, PMS2, which are mismatch repair enzymes, are mutated or the MLH1 gene is inhibited by methylation. Furthermore, it is known that due to mutations in the polymerase epsilon (POLE), which is a DNA replication enzyme, somatic mutations occur at an ultra-high frequency, resulting in hypermutation (Non-Patent Literature 1, 2).

[0004] On the other hand, cancer immune evasion mechanisms have been elucidated, and new cancer immunotherapies targeting these mechanisms are clinically available. Among them, the characteristic is the PD-1 (Programmed cell Death-1) / PD-L1 (PD-1 Ligand 1) pathway, also known as the immune checkpoint pathway. By blocking the immune suppressive co-signaling PD-1 / PD-L1 pathway, the immune suppression of T cells is lifted, and T cell activation occurs, leading to suppression of tumors expressing cancer-specific antigens. Furthermore, CTLA-4 is also expressed on activated T cells, and if the CD28 ligand of the antigen-presenting cell binds, it will inhibit the activation of T cells, so even if this pathway is blocked, the immune suppression of T cells can be lifted, leading to tumor suppression. Anti-cancer agents using this principle have been put into use (e.g., Nivolumab, Ipilimumab).

[0005] In addition, there are several such immune inhibitory mechanisms, and anti-tumor agents that block these immune inhibitory mechanisms are expected to be gradually developed and put into use in the future. Cancer of hypermutation has a higher number of cancer-specific antigens that become targets of immune mechanisms, and therefore exhibits a higher effect of therapies that block the signaling pathways of immune suppression, and there is a need for a method that can simply identify whether cancer is hypermutated (Non-Patent Literature 3).

[0006] In the past, when checking hypermutated cancer, a method of performing comprehensive gene analysis to calculate the number of mutations is known, but there is a problem that a large amount of effort and time is required for detection. Also, a method of detecting a defect or deficiency of a mismatch repair mechanism, which is one of the causes of hypermutation by cancer, by immunostaining of related genes or microsatellite instability test is known, but this method has a problem that all hypermutated cancers cannot be checked.

[0007] On the other hand, a pathological diagnosis support program as disclosed in Patent Literature 1 is known to exist.

[0008]

Prior Art Documents

[0009]

Non-Patent Literature

[0010]

Non-Patent Literature 1

[0011]

Non-Patent Literature 2

[0012]

Non-Patent Literature 3

[0013]

Patent Literature

[0014]

Patent Literature 1

Summary of Invention

[0015]

Problems to be Solved by the Invention

[0016] In Patent Literature 1, it is mentioned that the presence or absence of a tumor, benignity or malignancy can be determined, but there is no mention of a method of discriminating hypermutated cancer.

[0017] The present application was completed in view of such circumstances, and provides a method of discriminating hypermutated cancer with higher precision than in the past, a program, and a method.

[0018]

Means for Solving the Problems

[0019] According to the present invention, a system for identifying hypermutant cancer is provided, comprising an input unit, a holding unit, a machine learning execution unit, and a discrimination unit. The input unit is configured to input multiple first image data, multiple second image data, and multiple third image data. The first image data are image data representing pathological sections of stained hypermutant cancer. The second image data are image data representing pathological sections of non-hypermutant cancer, and are pathological sections with the same staining as the pathological sections of the cancer from which the first image data originates. The third image data are image data representing newly stained... The image data of a pathological section of cancer that is a source of supermutant cancer, and the pathological section of cancer with the same staining as the source of the first image data, are used to determine whether the cancer is a supermutant cancer. The holding unit is configured to hold the first image data and the second image data. The machine learning execution unit is configured to use the first image data and the second image data held by the holding unit as teaching data to generate a discrimination model for determining whether the cancer is a supermutant cancer. The discrimination unit is configured to input the third image data into the discrimination model and determine whether the third image data is a supermutant cancer.

[0020] According to the present invention, a discriminant model for determining whether a cancer is a hypermutant carcinoma is generated by using first image data and second image data as teaching data. Here, the first image data is image data representing a pathological section of a stained hypermutant carcinoma. The second image data represents image data of a pathological section that is not a hypermutant carcinoma and has the same staining as the pathological section of the cancer from which the first image data originated. Furthermore, the system is configured to input third image data into the discriminant model to determine whether the third image data represents a hypermutant carcinoma. Here, the third image data is image data of a pathological section of a newly determined cancer for which the determination of whether it is a hypermutant carcinoma is being performed, and has the same staining as the pathological section from which the first image data originated. Therefore, the determination of hypermutant carcinoma can be performed rapidly and with high accuracy, and drugs that are effective for treatment can be easily administered. Previously, the determination of hypermutant carcinoma was very difficult without gene analysis using next-generation sequencing instruments or the like.

[0021] The following describes various embodiments of the present invention. These embodiments can be combined with each other.

[0022] Preferably, the staining method for the pathological sections is hematoxylin and eosin staining.

[0023] Preferably, the input section is configured to also be able to input non-cancer image data that is image data of a non-cancer pathological section, the holding section is configured to also be able to hold the non-cancer image data, the machine learning execution section is configured to also be able to use the non-cancer image data held by the holding section as teaching data, generate a discrimination model that discriminates image data of a pathological section that is cancer, and the discrimination section is configured to also be able to discriminate whether the third image data is image data of a cancer.

[0024] Preferably, an image processing section is provided, which is configured to perform Z value conversion processing that converts each color of RGB in each pixel into a Z value in CIE color system, on at least one of the first image data, the second image data, and the non-cancer image data, based on a color distribution of the entire first image data, the entire second image data, or the entire non-cancer image data.

[0025] Preferably, the image processing section is configured to be able to perform split processing that splits at least one of the first image data, the second image data, and the non-cancer image data input to the input section.

[0026] Preferably, the split processing is configured to perform split processing that splits image data of the same pathological section, on at least one of the first image data and the second image data.

[0027] Preferably, the image processing section performs the split processing in a manner in which a partial region is repeated in the split image.

[0028] Preferably, the image processing section is configured to also be able to perform split processing that splits the third image data input to the input section.

[0029] Preferably, the discrimination section discriminates whether the third image data is image data of a pathological section that is cancer, and further discriminates whether it is a hypermutated cancer, for image data that is discriminated as being image data of a pathological section that is cancer.

[0030] Preferably, the discrimination section discriminates whether the cancer is a hypermutated cancer, based on a ratio of image data that is discriminated as being image data of a hypermutated cancer, in image data that is discriminated as being image data of a pathological section that is cancer.

[0031] According to another aspect, there is provided a program that causes a computer to function as an input unit configured to be able to input a plurality of first image data and a plurality of second image data, the first image data being image data representing a pathological section of a hypermutated cancer, the second image data being image data representing a pathological section of a cancer that is not a hypermutated cancer and that is the same pathological section as a cancer of which the first image data is derived, a holding unit configured to be able to hold the first image data and the second image data, and a machine learning execution unit configured to be able to generate a discrimination model that discriminates whether or not a cancer is a hypermutated cancer using the first image data and the second image data held by the holding unit as teaching data.

[0032] According to another aspect, there is provided a method of discriminating a hypermutated cancer using the system according to any one of the above aspects.

[0033] According to another aspect, there is provided a method of discriminating a hypermutated cancer using the program according to any one of the above aspects.

[0034] Preferably, the process includes judging the effectiveness of an immune checkpoint inhibitor. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 is a functional block diagram of the system 10 according to the first embodiment of the present application.

[0036] Figure 2 is a flowchart showing a flow of generating a discrimination model that discriminates whether or not a cancer is a hypermutated cancer according to the first embodiment of the present application.

[0037] Figure 3 is a schematic diagram of a flow of the third image data when the third image data is discriminated whether or not it is a hypermutated cancer using the discrimination model.

[0038] Figure 4 is a flowchart showing a flow of discriminating whether or not a cancer is a hypermutated cancer according to the first embodiment of the present application.

[0039] Figure 5 is a schematic diagram for explaining the analysis in S13 of Figure 4

[0040] Figure 6 is a schematic diagram for explaining the weight w in the discrimination model.

[0041] Figure 7 is a schematic diagram of the execution of the machine learning in S5 of Figure 2

[0042] Figure 8 ​​is a functional block diagram of the system 20 according to the second embodiment.

[0043] Figure 9 is a description of the split processing of the input image of the image processing section 22.

[0044] Figure 10 is a diagram that describes the split processing of the input image of the image processing section 22.

[0045] Figure 11 is a processing flow of the discrimination processing of the third image data in the present embodiment.

[0046] Figure 12 is a diagram that describes the discrimination processing of the discrimination section 25.

DETAILED DESCRIPTION

[0047] Hereinafter, an embodiment of the present application will be described with reference to the drawings. The various features shown in the following embodiment can be combined with each other.

[0048] <1. First Embodiment>

[0049] 1.1. Discrimination of whether or not a cancer is hypermutated

[0050] Hereinafter, a system 10 according to one embodiment of the present application will be described with reference to the drawings. Figures 1-4 A system 10 according to one embodiment of the present application will be described with reference to the drawings.

[0051] (1.1.1. System 10)

[0052] As shown in FIG. 1, the system 10 is provided with an input section 1, an image processing section 2, a holding section 3, a machine learning execution section 4, and a discrimination section 5. Figure 1

[0053] The input section 1 is configured to be able to output a plurality of first image data, a plurality of second image data, and a plurality of third image data. Here, the first image data is image data of a pathological section of a cancer that is hypermutated. Also, the second image data is image data of a pathological section of a cancer that is not hypermutated, and is the same pathological section as that of which the image data is the source of the first image data. In addition, the third image data is image data of a pathological section of a cancer for which discrimination of whether or not the cancer is hypermutated is newly performed, and is the same pathological section as that of which the image data is the source of the first image data. Here, in the present embodiment, the RGB values of these image data can take values of 0 to 255.

[0054] ​In the present embodiment, 17 cases each of pathological tissue staining specimens of large intestine cancer samples judged as Hypermutation type (hypermutation type) and Non-Hypermutation type (non-hypermutation type) by analysis of cancer genome DNA sequences were obtained. Here, such 17 cases are the results of cancer genome sequencing of 201 Japanese large intestine cancer patients, and 17 cases judged as Hypermutation (reference: Nagahashi et al Genome Med 2017). In addition, the pathological tissue staining specimens of large intestine cancer subjected to hematoxylin and eosin staining of such specimens were made as the first image data and the second image data using a digital pathology technique. Here, in the present embodiment, the first image data and the second image data were saved as digital pathology image data according to the MIRAX format. Here, the above conditions are not limited thereto, and a cancer sample other than large intestine cancer of a predetermined number of cases can be obtained.

[0055] Thus, in the present embodiment, since the image data of hematoxylin and eosin staining with a large number of clinical cases is used as the first image data and the second image data, a discrimination system with high versatility can be realized.

[0056] Among them, other staining methods can be used according to the conditions. In addition, the saving format of the image data is not limited thereto.

[0057] The image processing section 2 is configured to be able to perform a split processing of splitting the plurality of first image data, second image data, and third image data input to the input section 1. In the present embodiment, the image processing section 2 has a function of splitting the first image data, the second image data, and the third image data into a predetermined tile. As one example, the first image data, the second image data, and the third image data are split into images of 300 pixel x 300 pixel size by the image processing section 2. It should be noted that such a split size is not particularly limited, and a size in which whether the image data is a cancer tissue portion can be recognized is preferable. Furthermore, in the present embodiment, the first image data and the second image data are each split into 1000 or more by the split processing. In addition, in the present embodiment, the image processing section 2 is configured to be able to perform a split processing of splitting the image data of the same pathological section for at least one of the first image and the second image data. It should be noted that the split size and the number of splits are not particularly limited, and any conditions can be used.

[0058] Thus, by splitting the image data input to the input section 1, it is possible to increase the number of teaching data for subsequent machine learning, and it is possible to improve the accuracy of machine learning.

[0059] In the present embodiment, the image processing section 2 is configured to also be able to perform conversion processing that converts each of the colors of RGB to a Z value in the CIE color system based on the color distribution of the entire first image data and second image data for the split first image data and second image data. Specifically, since the Z value takes a normal distribution centered on 0 and the RGB values of the image data are values of 0 to 255, it is preferable to control the values of the Z values of the respective colors of RGB to be within a range of 2 times the standard deviation (σ). Therefore, the image processing section 2 has a function of correcting values of 2σ or more to 2σ and values of -2σ or less to -2σ. Further, the image processing section 2 has a function of multiplying these values by 2 to convert all of the values to values of 0 or more and then dividing by 4 to normalize to values of 0 to 1. In addition, the image processing section 2 has a function of converting to values of normal color expression by multiplying these values by 255. In addition to this, the image processing section 2 also performs processing of rounding off the decimal places to make these values integer values. Note that the method of normalization is not limited to this.

[0060] Here, if defined as "x = int(((min(max(xz, -2), 2) + 2) / 4) x 255)", "xz = the value of the Z value of RGB" holds.

[0061] In this way, by converting each of the colors of RGB of the first image data and second image data to a Z value, it is possible to reduce the unevenness of the color tone (the gradation of the color) in the dyeing processing and to suppress the influence of the degree of dyeing on the subsequent machine learning. As a result, it is possible to improve the accuracy of the machine learning.

[0062] The holding section 3 is configured to be able to hold the first image data and second image data. Here, the holding section 3 is constituted by any of a storage, a latch, an HDD, an SSD, or the like.

[0063] The machine learning execution section 4 is configured to generate a discrimination model that discriminates whether or not it is a hypermutated cancer using the first image data and second image data held by the holding section 3 as teaching data. The machine learning execution section 4 is constituted by any of a CPU, a GPU, a DSP, or the like. Figure 3 The discrimination model will be described in detail below.

[0064] The machine learning algorithm of the machine learning execution section 4 is not particularly limited, and for example, a neural network or deep learning can be used. Also, for example, a CNN (Convolutional Neural Network) for image recognition called "Inception-v3" developed by Google Inc. can be used. In addition, the CNN can be executed using a "Keras" framework. Also, regarding the machine learning itself, to prevent overlearning, an "Early Stopping" method can be used, which is a method of calculating the accuracy of the model in learning per 1 epoch using images of a validation set, and stopping the learning at the epoch at which the variation in the accuracy index converges. Note that in the present embodiment, the machine learning is repeatedly performed for 14 epochs in the learning of Z values.

[0065] The discrimination section 5 is configured to be able to input the third image data to the discrimination model, and to discriminate whether the third image data is a hypermutated cancer or not.

[0066] (1.1.2. Flowchart)

[0067] Next, the machine learning is performed using the first image data and the second image data as the learning data. Figure 2 The flow of generating the discrimination model for discriminating whether or not it is a hypermutated cancer according to one embodiment of the present application will be described.

[0068] First, the first image data and the second image data are input to the input section 1 in S1.

[0069] Next, in S2, the split processing of splitting the first image data and the second image data is performed by the image processing section 2. In the present embodiment, the first image data and the second image data are each split into 1000 or more tiles. Note that the number of such splits can be appropriately set, for example, to 1000 to 3000, preferably 1000 to 2000, and more preferably 1000 to 1500. Specifically, for example, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, or any range between any two of the values exemplified here can be used.

[0070] In addition, Z-value processing is performed by the image processing section 2, which Z-values the split first image data and the second image data.

[0071] Next, in S3, for the Z-valued first and second image data, labels are added to each image data according to whether the image data is a hypermutation type or a non-hypermutation type of cancer tissue.

[0072] Next, in S4, the first and second image data of the 17 case samples input to input unit 1 are selected to utilize the image data of 13 case samples for machine learning by machine learning execution unit 4. This selection can be performed randomly or by a pathologist specializing in cancer. Furthermore, the tagged first and second image data are held in holding unit 3. This first and second image data becomes the "positive solution set" in machine learning.

[0073] Next, in S5, the machine learning execution unit 4 uses the first image data and the second image data held by the holding unit 3 as teaching data to perform machine learning in order to generate a discriminative model for determining whether a cancer is a hypermutant cancer. Specifically, using the first image data and the second image data of 13 case samples with labels added in S4, machine learning is performed to determine whether such image data is a hypermutant cancer.

[0074] Next, in S6, it is determined whether the discrimination model's accuracy is above the specified accuracy. If the discrimination model's accuracy is not above the specified accuracy (NO), the process returns to S4, and image data from 13 different combinations of the first and second image data from the 17 case samples are selected, and the processing in S5 is specified. On the other hand, if the discrimination model's accuracy is above the specified accuracy (YES), such a discrimination model is adopted, and S7 is executed.

[0075] Finally, in S7, the discrimination unit 5 outputs the discrimination model determined in S6 and stores it in the holding unit 3 or a storage unit not shown.

[0076] (1.1.3. Determination of whether it is a hypermutant cancer)

[0077] Next, using Figure 3 and Figure 4 This describes the process of using a discriminant model to determine whether a third image is a hypermutant cancer.

[0078] like Figure 3 As shown, in this embodiment, the third image data input to the input unit 1 is output to the image processing unit 2, and the third image data that has undergone the above-described image processing (splitting processing and Z-value processing) is output to the discrimination unit 5. Furthermore, the discrimination unit 5 utilizes... Figure 2 The discriminant model output from S7 determines whether the third image data is a hypermutant cancer.

[0079] Thus, by also splitting the third image data, the size of the image data of the object to be judged is made to be consistent with the size of the first and second image data, thereby improving the judgment accuracy of the judgment unit 5.

[0080] The flowchart at this point is as follows.

[0081] like Figure 4 As shown, in S11, the third image data is first input to the input unit 1.

[0082] Next, in S12, image processing (splitting processing and Z-value conversion processing) is performed by the image processing unit 2.

[0083] Next, in S13, the discrimination unit 5 uses the above-mentioned discrimination model to determine whether the third image data is a supermutant cancer.

[0084] Finally, in S14, the discrimination result of the discriminant unit 5 is output. There is no particular limitation on the way such discrimination result is output, and it can be "is a supermutant cancer", "is not a supermutant cancer", "the probability of being a supermutant cancer is X%", etc.

[0085] (1.1.4. Discrimination by the discriminant model)

[0086] Next, using Figure 5 and Figure 6 Instructions for use Figure 4 The discrimination model in S13 makes the discrimination. It should be noted that in this embodiment, the machine learning algorithm is not particularly limited, and neural networks or deep learning can be used. Hereinafter, for ease of explanation, an example using a neural network will be described.

[0087] like Figure 5 As shown, the neural network (hereinafter labeled NN in the diagram) consists of multiple layers (layer 1 L1 to layer 3 L3) and multiple computing nodes N (N11 to N31). Here, Nij represents the j-th computing node N in the i-th layer. In this embodiment, the neural network is constructed with i = 3 and j = 5. It should be noted that the values ​​of i and j are not limited to these; for example, they can be integers between i = 1 and 100 and j = 1 and 100, or integers greater than 100.

[0088] Furthermore, each computing node N is assigned a predetermined weight w. For example... Figure 4 As shown, for example, when focusing on the computation node N23 of the second layer, a weight w is set between computation node N23 and the full computation nodes N11 to N15 of the first layer, which are the preceding layers. The weight w is set to a value of -1 to 1, for example.

[0089] The machine learning execution unit 4 outputs various parameters to the neural network. In this embodiment, the parameters input to the neural network are the Z-value of the third image data, the distribution of the Z-value of the third image data, the difference between the Z-value of the third image data and the Z-value of the first image data, the difference between the Z-value of the third image data and the Z-value of the second image data, and the difference between the Z-value of the third image data and the distribution of the Z-values ​​of the first and second image data. Here, the Z-values ​​of the first to third image data are Z-values ​​in pixel units. Furthermore, the distribution of the Z-values ​​of the first to third image data is the distribution of Z-values ​​in the image data (300 pixels × 300 pixels). Furthermore, the difference between the distribution of the Z-values ​​of the third image data and the distribution of the Z-values ​​of the first and second image data is the sum of the differences between the distribution of the Z-values ​​of the third image data and the distribution of the Z-values ​​of the corresponding pixels in the first and second image data, or the sum of the differences of the Z-values ​​of the corresponding pixels within the image data.

[0090] Here, as described above, when the parameters are input into the neural network, they are normalized to values ​​between 0 and 1. For example, when the input parameter is 0, it is input as the input signal 0. And when the input parameter is 1, it is input as the input signal 1.

[0091] Furthermore, the discrimination unit 5 inputs input signals defined by various parameters to the first layer L1. These input signals are output from the first layer's computing nodes N11 to N15 to the second layer's computing nodes N21 to N25. At this time, the value output from computing nodes N11 to N15 is multiplied by the weight w set for each computing node N, and the resulting value is input to computing nodes N21 to N25. Computing nodes N21 to N25 sum the input values ​​and add the sum to the input values. Figure 6 The value of the bias b shown is input to the activation function f(). Additionally, the output value of the activation function f() (in...) Figure 6 In the example, the output value from the virtual computing node N'23 is transmitted to the next computing node N31. At this time, the value obtained by multiplying the weight w set between computing nodes N21-N25 and computing node N31 by the aforementioned output value is input to computing node N31. Computing node N31 adds the input values ​​and outputs the total as an output signal. Alternatively, computing node N31 can add the input values ​​and input the biased value to the activation function f(), outputting its output as an output signal. Here, in this embodiment, the output signal value is adjusted to a value between 0 and 1. Furthermore, the machine learning execution unit 4 outputs the value corresponding to the output signal value as a probability of determining whether it is a hypermutant cancer.

[0092] As described above, in this embodiment, the system 10 uses the first image data and the second image data as teaching data and performs machine learning by the machine learning execution unit 4 to generate a discrimination model (neural network and weights w) for determining whether a cancer is a hypermutant cancer. Furthermore, the discrimination unit 5 uses this discrimination model to determine whether the third image data is a hypermutant cancer.

[0093] 1.2. Generation of the discriminant model

[0094] Next, using Figure 7 illustrate Figure 2 The generation of the discriminant model in S5 to S6.

[0095] like Figure 7 As shown, the machine learning execution unit has 4 pairs of components and Figure 5 Each computation node N of the neural network with the same structure as the neural network shown is assigned a weight w, for example, from -1 to 1. In this case, to reduce the influence of the weight w, it is preferable that the absolute value of the initially set weight w is small. Furthermore, five types of parameter sets are input into the neural network. In this embodiment, the parameters input into the neural network are the Z-values ​​of the first image data, the Z-values ​​of the second image data, the distribution of the Z-values ​​of the first image data, the distribution of the Z-values ​​of the second image data, and the difference between the Z-values ​​of the first and second image data. Here, the Z-values ​​of the first and second image data are the Z-values ​​of pixel units. Furthermore, the distribution of the Z-values ​​of the first and second image data are the distribution of Z-values ​​within the image data (300 pixels × 300 pixels). Furthermore, the difference between the Z-values ​​of the first and second image data is either the difference between the Z-values ​​of each corresponding pixel in the first and second image data or the sum of the differences between the Z-values ​​of each corresponding pixel within the image data.

[0096] Furthermore, the output signal from the neural network is compared with the teaching data. When the difference between the output signal and the teaching data (hereinafter referred to as the error) is above a predetermined threshold, the weights w are changed, and the five sets of parameters are input into the neural network. The change of weights w is performed using a known error propagation method. This calculation (machine learning) is repeated to minimize the error between the output signal from the neural network and the pre-given teaching data. The number of times machine learning is performed is not particularly limited; for example, it can be from 1000 to 20000 times. Even if the error between the actual output signal and the pre-given teaching data is not minimized, machine learning can be terminated when the error falls below the predetermined threshold or at any time by the developer.

[0097] Furthermore, if machine learning by the machine learning execution unit 4 ends, the machine learning execution unit 4 sets the weights of each computing node N in the neural network at that time. That is, in this embodiment, the weights w are stored in the set... Figure 5 The storage unit, such as the memory on the neural network, is used for storage. Additionally, the weights w set by the machine learning execution unit 4 are sent to a storage unit (not shown) provided in the system 10, becoming... Figure 5 The weights w of each computation node N in the neural network. Here, let... Figure 7 The structure of neural networks and Figure 5 The structure of the neural network is the same, so the weights w set by the machine learning execution unit 4 can be used.

[0098] <2. Second Implementation>

[0099] use Figures 8-12 The second embodiment of the present invention will be described below. It should be noted that structures and functions identical to those in Embodiment 1 will not be described again.

[0100] like Figure 8 As shown, in the system 20 of the second embodiment, in addition to the first image data and the second image data, the input unit 21 is configured to also input non-cancer image data. Here, non-cancer image data refers to image data other than pathological sections of cancer. The image processing unit 22 performs segmentation processing on the input image data. The segmentation processing will be described in detail later.

[0101] In addition to the split first and second image data, the holding unit 23 is configured to also hold the split non-cancer image data. The machine learning execution unit 24 is configured to use the first image data, the second image data, and the non-cancer image data held by the holding unit 3 as teaching data to generate a discrimination model for determining whether an image is cancerous (hereinafter referred to as the first discrimination model) and a discrimination model for determining whether a cancerous image is a hypermutant cancer (hereinafter referred to as the second discrimination model). The discrimination unit 25 is configured to input the third image data into the first and second discrimination models to determine whether the third image data is cancerous image data and whether it is hypermutant cancerous image data.

[0102] Figure 9 Image data P is represented as an example of input to input unit 21. Image data P has a tissue region T and a blank region BL (e.g., a region on a glass slide). The tissue region T includes a cancer region C1 of non-hypermutant carcinoma, a cancer region C2 of hypermutant carcinoma, and a non-cancerous tissue region NC.

[0103] The image processing unit 22 performs splitting processing on the image data P input to the input unit 21. Figure 9In the illustrated example, the tissue region T is split into 100 pieces vertically 10 times and horizontally 10 times. That is, pieces D of 100 pieces are set in a manner that includes the tissue region T 00 ~D 99 .

[0104] In this example, pieces corresponding to the region C2 of the hypermutated cancer (for example, piece D54) correspond to the first image data, and pieces corresponding to the region Cl of the cancer of the non-hypermutated cancer (for example, piece D34) correspond to the second image data. Also, pieces corresponding to only the non-cancerous tissue region NC (for example, piece D15), pieces corresponding to only the blank region BL (for example, piece D49), and pieces including the non-cancerous tissue region NC and the blank region BL (for example, piece D04) all correspond to non-cancer image data.

[0105] Thus, in the present embodiment, various images such as pieces corresponding to the non-cancerous tissue region NC, pieces corresponding to only the blank region BL, and pieces including the non-cancerous tissue region NC and the blank region BL are input as non-cancer image data to perform machine learning. By thus increasing the diversity of the non-cancer image, the accuracy of determining whether the examination target data is a cancer image is improved.

[0106] Also, in the present embodiment, further split processing (hereinafter referred to as second split processing) can be performed on the image data after the above-described split processing (hereinafter referred to as first split processing). In the present embodiment, the second split processing is performed on the pieces Dnm after the first split processing. Figure 10 In the present embodiment, the split pieces Dnm are further split into five pieces by the first split processing. Here, in the second split processing, the split processing is performed in a manner that the split pieces have partially overlapping regions. That is, the pieces Dnm1 after the second split processing partially overlap with the pieces Dnm2. Also, the pieces Dnm2 partially overlap with the pieces Dnm3.

[0107] Thus, by performing the split processing in a manner that the split pieces have partially overlapping regions, the number of images can be increased, and the learning efficiency in the subsequent machine learning can be improved.

[0108] Figure 11 A processing flow of the discrimination processing of the third image data in the present embodiment is shown. As shown in FIG. 27, in the present embodiment, the discrimination unit 25 performs discrimination of whether the third image data is a cancer image and discrimination of whether it is a hypermutated cancer. Figure 11

[0109] Specifically, in step S231 in step S23, the discrimination unit 25 performs discrimination of whether the third image data is a cancer image. In the case where it is not a cancer image (No in step S231), the third image data is discriminated as a non-cancer image in step S233. ​

[0110] On the other hand, in the case of being a cancer image (Yes in step S231), the discrimination unit 25 discriminates whether the 3rd image data is an image of a hypermutated cancer in step S232. In the case of not being a hypermutated cancer (No in step S232), it is discriminated that the 3rd image data is not an image of a hypermutated cancer in step S235. On the other hand, in the case of being a hypermutated cancer (Yes in step S232), it is discriminated that the 3rd image data is an image of a hypermutated cancer in step S234.

[0111] Thus, in the present embodiment, discrimination whether the 3rd image data is a cancer image and discrimination whether it is a hypermutated cancer are performed. Therefore, it is not necessary to previously diagnose whether the image data is a cancer image by a pathologist or the like, and the work efficiency in the discrimination process can be improved.

[0112] Here, the discrimination unit 25 can discriminate whether the cancer is a hypermutated cancer based on the ratio of the image data discriminated to be an image of a hypermutated cancer to the image data discriminated to be a cancer.

[0113] In Figure 12 In the example shown, in the 3rd image data P2, there is an image E2 discriminated to be an image of a hypermutated cancer within an image E1 discriminated to be an image of a cancer. At this time, when the ratio defined by the number of tiles of E2 / (the number of tiles of E1) is larger than a predetermined threshold value, the discrimination unit 25 discriminates that the region shown by the image E1 is a hypermutated cancer.

[0114] Thus, it is possible to exclude false positives of discriminating a hypermutated cancer as being local as interference information, and thus it is possible to improve the accuracy of discrimination.

[0115] Thus, in the 2nd embodiment, the input unit 21 is configured to be able to further input non-cancer image data and cancer image data, and the machine learning execution unit 24 is configured to be able to also use the non-cancer image data and the cancer image data as teaching data, and further generate a discrimination model that discriminates image data of a pathological section whether it is a cancer. In addition, the discrimination unit 25 is configured to be able to further discriminate whether the 3rd image data is cancer image data. With such a configuration, with respect to the 3rd image data, it is not necessary to diagnose whether it is a cancer by a pathologist or the like, and thus the work efficiency in the discrimination process is improved.

[0116] <3. Other Embodiments>

[0117] The above describes various embodiments, but the present application can also be implemented in the following manner.

[0118] A program which causes a computer to function as an input section, a holding section, a machine learning execution section, and a discrimination section, the input section being configured to be able to input a plurality of first image data and a plurality of second image data, the first image data being image data representing a pathological section of a hypermutated cancer, the second image data being image data representing a pathological section of a cancer which is not a hypermutated cancer and which is the same pathological section as a cancer of which the first image data is derived, the holding section being configured to be able to hold the first image data and the second image data, the machine learning execution section being configured to be able to generate a discrimination model which discriminates whether or not a cancer is a hypermutated cancer using the first image data and the second image data held by the holding section as teaching data.

[0119] A method of discriminating a hypermutated cancer using the system described above. Note that the hypermutated cancer described here includes any cancer, such as, for example, solid cancers such as brain tumors, head and neck cancers, breast cancers, lung cancers, esophageal cancers, stomach cancers, duodenal cancers, appendix cancers, large intestinal cancers, rectal cancers, liver cancers, pancreatic cancers, gallbladder cancers, bile duct cancers, anal cancers, kidney cancers, ureter cancers, bladder cancers, prostate cancers, penile cancers, testicular cancers, uterine cancers, ovarian cancers, vulvar cancers, vaginal cancers, skin cancers, and the like, but is not limited thereto. For the purposes of the present application, the hypermutated cancer is preferably a large intestinal cancer, a lung cancer, a stomach cancer, a melanoma (malignant melanoma), a head and neck cancer, or an esophageal cancer.

[0120] A method of discriminating a hypermutated cancer using the program described above.

[0121] The discrimination method according to any one of the above described aspects, including a process of judging the effectiveness of an immune checkpoint inhibitor. Such a discrimination method can further include a process of judging that a patient who is judged to have a hypermutated cancer exhibits high curative effectiveness of an immune checkpoint inhibitor. It has been proven that a hypermutated cancer has many cancer-specific antigens that become targets of an immune mechanism, and thus a therapy that blocks a signal pathway of immune suppression is very effective. In such a discrimination method, it is possible to easily judge that the cancer is hypermutated and thus is very advantageous. The "immune checkpoint" referred to herein is well known in the art (Naidoo et al. British Journal of Cancer (2014) 111, 2214-2219), and known examples include CTLA4, PD1, and its ligand PDL-1. Other examples include TIM-3, KIR, LAG-3, VISTA, and BTLA. Inhibitors of immune checkpoints inhibit these normal immune functions. For example, the expression of a molecule that controls an immune checkpoint is negatively controlled, or it is bound to a molecule, thereby inhibiting normal receptor / ligand interactions. Since the function of an immune checkpoint is to brake an immune system response to an antigen, an inhibitor thereof reduces this immune suppression effect and enhances an immune response. Inhibitors of immune checkpoints are well known in the art, and preferred examples include anti-CTLA-4 antibodies (e.g., ipilimumab and tremelimumab), anti-PD-1 antibodies (e.g., nivolumab, pembrolizumab, pidilizumab, and RG7446 (Roche)), and anti-PDL-1 antibodies (e.g., BMS-936559 (Bristol-Myers Squibb), MPDL3280A (Genentech), MSB0010718C (EMD-Serono), and MEDI4736 (AstraZeneca)), and the like.

[0122] Also, the holding unit 3 can be a cloud computing method of an information processing device such as a PC or a server provided outside. In this case, the information processing device outside transmits necessary data to the system 10 each time of computation.

[0123] Also, it can be provided as a computer-readable non-transitory recording medium in which the above-described program is stored. Further, an ASIC (application specific integrated circuit), an FPGA (field-programmable gate array), and a DRP (Dynamic Reconfigurable Processor) that realize the functions of the above-described program can be provided.

[0124] [Explanation of symbols]

[0125] 1, 21: input section

[0126] 2, 22: image processing section

[0127] 3, 23: holding section

[0128] 4, 24: machine learning execution section

[0129] 5, 25: discrimination section

[0130] 10, 20: system

Claims

1. A cancer detection system, characterized in that, It includes an input unit, a holding unit, a machine learning execution unit, a discrimination unit, and an image processing unit. The input unit is configured to accept multiple cancer image data, multiple non-cancer image data, and multiple discrimination object image data. The cancer image data refers to image data representing pathological sections of stained cancer cells. The non-cancerous image data refers to image data of pathological sections that represent non-cancerous tissue sections and have the same staining as the pathological sections from which the cancerous image data originated. The image data for discrimination is image data representing a newly determined pathological section for cancer discrimination, and the pathological section with the same staining as the pathological section from which the cancer image data originated. The holding unit is configured to hold the cancer image data, the non-cancer image data, and the discrimination object image data. The image processing unit is configured to perform a conversion process on the cancer image data, the non-cancer image data, and the discrimination object image data, based on the overall color distribution of the cancer image data, the non-cancer image data, and the discrimination object image data, converting each RGB color in each pixel into a Z value in the CIE color system. The machine learning execution unit is configured to use the cancer image data and the non-cancer image data, which have undergone the transformation process and are held by the holding unit, as teaching data to generate a discrimination model for determining whether a pathological slice image data is cancerous. The discrimination unit is configured to input the image data of the discrimination object that has undergone the transformation process into the discrimination model, and to determine whether the image data of the discrimination object is image data of a pathological section of cancer.

2. The system according to claim 1, characterized in that, The pathological sections were stained with hematoxylin and eosin.

3. The system according to claim 1 or 2, characterized in that, The image processing unit is configured such that, It is capable of performing a splitting process that splits at least one of the cancer image data, the non-cancer image data, and the discrimination object image data input to the input unit.

4. The system according to claim 3, characterized in that, The image processing unit performs the splitting process by repeating the splitting process in a partial region of the split image.

5. A computer-readable recording medium containing a program that enables a computer to function as an input unit, a storage unit, a machine learning execution unit, a discrimination unit, and an image processing unit. The input unit is configured to accept multiple cancer image data, multiple non-cancer image data, and multiple discrimination object image data. The cancer image data refers to image data representing pathological sections of stained cancer cells. The non-cancerous image data refers to image data of non-cancerous pathological sections, and of pathological sections with the same staining as the source pathological sections from which the cancerous image data originates. The image data for discrimination is image data representing a newly determined pathological section for cancer discrimination, and the pathological section with the same staining as the pathological section from which the cancer image data originated. The holding unit is configured to hold the cancer image data, the non-cancer image data, and the discrimination object image data. The image processing unit is configured to perform a conversion process on the cancer image data, the non-cancer image data, and the discrimination object image data, based on the overall color distribution of the cancer image data, the non-cancer image data, and the discrimination object image data, converting each RGB color in each pixel into a Z value in the CIE color system. The machine learning execution unit is configured to use the cancer image data and the non-cancer image data, which have undergone the transformation process and are held by the holding unit, as teaching data to generate a discrimination model for determining whether a pathological slice image data is cancerous. The discrimination unit is configured to input the image data of the discrimination object that has undergone the transformation process into the discrimination model, and to determine whether the image data of the discrimination object is image data of a pathological section of cancer.

Citation Information

Patent Citations

  • Method for controlling CNG engine based on fuel properties

    JP2004346911A

  • Image-based risk score-a prognostic predictor of survival and outcome from digital histopathology

    CN104376147A