Learning device, prediction device, learning method, and program

JPWO2024111084A5Active Publication Date: 2025-07-29NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024559796
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-29
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

Large image sizes in deep learning image classification and prediction tasks lead to inefficient learning due to small effective learning areas, resulting in reduced data diversity and model accuracy, especially when detailed labels are not provided for the entire image.

Method used

A learning device and method that generates partial images from input images, maps them into a feature space, selects effective regions for learning data, and updates probability distributions to iteratively improve the prediction model, allowing for accurate model learning and prediction while selecting effective areas as learning data.

Benefits of technology

This approach enables efficient and accurate model learning and prediction by selecting effective regions in the feature space, improving data diversity and model accuracy even without detailed labels for the entire image.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided is a training device, wherein a partial image generation means generates, from an input image, a partial image smaller than the input image. A feature space generation means generates, for each of input images, a feature space to which feature amounts of a plurality of partial images are mapped. A training data generation means, on the basis of the feature space, acquires a plurality of partial training images from the plurality of partial images and generates training data. A training means uses the training data to train a prediction model which predicts a probability that a prescribed feature is included in the partial training image. A prediction means uses the trained prediction model to perform the prediction on all of or a portion of partial images included in the input image. Furthermore, the training data generation means acquires, on the basis of the predicted values for all of or a portion of the partial images in the feature space, a plurality of partial training images to be used as the training data in the next training. In this way, the partial training images are updated and the training for the prediction model is repeated by the training means.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, prediction device, learning method, and recording medium

[0001] The present disclosure relates to a technique for predicting characteristic portions contained in an image.

[0002] A technique for classifying and predicting images using deep learning with neural networks is known. In so-called supervised learning, a model is trained using training data in which input images are labeled. Patent Literature 1 describes a method for tiling digital images of biological samples into a group of image patches and applying a classifier to classify the images.

[0003] Special Publication No. 2020-533725

[0004] When using deep learning to classify or predict images, if the image size is large and the area of ​​the image that is effective for learning is small, the probability of sampling effective data as training data decreases, resulting in inefficient learning. Also, if the image size is large and most of it is effective for learning, a large amount of similar data will be sampled as training data, reducing the diversity of the training data and potentially reducing the accuracy of the model obtained by learning.

[0005] One object of the present disclosure is to accurately train a model while selecting regions that are effective as training data in a situation where detailed labels are not given to the entire image.

[0006] In one aspect of the present disclosure, a learning device comprises: partial image generation means for generating, from an input image, partial images smaller than the input image; feature space generation means for generating, for each input image, a feature space into which feature amounts of a plurality of partial images are mapped; learning data generation means for acquiring a plurality of learning partial images from the plurality of partial images based on the feature space to generate learning data; learning means for using the learning data to train a prediction model that predicts the probability that a predetermined feature is included in the learning partial images; and prediction means for making predictions for all or some of the partial images included in the input image using the trained prediction model, wherein the learning data generation means acquires a plurality of learning partial images to be used as learning data in the next learning, based on predicted values ​​for all or some of the partial images in the feature space.

[0007] In another aspect of the present disclosure, a learning method is a computer-executed learning method, which includes: generating, from an input image, partial images smaller than the input image; generating, for each input image, a feature space in which feature amounts of multiple partial images are mapped; acquiring multiple partial training images from the multiple partial images based on the feature space to generate learning data; using the learning data, learning a prediction model that predicts the probability that the partial training images contain a predetermined feature; using the trained prediction model to make predictions for all or some of the partial images included in the input image; and generating the learning data by acquiring multiple partial training images to be used as learning data in the next learning based on predicted values ​​for all or some of the partial images in the feature space.

[0008] In yet another aspect of the present disclosure, a recording medium records a program that causes a computer to execute a process of: generating, from an input image, partial images that are smaller than the input image; generating, for each input image, a feature space in which feature amounts of multiple partial images are mapped; acquiring multiple training partial images from the multiple partial images based on the feature space to generate training data; using the training data to train a prediction model that predicts the probability that a predetermined feature is included in the training partial images; using the trained prediction model to make predictions for all or some of the partial images included in the input image; and generating the training data based on predicted values ​​for all or some of the partial images in the feature space, and acquiring multiple training partial images to be used as training data in the next training.

[0009] According to the present disclosure, even in a situation where detailed labels are not given to the entire image, it is possible to accurately train a model while selecting areas that are effective as training data.

[0010] 1 shows a learning device according to a first embodiment; FIG. 2 is a block diagram showing the hardware configuration of the learning device according to the first embodiment; FIG. 3 is a block diagram showing the functional configuration of the learning device; FIG. 4 is an explanatory diagram of the processing of an image division unit and a learning data generation unit; FIG. 5 shows an example of a partial image; FIG. 6 shows an example of dividing a feature space into a plurality of spatial regions by grid division; FIG. 7 is an explanatory diagram of a method of generating learning data; FIG. 8 is a flowchart of learning processing by the learning device; FIG. 9 shows an example of displaying region selection probabilities at the end of learning; FIG. 10 is a block diagram showing the functional configuration of a learning device according to a modified example; FIG. 11 is a diagram explaining processing by a learning device according to a modified example; FIG. 12 is a block diagram showing the functional configuration of a prediction device; FIG. 13 is a flowchart of prediction processing by the prediction device; FIG. 14 is a block diagram showing the functional configuration of a learning device according to a second embodiment; FIG. 15 is a flowchart of processing by the learning device according to the second embodiment.

[0011] Preferred embodiments of the present disclosure will be described below with reference to the drawings. <Basic Principle> The present disclosure provides an apparatus for predicting predetermined image features contained in an input image, in which effective partial images are selected based on a probability distribution in a feature space, thereby enabling efficient model learning. In the following embodiment, an example will be described in which a pathological tissue image of a patient receiving medication is input, and the medication effect is predicted based on the morphological characteristics of the pathological tissue.

[0012] Specifically, during training, the training device divides an input image into multiple partial images and trains a prediction model that makes predictions for each partial image. Here, the training device maps the multiple partial images onto a feature space and divides the feature space into multiple regions (hereinafter referred to as "spatial regions") with similar morphological features. The training device then acquires partial images from the multiple spatial regions according to probability distributions assigned to the multiple spatial regions to generate training data and trains the prediction model. This allows for efficient selection of effective partial images for training from the input image to train the prediction model. After training is performed and a trained model is obtained, the training device uses the trained prediction model to make predictions for each partial image included in the input image and updates the probability distributions for the multiple spatial regions based on the prediction results. The training device then acquires partial images from the spatial regions according to the updated probability distributions to generate training data and further trains the prediction model. In this way, a highly accurate prediction model is generated by repeatedly training the prediction model while updating the probability distributions assigned to the multiple spatial regions.

[0013] On the other hand, when making predictions (inferences) using a trained prediction model, the prediction device divides the input image into multiple partial images, performs predictions on the partial images using the prediction model, and integrates the prediction results for each partial image to obtain a prediction result for the input image. Furthermore, the prediction device can present parts of the input image that were important for the prediction based on the prediction results for each partial image.

[0014] 1 shows a learning device according to Embodiment 1. The learning device 100 learns a prediction model based on input image data (hereinafter also referred to as "input image").

[0015] 2 is a block diagram showing the hardware configuration of the learning device 100 according to the first embodiment. As shown in the figure, the learning device 100 includes an interface (IF) 12, a processor 13, a memory 14, a recording medium 15, a database (DB) 16, and a display device 17.

[0016] The IF 12 inputs image data used for training the prediction model. The processor 13 is a computer such as a CPU (Central Processing Unit) that controls the entire learning device 100 by executing a pre-prepared program. The processor 13 may be a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array). Specifically, the processor 13 executes the training process and prediction process described below.

[0017] The memory 14 is composed of a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The memory 14 stores various programs executed by the processor 13. The memory 14 is also used as a working memory while the processor 13 is executing various processes.

[0018] Recording medium 15 is a non-volatile, non-transitory recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from learning device 100. Recording medium 15 records various programs executed by processor 13. When learning device 100 executes various processes, the programs recorded on recording medium 15 are loaded into memory 14 and executed by processor 13.

[0019] DB 16 stores image data input via IF 12. Specifically, DB 16 stores data on input images used for learning by learning device 100. Display device 17 is, for example, a liquid crystal display device or a projector, and displays the prediction results made by learning device 100. In addition to the above, learning device 100 may also be equipped with input devices such as a keyboard and a mouse for the user to enter instructions and input.

[0020] 3 is a block diagram showing the functional configuration of the learning device 100. Functionally, the learning device 100 includes an image segmentation unit 21, a learning data generation unit 22, a learning unit 23, a prediction unit 24, a selection probability update unit 25, and an integration unit 26. The output of the integration unit 26 is supplied to the display device 17.

[0021] During learning, a learning dataset is prepared. In the following description, a set of one input image and a teacher label (hereinafter simply referred to as a "label") for that input image is referred to as learning data. A label for an input image is assigned to each input image. That is, one label is assigned to an entire input image, such that the label for one input image is a positive example and the label for another input image is a negative example. A collection of learning data for multiple input images is referred to as a learning dataset.

[0022] The learning data generation unit 22, the learning unit 23, the prediction unit 24, and the selection probability update unit 25 are configured to repeat a loop process (hereinafter referred to as a "learning loop") of generating learning data and learning a prediction model a predetermined number of times.

[0023] The image dividing unit 21 divides an input image included in a training dataset into partial images smaller than the input image. Hereinafter, the entire input image will also be referred to as the "whole image." FIG. 4 is an explanatory diagram of the processing performed by the image dividing unit 21 and the training data generating unit 22. As shown in FIG. 4, the image dividing unit 21 divides the whole image WI into multiple partial images PI and outputs them to the training data generating unit 22. In the example of FIG. 4, the whole image WI is an image of a patient's pathological tissue. However, the image dividing unit 21 only needs to generate partial images for the pathological tissue region, and does not need to generate partial images for the background region. In the example of FIG. 4, the image dividing unit 21 divides the whole image WI into multiple partial images PI using grid division. Alternatively, the whole image WI may be divided into multiple partial images PI centered on characteristic points.

[0024] The training data generation unit 22 uses the input partial images PI to generate training data for training a prediction model. Specifically, the training data generation unit 22 first converts each of the input partial images PI into feature quantities, performs dimensionality reduction, and maps the resulting feature quantities onto a two-dimensional feature space. FIG. 4 shows an example of mapping each partial image PI onto the feature space. One partial image PI is mapped as one point on the two-dimensional feature space. Examples of dimensionality reduction algorithms that can be used include t-distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP).

[0025] Next, the training data generation unit 22 divides the two-dimensional feature space into a plurality of regions (hereinafter also referred to as "spatial regions" to distinguish them from regions in the image). There are two methods for dividing the feature space into a plurality of spatial regions. In the first method, the training data generation unit 22 divides the feature space into a plurality of spatial regions by clustering. In the example of FIG. 4, the training data generation unit 22 divides the feature space into four spatial regions SA1 to SA4 by clustering. Note that a known clustering method such as kmeans can be used for the clustering.

[0026] Points located close to each other in the feature space mean that the image features of the partial images corresponding to those points are similar, and therefore, multiple partial images corresponding to points included in the same spatial region SA are images of morphologically similar tissues. Figure 5 shows an example of partial images PI. In Figure 5, if a partial image belonging to spatial region SA1 is a tissue image of interstitial tissue, it is highly likely that the other partial images belonging to spatial region SA1 are also tissue images of interstitial tissue. Similarly, if a partial image belonging to spatial region SA2 is a tissue image of a glandular duct, it is highly likely that the other partial images belonging to spatial region SA2 are also tissue images of glandular ducts. Furthermore, if a partial image belonging to spatial region SA4 is a tissue image of a tumor, it is highly likely that the other partial images belonging to spatial region SA4 are also tissue images of interstitial tissue.

[0027] In the second method, the training data generation unit 22 divides the feature space into multiple spatial regions by simple grid division instead of clustering. Fig. 6 shows an example of dividing the feature space into multiple spatial regions by grid division. In this example, the entire image WI is divided into n x m grids. In this case, each grid corresponds to one spatial region.

[0028] Next, the training data generation unit 22 acquires partial images from the plurality of spatial regions obtained by the first or second method described above, and generates training data. FIG. 7 is an explanatory diagram of the training data generation method. Note that the following description assumes that the first method described above is used. First, the training data generation unit 22 sets a probability of selecting one spatial region to be used to generate training data from the plurality of spatial regions (hereinafter referred to as the "region selection probability"). For example, the region selection probability of spatial region SA1 indicates the probability that spatial region SA1 is selected from the four spatial regions SA1 to SA4.

[0029] 7, in the initial state, the training data generation unit 22 sets the region selection probability of each spatial region to the same value. For example, as shown in probability distribution graph 51, the training data generation unit 22 sets the region selection probability of each spatial region SA1 to SA4 to the same value. Then, the training data generation unit 22 selects one spatial region from the spatial regions SA1 to SA4 based on the set region selection probability (first selection). In the example of FIG. 7, as shown by arrow 52, ​​the training data generation unit 22 selects spatial region SA4.

[0030] Next, the training data generation unit 22 selects a predetermined number of partial images PI from the multiple partial images PI belonging to the selected spatial region (second selection). The training data generation unit 22 then generates partial images (hereinafter also referred to as "patch images") by further subdividing each of the selected multiple partial images PI using grid division or the like. The training data generation unit 22 then acquires training patch images (hereinafter referred to as "training patch images") from each partial image PI. Specifically, the training data generation unit 22 may acquire all patch images obtained by subdividing the partial images PI as training patch images, or may select a predetermined number of these patch images by random sampling and acquire them as training patch images LI. Note that the training patch images are an example of training partial images. In this way, multiple training patch images LI are generated based on the multiple partial images corresponding to the selected spatial region (spatial region SA4 in FIG. 7 ).

[0031] Next, the training data generation unit 22 assigns a label to each of the obtained training patch images LI. As described above, in each piece of training data included in the training dataset, a label is assigned to each input image, i.e., each entire image WI. Therefore, the training data generation unit 22 uses the label assigned to the entire image WI to which each partial image belongs as the label of each training patch image LI generated from that partial image. In this way, the training data generation unit 22 assigns labels to all training patch images LI and ends the generation of training data. Note that the training data generation unit 22 generates training data for multiple input images.

[0032] The learning unit 23 learns a prediction model using the learning data input from the learning data generation unit 22. The prediction model predicts the probability that a predetermined image feature is included in the input image (whole image WI). In this embodiment, the prediction model predicts the probability that a predetermined feature such as the above-mentioned tumor, stroma, or duct is included in the patient's pathological tissue image, and outputs a confidence score for each feature. For example, a deep learning model such as a CNN (Convolutional Neural Network) can be used as the prediction model. The learning unit 23 performs a first learning round using the input learning data and outputs the learned prediction model to the prediction unit 24.

[0033] The prediction unit 24 uses the trained prediction model to make predictions for all partial images that make up the input image, calculates predicted values, and outputs them to the selection probability update unit 25 and the integration unit 26. Note that the prediction unit 24 performs this process for all input images included in the training dataset.

[0034] The selection probability update unit 25 calculates a prediction result corresponding to each spatial region based on the predicted value output by the prediction unit 24 using the trained prediction model. Specifically, the selection probability update unit 25 calculates the average value of the predicted values ​​for multiple subregions included in the spatial region SA1 as the prediction result for the spatial region SA1. For example, if the average value is equal to or greater than a predetermined threshold, the selection probability update unit 25 sets the prediction result to "+1," and if the average value is less than the threshold, the selection probability update unit 25 sets the prediction result to "-1." The selection probability update unit 25 similarly performs this process on the other spatial regions SA2 to SA4, generating prediction results for each spatial region.

[0035] Next, the selection probability update unit 25 updates the region selection probability of each spatial region using the prediction result for each spatial region. Specifically, the selection probability update unit 25 updates the region selection probability so that a spatial region with a high degree of certainty in the prediction result by the prediction model is more likely to be selected in the next training data.

[0036] In a preferred example, the selection probability update unit 25 calculates the selection probability D of the i-th spatial region SAi using the following equation (1): t+1 (i) is calculated.

[0037] Here, "t" is the number of iterations of the learning loop. t " indicates the weight of the model that has completed the tth learning, and is expressed by the following equation (2).

[0038] Also, "ε t " is the error probability of the model that has completed the tth learning, and is expressed by the following equation (3).

[0039] Also, "h t (x i ) is the model prediction result (±1) for the i-th spatial region SAi, and "y i " is the label (±1) in the i-th spatial region SAi. By repeating the learning loop while updating the region selection probability using the above formula (1), the error probability of the model decreases.

[0040] The selection probability update unit 25 outputs the updated region selection probabilities for each spatial region to the training data generation unit 22. The training data generation unit 22 uses the input updated region selection probabilities to generate training data to be used in the next training.

[0041] In this way, the learning loop is repeated a predetermined number of times while updating the region selection probability for each spatial region. Then, when the predetermined number of learning loops is completed, the prediction unit 24 performs predictions for all partial images included in the input image using the trained model at the time of completion of learning (also referred to as the “final model”), and outputs the predicted values ​​to the integrating unit 26.

[0042] The integrating unit 26 integrates the predicted values ​​for each partial image output by the predicting unit 24 to generate a prediction result for the entire input image. For example, for each input image, the integrating unit 26 determines the average value of the predicted values ​​for all partial images that make up the input image as the predicted value for the entire input image. The integrating unit 26 then displays the predicted value for the entire input image together with the input image on the display device 17.

[0043] In the above configuration, the image dividing unit 21 is an example of a partial image generating means, the learning data generating unit 22 is an example of a learning data generating means, the learning unit 23 is an example of a learning means, and the prediction unit 24 is an example of a prediction means.

[0044] (Learning Process) Fig. 8 is a flowchart of the learning process by the learning device 100. This process is realized by the processor 13 shown in Fig. 2 executing a program prepared in advance and operating as each element shown in Fig. 3.

[0045] First, the image segmentation unit 21 receives a training dataset and segments each input image included in the training dataset into partial images (step S10). Next, the training data generation unit 22 maps the partial images for each input image into a feature space (step S11). Next, the training data generation unit 22 segments the feature space into multiple spatial regions (step S12). In this case, the training data generation unit 22 may segment the feature space by clustering as described above, or by grid segmentation.

[0046] Next, the training data generation unit 22 selects one spatial region based on the region selection probability set for each spatial region (step S13). Note that, in the initial state, the region selection probabilities for multiple spatial regions are set to the same value. Next, the training data generation unit 22 generates training patch images from a predetermined number of partial images belonging to the selected spatial region (step S14). Next, the training data generation unit 22 assigns a label to each training patch image to generate training data (step S15).

[0047] Next, the learning unit 23 learns a prediction model using the training data generated in step S15, generating a trained prediction model (step S16). Next, the prediction unit 24 determines whether the learning unit 23 has trained a predetermined number of times (step S14). That is, the prediction unit 24 determines whether the above-mentioned training loop has been repeated a predetermined number of times. If the learning unit 23 has not trained a predetermined number of times (step S17: No), the prediction unit 24 uses the trained prediction model generated in step S16 to make predictions for partial images included in all input images in the training dataset (step S18).

[0048] Next, the selection probability update unit 25 updates the region selection probability of each spatial region based on the prediction result of each spatial region by the prediction unit 24 (step S19). As a result, the spatial region with a higher reliability of the prediction result by the trained prediction model has a higher region selection probability and is more likely to be selected in the next and subsequent generation of training data. Then, the process returns to step S13.

[0049] In this way, the learning loop of steps S13 to S19 is repeated until the learning unit 23 has performed learning a predetermined number of times. Then, when the learning unit 23 has performed learning a predetermined number of times (step S14: Yes), the learning process ends.

[0050] (Presenting the Basis for Prediction by the Model) By displaying the region selection probability at the end of the learning on the display device 17 or the like, it is possible to present the basis for prediction by the prediction model. FIG. 9 shows an example of a display of the region selection probability at the end of learning. In this example, for each spatial region, the tissue corresponding to that spatial region and the region selection probability for that spatial region are displayed. This display makes it possible to show the frequency with which the tissue corresponding to each spatial region was selected in learning the prediction model. For example, if the prediction model is a model that predicts the effect of medication on each tissue, it is possible to know which tissues have the greatest effect from medication.

[0051] (Modification) Next, a modification of the learning device according to the first embodiment will be described. As shown in Fig. 4, the learning device 100 divides the entire pathological tissue region included in the entire image WI into multiple partial images PI and maps them in a feature space. This is effective in ensuring the diversity of the pathological tissue represented by the partial images PI, but since normal regions and regions clearly unrelated to cancer are also mapped, there is a possibility that unnecessary processing will increase.

[0052] Therefore, in this modified example, the learning device detects tumor cells, lymphocytes, etc. from the pathological tissue region included in the entire image WI using a tumor cell detection method, selects partial images PI with a high proportion of tumor cells, and maps them into the feature space. This makes it possible to narrow down the targets for mapping into the feature space to the segmented images PI of the cancer and its surrounding area (regions thought to be highly related to drug efficacy), thereby reducing the amount of calculation.

[0053] FIG. 10 shows the functional configuration of a learning device 100x according to a modified example. As can be seen by comparing with FIG. 3, the learning device 100x includes an image filtering unit 28 between the image segmentation unit 21 and the training data generation unit 22. The image filtering unit 28 performs tumor cell detection processing on the entire image WI to detect tumor cells and other images. Existing tumor cell detection methods can be used to detect tumor cells. An example of a tumor cell detection method is described in the following literature: Cosatto, E., Gerard, K., Graf, HP, Ogura, M., Kiyuna, T., Hatanaka, KC, ... & Hatanaka, Y. (2021, February). A multi-scale conditional deep model for tumor cell ratio counting. In Medical Imaging 2021: Digital Pathology (Vol. 11603, pp. 31-42). SPIE.

[0054] FIG. 11(A) shows an example of tumor cell detection using an existing tumor cell detection method. In FIG. 11(A), the density of tumor cells is represented by shades of gray, with higher brightness indicating a higher proportion of tumor cells. Based on the tumor cell detection results shown in FIG. 11(A), the image narrowing unit 28 selects a segmented image PI from a region with a high local peak in the tumor cell density distribution and outputs the segmented image PI to the training data generation unit 22. Specifically, as shown in FIG. 11(B), the image narrowing unit 28 selects a partial image PI from a region with a high proportion of tumor cells, i.e., a region with high brightness, in FIG. 11(A), and outputs the partial image PI to the training data generation unit 22 as a target for mapping to the feature space. This reduces the computational load while improving the accuracy of the model.

[0055] [Prediction Device] Next, a prediction device that performs prediction using the trained model generated by the above-described learning device will be described.

[0056] (Hardware Configuration) The hardware configuration of the prediction device is basically the same as that of the learning device 100 shown in Fig. 2. However, unlike the learning device 100, the IF 12 acquires the input image that is actually the target of prediction from a database or the like, and the display device 17 displays the prediction result for the input image that is the target of prediction.

[0057] 12 is a block diagram showing the functional configuration of the prediction device 200. Functionally, the prediction device 200 includes an image dividing unit 31, a prediction unit 32, and an integration unit 33. The output of the integration unit 33 is supplied to the display device 17.

[0058] During actual prediction, an input image to be predicted is prepared. In this embodiment, a pathological tissue image of a patient receiving medication is input to the image segmentation unit 31 as the input image to be predicted. The image segmentation unit 31 divides the input image, i.e., the entire image WI, into multiple partial images PI using the same method as during learning, and further divides each partial image PI into patch images (hereinafter also referred to as "prediction patch images") of the same size as the learning patch images LI. That is, the image segmentation unit 31 divides the input image to be predicted into sizes corresponding to the input to the trained prediction model. The image segmentation unit 31 then outputs the prediction patch images obtained by the division to the prediction unit 32.

[0059] The prediction unit 32 performs prediction for the input image using the trained prediction model obtained by the above-described learning process. Specifically, the prediction unit 32 performs prediction for each prediction patch image obtained by the image division unit 31 using the trained prediction model, and outputs the predicted value to the integration unit 33.

[0060] The integrating unit 33 integrates the predicted values ​​calculated by the predicting unit 32 for each prediction patch image to calculate a predicted value for the entire input image, and outputs the calculated predicted value to the display device 17. The integrating unit 33 may output the calculated predicted value to an external device. This provides the probability that the input pathological tissue image contains tissue with a predetermined characteristic (such as a tumor, stroma, or duct in the previous example). The display device 17 then displays the predicted value and an important region on the input image. Furthermore, the integrating unit 33 may extract, in the prediction for the input image, a region of a partial image whose predicted value is equal to or greater than a predetermined reference value as an important region for prediction, and display the extracted region on the display device 17. In this manner, the prediction result for the input image is output.

[0061] The prediction model used by the prediction unit 32 is basically the prediction model at the end of a predetermined number of learning loops, i.e., the final model. However, since the final model does not necessarily have the highest accuracy, the prediction unit 32 may use the prediction model obtained in each learning loop that has the smallest prediction error. Specifically, the prediction results of the prediction model obtained at the end of each learning loop may be compared with the learning data set to calculate the prediction error, and the prediction model with the smallest prediction error may be adopted.

[0062] Furthermore, in the above example, one of the multiple prediction models obtained through multiple learning loops is used for prediction in the prediction unit 32, but instead, several of the multiple prediction models obtained may be used in combination. Specifically, the prediction unit 32 may make predictions for an input image using multiple prediction models obtained through multiple learning loops, and output a final prediction result by weighting and adding the multiple prediction results obtained. In this case, a greater weight may be assigned to the output of a prediction model with higher accuracy.

[0063] (Inference Processing) Fig. 13 is a flowchart of the prediction processing by the prediction device 200. This processing is realized by the processor 13 shown in Fig. 2 executing a program prepared in advance and operating as each element shown in Fig. 12.

[0064] First, the image dividing unit 31 divides the input image into partial images, and further divides each partial image into prediction patch images (step S21). Next, the prediction unit 32 uses the trained prediction model obtained by the training process to make predictions for each prediction patch image and output predicted values ​​(step S22). Next, the integration unit 33 integrates the predicted values ​​for each prediction patch image to calculate a prediction result for the entire input image (step S23). Then, the display device 17 displays the prediction result (step S24). Then, the prediction process ends.

[0065] [Application Example] A detailed description will be given of an example in which the prediction device of this embodiment is applied to predicting the effect of medication (drug efficacy or success, hereinafter referred to as "drug efficacy"). In the medical field, a pathological tissue image is used as input, and drug efficacy is predicted based on the image. For example, there is a method for predicting drug efficacy based on the staining rate of an immune image, which is based on an image in which cells including pathological tissue are stained. In this case, there are problems such as variations in the determination of the staining rate depending on the examiner, and it is not clear which part of the image is the characteristic part that reflects the drug efficacy.

[0066] Therefore, using the prediction device of this embodiment, a pathological tissue image is input, and a prediction model obtained by learning is used to predict drug efficacy, output a drug efficacy prediction score, and display areas that are important for determining drug efficacy.

[0067] Specifically, during training, training data is prepared in which labels indicating the presence or absence of drug efficacy are attached to pathological tissue images, and the above-mentioned training process is performed to train a drug efficacy prediction model. In this case, the method of this embodiment only requires labeling the entire pathological tissue image, so labeling is possible even if it is not clear which part of the pathological tissue image affects drug efficacy. Then, during prediction, the pathological tissue image to be predicted is input, and a drug efficacy score can be predicted using the trained prediction model. Furthermore, regions of the partial image that show high scores in the prediction process can be extracted as important regions that have a significant impact on drug efficacy and displayed on a display device.

[0068] When this embodiment is applied to predicting drug efficacy, the image segmentation unit 21, during learning and prediction, may utilize background knowledge, such as that information around the cell nucleus is particularly important for diagnosis, and generate a partial image centered on the position of the cell nucleus in the input image.

[0069] 14 is a block diagram showing the functional configuration of a learning device according to Embodiment 2. The learning device 70 includes a partial image generation unit 71, a feature space generation unit 72, a learning data generation unit 73, a learning unit 74, and a prediction unit 75.

[0070] FIG. 15 is a flowchart of processing by the learning device 70 of the second embodiment. The partial image generation means 71 generates partial images smaller than the input image from the input image (step S71). The feature space generation means 72 generates a feature space in which feature quantities of multiple partial images are mapped for each input image (step S72). The training data generation means 73 acquires multiple training partial images from the multiple partial images based on the feature space to generate training data (step S73). The learning means 74 uses the training data to train a prediction model that predicts the probability that a specific feature is included in the training partial images (step S74). The prediction means 75 uses the trained prediction model to make predictions for all or some of the partial images included in the input image (step S75). Furthermore, the training data generation means 73 acquires multiple training partial images to be used as training data in the next training session based on the predicted values ​​for all or some of the partial images in the feature space (step S76). In this way, the training partial images are updated, and the learning means 74 repeatedly trains the prediction model.

[0071] According to the learning device 70 of the second embodiment, even in a situation where detailed labels are not given to the entire image, it is possible to accurately learn a model while selecting areas that are effective as learning data.

[0072] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.

[0073] (Supplementary Note 1) A learning device comprising: partial image generation means for generating, from an input image, a partial image smaller than the input image; feature space generation means for generating, for each of the input images, a feature space into which feature amounts of a plurality of partial images are mapped; learning data generation means for acquiring a plurality of learning partial images from the plurality of partial images based on the feature space and generating learning data; learning means for using the learning data to train a prediction model that predicts the probability that a predetermined feature will be included in the learning partial images; and prediction means for making predictions for all or some of the partial images included in the input image using the trained prediction model, wherein the learning data generation means acquires a plurality of learning partial images to be used as learning data in the next learning, based on predicted values ​​for all or some of the partial images in the feature space.

[0074] (Supplementary Note 2) The learning device according to Supplementary Note 1, wherein the learning data generation means comprises: spatial domain division means for dividing the feature space into a plurality of spatial domains; and acquisition means for determining selection probabilities of the plurality of spatial domains and acquiring the plurality of learning partial images from a plurality of partial images corresponding to the spatial domains selected in accordance with the selection probabilities.

[0075] (Supplementary Note 3) The learning device according to Supplementary Note 2, wherein the learning data generating means updates the selection probabilities of the plurality of spatial regions based on predicted values ​​for all or some of the partial images included in the input image.

[0076] (Supplementary Note 4) The learning device according to Supplementary Note 3, wherein the learning data generation means updates the selection probabilities of the plurality of spatial regions so that the higher the confidence of the predicted value for the partial image, the higher the selection probability of the spatial region corresponding to the partial image.

[0077] (Supplementary Note 5) The learning device according to Supplementary Note 2, wherein the spatial region division means divides the feature space into a plurality of spatial regions by mapping feature amounts of the plurality of partial images onto the feature space and clustering the distribution of the feature amounts.

[0078] (Supplementary Note 6) The learning device according to Supplementary Note 1, wherein the learning data generating means uses a label previously assigned to the input image as a label for each learning partial image acquired from the input image.

[0079] (Supplementary Note 7) The learning device according to Supplementary Note 1, further comprising an output unit for outputting the selection probabilities of the plurality of spatial regions at the time of completion of learning of the trained prediction model. (Supplementary Note 8) The learning device according to Supplementary Note 1, wherein the feature space generation unit selects a partial image with a high proportion of tumor cells from a plurality of partial images generated from the input image and maps the partial image to the feature space.

[0080] (Supplementary Note 9) A prediction device comprising: a partial image generation means for generating, from an input image, partial images smaller than the input image; a prediction means for predicting, for all generated partial images, the probability that the partial images contain a predetermined feature using a prediction model trained by the learning device described in any one of Supplementary Notes 1 to 7; and an output means for integrating prediction results for all the partial images and outputting a prediction score indicating the probability that the input image contains the predetermined feature.

[0081] (Supplementary Note 10) A learning method executed by a computer, comprising: generating, from an input image, partial images smaller than the input image; generating, for each input image, a feature space in which feature amounts of a plurality of partial images are mapped; acquiring a plurality of partial images for training from the plurality of partial images based on the feature space to generate training data; using the training data, training a prediction model that predicts the probability that a predetermined feature is included in the partial images for training; making predictions for all or some of the partial images included in the input image using the trained prediction model; and generating the training data based on predicted values ​​for all or some of the partial images in the feature space, and acquiring a plurality of partial images for training to be used as training data in the next learning.

[0082] (Supplementary Note 11) A recording medium having recorded thereon a program that causes a computer to execute the following process: generating, from an input image, partial images smaller than the input image; generating, for each input image, a feature space in which feature amounts of multiple partial images are mapped; acquiring multiple training partial images from the multiple partial images based on the feature space to generate training data; using the training data, training a prediction model that predicts the probability that a predetermined feature is included in the training partial images; making predictions for all or some of the partial images included in the input image using the trained prediction model; and generating the training data based on predicted values ​​for all or some of the partial images in the feature space, and acquiring multiple training partial images to be used as training data in the next training.

[0083] Although the present disclosure has been described above with reference to the embodiments and examples, the present disclosure is not limited to the above-described embodiments and examples. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure.

[0084] 13 Processor 17 Display device 21, 31 Image segmentation unit 22 Learning data generation unit 23 Learning unit 24, 32 Prediction unit 25 Selection probability update unit 26, 33 Integration unit 100 Learning device 200 Prediction device

Claims

1. Sub-image generation means for generating a sub-image smaller than the input image from the input image, Feature space generation means for generating a feature space in which the feature amounts of a plurality of sub-images are mapped for each of the input images, Learning data generation means for obtaining a plurality of learning sub-images from the plurality of sub-images based on the feature space and generating learning data, Learning means for learning a prediction model that predicts the probability that a predetermined feature is included in the learning sub-image using the learning data, Prediction means for performing prediction on all or some of the sub-images included in the input image using the learned prediction model, comprising, The learning data generation means is a learning device that obtains a plurality of learning sub-images to be used as learning data in the next learning based on prediction values for all or some of the sub-images in the feature space.

2. The learning data generation means, Space region division means for dividing the feature space into a plurality of space regions, Acquisition means for determining selection probabilities of the plurality of space regions and acquiring the plurality of learning sub-images from a plurality of sub-images corresponding to the space regions selected according to the selection probabilities, The learning device according to claim 1, comprising.

3. The learning data generation means is the learning device according to claim 2, which updates the selection probabilities of the plurality of space regions based on prediction values for all or some of the sub-images included in the input image.

4. The learning data generation means is the learning device according to claim 3, which updates the selection probabilities of the plurality of space regions such that the higher the confidence level of the prediction value for the sub-image, the higher the selection probability of the space region corresponding to the sub-image.

5. The space region division means is the learning device according to claim 2, which maps the feature amounts of the plurality of sub-images onto the feature space and divides the feature space into a plurality of space regions by clustering the distribution of the feature amounts.

6. The learning device according to claim 1, comprising output means for outputting the selection probabilities of the plurality of space regions at the end of learning of the learned prediction model.

7. The feature space generation means is the learning device according to claim 1, which selects a sub-image with a high tumor cell ratio from among the plurality of sub-images generated from the input image and maps it to the feature space.

8. Sub-image generation means for generating a sub-image smaller than the input image from the input image, For all of the generated partial images, prediction means for predicting the probability that a predetermined feature is included in the partial image using a prediction model learned by the learning device according to any one of claims 1 to 7; output means for integrating the prediction results for all of the partial images and outputting a prediction score indicating the probability that the predetermined feature is included in the input image; A prediction device comprising the above.

9. A learning method executed by a computer, generating a partial image smaller than the input image from the input image, generating a feature space in which feature amounts of a plurality of partial images are mapped for each of the input images, generating learning data by acquiring a plurality of learning partial images from the plurality of partial images based on the feature space, learning a prediction model for predicting the probability that a predetermined feature is included in the learning partial image using the learning data, performing prediction on all or some of the partial images included in the input image using the learned prediction model, The generation of the learning data is a learning method for acquiring a plurality of learning partial images to be used as learning data in the next learning based on prediction values for all or some of the partial images of the feature space.

10. generating a partial image smaller than the input image from the input image, generating a feature space in which feature amounts of a plurality of partial images are mapped for each of the input images, generating learning data by acquiring a plurality of learning partial images from the plurality of partial images based on the feature space, learning a prediction model for predicting the probability that a predetermined feature is included in the learning partial image using the learning data, causing a computer to execute a process of performing prediction on all or some of the partial images included in the input image using the learned prediction model, The generation of the learning data is a program for acquiring a plurality of learning partial images to be used as learning data in the next learning based on prediction values for all or some of the partial images of the feature space.