Data set generation method, trained model evaluation method, trained model generation method, program, and data set generation device
The method addresses the challenge of inconsistent cell region detection in pathological tissue images by generating a dataset that distinguishes between cell regions and unclear regions, improving detection accuracy and subsequent analysis quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SCREEN HOLDINGS CO LTD
- Filing Date
- 2025-11-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods for detecting cell regions in pathological tissue images face challenges due to variations in tissue condition and staining, leading to inaccurate cell region detection and low model accuracy when dealing with unclear or poorly focused regions.
A method for generating a dataset that distinguishes between regions to be detected as cell regions and regions that may or may not be detected, using a combination of target images and annotation images with designated first and second regions, and evaluating the trained model by comparing these regions.
Improves the accuracy of cell region detection by distinguishing between essential and optional cell regions, reducing over-detection and under-detection, and enhancing the quality of subsequent analysis.
Smart Images

Figure JP2025039351_15052026_PF_FP_ABST
Abstract
Description
Dataset generation method, trained model evaluation method, trained model generation method, program, and dataset generation device.
[0001] The present invention relates to a technique for generating or evaluating a trained model for detecting cellular regions in a target image. [Reference to related applications] This application claims priority from Japanese Patent Application JP2024-196691 filed on 11 November 2024 and Japanese Patent Application JP2025-189161 filed on 10 November 2025, and all disclosures of those applications are incorporated herein.
[0002] Traditionally, single-cell analysis has been performed to obtain the expression levels of biomolecules, including proteins, at the cellular level from digital images of stained pathological tissue specimens. In single-cell analysis, individual cell regions are detected in the image, and staining intensity and other information for each cell are obtained based on the information of the cell region. This allows for the identification of, for example, the presence or absence of protein expression in each cell, and the analysis results are obtained.
[0003] Images of pathological tissues differ in tissue condition and staining even for the same area and staining type, depending on the site from which the tissue sample was taken and the type of staining used. Therefore, it is difficult to appropriately detect cell regions using so-called rule-based image processing, where processing conditions are predetermined by a human. On the other hand, it has been proposed to output cell regions using deep learning. For example, Japanese Patent Publication No. 2021-506022 discloses a method for training a multilayer neural network to detect and classify different cell types and regions within a sample image from a set of training images.
[0004] In addition, in Japanese Patent Publication No. 2023-544967, a technique is disclosed in which a target element region is detected from a processing target image based on artificial intelligence, and further, a target element envelope region is searched, and the target element region contour is obtained by fusing the target element region and the target element envelope region. A stained image is presented as an example of the processing target image, and the target elements are different types of cells. In Japanese Patent Publication No. 2024-528609, a machine learning model is used in pathological image analysis, and a training dataset is generated based on a first pathological dataset and a second pathological dataset, and the generated training dataset is used to train the machine learning model.
[0005] By the way, in cell recognition in a pathological tissue stained image, a policy of detecting some cells instead of all cells may be adopted. Even in this case, in order to train and evaluate the model, it is necessary to annotate the information of the cell region with reference to the staining property etc. from the pathological tissue stained image, that is, to teach the cell region. However, due to insufficient staining or poor focus of the image, there are cell regions where it is unclear whether they should be defined as cells to be detected. Such regions are considered to be acceptable from the user's perspective whether they are detected (i.e., identified) as cells or not. Therefore, if detection omission or over-detection of such regions is included in the evaluation of the trained model, the detection accuracy by the trained model will be evaluated lower than the accuracy originally desired. Therefore, a better training method and evaluation method are required when there are regions in the image that may or may not be detected as cell regions.
[0006] An object of the present invention is to distinguish and evaluate a region to be detected as a cell region and a region that may or may not be detected as a cell region when evaluating a trained model.
[0007] One aspect of the present invention is a method for generating a dataset used to generate or evaluate a trained model for detecting cell regions in a target image, comprising: a) preparing a plurality of target images; and b) generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each of the plurality of data elements includes a combination of a target image and an annotation image, and in step b), an annotation image corresponding to each of the plurality of target images is acquired by specifying a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.
[0008] According to the present invention, when evaluating a trained model, it is possible to distinguish between a first designated region that should be detected as a cell region and a second designated region that may or may not be detected as a cell region.
[0009] Aspect 2 of the present invention is a method for generating a dataset according to aspect 1, wherein in step b), each of the target images is displayed on a display unit, and the first designated area and the second designated area are designated as areas of different colors.
[0010] A third aspect of the present invention is a method for evaluating a trained model that detects cell regions in a target image, comprising: c) generating a dataset as an evaluation dataset using the dataset generation method described in claim 1; d) obtaining a plurality of detection images in which cell regions have been detected by inputting a plurality of target images included in the evaluation dataset into the trained model; and e) evaluating the trained model by comparing a plurality of annotation images included in the evaluation dataset with the plurality of detection images, wherein in step e), the evaluation is lowered when the second designated region in each of the plurality of annotation images is appropriately detected in the corresponding detection image, and when the second designated region is inappropriately detected compared to when it is not detected.
[0011] Aspect 4 of the present invention is a method for evaluating a trained model according to aspect 3, wherein the trained model is a model generated by machine learning using a training dataset, the training dataset is a collection of multiple data elements, each of the multiple data elements includes a combination of a target image and an annotation image, and the annotation image is an image in which regions to be detected as cell regions and regions that may or may not be detected as cell regions are designated as the first designated regions in relation to the target image.
[0012] Aspect 5 of the present invention is a method for generating a trained model that generates a trained model for detecting cell regions in a target image, comprising: f) generating a dataset as a training dataset using the dataset generation method described in claim 1; g) changing the second designated region of the annotation image of all data elements of the training dataset to a first designated region, or deleting the second designated region from the annotation image; and h) generating a trained model by performing machine learning using the training dataset.
[0013] Aspect 6 of the present invention is a method for generating a trained model that detects cell regions in a target image, comprising: i) a step of generating a dataset as a training dataset using the dataset generation method described in claim 1, while classifying the second designated region into a plurality of classes according to the degree to which it should be detected as a cell region when specifying the second designated region; j) a step of changing the second designated region belonging to some of the plurality of classes into a first designated region and deleting the second designated region belonging to other classes from the annotation image of the annotation image of all data elements of the training dataset; and k) a step of generating a trained model by performing machine learning using the training dataset.
[0014] Aspect 7 of the present invention is a computer-readable program that causes a computer to generate a dataset used for generating or evaluating a trained model for detecting cell regions in a target image, wherein the execution of the program by the computer causes the computer to perform the steps of a) preparing a plurality of target images and b) generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each of the plurality of data elements includes a combination of a target image and an annotation image, and in step b), an annotation image corresponding to each of the plurality of target images is acquired by accepting the designation of a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.
[0015] Aspect 8 of the present invention is a dataset generation device for generating or evaluating a trained model for detecting cell regions in a target image, comprising: a storage unit for storing a plurality of target images; and a dataset generation unit for generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each data element of the plurality of data elements includes a combination of a target image and an annotation image, and the dataset generation unit acquires annotation images corresponding to each of the plurality of target images by accepting the designation of a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.
[0016] The aforementioned objectives, as well as other objectives, features, embodiments, and advantages, will be revealed by the detailed description of the present invention below, with reference to the attached drawings.
[0017] This is a diagram showing the computer configuration. This is a block diagram showing the functional configuration of the dataset generation device. This is a diagram showing the flow of the dataset generation method. This is a diagram showing the target image and the annotation image. This is a diagram showing the dataset. This is a block diagram showing the functional configuration of the trained model generation device. This is a diagram showing the flow of the trained model generation method. This is a diagram showing the modified annotation image. This is a diagram showing the modified annotation image. This is a block diagram showing the functional configuration of the trained model evaluation device. This is a diagram showing the flow of the trained model evaluation method. This is a diagram showing the target image and the annotation image. This is a diagram showing the detected image. This is a diagram showing the annotation image and the detected image superimposed. This is a diagram showing the modified detected image. This is a diagram showing the target image and the annotation image. This is a diagram showing the modified annotation image. This is a diagram showing the modified annotation image.
[0018] Figure 1 shows the configuration of a computer 4 that functions as a dataset generation device according to one embodiment of the present invention. The computer 4 also functions as a trained model generation device and a trained model evaluation device. That is, the computer 4 executes the dataset generation method, trained model generation method, and trained model evaluation method described below.
[0019] The dataset generated by the dataset generator is used to generate or evaluate a trained model for detecting cell regions in the target image. "Detection of cell regions" refers to segmentation, which separates regions indicating cells from other regions in an image. Cell region detection using a trained model is used, for example, in the field of image cytometry. In this embodiment, the target image is assumed to be an image of a specimen in multiplex immunohistochemistry analysis, but the use of computer 4 is not limited to such specimen analysis.
[0020] Computer 4 has the configuration of a general computer system, including a CPU 41, a GPU 42, a ROM 43, a RAM 44, a fixed disk 45, a display 46, an input unit 47, a reader 48, a communication unit 49, and a bus 40. The CPU 41 performs various arithmetic operations. The GPU 42 performs various image processing operations at high speed. The ROM 43 stores basic programs. The RAM 44 stores various information. The fixed disk 45 stores information. The display 46 displays various information such as images.
[0021] The input unit 47 includes a keyboard 47a and a mouse 47b that receive input from the operator. The reader 48 reads information from a computer-readable recording medium 9 such as an optical disk, magnetic disk, magneto-optical disk, or memory card. The display 46, keyboard 47a, mouse 47b, and reader 48 are connected to the bus 40 via an interface (I / F). The communication unit 49 sends and receives signals to and from external devices of the computer 4. The bus 40 is a signal circuit that connects the CPU 41, GPU 42, ROM 43, RAM 44, fixed disk 45, display 46, input unit 47, reader 48, and communication unit 49.
[0022] In computer 4, program 91 is read in advance from recording medium 9 via reader 48 and stored in fixed disk 45. Program 91 may also be stored in fixed disk 45 via a network. The CPU 41 and GPU 42 perform arithmetic processing using RAM 44 and fixed disk 45 according to the computer-readable program 91. The CPU 41 and GPU 42 function as arithmetic units. Other configurations besides the CPU 41 and GPU 42 that function as arithmetic units may also be employed.
[0023] Figure 2 is a block diagram showing the functional configuration of the dataset generation device 1, which is realized when the computer 4 described above performs calculations and other processing according to the program 91. All or part of each functional configuration may be realized by dedicated electrical circuits. Furthermore, these functional configurations may be realized by multiple computers. The program 91 is a computer-readable program that causes the computer 4 to generate a dataset used for generating or evaluating a trained model that detects cell regions in a target image.
[0024] Of the functional configurations shown in Figure 2, the display unit 11 mainly performs the functions of the display 46 in Figure 1 and displays information to the operator. The dataset generation unit 12 is realized by the CPU 41, GPU 42, ROM 43, RAM 44, fixed disk 45 and their peripheral configurations. The operation unit 13 receives input from the operator and is mainly realized by the input unit 47 in Figure 1. The storage unit 14 is mainly realized by the RAM 44 and fixed disk 45. Various devices may be used as the storage unit 14 as long as they can store information.
[0025] The storage unit 14 is pre-prepared with data for multiple target images. In Figure 2, the data for one target image is shown as "target image data 51". The processing of target images in the following description is, more precisely, processing of target image data 51. The target image data 51 corresponding to multiple target images is input to the dataset generation unit 12, and a dataset is generated and stored in the storage unit 14. A training dataset 61 and an evaluation dataset 62 are generated as datasets, but these datasets do not need to be clearly distinguished. That is, the training dataset 61 and the evaluation dataset 62 may or may not have common parts.
[0026] Figure 3 shows the flow of the dataset generation method executed by the computer 4. When multiple target images are prepared in the storage unit 14 (step S11), the first target image is selected (step S12) and displayed on the display unit 11. By operating the operation unit 13, the operator allows the dataset generation unit 12 to accept the designation of a first designated region and a second designated region for the target image (step S13). The first designated region is the region in the target image that should be detected as a cell region by the trained model described later, and the second designated region is a region that may or may not be detected as a cell region. Whether to designate a region as the first or second designated region depends on the operator's judgment.
[0027] By specifying the first and second designated regions, a training image (corresponding to what is known as ground truth) representing these regions is generated. Hereafter, the training image will be referred to as the "annotation image." Cells corresponding to regions that should be detected as cellular regions are called "essential cells," and cells corresponding to regions that may or may not be detected as cellular regions are called "optional cells."
[0028] Figure 4 shows the display unit 11 with the target image 511 and the annotation image 521 displayed. The target image 511 on the left is an image of a stained specimen, with clearly stained cell regions 512 indicated by thick solid lines and faintly stained cell regions 513 indicated by thick dashed lines.
[0029] The annotation image 521 on the right shows the target image 511 displayed on the display unit 11, with the operator specifying the first designated region 522 corresponding to region 512 and the second designated region 523 corresponding to region 513. That is, the operator recognizes that region 512 corresponds to essential cells and performs the task of specifying the first designated region 522, and recognizes that region 513 corresponds to arbitrary cells and performs the task of specifying the second designated region 523. The first designated region 522 and the second designated region 523 are designated as regions of different colors. This makes it easy to specify the first designated region 522 and the second designated region 523. Once the first designated region 522 and the second designated region 523 are specified, the dataset generation unit 12 acquires the annotation image 521 (more precisely, the data of the annotation image 521) corresponding to each target image 511.
[0030] In Figure 4, the fact that the first designated area 522 and the second designated area 523 are displayed in different colors is represented by adding a solid parallel diagonal line to the first designated area 522 and a dashed parallel diagonal line to the second designated area 523. The first designated area 522 and the second designated area 523 can be designated by various methods. For example, this can be done by generating an image in which edge extraction and closed area extraction have been performed on the target image 511, and then using the mouse 47b to designate the closed area as either the first designated area 522 or the second designated area 523.
[0031] Figure 5 shows a dataset 60 generated by the dataset generation unit 12. The dataset generation unit 12 associates target image data 51 with annotation image data 52, which is the data of the annotation image 521 generated from the target image data 51, and generates data elements 50 that include these (step S14). That is, each data element 50 includes a combination of the target image 511 (more precisely, the target image data 51) and the annotation image 521 (more precisely, the annotation image data 52). Each data element 50 may be the combination of the target image 511 and the annotation image 521 itself, or it may include information related to them.
[0032] The dataset generation unit 12 generates a set of multiple data elements 50 by performing the above processing on multiple target image data 51 (step S15), and stores this as a dataset 60 in the storage unit 14. The dataset 60 is prepared in the storage unit 14 as a training dataset 61 or as an evaluation dataset 62. As described above, the dataset generation unit 12 generates the dataset 60 by acquiring multiple annotation images 521 that include information on the cell regions of multiple target images 511, namely the first designated region 522 and the second designated region 523.
[0033] Figure 6 is a block diagram showing the functional configuration of the trained model generation device 2, which is realized when the computer 4 performs calculations and other processing according to the program 91. All or part of each functional configuration may be realized by dedicated electrical circuits. Alternatively, these functional configurations may be realized by multiple computers. The program 91 is a computer-readable program that causes the computer 4 to generate a trained model for detecting cell regions in a target image.
[0034] In the functional configuration shown in Figure 6, the annotation image modification unit 21 and the learning unit 22 are realized by the CPU 41, GPU 42, ROM 43, RAM 44, fixed disk 45, and their peripheral configurations. Figure 7 is a diagram showing the flow of the trained model generation method executed by the computer 4 using the functional configuration shown in Figure 6.
[0035] When the training dataset 61 is generated using the dataset generation method shown in Figure 3 and prepared in the storage unit 14 (step S21), the training dataset 61 is input to the annotation image modification unit 21, and the second designated area 523 of the annotation image 521 in each data element 50 of the training dataset 61 is changed to the first designated area 522 (step S22). For example, if the first designated area 522 in the annotation image 521 is a blue area and the second designated area 523 is a green area, the green area is repainted blue. The reason for making such a change will be explained later. In the case of the annotation image 521 in Figure 4, as shown in Figure 8, both the first designated area 522 and the second designated area 523 in Figure 4 become the first designated area 522.
[0036] As another example of step S22, as shown in Figure 9, the second designated area 523 of the annotation image 521 in Figure 4 may be deleted, that is, the second designated area 523 may be filled with the background color. This modification will be described later.
[0037] The modified training dataset 61 is input to the training unit 22, and a trained model 7 is generated from the machine learning model through training (step S23). That is, training is performed using the target image 511 (target image data 51) as input information and the annotation image 521 (annotation image data 52) as training information. The structure of the model before training may be fixed or variable. The trained model 7 is generated by determining the values of each variable of the model through training. The trained model 7 is output to the storage unit 14 and stored (step S24).
[0038] Figure 10 is a block diagram showing the functional configuration of the trained model evaluation device 3, which is realized when a computer 4 performs calculations and other processing according to program 91. All or part of each functional configuration may be realized by dedicated electrical circuits. Alternatively, these functional configurations may be realized by multiple computers. Program 91 is a computer-readable program that causes the computer 4 to evaluate a trained model that detects cell regions in a target image.
[0039] In the functional configuration shown in Figure 10, the detection unit 31, the detected image modification unit 32, and the evaluation unit 33 are realized by a CPU 41, GPU 42, ROM 43, RAM 44, a fixed disk 45, and their peripheral configurations. The detection unit 31 has the function of outputting detected image data when target image data 51 is input, and the detection unit 31 functions as a detection device that detects cell regions from the target image 511. That is, when the trained model evaluation device 3 evaluates that an appropriate trained model 7 has been generated, the detection unit 31 with this trained model 7 set is used to detect cell regions from a new target image 511 (not included in the dataset 60).
[0040] Figure 11 shows the flow of the trained model evaluation method executed by computer 4 with the functional configuration shown in Figure 10. "Evaluation of trained model 7" means measuring the accuracy of how closely the trained model 7 can detect (i.e., identify) a cell region that is close to the corresponding annotation image 521 when the target image 511 is input, and obtaining an evaluation value.
[0041] When the evaluation dataset 62 is generated using the dataset generation method shown in Figure 3 and prepared in the storage unit 14 (step S31), the target image data 51 of multiple data elements 50 included in the evaluation dataset 62 is sequentially input to the trained model 7 of the detection unit 31, and image data of images in which cell regions are detected is acquired (step S32). In other words, an inference process using the trained model 7 is executed. Hereinafter, images in which cell regions are detected as an inference result by the detection unit 31 will be referred to as "detected images," and the data of the detected images will be referred to as "detected image data." However, as in the above explanation, processing on detected images is more accurately processing on detected image data, but in the following explanation, these terms will not be clearly distinguished.
[0042] Figure 12 illustrates the target image 511 and the corresponding annotation image 521. In Figure 12, reference numerals are assigned according to Figure 4. Figure 13 illustrates the detected image 531 obtained when the target image 511 from Figure 12 is input to the detection unit 31. During learning, in step S22 of Figure 7, the second designated region 523 is changed to the first designated region 522, so all cell regions shown in the detected image 531 become regions that have the attributes of the first designated region 522. Hereinafter, the detected cell regions will be referred to as "detected regions" and will be assigned reference numeral 532 as needed. The detected image data and annotation image data 52 are input to the detected image modification unit 32.
[0043] Figure 14 shows the annotation image 521 and the detection image 531 superimposed. In Figure 14, the upper second designated region 523a and its corresponding detection region 532a almost coincide. On the other hand, the detection region 532b is smaller than the second designated region 523b on the left. Also, the detection region 532c is larger than the second designated region 523c on the lower right. Note that the "detection region 532 corresponding to the second designated region 523" can be defined in various ways. For example, a detection region 532 that overlaps with at least a part of the second designated region 523 may be identified as the "detection region 532 corresponding to the second designated region 523". Alternatively, if the distance between the centroid of the second designated region 523 and the centroid of the detection region 532 is less than or equal to a predetermined distance, the detection region 532 may be identified as the "detection region 532 corresponding to the second designated region 523". The detection region 532 corresponding to the second designated region 523 may be identified under various conditions.
[0044] Here, the detection image changing unit 32 deletes a detection area 532 that substantially matches the second designated area 523 from the detection image 531 (e.g., fills it with a background color), and changes the detection image 531 so as to leave a detection area 532 that corresponds to the second designated area 523 but is significantly different from the second designated area 523 (step S33). As a result, in the next step, a detection area 532 that corresponds to the second designated area 523 but is significantly different from the second designated area 523 will be evaluated as an over-detection area. The "detection area 532 that substantially matches the second designated area 523" may be defined in various ways, but typically means that the two areas overlap at a ratio of a predetermined threshold or more.
[0045] For example, as the threshold, when the ratio of the detection area 532 overlapping with the second designated area 523 in the second designated area 523 is a predetermined threshold or more, and the ratio of the second designated area 523 overlapping with the detection area 532 in the detection area 532 is a predetermined threshold or more, the detection area 532 is determined to substantially match the second designated area 523. Alternatively, when the area belonging to only one of the second designated area 523 and the detection area 532 is below a certain ratio (threshold) of the common area of the two, the detection area 532 is determined to substantially match the second designated area 523. The detection area 532 that is (significantly) different from the second designated area 523 means that the requirement for determining that the detection area 532 substantially matches the second designated area 523 is not satisfied, and there are cases where the sizes of the two areas are different, the positions are different, and the shapes are different.
[0046] In the case of the example in FIG. 14, as shown in FIG. 15, the detection area 532a is deleted, and the detection areas 532b and 532c are left. The detection area 532 corresponding to the first designated area 522 of the annotation image 521, and the detection area 532 that does not correspond to either the first designated area 522 or the second designated area 523 (the area denoted by reference numeral 532d in FIG. 15) are maintained.
[0047] Then, the detection accuracy of the cell region by the learned model 7 is evaluated using the changed detection image 531. That is, the target images 511 of a large number of data elements 50 included in the evaluation dataset 62 are sequentially input into the learned model 7 to obtain a large number of detection images 531, and the evaluation unit 33 compares the plurality of annotation images 521 and the plurality of changed detection images 531 to evaluate the learned model 7.
[0048] The evaluation result is obtained as an evaluation value. Various things may be adopted as the evaluation value. For example, IoU (Intersection over Union) is used as the evaluation value. That is, when the annotation image 521 and the detection image 531 of each data element 50 are overlapped, the number of pixels in the region overlapping both the first designated region 522 and the detection region 532 is divided by the number of pixels in the region overlapping at least either the first designated region 522 or the detection region 532 to obtain an evaluation value element, and the average of the evaluation value elements of all the data elements 50 is obtained as the final evaluation value.
[0049] As another example of the evaluation value, when the two images are overlapped, the ratio of the number of pixels in the region overlapping both the first designated region 522 and the detection region 532 to the number of pixels in the first designated region 522 is obtained as an evaluation value element. The ratio of the number of pixels in the region overlapping both the first designated region 522 and the detection region 532 to the number of pixels in the detection region 532 may be adopted as an evaluation value element. Then, the average of the evaluation value elements of all the data elements 50 is obtained as the final evaluation value.
[0050] Furthermore, when the two images are overlapped, the value obtained by dividing the number of pixels in the region overlapping only either the first designated region 522 or the detection region 532 by the number of pixels in the region overlapping at least either the first designated region 522 or the detection region 532 is obtained as an evaluation value element, and the average of the evaluation value elements of all the data elements 50 may be obtained as the final evaluation value. In this case, the smaller the evaluation value, the higher the evaluation.
[0051] By deleting the detection region 532 that corresponds to the second designated region 523, that is, the detection region 532 that was appropriately detected in relation to an arbitrary cell, it is possible to prevent the evaluation from becoming high regardless of the detection accuracy of the first designated region 522 when the second designated region 523 is appropriately detected, and to lower the evaluation of the trained model 7 when there is a detection region 532 that was not appropriately detected in relation to the second designated region 523.
[0052] In other words, the evaluation of the trained model 7 is lower when the second designated region 523 in each annotation image 521 is improperly detected than when it is properly detected in the corresponding detection image 531, and when it is (substantially) not detected. Whether the second designated region 523 is properly detected or not detected at all does not affect the evaluation. As a result, it is possible to obtain an appropriate evaluation value while appropriately handling the detection region 532 corresponding to the second designated region 523.
[0053] As described above, by specifying the first designated region 522 and the second designated region 523 in the annotation image 521, it becomes possible to distinguish and evaluate the first designated region 522 and the second designated region 523 when evaluating the trained model 7. As a result, the quality of analysis, which is a subsequent step that requires appropriate detection of the cell region, can be improved.
[0054] The higher the evaluation of the trained model 7, the smaller the evaluation value may be, and the lower the evaluation, the larger the evaluation value may be. For example, the degree of over-detection may be obtained as one of the evaluation values. The degree of over-detection is determined, for example, as the ratio of the area of the detection area 532 that does not correspond to either the first designated area 522 or the second designated area 523, and the area of the detection area 532 that is inappropriately detected corresponding to the second designated area 523, to the total area of the detection area 532. In this case, the smaller the evaluation value, the higher the evaluation. In the following explanation, unless otherwise specified, a higher evaluation is assumed to correspond to a larger evaluation value.
[0055] The occurrence of many over-detections may be due to the inappropriate detection of the second designated region 523. Therefore, if there are many over-detections, the inappropriate detection of arbitrary cell regions can be suppressed by modifying the training dataset 61 to remove the second designated region 523 and then performing training. Specifically, the second designated region 523 of the annotation image 521 in Figure 4 is removed as shown in Figure 9. After that, the trained model 7 is generated and evaluated. The evaluation method is the same as described above, and the evaluation value is obtained by comparing the first designated region 522 and the detected region 532 after removing the detected region 532 that was appropriately detected in relation to the second designated region 523.
[0056] Generally speaking, training using the annotation image 521 in Figure 8 or Figure 9 can be described as follows: After generating a training dataset 61 using the dataset generation method in Figure 3 (step S21 in Figure 7), the second designated region 523 of the annotation image 521 for all data elements 50 of the training dataset 61 is changed to the first designated region 522, or the second designated region 523 is deleted from the annotation image 521 (step S22), and a trained model 7 is generated by performing machine learning using the modified training dataset 61 (steps S23, S24). This makes it possible to obtain a trained model 7 that can suppress over-detection or under-detection.
[0057] In the examples in Figures 8 and 9, the second designated area 523 is either changed to the first designated area 522 or deleted. However, a portion of the second designated area 523 may be changed to the first designated area 522 and the rest deleted. Figure 16 shows an example of the designation of the second designated area 523 in such a case.
[0058] In the example shown in Figure 16, the second designated region 523 is designated while being divided into two classes. Here, the two classes are referred to as "Class A" and "Class B". The second designated region 523 is classified into two classes at the time of designation according to the degree to which it should be detected as a cell region. The "degree to which it should be detected" can be subjective. For example, if multiple second designated regions 523 are classified into either Class A or Class B from two states of a region 513 of an arbitrary cell, one worker may judge that Class A should be detected as a cell region more highly than Class B, while another worker may judge that Class B should be detected as a cell region more highly than Class A. Hereafter, we will assume that the second designated region 523 belonging to Class A should be detected as a cell region more highly than the second designated region 523 belonging to Class B. Furthermore, the second designated region 523 belonging to Class A will be referred to as "Second Designated Region 523A", and the second designated region 523 belonging to Class B will be referred to as "Second Designated Region 523B".
[0059] Figure 16 illustrates an annotation image 521 in which, in step S13 of Figure 3, the operator has identified a first designated region 522 corresponding to essential cells (regions labeled with reference numeral 512) and second designated regions 523A and 523B corresponding to arbitrary cells (regions labeled with reference numeral 513) in the target image 511. The difference in color between the second designated region 523A and the second designated region 523B is represented by varying the spacing of the parallel diagonal lines. For example, the second designated region 523A is dark green, and the second designated region 523B is light green. By creating such annotation images 521 for a large number of target images 511, a training dataset 61 and an evaluation dataset 62 are generated (steps S11 to S15).
[0060] During training, as illustrated in Figure 17, all second designated regions 523A and 523B are changed to the first designated region 522 (step S22 in Figure 7), and then training is performed to generate a trained model 7 (steps S23 and S24). Subsequently, as already explained, multiple detected images 531 are obtained by inputting the target images 511 of the evaluation dataset 62 into the trained model 7, and after the detection regions 532 that were appropriately detected corresponding to the second designated regions 523A and 523B are deleted, an evaluation value is obtained by comparing the first designated region 522 of the annotation image 521 with the detected region 532.
[0061] If a suitable pre-trained model 7 with a high evaluation is obtained through the above process, the generation of the pre-trained model 7 is terminated. If the pre-trained model 7 has many over-detections of cell regions, the training dataset 61 is modified and training is performed again. At this time, in step S13 of Figure 3, the second designated region 523A of class A is changed to the first designated region 522, and the second designated region 523B of class B is deleted. In the example of Figure 16, the annotation image 521 shown in Figure 18 is obtained. Considering the fluctuations in the operator's judgment of the degree to which a region should be detected as a cell region, the second designated region 523B may be changed to the first designated region 522 and the second designated region 523A may be deleted. Then, training and evaluation of the pre-trained model 7 are performed in the same manner as above. In the evaluation, the detection region 532 that was appropriately detected corresponding to the second designated regions 523A and 523B is deleted, and then the first designated region 522 of the annotation image 521 and the detection region 532 are compared to obtain an evaluation value.
[0062] If a suitable pre-trained model 7 is obtained, the pre-trained model 7 generation process is completed. If the pre-trained model 7 has many over-detections of cell regions, the training dataset 61 is further modified and training is performed again. At this time, in step S13 of Figure 3, the second designated regions 523A and 523B of classes A and B are deleted. In the example of Figure 16, the annotation image 521 shown in Figure 19 is obtained. Subsequently, training and evaluation of the pre-trained model 7 are performed in the same manner as above. That is, training is performed using the annotation image 521 which contains only the first designated region 522, and in the evaluation of the pre-trained model 7, the detection regions 532 which were appropriately detected corresponding to the second designated regions 523A and 523B are deleted, and then the first designated region 522 of the annotation image 521 and the detection regions 532 are compared.
[0063] In cell region detection, it is generally desirable to detect as many cell regions as possible. Therefore, as described above, it is preferable to generate the initial trained model 7 by changing all second designated regions 523 to first designated regions 522. In this case, the annotation image 521 is an image in which the regions to be detected as cell regions, and regions that may or may not be detected as cell regions, are specified as first designated regions 522 relative to the target image 511.
[0064] The number of classes used to classify the second designated region 523 is not limited to two, but may be three or more. In other words, when designating the second designated region 523, the second designated region 523 is classified into multiple classes according to the degree to which it should be detected as a cell region (step S13 in Figure 3), and a training dataset 61 is generated by the dataset generation method in Figure 3. Then, for the annotation images 521 of all data elements 50 of the training dataset 61, modifications are made to the training dataset 61, such that some of the second designated regions 523 belonging to some of the multiple classes are changed to the first designated region 522, and the second designated regions 523 belonging to other classes are deleted from the annotation images 521 (step S22 in Figure 7). By performing machine learning using the modified training dataset 61, a trained model 7 that can suppress over-detection and under-detection of cell regions can be generated (step S23).
[0065] In actual operation, as already explained, a suitable trained model 7 is obtained by repeatedly generating and evaluating the trained model 7 while changing the selection of classes in the second designated region 523, which is changed to the first designated region 522 in step S21 during training. If a trained model 7 with a favorable evaluation value is not obtained, the number of target images 511 is increased, or the number of designations in the second designated region 523 is increased or decreased, and the trained model 7 is generated again.
[0066] As shown in Figures 8 and 9, changing or deleting the second designated region 523 of the annotation image 521 to the first designated region 522 during training can be considered equivalent to the case where the number of classes is 1 in the above explanation.
[0067] In the above description, the training dataset 61 and the evaluation dataset 62 have similar structures, and the annotation image 521 includes a first designated region 522 and a second designated region 523. However, the trained model 7 to be evaluated does not necessarily need to be generated using an annotation image 521 that includes both the first designated region 522 and the second designated region 523.
[0068] In other words, when evaluating a trained model 7 generated by another method, an evaluation dataset 62 including an annotation image 521 containing a first designated region 522 and a second designated region 523 may be used. In this case as well, the evaluation method is the same as described above. However, from the viewpoint of simplifying the work, it is preferable to create a dataset 60 containing a large number of data elements 50 as described above, and use a part of this dataset 60 as a training dataset 61 and another part as an evaluation dataset 62.
[0069] On the other hand, if the annotation image 521, which includes the first designated region 522 and the second designated region 523, is used directly for training to generate a trained model 7, the detection image 531 obtained by inputting the target image 511 into the trained model 7 will show a distinction between the detection region 532 corresponding to the first designated region 522 and the detection region 532 corresponding to the second designated region 523. For example, if the first designated region 522 is blue and the second designated region 523 is green in the annotation image 521, the detection image 531 will show both a blue detection region 532 and a green detection region 532. Furthermore, in order to train the first designated region 522 and the second designated region 523 to be distinguished, the number of data elements 50 required in the training dataset 61 increases, and the amount of computation required for training also increases. Note that the colors of the first designated region 522 and the second designated region 523 when specifying them in step S13 of Figure 3 do not need to be different, and can be determined as appropriate.
[0070] Since such distinctions are unnecessary in the evaluation of the trained model 7, post-processing is required to change the green detection region 532 to blue, or to perform further additional post-processing. The amount of training can also be reduced. Therefore, in the above embodiment, in step S22 of Figure 7, the second designated region 523 of the annotation image 521 is changed to the first designated region 522, and the annotation image 521 is an image in which the region that should be detected as a cell region and the region that may or may not be detected as a cell region are designated as the first designated region 522 with respect to the target image 511.
[0071] In the above embodiment, the trained model 7 is evaluated by comparing the first designated region 522 of the annotation image 521 with the detection region 532 of the modified detection image 531. This suppresses the influence of the detection of arbitrary cell regions 513 on the evaluation while giving greater weight to the detection of essential cell regions 512 that should be detected. However, cell regions can exist in complex states, and if the first designated region 522 of the annotation image 521 and the detection region 532 of the modified detection image 531 are the main comparisons, then it is not necessary to completely exclude the consideration of the second designated region 523 when comparing the annotation image 521 and the detection image 531.
[0072] The target image 511 can be any image that shows cells. The target image 511 is not limited to images showing pathological tissue. For example, it may be an image showing a collection of cells that do not bind to each other. The purpose of training by specifying regions that may or may not be detected as cell regions is to suppress the detection of incomplete cell regions. With the trained model 7 obtained through such training, a detection region 532 can be obtained in which the extent of individual cell regions is accurately detected.
[0073] The detection of the cell regions described above is suitable when subsequent analyses are performed based on segmented cell regions. For example, it is suitable for quantifying the amount of a specific substance (e.g., a specific protein) present within a cell. It is also suitable when a strategy is adopted to detect and analyze only a portion of the cells in the target image. The trained model 7 obtained as described above is particularly suitable when the target image 511 is an image of stained cells and the staining intensity of each cell is to be determined.
[0074] The machine learning performed in step S23 of Figure 7 is not limited to deep learning using neural networks. Other supervised machine learning methods such as regression or decision trees may also be employed.
[0075] The configurations in the above embodiments and each modified example may be combined as appropriate, as long as they do not contradict each other.
[0076] Although the invention has been described in detail, the above description is illustrative and not limiting. Therefore, it can be said that numerous modifications and embodiments are possible as long as they do not deviate from the scope of the present invention.
[0077] 1 Dataset generation device 5 Computer 7 Trained model 9 Program 11 Display unit 12 Dataset generation unit 14 Storage unit 50 Data elements 60 Dataset 61 Training dataset 62 Evaluation dataset 511 Target image 521 Annotation image 522 First designated area 523, 523A, 523B Second designated area 531 Detected image S11-S15, S21-S24, S31-S34 Generation process
Claims
1. A method for generating a dataset used to generate or evaluate a trained model for detecting cell regions in target images, comprising: a) a step of preparing a plurality of target images; and b) a step of generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each data element of the plurality of data elements includes a combination of a target image and an annotation image, and in step b), an annotation image corresponding to each of the plurality of target images is acquired by specifying a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.
2. A dataset generation method according to claim 1, wherein in step b), each of the target images is displayed on a display unit, and the first designated area and the second designated area are designated as areas of different colors.
3. A method for evaluating a trained model that detects cell regions in a target image, comprising: c) a step of generating a dataset as an evaluation dataset using the dataset generation method described in claim 1; d) a step of obtaining a plurality of detection images in which cell regions have been detected by inputting a plurality of target images included in the evaluation dataset into the trained model; and e) a step of evaluating the trained model by comparing a plurality of annotation images included in the evaluation dataset with the plurality of detection images, wherein in step e), the evaluation is lowered when the second designated region in each of the plurality of annotation images is appropriately detected in the corresponding detection image, and when the second designated region is inappropriately detected compared to when it is not detected.
4. A method for evaluating a trained model according to claim 3, wherein the trained model is a model generated by machine learning using a training dataset, the training dataset is a collection of a plurality of data elements, each of the plurality of data elements includes a combination of a target image and an annotation image, and the annotation image is an image in which the regions to be detected as cell regions and regions that may or may not be detected as cell regions are designated as the first designated regions in relation to the target image.
5. A method for generating a trained model for detecting cell regions in a target image, comprising: f) generating a dataset as a training dataset using the dataset generation method described in claim 1; g) changing the second designated region of the annotation image of all data elements of the training dataset to the first designated region, or deleting the second designated region from the annotation image; and h) generating a trained model by performing machine learning using the training dataset.
6. A method for generating a trained model for detecting cell regions in a target image, comprising: i) a step of generating a dataset as a training dataset by the dataset generation method described in claim 1, classifying the second designated region into a plurality of classes according to the degree to which it should be detected as a cell region when specifying the second designated region; j) a step of changing the second designated region belonging to some of the plurality of classes into a first designated region and deleting the second designated region belonging to other classes from the annotation image of the annotation image of all data elements of the training dataset; and k) a step of generating a trained model by performing machine learning using the training dataset.
7. A computer-readable program that causes a computer to generate a dataset used for generating or evaluating a trained model for detecting cell regions in a target image, wherein the execution of the program by the computer causes the computer to perform the steps of: a) preparing a plurality of target images; and b) generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each of the plurality of data elements includes a combination of a target image and an annotation image, and in step b), the program acquires an annotation image corresponding to each of the plurality of target images by accepting the designation of a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.
8. A dataset generation device for generating a dataset used to generate or evaluate a trained model for detecting cell regions in a target image, comprising: a storage unit for storing a plurality of target images; and a dataset generation unit for generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each data element of the plurality of data elements includes a combination of a target image and an annotation image, and the dataset generation unit acquires annotation images corresponding to each of the plurality of target images by accepting the designation of a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.