Dataset generation method, trained model evaluation method, trained model generation method, program, and dataset generation device.

By generating a dataset with designated regions for cell detection in pathological tissue images, the method addresses the challenge of distinguishing between necessary and unnecessary cell regions, improving model evaluation and detection accuracy.

JP2026084688APending Publication Date: 2026-05-21SCREEN HOLDINGS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SCREEN HOLDINGS CO LTD
Filing Date
2025-11-10
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing methods for cell region detection in pathological tissue images struggle with accurately distinguishing between regions that should be detected as cells and those that may or may not be detected due to insufficient staining or poor image focus, leading to inaccurate model evaluation.

Method used

A method for generating and evaluating a trained model by creating a dataset with designated regions, including a first region to be detected as a cell region and a second region that may or may not be detected, using machine learning and image annotation to improve detection accuracy.

Benefits of technology

Enables accurate differentiation between cell regions that should be detected and those that may not, enhancing the evaluation and training of the model to improve cell region detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026084688000001_ABST
    Figure 2026084688000001_ABST
Patent Text Reader

Abstract

When evaluating a trained model, regions that should be detected as cellular regions and regions that may or may not be detected as cellular regions are distinguished and evaluated accordingly. [Solution] A dataset generation method for generating a dataset used to generate or evaluate a trained model for detecting cell regions in a target image comprises the steps of preparing multiple target images (step S11) and generating a dataset by acquiring multiple annotation images containing cell region information from the multiple target images (steps S12 to S15). The dataset is a collection of multiple data elements, each data element containing a combination of a target image and an annotation image. In the dataset generation step, for each target image, a first designated region is specified, which is a region that should be detected as a cell region, and a second designated region is specified, which is a region that may or may not be detected as a cell region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for generating or evaluating a learned model for detecting cell regions in a target image.

Background Art

[0002] Conventionally, single-cell analysis has been performed to obtain the expression levels of biological substances such as proteins at the cell unit from digital images of stained pathological tissue specimens. In single-cell analysis, individual cell regions in an image are detected, and staining intensities and the like in each cell are obtained based on the information of the cell regions. Thereby, for example, the presence or absence of protein expression in each cell is specified, and an analysis result is obtained.

[0003] Images of pathological tissues differ in tissue state and staining conditions even for the same site and the same staining due to differences in the site where the tissue specimen was taken and the type of staining. Therefore, it is difficult to appropriately detect cell regions by so-called rule-based image processing in which a person determines processing conditions in advance. On the other hand, it has also been proposed to output cell regions using deep learning. For example, Patent Document 1 discloses a method of training a multi-layer neural network to detect and classify different cell types and regions in a sample image from a set of training images.

[0004] Further, Patent Document 2 discloses a technique for detecting a target element region from a processing target image based on artificial intelligence, further searching for a target element envelope region, and obtaining a target element region contour by fusing the target element region and the target element envelope region. A stained image is presented as an example of the processing target image, and the target elements are different types of cells. In Patent Document 3, a machine learning model is used in pathological image analysis, a learning data set is generated based on a first pathological data set and a second pathological data set, and the machine learning model is trained using the generated learning data set.

Prior Art Documents

Patent Documents

[0005] [Patent Document 1] Special Publication No. 2021-506022 [Patent Document 2] Special Publication No. 2023-544967 [Patent Document 3] Special Publication No. 2024-528609 [Overview of the project] [Problems that the invention aims to solve]

[0006] Incidentally, in cell recognition of pathological tissue stained images, a strategy is sometimes adopted to detect only some cells rather than all of them. Even in this case, in order to train and evaluate the model, it is necessary to annotate information about cell regions based on staining properties, etc., from the pathological tissue stained image, that is, to teach the model the cell regions. However, there are cell regions that are unclear whether they should be defined as cells to be detected due to insufficient staining or poor image focus. From a user's perspective, it is considered acceptable whether such regions are detected (i.e., identified) as cells or not. Therefore, if the evaluation of the trained model includes the detection failure or overdetection of such regions, the detection accuracy of the trained model will be evaluated as lower than the accuracy that is originally desired. Thus, a better training method and evaluation method is needed when there are regions in the image that do not necessarily need to be detected as cell regions.

[0007] This invention has been made in view of the above problems, and aims to distinguish and evaluate regions that should be detected as cellular regions and regions that may or may not be detected as cellular regions when evaluating a trained model. [Means for solving the problem]

[0008] One aspect of the present invention is a method for generating a dataset used to generate or evaluate a trained model for detecting cell regions in a target image, comprising: a) preparing a plurality of target images; and b) generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each of the plurality of data elements includes a combination of a target image and an annotation image, and in step b), an annotation image corresponding to each of the plurality of target images is acquired by specifying a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.

[0009] Aspect 2 of the present invention is a method for generating a dataset according to aspect 1, wherein in step b), each of the target images is displayed on a display unit, and the first designated area and the second designated area are designated as areas of different colors.

[0010] A third aspect of the present invention is a method for evaluating a trained model that detects cell regions in a target image, comprising: c) generating a dataset as an evaluation dataset using the dataset generation method described in claim 1; d) obtaining a plurality of detection images in which cell regions have been detected by inputting a plurality of target images included in the evaluation dataset into the trained model; and e) evaluating the trained model by comparing a plurality of annotation images included in the evaluation dataset with the plurality of detection images, wherein in step e), the evaluation is lowered when the second designated region in each of the plurality of annotation images is appropriately detected in the corresponding detection image, and when the second designated region is inappropriately detected compared to when it is not detected.

[0011] Aspect 4 of the present invention is a method for evaluating a trained model according to aspect 3, wherein the trained model is a model generated by machine learning using a training dataset, the training dataset is a collection of multiple data elements, each of the multiple data elements includes a combination of a target image and an annotation image, and the annotation image is an image in which regions to be detected as cell regions and regions that may or may not be detected as cell regions are designated as the first designated regions in relation to the target image.

[0012] Aspect 5 of the present invention is a method for generating a trained model that generates a trained model for detecting cell regions in a target image, comprising: f) generating a dataset as a training dataset using the dataset generation method described in claim 1; g) changing the second designated region of the annotation image of all data elements of the training dataset to a first designated region, or deleting the second designated region from the annotation image; and h) generating a trained model by performing machine learning using the training dataset.

[0013] Aspect 6 of the present invention is a method for generating a trained model that detects cell regions in a target image, comprising: i) a step of generating a dataset as a training dataset using the dataset generation method described in claim 1, while classifying the second designated region into a plurality of classes according to the degree to which it should be detected as a cell region when specifying the second designated region; j) a step of changing the second designated region belonging to some of the plurality of classes into a first designated region and deleting the second designated region belonging to other classes from the annotation image of the annotation image of all data elements of the training dataset; and k) a step of generating a trained model by performing machine learning using the training dataset.

[0014] Aspect 7 of the present invention is a computer-readable program that causes a computer to generate a dataset used for generating or evaluating a trained model for detecting cell regions in a target image, wherein the execution of the program by the computer causes the computer to perform the steps of a) preparing a plurality of target images and b) generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each of the plurality of data elements includes a combination of a target image and an annotation image, and in step b), an annotation image corresponding to each of the plurality of target images is acquired by accepting the designation of a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region.

[0015] Aspect 8 of the present invention is a dataset generation device for generating or evaluating a trained model for detecting cell regions in a target image, comprising: a storage unit for storing a plurality of target images; and a dataset generation unit for generating a dataset by acquiring a plurality of annotation images containing information on cell regions of the plurality of target images, wherein the dataset is a collection of a plurality of data elements, each of the plurality of data elements includes a combination of a target image and an annotation image, and the dataset generation unit acquires annotation images corresponding to each of the plurality of target images by accepting the designation of a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region. [Effects of the Invention]

[0016] According to the present invention, when evaluating a trained model, it is possible to distinguish between a first designated region that should be detected as a cell region and a second designated region that may or may not be detected as a cell region. [Brief explanation of the drawing]

[0017] [Figure 1] It is a diagram showing the configuration of a computer. [Figure 2] It is a block diagram showing the functional configuration of a dataset generation device. [Figure 3] It is a diagram showing the flow of a dataset generation method. [Figure 4] It is a diagram showing a target image and an annotation image. [Figure 5] It is a diagram showing a dataset. [Figure 6] It is a block diagram showing the functional configuration of a trained model generation device. [Figure 7] It is a diagram showing the flow of a trained model generation method. [Figure 8] It is a diagram showing the changed annotation image. [Figure 9] It is a diagram showing the changed annotation image. [Figure 10] It is a block diagram showing the functional configuration of a trained model evaluation device. [Figure 11] It is a diagram showing the flow of a trained model evaluation method. [Figure 12] It is a diagram showing a target image and an annotation image. [Figure 13] It is a diagram showing a detected image. [Figure 14] It is a diagram showing an annotation image and a detected image superimposed. [Figure 15] It is a diagram showing the changed detected image. [Figure 16] It is a diagram showing a target image and an annotation image. [Figure 17] It is a diagram showing the changed annotation image. <00​​​​​​​​​​

[0018] Figure 1 shows the configuration of a computer 4 that functions as a dataset generation device according to one embodiment of the present invention. The computer 4 also functions as a trained model generation device and a trained model evaluation device. That is, the computer 4 executes the dataset generation method, trained model generation method, and trained model evaluation method described below.

[0019] The dataset generated by the dataset generator is used to generate or evaluate a trained model for detecting cell regions in the target image. "Detection of cell regions" refers to segmentation, which separates regions indicating cells from other regions in an image. Cell region detection using a trained model is used, for example, in the field of image cytometry. In this embodiment, the target image is assumed to be an image of a specimen in multiplex immunohistochemistry analysis, but the use of computer 4 is not limited to such specimen analysis.

[0020] Computer 4 has the configuration of a general computer system, including a CPU 41, a GPU 42, a ROM 43, a RAM 44, a fixed disk 45, a display 46, an input unit 47, a reader 48, a communication unit 49, and a bus 40. The CPU 41 performs various arithmetic operations. The GPU 42 performs various image processing operations at high speed. The ROM 43 stores basic programs. The RAM 44 stores various information. The fixed disk 45 stores information. The display 46 displays various information such as images.

[0021] The input unit 47 includes a keyboard 47a and a mouse 47b that receive input from the operator. The reader 48 reads information from a computer-readable recording medium 9 such as an optical disk, magnetic disk, magneto-optical disk, or memory card. The display 46, keyboard 47a, mouse 47b, and reader 48 are connected to the bus 40 via an interface (I / F). The communication unit 49 sends and receives signals to and from external devices of the computer 4. The bus 40 is a signal circuit that connects the CPU 41, GPU 42, ROM 43, RAM 44, fixed disk 45, display 46, input unit 47, reader 48, and communication unit 49.

[0022] In computer 4, program 91 is read in advance from recording medium 9 via reader 48 and stored in fixed disk 45. Program 91 may also be stored in fixed disk 45 via a network. The CPU 41 and GPU 42 perform arithmetic processing using RAM 44 and fixed disk 45 according to the computer-readable program 91. The CPU 41 and GPU 42 function as arithmetic units. Other configurations besides the CPU 41 and GPU 42 that function as arithmetic units may also be employed.

[0023] Figure 2 is a block diagram showing the functional configuration of the dataset generation device 1, which is realized when the computer 4 described above performs calculations and other processing according to program 91. All or part of each functional configuration may be realized by dedicated electrical circuits. Furthermore, these functional configurations may be realized by multiple computers. Program 91 is a computer-readable program that causes the computer 4 to generate a dataset used for generating or evaluating a trained model that detects cell regions in a target image.

[0024] Of the functional configurations shown in Figure 2, the display unit 11 primarily performs the functions of the display 46 in Figure 1, displaying information to the operator. The dataset generation unit 12 is realized by the CPU 41, GPU 42, ROM 43, RAM 44, fixed disk 45, and their peripheral configurations. The operation unit 13 receives input from the operator and is primarily realized by the input unit 47 in Figure 1. The storage unit 14 is primarily realized by the RAM 44 and fixed disk 45. Various devices may be used as the storage unit 14 as long as they are capable of storing information.

[0025] The memory unit 14 is pre-prepared with data for multiple target images. In Figure 2, the data for one target image is shown as "target image data 51". The processing of target images in the following description is, more precisely, processing of target image data 51. The target image data 51 corresponding to multiple target images is input to the dataset generation unit 12, and a dataset is generated and stored in the memory unit 14. A training dataset 61 and an evaluation dataset 62 are generated as datasets, but these datasets do not need to be clearly distinguished. That is, the training dataset 61 and the evaluation dataset 62 may or may not have common parts.

[0026] Figure 3 shows the flow of the dataset generation method executed by the computer 4. When multiple target images are prepared in the storage unit 14 (step S11), the first target image is selected (step S12) and displayed on the display unit 11. By operating the operation unit 13, the operator allows the dataset generation unit 12 to accept the designation of a first designated region and a second designated region for the target image (step S13). The first designated region is the region in the target image that should be detected as a cell region by the trained model described later, and the second designated region is a region that may or may not be detected as a cell region. Whether to designate a region as the first or second designated region depends on the operator's judgment.

[0027] By specifying the first and second designated regions, a training image (corresponding to what is known as ground truth) representing these regions is generated. Hereafter, the training image will be referred to as the "annotation image." Cells corresponding to regions that should be detected as cellular regions are called "essential cells," and cells corresponding to regions that may or may not be detected as cellular regions are called "optional cells."

[0028] Figure 4 shows the display unit 11 with the target image 511 and the annotation image 521 displayed. The target image 511 on the left is an image of a stained specimen, with clearly stained cell regions 512 indicated by thick solid lines and faintly stained cell regions 513 indicated by thick dashed lines.

[0029] The annotation image 521 on the right shows the target image 511 displayed on the display unit 11, with the operator specifying the first designated region 522 corresponding to region 512 and the second designated region 523 corresponding to region 513. That is, the operator recognizes that region 512 corresponds to essential cells and performs the task of specifying the first designated region 522, and recognizes that region 513 corresponds to arbitrary cells and performs the task of specifying the second designated region 523. The first designated region 522 and the second designated region 523 are designated as regions of different colors. This makes it easy to specify the first designated region 522 and the second designated region 523. Once the first designated region 522 and the second designated region 523 are specified, the dataset generation unit 12 acquires the annotation image 521 (more precisely, the data of the annotation image 521) corresponding to each target image 511.

[0030] In Figure 4, the fact that the first designated area 522 and the second designated area 523 are displayed in different colors is represented by adding a solid parallel diagonal line to the first designated area 522 and a dashed parallel diagonal line to the second designated area 523. The designation of the first designated area 522 and the second designated area 523 can be performed by various methods. For example, this can be done by generating an image in which edge extraction and closed area extraction have been performed on the target image 511, and then using the mouse 47b to designate the closed area as either the first designated area 522 or the second designated area 523.

[0031] Figure 5 shows the dataset 60 generated by the dataset generation unit 12. The dataset generation unit 12 associates the target image data 51 with the annotation image data 52, which is the data of the annotation image 521 generated from the target image data 51, and generates data elements 50 that include these (step S14). That is, each data element 50 includes the combination of the target image 511 (more precisely, the target image data 51) and the annotation image 521 (more precisely, the annotation image data 52). Each data element 50 may be the combination of the target image 511 and the annotation image 521 itself, or it may include information related to them.

[0032] The dataset generation unit 12 generates a set of multiple data elements 50 by performing the above processing on multiple target image data 51 (step S15), and stores this as a dataset 60 in the storage unit 14. The dataset 60 is prepared in the storage unit 14 as a training dataset 61 or as an evaluation dataset 62. As described above, the dataset generation unit 12 generates the dataset 60 by acquiring multiple annotation images 521 that include information on the cell regions of multiple target images 511, namely the first designated region 522 and the second designated region 523.

[0033] Figure 6 is a block diagram showing the functional configuration of the trained model generation device 2, which is realized when computer 4 performs calculations and other processing according to program 91. All or part of each functional configuration may be realized by dedicated electrical circuits. Furthermore, these functional configurations may be realized by multiple computers. Program 91 is a computer-readable program that causes computer 4 to generate a trained model for detecting cell regions in a target image.

[0034] Of the functional configurations shown in Figure 6, the annotation image modification unit 21 and the learning unit 22 are realized by the CPU 41, GPU 42, ROM 43, RAM 44, fixed disk 45, and their peripheral configurations. Figure 7 is a diagram showing the flow of the trained model generation method executed by the computer 4 using the functional configuration shown in Figure 6.

[0035] When the training dataset 61 is generated using the dataset generation method shown in Figure 3 and prepared in the storage unit 14 (step S21), the training dataset 61 is input to the annotation image modification unit 21, and the second designated region 523 of the annotation image 521 in each data element 50 of the training dataset 61 is changed to the first designated region 522 (step S22). For example, if the first designated region 522 in the annotation image 521 is a blue region and the second designated region 523 is a green region, the green region is repainted blue. The reason for making such a change will be explained later. In the case of the annotation image 521 in Figure 4, as shown in Figure 8, both the first designated region 522 and the second designated region 523 in Figure 4 become the first designated region 522.

[0036] As another example of step S22, as shown in Figure 9, the second designated area 523 of the annotation image 521 in Figure 4 may be deleted, that is, the second designated area 523 may be filled with the background color. This modification will be discussed later.

[0037] The modified training dataset 61 is input to the training unit 22, and a trained model 7 is generated from the machine learning model through training (step S23). That is, training is performed using the target image 511 (target image data 51) as input information and the annotation image 521 (annotation image data 52) as training information. The structure of the model before training may be fixed or variable. The trained model 7 is generated by determining the values ​​of each variable of the model through training. The trained model 7 is output to the storage unit 14 and stored (step S24).

[0038] Figure 10 is a block diagram showing the functional configuration of the trained model evaluation device 3, which is realized when computer 4 performs calculations and other processing according to program 91. All or part of each functional configuration may be realized by dedicated electrical circuits. Alternatively, these functional configurations may be realized by multiple computers. Program 91 is a computer-readable program that instructs computer 4 to evaluate a trained model that detects cell regions in a target image.

[0039] Of the functional configurations shown in Figure 10, the detection unit 31, the detected image modification unit 32, and the evaluation unit 33 are realized by a CPU 41, a GPU 42, a ROM 43, a RAM 44, a fixed disk 45, and their peripheral configurations. The detection unit 31 has the function of outputting detected image data when target image data 51 is input, and the detection unit 31 functions as a detection device that detects cell regions from the target image 511. That is, if the trained model evaluation device 3 evaluates that an appropriate trained model 7 has been generated, the detection unit 31 with this trained model 7 set is used to detect cell regions from a new target image 511 (not included in the dataset 60).

[0040] Figure 11 shows the flow of the trained model evaluation method executed by computer 4 with the functional configuration shown in Figure 10. "Evaluation of trained model 7" means measuring the accuracy of how closely the trained model 7 can detect (i.e., identify) cell regions that are close to the corresponding annotation image 521 when the target image 511 is input, and obtaining an evaluation value.

[0041] When the evaluation dataset 62 is generated using the dataset generation method shown in Figure 3 and prepared in the storage unit 14 (step S31), the target image data 51 of multiple data elements 50 included in the evaluation dataset 62 is sequentially input to the trained model 7 of the detection unit 31, and image data of images in which cell regions are detected is acquired (step S32). In other words, an inference process using the trained model 7 is performed. Hereinafter, images in which cell regions are detected as an inference result by the detection unit 31 will be referred to as "detected images," and the data of the detected images will be referred to as "detected image data." However, as in the above explanation, processing on detected images is more accurately processing on detected image data, but in the following explanation, these terms will not be clearly distinguished.

[0042] Figure 12 illustrates the target image 511 and the corresponding annotation image 521. In Figure 12, reference numerals are assigned according to Figure 4. Figure 13 illustrates the detected image 531 obtained when the target image 511 from Figure 12 is input to the detection unit 31. During training, in step S22 of Figure 7, the second designated region 523 is changed to the first designated region 522, so all cell regions shown in the detected image 531 become regions that have the attributes of the first designated region 522. Hereinafter, the detected cell regions will be referred to as "detected regions" and will be assigned the reference numeral 532 as needed. The detected image data and annotation image data 52 are input to the detected image modification unit 32.

[0043] Figure 14 shows the annotation image 521 and the detection image 531 superimposed. In Figure 14, the upper second designated region 523a and its corresponding detection region 532a almost coincide. On the other hand, the detection region 532b is smaller than the second designated region 523b on the left. Also, the detection region 532c is larger than the second designated region 523c on the lower right. Note that the "detection region 532 corresponding to the second designated region 523" can be defined in various ways. For example, a detection region 532 that overlaps with at least a part of the second designated region 523 may be identified as the "detection region 532 corresponding to the second designated region 523". Alternatively, if the distance between the centroid of the second designated region 523 and the centroid of the detection region 532 is less than or equal to a predetermined distance, the detection region 532 may be identified as the "detection region 532 corresponding to the second designated region 523". The detection region 532 corresponding to the second designated region 523 may be identified under various conditions.

[0044] Here, the detection image modification unit 32 removes the detection region 532 that closely matches the second designated region 523 from the detection image 531 (for example, by filling it with the background color), and modifies the detection image 531 so that it leaves the detection region 532 that corresponds to the second designated region 523 but is significantly different from it (step S33). As a result, in the next step, the detection region 532 that corresponds to the second designated region 523 but is significantly different from it will be evaluated as an over-detection region. The "detection region 532 that closely matches the second designated region 523" can be defined in various ways, but typically it means that the two regions overlap to a degree greater than or equal to a predetermined threshold.

[0045] For example, as a threshold, if the proportion of the second designated area 523 that overlaps with the detection area 532 is greater than or equal to a predetermined threshold, and the proportion of the detection area 532 that overlaps with the second designated area 523 is greater than or equal to a predetermined threshold, then the detection area 532 is determined to be approximately the same as the second designated area 523. Alternatively, if the area belonging only to either the second designated area 523 or the detection area 532 is less than or equal to a certain proportion (threshold) of the common area between the two, then the detection area 532 is determined to be approximately the same as the second designated area 523. A detection area 532 that is (significantly) different from the second designated area 523 means that the requirements for determining that the detection area 532 is approximately the same as the second designated area 523 are not met, and this may be due to differences in the size, position, or shape of the two areas.

[0046] In the example shown in Figure 14, as shown in Figure 15, detection region 532a is deleted, while detection regions 532b and 532c remain. Detection region 532 corresponding to the first designated region 522 of the annotation image 521, as well as detection region 532 that does not correspond to either the first designated region 522 or the second designated region 523 (the region labeled with code 532d in Figure 15), are maintained.

[0047] Then, the accuracy of cell region detection by the trained model 7 is evaluated using the modified detection images 531. Specifically, the target images 511 of numerous data elements 50 included in the evaluation dataset 62 are sequentially input to the trained model 7 to obtain numerous detection images 531, and the evaluation unit 33 evaluates the trained model 7 by comparing multiple annotation images 521 with multiple modified detection images 531.

[0048] The evaluation result is obtained as an evaluation value. Various values ​​may be used as evaluation values. For example, IoU (Intersection over Union) can be used as an evaluation value. That is, when the annotation image 521 and the detection image 531 of each data element 50 are superimposed, the number of pixels in the region that overlaps with both the first designated region 522 and the detection region 532 is divided by the number of pixels in the region that overlaps with at least one of the first designated region 522 and the detection region 532 to obtain an evaluation value element, and the average of the evaluation value elements of all data elements 50 is obtained as the final evaluation value.

[0049] Other examples of evaluation values ​​include the ratio of the number of pixels in the region overlapping with both the first designated region 522 and the detection region 532 to the number of pixels in the first designated region 522 when the two images are superimposed, which is determined as an evaluation value element. Alternatively, the ratio of the number of pixels in the region overlapping with both the first designated region 522 and the detection region 532 to the number of pixels in the detection region 532 may be adopted as an evaluation value element. The average of the evaluation value elements of all data elements 50 is then determined as the final evaluation value.

[0050] Furthermore, when both images are superimposed, the number of pixels in the region that overlaps with only one of the first designated region 522 and the detection region 532 may be divided by the number of pixels in the region that overlaps with at least one of the first designated region 522 and the detection region 532 to obtain an evaluation value element, and the average of the evaluation value elements of all data elements 50 may be obtained as the final evaluation value. In this case, a smaller evaluation value indicates a higher evaluation.

[0051] By deleting the detection region 532 that corresponds to the second designated region 523, that is, the detection region 532 that was appropriately detected in relation to an arbitrary cell, it is possible to prevent the evaluation from becoming high regardless of the detection accuracy of the first designated region 522 when the second designated region 523 is appropriately detected, and to lower the evaluation of the trained model 7 when there is a detection region 532 that was not appropriately detected in relation to the second designated region 523.

[0052] In other words, the evaluation of the trained model 7 is lower when the second designated region 523 in each annotation image 521 is improperly detected than when it is properly detected in the corresponding detection image 531, and when it is (effectively) not detected. Whether the second designated region 523 is properly detected or not detected at all does not affect the evaluation. As a result, it is possible to obtain an appropriate evaluation value while appropriately handling the detection region 532 corresponding to the second designated region 523.

[0053] As described above, by specifying the first designated region 522 and the second designated region 523 in the annotation image 521, it becomes possible to distinguish and evaluate the first designated region 522 and the second designated region 523 when evaluating the trained model 7. As a result, the quality of analysis, which is a subsequent step that requires appropriate detection of the cell region, can be improved.

[0054] The higher the evaluation of the trained model 7, the smaller the evaluation value may be, and the lower the evaluation, the larger the evaluation value may be. For example, the degree of over-detection may be obtained as one of the evaluation values. The degree of over-detection is determined, for example, as the ratio of the area of ​​the detection region 532 that does not correspond to either the first designated region 522 or the second designated region 523, and the area of ​​the detection region 532 that is inappropriately detected corresponding to the second designated region 523, to the total area of ​​the detection region 532. In this case, the smaller the evaluation value, the higher the evaluation. In the following explanation, unless otherwise specified, a higher evaluation is assumed to correspond to a larger evaluation value.

[0055] The occurrence of many over-detections may be due to the inappropriate detection of the second designated region 523. Therefore, if there are many over-detections, the inappropriate detection of arbitrary cell regions can be suppressed by modifying the training dataset 61 to remove the second designated region 523 and then performing training. Specifically, the second designated region 523 of the annotation image 521 in Figure 4 is removed as shown in Figure 9. Subsequently, the trained model 7 is generated and evaluated. The evaluation method is the same as described above, and the evaluation value is obtained by comparing the first designated region 522 and the detected region 532 after removing the detected region 532 that was appropriately detected in relation to the second designated region 523.

[0056] Generally speaking, training using the annotation image 521 in Figure 8 or Figure 9 can be described as follows: After generating a training dataset 61 using the dataset generation method in Figure 3 (step S21 in Figure 7), the second designated region 523 of the annotation image 521 for all data elements 50 of the training dataset 61 is changed to the first designated region 522, or the second designated region 523 is deleted from the annotation image 521 (step S22), and a trained model 7 is generated by performing machine learning using the modified training dataset 61 (steps S23, S24). This makes it possible to obtain a trained model 7 that can suppress over-detection or under-detection.

[0057] In the examples in Figures 8 and 9, the second designated area 523 is either changed to the first designated area 522 or deleted. However, a portion of the second designated area 523 may be changed to the first designated area 522 and the rest deleted. Figure 16 shows an example of how the second designated area 523 is designated in such a case.

[0058] In the example in Figure 16, the second designation region 523 is designated while being divided into two classes. Here, the two classes are referred to as "Class A" and "Class B". The second designation region 523 is classified into two classes at the time of designation according to the degree to which it should be detected as a cellular region. The "degree to which it should be detected" can be subjective. For example, if multiple second designation regions 523 are classified into either Class A or Class B from two states of a region 513 of an arbitrary cell, one worker may judge that Class A should be detected as a cellular region more highly than Class B, while another worker may judge that Class B should be detected as a cellular region more highly than Class A. Hereafter, we will assume that the second designation region 523 belonging to Class A should be detected as a cellular region more highly than the second designation region 523 belonging to Class B. Furthermore, the second designation region 523 belonging to Class A will be referred to as "Second Designation Region 523A", and the second designation region 523 belonging to Class B will be referred to as "Second Designation Region 523B".

[0059] Figure 16 illustrates an annotation image 521 in which, in step S13 of Figure 3, the operator has identified the first designated region 522 corresponding to essential cells (regions labeled with code 512) and the second designated regions 523A and 523B corresponding to arbitrary cells (regions labeled with code 513) in the target image 511. The difference in color between the second designated region 523A and the second designated region 523B is represented by varying the spacing of the parallel diagonal lines. For example, the second designated region 523A is dark green, and the second designated region 523B is light green. By creating such annotation images 521 for a large number of target images 511, a training dataset 61 and an evaluation dataset 62 are generated (steps S11 to S15).

[0060] During training, as illustrated in Figure 17, all second designated regions 523A and 523B are changed to the first designated region 522 (step S22 in Figure 7), and then training is performed to generate a trained model 7 (steps S23 and S24). Subsequently, as already explained, multiple detected images 531 are obtained by inputting the target images 511 of the evaluation dataset 62 into the trained model 7. After the detection regions 532 that were appropriately detected corresponding to the second designated regions 523A and 523B are deleted, the first designated region 522 of the annotation image 521 and the detected region 532 are compared to obtain an evaluation value.

[0061] If a suitable pre-trained model 7 with a high evaluation is obtained through the above process, the generation of pre-trained model 7 is terminated. If the pre-trained model 7 has many over-detections of cell regions, the training dataset 61 is modified and training is performed again. At this time, in step S13 of Figure 3, the second designated region 523A of class A is changed to the first designated region 522, and the second designated region 523B of class B is deleted. In the example of Figure 16, the annotation image 521 shown in Figure 18 is obtained. Considering the fluctuations in the operator's judgment of the degree to which a region should be detected as a cell region, the second designated region 523B may be changed to the first designated region 522 and the second designated region 523A may be deleted. Then, training and evaluation of the pre-trained model 7 are performed in the same manner as above. In the evaluation, the detection region 532 that was appropriately detected corresponding to the second designated regions 523A and 523B is deleted, and the evaluation value is obtained by comparing the first designated region 522 and the detection region 532 of the annotation image 521.

[0062] If a suitable pre-trained model 7 is obtained, the generation of the pre-trained model 7 is completed. If the pre-trained model 7 has many over-detections of cell regions, the training dataset 61 is further modified and training is performed again. At this time, in step S13 of Figure 3, the second designated regions 523A and 523B of classes A and B are deleted. In the example of Figure 16, the annotation image 521 shown in Figure 19 is obtained. Subsequently, training and evaluation of the pre-trained model 7 are performed in the same manner as above. That is, training is performed using the annotation image 521 which contains only the first designated region 522, and in the evaluation of the pre-trained model 7, the detection regions 532 which were appropriately detected corresponding to the second designated regions 523A and 523B are deleted, and then the first designated region 522 of the annotation image 521 and the detection region 532 are compared.

[0063] In cell region detection, it is generally desirable to detect as many cell regions as possible. Therefore, as described above, it is preferable to generate the initial trained model 7 by changing all second designated regions 523 to first designated regions 522. In this case, the annotation image 521 is an image in which the regions to be detected as cell regions, and regions that may or may not be detected as cell regions, are specified as first designated regions 522 relative to the target image 511.

[0064] The number of classes used to classify the second designated region 523 is not limited to two, but may be three or more. In other words, when designating the second designated region 523, the second designated region 523 is classified into multiple classes according to the degree to which it should be detected as a cell region (step S13 in Figure 3), and a training dataset 61 is generated using the dataset generation method in Figure 3. Then, for the annotation images 521 of all data elements 50 of the training dataset 61, modifications are made to the training dataset 61, such that some of the second designated regions 523 belonging to some of the multiple classes are changed to the first designated region 522, and the second designated regions 523 belonging to other classes are deleted from the annotation images 521 (step S22 in Figure 7). By performing machine learning using the modified training dataset 61, a trained model 7 that can suppress over-detection and under-detection of cell regions can be generated (step S23).

[0065] In actual operation, as already explained, a suitable trained model 7 is obtained by repeatedly generating and evaluating the trained model 7 while changing the selection of classes in the second designated region 523, which is changed to the first designated region 522 in step S21 during training. If a trained model 7 with a favorable evaluation value is not obtained, the trained model 7 is generated again after increasing the number of target images 511 or increasing or decreasing the number of designations in the second designated region 523.

[0066] As shown in Figures 8 and 9, changing or deleting the second designated region 523 of the annotation image 521 to the first designated region 522 during training can be considered equivalent to the case where the number of classes is 1 in the above explanation.

[0067] In the above description, the training dataset 61 and the evaluation dataset 62 have similar structures, and the annotation image 521 includes a first designated region 522 and a second designated region 523. However, the trained model 7 to be evaluated does not necessarily need to be generated using an annotation image 521 that includes both the first designated region 522 and the second designated region 523.

[0068] In other words, when evaluating a trained model 7 generated by another method, an evaluation dataset 62 including an annotation image 521 containing a first designated region 522 and a second designated region 523 may be used. In this case as well, the evaluation method is the same as described above. However, from the viewpoint of simplifying the work, it is preferable to create a dataset 60 containing a large number of data elements 50 as described above, and use a part of this dataset 60 as a training dataset 61 and another part as an evaluation dataset 62.

[0069] On the other hand, if the annotation image 521, which includes the first designated region 522 and the second designated region 523, is used directly for training to generate a trained model 7, the detection image 531 obtained by inputting the target image 511 into the trained model 7 will show a distinction between the detection region 532 corresponding to the first designated region 522 and the detection region 532 corresponding to the second designated region 523. For example, if the first designated region 522 is blue and the second designated region 523 is green in the annotation image 521, the detection image 531 will show both a blue detection region 532 and a green detection region 532. Furthermore, in order to train the model to distinguish between the first designated region 522 and the second designated region 523, the number of data elements 50 required in the training dataset 61 increases, and the amount of computation required for training also increases. Note that the colors of the first designated region 522 and the second designated region 523 when specifying them in step S13 of Figure 3 do not need to be different, and can be determined as appropriate.

[0070] Since such distinctions are unnecessary in the evaluation of the trained model 7, post-processing is required to change the green detection region 532 to blue, or to perform further additional post-processing. The amount of training can also be reduced. Therefore, in the above embodiment, in step S22 of Figure 7, the second designated region 523 of the annotation image 521 is changed to the first designated region 522, and the annotation image 521 is an image in which the region that should be detected as a cell region and the region that may or may not be detected as a cell region are designated as the first designated region 522 in relation to the target image 511.

[0071] In the above embodiment, the trained model 7 is evaluated by comparing the first designated region 522 of the annotation image 521 with the detection region 532 of the modified detection image 531. This suppresses the influence of the detection of arbitrary cell regions 513 on the evaluation while giving greater weight to the detection of essential cell regions 512 that should be detected. However, cell regions can exist in complex states, and if the first designated region 522 of the annotation image 521 and the detection region 532 of the modified detection image 531 are the primary comparisons, then it is not necessary to completely exclude the consideration of the second designated region 523 when comparing the annotation image 521 and the detection image 531.

[0072] The target image 511 can be any image that shows cells. The target image 511 is not limited to images showing pathological tissue. For example, it may be an image showing a collection of cells that do not bind to each other. The purpose of training by specifying regions that may or may not be detected as cell regions is to suppress the detection of incomplete cell regions. With the trained model 7 obtained through such training, a detection region 532 can be obtained in which the extent of individual cell regions is accurately detected.

[0073] The detection of the cell regions described above is suitable when subsequent analyses are performed based on segmented cell regions. For example, it is suitable for quantifying the amount of a specific substance (e.g., a specific protein) present within a cell. It is also suitable when a strategy is adopted to detect and analyze only a portion of the cells in the target image. The trained model 7 obtained as described above is particularly suitable when the target image 511 is an image of stained cells and the staining intensity of each cell is to be determined.

[0074] The machine learning performed in step S23 of Figure 7 is not limited to deep learning using neural networks. Other supervised machine learning methods such as regression or decision trees may also be employed.

[0075] The configurations in the above embodiments and each modified example may be combined as appropriate, as long as they do not contradict each other. [Explanation of symbols]

[0076] 1. Dataset Generator 5 Computers 7. Pre-trained models 9 Programs 11 Display section 12. Dataset Generation Unit 14 Storage section 50 data elements 60 datasets 61 Training datasets 62 Evaluation datasets 511 Target Images 521 Annotated Images 522 1st designated area 523,523A,523B 2nd designated area 531 Detected Images Processes for generating products: S11-S15, S21-S24, S31-S34

Claims

1. A method for generating a dataset used to generate or evaluate a trained model for detecting cell regions in a target image, a) A step of preparing multiple target images, b) A step of generating a dataset by acquiring multiple annotation images containing information on the cell regions of the multiple target images, Equipped with, The dataset is a collection of multiple data elements, and each of the multiple data elements includes a combination of a target image and an annotation image. A method for generating a dataset in which, in step b) above, annotation images corresponding to each of the multiple target images are obtained by specifying a first designated region which is a region that should be detected as a cell region and a second designated region which is a region that may or may not be detected as a cell region for each of the multiple target images.

2. A method for generating a dataset according to claim 1, A method for generating a dataset in which, in step b) above, each of the target images is displayed on a display unit, and the first designated area and the second designated area are designated as areas of different colors.

3. A method for evaluating a trained model that detects cell regions in a target image, c) A step of generating the dataset as an evaluation dataset using the dataset generation method described in claim 1, d) A step of obtaining multiple detection images in which cell regions are detected by inputting multiple target images included in the evaluation dataset into a trained model, e) A step of evaluating the trained model by comparing a plurality of annotation images included in the evaluation dataset with the plurality of detection images, Equipped with, A trained model evaluation method in step e) above, wherein the evaluation is lower when the second designated region in each of the plurality of annotation images is appropriately detected in the corresponding detection image, and when the second designated region is inappropriately detected compared to when it is not detected.

4. A method for evaluating a trained model according to claim 3, The aforementioned trained model is a model generated by machine learning using a training dataset. The training dataset is a collection of multiple data elements, and each of the multiple data elements includes a combination of a target image and an annotation image. A method for evaluating a trained model, wherein the annotation image is an image in which the regions to be detected as cell regions and regions that may or may not be detected as cell regions are designated as the first designated regions in relation to the target image.

5. A method for generating a trained model that generates a trained model for detecting cell regions in a target image, f) A step of generating the dataset as a training dataset using the dataset generation method described in claim 1, g) A step of changing the second designated region of the annotation image of all data elements of the training dataset to the first designated region, or deleting the second designated region from the annotation image, h) A step of generating a trained model by performing machine learning using the aforementioned training dataset, A method for generating a pre-trained model that includes the following features.

6. A method for generating a trained model that generates a trained model for detecting cell regions in a target image, i) The method for generating a dataset according to claim 1, wherein when specifying the second designated region, the second designated region is classified into a plurality of classes according to the degree to which it should be detected as a cell region, and the dataset is generated as a training dataset by the dataset generation method, j) For annotation images of all data elements of the training dataset, the steps of changing the second designated region belonging to some of the multiple classes to the first designated region, and deleting the second designated region belonging to other classes from the annotation image, k) A step of generating a trained model by performing machine learning using the aforementioned training dataset, A method for generating a pre-trained model that includes the following features.

7. A computer-readable program that causes a computer to generate a dataset used for generating or evaluating a trained model for detecting cell regions in a target image, wherein the execution of the program by the computer is performed by the computer, a) A step of preparing multiple target images, b) A step of generating a dataset by acquiring multiple annotation images containing information on the cell regions of the multiple target images, Make it run, The dataset is a collection of multiple data elements, and each of the multiple data elements includes a combination of a target image and an annotation image. A program that, in step b) above, accepts the designation of a first designated region which is a region to be detected as a cell region and a second designated region which is a region which may or may not be detected as a cell region for each of the plurality of target images, thereby acquiring an annotation image corresponding to each of the target images.

8. A dataset generator that generates a dataset used for generating or evaluating a trained model for detecting cell regions in a target image, A memory unit that stores multiple target images, A dataset generation unit generates a dataset by acquiring multiple annotation images containing information on the cell regions of the multiple target images, Equipped with, The dataset is a collection of multiple data elements, and each of the multiple data elements includes a combination of a target image and an annotation image. A dataset generation device that acquires annotation images corresponding to each of the multiple target images by receiving a designation for each of the multiple target images of a first designated region which is a region that should be detected as a cell region and a second designated region which is a region that may or may not be detected as a cell region.