Image processing device, image processing method, and storage medium
Patent Information
- Application Number
- PCT/JP2024/008087
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-10-02
AI Technical Summary
Existing systems require significant time and effort to input prompts for multiple images when using generative AI to extract specific image areas, especially in applications like endoscopic imaging where multiple images of the same subject need to be analyzed for abnormalities.
An image processing device that automatically generates prompts for multiple images based on user input in one image, using image registration and a machine-learned model to determine the relationship between input images and prompts, reducing the need for manual input across a series of images.
Enables efficient and accurate setting of prompts for multiple images showing the same subject, thereby reducing the effort required for prompt input and improving the accuracy of image analysis, particularly in endoscopic imaging for detecting lesions.
Smart Images

Figure JP2024008087_02102025_PF_FP_ABST
Abstract
Description
Image processing device, image processing method, and storage medium
[0001] The present disclosure relates to the technical fields of an image processing device, an image processing method, and a storage medium that process images.
[0002] Conventionally, systems that use machine learning models to detect image regions that represent abnormal conditions of an object in an image have been known. For example, Patent Literature 1 discloses an information processing device that uses visible light images and infrared images to detect abnormalities in an object based on segmentation processing using a segmentation model that has been machine-learned in advance.
[0003] Japanese Patent Application Laid-Open No. 2023-174268
[0004] When using generative AI that can accept prompts, which are information suggesting the area to be extracted, to accurately extract a specific image area, the task of inputting the prompts is required for each image to be input. Therefore, in this case, the more images there are to detect the area, the more time and effort is required to input the prompts.
[0005] In view of the above-mentioned problems, one of the objects of the present disclosure is to provide an image processing device, an image processing method, and a storage medium that are capable of appropriately setting prompts for multiple images showing the same subject.
[0006] One aspect of the image processing device is an image processing device comprising: an image acquisition means for acquiring a plurality of images showing a subject; a first prompt acquisition means for acquiring a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image; a second prompt generation means for generating information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image; and an inference means for generating an inference result of the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that has been machine-learned to determine the relationship between an input image input to the machine learning model, the prompt suggesting the region of interest in the input image, and the region of interest in the input image.
[0007] One aspect of the image processing method is an image processing method characterized in that a computer acquires a plurality of images depicting a subject, acquires a first prompt indicating a position or area in a first image included in the plurality of images and suggesting an area of interest in the first image, generates information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the area of interest in the second image, and generates an inference result of the area of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that has learned by machine learning the relationship between an input image input to the machine learning model and the prompt suggesting the area of interest in the input image, and the area of interest in the input image.
[0008] One aspect of the storage medium is a storage medium storing a program that acquires a plurality of images depicting a subject, acquires a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image, generates information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image, and causes a computer to execute a process of generating an inference result for the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that machine-learns the relationship between an input image input to the machine learning model and the prompt suggesting the region of interest in the input image, and the region of interest in the input image.
[0009] As an example of an effect of the present disclosure, it becomes possible to appropriately set prompts for multiple images showing the same subject.
[0010] 8A , 8B, 8C, 8D, 8E, 8F, 8G, 8H, 8G, 8H, 8I, 8K ...
[0011] Hereinafter, embodiments of an image processing device, an image processing method, and a storage medium will be described with reference to the drawings.
[0012] <First Embodiment> (1) System Configuration Fig. 1 shows a schematic configuration of an image processing system 100. The image processing system 100 is a system that generates an image showing a region of interest from each of a plurality of images taken of the same subject, and mainly includes an image processing device 1, a storage device 2, a display device 3, and an input device 4. The "region of interest" is an image region that shows a notable state of the subject (such as an abnormal state), and the "subject" is any object that is the target for detecting the region of interest.
[0013] In this embodiment, as an example, the "subject" refers to an organ to be examined using an endoscope, the "image of the subject" refers to an image captured by the endoscope, and the "region of interest" refers to an image region (also referred to as a "lesion region") that represents a lesion or a suspected lesion in the subject. In this case, the subject may be any organ that can be examined using an endoscope, such as the large intestine, esophagus, stomach, or pancreas. Examples of endoscopes that are applicable to the present disclosure include pharyngeal endoscopes, bronchoscopes, upper gastrointestinal endoscopes, duodenoscopes, small intestinal endoscopes, colonoscopes, capsule endoscopes, thoracoscopes, laparoscopes, cystoscopes, cholangioscopes, arthroscopes, spinal endoscopes, angioscopes, and epidural endoscopes. Furthermore, examples of pathological conditions at lesion sites that are the target of endoscopic examination include the following (a) to (f).
[0014] (a) Head and neck: pharyngeal cancer, malignant lymphoma, papilloma (b) Esophagus: esophageal cancer, esophagitis, hiatal hernia, esophageal varices, esophageal achalasia, esophageal submucosal tumor, benign esophageal tumor (c) Stomach: gastric cancer, gastritis, gastric ulcer, gastric polyp, gastric tumor (d) Duodenum: duodenal cancer, duodenal ulcer, duodenitis, duodenal tumor, duodenal lymphoma (e) Small intestine: small intestine cancer, small intestine neoplastic disease, small intestine inflammatory disease, small intestine vascular disease (f) Large intestine: large intestine cancer, large intestine neoplastic disease, large intestine inflammatory disease, large intestine polyp, large intestine polyposis, Crohn's disease, colitis, intestinal tuberculosis, hemorrhoids
[0015] The image processing device 1 acquires multiple images of the same subject from the storage device 2 and generates an image representing a lesion area from each of the acquired images. In this embodiment, the image processing device 1 generates, as an example of an image representing a lesion area, a mask image that indicates, using a binary value, whether or not an area corresponds to a lesion area.
[0016] The storage device 2 is a memory that stores various information necessary for processing by the image processing device 1 , and includes an image storage unit 21 and a base model information storage unit 22 .
[0017] The image storage unit 21 stores a plurality of images (also referred to as a "target image group") of the same subject. The target image group is, for example, a group of images consisting of two or more images of a common lesion area of the subject. When the target image group is designated based on a user input via the input device 4, the image storage unit 21 may store images (e.g., an image database) that are candidates to be designated as the target image group by the user input.
[0018] The base model information D2 is information about the base model, which is a machine learning model, and includes trained parameters of the base model. The base model is a large-scale deep learning neural network trained using a diverse and large-scale dataset, and is an AI model known as a generative AI. The base model in this embodiment is a generative AI that generates an image (e.g., a mask image) showing a segmentation result from an image input to the base model based on a prompt, which is information representing an instruction (suggestion) input by a user. Examples of such generative AI include SAM (Segment Anything Model) and SEEM (Segment Everything Everywhere All at Once). For example, SAM is a generative AI that receives an image and a prompt (segmentation prompt) suggesting an object to be segmented in the image as input, and generates a mask image (segmentation mask) representing the region of an arbitrary object as an inference result. SAM accepts prompts such as points on an image suggesting an object to be segmented, a bounding box (rectangular frame), a mask image, text, or any combination thereof. The base model in this embodiment accepts prompts such as points or frames on an image suggesting an area to be segmented. Details of SAM are described in, for example, the following literature: Kirillov A, Mintun E, Ravi N, Mao HZ, Rolland C, Gustafson L, Xiao TT, Whitehead S, Berg AC, Lo WY, Dollar P, Girshick R. Segment anything. arXiv preprint arXiv:2304.02643, 2023
[0019] The display device 3 displays information under the control of the image processing device 1. Examples of the display device 3 include a display, a projector, etc. The display device 3 displays information based on display information supplied from the image processing device 1.
[0020] The input device 4 is an interface that accepts user input, which is external input based on user operation, and includes, for example, a touch panel, buttons, a keyboard, and a voice input device. The input device 4 supplies input information generated based on the user input to the image processing device 1. Note that the "user" in this embodiment is, for example, a doctor or other medical professional who inputs a prompt regarding the lesion area.
[0021] The configuration of the image processing system 100 shown in FIG. 1 is an example, and various modifications may be made to the configuration. For example, the image processing device 1, the storage device 2, the display device 3, and the input device 4 may be integrated into any combination. The image processing system 100 may also include a sound output device such as a speaker. The image processing device 1 may also be composed of multiple devices. In this case, the multiple devices that make up the image processing device 1 exchange information required to execute pre-assigned processes between these multiple devices.
[0022] (2) Hardware Configuration Fig. 2 shows the hardware configuration of the image processing device 1. The image processing device 1 includes, as hardware, a processor 11, a memory 12, and an interface 13. The processor 11, the memory 12, and the interface 13 are connected via a data bus 19.
[0023] The processor 11 executes predetermined processes by executing programs stored in the memory 12. The processor 11 is a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.
[0024] The memory 12 is composed of various types of volatile and non-volatile memories, such as RAM (Random Access Memory) and ROM (Read Only Memory). The memory 12 also stores programs for the image processing device 1 to execute various processes. The memory 12 is also used as a working memory, and temporarily stores information obtained from the storage device 2. The memory 12 may also function as the storage device 2. Similarly, the storage device 2 may also function as the memory 12 of the image processing device 1. The programs executed by the image processing device 1 may be stored in a storage medium other than the memory 12.
[0025] The interface 13 is an interface for electrically connecting the image processing device 1 to other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data to and from other devices, or may be hardware interfaces for connecting to other devices via cables or the like.
[0026] The hardware configuration of the image processing device 1 is not limited to the configuration shown in Fig. 2. For example, the image processing device 1 may include at least one of the display device 3 and the input device 4. The image processing device 1 may also be connected to or have a built-in sound display device such as a speaker.
[0027] (3) Overview of Mask Image Generation Process An overview of the mask image generation process, which is a process for generating mask images for each image from a group of target images, is described below. The image processing device 1 generally accepts user input specifying a prompt for a lesion area in an image included in the group of target images, and automatically generates prompts to be used for other images based on image alignment (so-called image registration) from the specified prompt. The image processing device 1 then generates mask images for each image by executing a base model using the prompt specified by the user input and the automatically generated prompt. This allows the image processing device 1 to generate mask images that accurately represent the lesion area in each image of the group of target images, while reducing the effort required for entering prompts.
[0028] Hereinafter, a prompt specified by a user input will be referred to as a "first prompt," and a prompt automatically generated by the image processing device 1 from the first prompt will be referred to as a "second prompt."
[0029] 3 is a diagram showing an outline of the mask image generation process, in which the image processing device 1 acquires endoscopic images Im1 to Im3 of the same lesion as a target image group.
[0030] In this case, the image processing device 1 first accepts a user input of a prompt for any image in the target image group. In this example, the image processing device 1 acquires three user inputs specifying points on the endoscopic image Im1 that correspond to the lesion area, and generates a first prompt specifying three locations on the endoscopic image Im1. Note that in this example, the image processing device 1 generates a first prompt indicating the points that constitute the lesion area. However, instead, the image processing device 1 may generate a first prompt indicating a frame (bounding box) that surrounds the lesion area. The endoscopic image Im1 is an example of a "first image."
[0031] Furthermore, the image processing device 1 performs image registration (i.e., a process for determining the correspondence between pixels on the images) on the endoscopic images Im1 to Im3 to recognize the correspondence between pixels between the endoscopic images Im1 to Im3. The image processing device 1 may perform image registration before generating the first prompt. In this case, image registration may be performed as preprocessing, and information indicating the correspondence between pixels between the endoscopic images Im1 to Im3 may be stored in the storage device 2 as prior information. In this case, the image processing device 1 acquires information indicating the correspondence between pixels between the endoscopic images Im1 to Im3 from the storage device 2. The image processing device 1 may also use any method for image registration. For example, the image processing device 1 may perform image registration by matching features between images and image warping, or may perform image registration based on deep learning (homography learning).
[0032] Next, the image processing device 1 identifies three locations on the endoscopic image Im2 and three locations on the endoscopic image Im3 that correspond to the three locations on the endoscopic image Im1, based on the correspondences between the endoscopic images Im1 to Im3 recognized by image registration. The image processing device 1 then generates second prompts indicating the identified three locations on the endoscopic image Im2 and three locations on the endoscopic image Im3, respectively. The endoscopic images Im2 and Im3 are examples of "second images."
[0033] The image processing device 1 then generates mask images of lesion areas corresponding to the endoscopic images Im1 to Im3 using a base model constructed with reference to parameters stored in the base model information storage unit 22. In this case, when the endoscopic image Im1 is used as input to the base model, the image processing device 1 generates a mask image corresponding to the endoscopic image Im1 by inputting three first prompts specified by user input to the base model as prompts suggesting the lesion area. On the other hand, when the endoscopic image Im2 is used as input to the base model, the image processing device 1 generates a mask image corresponding to the endoscopic image Im2 by inputting second prompts indicating three locations on the endoscopic image Im2 that correspond to the three first prompts as prompts suggesting the lesion area. Furthermore, when the endoscopic image Im3 is used as input to the base model, the image processing device 1 generates a mask image corresponding to the endoscopic image Im3 by inputting second prompts indicating three locations on the endoscopic image Im3 that correspond to the three first prompts as prompts suggesting the lesion area.
[0034] In the example of Figure 3, the image processing device 1 can generate a mask image that represents the lesion area in each image of the target image group with high accuracy while reducing the effort required to input prompts into endoscopic image Im2 and endoscopic image Im3.
[0035] The image processing device 1 may set the first prompt for a plurality of images instead of setting the first prompt for one image (endoscopic image Im1 in the example of FIG. 3 ). A specific example of this will be described later in the display example.
[0036] Furthermore, after setting a first prompt for an arbitrary image in the target image group, the image processing device 1 may accept user input specifying a location (point or area) in another image that corresponds to the first prompt. In this case, the image processing device 1 considers the first prompt and the location specified by the user input as a pair of corresponding locations between the two images, and performs image registration on the premise that this pair of corresponding locations matches. In this case, for example, when the image processing device 1 accepts a correction to the second prompt, it considers the corrected location and the first prompt as the pair of corresponding locations. A specific example of this will also be described in the display example below.
[0037] (4) Functional Blocks Fig. 4 shows an example of functional blocks of the image processing device 1. As shown in Fig. 4, the processor 11 of the image processing device 1 functionally includes an image acquisition unit 31, a first prompt acquisition unit 32, a second prompt generation unit 33, a mask image generation unit 34, and a display control unit 35. Note that in the figure, blocks between which data is exchanged are connected by solid lines, but the combination of blocks between which data is exchanged is not limited to Fig. 4. The same applies to other functional block diagrams described later.
[0038] The image acquisition unit 31 acquires a group of target images for generating a mask image from the image storage unit 21 via the interface 13. The image acquisition unit 31 may specify images to be acquired as the group of target images from the images held in the image storage unit 21 based on input information supplied from the input device 4. The image acquisition unit 31 supplies the group of target images to the first prompt acquisition unit 32.
[0039] The first prompt acquisition unit 32 accepts a user input specifying a prompt for an arbitrary image in the target image group supplied from the image acquisition unit 31, and acquires a first prompt for the arbitrary image based on the user input. In this case, the first prompt acquisition unit 32, for example, supplies the target image group to the display control unit 35 and causes the display control unit 35 to display the target image group. Then, the first prompt acquisition unit 32 detects a user input specifying a location (point or area) suggesting a lesion area on an arbitrary image in the displayed target image group based on input information generated by the input device 4. In this case, the first prompt acquisition unit 32 generates a first prompt indicating the location specified by the user input and associates the first prompt with the image for which the user input was made. The first prompt acquisition unit 32 supplies the target image group and the first prompt to the second prompt generation unit 33.
[0040] The second prompt generation unit 33 generates a second prompt based on the target image group and the first prompt. In this case, the second prompt generation unit 33 recognizes pixel correspondences between the target images by performing image registration of the target images and identifies a location in another image that corresponds to a location indicated by the first prompt set in one image. The second prompt generation unit 33 then generates a second prompt indicating the identified location in the other image and associates the other image with the generated second prompt. The second prompt generation unit 33 then supplies the target image group, the first prompt, and the second prompt to the mask image generation unit 34.
[0041] When the second prompt generation unit 33 receives input information from the input device 4 specifying a pair of corresponding locations between two images in the target image group, the second prompt generation unit 33 performs image registration so that the corresponding locations match between the two images. In this case, for example, the second prompt generation unit 33 performs optimization to maximize the evaluation index of image registration, with the matching pair of corresponding locations as a constraint. This allows the second prompt generation unit 33 to use the pair of matching corresponding locations between the two images as prior knowledge for image registration, thereby improving the accuracy of image registration. The above-described pair of corresponding locations is also used to generate mask images as prompts (first prompts) for the two images after image registration.
[0042] The mask image generation unit 34 generates a mask image showing the lesion area in each image of the target image group based on each image of the target image group and the first prompt and second prompt associated with each image of the target image group. In this case, the image processing device 1 uses a base model configured with reference to parameters stored in the base model information storage unit 22 to generate a mask image showing the lesion area in each image. In this case, the first prompt and / or second prompt associated with the image to be input to the base model is input to the base model as a prompt suggesting the lesion area in the input image. Then, the mask image generation unit 34 acquires a mask image for each image based on the inference result output by the base model upon input of each image and prompt to the base model. The mask image generation unit 34 supplies each image of the target image group, the corresponding mask image, and the prompt (including the first prompt and the second prompt) used to generate the mask image to the display control unit 35. Note that the mask image generation unit 34 may store the generated mask image in the image storage unit 21 in association with the corresponding image.
[0043] The display control unit 35 displays information regarding at least one of the target image group, the prompt (including the first prompt and the second prompt), and the mask image on the display device 3. In this case, the display control unit 35 generates display information and supplies the generated display information to the display device 3, thereby causing the display device 3 to display predetermined information. Examples of the display that the display control unit 35 causes the display device 3 to display will be described later.
[0044] Here, each of the components, including the image acquisition unit 31, the first prompt acquisition unit 32, the second prompt generation unit 33, the mask image generation unit 34, and the display control unit 35, can be realized, for example, by the processor 11 executing a program. Alternatively, each component may be realized by recording the necessary program on any non-volatile storage medium and installing it as needed. Note that at least a portion of these components may not necessarily be realized by software programs, but may also be realized by any combination of hardware, firmware, and software. Also, at least a portion of these components may be realized using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, the integrated circuit may be used to realize a program consisting of the above components. Also, at least a portion of the components may be configured by an ASSP (Application Specific Standard Produce), an ASIC (Application Specific Integrated Circuit), or a quantum processor (quantum computer control chip). In this way, each component may be realized by various hardware. The same applies to other embodiments described below. Furthermore, each of these components may be realized by the cooperation of multiple computers, for example, using cloud computing technology.
[0045] (5) Display Example Next, the display control of the display device 3 executed by the display control unit 35 will be described. Fig. 5 is an example of a display screen showing a target image group, a prompt, and a mask image. The display control unit 35 of the image processing device 1 transmits display information generated based on the mask image, the target image group used to generate the mask image, and the prompt (including the first prompt and the second prompt) to the display device 3, thereby causing the display screen shown in Fig. 5 to be displayed on the display device 3.
[0046] The display control unit 35 displays endoscopic images 70A and 70B (here, "captured images"), which are the target image group. Furthermore, the display control unit 35 clearly displays a first prompt (here, "user-specified prompt") and a second prompt (here, "automatically generated prompt") on the endoscopic images 70A and 70B. Furthermore, the display control unit 35 displays mask images 71A and 71B (here, "lesion detection results (mask images)") generated from the endoscopic images 70A and 70B, respectively, in association with the endoscopic images 70A and 70B.
[0047] The display control unit 35 also displays on the display screen an add prompt button 72, a modify prompt button 73, and an input completion button 74. The add prompt button 72 is a button for adding a first prompt ("user-specified prompt" in FIG. 5), and the modify prompt button 73 is a button for modifying a second prompt ("automatically generated prompt" in FIG. 5). The input completion button 74 is a button for instructing completion of input related to the prompt.
[0048] After detecting that the prompt addition button 72 has been selected, if an arbitrary location on the endoscopic image 70A or 70B is designated, the display control unit 35 supplies information indicating the designated location to the first prompt acquisition unit 32. In this case, the first prompt acquisition unit 32 acquires a first prompt indicating the designated location.
[0049] When a new first prompt is acquired, the second prompt generator 33 generates a second prompt corresponding to the acquired first prompt, and the mask image generator 34 generates mask images for the endoscopic images 70A and 70B based on the first and second prompts. The display controller 35 then displays the first and second prompts on the endoscopic images 70A and 70B, and updates the mask images 71A and 71B with the generated mask images. In this way, the display controller 35 updates the display on the display screen in real time in accordance with the added first prompt each time a first prompt is added.
[0050] 5, first prompts (here, "user-specified prompts") are set on two images (endoscopic image 70A and endoscopic image 70B). In this way, the display control unit 35 can accept a user's operation to specify a location that is clearly a lesion area from any image in the target image group, and can appropriately generate the first prompt. Endoscopic image 70A and endoscopic image 70B are examples of a first image and a second image, and when any one of the images is considered to be the first image, the other becomes the second image.
[0051] Furthermore, when the display control unit 35 detects that the prompt correction button 73 has been selected, it accepts a correction of the position of the second prompt ("automatically generated prompt" in FIG. 5 ) displayed on the endoscopic image 70A and the endoscopic image 70B. For example, the display control unit 35 accepts a user input to move the mark of the automatically generated prompt to an arbitrary position by a drag-and-drop operation. In this case, the display control unit 35 supplies the corrected second prompt to the second prompt generation unit 33. In this case, the second prompt generation unit 33 considers the position of the corrected second prompt and the position of the corresponding first prompt as a pair of corresponding locations between the images and re-performs image registration so that these corresponding locations match. Then, the second prompt generation unit 33 corrects the positions of the other second prompts based on the result of image registration, and the mask image generation unit 34 updates the mask image based on the corrected second prompt.
[0052] In this way, if the user thinks that the accuracy of the automatically generated second prompt position is low, the user can correct the position of the second prompt based on the user input, etc. This allows the image processing device 1 to perform more accurate image registration that reflects the correction, and obtain a more accurate mask image.
[0053] (6) Processing Flow FIG. 6 is an example of a flowchart showing an outline of the processing executed by the image processing device 1.
[0054] First, the image acquisition unit 31 of the image processing device 1 acquires a group of target images of the same subject from the image storage unit 21 (step S11). In this case, the image acquisition unit 31 may regard a plurality of images specified by a user input indicated by the input information generated by the input device 4 as the group of target images, and acquire the plurality of images from the image storage unit 21.
[0055] Next, the first prompt acquisition unit 32 of the image processing device 1 accepts input related to a prompt (step S12). In this case, the display control unit 35 displays each image of the target image group as shown in FIG. 5, for example, and accepts user input specifying an arbitrary location in an arbitrary image. Then, the first prompt acquisition unit 32 sets a first prompt indicating the location specified by the accepted user input to the image in which the location is specified. The first prompt acquisition unit 32 may also accept input related to modifying an already generated second prompt.
[0056] The second prompt generation unit 33 of the image processing device 1 then performs image registration between the target images to generate a second prompt (step S13). In this case, the second prompt generation unit 33 recognizes the correspondence between pixels by image registration between the image for which the first prompt is set and each of the other images, and generates a second prompt that indicates the portion of the other images that corresponds to the first prompt.
[0057] Next, the mask image generation unit 34 of the image processing device 1 uses the prompts obtained in steps S12 and S13 to obtain a mask image inferred by the base model from each image of the target image group (step S14). In this case, each input image and the corresponding prompt are input to the base model to which the parameters stored in the base model information storage unit 22 are applied, and a mask image is obtained, which is the inference result output by the base model in response to the input.
[0058] The image processing device 1 then determines whether the input related to the prompt has been completed (step S15). If the input related to the prompt has been completed (step S15; No), the image processing device 1 ends the processing of the flowchart. For example, if the image processing device 1 detects a user input indicating that the input related to the prompt has been completed, the image processing device 1 determines that the input related to the prompt has been completed. On the other hand, if the input related to the prompt has not been completed (step S15; Yes), the image processing device 1 returns the processing to step S12 and accepts the input related to the prompt.
[0059] (7) Modifications Next, suitable modifications of the above-described embodiment will be described. The following modifications may be applied in combination to the above-described embodiment.
[0060] (Modification 1) The base model may be executed by a device other than the image processing device 1.
[0061] 7 is a schematic diagram of the image processing system 100A. For simplicity, the storage device 2 and the display device 3 are not shown. The image processing system 100A includes a server device 5 that stores base model information D2. The server device 5 may be configured from multiple devices. The image processing system 100A also includes multiple image processing devices 1 (1A, 1B, ...) that are capable of data communication with the server device 5 via a network.
[0062] Each image processing device 1 transmits each image of the target image group and the corresponding prompt to the server device 5, and in response obtains the inference result of the base model from the server device 5. When the server device 5 receives each image of the target image group and the corresponding prompt from the image processing device 1, it inputs each image and the corresponding prompt to the base model constructed based on the base model information D2, and transmits the inference result output by the base model in response to the input to the image processing device 1. In this case, the interface 13 of each image processing device 1 includes a communication interface such as a network adapter for communication.
[0063] In this way, even when the base model is executed by a device other than the image processing device 1, the image processing device 1 can preferably obtain the inference results of the base model.
[0064] (Modification 2) The target images from which a mask image is generated may be any medical images.
[0065] FIG. 8(A) is a pathological tissue image of excised stomach cells, and FIG. 8(B) is a partial image of the pathological tissue image shown in FIG. 8(A). The images shown in FIG. 8(A) and FIG. 8(B) represent the same subject (specimen) and form a target image group. When the image processing device 1 sets a first prompt for the image shown in FIG. 8(A) based on user input, it sets a second prompt indicating the corresponding portion of the other image (see FIG. 8(B)) based on the results of image registration. In this way, the image processing device 1 may regard a group of medical images of the same subject other than endoscopic images as a target image group and generate a mask image for each image in the target image group.
[0066] Furthermore, the target image group is not limited to a group of medical images, but may be a group of images of the same subject (more specifically, a region of interest of the same subject). In this case, a mask image representing a lesion region may be generated from the target image group, and a mask image representing a region of interest that represents an abnormal region other than the lesion region may be generated. In this case, the subject may be any object from which a region of interest is detected.
[0067] (Variation 3) The image processing device 1 may set a mask image generated from an image to which a first prompt is set as a prompt for a base model when generating a mask image from another image in a target image group.
[0068] 9 is a diagram showing an overview of a process for setting a mask image generated from an image in which a first prompt has been set as a prompt for a base model when generating a mask image from another image. In the example of FIG. 9, endoscopic image Im11 and endoscopic image Im12 are present as a target image group, three first prompts have been set for endoscopic image Im11 based on user input, and three second prompts have been generated for endoscopic image Im12 in accordance with the above-mentioned three first prompts. The image processing device 1 then inputs endoscopic image Im11 and the three first prompts into the base model to generate a mask image corresponding to endoscopic image Im11.
[0069] In this modification, the image processing device 1 generates a mask image corresponding to the endoscopic image Im12 based on the endoscopic image Im12, three second prompts, a mask image corresponding to the endoscopic image Im11, and a base model. In this case, the image processing device 1 inputs the endoscopic image Im12 to the base model, and also inputs the three second prompts and the mask image corresponding to the endoscopic image Im11 as prompts to the base model. As a result, the base model generates a mask image of the endoscopic image Im12 as an inference result, taking into account the mask image corresponding to the endoscopic image Im11. Therefore, in this case, the image processing device 1 can generate mask images of other images so that they resemble the mask image of the image for which the first prompt is set. This modification is particularly effective when the images in a group of target images are similar (for example, when the images are generated consecutively in chronological order, such as endoscopic images).
[0070] (Modification 4) The image processing apparatus 1 may use a mask image generated from an image to which the first prompt is set to correct a mask image generated from another image in the target image group.
[0071] 10 is a diagram illustrating an overview of a process in which a mask image generated from an image in which a first prompt is set is used to correct the mask image from another image. In the example of FIG. 10 , endoscopic image Im11 and endoscopic image Im12 are included as a target image group. Three first prompts are set for endoscopic image Im11 based on user input, and three second prompts are generated for endoscopic image Im12 in response to the three first prompts. In this case, the image processing device 1 generates a mask image corresponding to endoscopic image Im11 by inputting endoscopic image Im11 and the first prompt into a base model. Furthermore, the image processing device 1 generates a mask image corresponding to endoscopic image Im12 by inputting endoscopic image Im12 and the second prompt into the base model.
[0072] Furthermore, the image processing device 1 according to this modification corrects the mask image corresponding to the endoscopic image Im12 using the mask image corresponding to the endoscopic image Im11. In this case, for example, the image processing device 1 performs image registration between the mask images and transforms the mask image (image warping) so that the feature points of the mask image of the endoscopic image Im12 are closer to the feature points of the mask image of the corresponding endoscopic image Im11. Alternatively, the image processing device 1 may generate a mask image by integrating two mask images based on any statistical processing as a corrected mask image corresponding to the endoscopic image Im12. This allows the image processing device 1 to generate the mask image of the endoscopic image Im12 so that it resembles the mask image of the endoscopic image Im11. Therefore, the image processing device 1 can correct the mask image of another image so that it resembles the mask image corresponding to the image for which the first prompt is set. This modification is particularly effective when the images in the target image group are similar (for example, when the images are generated consecutively in chronological order, such as endoscopic images).
[0073] (Variation 5) The image processing apparatus 1 may select mask images generated from other images in the target image group based on a mask image generated from an image to which a first prompt is set.
[0074] 11 is a diagram showing an overview of a process in which a mask image generated from an image in which a first prompt has been set is used to select mask images generated from other images. In the example of FIG. 11, endoscopic images Im11 to Im13 are included as a group of target images, and three first prompts have been set for endoscopic image Im11 based on user input, and three second prompts have been generated for endoscopic image Im12 and endoscopic image Im13, respectively, in accordance with the three first prompts.
[0075] In this case, the image processing device 1 generates a mask image Im21 corresponding to the endoscopic image Im11 by inputting the endoscopic image Im11 and the first prompt into the base model. The image processing device 1 also generates a mask image Im22 corresponding to the endoscopic image Im12 by inputting the endoscopic image Im12 and the corresponding second prompt into the base model. The image processing device 1 also generates a mask image Im23 corresponding to the endoscopic image Im13 by inputting the endoscopic image Im13 and the corresponding second prompt into the base model.
[0076] Next, the image processing device 1 calculates the similarity between the mask image Im21 generated from the image to which the first prompt is set and other mask images (mask images Im22 and Im23), and excludes mask images for which the similarity is equal to or less than a predetermined level, as mask images with low accuracy. In the example of FIG. 11 , the image processing device 1 determines that the similarity between the mask image Im21 and the mask image Im23 is equal to or less than a predetermined level, and excludes the mask image Im23. The predetermined level is, for example, a default value stored in advance in the storage device 2 or the memory 12. The similarity may be any index value used to measure the similarity between images, such as SSIM (Structural Similarity), MSE (Mean Squared Error), or NCC (Normalized Cross Correlation). The similarity may be an index value indicating the degree of overlap of the lesion areas between two images, such as Intersection over Union (IoU). Alternatively, the similarity may be an index value calculated based on a deep learning model.
[0077] Here, a specific example of excluding a mask image will be described. For example, the image processing device 1 does not display the mask image to be excluded on the display screen, and displays the other mask images that were not excluded on the display screen. In another example, the image processing device 1 does not store the mask image to be excluded in the image storage unit 21, and stores the other mask images that were not excluded in the image storage unit 21.
[0078] According to this modification, the image processing device 1 can preferably exclude mask images that are not similar to the mask image of the image to which the first prompt is set.
[0079] (Variation 6) The image processing device 1 may select a mask image generated from another image based on an index value that evaluates the accuracy of image registration between the image for which the first prompt is set and the other image. In this case, the index value that evaluates the accuracy of image registration is an evaluation function value that indicates the degree of alignment (matching degree) between images used in optimization of image registration, and any index value such as cross-correlation or mutual information may be used.
[0080] For example, when an index value for evaluating the accuracy of image registration between the image for which the first prompt is set and the other image is equal to or lower than a predetermined level, the image processing device 1 excludes the mask image of the other image. The predetermined level is a default value pre-stored in the storage device 2 or the memory 12. According to this aspect, the image processing device 1 can preferably exclude the mask image of an image that is not similar to the image for which the first prompt is set.
[0081] 12 is a block diagram of an image processing device 1X. The image processing device 1X includes an image acquisition unit 31X, a first prompt acquisition unit 32X, a second prompt generation unit 33X, and an inference unit 34X. The image processing device 1X may be composed of multiple devices.
[0082] The image acquisition means 31X acquires a plurality of images depicting the subject. The "plurality of images depicting the subject" is, for example, the target image group in the first embodiment. The image acquisition means 31X can be, for example, the image acquisition unit 31 in the first embodiment.
[0083] The first prompt acquisition unit 32X acquires a first prompt indicating a position or area within a first image included in the plurality of images and indicating a region of interest in the first image. The first prompt acquisition unit 32X may be, for example, the first prompt acquisition unit 32 in the first embodiment.
[0084] The second prompt generating unit 33X generates information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting a region of interest in the second image. The second prompt generating unit 33X may be, for example, the second prompt generating unit 33 in the first embodiment.
[0085] The inference means 34X generates an inference result of the region of interest in the second image based on the second image, the second prompt, and the machine learning model. Here, the machine learning model is a model obtained by machine learning of the relationship between the input image input to the machine learning model, the prompt suggesting the region of interest in the input image, and the region of interest in the input image. The machine learning model is, for example, the base model in the first embodiment. The inference means 34X can be, for example, the mask image generation unit 34 in the first embodiment.
[0086] 13 is an example of a flowchart showing the processing procedure of the image processing device 1X. First, the image acquisition unit 31X acquires multiple images depicting a subject (step S21). Next, the first prompt acquisition unit 32X acquires a first prompt indicating a position or area in a first image included in the multiple images and suggesting a region of interest in the first image (step S22). The second prompt generation unit 33X generates information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the multiple images as a second prompt suggesting a region of interest in the second image (step S23). The inference unit 34X generates an inference result for the region of interest in the second image based on the second image, the second prompt, and the machine learning model (step S24).
[0087] According to the second embodiment, the image processing device 1X can suitably acquire an inference result regarding the region of interest of the subject for each of a plurality of images showing the subject.
[0088] In each of the above-described embodiments, the program can be stored using various types of non-transitory computer-readable media and supplied to a computer processor, etc. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, semiconductor memories (e.g., mask ROMs, programmable ROMs (PROMs), erasable PROMs (EPROMs), flash ROMs, and random access memories (RAMs). The program may also be supplied to a computer by various types of transient computer-readable media. Examples of transient computer-readable media include electric signals, optical signals, and electromagnetic waves. The transient computer-readable medium can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.
[0089] In addition, part or all of the above-described embodiments (including variations, the same applies below) may also be described as, but are not limited to, the following supplementary notes. Furthermore, not only the devices, methods, and storage media described in the supplementary notes, but also various hardware, software, various recording means (including storage media) for recording software, or systems may be made to depend on part or all of the configurations described in the supplementary notes, as long as they do not deviate from the above-described embodiments.
[0090] [Supplementary Note 1] An image processing device comprising: an image acquisition means for acquiring a plurality of images depicting a subject; a first prompt acquisition means for acquiring a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image; a second prompt generation means for generating information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image; and an inference means for generating an inference result of the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that has learned by machine learning a relationship between an input image input to the machine learning model, the prompt suggesting the region of interest in the input image, and the region of interest in the input image. [Supplementary Note 2] The image processing device according to Supplementary Note 1, wherein the second prompt generation means recognizes a correspondence between pixels in the first image and the second image by aligning the first image and the second image, and generates the second prompt based on the correspondence. [Supplementary Note 3] The image processing device according to Supplementary Note 2, wherein the second prompt generation means, when a pair of corresponding locations in the first image and the second image is specified, performs the image registration such that the pair of corresponding locations matches. [Supplementary Note 4] The image processing device according to Supplementary Note 3, wherein the corresponding location in the first image is a position or area indicated by the first prompt, and the corresponding location in the second image is a position or area where the position or area indicated by the second prompt has been corrected by external input. [Supplementary Note 5] The image processing device according to Supplementary Note 2, wherein the inference means selects an inference result for the region of interest in the second image based on an index for evaluating the accuracy of the image registration. [Supplementary Note 6] The image processing device according to Supplementary Note 1, wherein, when a position or area on any one of the multiple images displayed on a display device is specified by external input, the first prompt acquisition means acquires information indicating the specified position or area as the first prompt suggesting the region of interest in the any one of the multiple images displayed on a display device.[Supplementary Note 7] The image processing device according to Supplementary Note 1, wherein the inference means generates an inference result for the region of interest in the first image based on the first image, the first prompt, and the machine learning model. [Supplementary Note 8] The image processing device according to Supplementary Note 7, wherein, when the inference means generates an inference result for the region of interest in the second image using the machine learning model, the inference result for the region of interest in the first image and the second prompt are input as the prompt. [Supplementary Note 9] The image processing device according to Supplementary Note 7, wherein the inference means corrects the inference result for the region of interest in the second image based on the inference result for the region of interest in the first image. [Supplementary Note 10] The image processing device according to Supplementary Note 7, wherein the first prompt acquisition means acquires a plurality of first prompts respectively corresponding to a plurality of first images, the second prompt generation means generates, for each of the plurality of first images, a second prompt based on a first prompt corresponding to another of the first images, and the inference means generates an inference result for the region of interest in each of the plurality of first images based on each of the plurality of first images, the first prompt and the second prompt corresponding to each of the plurality of first images, and the machine learning model. [Supplementary Note 11] The image processing device according to Supplementary Note 1, further comprising display control means for displaying, on a display device, the second image showing the position or area indicated by the second prompt and the inference result for the region of interest in the second image in association with each other. [Supplementary Note 12] The image processing device according to Supplementary Note 1, wherein the inference means generates, as the inference result, a mask image representing the region of interest. [Supplementary Note 13] The image processing device according to Supplementary Note 1, wherein the plurality of images are medical images showing the same lesion, and the region of interest is an image region representing the lesion.[Supplementary Note 14] An image processing method, in which a computer acquires a plurality of images depicting a subject, acquires a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image, generates information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image, and generates an inference result of the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that has learned by machine learning a relationship between an input image input to the machine learning model and the prompt suggesting the region of interest in the input image, and the region of interest in the input image. [Supplementary Note 15] A storage medium storing a program, comprising: acquiring a plurality of images depicting a subject; acquiring a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image; generating information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image; causing a computer to execute a process of generating an inference result for the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that has learned by machine learning a relationship between an input image input to the machine learning model, the prompt suggesting the region of interest in the input image, and the region of interest in the input image. [Supplementary Note 16] The image processing method according to Supplementary Note 14, comprising: recognizing a correspondence between pixels in the first image and the second image by aligning the first image and the second image; and generating the second prompt based on the correspondence. [Supplementary Note 17] The image processing method according to Supplementary Note 16, wherein when a pair of corresponding portions in the first image and the second image is specified, the image registration is performed such that the pair of corresponding portions matches.[Supplementary Note 18] The image processing method according to Supplementary Note 17, wherein the corresponding location in the first image is a position or area indicated by the first prompt, and the corresponding location in the second image is a position or area where the position or area indicated by the second prompt has been corrected by external input. [Supplementary Note 19] The image processing method according to Supplementary Note 16, wherein an inference result for the region of interest in the second image is selected based on an index for evaluating the accuracy of the image alignment. [Supplementary Note 20] The image processing method according to Supplementary Note 14, wherein, when a position or area is specified by external input on any one of the multiple images displayed on a display device, information indicating the specified position or area is obtained as the first prompt suggesting a region of interest in the any one of the multiple images. [Supplementary Note 21] The image processing method according to Supplementary Note 14, wherein an inference result for the region of interest in the first image is generated based on the first image, the first prompt, and the machine learning model. [Supplementary Note 22] The image processing method according to Supplementary Note 21, wherein, when the inference means generates an inference result for the region of interest in the second image using the machine learning model, the inference means inputs the inference result for the region of interest in the first image and the second prompt as the prompt. [Supplementary Note 23] The image processing method according to Supplementary Note 21, wherein the inference result for the region of interest in the second image is corrected based on the inference result for the region of interest in the first image. [Supplementary Note 24] The image processing method according to Supplementary Note 21, further comprising: acquiring a plurality of first prompts respectively corresponding to a plurality of first images; generating, for each of the plurality of first images, the second prompt based on a first prompt corresponding to another of the first images; and generating an inference result for the region of interest in each of the plurality of first images based on each of the plurality of first images, the first prompt and the second prompt corresponding to each of the plurality of first images, and the machine learning model. [Supplementary Note 25] The image processing method according to Supplementary Note 14, wherein the second image showing the position or area indicated by the second prompt and an inference result of the region of interest in the second image are displayed in association with each other on a display device. [Supplementary Note 26] The image processing method according to Supplementary Note 14, wherein a mask image representing the region of interest is generated as the inference result.[Supplementary Note 27] The image processing method according to Supplementary Note 14, wherein the plurality of images are medical images representing the same lesion, and the region of interest is an image region representing the lesion.
[0091] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications within the scope of the present invention that would be understood by those skilled in the art can be made to the configuration and details of the present invention. In other words, the present invention naturally includes various modifications and alterations that would be possible for those skilled in the art based on the entire disclosure, including the claims, and the technical ideas. Furthermore, the disclosures of the above-cited patent and non-patent documents are incorporated herein by reference.
[0092] REFERENCE SIGNS LIST 1, 1X image processing device 2 storage device 3 display device 4 input device 5 server device 11 processor 12 memory 13 interface 21 image storage unit 22 base model information storage unit 100 image processing system
Claims
1. An image processing device comprising: an image acquisition means for acquiring a plurality of images depicting a subject; a first prompt acquisition means for acquiring a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image; a second prompt generation means for generating information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image; and an inference means for generating an inference result of the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model obtained by machine learning of the relationship between an input image input to the machine learning model and the prompt suggesting the region of interest in the input image, and the region of interest in the input image.
2. An image processing device as described in claim 1, wherein the second prompt generation means recognizes the correspondence between pixels of the first image and the second image by aligning the first image and the second image, and generates the second prompt based on the correspondence.
3. The image processing device according to claim 2, wherein the second prompt generating means, when a pair of corresponding portions in the first image and the second image is specified, performs the image registration such that the pair of corresponding portions matches.
4. An image processing device as described in claim 3, wherein the corresponding location in the first image is a position or area indicated by the first prompt, and the corresponding location in the second image is a position or area indicated by the second prompt that has been modified by external input.
5. The image processing device according to claim 2, wherein the inference means selects an inference result for the region of interest in the second image based on an index for evaluating the accuracy of the image alignment.
6. An image processing device as described in claim 1, wherein when a position or area on any one of the plurality of images displayed on the display device is specified by external input, the first prompt acquisition means acquires information indicating the specified position or area as the first prompt suggesting a region of interest in the any one of the images.
7. The image processing device according to claim 1, wherein the inference means generates an inference result for the region of interest in the first image based on the first image, the first prompt, and the machine learning model.
8. The image processing device described in claim 7, wherein when the inference means generates an inference result for the region of interest in the second image using the machine learning model, the inference result for the region of interest in the first image and the second prompt are input as the prompt.
9. The image processing device according to claim 7, wherein said inference means corrects the inference result of said region of interest in said second image based on the inference result of said region of interest in said first image.
10. The image processing device described in claim 7, wherein the first prompt acquisition means acquires a plurality of first prompts corresponding respectively to a plurality of first images; the second prompt generation means generates, for each of the plurality of first images, a second prompt based on a first prompt corresponding to another of the first images; and the inference means generates an inference result for the region of interest in each of the plurality of first images based on each of the plurality of first images, the first prompt and the second prompt corresponding to each of the plurality of first images, and the machine learning model.
11. An image processing device as described in claim 1, further comprising a display control means for displaying on a display device the second image showing the position or area indicated by the second prompt and the inference result for the area of interest in the second image in correspondence with each other.
12. The image processing device according to claim 1, wherein said inference means generates a mask image representing said region of interest as the inference result.
13. The image processing device according to claim 1, wherein the plurality of images are medical images showing the same lesion, and the region of interest is an image region showing the lesion.
14. An image processing method in which a computer acquires a plurality of images depicting a subject, acquires a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image, generates information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image, and generates an inference result for the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that has learned by machine learning the relationship between an input image input to the machine learning model and the prompt suggesting the region of interest in the input image, and the region of interest in the input image.
15. A storage medium storing a program that includes: acquiring a plurality of images depicting a subject; acquiring a first prompt indicating a position or area in a first image included in the plurality of images and suggesting a region of interest in the first image; generating information indicating a position or area corresponding to the first prompt in a second image other than the first image included in the plurality of images as a second prompt suggesting the region of interest in the second image; and causing a computer to execute a process of generating an inference result for the region of interest in the second image based on the second image, the second prompt, and a machine learning model, wherein the machine learning model is a model that machine-learns the relationship between an input image input to the machine learning model and the prompt suggesting the region of interest in the input image, and the region of interest in the input image.
16. The image processing method according to claim 14, wherein the correspondence between pixels of the first image and the second image is recognized by aligning the first image and the second image, and the second prompt is generated based on the correspondence.
17. The image processing method according to claim 16, wherein, when a pair of corresponding portions in the first image and the second image is designated, the image registration is performed such that the pair of corresponding portions matches.
18. An image processing method as described in claim 17, wherein the corresponding location in the first image is a position or area indicated by the first prompt, and the corresponding location in the second image is a position or area indicated by the second prompt that has been modified by external input.
19. The image processing method according to claim 16, wherein the inference result of the region of interest in the second image is selected based on an index for evaluating the accuracy of the image registration.
20. An image processing method as described in claim 14, wherein, when a position or area is specified by external input on any one of the plurality of images displayed on a display device, information indicating the specified position or area is obtained as the first prompt suggesting an area of interest in the any one of the images.