Image processing apparatus, image processing method, and program

JPWO2024084838A5Pending Publication Date: 2025-06-20
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024551299
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-04-08
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Endoscopic images often contain diverse lesions and varying environments, making it difficult for doctors to accurately detect lesion areas, leading to inconsistencies in identifying candidate biopsy sites.

Method used

An image processing device that acquires endoscopic images, generates multiple inference results using a lesion area inference model, and integrates these results using weighted averaging to accurately detect lesion areas, providing a consistent candidate site for biopsy.

Benefits of technology

The solution enables accurate detection of lesion areas within endoscopic images, improving consistency among doctors and supporting decision-making during endoscopic examinations.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

An image processing device 1X is provided with an acquisition means 30X, an inference means 32X, and an integration means 33X. The acquisition means 30X acquires an endoscopic image taken of a subject. The inference means 32X generates a plurality of inference results related to a site of interest of the subject in the endoscopic image, on the basis of the endoscopic image. The integration means 33X integrates the plurality of inference results.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method, and storage medium

[0001] The present disclosure relates to the technical fields of an image processing device, an image processing method, and a storage medium that process images acquired during an endoscopic examination.

[0002] Conventionally, endoscopic examination systems that display images of the inside of organ lumens have been known. For example, Patent Literature 1 discloses an endoscopic examination system that detects a region of interest based on an endoscopic image and a region of interest detection threshold, and determines whether the region of interest is a flat lesion or a raised lesion.

[0003] International Publication WO2019 / 146077

[0004] In general, endoscopic images may contain a wide variety of lesions, and the shooting environment of endoscopic images is also diverse, making it very difficult to accurately detect the lesion area. Therefore, even among doctors, there is sometimes disagreement about the lesion area that is a candidate site for biopsy.

[0005] In view of the above-mentioned problems, one object of the present disclosure is to provide an image processing device, an image processing method, and a storage medium that are capable of accurately detecting an area of ​​interest included in an endoscopic image.

[0006] One aspect of the image processing device is an image processing device having: an acquisition means for acquiring an endoscopic image of a subject; an inference means for generating a plurality of inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image; and an integration means for integrating the plurality of inference results.

[0007] One aspect of the image processing method is an image processing method in which a computer acquires an endoscopic image of a subject, generates a plurality of inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image, and integrates the plurality of inference results.

[0008] One aspect of the storage medium is a storage medium that stores a program that causes a computer to execute the following processes: acquire an endoscopic image of a subject; generate multiple inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image; and integrate the multiple inference results.

[0009] As an example of an effect of the present disclosure, it becomes possible to accurately detect a region of interest included in an endoscopic image.

[0010] 1 shows a schematic configuration of an endoscopic examination system. 2 shows a hardware configuration of an image processing device. 3 shows an overview of lesion detection processing performed by an image processing device in the first embodiment. 4 shows an example of functional blocks of lesion detection processing in the first embodiment. 5 shows (A) an example of calculating the similarity between a model input image and a representative image. 6 shows (B) an example of calculating the similarity between a lesion confidence map of a model input image and a representative image. 7 shows an example of a display screen displayed by a display device in endoscopic examination. 8 shows an example of a flowchart illustrating an overview of processing performed by an image processing device during endoscopic examination in the first embodiment. 9 shows an overview of lesion detection processing performed by an image processing device in the second embodiment. 10 is a functional block diagram of an image processing device related to lesion detection processing in the second embodiment. 11 is an example of a flowchart illustrating an overview of processing performed by an image processing device during endoscopic examination in the second embodiment. 12 is a diagram illustrating an overview of lesion detection processing performed by an image processing device in the third embodiment. 13 is an example of a flowchart illustrating an overview of processing performed by an image processing device during endoscopic examination in the third embodiment. 14 is a block diagram of an image processing device in the fourth embodiment. 15 is an example of a flowchart performed by an image processing device in the fourth embodiment.

[0011] Hereinafter, embodiments of an image processing device, an image processing method, and a storage medium will be described with reference to the drawings.

[0012] <First Embodiment> (1) System Configuration Fig. 1 shows a schematic configuration of an endoscopic examination system 100. As shown in Fig. 1, the endoscopic examination system 100 is a system that detects a region of a subject suspected of having a lesion (also referred to as a "lesion region") for an examiner such as a physician who performs an examination or treatment using an endoscope, and presents the region as a candidate region for cell sampling (biopsy). As a result, the endoscopic examination system 100 can support the examiner such as a physician in making decisions regarding, for example, determining a biopsy site, operating the endoscope, and determining a treatment plan for the subject of the examination. The endoscopic examination system 100 mainly includes an image processing device 1, a display device 2, and an endoscope 3 connected to the image processing device 1.

[0013] The image processing device 1 acquires images (also referred to as "endoscopic images Ia") captured by the endoscope 3 in time series from the endoscope 3, and displays a screen based on the endoscopic images Ia on the display device 2. The endoscopic images Ia are images captured at a predetermined frame rate during at least one of the steps of inserting or ejecting the endoscope 3 into the subject. In this embodiment, the image processing device 1 analyzes the endoscopic images Ia to detect the area of ​​the lesion site (also referred to as the "lesion area") in the endoscopic images Ia, and displays information related to the detection results on the display device 2. The lesion area is an example of a "region of interest."

[0014] The display device 2 is a display or the like that displays a predetermined image based on a display signal supplied from the image processing device 1 .

[0015] The endoscope 3 mainly comprises an operation unit 36 ​​for the examiner to input predetermined information, a flexible shaft 37 that is inserted into the subject's organ to be photographed, a tip 38 that incorporates an imaging unit such as a micro-imaging element, and a connection unit 39 for connecting to the image processing device 1.

[0016] 1 is an example, and various modifications may be made. For example, the image processing device 1 may be configured integrally with the display device 2. In another example, the image processing device 1 may be configured from multiple devices.

[0017] The subject of endoscopic examination in the present disclosure may be any organ that can be examined endoscopically, such as the large intestine, esophagus, stomach, pancreas, etc. Examples of endoscopes that are applicable in the present disclosure include pharyngoscopes, bronchoscopes, upper gastrointestinal endoscopes, duodenoscopes, small intestinal endoscopes, colonoscopes, capsule endoscopes, thoracoscopes, laparoscopes, cystoscopes, cholangioscopes, arthroscopes, spinal endoscopes, angioscopes, and epidural endoscopes. Examples of pathological conditions at lesion sites that are the target of endoscopic examination include the following (a) to (f).

[0018] (a) Head and neck: pharyngeal cancer, malignant lymphoma, papilloma (b) Esophagus: esophageal cancer, esophagitis, hiatal hernia, esophageal varices, esophageal achalasia, esophageal submucosal tumor, benign esophageal tumor (c) Stomach: gastric cancer, gastritis, gastric ulcer, gastric polyp, gastric tumor (d) Duodenum: duodenal cancer, duodenal ulcer, duodenitis, duodenal tumor, duodenal lymphoma (e) Small intestine: small intestine cancer, small intestine neoplastic disease, small intestine inflammatory disease, small intestine vascular disease (f) Large intestine: large intestine cancer, large intestine neoplastic disease, large intestine inflammatory disease, large intestine polyp, large intestine polyposis, Crohn's disease, colitis, intestinal tuberculosis, hemorrhoids

[0019] (2) Hardware Configuration Fig. 2 shows the hardware configuration of the image processing device 1. The image processing device 1 mainly includes a processor 11, a memory 12, an interface 13, an input unit 14, a light source unit 15, and a sound output unit 16. These elements are connected via a data bus 19.

[0020] The processor 11 performs predetermined processing by executing programs stored in the memory 12. The processor 11 is a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.

[0021] The memory 12 is composed of various volatile memories used as working memories, such as RAM (Random Access Memory) and ROM (Read Only Memory), and non-volatile memories that store information necessary for processing by the image processing device 1. The memory 12 may include an external storage device such as a hard disk connected to or built into the image processing device 1, or may include a storage medium such as a removable flash memory. The memory 12 stores programs for the image processing device 1 to execute each process in this embodiment.

[0022] The memory 12 also stores lesion region inference model information D1, which is information related to the lesion region inference model. The lesion region inference model is a machine learning model that generates inference results related to a lesion region corresponding to a disease to be detected in an endoscopic examination, and parameters required for the model are stored in the lesion region inference model information D1. For example, when an endoscopic image is input, the lesion region inference model outputs an inference result indicating a lesion region in the input endoscopic image. The lesion region inference model may be a model (including a statistical model, the same applies hereinafter) including an architecture adopted in any machine learning such as a neural network or a support vector machine. Representative models of such neural networks include, for example, Fully Convolutional Network, SegNet, U-Net, V-Net, Feature Pyramid Network, Mask R-CNN, and DeepLab. When the lesion region inference model is constructed using a neural network, the lesion region inference model information D1 includes various parameters such as the layer structure, the neuron structure of each layer, the number of filters and filter size in each layer, and the weight of each element of each filter.

[0023] Here, the inference result output by the lesion region inference model is, for example, a map of scores (also referred to as "lesion confidence scores") representing the confidence that each unit region of the input endoscopic image is a lesion region. Hereinafter, the above map will also be referred to as a "lesion confidence map." For example, the lesion confidence map is an image showing the lesion confidence score for each unit pixel (which may include subpixels) or pixel group. Note that the higher the lesion confidence score, the higher the confidence that the region is a lesion region. Note that the lesion confidence map may also be a mask image that shows the lesion region using binary values. In this way, the lesion region inference model is a model that learns the relationship between the image input to the lesion region inference model and the lesion region in that image.

[0024] The lesion region inference model is trained in advance based on a combination of an input image conforming to the input format of the lesion region inference model and correct answer data (in this embodiment, a correct lesion reliability map) indicating the correct inference result that the lesion region inference model should output when the input image is input. Then, the parameters of each model obtained by training are stored in memory 12 as lesion region inference model information D1.

[0025] The memory 12 may also contain other information that is necessary for the image processing device 1 to execute each process in this embodiment.

[0026] The lesion region inference model information D1 may be stored in a storage device separate from the image processing device 1. In this case, the image processing device 1 receives the lesion region inference model information D1 from the storage device.

[0027] The interface 13 acts as an interface between the image processing device 1 and an external device. For example, the interface 13 supplies the display information "Ib" generated by the processor 11 to the display device 2. The interface 13 also supplies light generated by the light source unit 15 to the endoscope 3. The interface 13 also supplies an electrical signal indicating the endoscopic image Ia supplied from the endoscope 3 to the processor 11. The interface 13 may be a communication interface such as a network adapter for wired or wireless communication with an external device, or may be a hardware interface compliant with USB (Universal Serial Bus), SATA (Serial AT Attachment), or the like.

[0028] The input unit 14 generates an input signal based on an operation by the examiner. The input unit 14 is, for example, a button, a touch panel, a remote controller, or a voice input device. The light source unit 15 generates light to be supplied to the tip 38 of the endoscope 3. The light source unit 15 may also incorporate a pump or the like for sending water or air to be supplied to the endoscope 3. The sound output unit 16 outputs sound based on the control of the processor 11.

[0029] (3) Lesion Detection Processing The lesion detection processing, which is processing related to the detection of lesion areas, will be described. In summary, the image processing device 1 expands the endoscopic image Ia to "N" images (N is an integer equal to or greater than 2) through data expansion (i.e., data augmentation), inputs each of the N images into a lesion area inference model, and integrates the inference results obtained. In this way, the image processing device 1 accurately detects lesion areas that are candidate sites for biopsy.

[0030] (3-1) Overview FIG. 3 is a diagram showing an overview of the lesion detection process executed by the image processing device 1 in the first embodiment.

[0031] First, the image processing device 1 generates N images (also referred to as "model input images") to be input into the lesion region inference model from each endoscopic image Ia obtained from the endoscope 3 at a predetermined frame period by data expansion. In the example of FIG. 3, the image processing device 1 generates four model input images (i.e., N = 4) by rotating the endoscopic image Ia clockwise by 0 degrees, 90 degrees, 180 degrees, and 270 degrees. In addition to rotation, any operation such as image resizing, brightness resizing (including specifying whether or not brightness normalization is performed), color resizing (including adjusting the intensity of redness), or a combination thereof may be used as a data expansion method.

[0032] Next, the image processing device 1 inputs each model input image to a lesion region inference model, and obtains a lesion confidence map, which is an inference result regarding a lesion region output by the lesion region inference model. In Figure 3, as an example, the lesion confidence map is a mask image that indicates whether or not a region is a lesion region using a binary value (here, white indicates a lesion region).

[0033] The image processing device 1 then generates an image (also referred to as an "integrated image") by integrating N (N = 4) lesion confidence maps using a weighted average. Here, weighting coefficients "wi" (i is an index of the inference result, i = 1, ..., N) whose sum total is 1 are used. The lesion confidence scores for each pixel of the N lesion confidence maps are averaged using the weighting coefficients to determine the lesion confidence score for each pixel of the integrated image. In Figure 3, the closer to white an area in the integrated image is, the higher the lesion confidence score (i.e., the higher the confidence that it is a lesion area). Note that if geometric image transformations such as image rotation or size enlargement / reduction are performed during data expansion, the images are integrated after being restored to the original angle and size. Therefore, in the example of Figure 3, the image processing device 1 performs a rotation operation on each lesion confidence map in the reverse direction (counterclockwise) of 0 degrees, 90 degrees, 180 degrees, and 270 degrees (i.e., a rotation operation that reverses the rotation operation performed during generation of the model input image), and then integrates the images obtained by the rotation operation using a weighted average.

[0034] The image processing device 1 then regards pixels in the integrated image that have pixel values ​​indicating a certain degree of certainty that they are lesion areas as lesion areas, and generates an image showing the final lesion detection result (here, a mask image representing the lesion area). The image processing device 1 displays this mask image together with the endoscopic image Ia.

[0035] Generally, the endoscopic image Ia used for lesion detection may contain a wide variety of lesions, and the imaging environment of the endoscopic image Ia may also vary widely, making accurate detection of the lesion area extremely difficult. For example, lesions contained in the endoscopic image Ia may be of various types, such as raised, flat, and depressed, and their shapes may change over time. Furthermore, the imaging environment may vary depending on the lesion location, lighting conditions, the presence or absence of water splashes, and the presence or absence of blur or blur. Therefore, even among physicians, there may be disagreement over the lesion area that is a candidate site for biopsy.

[0036] Taking the above into consideration, the image processing device 1 generates multiple inference results and integrates them to ultimately identify the lesion area, thereby making it possible to present the lesion area, which is a candidate for biopsy, to the examiner in an appropriate manner.

[0037] (3-2) Functional Blocks Figure 4 shows an example of functional blocks for lesion detection processing in the first embodiment. Functionally, the processor 11 of the image processing device 1 has an endoscopic image acquisition unit 30, a conversion unit 31, an inference unit 32, an integration unit 33, a lesion detection unit 34, and a display control unit 35. Note that in Figure 4, blocks between which data is exchanged are connected by solid lines, but the combination of blocks between which data is exchanged is not limited to this. The same applies to other functional block diagrams described below.

[0038] The endoscopic image acquisition unit 30 acquires endoscopic images Ia captured by the endoscope 3 via the interface 13 at predetermined intervals. The endoscopic image acquisition unit 30 then supplies the acquired endoscopic images Ia to the conversion unit 31 and the display control unit 35. Each processing unit in the subsequent stages then performs the processing described below, with the time interval at which the endoscopic image acquisition unit 30 acquires the endoscopic images Ia being used as a cycle. Hereinafter, the time for each frame cycle will also be referred to as the "processing time."

[0039] The converter 31 generates N model input images from the endoscopic image Ia through data expansion. In this case, the converter 31 generates N different model input images by, for example, performing a rotation operation, an image size change operation, a brightness change operation, a color change operation, or any combination of these operations on the endoscopic image Ia. Note that the data expansion method is not limited to the various operations exemplified above, and may be any operation used for data expansion. The converter 31 supplies the generated N model input images to the inference unit 32.

[0040] The inference unit 32 obtains N lesion confidence maps, which are inference results related to the lesion region, based on the N model input images and the lesion region inference model constructed by referencing the lesion region inference model information D1. In this case, the inference unit 32 inputs each of the N model input images to the lesion region inference model and obtains N lesion confidence maps output by the lesion region inference model. The inference unit 32 supplies the N lesion confidence maps to the integrating unit 33.

[0041] The integration unit 33 generates an integrated image by integrating N lesion confidence maps by weighted averaging. In this case, the inference unit 32 converts the angle and image size of each lesion confidence map to reverse the conversion operation performed by the conversion unit 31, then sets weighting coefficients wi (i = 1, ..., N) whose total value is 1 for each lesion confidence map. The inference unit 32 multiplies the lesion confidence scores for each pixel of the N lesion confidence maps by the corresponding weighting coefficient wi and adds the resulting values ​​to determine the lesion confidence score for each pixel of the integrated image. The weighting coefficient wi is determined, for example, based on the similarity between the corresponding model input image or lesion confidence map and an image representative of the input image or ground truth data used to train the lesion region inference model (also referred to as a "representative image"). In another example, the weighting coefficient wi is set to the same value (i.e., "1 / N") regardless of the index i so that the weights are uniform. The method for determining the weighting coefficient wi will be described later. The integration unit 33 supplies the generated integrated image to the lesion detection unit 34.

[0042] The lesion detection unit 34 determines the presence or absence of a lesion area based on the integrated image and identifies the lesion area if a lesion area exists. In this case, for example, the lesion detection unit 34 determines the presence of a lesion area if a predetermined number or more of pixels in the integrated image have a lesion confidence score equal to or greater than a predetermined threshold, and identifies pixels in the integrated image having a lesion confidence score equal to or greater than the predetermined threshold as a lesion area. Note that the lesion detection unit 34 may perform clustering on pixels having a lesion confidence score equal to or greater than the predetermined threshold, classifying adjacent pixels into the same cluster, and may consider a cluster having a predetermined number or more pixels to be a lesion area. The lesion detection unit 34 supplies the determination result of the presence or absence of a lesion area and information indicating the identified lesion area to the display control unit 35 as a lesion detection result.

[0043] The display control unit 35 generates display information Ib based on the latest endoscopic image Ia supplied from the endoscopic image acquisition unit 30 and the lesion detection result supplied from the lesion detection unit 34, and supplies the generated display information Ib to the display device 2, thereby displaying the latest endoscopic image Ia, the lesion detection result, etc. on the display device 2. Note that the display control unit 35 may also control the sound output of the sound output unit 16 so as to output a warning sound, voice guidance, etc., to notify the user that a lesion site has been detected, based on the lesion detection result.

[0044] The components of the endoscopic image acquisition unit 30, conversion unit 31, inference unit 32, integration unit 33, lesion detection unit 34, and display control unit 35 can be realized, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded on any non-volatile storage medium and installed as needed to realize each component. At least some of these components may not necessarily be realized by software programs, but may be realized by any combination of hardware, firmware, and software. At least some of these components may be realized using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, the integrated circuit may be used to realize a program consisting of the above components. Furthermore, at least a portion of each component may be configured by an ASSP (Application Specific Standard Product), an ASIC (Application Specific Integrated Circuit), or a quantum processor (quantum computer control chip). In this way, each component may be realized by various hardware. The same applies to other embodiments described below. Furthermore, each of these components may be realized by the cooperation of multiple computers, for example, using cloud computing technology.

[0045] (3-3) Example of Setting Weighting Coefficients Next, a description will be given of an example of setting the weighting coefficients wi by the integrating unit 33 when setting the weighting coefficients wi for each index i. In this case, the integrating unit 33 determines the weighting coefficients wi based on the similarity between each model input image or its lesion confidence map and a representative image that represents images including the lesion region to be detected.

[0046] 5A shows an example of calculating the similarity between a model input image of index i and a representative image. In this example, the integration unit 33 uses a training endoscopic image (also referred to as a “training lesion image”) that is used to train the lesion area inference model as the representative image, and calculates the similarity between the model input image and the representative image for each index i.

[0047] 5A, one arbitrary training lesion image is determined as the representative image, but the invention is not limited to this. For example, the integrating unit 33 may determine an average image of multiple training lesion images or an image integrated using any statistical method other than averaging as the representative image. In another example, the integrating unit 33 may determine each of multiple training lesion images as a representative image, and determine the average of the similarities between each training lesion image and the model input image of index i as the similarity used to determine the weighting coefficient wi.

[0048] Note that, as the similarity between the model input image and the representative image, any similarity index based on a comparison between the images (i.e., a comparison between images) may be calculated. In this case, examples of the similarity index include a correlation coefficient, a structural similarity (SSIM) index, a peak signal-to-noise ratio (PSNR) index, and the squared error between corresponding pixels. Furthermore, the integration unit 33 may vectorize the model input image and the representative image after normalizing their sizes, and calculate the cosine similarity of these vectors as the similarity.

[0049] The integrating unit 33 then calculates the above-described similarity for each index i, and sets a larger value for the weighting factor wi for the index i with a higher similarity. For example, if the similarity of index i is "Si", the integrating unit 33 sets the weighting factor wi using the following formula using "ΣSi" representing the total value of the similarities Si: wi = Si / ΣSi According to this example, it is possible to set a weighting factor wi such that the total value Σwi of all indexes i is 1, and the higher the corresponding similarity Si is, the higher the value of the weighting factor wi.

[0050] 5(B) shows an example of calculating the similarity between the lesion confidence map (here, a mask image) of the model input image with index i and the representative image. In this example, the integrating unit 33 uses, as the representative image, the correct lesion confidence map (here, a mask image indicating a lesion area) that should be output by the lesion region inference model annotated to the learning lesion image. The integrating unit 33 then calculates, for each index i, the similarity between the correct lesion confidence map used in learning and the lesion confidence map generated by the lesion region inference model from the model input image.

[0051] In the example of Figure 5(B), the correct lesion reliability map for any one learning lesion image is defined as the representative image, but this is not limited to this. The representative image may also be an average image of the correct lesion reliability maps for multiple learning lesion images or an image obtained by integrating lesion reliability maps using any statistical method other than averaging.

[0052] The integrating unit 33 then calculates an arbitrary similarity index based on a comparison between the images as the similarity between the lesion confidence map of the learning lesion image and the lesion confidence map of the model input image. The integrating unit 33 then sets the weighting coefficient wi so that the total value Σwi for all indexes i is 1 and the higher the corresponding similarity, the higher the value. Specific examples of the method for calculating the similarity and the method for setting the weighting coefficient wi based on the similarity are the same as those shown in FIG. 5(A).

[0053] (3-4) Display Example Fig. 6 shows an example of a display screen displayed by the display device 2 during an endoscopic examination. The display control unit 35 of the image processing device 1 outputs display information Ib generated based on the endoscopic image Ia acquired by the endoscopic image acquisition unit 30 and the lesion detection result by the lesion detection unit 34 to the display device 2. The display control unit 35 transmits the endoscopic image Ia and the display information Ib to the display device 2, thereby causing the display device 2 to display the above-mentioned display screen. In the display screen example shown in Fig. 6, the display control unit 35 of the image processing device 1 provides a real-time image display area 70 and a lesion detection result display area 71 on the display screen.

[0054] Here, the display control unit 35 displays a moving image representing the latest endoscopic image Ia in the real-time image display area 70. Furthermore, the display control unit 35 displays the lesion detection result by the lesion detection unit 34 in the lesion detection result display area 71. Note that, since the lesion detection unit 34 determined that a lesion area was present at the time the display screen shown in FIG. 6 was displayed, the display control unit 35 displayed a text message indicating the high possibility of the presence of a lesion and a mask image indicating the lesion area in the lesion detection result display area 71 based on the lesion detection result. Note that instead of or in addition to displaying the text message indicating the high possibility of the presence of a lesion in the lesion detection result display area 71, the display control unit 35 may output a sound (including audio) notifying the high possibility of the presence of a lesion from the sound output unit 16. Furthermore, the display control unit 35 may determine and output a countermeasure based on, for example, a model generated by machine learning the correspondence between lesion detection results and countermeasures and the lesion detection result of the subject. The method for determining a countermeasure is not limited to the above-described method. Outputting a countermeasure can further assist the examiner in making a decision.

[0055] (3-5) Processing Flow FIG. 7 is an example of a flowchart showing an outline of the processing executed by the image processing device 1 during endoscopic examination in the first embodiment.

[0056] First, the image processing device 1 acquires an endoscopic image Ia (step S11). In this case, the endoscopic image acquisition unit 30 of the image processing device 1 receives the endoscopic image Ia from the endoscope 3 via the interface 13.

[0057] Next, the image processing device 1 generates N different model input images from the endoscopic image Ia acquired in step S11 by data expansion (step S12).The image processing device 1 then generates a lesion confidence map from each model input image using a lesion region inference model configured with reference to the lesion region inference model information D1 (step S13).In this case, the image processing device 1 inputs each model input image to the lesion region inference model to obtain a lesion confidence map output from the lesion region inference model.

[0058] The image processing device 1 then calculates a weighting factor wi for each lesion reliability map (step S14). In this case, for example, the image processing device 1 sets the weighting factor wi for each index i (i = 1, ..., N) based on the similarity between the model input image or the lesion reliability map and the corresponding representative image. The image processing device 1 also converts the angle and size of each lesion reliability map to reverse the conversion operation due to data expansion in step S12.

[0059] Next, the image processing device 1 generates an integrated image by integrating the lesion reliability maps using the weighting coefficient wi (step S15).The image processing device 1 then generates a lesion detection result based on the integrated image (step S16).The image processing device 1 then displays information based on the endoscopic image Ia obtained in step S11 and the lesion detection result generated in step S16 on the display device 2 (step S17).

[0060] After step S17, the image processing device 1 determines whether the endoscopic examination has ended (step S18). For example, the image processing device 1 determines that the endoscopic examination has ended when it detects a predetermined input to the input unit 14 or the operation unit 36. If the image processing device 1 determines that the endoscopic examination has ended (step S18; Yes), it ends the processing of the flowchart. On the other hand, if the image processing device 1 determines that the endoscopic examination has not ended (step S18; No), it returns the processing to step S11. Then, the image processing device 1 performs the processing of steps S11 to S17 on the endoscopic image Ia newly generated by the endoscope 3.

[0061] (4) Modification The image processing device 1 may process a video image composed of endoscopic images Ia generated during an endoscopic examination after the examination.

[0062] For example, at any timing after an examination, when an image to be processed is designated based on user input via the input unit 14, the image processing device 1 sequentially performs the processing of the flowchart in Fig. 7 on the time-series endoscopic images Ia that constitute the designated image. Then, when it is determined in step S18 that the target image has ended, the image processing device 1 ends the processing of the flowchart, but when the target image has not ended, the image processing device 1 returns to step S11 and performs the processing of the flowchart on the next endoscopic image Ia in the time series.

[0063] Furthermore, the detection target is not limited to a lesion area, but may be an area in the endoscopic image Ia that represents an arbitrary area of ​​interest that the examiner needs to pay attention to (also referred to as a "area of ​​interest"). Such an area of ​​interest may be a lesion area, an area where inflammation has occurred, an area where a surgical scar or other incision has occurred, an area where a fold or protrusion has occurred, or an area where the tip 38 of the endoscope 3 is likely to come into contact (be easily pinched) with the wall surface inside the lumen.

[0064] This modification is also similarly applied to the second and third embodiments described later.

[0065] Second Embodiment An image processing device 1 according to the second embodiment differs from the first embodiment in that, instead of generating N lesion reliability maps from N model input images generated from an endoscopic image Ia, N lesion reliability maps are generated from the endoscopic image Ia using N different lesion region inference models. Hereinafter, components similar to those in the first embodiment will be appropriately designated by the same reference numerals, and their description will be omitted. The hardware configuration of the image processing device 1 according to the second embodiment is assumed to be the same as the configuration shown in FIG. 2 described in the first embodiment.

[0066] FIG. 8 is a diagram showing an outline of the lesion detection process executed by the image processing device 1 in the second embodiment.

[0067] First, the image processing device 1 inputs each endoscopic image Ia obtained at a predetermined frame cycle from the endoscope 3 into N lesion region inference models (here, models A to D). As a result, the image processing device 1 obtains a total of N lesion confidence maps from the N lesion region inference models.

[0068] Here, the N lesion region inference models differ from the other lesion region inference models in at least one of their architectures or the training data used for training, so that even when the same endoscopic image Ia is input, the N lesion region inference models generate different inference results.

[0069] An example of a different architecture is, for example, in the case of a deep learning model, when at least one of the layer structure, the neuron structure of each layer, the number and size of filters in each layer, and the weight of each element of each filter is different. Furthermore, the N lesion region inference models may include a model other than a deep learning model (e.g., a model based on a support vector machine) or a combination of a deep learning model and a model other than a deep learning model.

[0070] In an example where the training data is different, a set of training data consisting of pairs of endoscopic images and correct answer data for lesion regions (i.e., sets of training data corresponding to N vendors) is prepared for each endoscope vendor, and N lesion region inference models are trained using the training data set for each vendor. In another example where the training data is different, a set of training data consisting of pairs of endoscopic images and correct answer data for lesion regions (i.e., sets of training data corresponding to N lesion types) is prepared for each lesion type (protruding, flat, depressed, etc.), and N lesion region inference models are trained using the training data set for each lesion type.

[0071] The image processing device 1 then generates an integrated image by weighting and integrating the N lesion reliability maps using a weighting coefficient wi. In this case, the image processing device 1 determines the lesion reliability score for each pixel of the integrated image by adding values ​​obtained by multiplying the lesion reliability score for each pixel of the N lesion reliability maps by the corresponding weighting coefficient.

[0072] The image processing device 1 then regards pixels in the integrated image that have a lesion confidence score indicating that the confidence level of the lesion area is equal to or higher than a predetermined level as the lesion area, and generates an image showing the final lesion detection result (here, a mask image representing the lesion area). The image processing device 1 displays this mask image together with the endoscopic image Ia.

[0073] In this way, the image processing device 1 in the second embodiment generates multiple inference results and integrates the inference results to ultimately identify the lesion area, thereby making it possible to present the lesion area that is a candidate for biopsy to the examiner in an appropriate manner.

[0074] 9 is a functional block diagram of the image processing device 1 relating to the lesion detection process in the second embodiment. The processor 11 of the image processing device 1 according to the second embodiment functionally includes an endoscopic image acquisition unit 30A, an inference unit 32A, an integration unit 33A, a lesion detection unit 34A, and a display control unit 35A. Furthermore, the memory 12 stores lesion region inference model information D1 that includes at least trained parameters of N lesion region inference models.

[0075] The endoscopic image acquisition unit 30A acquires, at predetermined intervals, endoscopic images Ia captured by the endoscope 3 via the interface 13. Then, the endoscopic image acquisition unit 30A supplies the acquired endoscopic images Ia to the inference unit 32A and the display control unit 35A, respectively.

[0076] The inference unit 32A obtains N lesion reliability maps, which are inference results related to the lesion region, based on the endoscopic image Ia and N lesion region inference models constructed by referencing the lesion region inference model information D1. In this case, the inference unit 32A inputs the endoscopic image Ia to each of the N lesion region inference models and obtains N lesion reliability maps output by the lesion region inference models. The inference unit 32A supplies the N lesion reliability maps to the integrating unit 33A.

[0077] The integrating unit 33A generates an integrated image by integrating N lesion reliability maps by weighted averaging. In this case, for example, the integrating unit 33A sets the weighting coefficient wi to the same value (i.e., "1 / N") regardless of the index i. In another example, the integrating unit 33A sets the weighting coefficient wi based on the similarity between the lesion reliability map for each index i and the representative image. In this case, the representative image is, for example, the correct lesion reliability map used in training the lesion region inference model corresponding to index i. Note that the "correct lesion reliability map" includes an average image of lesion reliability maps indicated by correct data corresponding to multiple training lesion images or an image obtained by integrating the lesion reliability maps using any statistical method other than averaging. In this way, the representative image may be prepared in advance according to the training data used in the lesion region inference model corresponding to each index i.

[0078] Based on the integrated image generated by integration unit 33A, lesion detection unit 34A determines whether or not a lesion area exists and identifies the lesion area if one exists, and supplies the determination result of whether or not a lesion area exists and information indicating the identified lesion area to display control unit 35A as lesion detection results. Note that the processing performed by lesion detection unit 34A is the same as the processing performed by lesion detection unit 34.

[0079] The display control unit 35A generates display information Ib based on the latest endoscopic image Ia supplied from the endoscopic image acquisition unit 30A and the lesion detection result supplied from the lesion detection unit 34A, and supplies the generated display information Ib to the display device 2, thereby displaying the latest endoscopic image Ia, the lesion detection result, etc. on the display device 2. The processing executed by the display control unit 35A is the same as the processing executed by the display control unit 35.

[0080] FIG. 10 is an example of a flowchart showing an outline of the processing executed by the image processing device 1 during endoscopic examination in the second embodiment.

[0081] First, the image processing device 1 acquires an endoscopic image Ia (step S21). Next, the image processing device 1 generates N lesion reliability maps from the endoscopic image Ia acquired in step S11 using N lesion region inference models configured with reference to the lesion region inference model information D1 (step S22). In this case, the image processing device 1 inputs the endoscopic image Ia to each lesion region inference model to acquire lesion reliability maps output from each lesion region inference model.

[0082] The image processing device 1 then calculates a weighting factor wi for each lesion reliability map (step S23). In this case, for example, the image processing device 1 sets the weighting factor wi based on the similarity between the representative image prepared for each index i (i = 1, ..., N) and the lesion reliability map corresponding to the index i.

[0083] Next, the image processing device 1 generates an integrated image by integrating the lesion reliability maps using the weighting coefficient wi (step S24).The image processing device 1 then generates a lesion detection result based on the integrated image (step S25).The image processing device 1 then displays information based on the endoscopic image Ia obtained in step S11 and the lesion detection result generated in step S25 on the display device 2 (step S26).

[0084] After step S26, the image processing device 1 determines whether the endoscopic examination has ended (step S27). If the image processing device 1 determines that the endoscopic examination has ended (step S27; Yes), it ends the processing of the flowchart. On the other hand, if the image processing device 1 determines that the endoscopic examination has not ended (step S27; No), it returns the processing to step S21. Then, the image processing device 1 performs the processing of steps S21 to S26 on the endoscopic image Ia newly generated by the endoscope 3.

[0085] Third Embodiment An image processing device 1 according to the third embodiment differs from the first or second embodiment in that it applies setting conditions for N different patterns (N patterns) to one lesion region inference model to generate N lesion reliability maps from an endoscopic image Ia. Hereinafter, components similar to those in the first or second embodiment will be appropriately designated by the same reference numerals, and their description will be omitted.

[0086] The hardware configuration of the image processing device 1 according to the third embodiment is the same as the configuration shown in Fig. 2 described in the first embodiment. The functional blocks of the image processing device 1 related to the lesion detection process in the third embodiment are the same as the configuration shown in Fig. 9 described in the second embodiment, for example.

[0087] FIG. 11 is a diagram showing an outline of the lesion detection process executed by the image processing device 1 in the third embodiment.

[0088] First, the image processing device 1 inputs each endoscopic image Ia obtained from the endoscope 3 at a predetermined frame cycle to a lesion area inference model (here, model A) to which N patterns of setting conditions (here, setting conditions a to d) have been applied. As a result, the image processing device 1 obtains a total of N lesion reliability maps from the lesion area inference model to which the N patterns of setting conditions have been applied. In other words, the image processing device 1 inputs the endoscopic image Ia obtained at each processing time to the lesion area inference model N times while changing the setting conditions of the lesion area inference model, thereby obtaining N inference results output from the lesion area inference model.

[0089] Here, the setting condition may be, for example, a setting parameter of the lesion area inference model that can be adjusted by user input, and may be a threshold parameter that determines whether a pixel is a lesion area or not depending on the lesion confidence score of each pixel. Specifically, the lesion confidence score is set to a value between 0 and 1 (where 1 indicates that the pixel is most likely to be a lesion). When the lesion confidence score of a pixel is smaller than the threshold parameter, the lesion confidence score of the pixel can be set to 0, thereby classifying the pixel as a non-lesion area. In this case, for example, when the threshold parameter is set to a value close to 1, only the area that the inference model infers to be more likely to be a lesion will be classified as a lesion area, and the other areas will be classified as non-lesion areas. Conversely, when the threshold parameter is set to a value close to 0, areas that the inference model infers to be non-lesion will also be classified as lesion areas. The former setting emphasizes that the estimated lesion area is correct and prevents erroneous estimation of non-lesion areas as lesion areas (emphasis on precision), while the latter setting allows non-lesion areas to be included as lesion areas while preventing missed lesion detection (emphasis on recall). In this way, it is possible to generate respective reliability maps using the setting parameters of a plurality of lesion region inference models with different intentions (for example, whether emphasis is placed on recall or precision).

[0090] The image processing device 1 then generates an integrated image by weighting and integrating the N lesion reliability maps using a weighting factor wi. In this case, the image processing device 1 determines the pixel value of the integrated image by adding values ​​obtained by multiplying the lesion reliability scores for each pixel of the N lesion reliability maps by the corresponding weighting factor.

[0091] The image processing device 1 then regards pixels in the integrated image that have a lesion confidence score indicating that the confidence level of the lesion area is equal to or higher than a predetermined level as the lesion area, and generates an image showing the final lesion detection result (here, a mask image representing the lesion area). The image processing device 1 displays this mask image together with the endoscopic image Ia.

[0092] In this way, the image processing device 1 in the third embodiment generates multiple inference results and integrates the inference results to ultimately identify the lesion area, thereby making it possible to present the lesion area that is a candidate for biopsy to the examiner in an appropriate manner.

[0093] FIG. 12 is an example of a flowchart showing an outline of the processing executed by the image processing device 1 during endoscopic examination in the third embodiment.

[0094] First, the image processing device 1 acquires an endoscopic image Ia (step S31). Next, the image processing device 1 applies N patterns of setting conditions to one lesion area inference model configured with reference to the lesion area inference model information D1, and generates N lesion reliability maps from the endoscopic image Ia acquired in step S11 (step S32). In this case, the image processing device 1 inputs the endoscopic image Ia obtained at each processing time into the lesion area inference model N times while changing the setting conditions of the lesion area inference model, thereby acquiring N lesion reliability maps (i.e., inference results) output from the lesion area inference model.

[0095] Then, the image processing device 1 calculates a weighting factor wi for each lesion reliability map (step S33). In this case, for example, the image processing device 1 sets the weighting factor wi based on the similarity between a representative image common to all indexes i and the lesion reliability map corresponding to index i.

[0096] Next, the image processing device 1 generates an integrated image by integrating the lesion reliability maps using the weighting coefficient wi (step S34).The image processing device 1 then generates a lesion detection result based on the integrated image (step S35).The image processing device 1 then displays information based on the endoscopic image Ia obtained in step S11 and the lesion detection result generated in step S25 on the display device 2 (step S36).

[0097] After step S36, the image processing device 1 determines whether the endoscopic examination has ended (step S37). If the image processing device 1 determines that the endoscopic examination has ended (step S37; Yes), it ends the processing of the flowchart. On the other hand, if the image processing device 1 determines that the endoscopic examination has not ended (step S37; No), it returns the processing to step S31. Then, the image processing device 1 performs the processing of steps S31 to S36 on the endoscopic image Ia newly generated by the endoscope 3.

[0098] 13 is a block diagram of an image processing device 1X according to a fourth embodiment. The image processing device 1X includes an acquisition unit 30X, an inference unit 32X, and an integration unit 33X. The image processing device 1X may be composed of multiple devices.

[0099] The acquisition means 30X acquires an endoscopic image of a subject. The acquisition means 30X can be the endoscopic image acquisition unit 30 in the first embodiment or the endoscopic image acquisition unit 30A in the second or third embodiment. The acquisition means 30X may immediately acquire an endoscopic image generated by the imaging unit, or may acquire an endoscopic image generated in advance by the imaging unit and stored in a storage device at a predetermined timing.

[0100] The inference unit 32X generates a plurality of inference results related to the region of interest of the subject in the endoscopic image based on the endoscopic image. The inference unit 32X can be the inference unit 32 in the first embodiment or the inference unit 32A in the second or third embodiment.

[0101] The integration unit 33X integrates a plurality of inference results, and may be the integration unit 33 in the first embodiment or the integration unit 33A in the second or third embodiment.

[0102] 14 is an example of a flowchart showing a processing procedure in the fourth embodiment. The acquisition means 30X acquires an endoscopic image of a subject. The acquisition means 30X acquires an endoscopic image of a subject (step S41). Next, the inference means 32X generates multiple inference results related to the region of interest of the subject in the endoscopic image based on the endoscopic image (step S42). Then, the integration means 33X integrates the multiple inference results (step S43).

[0103] According to the fourth embodiment, the image processing device 1X can accurately detect the region of interest from an endoscopic image of a subject.

[0104] In each of the above-described embodiments, the program can be stored using various types of non-transitory computer-readable media and supplied to a computer processor, etc. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, semiconductor memories (e.g., mask ROMs, programmable ROMs (PROMs), erasable PROMs (EPROMs), flash ROMs, and random access memories (RAMs). The program may also be supplied to a computer by various types of transient computer-readable media. Examples of transient computer-readable media include electric signals, optical signals, and electromagnetic waves. The transient computer-readable medium can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0105] In addition, part or all of the above-described embodiments (including modified examples, the same applies below) can be described as, but are not limited to, the following supplementary notes.

[0106] [Supplementary Note 1] An image processing device comprising: an acquisition means for acquiring an endoscopic image of a subject; an inference means for generating a plurality of inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image; and an integration means for integrating the plurality of inference results. [Supplementary Note 2] The image processing device according to Supplementary Note 1, further comprising: a conversion means for converting the endoscopic image into a plurality of images by data augmentation, wherein the inference means generates an inference result regarding the region of interest from each of the plurality of images. [Supplementary Note 3] The image processing device according to Supplementary Note 2, wherein the inference means inputs each of the plurality of images into an inference model to acquire the inference result output from the inference model, the inference model being a model obtained by machine learning of the relationship between the image input to the inference model and the region of interest in the image. [Supplementary Note 4] The image processing device according to Supplementary Note 1, wherein the inference means inputs the endoscopic image into a plurality of inference models to acquire the plurality of inference results output from the plurality of inference models, the plurality of inference models being models obtained by machine learning of the relationship between the image input to the inference model and the region of interest in the image. [Supplementary Note 5] The image processing device according to Supplementary Note 4, wherein the multiple inference models are models that differ from each other in at least one of model architecture or learning data used for machine learning. [Supplementary Note 6] The image processing device according to Supplementary Note 1, wherein the inference means acquires the multiple inference results output from the inference model by inputting the endoscopic image into the inference model multiple times while changing the setting conditions of the inference model, and the multiple inference models are models that have undergone machine learning to determine the relationship between the image input to the inference model and the region of interest in the image. [Supplementary Note 7] The image processing device according to Supplementary Note 6, wherein the setting conditions are threshold parameters that determine whether or not the region is the region of interest. [Supplementary Note 8] The image processing device according to Supplementary Note 7, wherein the inference means acquires at least the inference results obtained when the threshold parameter that emphasizes recall and the threshold parameter that emphasizes precision are respectively set in the inference model.[Supplementary Note 9] The image processing device according to Supplementary Note 3, wherein the integration means weights and integrates each of the plurality of inference results based on the similarity between each of the plurality of images and a learning image including the region of interest used in machine learning of the inference model. [Supplementary Note 10] The image processing device according to any one of Supplementary Notes 3 to 8, wherein the integration means weights and integrates each of the plurality of inference results based on the similarity between each of the plurality of inference results and ground truth data used in machine learning of the inference model. [Supplementary Note 11] The image processing device according to Supplementary Note 1, further comprising detection means for detecting the region of interest based on an image obtained by integrating the plurality of inference results. [Supplementary Note 12] The image processing device according to Supplementary Note 9, further comprising output control means for displaying or outputting information related to the detection result as audio. [Supplementary Note 13] The image processing device according to Supplementary Note 12, wherein the output control means outputs information related to the detection result to support an examiner in making a decision. [Supplementary Note 14] An image processing method, in which a computer acquires an endoscopic image of a subject, generates a plurality of inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image, and integrates the plurality of inference results. [Supplementary Note 15] A storage medium storing a program that causes a computer to execute a process of acquiring an endoscopic image of a subject, generating a plurality of inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image, and integrating the plurality of inference results.

[0107] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications within the scope of the present invention that would be understood by those skilled in the art can be made to the configuration and details of the present invention. In other words, the present invention naturally includes various modifications and alterations that would be possible for those skilled in the art based on the entire disclosure, including the claims, and the technical ideas. Furthermore, the disclosures of the above-cited patent and non-patent documents are incorporated herein by reference.

[0108] REFERENCE SIGNS LIST 1, 1X image processing device 2 display device 3 endoscope 11 processor 12 memory 13 interface 14 input unit 15 light source unit 16 sound output unit 100 endoscopic examination system

Claims

1. An acquisition means for acquiring an endoscopic image of a subject; an inference means for generating a plurality of inference results related to the region of interest of the subject in the endoscopic image based on the endoscopic image; An integration means for integrating the plurality of inference results; An image processing device comprising:

2. A conversion means for converting the endoscopic image into a plurality of images by data expansion, The image processing apparatus according to claim 1 , wherein the inference means generates an inference result regarding the region of interest from each of the plurality of images.

3. the inference means inputs each of the plurality of images into an inference model to obtain the inference result output from the inference model; The image processing device according to claim 2 , wherein the inference model is a model that is machine-learned to understand the relationship between an image input to the inference model and the region of interest in the image.

4. the inference means inputs the endoscopic image into a plurality of inference models to obtain the plurality of inference results output from the plurality of inference models; The image processing device according to claim 1 , wherein the plurality of inference models are models that are machine-learned to understand the relationship between an image input to the inference model and the region of interest in the image.

5. The image processing device according to claim 4 , wherein the multiple inference models are models that differ from each other in at least one of the model architecture or the learning data used in machine learning.

6. the inference means inputs the endoscopic image into an inference model a plurality of times while changing a setting condition of the inference model, thereby obtaining the plurality of inference results output from the inference model; The image processing device according to claim 1 , wherein the inference model is a model that is machine-learned to understand the relationship between an image input to the inference model and the region of interest in the image.

7. The image processing device according to claim 6 , wherein the setting condition is a threshold parameter for determining whether or not the region is the region of interest.

8. The image processing device according to claim 7 , wherein the inference means at least acquires the inference result obtained when the threshold parameter that emphasizes recall and the threshold parameter that emphasizes precision are respectively set in the inference model.

9. The image processing device described in claim 3, wherein the integration means weights and integrates each of the multiple inference results based on the similarity between each of the multiple images and a learning image including the region of interest used in machine learning of the inference model.

10. The image processing device according to any one of claims 3 to 8, wherein the integration means weights and integrates each of the multiple inference results based on the similarity between each of the multiple inference results and correct answer data used for machine learning of the inference model.

11. The image processing apparatus according to claim 1 , further comprising a detection unit that detects the region of interest based on an image obtained by integrating the plurality of inference results.

12. The image processing device according to claim 9 , further comprising an output control unit that displays or outputs information related to the result of the detection.

13. The image processing apparatus according to claim 12 , wherein the output control means outputs information about the result of the detection to assist an inspector in making a decision.

14. The computer An endoscopic image of the subject is acquired, generating a plurality of inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image; Integrating the multiple inference results; Image processing methods.

15. An endoscopic image of the subject is acquired, generating a plurality of inference results regarding a region of interest of the subject in the endoscopic image based on the endoscopic image; A program that causes a computer to execute a process of integrating the multiple inference results.