Data processing device, data processing method, data processing program, endoscope system, medical system, display control method, insertion guide method, learning method for inference model, and learning model
The described system addresses the challenge of endoscope insertion by using AI to analyze endoscopic images and construct an inference model that assists in navigating difficult-to-insert areas, enhancing insertion success for both experts and novices.
Patent Information
- Application Number
- PCT/JP2024/006332
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-21
- Publication Date
- 2025-08-28
AI Technical Summary
Existing endoscope insertion methods are challenging, particularly for non-experts, due to the need for pre-created 3D models of lumens which may not accurately represent dynamic or collapsed lumens, and lack of real-time assistance during insertion.
A data processing device and method that utilizes endoscopic images to construct an inference model by learning from insertion processes, identifying difficult-to-insert portions and providing guidance through an inference model trained on chronological image groups.
Enables effective insertion assistance by constructing an inference model that supports insertion into difficult areas using AI, even for beginners, by analyzing image features and providing real-time guidance.
Smart Images

Figure JP2024006332_28082025_PF_FP_ABST
Abstract
Description
Data processing device, data processing method, data processing program, endoscope system, medical system, display control method, insertion guide method, inference model learning method and learning model
[0001] The present invention relates to a data processing device, a data processing method, a data processing program, a medical system, a display control method, an insertion guide method, an inference model learning method, and a learning model for assisting insertion into a lumen.
[0002] An endoscope is a device that enables observation of diseased areas that cannot be seen from outside the body by inserting an imaging unit (composed of an imaging device) into the body or the like and irradiating illumination light from an attached illumination unit to obtain imaging results (endoscopically acquired images, endoscopic images) from the imaging unit. When a physician visually checks the imaging results on a display or the like, white light is typically used as the light source for the illumination unit, but other light sources may also be used. An endoscope has an insertion section that is inserted into a body cavity, and the imaging device is provided, for example, at the tip of the insertion section. During an examination using an endoscope, a physician sequentially displays images acquired by the imaging device at the tip of the insertion section inserted into the body, and adjusts the position of the tip of the insertion section while checking the displayed images to diagnose the body's health and disease state. Image information acquired from the beginning of insertion of the endoscope into the body to its removal from the body may be digitized and recorded.
[0003] However, inserting an endoscope into the body is a relatively difficult task, and unless the person is an expert, it may take a long time to insert the endoscope into the target area.
[0004] Japanese Patent Application Publication No. 2022-105685 (hereinafter referred to as Patent Document 1) discloses a technology for assisting the insertion of an endoscope by creating a 3D model of a lumen and referring to the modeled lumen.
[0005] Japanese Patent Application Laid-Open No. 2022-105685
[0006] However, to utilize the proposal in Patent Document 1, there is a problem in that a 3D model of the lumen must be created in advance. Furthermore, the condition of the lumen may change between the time of modeling and the time of endoscopic examination. Furthermore, narrow or collapsed lumens may not be reproducible through modeling due to limitations in the resolution of the information used for modeling. For example, the portion from the duodenum through the papilla to the bile duct or pancreatic duct may be collapsed by the sphincter, making accurate 3D modeling difficult. The present invention aims to provide a data processing device, a data processing method, a data processing program, a medical system, a display control method, an insertion guide method, a learning method for an inference model, and a learning model that enable the construction of an inference model that effectively assists the insertion of an insertion instrument by performing learning using endoscopic images obtained during difficult endoscope insertion.
[0007] A data processing device according to one aspect of the present invention includes a difficult-to-insert portion image determination unit that receives imaged images obtained in chronological order by imaging a lumen portion into which an insertion instrument is inserted and determines whether the input imaged images are difficult-to-insert portion images obtained at a difficult-to-insert portion of the insertion instrument; a success timing acquisition unit that acquires, as success timing, the timing at which a successful insertion image obtained when the insertion instrument is successfully inserted into the lumen; a feature determination unit that receives the imaged images and determines features of an intermediate image group that is a group of images captured between the imaging timing of the difficult-to-insert portion image and the successful timing; and an association unit that associates and records the difficult-to-insert portion image with information on the features of the intermediate image group corresponding to the difficult-to-insert portion image.
[0008] An endoscopic system according to one aspect of the present invention includes an endoscope including an imaging device that images the lumen into which the insertion instrument is inserted and obtains the captured images in chronological order; an endoscope that provides the captured images obtained by the endoscope to the inference model; and a control unit that provides the captured images obtained by the endoscope to the inference model and causes the inference model to output, as an inference result, the features of the intermediate image group corresponding to the difficult-to-insert portion.
[0009] A medical system according to one aspect of the present invention includes an inference model obtained by learning based on an image of a difficult-to-insert portion obtained by imaging the difficult-to-insert portion of an insertion instrument into a lumen and feature information about a group of intermediate images acquired in chronological order between the acquisition of the image of the difficult-to-insert portion and the acquisition of an insertion success image acquired when the insertion instrument is successfully inserted into the lumen; an endoscope including an imaging device that images the lumen into which the insertion instrument is inserted to obtain images in chronological order; a control unit that provides the images acquired by the endoscope to the inference model and causes the inference model to output, as inference results, features of the group of intermediate images corresponding to the difficult-to-insert portion; an external medical device that images the endoscope, the insertion instrument, and the lumen; and a display control unit that displays the images acquired by the endoscope, images acquired by imaging the external medical device, and the inference result.
[0010] A data processing method according to one aspect of the present invention inputs captured images obtained in chronological order by imaging a lumen into which an insertion instrument is inserted, determines whether the input captured images are difficult-to-insert portion images acquired at a difficult-to-insert portion of the insertion instrument, inputs the captured images and determines whether the input captured images are successful insertion images acquired when the insertion instrument is successfully inserted into the lumen, inputs the captured images and determines the characteristics of an intermediate image group, which is a group of images captured between the capture of the difficult-to-insert portion image and the capture of the successful insertion image, and records the difficult-to-insert portion image in association with information on the characteristics of the intermediate image group corresponding to the difficult-to-insert portion image.
[0011] A data processing program according to one aspect of the present invention causes a computer to execute the following steps: inputting captured images obtained in chronological order by imaging a lumen into which an insertion instrument is inserted, determining whether the input captured images are difficult-to-insert portion images acquired at a difficult-to-insert portion of the insertion instrument, inputting the captured images and determining whether the input captured images are successful insertion images acquired when the insertion instrument was successfully inserted into the lumen, inputting the captured images and determining the characteristics of an intermediate image group that is a group of images captured between the capture of the difficult-to-insert portion image and the capture of the successful insertion image, and recording the difficult-to-insert portion image and information on the characteristics of the intermediate image group that correspond to the difficult-to-insert portion image in association with each other.
[0012] A display control method according to one aspect of the present invention involves using an imaging device provided in an endoscope to image a lumen into which an insertion instrument is inserted, obtaining imaged images in chronological order; providing the images obtained by the endoscope to an inference model obtained by learning based on an image of the difficult-to-insert portion obtained by imaging the difficult-to-insert portion of the insertion instrument into the lumen, and information on the features of a group of intermediate images obtained in chronological order between the acquisition of the image of the difficult-to-insert portion and the acquisition of an image of successful insertion obtained by imaging when the insertion instrument is successfully inserted into the lumen; and outputting and displaying the features of the group of intermediate images corresponding to the difficult-to-insert portion from the inference model as inference results.
[0013] A display control method according to one aspect of the present invention involves using an imaging device provided in an endoscope to image a lumen into which an insertion instrument is inserted to obtain imaged images in chronological order, using an external medical device to image the endoscope, the insertion instrument, and the lumen, and providing the images acquired by the endoscope to an inference model obtained by learning based on an image of the difficult-to-insert portion obtained by image-capturing an image of the insertion instrument at a portion where insertion into the lumen is difficult and feature information about a group of intermediate images acquired in chronological order between acquisition of the image of the difficult-to-insert portion and acquisition of an image of successful insertion obtained by image-capturing an insertion instrument successfully inserted into the lumen, outputting the features of the group of intermediate images corresponding to the difficult-to-insert portion as inference results from the inference model, and displaying the images acquired by the endoscope, the inference results, and images acquired by the external medical device.
[0014] An insertion guide method according to one aspect of the present invention is a guiding method for providing guidance using images obtained from an imaging unit when inserting an insertion instrument into a lumen, the method comprising the steps of: acquiring captured images obtained in time series by starting imaging before the insertion of the insertion instrument into the lumen; determining images of a difficult-to-insert portion from the captured images, the difficult-to-insert portion being an image of the difficult-to-insert portion of the insertion instrument; and providing images obtained from the imaging unit when the insertion instrument is inserted into the lumen to an inference model trained with training data obtained by annotating features of a group of intermediate images obtained from the time series of captured images obtained in each of a plurality of different cases between the imaging timing of the image of the difficult-to-insert portion and the determination of the success or failure of the insertion into the lumen, and displaying the images based on guide information obtained from the inference model.
[0015] In one aspect of the present invention, a method for training an inference model uses training data that uses images obtained from an imaging unit when an insertion instrument is inserted into a lumen for each of a plurality of cases, and uses a group of training data in which guide information created based on characteristic information for each case obtained from subsequent images taken during insertion for images of difficult-to-insert parts selected for each case is annotated, so that the image obtained from the imaging unit when the insertion instrument is inserted into the lumen is used as input and guide information is used as output.
[0016] A learning model according to one aspect of the present invention is configured by a neural network in which weighting coefficients are learned using a group of captured images obtained in time series by capturing images during insertion of an insertion instrument into a lumen and annotation information in which an insertion difficulty status is assigned to each image included in the group of captured images, and causes a computer to function in such a way that calculations based on the learned weighting coefficients are performed on the group of captured images in time series input to an input layer of the neural network, and guide information indicating an insertion difficulty status is output from an output layer of the neural network.
[0017] According to the present invention, by performing learning using endoscopic images obtained when endoscope insertion is difficult, it is possible to construct an inference model that effectively supports the insertion of an insertion instrument.
[0018] FIG. 1 is a block diagram showing a data processing device according to a first embodiment of the present invention. FIG. 2 is an explanatory diagram showing an endoscope system that uses an inference model constructed based on training data obtained by the data processing device of FIG. 1. FIG. 3 is an explanatory diagram for explaining learning for constructing an inference model. FIG. 4 is a flowchart for explaining the creation of training data. FIG. 5 is a block diagram showing the configuration of the endoscope system 30 of FIG. 2. FIG. 6 is a flowchart for explaining the operation of the endoscope system of FIG. 5. FIG. 7 is an explanatory diagram showing an example of a display on the display screen 33a of the monitor 33. FIG. 8 is an explanatory diagram showing a second embodiment. FIG. 9 is a flowchart showing the operation of an endoscope system. FIG. 10 is a flowchart showing the operation of an external medical device such as an X-ray diagnostic device 70 that cooperates with the endoscope system. FIG. 11 is an explanatory diagram showing an example of a display resulting from cooperation with an external medical device.
[0019] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0020] First Embodiment FIG. 1 is a block diagram illustrating a data processing device according to a first embodiment of the present invention. FIG. 2 is an explanatory diagram illustrating an endoscopic system that utilizes an inference model constructed based on training data obtained by the data processing device of FIG. 1. A body lumen inspected with an endoscope may be constricted by a sphincter or the like, bent three-dimensionally, clogged with metabolites, or have a luminal wall protruding due to a lesion. For this reason, the state of the lumen ahead in the insertion direction along its length (the direction of the connecting holes) may not always be clearly visible from an endoscopic image acquired by an imaging unit configured with an imaging device provided at the tip of the insertion section of the endoscope. Even when a lumen is entered and continues, there may be difficult-to-insert portions along the way, as described above, or when the lumen enters from outside the body, enters another lumen from there, or branches off.
[0021] Therefore, the endoscope operator may need to proceed with insertion while searching for a path for the endoscope insertion section while viewing an image in which the cavity ahead in the insertion direction is not necessarily visible (a dead end that appears to be a difficult-to-insert area). Furthermore, the condition of the lumen varies from subject to subject. Even for the same subject, the condition of the lumen may vary depending on changes in the position of the examination site, changes over time, and the current situation, making insertion of the endoscope insertion section difficult. Similar problems arise not only when inserting the endoscope insertion section into a lumen, but also when inserting a catheter or other device. In the following explanation, the term "insertion section" will be used to refer to the endoscope insertion section and tubular medical instruments such as catheters, guidewires, drainage drains, lithotripsy baskets, snares, balloon dilators, and other medical instruments inserted into a lumen (insertion medical instruments), such as catheters.
[0022] Considering differences in testing equipment and usage methods, in many cases, smooth insertion of the insertion part into the lumen requires medical professionals to gain sufficient experience. Therefore, when a beginner performs an examination, support from an experienced physician is also required. To address these issues, the use of AI (artificial intelligence) to assist in the insertion of the insertion part can be considered as a more efficient technique transfer and an examination method that even beginners can easily perform. However, building an inference model for such AI has not been easy until now.
[0023] Therefore, in this embodiment, by utilizing image information acquired during the process from the start of insertion of the endoscope into the body until successful insertion, it is possible to construct an inference model that supports the insertion of the endoscope into difficult-to-insert areas, such as areas where the cavity in the insertion direction cannot be confirmed using endoscopic images. Note that in this embodiment, an inference model that more effectively supports the insertion of the endoscope into difficult-to-insert areas may also be constructed by utilizing image information acquired when insertion cannot be considered successful, i.e., when insertion into the lumen is difficult and insertion is not possible. Then, in this embodiment, the constructed inference model enables insertion assistance for the insertion instrument.
[0024] 1, medical images such as endoscopic images acquired by an endoscope (not shown) are input to a data processing device 1. The endoscope has an insertion section that is inserted into a body cavity, and an imaging device is provided at the tip of the insertion section. This imaging device includes an imaging element such as a CCD or CMOS sensor, and photoelectrically converts an optical image from a subject to obtain an imaging signal.
[0025] There are also endoscopes that have an ultrasonic transmitter and receiver near the tip and use ultrasonic images for observation and diagnosis, but the present application can be applied to any endoscope that can obtain images without using an imaging element. In this case, the images are generated by imaging the reflection of ultrasonic signals. Furthermore, images obtained by MRI, X-ray CT, etc. can also be used, not limited to endoscopes.
[0026] During an endoscopic examination, a doctor operates an endoscope to insert the insertion section of the endoscope into the human body. The tip of the insertion section is provided with a bending section, and the doctor operates a bending knob or the like provided on the control section of the endoscope to bend the bending section or move the insertion section forward or backward, thereby inserting the tip of the insertion section to the site to be observed. Endoscopic images are acquired during the process from inserting the insertion section to removing it.
[0027] Images of the inside of the body acquired by such an endoscopic examination or the like are input to the data processing device 1. Note that medical images other than images acquired by an ultrasonic endoscope or endoscopic images may also be input to the data processing device 1. While FIG. 1 shows an example of a moving image as the image (input image) input to the data processing device 1, not only a series of continuously acquired images but also still images may be input. Furthermore, images from an external medical device 2 that observes the insertion of the insertion portion are also input to the data processing device 1.
[0028] 1 shows an example in which a plurality of images P1A, P2A, ... (hereinafter referred to as image PA when these images are not distinguished) are input from medical institution A, a plurality of images P1B, P2B, ... (hereinafter referred to as image PB when these images are not distinguished) are input from medical institution B, and an image PC is input from external medical device 2. For example, images PA and PB are captured images (endoscopic images) acquired in time series by imaging the inside of a lumen. Furthermore, various imaging devices such as an ultrasound endoscope or an X-ray device are possible examples of external medical device 2, and image PC is, for example, an X-ray image.
[0029] Each of the images PA to PC may include accompanying information added to the image in addition to the image (moving or still image) portion. The accompanying information may include operation information, sensor information, etc. For example, various switches are provided on the operation section of the endoscope (not shown), and operation information such as the operation of these switches may be added to the image data of the endoscopic image as accompanying information and input. These images PA to PC are supplied as input images to the information processing section 10 of the data processing device 1.
[0030] The data processing device 1 includes an information processing unit 10 and a recording unit 20. The information processing unit 10 may be configured by a processor using a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), etc. The information processing unit 10 may operate according to a program stored in a memory (not shown) to control each unit, or may realize some or all of its functions using hardware electronic circuits.
[0031] To create training data necessary for training an inference model that effectively assists insertion through difficult-to-insert sections, the difficult-to-insert section image determination unit 11 of the information processing unit 10 determines whether each frame of the input image is an image of a difficult-to-insert section, for example, by image analysis of the input image. An image of a difficult-to-insert section may be an image captured at a location where insertion of the endoscope insertion section begins to become relatively difficult, i.e., at the entrance of a difficult-to-insert section where insertion of the insertion section is difficult. For example, the difficult-to-insert section image determination unit 11 may determine an image of a blocked lumen as an image of a difficult-to-insert section. Furthermore, for example, the difficult-to-insert section image determination unit 11 may determine an image that does not include an image portion of a cavity in the insertion direction of the insertion section as an image of a difficult-to-insert section. For example, when the insertion section advances through the lumen, the image portion of the lumen deep inside (deep in the lumen length direction) where illumination light from the endoscope tip does not reach has a low-brightness lumen cross-sectional shape (often approximately circular). As the insertion section advances through the lumen, this image portion is located approximately at the center of the endoscopic image, and continuous images are obtained in which the lumen wall pattern moves toward the periphery of the image. The difficult-to-insert portion image determination unit 11 can determine the state in which the insertion portion is advancing through the lumen by analyzing the input image, and conversely, can determine the state in which the insertion portion is not advancing through the lumen due to difficulty in insertion, such as when there is no cavity in front of the insertion portion. The difficult-to-insert portion image determination unit 11 determines an image that does not include an image portion of such a cavity as an image of a difficult-to-insert portion. Note that the difficult-to-insert portion image determination unit 11 may determine an image obtained by imaging the inside of the lumen of the lumen as an image of a difficult-to-insert portion, or may determine an image obtained by imaging the lumen from outside the lumen as an image of a difficult-to-insert portion.
[0032] Furthermore, in a difficult-to-insert portion, the movement of the imaging device at the tip of the insertion portion is restricted, causing the insertion portion to remain at the entrance of the difficult-to-insert portion for a predetermined period of time. Therefore, the difficult-to-insert portion image determination unit 11 may determine an image as an image of a difficult-to-insert portion when an image that does not include an image portion of a cavity is detected for a predetermined period of time or more. Furthermore, in a difficult-to-insert portion, the insertion portion may be moved back and forth or its bending direction may be frequently changed for insertion. Therefore, the difficult-to-insert portion image determination unit 11 may determine that an input image is an image of a difficult-to-insert portion when it detects that the number of times the insertion portion is moved back and forth or its bending direction is changed exceeds a predetermined value.
[0033] Furthermore, an image of the state in which the insertion instrument to be inserted contacts the living body, such as the affected part, may be used as the image of the difficult-to-insert part, or an image of the living body, such as the affected part, before the insertion instrument contacts may be used as the image of the difficult-to-insert part, or an image that indicates that insertion has actually been performed may be used to determine the image of the difficult-to-insert part. An image that includes the insertion position may also be used as the image of the difficult-to-insert part, and whether or not it is an image of the difficult-to-insert part may be determined using an inference model that has been trained as training data by annotating image frames during the examination as images of the difficult-to-insert part, or the image pattern determination technology described above.
[0034] In addition to using the information from the image itself, the image of the difficult-to-insert portion may be determined in cooperation with an external medical device. For example, the difficulty of insertion may be inferred from data such as the shape of the insertion portion of the endoscope determined by X-rays or a UPD (a magnetic shape determination device), the murmurings of the doctor or assistant, body changes (electromyography data, camera images of the surgeon's posture), gaze, impatience (heart rate, sweating, and brain waves), or the patient's discomfort (groans, heart rate, sweating, movement, and brain waves), and the image corresponding to the timing at which insertion is inferred to be difficult may be used as the image of the difficult-to-insert portion.
[0035] The accompanying information added to the input image is provided to the accompanying information reflecting unit 16. The accompanying information reflecting unit 16 extracts accompanying information from the input image and reflects the extracted accompanying information in each determination of the information processing unit 10. For example, when a surgeon inserting the insertion portion into a lumen determines that the insertion portion has reached a difficult-to-insert portion, the surgeon may operate a predetermined switch on the operation unit, and this operation information may be added to the image as accompanying information. In this case, the accompanying information reflecting unit 16 extracts information from the accompanying information added to the image and provides it to the difficult-to-insert portion image determining unit 11. As a result, the difficult-to-insert portion image determining unit 11 may determine that an input image corresponding to the timing of a difficult-to-insert portion determined by the surgeon himself / herself is an image of a difficult-to-insert portion.
[0036] The success timing acquisition unit 12 acquires the timing at which the insertion instrument passes through the difficult-to-insert area. For example, the success timing acquisition unit 12 determines whether each frame of the input image is a successful insertion image through image analysis of the input image, and acquires the timing at which the successful insertion image is obtained as the successful insertion timing. Note that the success timing acquisition unit 12 does not necessarily need to acquire a successful insertion image; it is sufficient to know the timing at which the successful insertion image is obtained. Furthermore, the success timing acquisition unit 12 may be configured to detect incomplete insertion and acquire timing information thereof. Information on the timing at which the successful insertion image is obtained and information on the timing at which incomplete insertion is detected may be referred to as insertion success / failure information indicating the timing at which the insertion success / failure is determined.
[0037] A successful insertion image is an image obtained by the imaging device at the tip of the insertion portion when the insertion of the insertion portion becomes easy and it becomes possible to insert the insertion portion into the lumen, or when the tip of the insertion portion reaches the observation target site. For example, the successful timing acquisition unit 12 may determine an image as a successful insertion image when an image portion of the cavity is included in the direction of insertion of the insertion portion. Furthermore, the successful timing acquisition unit 12 may determine an input image as a successful insertion image when a series of images are obtained in which an image portion of the cavity is located approximately in the center of the endoscopic image and the pattern of the lumen wall moves toward the periphery of the image. Furthermore, the successful timing acquisition unit 12 may determine an input image as a successful insertion image when it detects that the pattern of the observation target site has begun to be included in the input image.
[0038] For example, when the surgeon inserting the insertion portion into the lumen determines that the insertion of the insertion portion has become easy, the surgeon may operate a predetermined switch on the operation unit, and this operation information may be added to the image as accompanying information. In this case, the accompanying information reflecting unit 16 extracts the information of the accompanying information added to the image and provides it to the success timing acquiring unit 12. As a result, the success timing acquiring unit 12 may determine that the input image corresponding to the timing at which the surgeon himself / herself determined that insertion was easy is the successful insertion image.
[0039] Furthermore, cooperation with an external medical device such as the X-ray device described above may assist in determining situations such as insertion difficulty, a feature of the present application. For example, whether insertion was successful or the timing of success can often be determined by an external medical device from a bird's-eye view of the lumen and treatment tool position. Therefore, the success timing acquisition unit 12 may acquire such success timing from the external medical device rather than acquiring a successful insertion image. The technology of the present application can also be applied by the success timing acquisition unit 12 requesting information such as an image from the external medical device and treating the information acquired in response as a successful insertion image. In this case, the success timing acquisition unit 12 of the data processing device 1 may determine the success timing based on the acquired successful insertion image. Furthermore, the success timing may be acquired by having another device equipped with the functionality of the success timing acquisition unit 12 and cooperating with it.
[0040] The intermediate image group determination unit 13 determines an intermediate image group, which is a group of images captured between the capture timing of the image of the difficult-to-insert portion and the successful insertion timing. For example, the intermediate image group determination unit 13 may determine an intermediate image group including each of the input image frames from the image (frame) following the image of the difficult-to-insert portion to the image (frame) immediately preceding the image of successful insertion (hereinafter referred to as intermediate images). In other words, these chronologically consecutive intermediate images record the surgeon's consideration of how to insert the difficult-to-insert portion and how to determine the insertion position and the angle at which the endoscope insertion portion is placed against the difficult-to-insert portion. That is, the characteristics of the intermediate image group include characteristics corresponding to the surgeon's insertion operation (hereinafter referred to as insertion characteristics). Furthermore, insertion characteristics may be obtained using information other than images, or a means for monitoring the movement of the operation portion at that time may be provided and the monitoring results may be used as a reference. Alternatively, insertion characteristics may be obtained by analyzing X-ray images, etc., to determine whether the surgeon is moving in the same direction or whether other operations are being performed when passing through the difficult-to-insert portion. Furthermore, this series of chronologically consecutive intermediate images may also include a situation in which the endoscope insertion section finally passes through the difficult-to-insert section and begins to advance within the lumen. The insertion characteristics of this advancement state can also be determined from the manner in which the control section is moved, monitor results, X-ray images, etc. The intermediate image group determination unit 13 may use such auxiliary information, and does not necessarily need to strictly detect the image (frame) next to the image of the difficult-to-insert section or the image (frame) immediately before the image of successful insertion among the frames of the input images.
[0041] The feature determination unit 14 determines the features of the intermediate image group and sets the determined features as insertion features. For example, the feature determination unit 14 analyzes the intermediate images to obtain information representing insertion features, such as the number of frames in the intermediate image group, the imaging period for the intermediate image group, and various information such as the bending direction of the insertion portion and the movement amount of the insertion portion obtained by analyzing the intermediate image group. In other words, it is sufficient to determine the characteristics of the difficult-to-insert portion from the images of the difficult-to-insert portion in each frame of the input images, and the subsequent images (intermediate image group) are mainly images that record the process of the surgeon's efforts to successfully insert the catheter based on the appearance of the difficult-to-insert image, so changes over time are recorded and the information provides information that can determine the process of first doing this and then doing that.
[0042] The feature determination unit 14 may classify and use the insertion features determined based on the intermediate image group. For example, the number of frames in the intermediate image group can be considered to correspond to the time from when the insertion part reaches the difficult-insertion part image to when the successful insertion image is obtained, i.e., the difficulty of insertion. The fewer the number of frames, the easier the insertion, and the greater the number of frames, the more difficult the insertion. Therefore, the feature determination unit 14 may classify the number of frames in the intermediate image group into, for example, three classes: "easy insertion" when the number of frames is less than a first threshold, "difficult insertion" when the number of frames is greater than a second threshold, and "normal insertion" when the number of frames is greater than or equal to the first threshold and less than or equal to the second threshold. Even such classification alone can help the surgeon prepare mentally.
[0043] For example, the surgeon inserting the insertion section into the lumen may specify his / her own judgment of the difficulty of inserting the insertion section by operating the operation section, and this operation information may be added to the image as accompanying information. In this case, the accompanying information reflecting unit 16 extracts the information of the accompanying information added to the image and provides it to the feature determining unit 14. In this way, the feature determining unit 14 may acquire the result of the surgeon's own judgment of the difficulty of insertion as the insertion feature.
[0044] It is also possible that insertion is not successful and results in an incomplete insertion. In this case, the group of images taken from the acquisition of the image of the difficult-to-insert portion until a predetermined time has elapsed or the group of images taken until it is determined that insertion has been abandoned may be treated as an intermediate group of images. Although such an intermediate group of images cannot determine the ingenuity used to achieve success, it can be used to determine whether the insertion was difficult.
[0045] For example, by using the insertion features of the intermediate images in the case of incomplete insertion, it is possible to provide guidance such as "In that case, you will need to try again." Whether successful or not, "insertion success / failure information" indicating the timing of the insertion success / failure judgment can be obtained, and the contents can be analyzed to distinguish between successful insertion and incomplete insertion.
[0046] In other words, even cases of failed insertion can provide extremely useful support information. For example, cases that do not result in a successful insertion image (goal) are defined as "failed insertion cases," and the image of the difficult insertion area (start) at that time is annotated as the most difficult insertion case with the highest level of difficulty. By using an inference model trained using such training data, when an image of such a difficult insertion area is obtained clinically, it is possible to display that it is the most difficult case.
[0047] Furthermore, for insertion failure cases, the image of the difficult-to-insert portion (start) at that time and the group of intermediate images may be compared with the trained inference model, and the following annotations (1) to (3) may be applied to classify the case. (1) If the image of the difficult-to-insert portion closely matches one of the example images of the difficult-to-insert portion in the inference model, and one image in the group of intermediate images closely matches the intermediate image that is the solution in the inference model, the case is classified as being of the highest difficulty level based on features other than the image, and additional learning is performed. (2) If the image of the difficult-to-insert portion closely matches one of the example images of the difficult-to-insert portion in the inference model, and none of the images in the group of intermediate images matches the intermediate image that is the solution in the inference model, the case is determined to be one in which the correct answer could not be reached, and the data is discarded. (3) If the image of the difficult-to-insert portion does not match the example image of the difficult-to-insert portion in the inference model, the image is a new image of the difficult-to-insert portion with the highest difficulty level, and learning is performed with an annotation indicating this.
[0048] While the above description describes an example of acquiring one image of a difficult-to-insert portion, the image of the difficult-to-insert portion does not need to be limited to one frame. For example, a group of captured images obtained in chronological order by capturing images of the insertion of an insertion instrument into a lumen can be obtained, and images of the insertion instrument being inserted into the lumen can be selected from the captured image group. The selected images (which are also included in the difficult-to-insert images) can then be annotated to indicate a difficult-to-insert state. By saving the captured image group and the annotation data, an organized dataset can be created in which images taken during insertion are annotated. Learning with such a dataset can produce inference results even during insertion. Of course, the difficult state is a concept that includes the ease of insertion. That is, the inference model trained with such a data set is composed of a neural network in which weighting coefficients are trained using a group of captured images acquired in chronological order during the insertion of the insertion instrument into the lumen and annotation information in which an insertion difficulty status is assigned to each image included in the group of captured images. This learning model (inference model) causes a computer to function by performing calculations based on the trained weighting coefficients on the group of captured images in chronological order input to the input layer of the neural network and outputting guide information indicating an insertion difficulty status from the output layer of the neural network. As a result, during the insertion of the insertion instrument into the lumen, the group of captured images acquired in chronological order up to the present are input to the learning model, and an inference result of the learning model is obtained. Based on this inference result, when the insertion instrument is being inserted into the lumen, a guide indicating an insertion difficulty status can be output during the insertion procedure.
[0049] The associating unit 15 can organize information on the difficult-to-insert portion images determined by the difficult-to-insert portion image determining unit 11, the successful insertion images determined by the success timing obtaining unit 12, the intermediate image groups determined by the intermediate image group determining unit 13, and the insertion features determined by the feature determining unit 14 for the large number of images input to the data processing device 1, and record the information on the difficult-to-insert portion images determined by the difficult-to-insert portion image determining unit 11 and the insertion features determined by the feature determining unit 14 corresponding to the difficult-to-insert portion images, and record the information on the difficult-to-insert portion images for each insertion feature or for each class of insertion features in the recording unit 20. The associating unit 15 may also record the intermediate image groups and successful-insertion images in the recording unit 20 for creating and updating an inference model.
[0050] When acquiring intermediate image groups, the imaging unit or light source unit located at the tip of the endoscope insertion section may encounter the entrance to a lumen or a portion that appears to be a dead end, resulting in poor image quality. Even in this case, because the intermediate image groups are a time-series image group, it is possible to determine how the endoscope tip is being moved based on changes in the reflection and leakage of light from the light source in the images in the intermediate image groups. Furthermore, if good image quality is obtained as intermediate images, it is possible to determine the angle and direction of the insertion instrument from the intermediate image groups. In other words, the intermediate image groups are considered to contain enough information to obtain insertion characteristic information. As mentioned above, an increase in the number of frames in the intermediate image group indicates difficulty in insertion, and the number of frames also serves as insertion characteristic information. Furthermore, rapid changes in the images in the intermediate image groups can indicate a fast insertion speed. Since prolonged procedures in internal lumens can be painful, appropriate treatment speed must also be considered. Therefore, information such as image changes in the intermediate image groups is useful. Of course, some of the frames constituting the intermediate image group may be unnecessary, such as an image of the moment when insertion is about to be successful, and it is not necessary to use all of the intermediate image group.
[0051] The recording unit 20 includes a recording medium (not shown) and multiple areas for storing images provided by the associating unit 15. For example, the recording unit 20 may include multiple recording areas for each insertion feature or each insertion feature class. The example in FIG. 1 shows an example in which insertion features are classified by class, with a first insertion feature, a second insertion feature, etc. indicating the respective insertion feature classes. The recording unit 20 includes multiple recording areas R1, R2, etc. for each insertion feature class (hereinafter, referred to as recording area R when there is no need to distinguish between recording areas R1, R2, etc.). Information indicating the insertion feature class, multiple difficult-insertion portion images classified into the class, and information on the insertion features corresponding to each difficult-insertion portion image are recorded in association with each other in the recording area R. For example, if information on the number of frames in the intermediate image group is recorded as the insertion feature corresponding to each difficult-insertion portion image, the first to third insertion features may be classes of easy insertion, normal insertion, and difficult insertion, respectively. The difficult-to-insert portion image and the annotation information are used as annotation information, and training data are created from the difficult-to-insert portion image and the annotation information. In other words, in this case, the recording unit 20 records training data for constructing an inference model that inputs the difficult-to-insert portion image and outputs an inference result based on the features of the intermediate image group, i.e., the insertion feature corresponding to the surgeon's insertion operation. As described above, the recording unit 20 may also store information on the difficult-to-insert portion image, the intermediate image group, the successful insertion image, and the insertion feature in an organized manner.
[0052] In the above description, the insertion characteristics are classified into easy insertion, normal insertion, and difficult insertion, but the insertion characteristics may specifically indicate the state of insertion. For example, the insertion characteristics may be information collected in time series, such as the position adjustment of the insertion unit (imaging unit) when it moves toward the target part based on the insertion operation, the insertion angle, speed information, bending direction, turning direction, the number of frames advanced, and the number of frames removed.
[0053] While specific information regarding insertion can be determined from images, the acquisition of such specific information may be assisted by the output of an external device. For example, when inserting a bile duct through the papilla, it is conceivable that the direction of duct extension can be determined by inferring the extent of the bile duct and pancreatic duct behind the papilla using images obtained by capturing the papilla surface. However, it would be preferable to be able to provide specific assistance on the actual insertion angle of the insertion instrument to match the orientation of the duct. While such insertion angle can be determined from images, X-ray images from an external medical device can be used to observe the procedure from a bird's-eye view and obtain information on the angle and position of the insertion instrument. Furthermore, if a dedicated sensor is built into or attached to the insertion instrument, such as a gravity sensor, it can obtain tilt information and provide guidance to experienced users, such as a recommended angle. Annotation of the insertion angle may also be provided from the learning stage. This would enable guidance regarding the difference between the current insertion angle and the inferred result when inserting the bile duct through the papilla. The present application also contemplates providing an actuator or other device in the insertion section or a drive member driven by the actuator to automatically control the position and angle. In other words, the part written as a guide contemplates not only transmission to the operator but also transmission to machines such as robots.
[0054] In the above explanation, the insertion features obtained from the intermediate image group may be classified as easy or difficult, for example, using information on the period (number of frames) from the acquisition of the image of the difficult-to-insert area to the acquisition of the image of the successful insertion. However, the same difficult-to-insert area may be classified differently depending on the surgeon's skill, such as being judged as easy by an experienced surgeon but difficult by an inexperienced surgeon. Therefore, the classification of the insertion features may be corrected depending on the technician's class.
[0055] The classification of surgeons may reflect their skill level or the amount of experience they have with procedures. Insertion characteristics may be a numerical indicator, such as the probability of successful insertion for a doctor with average skill, or the indicator may be changed depending on the level and skill of the target doctor. In this case, an input unit may be provided for inputting doctor information such as the number of years of experience and the number of medical treatments, diagnoses, and examinations the doctor has performed, and the information may be weighted and quantified based on the skill level, stratified for control (inference, judgment, display switching, etc.), or the results of learning using a group of examination images obtained by the operations of a doctor with average experience (a doctor with average skill) may be prioritized.
[0056] In this way, the characteristics of the group of intermediate images obtained between the image of the difficult-to-insert area used for the annotation in this application and the determination of the success or failure of the insertion into the lumen can be expressed as characteristics obtained by classifying the skill of the person who performed the insertion procedure of the insertion device.
[0057] If it were possible to acquire, record, or input information about the operator's skill and experience as accompanying information in Figure 1 (even if it is not recorded, it may be possible to infer it from intermediate images, etc.), an inexperienced doctor could be able to guide the procedures of an experienced doctor with improved skills.
[0058] Images acquired by a less skilled surgeon, such as when there are too many frames in the intermediate image group, may not be treated as successfully inserted images, and in this case, an annotation "not successful" may be added.
[0059] The records in the recording unit 20 may be organized so that the training data is changed for each skill and learning is performed, such as "difficult for an inexperienced operator" or "easy for an experienced operator." By performing such learning and using the inference results obtained by the inference model, an inexperienced operator may be guided to learn the techniques of an experienced operator or may be guided to seek assistance from an experienced operator. Furthermore, if an operation is difficult for an inexperienced operator but relatively easy for an experienced operator, the procedure (insertion position and angle) of the inexperienced operator may be monitored and compared with the operation (insertion position and angle) of the experienced operator. The difference between the insertion position and angle may be displayed on an image, for example, by indicating the difference between the current position and the insertion point of the experienced doctor using text, graphics including arrows, audio, or the like, so that the inexperienced operator can understand the difference.
[0060] In other words, for each image of a similar difficult-to-insert area, depending on the physician's skill, we determine whether insertion was easy or difficult based on the number of intermediate images and the interval between the difficult times. Using the information from the intermediate images, we can then use the training data obtained by annotating the difficult-to-insert image to determine the insertion position and angle relative to the target area when successful. This allows us to build an inference model that guides the physician to insert the catheter at a specific position and angle. Since the technique of an inexperienced physician is also captured as an image, guidance is possible when beginning insertion, such as suggesting a different position instead of that, or a different angle instead of that. Since the image position is annotated, the position when successful can also be displayed, and differences can be converted into words or displayed with an arrow. Of course, classification by physician skill is not essential, as some may consider it acceptable to take longer if successful.
[0061] In other words, the present application is also an invention of a guiding method for guiding an insertion instrument using images obtained from an imaging unit when inserting the insertion instrument into the lumen, in which imaging is started before the insertion into the lumen to acquire captured images obtained in chronological order, and a difficult-to-insert portion image determination is performed to determine an image of the difficult-to-insert portion among the captured images. Images obtained during insertion of a plurality of separate cases are prepared, and features of a group of intermediate images obtained in advance from the image of the difficult-to-insert portion in the chronologically acquired images obtained between the image of the difficult-to-insert portion in the captured images and successful insertion into the lumen are annotated for the image of the difficult-to-insert portion. The image of the difficult-to-insert portion obtained prior to issuing a guide during the current insertion procedure is input into an inference model trained with the obtained training data, and guide information obtained is displayed. Furthermore, the features of the group of intermediate images obtained between the image of the difficult-to-insert portion used for the annotation and the image of the successful insertion into the lumen are also information indicating the position of the insertion instrument relative to the object into which the insertion instrument is inserted. Furthermore, the characteristics of the group of intermediate images obtained between the image of the difficult-to-insert area used for the annotation and the successful insertion into the lumen are also characteristics obtained by classifying the multiple cases according to the similarity of the images of the difficult-to-insert area and determining the ease of insertion for each similar image of the difficult-to-insert area.
[0062] FIG. 3 is an explanatory diagram for explaining learning for constructing an inference model.
[0063] FIG. 3 shows that the series of images acquired in cases a to c are classified by the information processing unit 10 into images of difficult-to-insert areas, intermediate images, and images of successful insertion. Cases a to c are not necessarily images of the same area; the images of difficult-to-insert areas may be different from one another, and the intermediate images and successful-insertion images may also be different from one another. In the example of FIG. 3, the intermediate images of cases a to c are shown as four, three, and seven frames, respectively. The feature determination unit 14 sets the first threshold to four frames and the second threshold to six frames, and classifies case b, which has fewer frames than the first threshold, into three categories: "easy insertion," case c, which has more frames than the second threshold, into "difficult insertion," and case a, which has a number of frames equal to or greater than the first threshold but less than the second threshold, into "normal insertion."
[0064] The associating unit 15 records, for example, images of difficult insertion in case a and information on insertion characteristics in recording area R1. The associating unit 15 also records, for example, images of difficult insertion in cases b and c and information on insertion characteristics in recording areas R2 and R3, respectively. In this case, images of difficult insertion areas (teacher data) annotated with an indication of normal insertion are recorded in recording area R1 of the recording unit 20. Similarly, images of difficult insertion areas (teacher data) annotated with an indication of easy insertion and difficult insertion are recorded in recording areas R2 and R3 of the recording unit 20. Similar processing is performed on a large number of images, and a huge amount of teacher data is recorded in the recording unit 20.
[0065] The training data recorded in the recording unit 20 is supplied to the network. FIG. 3 shows that training data is provided to the network N to perform training. Taking training data-based training as an example, the network design is determined so that network N can obtain an output corresponding to each input by training using a large amount of training data, such as deep learning. An inference model is constructed by each network N. Note that the inference model may be trained using training data without training data, or the difficult-to-insert part images, intermediate image groups, and successful insertion images organized and recorded in the recording unit 20 may be provided to the network N as is, and an inference model may be constructed by deep learning of what image changes result in success.
[0066] In Figure 3, as the simplest example, the results of classifying the image features of difficult-to-insert parts are annotated. Note that the image features correspond to the appearance of the difficult-to-insert parts, and the difficult-to-insert parts can be classified according to their appearance. If more detailed information is annotated, even more diverse inferences can be made.
[0067] In cases where a smaller number of intermediate images indicates easier insertion and a larger number of intermediate images indicates more difficult insertion, if images of areas that are difficult to insert are judged from endoscopic images obtained from examinations performed by doctors with various skills, even in cases with similar characteristics of areas that are difficult to insert, images of areas that are difficult to insert obtained by doctors with few years of experience and few cases of experience may be judged to be difficult to insert, while images of areas that are difficult to insert obtained by doctors with many years of experience and many cases of experience may be judged to be easy to insert.
[0068] In cases like this where there are differences in ease of insertion between similar cases, or where there are differences depending on the situation even between the same doctor, the results of comparing cases where insertion was difficult and cases where it was easy can be important information for annotation.
[0069] In other words, by comparing images of the insertion process when it was difficult with images of the insertion process when it was easy, and annotating and learning the insertion position and angle of the insertion part for each case, or by having the position and angle of the insertion tool when insertion was performed in a short time be designated as "best" and having the system compete for superiority, it becomes possible to infer a guide for the best insertion position and angle. The insertion position and angle of the insertion part can be determined from changes in the image during insertion. If the image sensor is capable of detecting and outputting distance distribution, this distance distribution information can be used to determine the insertion position and angle of the insertion part. Furthermore, it is possible to determine whether the insertion part is inserted perpendicular to the object or at an angle by looking at changes in the image. For example, when the magnifying glass is brought closer to the paper perpendicularly to the surface than when it is brought closer at an angle, image distortion is different, and the image change is greater at the edge of the screen when the magnifying glass is closer at an angle, making image determination easy.
[0070] In other words, the method for learning the inference model is a method for learning the inference model using training data that uses images obtained from the imaging unit when the insertion instrument is inserted into the lumen for each of a plurality of cases, and the inference model is trained to use images obtained from the imaging unit when the insertion instrument is inserted into the lumen as input and guide information as output using a group of training data in which images of difficult-to-insert parts selected for each of the above cases are annotated with guide information created based on characteristic information for each of the above cases obtained from images taken during subsequent insertion.
[0071] By annotating guide information created based on characteristic information for each case, it becomes possible to adjust the position of the insertion unit (the imaging unit or the insertion instrument observed by the imaging unit) as it moves toward the target area, as well as guide the insertion angle. Furthermore, learning can be performed that reflects information collected in chronological order on insertion speed information, bending direction, rotation direction, number of frames advanced, number of frames removed, etc., and guidance based on this information can be made possible.
[0072] Deep learning is a multilayered version of the machine learning process using neural networks. A typical example is a forward propagation neural network, which sends information from front to back and makes a judgment. In its simplest form, it requires three layers: an input layer consisting of m1 neurons, a hidden layer consisting of m2 neurons determined by parameters, and an output layer consisting of m3 neurons corresponding to the number of classes to be discriminated. The neurons in the input and hidden layers, and those in the hidden and output layers, are connected by connection weights, and a bias value is added between the hidden and output layers, making it easy to form logic gates. While three layers are sufficient for simple discrimination, increasing the number of hidden layers makes it possible to learn how to combine multiple features during the machine learning process. In recent years, neural networks with 9 to 152 layers have become practical due to their training time, judgment accuracy, and energy consumption.
[0073] The network N used for machine learning may be any of a variety of well-known networks. For example, R-CNN (Regions with CNN features) or FCN (Fully Convolutional Networks) using CNN (Convolution Neural Network) may be used. This involves a process called "convolution" that compresses image features, operates with minimal processing, and is strong in pattern recognition. Furthermore, a "recurrent neural network" (fully connected recurrent neural network) that can handle more complex information and allows information analysis whose meaning changes depending on the order or sequence of information may be used, allowing information to flow bidirectionally.
[0074] To realize these technologies, conventional general-purpose arithmetic processing circuits such as CPUs and FPGAs can be used, but because much of the processing in neural networks involves matrix multiplication, GPUs and Tensor Processing Units (TPUs), which are specialized for matrix calculations, may also be used.In recent years, such dedicated artificial intelligence (AI) hardware, called "neural network processing units (NPUs)," have been designed to be integrated and embeddable with CPUs and other circuits, and may even become part of the processing circuit.
[0075] In the example of Figure 3, an endoscopic image is provided to the inference model constructed in this way. When an image of a difficult-to-insert portion included in the endoscopic image is input to the inference model, the inference model outputs information indicating normal, easy, or difficult insertion. For example, the example of Figure 3 shows that the inference result can be displayed as "easy," indicating easy insertion.
[0076] Next, the operation of the data processing device configured as above will be described with reference to Fig. 4. Fig. 4 is a flow chart for explaining the creation of teacher data.
[0077] 4, the information processing unit 10 of the data processing device 1 accesses a group of endoscopic images such as images PA, PB, etc. The information processing unit 10 selects and captures one of the target images (S2). The difficult-to-insert portion image determining unit 11 and the success timing acquiring unit 12 of the information processing unit 10 determine whether the captured image is an image of a difficult-to-insert portion or an image of successful insertion (S3).
[0078] Here, we assume an image of the lumen into which the insertion instrument will be inserted, but difficulty of insertion may also be determined based on an image of the lumen interior, which may have a narrowing or bend after the insertion instrument has entered the lumen, or an image of the lumen entrance, where the entrance is blocked before the insertion instrument enters the lumen, i.e., an image taken from outside the lumen.
[0079] When the difficult-to-insert portion image determination unit 11 and the success timing acquisition unit 12 detect the difficult-to-insert portion image and the successful insertion image (YES in S3), the intermediate image group determination unit 13 acquires the intermediate image group (S4). Also, when there is accompanying information corresponding to the input image, the accompanying information reflection unit 16 acquires the accompanying information, and the intermediate image group determination unit 13 acquires the intermediate image group based on the accompanying information.
[0080] If the difficult-to-insert portion image determination unit 11 and the success timing acquisition unit 12 cannot detect the difficult-to-insert portion image and the successful insertion image (NO in S3), the information processing unit 10 requests accompanying information corresponding to the input image. If accompanying information corresponding to the input image exists, the accompanying information reflection unit 16 acquires the accompanying information, and the difficult-to-insert portion image determination unit 11 and the success timing acquisition unit 12 acquire the difficult-to-insert portion image and the successful insertion image based on the accompanying information. If the difficult-to-insert portion image and the successful insertion image can be detected based on the accompanying information (YES in S6), the process proceeds to S4 to acquire an intermediate image group, and if they cannot be detected (NO in S6), the process proceeds to S12.
[0081] In S7, the feature determination unit 14 determines insertion features from the intermediate image group and accompanying information. In S8, it is determined whether insertion feature information can be acquired. If insertion feature information is acquired (YES in S8), the associating unit 15 associates the insertion feature information with the difficult-to-insert portion image, groups the images by feature information, records them in the recording unit 20, and then proceeds to S12. In this way, the recording unit 20 records training data in which insertion features are annotated for the difficult-to-insert portion image.
[0082] If it is determined in S8 that the insertion feature information cannot be acquired, determination information indicating that the insertion feature cannot be acquired is added to the difficult-to-insert portion image as metadata. In S12, it is determined whether or not to select another image. If it is selected (YES in S12), the process returns to S2 to select the next image. If it is not selected (NO in S12), the process ends.
[0083] An inference model for inferring insertion characteristics in difficult-to-insert portions is constructed by learning using a large amount of training data recorded in the recording unit 20. This inference model is employed as an inference model 41 in Fig. 2, which will be described later.
[0084] In this embodiment, images of difficult-to-insert portions and images of successful insertion are determined from the input images, a group of intermediate images between these images is obtained, the insertion characteristics of the intermediate images are determined, and the images of difficult-to-insert portions and the insertion characteristics are classified and recorded by the insertion characteristics, thereby creating training data for constructing an inference model that infers the insertion characteristics when an image of a difficult-to-insert portion is input. This simplifies the creation of training data for constructing an inference model that supports insertion into difficult-to-insert portions.
[0085] Next, with reference to FIGS. 2, 5 and 6, an endoscope system that uses an inference model constructed based on training data obtained by the data processing device of FIG. 1 will be described.
[0086] The endoscope system 30 includes an endoscope 31, a processor 40, and a monitor 33. The endoscope 31 has an elongated and flexible insertion section 35 that is inserted into a body cavity of a patient P, who is a subject, an operation section 36 that is connected to the base end of the insertion section 35 and is provided with various operation devices, and a cable 37 that extends from the operation section 36 for connection to the processor 40. Fig. 2 shows a state in which the insertion section 35 is inserted into the large intestine from the anus of the patient P who is lying on an examination bed 38.
[0087] The insertion section 35 of the endoscope 31 emits illumination light from a light source (not shown) at its tip. The illumination light is irradiated onto the subject from the tip of the insertion section 35, and return light from the subject is received by an imaging device disposed at the tip of the insertion section 35. The imaging device is driven and controlled by a processor 40, converts an optical image of the subject into an image signal, and outputs the image signal to the processor 40. The processor 40 has an image signal processing section (not shown), which receives the image signal from the imaging device, processes the signal, and outputs the processed endoscopic image to the monitor 33. In this way, the endoscopic image is displayed on the screen of the monitor 33.
[0088] A bending portion (not shown) is provided at the tip of the insertion portion 35, and this bending portion is driven to bend by a bending knob 36a provided on the operation portion 36. The surgeon pushes the insertion portion 35 into the body cavity and inserts it while bending the bending portion by operating the bending knob 36a.
[0089] Fig. 5 is a block diagram showing the configuration of the endoscopic system 30 of Fig. 2. Note that Fig. 5 only shows the configuration related to the inference processing, but the endoscopic system 30 also includes, in addition to the image signal processing unit described above, an operation input unit that accepts user operations and various communication functions for communicating with the outside. The endoscopic system 30 of Fig. 5 functions as an inspection device that uses an inference model constructed by learning teacher data created by the data processing device 1 in an endoscopic examination, but can also be used as an image generation device that generates input images for the data processing device 1 of Fig. 1, and can also be used as the teacher data creation device of Fig. 1 and the inference model generation device of Fig. 3.
[0090] 5, an imaging device 35a is provided in the insertion section 35 of the endoscope 31. The imaging device 35a captures images of the inside of the lumen into which the insertion section 35 is inserted, acquires captured images in time series, and supplies the acquired captured images to the processor 40. Operation information of the operation section 36 of the endoscope 31 is also supplied to the processor 40. The operation information of the operation section 36 may be added to the image from the imaging device 35a as accompanying information and supplied to the processor 40.
[0091] The processor 40 includes a control unit 45, an inference model 41, a display control unit 42, and an insertion-difficult portion image determination unit 43. The control unit 45 and each unit of the processor 40 may be configured by a processor using a CPU, GPU, FPGA, etc., and may operate according to a program stored in a memory (not shown) to control each unit, or may realize some or all of its functions using a hardware electronic circuit. The control unit 45 comprehensively controls the entire processor 40. Note that, when creating training data in the processor 40, in addition to the insertion-difficult portion image determination unit 43, the processor 40 may be configured to realize the functions of each unit of the information processing unit 10. Also, for example, the control unit 45 may be configured to realize the functions of each unit of the information processing unit 10.
[0092] The inference model 41 is a model constructed based on learning of teacher data created by the data processing device 1 of FIG. 1 . The control unit 45 provides the inference model 41 with the captured images captured by the imaging device 35 a in time series. The inference model 41 is controlled by the control unit 45 to determine an insertion-difficult portion image from the input captured images and output information on insertion features corresponding to the determined insertion-difficult portion image. The insertion-difficult portion image determination unit 43 may also determine the insertion-difficult portion image. The insertion-difficult portion image determination unit 43 has the same function as the insertion-difficult portion image determination unit 11 of FIG. 1 , and determines the insertion-difficult portion image from the images captured by the imaging device 35 a. In this case, the insertion-difficult portion image determined by the insertion-difficult portion image determination unit 43 is supplied to the inference model 41, and the inference model 41 obtains information on insertion features corresponding to the insertion-difficult portion image.
[0093] The display control unit 42 provides the image from the imaging device 35a to the monitor 33, causing it to be displayed on the display screen 33a. The display control unit 42 also displays a guide (support display) M1 indicating information about insertion characteristics on the endoscopic image displayed on the display screen 33a. In the example of Fig. 5, the endoscopic image is displayed on the display screen 33a, and the support display M1 saying "This is a difficult case" is also displayed.
[0094] The support display M1 allows the surgeon inserting the endoscope to recognize that the insertion of the insertion section 35 will cause the tip of the insertion section 35 to reach a difficult-to-insert section, and that the insertion from this difficult-to-insert section will become easier until an image of successful insertion is obtained, and that the insertion will be relatively difficult. This allows the surgeon to take appropriate action, such as requesting assistance from an experienced surgeon, if necessary.
[0095] Next, the operation of the endoscope system configured as above will be described with reference to Fig. 6. Fig. 6 is a flowchart for explaining the operation of the endoscope system of Fig. 5.
[0096] 6, basic information is input. For example, basic information such as the patient's name, sex, age, medical condition, doctor's name, and medical instrument name is input. Next, imaging by the endoscope 31 is started, and the input endoscopic image is subjected to predetermined image processing by the processor 40 and then displayed on the display screen 33a of the monitor 33 (S22).
[0097] The endoscopic image acquired by the endoscope 31 is also provided to the difficult-to-insert portion image determination unit 43. The difficult-to-insert portion image determination unit 43 determines the difficult-to-insert portion image from the endoscopic image. In S23, it is determined whether or not the difficult-to-insert portion image has been determined. If the difficult-to-insert portion image has not been determined (NO in S23), the process proceeds to S32. If the difficult-to-insert portion image has been determined, the difficult-to-insert portion image is input to the inference model 41 (S24). In addition, the difficult-to-insert portion image may be provided to a recording device (not shown) so that the difficult-to-insert portion image can be recorded for creating training data.
[0098] The inference model 41 receives an image of a difficult-to-insert portion and outputs an inference result. This inference result indicates insertion characteristics corresponding to the image of a difficult-to-insert portion. The inference model 41 may be configured to output an inference result of the case type obtained from the image of a difficult-to-insert portion if inference is possible. When an inference result of the type is obtained, the display control unit 42 displays a type detection display indicating the inference result of the case type on the endoscopic image being displayed on the display screen 33a of the monitor 33. Also, FIG. 6 shows an example in which the image of a difficult-to-insert portion is determined by the image determination unit 43 of a difficult-to-insert portion and then the determined image of a difficult-to-insert portion is provided to the inference model 41. However, the input image may be provided directly to the inference model 41, and the inference model 41 may determine the image of a difficult-to-insert portion and infer its insertion characteristics.
[0099] 7 is an explanatory diagram showing an example of a display on the display screen 33a of the monitor 33. The upper part of FIG. 7 shows an example of a type detection display. An endoscopic image P1 is displayed in the center of the display screen 33a, and a type detection display is displayed as a support display M2 at the bottom of the display screen 33a. In practice, the support display M2 displays, for example, text indicating the case type.
[0100] In S25 of FIG. 6 , the display control unit 42 converts the output of the inference model 41 into display information. In S26, it is determined whether the converted display information is text information. If the display control unit 42 outputs the display information as text information (YES in S26), the display control unit 42 proceeds to S27, where it controls the display of the difficult-to-insert portion image by combining the text information with the difficult-to-insert portion image. If the display control unit 42 outputs the display information as information other than text information, for example, image information such as an arrow (NO in S26), the display control unit 42 proceeds to S28, where it controls the display of the difficult-to-insert portion image by combining the image information such as an arrow with the difficult-to-insert portion image.
[0101] The middle section of Fig. 7 shows a display example in this case. An endoscopic image P2, which is an image of the difficult-to-insert portion, is displayed in the center of the display screen 33a. An example is shown in which a message "Bend the tip more upward" based on text information is displayed as a support message M3 in the lower part of the display screen 33a. An example is shown in which an arrow image based on image information is displayed as a support message M4 in the center of the display screen 33a. Note that while Fig. 6 shows control to display either a message based on text information or a message based on information other than text information, as in the example in the middle section of Fig. 7, both a message based on text information and a message based on information other than text information may be displayed.
[0102] In addition, the above explanation describes an example in which the inference results of the inference model 41 are displayed by the display control unit 42, but the processor 40 may also be configured to notify an external device such as a smartphone of the inference results by utilizing a communication function (not shown) that the endoscopic system 30 has.
[0103] For example, in FIG. 7 , an insertion method is acquired as an inference result of the inference model 41, and the insertion method is displayed as a support display. However, the inference result of the inference model 41 may also be obtained as an inference result indicating the difficulty of insertion in a difficult-to-insert section, such as “easy,” “normal,” or “difficult.” In this case, the processor 40 may display information based on guide information indicating the difficulty of insertion, and may further perform processing according to the difficulty of insertion. For example, if an inference result indicating that insertion is “difficult” is obtained, the processor 40 may transmit information for “calling” to, for example, a smartphone. For example, if the doctor inserting the endoscope 31 is a relatively inexperienced doctor, “request for help” information may be transmitted to the smartphone of a veteran doctor. If the doctor inserting the endoscope 31 is an experienced doctor, “call for junior doctors” information may be transmitted to the smartphone of a junior doctor to provide insertion training. Furthermore, if an inference result indicating that insertion is “difficult” is obtained, the input image may be recorded as a video with captions for educational purposes in a recording device (not shown).
[0104] In S29 of Fig. 6, it is determined whether the insertion section 35 has passed through the difficult-to-insert section. For example, the control section 45 can determine whether the insertion section 35 has passed through the difficult-to-insert section by including the success timing acquisition section 12 of Fig. 1. Furthermore, the control section 45 can also determine whether the insertion section 35 has passed through the difficult-to-insert section by operating the operation section 36 by a doctor or the like by including the accompanying information reflection section 16 of Fig. 1. If it is determined that the insertion section 35 has passed through the difficult-to-insert section (YES in S29), the display control section 42 performs control to synthesize and display text information indicating that the insertion section 35 is currently passing through the lumen (S30).
[0105] 7 shows an example of a display in this case. An endoscopic image P3 obtained during insertion into the lumen is displayed on the display screen 33a, and a support display M3 saying "inserting" is displayed at the bottom of the display screen 33a, indicating that the endoscope is passing through the lumen.
[0106] In S31 of Fig. 6, the control unit 45 records a series of images from the image of the difficult-to-insert portion to the image of successful insertion or the entire input image as evidence data using a recording device (not shown). Next, in S32, it is determined whether the examination has ended. If an instruction to end the examination has not been issued, the process returns to S22; if an instruction to end the examination has been issued, the process ends. Note that if passage of the difficult-to-insert portion has not been confirmed in S29, the process returns from S32 to S22, and S22 to S28 are repeated. Therefore, a support display M3 is displayed to support the doctor's insertion operation according to the image of the difficult-to-insert portion acquired at that time.
[0107] As described above, in this embodiment, an inference model is used that outputs the insertion characteristics of the difficult-to-insert portion for the image of the difficult-to-insert portion, thereby enabling effective insertion assistance.
[0108] Although the above description has primarily focused on displaying the inference results, the inference results may also be presented to the user by voice. Furthermore, in recent years, automatic insertion endoscopes capable of automatically inserting an endoscope into a lumen have been developed. Therefore, in addition to presenting the inference results, such automatic insertion endoscopes may also be controlled based on the inference results. For example, smooth insertion into difficult-to-insert sections can be achieved by automatically adjusting the angle of the bending portion of the insertion section or the insertion depth of the insertion section based on the inference results.
[0109] Second Embodiment Fig. 8 is an explanatory diagram showing a second embodiment. This embodiment enables the use of not only endoscopic images but also images acquired by an external medical device, such as X-ray images, as training data used to create an inference model that presents insertion characteristics for difficult-to-insert sections. This embodiment also provides a medical system that effectively supports medical procedures by linking an endoscopic system with an external medical device. In this embodiment, a medical system that inserts (cannulates) an insertion section, such as a cannula tube or a guidewire, into a lumen, such as the bile duct or pancreatic duct, will be described as an example.
[0110] The medical system 50 in Fig. 8 includes an endoscope 60, a processor 80, and a monitor 90, which constitute an endoscopic system, and an X-ray diagnostic device 70. The endoscope 60 has an elongated and flexible insertion section 61 that is inserted into a body cavity of a patient P, who is a subject, an operation section 62 that is connected to the proximal end of the insertion section 61 and is provided with various operation devices, and a cable 63 that extends from the operation section 62 for connection to the processor 80. Note that, in the case of assuming cannulation of the bile duct, pancreatic duct, etc., the endoscope 60 is a side-viewing endoscope such as a duodenoscope, and the imaging device provided in the insertion section 61 is positioned so as to be able to capture images in a rearward oblique direction. Fig. 8 shows a state in which the insertion section 61 is inserted into the body through the mouth of the patient P, who is lying on an examination tabletop 81.
[0111] The insertion section 61 emits illumination light from a light source (not shown) at its tip. The illumination light is irradiated onto the subject from the tip of the insertion section 61, and return light from the subject is received by an imaging device disposed at the tip of the insertion section 61. The imaging device is driven and controlled by the processor 80, converts an optical image of the subject into an image signal, and outputs the image signal to the processor 80. The processor 80 receives the image signal from the imaging device, performs predetermined signal processing, and outputs the processed endoscopic image to the monitor 90. In this way, the endoscopic image is displayed on the screen of the monitor 90.
[0112] The X-ray diagnostic apparatus 70 includes an X-ray detector unit 71, a tabletop 81, a gantry 91, slide rails 92, a slider 93, a C-arm 94, and a control circuit (not shown). The tabletop 81 is supported by a support base (not shown) and is rotatable by a predetermined angle around an axis (hereinafter referred to as the longitudinal axis) extending in the longitudinal direction and passing through the center of the tabletop 81 in the short direction, as well as around an axis (hereinafter referred to as the short axis) extending in the short direction and passing through the center of the tabletop 81 in the long direction. This makes it possible to change the surface of the tabletop 81 on which a patient P lies from a horizontal state to a state inclined by a predetermined angle in any direction. A pair of X-ray detector units 71 are arranged above and below the surface of the tabletop 81, so as to sandwich the patient P when the patient P lies on the tabletop 81.
[0113] A gantry 91 is fixed near the tabletop 81. A slide rail 92 is attached to the side of the gantry 91. A slider 93 is slidably supported on the slide rail 92. An arm holder (not shown) is attached to the slider 93, and a C-arm 94 is slidably attached to the arm holder. An X-ray detector unit 71 is attached to the tip of the C-arm 94. The slider 93 slides along the longitudinal direction of the slide rail 92, thereby allowing the C-arm 94 to slide parallel to the longitudinal direction of the tabletop 81. The C-arm 94 has a semicircular arc shape, and is supported by the arm holder so that its tip moves along an arc-shaped path. The plane formed by the arc-shaped path of the tip of the C-arm 94 is perpendicular to the longitudinal direction of the tabletop 81. In other words, the X-ray detector unit 71 attached to the tip of the C-arm 94 is movable around the tabletop 81 along the arc-shaped path.
[0114] Therefore, the X-ray detector unit 71 is movable along the longitudinal direction of the tabletop 81 and can rotate around the tabletop 81 along a plane parallel to the plane including the short axis. The X-ray diagnostic apparatus 70 is connected to the processor 80 by a signal cable (not shown), and the tabletop 81, slider 93, and C-arm 94 of the X-ray diagnostic apparatus 70 are driven and controlled by the processor 80. That is, by controlling the drive of the tabletop 81, slider 93, and C-arm 94 by the processor 80, the patient P on the tabletop 81 can be oriented in any position and direction relative to the X-ray detector unit 71, and the X-ray detector unit 71 can take an X-ray image of any part of the patient P from any direction.
[0115] The X-ray detector unit 71 of the X-ray diagnostic apparatus 70 includes an irradiation section and a detection section, and is controlled by the processor 80 to irradiate X-rays onto the patient P and acquire an X-ray image of the patient P. The X-ray diagnostic apparatus 70 outputs the X-ray image acquired by the X-ray detector unit 71 to the processor 80.
[0116] The processor 80 may have the same configuration as the processor 40 in Fig. 5. That is, the processor 80 may include the control unit 45, the inference model 41, the display control unit 42, and the difficult-to-insert portion image determination unit 43 in Fig. 5, and may further have a configuration including the information processing unit 10 and the recording unit 20 in Fig. 1. Also, for example, the control unit 45 may be configured to realize the functions of each unit of the information processing unit 10. In this case, the difficult-to-insert portion image determination unit 43 in Fig. 5 can be omitted.
[0117] The processor 80 is equipped with a communication device (not shown) and is capable of transmitting and receiving data between the endoscope 60 and the X-ray diagnostic apparatus 70. This allows the processor 80 to display not only endoscopic images but also X-ray images from the X-ray detector unit 71 on the display screen 90a of the monitor 90. The processor 80 can also control the drive of the endoscope 60 and the X-ray diagnostic apparatus 70 in cooperation with each other. If the processor 80 has the functions of the information processing unit 10 and the recording unit 20 in FIG. 1 , it can create training data based on at least one of the endoscopic images acquired by the endoscope 60 and the X-ray images acquired by the X-ray detector unit 71.
[0118] That is, in this embodiment, the difficult-to-insert area image determination unit 11 of the processor 80 determines an image of a difficult-to-insert area for at least one of the endoscopic images and the X-ray images, the success timing acquisition unit 12 determines an image of a successful insertion for at least one of the endoscopic images and the X-ray images, the intermediate image group determination unit 13 determines an intermediate image group for at least one of the endoscopic images and the X-ray images, the feature determination unit 14 determines insertion features for at least one of the intermediate image groups for the endoscopic images and the X-ray images, and the association unit 15 associates the image of a difficult-to-insert area for at least one of the endoscopic images and the X-ray images with the insertion features and records them in the recording unit 20.
[0119] In this way, the recording unit 20 records, for each insertion feature or insertion feature class, multiple images of difficult-insertion areas, each of which is an endoscopic image or an X-ray image classified into each class, associated with information on the corresponding insertion features. Information on the insertion feature or insertion feature class corresponding to the image of the difficult-insertion area is used as annotation information, and training data is created from the image of the difficult-insertion area and the annotation information. That is, the recording unit 20 records training data for constructing an inference model that inputs the image of the difficult-insertion area and outputs information corresponding to the features of the intermediate image group, i.e., the insertion feature corresponding to the surgeon's insertion operation.
[0120] An inference model is constructed by learning such training data as shown in Figure 3. Such an inference model is implemented in the processor 80 as the inference model 41. Note that the functions of the information processing unit 10 and the recording unit 20 may not be built into the processor 80, but may be provided as an external data processing device that can operate in cooperation with the processor 80.
[0121] Next, the operation of the embodiment configured as above will be described with reference to Fig. 9 to Fig. 11. Fig. 9 is a flowchart showing the operation of the endoscope system, and Fig. 10 is a flowchart showing the operation of external medical devices such as the X-ray diagnostic apparatus 70 that cooperate with the endoscope system. In Fig. 9, the same steps as in Fig. 6 are assigned the same reference numerals, and their description will be omitted.
[0122] The examples of Figures 9 and 10 show a case where the functions of the information processing unit 10 and the recording unit 20 of Figure 1 are not included in the processor 80. The examples of Figures 9 and 10 also show an example in which the processor 80 operates in cooperation with an endoscopic system and an external medical device. The inference model 41 in the processor 80 is initially constructed by learning training data based only on endoscopic images. For example, in ERCP (endoscopic retrograde cholangiopancreatography), in which an insertion part such as a cannula tube or a guidewire is inserted into the papilla, the papilla corresponds to the difficult-to-insert part. The angle at which the insertion part should be positioned relative to the papilla is determined to some extent depending on the shape of the papilla, etc. Therefore, the data processing device of Figure 1 can obtain insertion features from an image of the difficult-to-insert part as an image of the difficult-to-insert part, and the inference model 41 constructed by learning using the obtained training data can provide effective insertion assistance into the papilla.
[0123] 9, basic information is input in S21, and then imaging is started by the endoscope 60. The endoscopic image input from the endoscope 60 is subjected to predetermined image processing by the processor 80 and then displayed on the display screen 90a of the monitor 90 (S22).
[0124] The endoscopic image acquired by the endoscope 31 is also provided to the inference model 41. The inference model 41 determines an image of a difficult-to-insert portion from the input image and outputs the determination result and information on its reliability. In S42, the control unit 45 determines whether the reliability is higher than a predetermined threshold. If the reliability is higher than the predetermined threshold (YES in S42), the control unit 45 causes the display control unit 42 to convert the output of the inference model 41 into display information.
[0125] In addition, in S43, the control unit 45 controls each unit in cooperation with an external medical device such as an X-ray diagnostic device 70 (S43). For example, in ERCP, for pancreatic duct intubation or bile duct intubation, it may be necessary to position an insertion section such as a guidewire near the center of the papilla opening at a right angle or at a predetermined angle. The guidewire used for intubation is inserted through the insertion section 61 of the endoscope 60 and protrudes from the tip of the insertion section 61. To position the tip of the guidewire at a desired position and in a desired direction, the insertion section 61 must be positioned and oriented appropriately. Therefore, to allow the surgeon to confirm the optimal position of the insertion section 61 inserted into the duodenum and positioned to observe the papilla, the control unit 45 performs cooperative control, such as starting imaging by the X-ray detector unit 71.
[0126] A signal for coordinated control from the control unit 45 is supplied to the X-ray diagnostic apparatus 70, and the X-ray diagnostic apparatus 70 controls the X-ray diagnostic apparatus 70 in accordance with instructions from the control unit 45 to drive the X-ray detector unit 71, slider 93, C-arm 94, etc.
[0127] In S61 of FIG. 10 , the X-ray diagnostic apparatus 70, which is an external medical device, is operated by a user or communicates with the processor 80. The X-ray diagnostic apparatus 70 determines whether an instruction to perform imaging or an instruction to adjust the imaging position have been received through operation or communication (S62). If an instruction to perform X-ray imaging has been received from the control unit 45, the X-ray detector unit 71 irradiates X-rays onto the patient P to perform imaging (imaging in S62) and outputs the imaging results to the processor 80. As a result, an X-ray image is displayed on the display screen 90a of the monitor 90 along with an endoscopic image (S63). Note that if an instruction to stop imaging is received in S62, the X-ray detector unit 71 stops imaging. If an instruction to adjust the imaging position has been received from the processor 80 (position S62), the X-ray diagnostic apparatus 70 drives the top plate 81, slider 93, C-arm 94, etc. to adjust the imaging position of the patient P by the X-ray detector unit 71 (S64). As a result, the X-ray detector unit 71 adjusts the position of the imaging range so as to image, for example, the insertion portion 61 of the endoscope 60, the guide wire which is an insertion tool, and the nipple portion. If there is no instruction in S62, or when the processes of S63 and S64 are completed, the X-ray diagnostic apparatus 70 determines whether there is an instruction to end (S65), and returns to S61 if there is no instruction to end, or terminates the process if there is an instruction to end.
[0128] Fig. 11 is an explanatory diagram showing an example of a display in cooperation with an external medical device. The upper part of Fig. 11 shows an example of a display when the insertion portion 61 has reached the duodenum. On the display screen 33a, an endoscopic image P11 is displayed on the right side, and an X-ray image P21 is displayed on the left side. A type detection display is displayed as a support display M11 below the endoscopic image P11. The X-ray image P21 includes an image PD1 of the duodenum and an image PS1 of the insertion portion 61 inserted into the duodenum. By checking the display on the display screen 90a, the surgeon can understand the state of the papilla and the position of the insertion portion 61 relative to the duodenum and papilla.
[0129] In S44 of FIG. 9 , the control unit 45 determines whether movement of the endoscope 60, treatment tools, etc. is detected. The endoscope 60 captures an image of the vicinity of the papilla, and the control unit 45 can detect movement of the insertion section 61, treatment tools, etc. through image analysis of the endoscopic image. If the control unit 45 detects movement, the process proceeds to S45. In S45, if movement of the insertion section 61 or other components is detected, the control unit 45 increases the imaging frame rate (FPS) of the X-ray detector unit 71, which is an external medical device, to facilitate observation of the insertion section 61 or other components (S45). Furthermore, if there is no movement of the insertion section 61 or other components, the control unit 45 stops imaging (imaging OFF) the X-ray detector unit 71, which is an external medical device. This minimizes radiation exposure to the patient P. After the process of S45 or if movement detection is not performed in S44, the display control unit 42 displays an assistance display based on the insertion characteristics obtained by inference of the inference model 41 in S46.
[0130] The middle section of Figure 11 shows a display example in this case. On the right side of the display screen 90a, an endoscopic image P12, which is an image of the difficult-to-insert portion, is displayed. Below the endoscopic image P12, a support message M12 saying "Bend the tip more upward" is displayed. Also, on the left side of the display screen 90a, an X-ray image P22 is displayed. The X-ray image P22 includes an image PD2 of the duodenum and an image PS2 of the insertion section 61 inserted into the duodenum.
[0131] By checking the support display M12 on the display screen 90a, the physician receives support for the insertion operation to position the insertion portion 61 relative to the papilla. Note that the X-ray image displayed on the display screen 90a also displays the insertion portion 61 and the insertion portion, such as the guide wire and cannula tube protruding from the tip of the insertion portion 61, and the physician can receive appropriate insertion support by checking these displays. That is, in this embodiment, appropriate insertion support can be obtained by cooperation between the endoscopic system and the X-ray diagnostic device 70, which is an external medical device.
[0132] In addition, when inserting the insertion part, an incision of the nipple may be required, and a support display such as "Incision Required" may be displayed based on the inference result of the inference model 41. Such a display allows the doctor to receive support regarding the procedures necessary for insertion. The control unit 45 may be capable of causing the display control unit 42 to display various displays as needed, or of giving necessary instructions to external devices.
[0133] Because the difficulty of insertion can change depending on pre-treatments such as incisions, the same training data and inference model creation as described above can also be performed for insertion after incisions. In other words, a training method for an inference model using training data using images obtained from an imaging unit when inserting an insertion instrument into a lumen from multiple cases can be performed by preparing a second training data set in which guide information created based on characteristic information for each case obtained from images of the difficult-to-insert portion after pre-treatment prepared for each case is annotated with the images (pre-pretreatment training data is considered to be the first) of the difficult-to-insert portion after pre-treatment. The training is performed by inputting the images obtained from the imaging unit when inserting the insertion instrument into the difficult-to-insert portion after pre-treatment and outputting the guide information. Furthermore, the guide information may be used for training to output guide information indicating whether or not pre-treatment is necessary.
[0134] Once the physician determines that the insertion portion 61 is positioned and oriented appropriately, the physician inserts an insertion portion, such as a guidewire, from the papilla into the pancreatic duct or bile duct. In S29 of FIG. 9 , the control unit 45 determines whether the insertion portion 61 has passed through the papilla, which is a difficult-to-insert portion. The control unit 45 can determine that the insertion portion has passed through the papilla, which is a difficult-to-insert portion, by analyzing the endoscopic image. For example, the surface of the guidewire may be patterned, and the control unit 45 can easily determine whether the guidewire has passed through the papilla based on changes in the image of the guidewire in the endoscopic image. Alternatively, the control unit 45 may determine that the guidewire has passed through the insertion duct by detecting that a post-insertion procedure, etc., has begun within a predetermined time after reaching the difficult-to-insert portion.
[0135] The control unit 45 may also be configured to perform image analysis of the X-ray image acquired by the X-ray detector unit 71. In this case, the image analysis of the X-ray image can determine whether the insertion portion has passed through the papilla, which is a difficult-to-insert region. Furthermore, by viewing both the endoscopic image and the X-ray image displayed on the display screen 90a, the physician can easily confirm whether the insertion portion has passed through the papilla. When the physician confirms the passage of the insertion portion and operates the operation unit 62, the control unit 45 has the function of the associated information reflecting unit 16 that acquires associated information based on the operation of the operation unit 62, thereby enabling the physician to determine that the insertion portion has passed through the papilla. The physician can also move the X-ray detector unit 71 and the top plate 81 of the X-ray diagnostic apparatus 70 to easily confirm the passage of the insertion portion through the papilla.
[0136] When the control unit 45 determines that the insertion unit has passed through the papilla, which is a difficult-to-insert portion (YES in S29), the control unit 45 may transmit a series of videos from the image of the difficult-to-insert portion to the image of successful insertion to an external data processing device in order to update the teacher data (S47). Next, the display control unit 42 performs control to synthesize and display text information indicating that the insertion unit is passing through the lumen (S48).
[0137] The lower part of Fig. 11 shows a display example when the insertion portion has passed through a difficult-to-insert portion. On the right side of the display screen 90a, the control unit 45 displays an endoscopic image P13 obtained while the guidewire is being inserted into the papilla. The endoscopic image P13 includes an image P14 of the guidewire inserted into the papilla. In addition, below the endoscopic image P13, a support display M13, "Inserting," is displayed to indicate that the guidewire is passing through the papilla. In addition, on the left side of the display screen 90a, the control unit 45 displays an X-ray image P23 obtained while the guidewire is being inserted into the papilla. The X-ray image P23 includes an image PD3 of the duodenum, an image PS3 of the insertion portion 61 inserted into the duodenum, and an image PW3 of the guidewire inserted into the papilla.
[0138] In addition, in S48, the control unit 45 also performs other processes besides display control. For example, the control unit 45 may perform control for contrast agent injection, control for cooperation with external medical devices, control for referring to various information, etc. For example, the control unit 45 may perform the processes of S44 and S45. Furthermore, for example, when inserting a cannula tube and injecting contrast agent, the control unit 45 can control the X-ray detector unit 71 to acquire an X-ray moving image and display it on the display screen 90a, and when determining that the injection of contrast agent is complete, the control unit 45 can control the X-ray image from the X-ray detector unit 71, which is an external medical device, to be displayed on the display screen 90a to facilitate confirmation by the doctor.
[0139] In S42, if the reliability of the input image being an image of a difficult-to-insert portion is equal to or less than a predetermined threshold, the control unit 45 proceeds to S51. For example, in the case of a first case, the current inference model 41 may produce an inference result with low reliability as an image of a difficult-to-insert portion, even if the input image is acquired in a difficult-to-insert portion. Therefore, the control unit 45 uses the function of the difficult-to-insert portion image determination unit 11 to determine whether the input image is an image of a difficult-to-insert portion (S51). If the control unit 45 determines that the input image is not an image of a difficult-to-insert portion (S51, YES), it proceeds to S32. If the control unit 45 determines that the input image is an image of a difficult-to-insert portion (S51, NO), it proceeds to S52.
[0140] The control unit 45 transmits the difficult-to-insert portion image determined in S51 to the external data processing device so that a new inference model for correctly inferring the acquired image can be created (S52). The control unit 45 cooperates with the external data processing device to make a learning request. The external data processing device detects successful insertion images related to the difficult-to-insert portion image received from the recording unit 20, determines insertion features for the intermediate image group corresponding to the detected successful insertion images, and constructs an inference model by learning using training data based on the insertion features and the received difficult-to-insert portion image.
[0141] The control unit 45 determines whether a new inference model has been created by the external data processing device as a result of the learning request (S53). If a new inference model has been created (YES in S53), the control unit 45 acquires information about the new inference model from the external data processing device, updates the inference model 41 based on the information, replaces the inference model (S54), and then returns to S41.
[0142] In addition, if the processor 80 has the functions of the information processing unit 10 and recording unit 20 of Figure 1, the control unit 45 may be configured to construct a new inference model by processing similar to that of an external data processing device and replace the inference model 41 with the constructed inference model.
[0143] If a new inference model cannot be created (NO in S53), the control unit 45 issues an instruction to the external medical device, such as an information request, in S55. As a result, the X-ray detector unit 71, which is an external medical device, starts X-ray imaging and outputs the captured X-ray image to the processor 80. The control unit 45 controls the display control unit 42 to display the X-ray image from the X-ray detector unit 71 on the display screen 90a of the monitor 90 as reference information (S56), and proceeds to S29. The doctor or other medical professional refers to the endoscopic image and X-ray image displayed on the display screen 90a of the monitor 90 and attempts to insert the insertion part into the papilla, which is a difficult-to-insert part. By referring to both the endoscopic image and the X-ray image, insertion becomes relatively easy.
[0144] As described above, an inference model can be constructed by learning training data using images of difficult-to-insert areas based on at least one of endoscopic images and X-ray images and information on the insertion features of the corresponding intermediate images. Therefore, in S55, the control unit 45 may request information on the inference model constructed based on the endoscopic images and X-ray images from an external device. In this case, the control unit 45 may acquire the inference model in S56, replace the inference model 41 with the acquired inference model, and then proceed to S41.
[0145] In this embodiment, by linking the endoscopic system with the external medical device, more effective insertion assistance is possible in difficult-to-insert areas. Furthermore, by performing learning using training data created using endoscopic images obtained by the endoscopic system and X-ray images obtained by an X-ray diagnostic device, which is an external medical device, it is possible to construct an inference model with improved inference accuracy for the insertion characteristics of the insertion part in difficult-to-insert areas.
[0146] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some of the components shown in the embodiments may be omitted. Furthermore, components from different embodiments may be appropriately combined.
[0147] Furthermore, while the above embodiments have primarily described assistance in difficult insertion situations, this is because successful insertion can be easily determined by, for example, checking the cavity beyond the difficult-to-insert portion, making it easy to explain assistance in inserting an insertion instrument into a lumen (through a hole). Furthermore, while an example of determining a successful insertion image is shown, this is because it is relatively easy to determine what was done between the start of insertion and the successful state. However, the present invention is not limited to assistance in difficult insertion situations. Even during removal or other access procedures, if there is information on the situation before the procedure and information that can determine whether the procedure was successful, difficulty information can be obtained to support difficult situations. The method described herein can be applied and generalized as a technology for dealing with difficult endoscopic operations. As mentioned above, the present invention can be applied to endoscopes that have an ultrasound transmitter and receiver near the tip and use ultrasound images for observation and diagnosis, as long as the endoscope can acquire images. In other words, the imaging unit may acquire optical images or ultrasound images.
[0148] Furthermore, among the technologies described herein, many of the controls and functions, mainly those described in the flowcharts, can be set by a program, and the above-described controls and functions can be realized by a computer reading and executing the program. The program can be recorded or stored, in whole or in part, as a computer program product on a portable medium such as a flexible disk, CD-ROM, or nonvolatile memory, or on a storage medium such as a hard disk or volatile memory, and can be distributed or provided at the time of product shipment, via a portable medium, or via a communication line. A user can easily realize the data processing device, data processing method, data processing program, medical system, display control method, insertion guide method, inference model learning method, and learning model of the present embodiments by downloading the program via a communication network and installing it on a computer, or by installing it on a computer from a recording medium.
Claims
1. A data processing device comprising: a difficult-to-insert portion image determination unit that receives imaged images obtained in time series by imaging a lumen portion into which an insertion instrument is inserted, and determines whether the received imaged images are difficult-to-insert portion images obtained at a difficult-to-insert portion of the insertion instrument; a success timing acquisition unit that acquires, as success timing, the timing at which a successful insertion image obtained when the insertion instrument is successfully inserted into the lumen; a feature determination unit that receives the received imaged images, and determines features of an intermediate image group that is a group of images captured between the imaging timing of the difficult-to-insert portion image and the successful timing; and an association unit that associates and records the difficult-to-insert portion image with information on the features of the intermediate image group corresponding to the difficult-to-insert portion image.
2. The data processing device according to claim 1, wherein the image of the difficult-to-insert portion is a captured image of the occluded lumen.
3. The data processing device according to claim 1, wherein the information on the characteristics of the intermediate image group is information on the number of images included in the intermediate image group.
4. The data processing device according to claim 1, wherein the associating unit also records the successfully inserted image and the intermediate image group corresponding to the difficult-to-insert portion image.
5. The data processing device according to claim 1, wherein the insertion instrument is an insertion portion of an endoscope, an insertion medical instrument, or a guide wire.
6. The data processing device according to claim 1, wherein the success timing acquisition unit determines the success timing based on an image obtained by an external device.
7. An endoscopic system comprising: an inference model obtained by learning based on an image of a difficult-to-insert portion obtained by imaging the difficult-to-insert portion of an insertion instrument into a lumen, and feature information about a group of intermediate images acquired in time series from the acquisition of the image of the difficult-to-insert portion to the acquisition of an insertion success image acquired when the insertion instrument is successfully inserted into the lumen; an endoscope including an imaging device that images the lumen into which the insertion instrument is inserted and obtains images in time series; and a control unit that provides the images acquired by the endoscope to the inference model and causes the inference model to output, as inference results, the features of the group of intermediate images corresponding to the difficult-to-insert portion.
8. A medical system comprising: an inference model obtained by learning based on an image of a difficult-to-insert portion obtained by imaging the difficult-to-insert portion of an insertion instrument into a lumen, and feature information about a group of intermediate images acquired in time series from the acquisition of the image of the difficult-to-insert portion to the acquisition of an insertion success image acquired when the insertion instrument is successfully inserted into the lumen; an endoscope including an imaging device that images the lumen into which the insertion instrument is inserted to obtain images in time series; a control unit that provides the images acquired by the endoscope to the inference model and causes the inference model to output, as inference results, features of the group of intermediate images corresponding to the difficult-to-insert portion; an external medical device that images the endoscope, the insertion instrument, and the lumen; and a display control unit that displays the images acquired by the endoscope, images acquired by the external medical device, and the inference results.
9. A data processing method comprising: inputting captured images obtained in time series by imaging a lumen into which an insertion instrument is inserted, and determining whether the input captured images are difficult-to-insert portion images obtained at a difficult-to-insert portion of the insertion instrument; inputting the captured images and determining whether the input captured images are successful insertion images obtained when the insertion instrument is successfully inserted into the lumen; inputting the captured images and determining characteristics of an intermediate image group which is a group of images captured between the capture of the difficult-to-insert portion image and the capture of the successful insertion image; and recording the difficult-to-insert portion image and information on the characteristics of the intermediate image group corresponding to the difficult-to-insert portion image in association with each other.
10. A data processing program for causing a computer to execute the following procedures: inputting captured images obtained in time series by imaging a lumen into which an insertion instrument is inserted, and determining whether the input captured images are difficult-to-insert portion images obtained at a difficult-to-insert portion of the insertion instrument; inputting the captured images, and determining whether the input captured images are successful insertion images obtained when the insertion instrument is successfully inserted into the lumen; inputting the captured images, and determining characteristics of an intermediate image group that is a group of images captured between the capture of the difficult-to-insert portion image and the capture of the successful insertion image; and recording the difficult-to-insert portion image and information on the characteristics of the intermediate image group corresponding to the difficult-to-insert portion image in association with each other.
11. A display control method comprising: using an imaging device provided in an endoscope to image a lumen into which an insertion instrument is inserted, thereby obtaining imaged images in chronological order; providing the images obtained by the endoscope to an inference model obtained by learning based on images of the difficult-to-insert portion obtained by imaging the difficult-to-insert portion of the insertion instrument into the lumen, and feature information about a group of intermediate images obtained in chronological order between the acquisition of the image of the difficult-to-insert portion and the acquisition of an image of successful insertion obtained by imaging when the insertion instrument is successfully inserted into the lumen; and outputting and displaying the features of the group of intermediate images corresponding to the difficult-to-insert portion from the inference model as inference results.
12. A display control method comprising: imaging a lumen into which an insertion instrument is inserted using an imaging device provided on an endoscope to obtain captured images in chronological order; imaging the endoscope, the insertion instrument, and the lumen using an external medical device; providing the captured images obtained by the endoscope to an inference model obtained by learning based on images of the difficult-to-insert portion obtained by imaging the difficult-to-insert portion of the lumen with the insertion instrument and feature information about a group of intermediate images obtained in chronological order between the acquisition of the image of the difficult-to-insert portion and the acquisition of an image of successful insertion obtained by imaging when the insertion instrument is successfully inserted into the lumen; outputting the features of the group of intermediate images corresponding to the difficult-to-insert portion as inference results from the inference model; and displaying the captured images obtained by the endoscope, the inference results, and images obtained by imaging using the external medical device.
13. A guiding method for providing guidance using images obtained from an imaging unit when inserting an insertion instrument into a lumen, comprising: a step of acquiring captured images obtained in time series by starting imaging before the insertion of the insertion instrument into the lumen; a difficult-to-insert portion image determination step of determining, from the captured images, images of a difficult-to-insert portion obtained at a difficult-to-insert portion of the insertion instrument; and a step of providing images obtained from the imaging unit when the insertion instrument is inserted into the lumen to an inference model trained with training data obtained by annotating, for the images of a difficult-to-insert portion, features of a group of intermediate images obtained from the time-series captured images acquired in each of a plurality of separate cases, between the capture timing of the image of the difficult-to-insert portion and the determination of the success or failure of the insertion into the lumen, and displaying based on guide information obtained from the inference model.
14. The insertion guide method according to claim 13, wherein the features of the group of intermediate images are information indicating the position of the insertion instrument relative to the object into which the insertion instrument is to be inserted.
15. The insertion guide method according to claim 13, wherein the features of the group of intermediate images are obtained by classifying the multiple cases according to the similarity of the images of difficult-to-insert areas, and determining the ease of insertion for each similar image of difficult-to-insert areas.
16. Corresponding inference model claim A method for training an inference model using training data that uses images obtained from an imaging unit when an insertion instrument is inserted into a lumen in each of a plurality of cases, the method training an inference model so that the images obtained from the imaging unit when an insertion instrument is inserted into a lumen are used as input and guide information is used as output, using a group of training data in which guide information created based on feature information for each case obtained from images taken during subsequent insertion of an image of a difficult-to-insert part selected for each of the cases is annotated.
17. A method for learning an inference model as described in claim 16, wherein the guide information is used for learning to output guide information including the difficulty of insertion or the need for pre-treatment for each case.
18. A method for training an inference model using training data that uses images obtained from an imaging unit when an insertion instrument is inserted into a lumen in each of a plurality of cases, wherein the method trains the inference model so that images obtained from the imaging unit when an insertion instrument is inserted into a pre-treated difficult-to-insert part are used as input and guide information is used as output using a second group of training data in which images of the difficult-to-insert part after pre-treatment prepared for each case are annotated with guide information created based on characteristic information for each case obtained from images taken during subsequent insertion.
19. The guiding method according to claim 13, wherein the features of the group of intermediate images are features obtained by classifying the skill of a person who performed the insertion procedure of the insertion instrument.
20. A learning model for causing a computer to function as follows: the learning model is composed of a neural network in which weighting coefficients are learned using a group of captured images obtained in time series during the insertion of an insertion instrument into a lumen and annotation information in which an insertion difficulty situation is assigned to each image included in the group of captured images; the learning model performs calculations based on the learned weighting coefficients on the group of captured images in time series input to the input layer of the neural network; and outputs guide information indicating an insertion difficulty situation from the output layer of the neural network.
Citation Information
Patent Citations
Endoscopic apparatus
WO2019107226A1
Examination moving image processing device, examination moving image processing method, and examination moving image processing program
WO2019216084A1
Movement assist system, movement assist method, and movement assist program
WO2020194472A1
Endoscopic image processing system and endoscopic image processing method
WO2021054419A1
Endoscope insertion control device and endoscope insertion control method
WO2021064861A1