Medical system and method of operating the medical system
The medical system uses endoscopic images and a trained model to detect contact states between a treatment tool and tissue, addressing cost and size issues of tactile sensors, enhancing surgical precision and safety.
Patent Information
- Application Number
- JP2023206876
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-06-23
- Filing Date
- 2023-12-07
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2043-12-07
AI Technical Summary
Existing medical instruments that detect contact between a treatment tool and tissue using tactile sensors face issues of increased device costs and size, as well as limitations on device applicability.
A medical system that utilizes an endoscopic image captured by an endoscope, combined with a trained model to detect the contact state between a treatment tool and tissue, allowing for accurate contact state recognition without increasing the device size or cost, and applicable to various treatment tools.
Enables accurate and timely detection of contact states between a treatment tool and tissue, reducing unnecessary information presentation to surgeons and improving treatment safety by optimizing energy output based on contact detection.
Smart Images

Figure 0007752161000002 
Figure 0007752161000003 
Figure 0007752161000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a medical system and How the medical system works etc. [Background technology]
[0002] Patent Document 1 discloses a medical instrument that uses a tactile sensor to detect contact between a treatment tool and an object during an endoscopic procedure. This medical instrument is used together with an endoscope and includes an insertion section that is inserted into a body cavity and a tactile sensor provided in the insertion section that can detect contact of an object with the insertion section outside the observation field of the endoscope. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-61979 Summary of the Invention [Problem to be solved by the invention]
[0004] In Patent Document 1, contact between the treatment tool and tissue is detected by a tactile sensor mounted on the treatment tool, which poses the problem of increased device costs, increased size of the treatment tool, and potential limitations on the device. [Means for solving the problem]
[0005] One aspect of the present disclosure relates to a medical system that includes: an endoscopic image captured by an endoscope that captures an image including a treatment tool and tissue to be treated; a memory that stores a trained model trained using training data including the endoscopic image, the training data including the contact state between the treatment tool and the tissue in the endoscopic image; and a processor, wherein the processor acquires the endoscopic image including the treatment tool and the tissue, and uses the trained model to detect the contact state between the treatment tool and the tissue from the endoscopic image as first information.
[0006] Another aspect of the present disclosure relates to a contact state detection method including: capturing an endoscopic image including a treatment tool and tissue to be treated using an endoscope; and detecting the contact state between the treatment tool and the tissue from the endoscopic image as first information using a trained model trained with training data including the endoscopic image and the contact state between the treatment tool and the tissue in the endoscopic image.
[0007] Yet another aspect of the present disclosure relates to a medical system including: a memory that stores a trained model trained using teacher data including an endoscopic image captured by an endoscope that captures an image including a treatment tool having jaws that open and close and tissue to be treated, and the open / closed state of the jaws in the endoscopic image; and a processor, wherein the treatment tool is an energy device that treats the tissue by outputting energy from the jaws; the processor acquires the endoscopic image including the treatment tool and the tissue, performs first detection using the trained model to detect whether the jaws are open or closed from the endoscopic image, performs second detection based on electrical information in the energy output to detect the presence or absence of tissue between the jaws, and detects the gripping state of the tissue by the jaws as a contact state based on the results of the first detection and the results of the second detection.
[0008] Yet another aspect of the present disclosure relates to a contact state detection method for detecting a contact state between a treatment tool that is an energy device having jaws that open and close and that treats tissue by outputting energy from the jaws and the tissue to be treated, the contact state detection method including: capturing an endoscopic image including the treatment tool and the tissue with an endoscope; performing first detection to detect whether the jaws are open or closed from the endoscopic image using a trained model trained with teacher data including the endoscopic image and the open / closed state of the jaws in the endoscopic image; performing second detection to detect whether the tissue is present between the jaws based on electrical information in the energy output; and detecting a gripping state of the tissue by the jaws as the contact state based on a result of the first detection and a result of the second detection. [Brief explanation of the drawings]
[0009] [Figure 1] An example of the support provided by the medical system to the surgeon. [Figure 2] An example of a medical system configuration. [Figure 3] An example of the processing flow performed by a medical system. [Figure 4] 10 shows a first configuration example of a treatment tool detection unit and a contact detection unit. [Figure 5] 10 shows a second configuration example of the contact detection unit. [Figure 6] FIG. 10 is a diagram for explaining movement vectors of a treatment tool and tissue. [Figure 7] 10A and 10B are diagrams illustrating the "difference in distance between the treatment tool and the tissue in the movement direction" acquired by the preprocessing unit. [Figure 8] 10A and 10B are diagrams illustrating the "difference in distance between the treatment tool and the tissue in the movement direction" acquired by the preprocessing unit. [Figure 9] 10A and 10B are diagrams illustrating the "directional difference between the movement directions of the treatment tool and the tissue" acquired by the preprocessing unit. [Figure 10] 10A and 10B are diagrams illustrating "statistics of the distribution of the amount of movement of tissue subjected to basis transformation based on the treatment tool movement direction" acquired by the preprocessing unit. [Figure 11]10A and 10B are diagrams illustrating "statistics of the distribution of the amount of movement of tissue subjected to basis transformation based on the treatment tool movement direction" acquired by the preprocessing unit. [Figure 12] FIG. 4 is a diagram for explaining “color information” acquired by a preprocessing unit. [Figure 13] First example of setting an area of interest. [Figure 14] Second example of setting the area of interest. [Figure 15] Third example of region of interest setting. [Figure 16] Third example of region of interest setting. [Figure 17] Fourth example of region of interest setting. [Figure 18] Fourth example of region of interest setting. [Figure 19] FIG. 10 is a diagram illustrating detection of tissue divisions. [Figure 20] FIG. 10 is a diagram illustrating detection of tissue divisions. [Figure 21] 10 shows a first configuration example of a contact detection unit in the second embodiment. [Figure 22] FIG. 10 is a diagram illustrating a second configuration example of the jaw open / close detection unit. [Figure 23] 10 shows a second example of the configuration of the inter-jog structure detection unit. [Figure 24] 10 shows a third example of the configuration of the inter-joint structure detection unit. [Figure 25] 10 is a fourth example of the configuration of the inter-jog structure detection unit. [Figure 26] 10 shows a second configuration example of the contact detection unit in the second embodiment. [Figure 27] 10 shows an example of contact detection using detection results stored in a memory. [Figure 28] 10 shows a third configuration example of the contact detection unit in the second embodiment. [Figure 29] 13 shows a first configuration example of a contact detection unit in the third embodiment. [Figure 30] 13 is a first example of a flow of processing performed by a contact detection unit in the third embodiment. [Figure 31] 13 shows a second example flow of processing performed by the contact detection unit in the third embodiment. [Figure 32] 13 shows a second configuration example of the contact detection unit in the third embodiment. [Figure 33]13 shows an example of the configuration of a contact detection unit in the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] The following disclosure provides many different embodiments and examples for implementing different features of the presented subject matter. Of course, these are merely examples and are not intended to be limiting. Furthermore, the present disclosure may repeat reference numerals and / or letters in various examples. Such repetition is for the purposes of brevity and clarity and does not, in itself, require a relationship between the various embodiments and / or configurations being described. Furthermore, when a first element is described as being "connected" or "coupled" to a second element, such a description includes embodiments in which the first and second elements are directly connected or coupled to each other, as well as embodiments in which the first and second elements are indirectly connected or coupled to each other with one or more other intervening elements therebetween.
[0011] 1.Method In a procedure using an endoscope and a treatment tool, it is desirable for the medical system to be able to detect the contact state between the treatment tool and tissue and provide information to the surgeon based on the detection results. There are various situations in which such contact detection is necessary, but here we will explain an example of providing support information to the surgeon in a procedure using an energy device.
[0012] Figure 1 shows an example of support provided to a surgeon by a medical system. The medical system displays an endoscopic image 232 on a monitor 231. The medical system acquires information on the tissue to be treated from the endoscopic image 232, determines an appropriate energy output according to the information, and presents the output setting 233 to the surgeon by displaying it on the monitor 231. The surgeon approves or rejects the output setting 233 recommended by the medical system. When the output setting 233 is approved, the medical system outputs energy from the energy device at that output setting 233.
[0013] Such automatic adjustment of the energy output can improve the safety of the treatment. However, constantly presenting the recommended output setting 233 to the surgeon is troublesome for the surgeon. Therefore, it is desirable to present information when the surgeon intends to perform treatment. For example, it is desirable to present information when the treatment tool comes into contact with tissue or when a forceps-type surgical treatment tool grasps tissue. To present information at the appropriate time, it is necessary to accurately recognize the contact state between the treatment tool and tissue.
[0014] Therefore, the medical system of this embodiment detects a treatment tool from an endoscopic image and recognizes the contact state between the treatment tool and tissue based on the detection result. Alternatively, the medical system improves the accuracy of contact detection by combining multiple machine learning models or by combining a machine learning model with a non-machine learning method. As a result, support information such as information about the contacted tissue is presented to the surgeon only when the treatment tool comes into contact with the tissue, thereby preventing unnecessary information presentation to the surgeon and reducing the inconvenience.
[0015] While the treatment tool in the above-mentioned Japanese Patent Application Laid-Open Publication No. 2003-61979 is equipped with a tactile sensor, the medical system of this embodiment can detect the contact state between the treatment tool and tissue through image recognition processing. This makes it possible to recognize the contact state without increasing the size of the treatment tool while keeping device costs low. Furthermore, while the above-mentioned Japanese Patent Application Laid-Open Publication No. 2003-61979 is limited to devices equipped with a tactile sensor, the method of this embodiment does not limit the device. In other words, the method of this embodiment can be applied to various existing treatment tools.
[0016] Examples of contact states detected in this embodiment are as follows. The contact state refers to the treatment tool being in contact with tissue or not being in contact with tissue. The contact state may also refer to the treatment tool grasping tissue or not grasping tissue. When the treatment tool grasps tissue, it is considered to be in contact with the tissue. The contact state may also refer to a transition from contact to non-contact, a transition from non-contact to contact, a transition from grasping to non-grasping, or a transition from non-grasping to grasping. For example, tissue separation, which will be described later, corresponds to a transition from contact to non-contact or a transition from grasping to non-grasping. The contact state may also refer to contact being maintained, non-contact being maintained, grasping being maintained, or non-grasping being maintained. The contact state may also refer to a combination of one or more of the above states. The treatment tool includes a shaft and an end effector provided at the tip of the shaft for treating the tissue. In this case, the detected contact state is the contact state between the end effector and the tissue. However, for example, in cases where it is desirable for the shaft not to come into contact with the tissue, the contact state between the shaft and the tissue may be detected.
[0017] The period or timing for detecting the contact state may be arbitrary. For example, when the treatment tool is an energy device, the contact state may be detected when energy is not being output, when energy is being output, or both when energy is being output and when energy is not being output.
[0018] 2. First embodiment FIG. 2 shows an example of the configuration of a medical system. The medical system 1 includes an endoscope system 200, a controller 100, and a device or system 230. Below, an example will be described in which the medical system 1 is a system using a rigid endoscope for surgical operations. The treatment tool is basically inserted into the abdominal cavity or the like as a device separate from the rigid endoscope. However, the detection method of this embodiment can also be applied to a system using a flexible endoscope for the digestive tract. In that case, the treatment tool passes through a treatment tool channel of the flexible endoscope and protrudes from the tip of the flexible endoscope.
[0019] The endoscope system 200 is a system for capturing images inside a body cavity, and includes an endoscope and a main body device.
[0020] An endoscope is a rigid scope that is inserted into a body cavity to capture images of the interior of the body cavity. An endoscope includes an insertion section that is inserted into the body cavity, an operating section that is connected to the base end of the insertion section, a universal cord that is connected to the base end of the operating section, and a connector section that is connected to the base end of the universal cord. An imaging device for capturing images of the interior of the body cavity and an illumination optical system for illuminating the interior of the body cavity are provided at the tip of the insertion section. The imaging device includes an objective optical system and an imaging element that captures an image of a subject formed by the objective optical system. The connector section detachably connects the transmission cable to the main device. Images captured by an endoscope are referred to as captured images or endoscopic images.
[0021] The main device includes a processing device that controls the endoscope and performs image processing and display processing of the endoscopic image, and a light source device that generates and controls illumination light. The processing device is composed of a processor such as a CPU, processes image signals transmitted from the endoscope to generate endoscopic images, and outputs the endoscopic images to the display and controller 100. The illumination light emitted from the light source device is guided by a light guide to the illumination optical system of the endoscope, and is emitted from the illumination optical system into the body cavity.
[0022] The device or system 230 is a device or system that operates based on the output of the controller 100. The device or system 230 is, for example, a display that displays a contact state or support information based on the contact state. This display may be provided separately from the display of the endoscope system 200, or may be the display of the endoscope system 200. Alternatively, the device or system 230 may include an energy device and a generator that drives the energy device. The energy device outputs energy from its distal end using high-frequency power, ultrasound, or the like to perform treatments such as coagulation, sealing, hemostasis, incision, excision, or ablation on tissue that comes into contact with the distal end. The energy device may be a monopolar device that outputs electrical energy from a single end effector, a bipolar device that applies electrical energy between jaws, an ultrasonic device that outputs ultrasonic energy, or a combination device that uses both ultrasonic energy and electrical energy. The generator controls the energy output based on the contact state or support information based on the contact state from the controller 100.
[0023] The medical system 1 may further include a treatment tool that is a non-energy device such as forceps or a spatula. The treatment tool that is the target of contact detection may be either an energy device or a non-energy device.
[0024] The controller 100 includes a processor 110, a memory 120, an I / O device 180, and an I / O device 190. Here, an example is shown in which the controller 100 is a device separate from the main body device of the endoscope system 200. The controller 100 may be an information processing device such as a personal computer, or may be a cloud system made up of multiple information processing devices connected via a network. However, the functions of the controller 100 may also be configured to be built into the main body device of the endoscope system 200.
[0025] The I / O device 180 receives endoscopic images from the endoscope system 200. The I / O device 180 is a cable connector for connecting to the main body of the endoscope system 200, or a communication circuit for performing communication processing with the main body.
[0026] The I / O device 190 outputs contact state information or support information based on the contact state output by the processor 110 to the device or system 230. The I / O device 190 is a cable connector for connecting to the device or system 230, or a communication circuit for performing communication processing with the device or system 230.
[0027] The processor 110 includes hardware. The processor 110 is, for example, a central processing unit (CPU), a graphics processing unit (GPU), a microcomputer, or a digital signal processor (DSP). Alternatively, the processor 110 may be an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). The processor 110 may be composed of one or more of a CPU, a GPU, a microcomputer, a DSP, an ASIC, and an FPGA. The memory 120 is, for example, a semiconductor memory such as a volatile memory or a nonvolatile memory. Alternatively, the memory 120 may be a magnetic storage device such as a hard disk drive, or an optical storage device such as an optical disk drive.
[0028] The memory 120 stores a program 121 in which the processing contents of the treatment tool detection unit 130 and the contact detection unit 140 are described. The processor 110 executes the program 121 to perform the processing of the treatment tool detection unit 130 and the contact detection unit 140. For example, the program 121 includes program modules in which the processing of each unit is described, and the processor 110 executes the program modules to perform the processing of each unit.
[0029] The program 121 includes a trained model 122 obtained by machine learning. The trained model is, for example, a neural network trained by deep learning. In this case, the trained model includes a program describing the neural network algorithm and weight parameters between nodes of the neural network. The neural network includes an input layer to which input data is input, an intermediate layer that performs arithmetic processing on the data input through the input layer, and an output layer that outputs an inference result based on the arithmetic result output from the intermediate layer. In the learning stage, a learning system configured by an information processing device or a cloud system executes the machine learning process. The learning system includes a processor and a memory that stores the model and training data. The processor generates the trained model 122 by training the model using the training data. This trained model 122 is stored in the memory 120 of the controller 100. Note that the controller 100 may also function as the learning system.
[0030] The program 121 may be stored in a non-transitory information storage medium that is a computer-readable medium. The information storage medium may be, for example, an optical disk, a memory card, a hard disk drive, or a semiconductor memory. The semiconductor memory may be, for example, a ROM or a non-volatile memory. The processor 110 loads the program 121 stored in the information storage medium into the memory 120 and performs various processes based on the program 121.
[0031] FIG. 3 shows an example of a processing flow performed by the medical system. In step S1, the surgeon operates a treatment tool. The endoscope system 200 captures an endoscopic image showing the treatment tool and the treatment target. The treatment target is the tissue to be treated by the treatment tool. The processor 110 acquires the endoscopic image from the endoscope system 200.
[0032] In step S2, the treatment tool detection unit 130 detects the treatment tool from the endoscopic image. One example of the detection method is detection or segmentation using AI.
[0033] In step S3, the contact detection unit 140 detects the contact state between the treatment tool and the tissue detected in step S2 from the endoscopic image. The image used for detection may be the entire endoscopic image or a part of it. For example, the contact detection unit 140 may set a region of interest including the treatment tool and the surrounding tissue in the endoscopic image based on the detection or segmentation results, and detect the contact state from the image within that region of interest.
[0034] Fig. 4 shows a first configuration example of the treatment tool detection unit and the contact detection unit. Fig. 4 shows the processing in the learning stage and the processing in the inference stage. As described above, the processing in the learning stage is executed by the learning system.
[0035] First, the inference stage will be described. The trained model 122 includes a trained model 122a for treatment tool detection and a trained model 122b for contact detection. The endoscope system 200 captures an endoscopic image showing the treatment tool 10 and the tissue 50.
[0036] The treatment tool detection unit 130 inputs an endoscopic image to the trained model 122a. The trained model 122a detects the position or area of the treatment tool 10 from the endoscopic image and outputs the detection result 130Q. The detection result 130Q is, for example, a bounding box detected by detection or an area detected by segmentation.
[0037] The contact detection unit 140 inputs the endoscopic image and the detection result 130Q to the trained model 122b, or inputs an image of a region of interest in the endoscopic image that is set based on the detection result 130Q to the trained model 122b. The trained model 122b detects the contact state between the treatment tool 10 and the tissue 50 from the input data and outputs the detection result 140Q.
[0038] The processor 110 performs output using the detection result 140Q. The processor 110 may output the contact state between the treatment tool 10 and the tissue 50 obtained as the detection result 140Q, or may perform information output using the contact state. An example of information output using the contact state will be given below.
[0039] In a first example, the processor 110 uses the contact state detection result 140Q as a trigger for information recognition or information presentation. When contact between the treatment tool 10 and the tissue 50 is detected, the processor 110 determines the tissue type or tissue state of the contact area and presents the information to the surgeon via a monitor display or the like. This allows support information to be provided to a non-expert surgeon at an appropriate time. Alternatively, when contact between the treatment tool 10 and the tissue 50 is detected, the processor 110 determines the output setting of the energy device according to the tissue type or tissue state of the contact area and recommends the output setting to the surgeon via a monitor display or the like. This allows support information to be provided at an appropriate time rather than being displayed constantly. As described with reference to FIG. 1, the processor 110 may instruct the generator to output at the recommended setting when an approval input is received from the surgeon, and the generator may output from the energy device at the recommended setting.
[0040] In a second example, the processor 110 controls the energy device using the contact state detection result 140Q. When the processor 110 detects contact between the treatment tool 10 and the tissue 50, the processor 110 may send an instruction to release the safety lock of the energy device to the generator. That is, while non-contact between the treatment tool 10 and the tissue 50 is detected, the energy device is in a safety locked state, and therefore, energy is not output even if an output operation of the energy device is performed.
[0041] Next, the learning stage will be described. The learning system generates a trained model 122a by training a model to detect the position or area of a treatment tool from endoscopic images IMG for training. The endoscopic images IMG for training may be images captured by an endoscopic system different from the endoscopic system 200 used in the inference stage. The training data is a plurality of endoscopic images IMG and annotations indicating the position or area of a treatment tool in each image. The annotations are a bounding box indicating the position of the end effector of the treatment tool in detection, or the area occupied by the end effector of the treatment tool in segmentation.
[0042] The learning system generates a trained model 122b by training a model to detect the contact state between the treatment tool and tissue from input data. The input data is an endoscopic image IMG and an annotation indicating the position or area of the treatment tool. Alternatively, the input data is an image of a region of interest in the endoscopic image IMG that is set based on the annotation. The training data is a plurality of endoscopic images IMG, an annotation indicating the position or area of the treatment tool in each image, and a label indicating the contact state between the treatment tool and tissue in each image. The label indicating the contact state here is mainly assumed to be contact, non-contact, grasped, or non-grasped, but labels for various contact states such as those described above may also be used.
[0043] 5 shows a second configuration example of the contact detection unit. The contact detection unit 140 performs preprocessing by a preprocessing unit 141 and contact detection using the output of the preprocessing and the trained model 122b. The program that realizes the preprocessing unit 141 may be a rule-based program or a trained model trained by machine learning.
[0044] The preprocessing unit 141 block shows the information acquired in preprocessing. The information is organized hierarchically according to its nature. The information acquired in preprocessing includes the "distance difference between the movement direction of the treatment tool and the tissue," "directional difference between the movement vector of the treatment tool and the tissue," and "detection of jaw opening and closing, and detection of tissue between the jaws" shown in the lowest layer. The preprocessing unit 141 acquires one or more pieces of information and inputs them to the trained model 122b. The trained model 122b detects the contact state between the treatment tool and the tissue from the input one or more pieces of information. Among the information listed in the preprocessing unit 141 block, the information shown in the layers under "Other interpolation information" and "Contact detection accuracy improvement techniques" is used in combination with the information shown in the layers under "Movement vector" and "Grip detection feature." During the training phase of the trained model 122b, the training data consists of one or more pieces of input data and the contact state labels corresponding to the input data.
[0045] The information acquired in the pre-processing will be described in detail below, although the "detection of jaw opening and closing, and detection of tissue between the jaws" will be described later in the second embodiment.
[0046] FIG. 6 is a diagram illustrating the movement vectors of the treatment tool and the tissue.
[0047] The preprocessing unit 141 acquires motion vectors from the endoscopic image using optical flow or image registration techniques. Motion vectors may be acquired for each pixel, for each grid point, for feature points of the image, or a representative motion vector for the treatment tool and a representative motion vector for the tissue. The representative motion vector for the treatment tool is, for example, the average motion vector in the treatment tool region detected by segmentation or the like. The representative motion vector for the tissue is, for example, the average motion vector in the region surrounding the treatment tool region detected by segmentation. The average motion vector is not limited to a simple average of the motion vectors, but may also be calculated by a weighted average. Taking the average motion vector of the treatment tool as an example, the reliability of each motion vector in the motion vector distribution in the treatment tool region is calculated, and the average motion vector is calculated by weighting the motion vectors based on the reliability. The same applies to the average motion vector of the tissue.
[0048] As shown in FIG. 6, it is assumed that the movement vector Vtool of the treatment tool 10 and the movement vectors Vtissue1 and Vtissue2 of the tissue 50 are acquired. In the endoscopic image IMGa taken in a non-contact state, the surrounding tissue does not move even when the treatment tool moves, so there is little correlation between the movement of the treatment tool and the surrounding tissue. That is, in a non-contact state, there is little correlation between the movement vector Vtool of the treatment tool 10 and the movement vectors Vtissue1 and Vtissue2 of the tissue 50. On the other hand, in the endoscopic image IMGb taken in a contact state, the surrounding tissue moves in a similar manner when the treatment tool moves, so there is a strong correlation between the movement of the treatment tool and the surrounding tissue. That is, in a contact state, there is a strong correlation between the movement vector Vtool of the treatment tool 10 and the movement vectors Vtissue1 and Vtissue2 of the tissue 50. From these facts, contact detection is possible using the movement vectors.
[0049] 7 and 8 are diagrams for explaining the "difference in distance between the treatment tool and the tissue in the moving direction" acquired by the preprocessing unit.
[0050] 7, the preprocessing unit 141 calculates the average movement vector v of the treatment tool and the average movement vector w of the tissue surrounding the treatment tool, and calculates the Euclidean distance |wv| between these average movement vectors. The smaller the Euclidean distance |wv|, the greater the similarity between the movement vectors v and w, and therefore it is considered that the treatment tool and tissue are in contact. In other words, the trained model 122b increases the probability of determining that the treatment tool and tissue are in contact as the Euclidean distance |wv| decreases.
[0051] As shown in Figure 8, the preprocessing unit 141 acquires the time-series change in the Euclidean distance |wv| and inputs it to the trained model 122b. The trained model 122b determines that there is no contact when the Euclidean distance |wv| is greater than a threshold, and determines that there is contact when the Euclidean distance |wv| is less than the threshold. However, because the trained model 122b makes its judgment using a network obtained by training, the result is not a simple threshold judgment.
[0052] FIG. 9 is a diagram illustrating the "directional difference between the movement directions of the treatment tool and the tissue" acquired by the preprocessing unit.
[0053] The preprocessing unit 141 calculates the directional difference between the average movement vector v of the treatment tool and the average movement vector w of the tissue surrounding the treatment tool. An example of the directional difference is cosine similarity. If the angle formed between the average movement vector v of the treatment tool and the average movement vector w of the tissue surrounding the treatment tool is θ, the cosine similarity is cosθ. The closer the cosine similarity cosθ is to 1, the greater the similarity between the movement vectors v and w, and therefore it is considered that the treatment tool and the tissue are in contact. In other words, the closer the cosine similarity cosθ is to 1, the higher the probability of the trained model 122b determining that the treatment tool and the tissue are in contact.
[0054] The "movement vector of the treatment tool and tissue" acquired by the preprocessing unit will now be described.
[0055] The trained model 122b uses the movement vector of the treatment tool and the movement vector of the surrounding tissue as features for contact detection. The movement vector of the treatment tool and the movement vector of the surrounding tissue can themselves be useful features for contact detection.
[0056] For example, if the absolute value of the movement vector is small, the influence of noise is considered to be large, and the reliability of contact recognition is considered to be low. For example, if the treatment tool or treatment tool is actually moving very little, the influence of hand shake is relatively large, and it is considered difficult to appropriately determine the contact state from the above-mentioned Euclidean distance or cosine similarity. For this reason, if the absolute value of the movement vector of the treatment tool or the absolute value of the movement vector of the surrounding tissue is small, the trained model 122b determines that the reliability of contact recognition is low. For example, if the above absolute value is small, the trained model 122b may not output the contact state recognition result, or may output a result of "unrecognizable."
[0057] Alternatively, when a treatment tool is in contact with tissue, particularly when the treatment tool is gripping tissue, the amount of movement of the treatment tool is considered to be smaller than when the treatment tool is free to move. This is because it is considered that camera shake is reduced or positioning is being performed. For example, when the absolute value of the treatment tool's movement vector is small, the trained model 122b may determine the contact state accordingly. As an example, when the absolute value of the treatment tool's movement vector is small, the trained model 122b may maintain the previously obtained recognition result of the contact state.
[0058] 10 and 11 are diagrams for explaining "statistics of the distribution of the amount of movement of tissue subjected to basis transformation based on the treatment tool movement direction" acquired by the preprocessing unit.
[0059] 10, the preprocessing unit 141 calculates the movement direction MDtool of the treatment tool 10 and the movement vector Vtissue of the tissue 50 surrounding the treatment tool. The preprocessing unit 141 decomposes the movement vector Vtissue of the surrounding tissue 50 into a component Vp parallel to the movement direction MDtool of the treatment tool 10 and a component Vv perpendicular to the movement direction MDtool of the treatment tool 10. The movement direction MDtool of the treatment tool 10 is the longitudinal axis direction of the treatment tool 10 or the direction of the movement vector of the treatment tool 10. The trained model 122b determines the contact state using the components Vp and Vv.
[0060] FIG. 11 shows histograms of the components Vp and Vv acquired at multiple positions of the surrounding tissue in a certain frame. In the non-contact state, the histograms of the parallel component Vp and the perpendicular component Vv are both approximately symmetrical around "0". Meanwhile, in the contact state, the histogram of the perpendicular component Vv is approximately symmetrical around "0", but the histogram of the parallel component Vp is clearly asymmetrical around "0". This indicates that the contact state can be determined using the components Vp and Vv. For example, the preprocessing unit 141 calculates the kurtosis and skewness of each histogram as statistics and inputs the kurtosis and skewness to the trained model 122b. Kurtosis is an index indicating the degree of peaking of a histogram. Skewness is an index indicating the degree of asymmetry of a histogram.
[0061] 12 is a diagram illustrating the "color information" acquired by the preprocessing unit. In this example, the treatment tool is an energy device.
[0062] The trained model 122b uses color information of the tissue surrounding the treatment tool as a feature for contact detection. By using color information, changes such as white scorch or mist that occur in the tissue during energy treatment are captured, and this information is used to complement the judgment of the contact state. The occurrence of white scorch or mist indicates that energy treatment is in progress, which means that "contact" has occurred.
[0063] As shown in the left diagram of FIG. 12, a region of interest ROIc is set at the tip of the treatment tool in the endoscopic image IMGc before whitening. In the region of interest ROIc, the color of the tissue has a strong red component. In other words, the red component is sufficiently larger than the blue component, and the red component is also sufficiently larger than the green component. As shown in the right diagram of FIG. 12, a region of interest ROId is set at the tip of the treatment tool in the endoscopic image IMGd after whitening. The hatched area indicates the whitened portion. In the region of interest ROId, the tissue has turned white, so the difference between the color components is almost eliminated. In other words, the red component is almost the same as the blue and green components. In this way, whitening can be determined based on the color components. For example, the preprocessing unit 141 calculates the O1 channel in the opponent color space shown in the following equation (1) and inputs it to the trained model 122b.
[0064]
number
[0065] The "ROI setting" acquired by the preprocessing unit will now be explained. ROI stands for Region of Interest.
[0066] The preprocessing unit 141 sets a region of interest in the endoscopic image, calculates features from the image within the region of interest, and inputs the features to the trained model 122b. The region of interest includes the portion of the treatment tool where contact is desired to be detected and the surrounding tissue. The trained model 122b detects the contact state between the treatment tool and the tissue from the input features. The features are features using the movement vectors described above, color information, or grip detection features described below. The movement vectors, etc. of the treatment tool and the surrounding tissue are basically correlated around the contact point. Therefore, by setting a region of interest, the accuracy of detecting the contact state can be improved.
[0067] FIG. 13 shows a first example of setting a region of interest (ROIe). The treatment tool 10 has a shaft 11 and an end effector 14. The preprocessing unit 141 sets a region of interest (ROIe) around the end effector 14. FIG. 13 shows an example in which the region of interest (ROIe) includes the entire end effector 14 and its surrounding area. For example, the end effector 14 is detected from the endoscopic image (IMGe) by detection processing using machine learning. The preprocessing unit 141 may perform the detection processing, or if the treatment tool detection unit 130 detects the treatment tool by detection, the preprocessing unit 141 may use the detection result of the treatment tool detection unit 130. The preprocessing unit 141 uses the detected region as is as the region of interest (ROIe), or expands or reduces the detected region and uses it as the region of interest (ROIe).
[0068] FIG. 14 shows a second example of region of interest setting. As shown in the left diagram, the preprocessing unit 141 sets a region of interest ROIf around the tip of the end effector 14. For example, the region of the treatment tool 10 is detected from the endoscopic image IMGf by segmentation processing using machine learning. The preprocessing unit 141 may perform the segmentation processing, or if the treatment tool detection unit 130 detects the treatment tool by segmentation, the preprocessing unit 141 may use the detection result of the treatment tool detection unit 130. The preprocessing unit 141 detects the tip of the end effector 14 from the detected region of the treatment tool 10 and sets the region of interest ROIf based on that tip. The right diagram shows an example of tip detection. The treatment tool mask MKtool is the region of the treatment tool 10 detected by segmentation. The preprocessing unit 141 calculates the midpoint PC of the intersection line between the edge of the image and the treatment tool mask MKtool, calculates the point PF within the treatment tool mask MKtool that is farthest from the midpoint PC, and determines that point PF to be the tip of the end effector 14.
[0069] 15 and 16 show a third example of setting the region of interest.
[0070] FIG. 15 shows example images with different optimal positions of the region of interest. Endoscopic image IMGg is an example image when tissue is grasped long, and endoscopic image IMGh is an example image when tissue is grasped short. The portion of the tissue where movement is highly correlated with movement of the treatment tool is the portion grasped by the jaws, which are the end effectors. Therefore, when tissue is grasped long, the tissue is grasped by more than half of the jaws, so the surrounding region RAg is an appropriate region of interest. On the other hand, when tissue is grasped short, the tissue is grasped only near the tip of the jaws, so the region RAh near the tip of the jaw is an appropriate region of interest.
[0071] As shown in FIG. 16, the pre-processing unit 141 generates a plurality of region of interest candidates CR1 to CR9 around the tip of the treatment tool 10. For example, the pre-processing unit 141 sets a region of a predetermined size based on the tip of the treatment tool 10, and generates a plurality of region of interest candidates CR1 to CR9 by dividing the region. The number of candidates is not limited to nine. The pre-processing unit 141 sets one or more of the plurality of region of interest candidates CR1 to CR9 as the region of interest ROIp. FIG. 16 shows an example in which candidate CR2 is set as the region of interest ROIp. For example, the pre-processing unit 141 selects a region of interest candidate having a movement vector close to the movement vector of the treatment tool as the region of interest ROIp. Note that although FIG. 15 illustrates an example in which the end effector of the treatment tool is a jaw, the technique of FIG. 16 can also be applied to cases in which the end effector is not a jaw.
[0072] 17 and 18 show a fourth example of setting a region of interest, in which the end effector of the treatment tool is a jaw.
[0073] Endoscopic image IMGp is an example of an image when tissue is grasped short. The part of the jaw that is grasping the tissue is indicated by HLp. If a region of interest were set that included the entire jaw, most of the tissue included in the region of interest would not be grasped, resulting in a low correlation between the movement of the jaw and the surrounding tissue. Endoscopic image IMGq is an example of an image when tissue is grasped long. The part of the jaw that is grasping the tissue is indicated by HLq. If a region of interest were set that included only the tip of the jaw, most of the area where there is a high correlation between the movement of the jaw and the surrounding tissue would not be included in the region of interest.
[0074] Therefore, the preprocessing unit 141 detects the center of the portion of the jaw that grasps the tissue, and sets a region of interest based on that center. That is, the preprocessing unit 141 detects the centers of the portions HLp and HLq of the jaw that grasp the tissue from the endoscopic images IMGp and IMGq, and sets regions of interest ROIp and ROIq based on those centers. This sets a region of interest that includes an area where there is a high correlation between the movements of the jaw and the surrounding tissue, thereby improving the accuracy of contact detection.
[0075] FIG. 18 shows an example of a method for detecting the center of the jaw that grasps tissue. The region of the treatment tool 10 is detected from the endoscopic image IMGf by segmentation processing. The pre-processing unit 141 may perform the segmentation processing, or if the treatment tool detection unit 130 detects the treatment tool by segmentation, the pre-processing unit 141 may use the detection result of the treatment tool detection unit 130. The detected region is defined as a treatment tool mask MKtool. The pre-processing unit 141 performs dilation processing on the treatment tool mask MKtool and obtains the difference by subtracting the treatment tool mask MKtool before dilation from the mask after dilation. This results in a peripheral region mask MKperi that follows the contour of the treatment tool.
[0076] The preprocessing unit 141 performs edge extraction on the image within the peripheral region mask MKperi of the endoscopic image, and detects the portion with a large edge component as the ridge of the grasped tissue. The preprocessing unit 141 detects the tip of the treatment tool using the method described in FIG. 14 or the like. The preprocessing unit 141 sets the midpoint of the line connecting the ridge of the grasped tissue and the tip of the treatment tool as the reference point for setting the region of interest. If no ridge is detected in the edge extraction, the preprocessing unit 141 may set the region of interest using the first example of region of interest setting or the like.
[0077] This section explains the "correction of the movement vector taking into account the movement of the camera" performed by the preprocessing unit. The movement of the camera is an apparent movement amount, so it affects the movement vector of the treatment tool and the movement vector of the tissue. Therefore, canceling this can be expected to improve the accuracy of contact detection.
[0078] The preprocessing unit 141 detects the amount of camera movement and uses the amount of camera movement to cancel the influence of the camera movement from the movement vector of the treatment tool and the movement vector of the tissue. For example, the preprocessing unit 141 subtracts the movement vector of the camera from the movement vector of the treatment tool and the movement vector of the tissue. The preprocessing unit 141 calculates feature quantities such as a distance difference using the movement vector after the subtraction. Various known techniques can be used to acquire the amount of camera movement, and the following two examples are possible. In a first example, a motion sensor or position sensor such as an optical sensor, a magnetic sensor, or an inertial sensor is installed on the endoscope camera. The preprocessing unit 141 detects the amount of camera movement based on the sensor output. In a second example, the preprocessing unit 141 estimates the amount of camera movement by image processing of the endoscopic image. For example, the preprocessing unit 141 may detect a global movement vector representing the movement of the entire image and estimate the amount of camera movement based on the global movement vector.
[0079] The "scaling of movement vectors" performed by the preprocessing unit will now be explained. The magnitude of the movement vector on the image is affected by both the magnitude of the movement in real space and the distance between the camera and the subject. Therefore, by scaling using depth distance information or information such as treatment tool width, the influence of the distance between the camera and the subject can be eliminated, and the accuracy of contact detection can be expected to improve.
[0080] The preprocessing unit 141 detects the distance between the camera and the subject. For example, if the endoscope has a 3D camera, the preprocessing unit 141 detects the distance to the subject using stereoscopic vision from the 3D camera. Alternatively, the preprocessing unit 141 detects the distance to the subject from the endoscopic image using a known length such as the shaft width of the treatment tool. Taking the shaft width as an example, the preprocessing unit 141 estimates the distance to the subject from the ratio between the width of the shaft shown in the endoscopic image and the known shaft width. The preprocessing unit 141 scales the movement vector of the treatment tool and the movement vector of the tissue to a movement vector at a reference distance using the detected distance. The preprocessing unit 141 calculates feature quantities such as a distance difference using the scaled movement vectors.
[0081] The "3D conversion of movement vectors" performed by the preprocessing unit will now be explained.
[0082] The preprocessing unit 141 uses the depth information to acquire the movement vector of the treatment tool and the movement vector of the tissue as three-dimensional vectors, and calculates feature quantities such as distance difference using the three-dimensional movement vectors. The preprocessing unit 141 calculates the three-dimensional movement vector by, for example, combining two-dimensional optical flow and depth information. This method is called scene flow. For example, if the endoscope has a 3D camera, the preprocessing unit 141 detects the amount of movement in the depth direction using stereo vision from the 3D camera.
[0083] The use of three-dimensional movement vectors offers the following advantages. Although movement in real space is three-dimensional, image information is projected onto a two-dimensional plane, so movement information in the depth direction is lost. By using three-dimensional movement vectors, it becomes possible to capture movement in the normal direction of the image with high sensitivity. This improves the reliability of the movement vectors, and therefore the accuracy of contact detection.
[0084] The "time series conversion of feature quantities" performed by the preprocessing unit will now be described.
[0085] The preprocessing unit 141 inputs time-series feature quantities obtained from multiple frames arranged in time series to the trained model 122b. The trained model 122b detects the contact state between the treatment tool and tissue from the time-series feature quantities. When performing contact detection, using feature quantities from the most recent multiple frames is expected to improve the accuracy of contact detection.
[0086] 19 and 20 are diagrams illustrating tissue separation detection. In this example, the treatment tool is an energy device and the end effector is a jaw.
[0087] The value of grasping state judgment is that it can detect tissue separation from images. Detecting tissue separation is expected to be effective in preventing damage to the ultrasound device probe during surgery. Contact detection methods using movement vectors make detection easy because the feature values change suddenly in tissue separation scenes. This is explained below.
[0088] As shown in FIG. 19 , in the endoscopic image IMGr, the jaws 12 are not grasping the tissue 50. In the subsequent endoscopic image IMGs, the jaws 12 are grasping the tissue, and energy is being applied to the tissue 50 between the jaws. At this point, the tissue grasped between the jaws is connected. In the subsequent endoscopic image IMGt, the tissue 50 between the jaws is cut by energy treatment. This cutting of connected tissue by energy treatment is called tissue severing. Note that while FIG. 19 shows an example in which the treatment tool is a bipolar device, severing can also be detected by the method of this embodiment when a hook knife-type monopolar device or the like is used.
[0089] FIG. 20 shows an example of the change over time in the Euclidean distance |wv| between the movement vector of the treatment tool and the movement vector of the tissue. As shown in IMGs in FIG. 19, |wv| is small when the tissue 50 is grasped by the jaws 12. When the tissue 50 is torn apart as shown in IMGt, the correlation between the movement vectors of the jaws 12 and the tissue 50 decreases. Therefore, as shown by the dot-dash circle in FIG. 20, |wv| suddenly increases at the moment of torn apart. The contact detection unit 140 detects the torn apart of the tissue by detecting the timing at which |wv| suddenly increases.
[0090] 3. Second embodiment In the second embodiment, the end effector of the treatment tool is a jaw. Note that a description of the same parts as in the first embodiment will be omitted. In the second embodiment, for example, the configuration of the medical system 1, the configuration of the controller 100, and the processing performed by the treatment tool detection unit 130 are the same as in the first embodiment.
[0091] Fig. 21 shows a first configuration example of the contact detection unit in the second embodiment. The contact detection unit 140 includes a jaw open / close detection unit 144 and an inter-jaw structure detection unit 146. Fig. 21 shows the processing in the learning stage and the processing in the inference stage.
[0092] First, the inference stage will be described. The trained model 122 described above in Fig. 2 includes a trained model 122c for detecting jaw opening and closing and a trained model 122d for detecting inter-jaw tissue.
[0093] The jaw open / closed state detection unit 144 detects the open / closed state of the jaw shown in the endoscopic image using the endoscopic image, the detection result 130Q of the treatment tool detection unit 130, and the trained model 122c. The jaw open / closed state detection unit 144 inputs the endoscopic image and the detection result 130Q to the trained model 122c, or inputs an image of a region of interest in the endoscopic image that is set based on the detection result 130Q to the trained model 122c. The trained model 122c detects the open / closed state of the jaw from the input data and outputs the detection result 144Q.
[0094] The inter-jaw tissue detection unit 146 detects the presence or absence of tissue between the jaws shown in the endoscopic image using the endoscopic image, the detection result 130Q of the treatment tool detection unit 130, and the trained model 122d. The inter-jaw tissue detection unit 146 inputs the endoscopic image and the detection result 130Q to the trained model 122d, or inputs an image of a region of interest in the endoscopic image that is set based on the detection result 130Q to the trained model 122d. The trained model 122d detects the presence or absence of tissue between the jaws from the input data and outputs the detection result 146Q. The condition for tissue to be present between the jaws is that "tissue exists that covers one of a pair of jaws but does not cover the other jaw." Tissue covering a jaw means that part or all of the jaw is hidden by the tissue in the image.
[0095] The contact detection unit 140 determines the contact state using the detection results 144Q and 146Q. When it is determined that the jaws are closed and tissue is present between the jaws, the contact detection unit 140 determines that the jaws are grasping tissue, that is, that the treatment tool and tissue are in contact. The program that determines the contact state using the detection results 144Q and 146Q may be either a rule-based program or a trained model using machine learning. The processor 110 performs output using the detection results of the contact state. The output content is the same as that of FIG. 4.
[0096] According to this embodiment, the complex image recognition for detecting a contact state from an image is divided into a jaw open / close detection unit 144 and an inter-jaw tissue detection unit 146. This improves the recognition accuracy of each trained model compared to when a contact state is detected using a single trained model, and therefore, an improvement in the recognition accuracy of contact detection as a whole can be expected.
[0097] Next, the learning stage will be described. The learning system generates a trained model 122c by training a model to detect the open / closed state of the jaw from input data. The input data is an endoscopic image IMG and an annotation indicating the position or area of the treatment tool. Alternatively, the input data is an image of a region of interest in the endoscopic image IMG that is set based on the annotation. The training data is a plurality of endoscopic images IMG, an annotation indicating the position or area of the treatment tool in each image, and a label indicating the open / closed state of the jaw in each image.
[0098] The learning system generates a trained model 122d by training a model to detect the presence or absence of tissue between the jaws from input data. The input data is an endoscopic image IMG and annotations indicating the position or area of a treatment tool. Alternatively, the input data is an image of a region of interest in the endoscopic image IMG that is set based on the annotations. The training data is multiple endoscopic images IMG, annotations indicating the position or area of a treatment tool in each image, and labels indicating the presence or absence of tissue between the jaws in each image. The labels indicating the presence or absence of tissue between the jaws are assigned based on the above-mentioned conditions for the presence of tissue between the jaws. By performing machine learning using such training data, an image in which "tissue exists that covers one of a pair of jaws but not the other" is determined to have tissue between the jaws.
[0099] 22 is a diagram illustrating a second configuration example of the jaw open / closed state detection unit 144. In this example, the jaw open / closed state detection unit 144 determines whether the jaw is open or closed by a rule-based program using three-dimensional information.
[0100] The treatment tool detection unit 130 acquires the three-dimensional shapes of the shaft 11 and the jaw 12 from the endoscopic images. The treatment tool detection unit 130 inputs stereo images acquired by, for example, a 3D camera to the trained model 122a, and the trained model 122a recognizes the three-dimensional shapes of the shaft 11 and the jaw 12 from the input stereo images.
[0101] The jaw open / close detection unit 144 calculates the shaft axis AX from the three-dimensional shape of the shaft 11, and defines a cylinder CY of a specified radius centered on the shaft axis AX. The jaw open / close detection unit 144 determines whether the jaws 12 are contained within the cylinder CY. If the jaws 12 are contained within the cylinder CY, the jaw open / close detection unit 144 determines that the jaws are closed, and if even a portion of the jaws 12 are outside the cylinder CY, the jaws are open.
[0102] According to this embodiment, whether the jaw is open or closed is determined using a rule-based program, so the basis for determining whether the jaw is open or closed is easier to understand than when machine learning is used.
[0103] Fig. 23 shows a second example of the configuration of the inter-jaw structure detection unit. As described above, the condition for the presence of a structure between the jaws is that "a structure exists that covers one of a pair of jaws and does not cover the other jaw." In this example, the inter-jaw structure detection unit 146 determines the presence or absence of a structure covering each jaw using a trained model. Fig. 23 shows the processing in the learning stage and the processing in the inference stage.
[0104] First, the inference stage will be described. The treatment tool detection unit 130 detects the first jaw 12a and the second jaw 12b from the endoscopic image. One of the pair of jaws connected to the shaft 11 is the first jaw 12a, and the other is the second jaw 12b.
[0105] 2 includes a trained model 122d1 for detecting a first jaw covering and a trained model 122d2 for detecting a second jaw covering. The trained model 122d1 judges the first jaw 12a, and the trained model 122d2 judges the second jaw 12b.
[0106] The inter-jaw tissue detection unit 146 inputs the endoscopic image and the area of the first jaw 12a detected by the treatment tool detection unit 130 to the trained model 122d1. The trained model 122d1 detects from the input data whether the first jaw 12a is covered with tissue 55, and outputs the detection result 122d1Q.
[0107] The inter-jaw tissue detection unit 146 inputs the endoscopic image and the area of the second jaw 12b detected by the treatment tool detection unit 130 to the trained model 122d2. The trained model 122d2 detects from the input data whether the second jaw 12b is covered with tissue 55, and outputs the detection result 122d2Q.
[0108] The inter-jaw tissue detection unit 146 determines that there is tissue between the jaws when one of the detection results 122d1Q and 122d2Q is "not covered by tissue" and the other is "covered by tissue." If the detection result is anything other than this, the inter-jaw tissue detection unit 146 determines that there is no tissue between the jaws.
[0109] Next, the learning stage will be described. The learning system generates a learned model 122d1 by training a model to detect whether or not the first jaw 12a is covered by tissue 55 from input data. The input data is an endoscopic image IMG and an annotation indicating the area of the first jaw 12a. The training data is a plurality of endoscopic images IMG, an annotation indicating the area of the first jaw 12a in each image, and a coverage label indicating whether or not the first jaw 12a in each image is covered by tissue 55.
[0110] The learning system generates a trained model 122d2 by training a model to detect whether or not the second jaw 12b is covered by tissue 55 from input data. The input data is an endoscopic image IMG and an annotation indicating the area of the second jaw 12b. The training data is a plurality of endoscopic images IMG, an annotation indicating the area of the second jaw 12b in each image, and a coverage label indicating whether or not the second jaw 12b in each image is covered by tissue 55.
[0111] Fig. 24 shows a third example of the configuration of the inter-jaw structure detection unit. Fig. 24 shows only the trained model 122d2 for detecting the second jaw covering, but the trained model 122d1 for detecting the first jaw covering also performs similar detection and learning for the first jaw 12a. Fig. 24 shows the processing in the learning stage and the processing in the inference stage.
[0112] First, the inference stage will be described. The treatment tool detection unit 130 acquires a two-dimensional normal image and a three-dimensional image of the endoscopic image. The three-dimensional image is, for example, a depth map obtained by stereo vision. The treatment tool detection unit 130 detects the first jaw 12a and the second jaw 12b from each of the normal image and the three-dimensional image.
[0113] The inter-jaw tissue detection unit 146 inputs the three-dimensional image, the area of the second jaw 12b detected from the three-dimensional image, the normal image, and the area of the second jaw 12b detected from the normal image to the trained model 122d2. The trained model 122d2 detects whether the second jaw 12b is covered with tissue 55 from the input data and outputs the detection result 122d2Q.
[0114] Next, the learning stage will be described. The learning system generates a trained model 122d2 by training a model to detect whether the second jaw 12b is covered by tissue 55 from input data. In the figure, the symbol "DMAP" indicates a three-dimensional image, and the symbol "IMG" indicates a normal image. The input data is a three-dimensional image DMAP, an annotation indicating the area of the second jaw 12b in the three-dimensional image DMAP, a normal image IMG, and an annotation indicating the area of the second jaw 12b in the normal image IMG. The training data is a plurality of three-dimensional images DMAP, an annotation indicating the area of the second jaw 12b in each three-dimensional image DMAP, a plurality of normal images IMG, an annotation indicating the area of the second jaw 12b in each normal image IMG, and a coverage label indicating whether the second jaw 12b in each image is covered by tissue 55. A coverage label may be attached to each of the three-dimensional image and the normal image, or one coverage label may be attached to a set of a three-dimensional image and a normal image captured at the same time.
[0115] Fig. 25 shows a fourth example of the configuration of the inter-jaw structure detection unit. Fig. 25 shows only the trained model 122d2 for detecting the second jaw covering, but the trained model 122d1 for detecting the first jaw covering also performs similar detection and learning for the first jaw 12a. Fig. 25 shows the processing in the learning stage and the processing in the inference stage.
[0116] First, the inference stage will be described. The trained model 122d2 for detecting the second jaw cover includes a trained model 122d2a for a three-dimensional image and a trained model 122d2b for a normal image.
[0117] The inter-jaw tissue detection unit 146 inputs the three-dimensional image and the area of the second jaw 12b detected from the three-dimensional image to the trained model 122d2a. The trained model 122d2a detects whether the second jaw 12b is covered with tissue 55 from the input data.
[0118] The inter-jaw tissue detection unit 146 inputs the normal image and the area of the second jaw 12b detected from the normal image to the trained model 122d2b. The trained model 122d2b detects whether the second jaw 12b is covered with tissue 55 from the input data.
[0119] The inter-jaw tissue detection unit 146 integrates the detection result of the trained model 122d2a and the detection result of the trained model 122d2b, and outputs a final detection result 122d2Q indicating whether the second jaw 12b is covered with tissue 55. If the detection result of the trained model 122d2a and the detection result of the trained model 122d2b are the same, the inter-jaw tissue detection unit 146 outputs them as the final detection result 122d2Q. If the detection result of the trained model 122d2a and the detection result of the trained model 122d2b are different, the inter-jaw tissue detection unit 146 outputs the detection result with the higher prediction probability or a detection result selected based on a preset weighting as the final detection result 122d2Q.
[0120] Next, the learning stage will be described. The learning system generates a learned model 122d2a by training a model to detect whether the second jaw 12b is covered with tissue 55 from input data. The input data is a three-dimensional image DMAP and annotations indicating the area of the second jaw 12b in the three-dimensional image DMAP. The training data is a plurality of three-dimensional images DMAP, annotations indicating the area of the second jaw 12b in each three-dimensional image DMAP, and a coverage label indicating whether the second jaw 12b in each image is covered with tissue 55.
[0121] The learning system generates a trained model 122d2b by training a model to detect whether the second jaw 12b is covered by tissue 55 from input data. The input data is a normal image IMG and an annotation indicating the area of the second jaw 12b in the normal image IMG. The training data is a plurality of normal images IMG, an annotation indicating the area of the second jaw 12b in each normal image IMG, and a coverage label indicating whether the second jaw 12b in each image is covered by tissue 55.
[0122] Fig. 26 shows a second configuration example of the contact detection unit in the second embodiment. The following mainly describes the differences from the first configuration example in Fig. 21.
[0123] There are some situations where the presence or absence of tissue between the jaws cannot be visually confirmed after the jaws are closed, for example, because the pair of jaws overlap in the depth direction. In such a situation, there is a possibility that the inter-jaw tissue detection unit 146 will not be able to detect the presence or absence of tissue between the jaws. Therefore, a memory function is added to the inter-jaw tissue detection unit 146. After the contact detection unit 140 has recognized that "tissue is present between the jaws," it will maintain the recognition result of "tissue is present between the jaws" using the memory function, even if the situation changes to one where it is no longer possible to determine the presence or absence of tissue.
[0124] Specifically, memory 120 stores detection result 140Q of inter-jaw tissue detection unit 146. Memory 120 holds, for example, detection result 140Q output by inter-jaw tissue detection unit 146 for a certain period of time from the timing of output. Contact detection unit 140 detects the contact state between the treatment tool and the tissue using detection result 144Q of jaw opening / closing detection unit 144, detection result 146Q of inter-jaw tissue detection unit 146, and detection result 146Q of inter-jaw tissue detection unit 146 stored in memory 120. Specifically, when inter-jaw tissue detection unit 146 cannot determine the presence or absence of tissue between the jaws, contact detection unit 140 detects the contact state between the treatment tool and the tissue from detection result 144Q of jaw opening / closing detection unit 144 and detection result 146Q of inter-jaw tissue detection unit 146 stored in memory 120.
[0125] 27 shows an example of contact detection using detection results stored in memory. As shown in the upper diagram, at an arbitrary first time point, the jaw open / close detection unit 144 determines that the jaws are open, and the inter-jaw tissue detection unit 146 determines that tissue is present between the jaws. This determination result of "tissue present between the jaws" is stored in memory 120. As shown in the lower diagram, at a second time point after the first time point, the jaw open / close detection unit 144 determines that the jaws are closed, and the inter-jaw tissue detection unit 146 determines that the presence or absence of tissue between the jaws is unknown. In this case, the contact detection unit 140 uses the determination result of "tissue present between the jaws" at the first time point stored in memory 120 to determine that the jaws are grasping tissue, that is, that the treatment tool and tissue are in contact.
[0126] According to this embodiment, if it is recognized that there is tissue between the jaws when the jaws are opened, the contact state can be detected from the results stored in memory even if it is not possible to determine whether there is tissue between the jaws from the endoscopic image after the jaws are closed.
[0127] Fig. 28 shows a third configuration example of the contact detection unit in the second embodiment. The following mainly describes the differences from the first configuration example in Fig. 21. The current frame is denoted as t, and the frame n frames before that is denoted as tn.
[0128] First, the inference stage will be described. The treatment tool detection unit 130 detects a treatment tool from the endoscopic image of frame t and outputs a detection result 130Q. Detection results 130Q are also output for frames tn, . . . , t-1, and are stored, for example, in the memory 120. The jaw open / close detection unit 144 detects the open / close state of the jaws from the endoscopic image of frame t and outputs a detection result 144Q. The trained model 122 in FIG. 2 includes a trained model 122e for detecting inter-jaw tissue. The inter-jaw tissue detection unit 146 inputs the time-series endoscopic images of frames tn, . . . , t-1, and t, and the detection results 130Q for each image, to the trained model 122e. The trained model 122e detects the presence or absence of tissue between the jaws from the input data and outputs a detection result 146Q. An example of a trained model that handles time-series data is a neural network using Long Short Term Memory (LSTM). The jaw open / closed state detection unit 144 may also be configured to detect the open / closed state of the jaw from time-series endoscopic images.
[0129] According to this embodiment, even if the situation changes such that it is no longer possible to determine whether tissue is present between the jaws after a tissue grasping operation, the grasping state can be recognized from the time-series information. By recognizing the grasping state including the time-series data, even if it is no longer possible to determine whether tissue is present between the jaws from the endoscopic image, it is possible to recognize whether the tissue is being grasped or not from the immediately preceding grasping detection result.
[0130] 4. Third embodiment In the third embodiment, the treatment tool is an energy device. The energy device may be a device that uses high-frequency power, ultrasound, or both. The energy device may also be a bipolar device or a monopolar device. In the third embodiment, contact detection using images and contact detection using electrical information from the energy device are integrated to detect the contact state between the treatment tool and tissue. Note that a description of parts similar to the first or second embodiment will be omitted. In the third embodiment, for example, the configuration of the medical system 1, the configuration of the controller 100, and the processing performed by the treatment tool detection unit 130 are similar to those in the first embodiment.
[0131] First, the generator 235 that drives the energy device 236 will be described. The generator 235 supplies energy to the energy device 236, controls the energy supply, and acquires electrical information. That is, the generator 235 outputs high-frequency power, and the energy device 236 outputs the high-frequency power from an end effector. Alternatively, the end effector of the energy device 236 has an ultrasonic device. The generator 235 outputs a drive signal for the ultrasonic device, and the ultrasonic device of the end effector receives the drive signal and outputs ultrasonic waves.
[0132] The electrical information is electrical impedance obtained when the energy device 236 outputs high-frequency power to tissue. However, the electrical information is not limited to this. The electrical information may be current, voltage, or the phase between current and voltage. Alternatively, the electrical information may be power, amount of power, impedance, resistance, reactance, admittance (the inverse of impedance), conductance (the real part of admittance), or susceptance (the imaginary part of admittance). Alternatively, the electrical information may be the above-mentioned changes over time, changes between each parameter, differential and integral calculations between each parameter (when P is a parameter, the differential over time is dP / dt, and the differential with respect to resistance is dP / dR), values derived by elementary calculations such as sums and differences of each block, or trigger information such as whether each threshold has been crossed.
[0133] Alternatively, the electrical information may be mechanical impedance obtained when the energy device 236 outputs ultrasound waves to tissue. However, the electrical information is not limited to this. The electrical information may be a change in mechanical impedance over time, or trigger information such as whether or not a respective threshold value has been crossed.
[0134] Hereinafter, electrical information will be referred to as electrical impedance or mechanical impedance, and will be simply referred to as impedance. Impedance in the following description can be read as electrical information.
[0135] 29 shows a first configuration example of the contact detection unit in the third embodiment. The contact detection unit 140 includes an image information determination unit 147, an electrical information determination unit 142, and a determination unit 143.
[0136] The image information determination unit 147 detects the contact state between the treatment tool and tissue from the endoscopic image using the trained model for contact detection 122b. The method for detecting contact from this image is as described in the first embodiment. Alternatively, the method for detecting contact from the image may be the method described in the second embodiment.
[0137] When the image information determination unit 147 determines that the treatment tool is in contact with the tissue, the electrical information determination unit 142 acquires impedance from the generator 235 and determines the contact state between the treatment tool and the tissue from the impedance. When the jaws grasp the tissue or when the monopolar is in contact with the tissue, an impedance corresponding to the tissue is observed. This makes it possible to determine whether contact is occurring or not based on the impedance. As an example, the electrical information determination unit 142 determines whether the impedance is equal to or greater than a specified value to determine whether contact is occurring or not. The electrical information determination unit 142 may be implemented by a rule-based program or a trained model using machine learning.
[0138] The determination unit 143 integrates the detection result of the image information determination unit 147 and the detection result of the electrical information determination unit 142 to detect the contact state between the treatment tool and the tissue, and outputs the detection result 140Q. The integration process will be described with reference to the flows in Figures 30 and 31.
[0139] According to this embodiment, contact detection is performed by combining images and impedance, which is expected to improve the accuracy of contact detection compared to contact detection using images alone. Furthermore, when contact detection is performed using only electrical or load information, contact detection can only begin when the handpiece switch is turned on. This delays the action taken after contact detection. The action may be, for example, tissue type recognition or automatic energy adjustment based on the recognition. Such a delay reduces the value of contact detection. According to this embodiment, contact detection is also performed using images, which makes it possible to speed up the action taken after contact detection.
[0140] 30 shows a first example flow of processing performed by the contact detection unit in the third embodiment. In this flow, when the trained model 122b of the image information determination unit 147 determines that the treatment tool is in contact with tissue, the determination unit 143 performs case classification according to the probability that contact has been determined.
[0141] In step S11, the image information determination unit 147 performs contact detection using the endoscopic image. This process is performed for each frame of the endoscopic image.
[0142] When it is determined in step S11 that the treatment tool and tissue are in contact with each other, in step S12, the determination unit 143 compares the probability of contact output by the trained model 122b of the image information determination unit 147 with a threshold value x. The threshold value x is, for example, 60%, but may be any value.
[0143] If the probability is equal to or less than the threshold value x in step S12, the electrical information determination unit 142 performs contact detection using impedance in step S16. The determination unit 143 uses the detection result of the contact state output by the electrical information determination unit 142 as the final detection result 140Q.
[0144] If the probability is greater than the threshold value x in step S12, the electrical information determination unit 142 performs contact detection using impedance in step S13. The determination unit 143 compares the detection result output by the image information determination unit 147 with the detection result output by the electrical information determination unit 142.
[0145] If the two detection results match in step S13, then in step S14, the determination unit 143 adopts the contact state detection result output by the image information determination unit 147 as the final detection result 140Q.
[0146] If the two detection results do not match in step S13, in step S15, the judgment unit 143 compares the probability of contact output by the trained model 122b of the image information judgment unit 147 with a threshold y. The threshold y is, for example, 95%, but may be any value greater than the threshold x.
[0147] If the probability is equal to or less than the threshold value y in step S15, then in step S16, the determination unit 143 adopts the contact state detection result output by the electrical information determination unit 142 as the final detection result 140Q.
[0148] If the probability is greater than the threshold value y in step S15, the determination unit 143 adopts the contact state detection result output by the image information determination unit 147 as the final detection result 140Q in step S14.
[0149] FIG. 31 shows a second example flow of processing performed by the contact detection unit in the third embodiment. In this flow, if the contact detection results of the image and impedance differ, the amount of electrical information used for detection is increased. By utilizing multiple pieces of electrical information rather than just a single piece of impedance information, detection accuracy is improved. In addition, by utilizing multiple pieces of electrical information only when necessary, the recognition flow is efficient.
[0150] Differences from the first flow in FIG. 30 will be described. If the two detection results do not match in step S13, in step S17, the electrical information determination unit 142 determines the contact state between the treatment tool and the tissue from the plurality of pieces of electrical information. As an example, the plurality of pieces of electrical information are impedance, impedance variation, and current value. In this case, the electrical information determination unit 142 determines that the treatment tool and the tissue are in contact when the impedance is less than a specified value, the impedance variation is less than a specified value, and the current value is equal to or greater than a specified value. After step S17, in step S16, the determination unit 143 adopts the detection result of the contact state output by the electrical information determination unit 142 as the final detection result 140Q.
[0151] In this flow, in normal operation, the electrical information determination unit 142 performs contact detection by determining whether the impedance is equal to or greater than a specified value. The normal operation refers to the path from S12 to S16, or the path from S12 to S13 and S14. Only in the path from S12 to S13, S17, and S16 where it is determined that the two detection results do not match, the electrical information determination unit 142 performs contact detection using multiple pieces of electrical information.
[0152] 32 shows a second configuration example of the contact detection unit in the third embodiment. The contact detection unit 140 includes an image information determination unit 147, an electrical information determination unit 142, and a separation determination unit 148.
[0153] 2 includes a trained model 122f for detecting tissue divisions from an image. First, the inference stage will be described.
[0154] The image information determination unit 147 inputs the endoscopic image and the detection result 130Q of the treatment tool detection unit 130 to the trained model 122f, or inputs an image of a region of interest in the endoscopic image that is set based on the detection result 130Q to the trained model 122f. The trained model 122f detects tissue separation from the input data.
[0155] The electrical information determination unit 142 acquires the detection result of tissue separation using the separation detection function of the generator 235. Alternatively, the electrical information determination unit 142 may acquire impedance from the generator 235 and detect tissue separation from the impedance. This separation detection may be performed only when the image information determination unit 147 determines that the treatment tool is in contact with the tissue, or may be performed continuously. As an example, the energy device is an ultrasonic device. The generator 235 may detect load fluctuations of the ultrasonic probe that occur when tissue separates, and the electrical information determination unit 142 may acquire the detection result of tissue separation.
[0156] The slice / segment determination unit 148 integrates the slice / segment detection result output by the image information determination unit 147 and the slice / segment detection result output by the electrical information determination unit 142 to output a final slice / segment detection result 140Q. The processor 110 performs output using the detection result 140Q. For example, when it is determined that the tissue has been sliced, the processor 110 transmits an instruction to the generator 235 to suppress or stop the energy output, and the generator 235 suppresses or stops the energy output.
[0157] Next, the learning stage of the trained model 122f will be described. The learning system generates the trained model 122f by training the model to detect tissue separation from input data. The input data is an endoscopic image IMG and annotations indicating the position or area of a treatment tool. Alternatively, the input data is an image of a region of interest in the endoscopic image IMG that is set based on the annotation. The training data is a plurality of endoscopic images IMG, annotations indicating the position or area of a treatment tool in each image, and labels indicating the tissue separation state in each image. The labels indicating the separation state are, for example, a label indicating before separation and a label indicating after separation.
[0158] The impedance-based detection function for tissue separation described above is affected by factors such as the state of the tissue, the type of tissue, and how the handle is gripped or the tissue is grasped. Furthermore, there may be situations where tissue separation cannot be detected using images alone due to the influence of tissue adhesion. By combining electrical information-based and image-based separation detection, the frequency of missed detections can be reduced. As a result, excessive temperature rise in the device after tissue incision can be reduced, reducing the risk of damage to the probe or tissue pad.
[0159] In the above-described method and the first to third embodiments, the medical system 1 includes a memory 120 storing a trained model 122, and a processor 110. The trained model 122 is a model trained using training data including an endoscopic image and a contact state between a treatment tool and tissue in the endoscopic image. The endoscopic image is captured by an endoscope that captures an image including the treatment tool and tissue to be treated. The processor 110 acquires the endoscopic image including the treatment tool and tissue. The processor 110 uses the trained model to detect the contact state between the treatment tool and tissue from the endoscopic image as first information.
[0160] According to this embodiment, the contact state between the treatment tool and tissue is detected from an endoscopic image, and the detection results can be used to provide various types of support to the surgeon. For example, support information, such as information about the tissue in contact, can be presented to the surgeon only when the treatment tool comes into contact with the tissue. This prevents unnecessary information presentation to the surgeon, reducing the surgeon's inconvenience. Furthermore, since the contact state between the treatment tool and tissue can be detected by image-based recognition processing, the contact state can be recognized without increasing the size of the treatment tool and keeping device costs low compared to when a tactile sensor or the like is provided on the treatment tool. Furthermore, since contact detection from an image does not limit the device used as the treatment tool, it can be applied to a variety of existing treatment tools.
[0161] In this embodiment, the processor 110 may detect a region of interest including a treatment tool from the endoscopic image. The processor 110 may detect a contact state between the treatment tool and tissue from an image of the region of interest in the endoscopic image using the trained model 122. The setting of the region of interest is described with reference to FIGS. 13 to 18, etc.
[0162] Among the tissues, the tissues that are affected by the treatment tool are the tissues surrounding the treatment tool that is in contact with the tissue. For example, there is a high correlation between the movement vectors of the treatment tool and the tissues surrounding it. Alternatively, it is the tissues surrounding the treatment tool that are affected by the energy treatment. According to this embodiment, a region of interest including the treatment tool is detected, and the contact state is detected from an image of that region of interest, so that the contact state can be detected with high accuracy using the tissues surrounding the treatment tool that are affected by the treatment tool.
[0163] In this embodiment, when the processor 110 detects first information indicating that the treatment tool is in contact with tissue, the processor 110 may perform support processing related to treatment using the treatment tool.
[0164] In this embodiment, the support process may be a process of presenting first information, a process of presenting support information related to treatment, a process of presenting support information related to energy output settings of a treatment tool that is an energy device, or a process of automatically controlling the treatment tool that is an energy device. An example of the support information is described in FIG. 4 etc.
[0165] According to this embodiment, the contact state between the treatment tool and tissue is detected from an endoscopic image, and the detection results can be used to provide various types of support to the surgeon. "Support information related to treatment" is, for example, the type or state of tissue with which the treatment tool has come into contact. "Support information related to energy output settings" is, for example, recommended output settings determined from images, etc. "Processing for automatically controlling the treatment tool" is, for example, processing for changing, suppressing, or stopping energy output. Note that the support processing is not limited to the examples given here, as long as it is support using the detection results of the contact state.
[0166] In this embodiment, the processor 110 may also obtain an average movement vector of the treatment tool and an average movement vector of the surrounding tissue, which is tissue around the treatment tool, from the endoscopic image. The processor 110 may detect the contact state between the treatment tool and the tissue from the average movement vector of the treatment tool and the average movement vector of the surrounding tissue, using the trained model 122. Contact detection using movement vectors is described with reference to FIGS. 5 to 11, etc.
[0167] 6 etc., when the treatment tool is in contact with the tissue, the correlation between the movement vectors of the treatment tool and the tissue is high, and when the treatment tool is not in contact with the tissue, the correlation between the movement vectors of the treatment tool and the tissue is low. According to this embodiment, this correlation can be used to detect the contact state between the treatment tool and the tissue.
[0168] Furthermore, in this embodiment, the processor 110 may detect the contact state between the treatment tool and tissue using at least one of the following (i), (ii), (iii), and (iv): (i) the average movement vector of the treatment tool; (ii) the average movement vector of the surrounding tissue; (iii) the Euclidean distance between the average movement vector of the treatment tool and the average movement vector of the surrounding tissue; and (iv) the directional difference between the average movement vector of the treatment tool and the average movement vector of the surrounding tissue. These methods are described with reference to FIGS. 7 to 9, etc.
[0169] In this embodiment, the processor 110 may also decompose the movement vector of the surrounding tissue into a component parallel to the movement direction of the treatment tool and a component perpendicular to the movement direction of the treatment tool, and detect the contact state between the treatment tool and the tissue using statistics of the parallel component and the perpendicular component. These methods are described in Figures 10 to 11, etc.
[0170] According to this embodiment, by using the average movement vector or a quantity calculated from the average movement vector, the correlation between the movement vectors of the treatment tool and the tissue can be evaluated and the contact state between the treatment tool and the tissue can be detected.
[0171] In this embodiment, the treatment tool may be an energy device that treats tissue by outputting energy from an end effector. The processor 110 may determine the contact state between the treatment tool and tissue using first information that is the contact state determined from an endoscopic image and second information that is the contact state determined from electrical information in the energy output. Such a method is described in Figures 29 to 32, etc.
[0172] According to this embodiment, contact detection is performed by combining image-based contact detection and electrical information-based contact detection, so that improved accuracy of contact detection can be expected compared to contact detection performed from only one of them.
[0173] In this embodiment, the contact state may be determined by determining whether the treatment tool is in contact with the tissue or not. The processor 110 may determine whether to adopt the first information or the second information depending on the estimated probability that the treatment tool is in contact with the tissue in the estimation of the first information using the trained model 122. Such a method is described with reference to FIGS. 30 to 31, etc.
[0174] According to this embodiment, it is possible to determine whether to adopt the result of contact detection based on an image or the result of contact detection based on electrical information, depending on the estimated probability of contact detection based on an image. For example, when the estimated probability of contact detection based on an image is high, it is possible to adopt the result of contact detection based on an image. When the estimated probability of contact detection based on an image is high, it is possible to determine whether to adopt the result of contact detection based on an image or the result of contact detection based on electrical information, while taking into account the result of contact detection based on electrical information.
[0175] In this embodiment, if the first information and the second information do not match, the processor 110 may determine the contact state from a plurality of pieces of electrical information in the energy output. Such a method is described with reference to FIG. 31 etc.
[0176] According to this embodiment, when the result of contact detection based on the image and the result of contact detection based on the electrical information do not match, it is assumed that the estimated probability of contact detection is low. In such a case, the contact state can be detected with high accuracy by determining the contact state from multiple pieces of electrical information.
[0177] In this embodiment, the treatment tool may have a pair of jaws that open and close. The memory 120 may store a first trained model and a second trained model as trained models 122. The first trained model may be a model trained to detect whether the jaws are open or closed from an endoscopic image. The second trained model may be a model trained to detect the presence or absence of tissue between the jaws from an endoscopic image. The processor 110 may perform first detection to detect whether the jaws are open or closed from an endoscopic image using the first trained model. The processor 110 may perform second detection to detect the presence or absence of tissue between the jaws from an endoscopic image using the second trained model. The processor 110 may detect the contact state between the treatment tool and tissue based on the results of the first detection and the second detection. Such a method is described with reference to FIGS. 21 to 28, etc. In the configuration examples of FIGS. 21 and 26, the trained model 122c is the first trained model, and the trained model 122d is the second trained model. In the configuration example of FIG. 28, trained model 122c is the first trained model, and trained model 122e is the second trained model.
[0178] According to this embodiment, the trained model for contact detection is separated into a first trained model for detecting jaw opening and closing and a second trained model for detecting inter-jaw tissue. This improves the estimation accuracy of each individual model compared to estimating contact detection with a single model, enabling contact detection with higher accuracy.
[0179] In addition, in this embodiment, the processor 110 may determine that the treatment tool is in contact with tissue when the first detection detects that the jaws are closed and the second detection detects that tissue is present between the jaws.
[0180] In this embodiment, the condition for determining that tissue is present between the jaws in the second detection is that one of the pair of jaws that opens and closes is covered with tissue, and the other jaw is not covered with tissue.
[0181] By using these judgment conditions, it is possible to determine whether the jaws of the treatment tool are grasping tissue, i.e., whether the treatment tool is in contact with tissue, based on the results of jaw opening / closing detection and the results of tissue detection between the jaws.
[0182] In this embodiment, the treatment tool may be an energy device that treats tissue by outputting energy from an end effector. When the processor 110 detects first information indicating contact between the treatment tool and tissue, the processor 110 may present an output setting for the energy output. When the processor 110 receives an output instruction after the presentation, the processor 110 may cause the energy device to output energy at the output setting. Such a method is described with reference to FIGS. 1 and 4, etc.
[0183] According to this embodiment, only when the treatment tool comes into contact with tissue, support information such as information about the tissue in contact is presented to the surgeon, and the surgeon is prompted to input approval or rejection. This prevents the presentation of information at times that are unnecessary for the surgeon, reducing the inconvenience.
[0184] In this embodiment, the treatment tool may be an energy device that treats tissue by outputting energy from an end effector. The processor 110 may make a decision regarding a change in the energy output setting when it detects first information indicating that the treatment tool and tissue have changed from contact to non-contact. Detection of tissue separation is described in Figures 19 and 20, etc. The "decision regarding a change in the energy output setting" is, for example, a decision regarding whether to change, suppress, or stop the energy output when tissue is separated.
[0185] According to this embodiment, by detecting a change from contact between the treatment tool and the tissue to non-contact, it is possible to detect that the tissue has been cut by the energy treatment. When the treatment tool and the tissue are no longer in contact, a decision can be made regarding changing the energy output settings, thereby assisting the surgeon.
[0186] This embodiment may also be implemented as a contact state detection method as follows. That is, the contact state detection method includes capturing an endoscopic image including a treatment tool and tissue to be treated by an endoscope. The contact state detection method includes detecting, from the endoscopic image, a contact state between the treatment tool and the tissue as first information using a trained model. The trained model is a model trained using training data including the endoscopic image and the contact state between the treatment tool and the tissue in the endoscopic image.
[0187] 5. Fourth embodiment In the fourth embodiment, the treatment tool is an energy device. The energy device may be a device that uses high-frequency power, ultrasound, or both. The energy device may also be a bipolar or monopolar device. In the fourth embodiment, the contact state between the treatment tool and tissue is detected by combining jaw open / close detection using images and inter-jaw tissue detection using electrical information. Note that a description of parts similar to the first, second, or third embodiment will be omitted. In the fourth embodiment, for example, the configuration of the medical system 1, the configuration of the controller 100, and the processing performed by the treatment tool detection unit 130 are similar to those in the first embodiment.
[0188] 33 shows an example of the configuration of the contact detection unit in the fourth embodiment. The contact detection unit 140 includes a jaw open / close detection unit 144 and an electrical information determination unit 142.
[0189] The jaw open / closed state detection unit 144 detects the open / closed state of the jaw shown in the endoscope image from the endoscope image using the trained model 122c. This detection method is as described with reference to FIG.
[0190] When the jaw open / close detection unit 144 determines that the jaws are closed, the electrical information determination unit 142 acquires electrical information from the generator 235, detects the contact state between the treatment tool and the tissue using the electrical information, and outputs the detection result 140Q. Specifically, the electrical information determination unit 142 determines the presence or absence of tissue between the jaws using the electrical information. This detection method is as described in the third embodiment.
[0191] When the electrical information determining section 142 determines that tissue is present between the jaws, the contact detecting section 140 determines that the treatment tool is gripping tissue, that is, that the treatment tool is in contact with tissue.
[0192] In the above-described fourth embodiment, the medical system 1 includes a memory 120 storing a trained model 122 and a processor 110. The trained model 122 is a model trained using training data. The training data includes an endoscopic image captured by an endoscope that captures an image including a treatment tool having opening and closing jaws and tissue to be treated, and the open / closed state of the jaws in the endoscopic image. The treatment tool is an energy device that treats tissue by outputting energy from the jaws. The processor 110 acquires an endoscopic image including the treatment tool and tissue. The processor 110 performs a first detection using the trained model to detect whether the jaws are open or closed from the endoscopic image. The processor 110 performs a second detection to detect the presence or absence of tissue between the jaws based on electrical information in the energy output. The processor 110 detects the state of tissue gripped by the jaws as a contact state based on the results of the first detection and the second detection.
[0193] When considering recognition from an image, it is assumed that jaw opening and closing is easier to detect than contact between a treatment tool and tissue. According to this embodiment, the contact state can be detected with high accuracy by combining jaw opening and closing detection, which is thought to be more accurate than contact state detection, with inter-jaw tissue detection using impedance.
[0194] In this embodiment, when the processor 110 detects in the first detection that the jaws are closed, it performs the second detection.
[0195] According to this embodiment, only jaw open / close detection is performed using images until jaw closure is detected, and then inter-jaw tissue detection is performed using impedance, thereby reducing the computational load.
[0196] In this embodiment, the processor 110 determines that the jaws are grasping tissue when it detects that the jaws are closed in the first detection and detects that there is tissue between the jaws in the second detection.
[0197] By using such judgment conditions, it is possible to determine whether the jaws of the treatment tool are grasping tissue, i.e., whether the treatment tool is in contact with tissue, based on the results of jaw opening / closing detection and the results of tissue detection between the jaws.
[0198] Furthermore, this embodiment may be implemented as a contact state detection method as follows. That is, the contact state detection method is a method for detecting a contact state between a treatment tool and tissue to be treated. The treatment tool is an energy device that has jaws that open and close and treats tissue by outputting energy from the jaws. The contact state detection method includes capturing an endoscopic image including the treatment tool and the tissue to be treated by an endoscope. The contact state detection method includes performing first detection to detect whether the jaws are open or closed from the endoscopic image using a trained model 122. The trained model 122 is a model trained with training data including the endoscopic image and the open / closed state of the jaws in the endoscopic image. The contact state detection method includes performing second detection to detect whether tissue is present between the jaws based on electrical information in the energy output. The contact state detection method includes detecting a state in which the tissue is gripped by the jaws as a contact state based on a result of the first detection and a result of the second detection.
[0199] Although the above describes embodiments and variations to which the present disclosure is applied, the present disclosure is not limited to the embodiments and variations thereof as they are. In practice, the components can be modified and embodied within the scope of the gist of the disclosure. Furthermore, various disclosures can be formed by appropriately combining multiple components disclosed in the above-described embodiments and variations. For example, some components may be omitted from all components described in each embodiment or variation. Furthermore, components described in different embodiments or variations may be appropriately combined. In this manner, various modifications and applications are possible within the scope of the gist of the disclosure. Furthermore, a term described at least once in the specification or drawings together with a different term having a broader or equivalent meaning may be replaced with that different term anywhere in the specification or drawings. [Explanation of symbols]
[0200] 1... medical system, 10... treatment tool, 11... shaft, 12... jaw, 14... end effector, 50... surrounding tissue, 100... controller, 110... processor, 120... memory, 121... program, 122... model, 130... treatment tool detection unit, 140... contact detection unit, 141... pre-processing unit, 142... electrical information determination unit, 143... determination unit, 144... jaw opening / closing detection unit, 146... inter-jaw tissue detection unit, 147... image information determination unit, 148... determination unit, 180, 190... I / O device, 200... endoscope system, 230... system, 231... monitor, 232... endoscopic image, 233... output setting, 235... generator, 236... energy device, IMG... endoscopic image, ROIc... region of interest, Vtissue1, Vtissue2... movement vector, Vtool... movement vector, v... average movement vector, w... average movement vector
Claims
1. a memory that stores a trained model trained using training data including an endoscopic image captured by an endoscope that captures an image including a treatment tool and a tissue to be treated, and a contact state between the treatment tool and the tissue in the endoscopic image; a processor; Including, The processor: acquiring the endoscopic image including the treatment tool and the tissue; detecting a region of interest including the treatment tool from the endoscopic image; A medical system characterized by using the trained model to detect the contact state between the treatment tool and the tissue as first information from an image of the region of interest in the endoscopic image.
2. a memory that stores a trained model trained using training data including information acquired from an endoscopic image captured by an endoscope that captures an image including a treatment tool and a tissue to be treated, and a contact state between the treatment tool and the tissue in the endoscopic image; a processor; Including, the information is at least one of an average movement vector of the treatment tool and an average movement vector of the surrounding tissue, which is the tissue around the treatment tool; a Euclidean distance between the average movement vector of the treatment tool and the average movement vector of the surrounding tissue; a directional difference between the average movement vector of the treatment tool and the average movement vector of the surrounding tissue; and a statistic of a parallel component and a perpendicular component when the movement vector of the surrounding tissue is decomposed into a component parallel to the movement direction of the treatment tool and a component perpendicular to the movement direction of the treatment tool; The processor: acquiring the endoscopic image including the treatment tool and the tissue; acquiring the information from the endoscopic image; A medical system characterized in that the contact state between the treatment tool and the tissue is detected as first information from the information using the trained model.
3. a memory that stores the first trained model and the second trained model as trained models; a processor; Including, The first trained model is The device is trained to detect opening and closing of the jaws from an endoscopic image captured by an endoscope that captures an image including a treatment tool having a pair of jaws that open and close and tissue to be treated, The second trained model is and learning to detect the presence or absence of the tissue between the jaws from the endoscopic image; The processor: acquiring the endoscopic image including the treatment tool and the tissue; performing a first detection of detecting opening and closing of the jaw from the endoscopic image using the first trained model; performing a second detection using the second trained model to detect the presence or absence of tissue between the jaws from the endoscopic image; A medical system comprising: a medical device for detecting a contact state between the treatment tool and the tissue as first information based on a result of the first detection and a result of the second detection.
4. 4. The medical system according to claim 3, The processor: A medical system characterized in that it is determined that the treatment tool and the tissue are in contact when the first detection detects that the jaws are closed and the second detection detects that the tissue is between the jaws.
5. 4. The medical system according to claim 1, The processor: A medical system characterized in that, when the first information indicating that the treatment tool is in contact with the tissue is detected, support processing related to treatment using the treatment tool is performed.
6. 6. The medical system according to claim 5, The medical system is characterized in that the support processing is a processing of presenting the first information, a processing of presenting support information regarding the treatment, a processing of presenting support information regarding energy output settings of the treatment tool which is an energy device, or a processing of automatically controlling the treatment tool which is an energy device.
7. 4. The medical system according to claim 1, the treatment tool is an energy device that treats the tissue by outputting energy from an end effector, The processor: A medical system characterized in that the contact state between the treatment tool and the tissue is determined using the first information and second information, which is the contact state determined from electrical information in the energy output.
8. 8. The medical system according to claim 7, The determination of the contact state is a determination of whether the treatment tool is in contact with the tissue or not, The processor: A medical system characterized by determining whether the first information or the second information is to be adopted depending on the estimated probability that the treatment tool and the tissue are in contact in estimating the first information using the trained model.
9. 8. The medical system according to claim 7, The processor: A medical system characterized in that, when the first information and the second information do not match, the contact state is determined from a plurality of pieces of electrical information in the energy output.
10. 4. The medical system according to claim 1, the treatment tool is an energy device that treats the tissue by outputting energy from an end effector, The processor: a medical system that, when the first information indicating contact between the treatment tool and the tissue is detected, presents an output setting for the energy output, and, when an output instruction is received after the presentation, causes the energy device to output the energy at the output setting.
11. 4. The medical system according to claim 1, the treatment tool is an energy device that treats the tissue by outputting energy from an end effector, The processor: A medical system characterized in that, when a change in the first information indicates that the treatment tool and the tissue have changed from contact to non-contact, a decision is made regarding a change in the energy output setting.
12. A medical system comprising: an endoscope capturing an endoscopic image including a treatment tool and tissue to be treated; the medical system detects a region of interest including the treatment tool from the endoscopic image; the medical system detects, as first information, the contact state between the treatment tool and the tissue from an image of the region of interest in the endoscopic image using a trained model trained with training data including the endoscopic image and a contact state between the treatment tool and the tissue in the endoscopic image; 10. A method of operating a medical system, comprising:
13. A medical system comprising: an endoscope capturing an endoscopic image including a treatment tool and tissue to be treated; The medical system acquires, from the endoscopic image, information that is at least one of the average movement vector of the treatment tool and the average movement vector of the surrounding tissue, which is the tissue around the treatment tool, the Euclidean distance between the average movement vector of the treatment tool and the average movement vector of the surrounding tissue, the directional difference between the average movement vector of the treatment tool and the average movement vector of the surrounding tissue, and statistics of the parallel component and the perpendicular component when the movement vector of the surrounding tissue is decomposed into a component parallel to the movement direction of the treatment tool and a component perpendicular to the movement direction of the treatment tool; the medical system detects, from the information acquired from the endoscopic image, the contact state between the treatment tool and the tissue as first information using a trained model trained with training data including the information acquired from the endoscopic image and the contact state between the treatment tool and the tissue in the endoscopic image; 10. A method of operating a medical system, comprising:
14. A medical system comprising: an endoscope capturing an endoscopic image including a treatment tool having a pair of jaws that open and close and a tissue to be treated; The medical system performs a first detection to detect opening and closing of the jaw from the endoscopic image using a first trained model that has been trained to detect opening and closing of the jaw from the endoscopic image; the medical system performs a second detection to detect the presence or absence of tissue between the jaws from the endoscopic image using a second trained model trained to detect the presence or absence of tissue between the jaws from the endoscopic image; the medical system detects a contact state between the treatment tool and the tissue as first information based on a result of the first detection and a result of the second detection; 10. A method of operating a medical system, comprising:
15. a memory that stores a trained model trained using training data including an endoscopic image captured by an endoscope that captures an image including a treatment tool having opening and closing jaws and tissue to be treated, and the open / closed state of the jaws in the endoscopic image; a processor; Including, the treatment tool is an energy device that treats the tissue by outputting energy from the jaws, The processor: acquiring the endoscopic image including the treatment tool and the tissue; performing a first detection of detecting opening and closing of the jaw from the endoscopic image using the trained model; performing a second detection based on the electrical information of the energy output to detect the presence or absence of tissue between the jaws; A medical system comprising: a medical device for detecting a state in which the jaws are gripping the tissue as a contact state based on a result of the first detection and a result of the second detection.
16. 16. The medical system of claim 15, The processor: A medical system, characterized in that the second detection is performed when it is detected in the first detection that the jaws are closed.
17. 16. The medical system of claim 15, The processor: a medical system that determines that the jaws are grasping the tissue when the first detection detects that the jaws are closed and the second detection detects that the tissue is between the jaws.
18. A method for operating a medical system that detects a contact state between a treatment tool, which is an energy device having jaws that open and close and that treats tissue by outputting energy from the jaws, and the tissue to be treated, comprising: the medical system captures an endoscopic image including the treatment tool and the tissue using an endoscope; the medical system performs a first detection of detecting whether the jaw is open or closed from the endoscopic image using a trained model trained with training data including the endoscopic image and the open / closed state of the jaw in the endoscopic image; the medical system performs a second detection to detect the presence or absence of tissue between the jaws based on electrical information of the energy output; the medical system detects a state in which the jaws grasp the tissue as the contact state based on a result of the first detection and a result of the second detection; 10. A method of operating a medical system, comprising:
Citation Information
Patent Citations
Medical instrument
JP2003061979A
Medical system, notification method, and method for operating medical system
WO2023149527A1
Medical assistance device, endoscope, medical assistance method, and program
WO2024095675A1
Medical assistance device, endoscope, and medical assistance method
WO2024095676A1