Information processing device, information processing method, and computer program
The information processing apparatus integrates multiple learning models to overcome the limitations of single-class inference, providing enhanced recognition results for medical images.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ANAUT INC
- Filing Date
- 2024-09-17
- Publication Date
- 2026-05-29
Smart Images

Figure 0007867294000001 
Figure 0007867294000002 
Figure 0007867294000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a computer program.
Background Art
[0002] In the medical field, since pathological diagnosis requires specialized knowledge and experience, a part of body tissue is excised and diagnosis is performed using a microscope or the like outside the body.
[0003] On the other hand, in recent years, inventions have been proposed in which the inside of the body is directly observed using an endoscope or the like, and pathological diagnosis is performed by analyzing an observation image with a computer (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In creating a learning model for recognizing lesions and the like from an observation image, knowledge of a specialist doctor is indispensable. However, since the specialist fields of doctors are subdivided, single-class teacher data is often created. A learning model created with single-class teacher data can only perform single-class inference.
[0006] An object of the present invention is to provide an information processing apparatus, an information processing method, and a computer program that derive an integrated recognition result from calculation results of a plurality of types of learning models.
Means for Solving the Problems
[0007] An information processing device according to one aspect of the present invention includes a first calculation unit that performs calculations by a first learning model in response to an input of a surgical field image, a second calculation unit that performs calculations by a second learning model in response to an input of a surgical field image, a derivation unit that derives an integrated recognition result for the surgical field image based on the calculation result output from the first calculation unit and the calculation result output from the second calculation unit, and an output unit that outputs information based on the derived recognition result.
[0008] An information processing method in one aspect of the present invention involves a computer performing calculations by a first learning model in response to an input of a surgical field image, performing calculations by a second learning model in response to the input of the surgical field image, deriving an integrated recognition result for the surgical field image based on the calculation results of the first learning model and the calculation results of the second learning model, and outputting information based on the derived recognition result.
[0009] A computer program in one aspect of the present invention is a computer program that causes a computer to perform the following processes: execute calculations by a first learning model in response to input of a surgical field image; execute calculations by a second learning model in response to input of the surgical field image; derive an integrated recognition result for the surgical field image based on the calculation results of the first learning model and the calculation results of the second learning model; and output information based on the derived recognition result. [Effects of the Invention]
[0010] According to this invention, an integrated recognition result can be derived from the computational results of multiple types of learning models. [Brief explanation of the drawing]
[0011] [Figure 1] This is a schematic diagram illustrating the general configuration of the surgical support system according to Embodiment 1. [Figure 2] This is a block diagram illustrating the internal configuration of an information processing device. [Figure 3] This is a schematic diagram showing an example of a surgical field image. [Figure 4]It is a schematic diagram showing a configuration example of the first learning model. [Figure 5] It is a schematic diagram showing a configuration example of the second learning model. [Figure 6] It is a flowchart showing the procedure of the processing executed by the information processing apparatus according to Embodiment 1. [Figure 7] It is a schematic diagram showing an example of display of recognition results in Embodiment 1. [Figure 8] It is a schematic diagram showing an example of display of warning information in Embodiment 1. [Figure 9] It is a schematic diagram showing an example of display of recognition results according to confidence level. [Figure 10] It is a schematic diagram showing a configuration example of the third learning model. [Figure 11] It is a flowchart showing the procedure of the processing executed by the information processing apparatus according to Embodiment 2. [Figure 12] It is a schematic diagram showing an example of display of recognition results in Embodiment 2. [Figure 13] It is a schematic diagram showing a configuration example of the fourth learning model. [Figure 14] It is a flowchart showing the procedure of the processing executed by the information processing apparatus according to Embodiment 3. [Figure 15] It is a flowchart showing the procedure for导出 dimensional information. [Figure 16] It is a schematic diagram showing a configuration example of the fifth learning model. [Figure 17] It is a flowchart showing the procedure of the processing executed by the information processing apparatus according to Embodiment 4. [Figure 18] It is a flowchart showing the processing procedure in Variation 4-1. [Figure 19] It is a conceptual diagram showing an example of a prior information table. [Figure 20] It is a flowchart showing the processing procedure in Variation 4-2. [Figure 21] It is a flowchart showing the processing procedure in Variation 4-3. [Figure 22] It is a flowchart showing the processing procedure in Variation 4-4. [Figure 23] It is a conceptual diagram showing an example of a proper name table. [Figure 24] It is a schematic diagram showing an example of the display of organ names. [Figure 25] It is a conceptual diagram showing an example of a structure table. [Figure 26] It is a schematic diagram showing an example of the display of a structure. [Figure 27] It is a conceptual diagram showing an example of a case table. [Figure 28] It is a schematic diagram showing an example of the display of an event. [Figure 29] It is a flowchart showing the procedure of the process executed by the information processing apparatus according to Embodiment 5. [Figure 30] It is a conceptual diagram showing an example of a scene recording table. [Figure 31] It is a flowchart showing the procedure of the process executed by the information processing apparatus according to Embodiment 6. [Figure 32] It is an explanatory diagram explaining the estimation method in Embodiment 7. [Figure 33] It is a flowchart showing the estimation procedure in Embodiment 7. [Figure 34] It is an explanatory diagram explaining the analysis method of the calculation result. [Figure 35] It is a diagram showing an example of an evaluation coefficient table. [Figure 36] It is a diagram showing an example of the calculation result of a score. [Figure 37] It is a flowchart showing the procedure of the process executed by the information processing apparatus according to Embodiment 8. [Figure 38] It is a sequence diagram showing an example of the process executed by the information processing apparatus according to Embodiment 9. [Figure 39] It is a flowchart showing the procedure of the process executed by the first arithmetic unit. [Figure 40] It is a flowchart showing the procedure of the process executed by the second arithmetic unit.
Mode for Carrying Out the Invention
[0012] The following describes in detail, with reference to the drawings, an embodiment of the present invention applied to a laparoscopic surgery support system. It should be noted that the present invention is not limited to laparoscopic surgery, but is applicable to all endoscopic surgeries using imaging devices such as thoracoscopy, gastrointestinal endoscopes, cystoscopes, arthroscopes, robot-assisted endoscopes, surgical microscopes, and exoscopy. (Embodiment 1) Figure 1 is a schematic diagram illustrating the general configuration of a surgical support system according to Embodiment 1. In laparoscopic surgery, instead of performing open surgery, multiple opening devices called troccas 10 are attached to the patient's abdominal wall, and instruments such as a laparoscope 11, energy treatment instruments 12, and forceps 13 are inserted into the patient's body through openings in the troccas 10. The surgeon performs procedures such as excising the affected area using the energy treatment instruments 12 while viewing images of the patient's body (surgical field images) captured by the laparoscope 11 in real time. The surgical instruments such as the laparoscope 11, energy treatment instruments 12, and forceps 13 are held by the surgeon or a robot. The term "surgeon" refers to medical professionals involved in laparoscopic surgery, including the surgeon, assistants, nurses, and physicians monitoring the surgery.
[0013] The laparoscope 11 includes an insertion section 11A that is inserted into the patient's body, an imaging device 11B built into the tip of the insertion section 11A, an operating section 11C provided at the rear end of the insertion section 11A, and a universal cord 11D for connecting to a camera control unit (CCU) 110 and a light source device 120.
[0014] The insertion section 11A of the laparoscope 11 is formed by a rigid tube. A curved section is provided at the tip of the rigid tube. The bending mechanism in the curved section is a well-known mechanism incorporated into general laparoscopes, and is configured to bend in four directions, for example, up, down, left, and right, by the pulling of an operating wire linked to the operation of the operating section 11C. Note that the laparoscope 11 is not limited to a flexible endoscope having a curved section as described above, but may also be a rigid endoscope without a curved section, or an imaging device without a curved section or rigid tube.
[0015] The imaging device 11B includes a solid-state image sensor such as a CMOS (Complementary Metal Oxide Semiconductor), a driver circuit equipped with a timing generator (TG), and an analog signal processing circuit (AFE). The driver circuit of the imaging device 11B acquires the RGB color signals output from the solid-state image sensor in synchronization with the clock signal output from the TG, and performs necessary processing such as noise reduction, amplification, and AD conversion in the AFE to generate digital image data. The driver circuit of the imaging device 11B transmits the generated image data to the CCU 110 via the universal code 11D.
[0016] The control unit 11C includes an angle lever and remote switches operated by the operator. The angle lever is a control device that receives input for bending the curved section. Instead of the angle lever, a bending operation knob, joystick, etc., may be provided. The remote switches include, for example, a toggle switch for switching the observed image between video display and still image display, and a zoom switch for enlarging or reducing the observed image. The remote switches may be assigned specific functions predetermined in advance, or functions set by the operator.
[0017] Furthermore, the operating unit 11C may have a built-in vibrator, such as a linear resonant actuator or a piezo actuator. If an event occurs that should be notified to the operator operating the laparoscope 11, the CCU 110 may activate the vibrator built into the operating unit 11C to vibrate the operating unit 11C and notify the operator of the occurrence of the event.
[0018] The insertion section 11A, operating section 11C, and universal cord 11D of the laparoscope 11 contain transmission cables for transmitting control signals output from the CCU 110 to the imaging device 11B and image data output from the imaging device 11B, as well as a light guide that guides illumination light emitted from the light source device 120 to the tip of the insertion section 11A. The illumination light emitted from the light source device 120 is guided to the tip of the insertion section 11A through the light guide and irradiated onto the surgical field through an illumination lens provided at the tip of the insertion section 11A. In this embodiment, the light source device 120 is described as an independent device, but the light source device 120 may be built into the CCU 110.
[0019] The CCU110 includes a control circuit that controls the operation of the imaging device 11B of the laparoscope 11, and an image processing circuit that processes image data from the imaging device 11B input via a universal code 11D. The control circuit includes a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc., and outputs control signals to the imaging device 11B in response to the operation of various switches on the CCU110 and the operation of the control unit 11C of the laparoscope 11, performing controls such as starting and stopping imaging and zooming. The image processing circuit includes a DSP (Digital Signal Processor) and image memory, etc., and performs appropriate processing such as color separation, color interpolation, gain correction, white balance adjustment, and gamma correction on the image data input via the universal code 11D. The CCU110 generates frame images for video from the processed image data and sequentially outputs each generated frame image to the information processing device 200 described later. The frame rate of the frame images is, for example, 30 FPS (Frames Per Second).
[0020] The CCU110 may generate video data compliant with predetermined standards such as NTSC (National Television System Committee), PAL (Phase Alternating Line), and DICOM (Digital Imaging and Communication in Medicine). By outputting the generated video data to the display device 130, the CCU110 can display surgical field images (video) on the display screen of the display device 130 in real time. The display device 130 is a monitor equipped with an LCD panel or an OLED (Electro-Luminescence) panel. The CCU110 may also output the generated video data to the recording device 140 and have the recording device 140 record the video data. The recording device 140 is equipped with a recording device such as an HDD (Hard Disk Drive) that records the video data output from the CCU110 along with an identifier to identify each surgery, the date and time of the surgery, the location of the surgery, the patient's name, the surgeon's name, etc.
[0021] The information processing device 200 acquires image data of the surgical field from the CCU 110 and inputs the acquired image data of the surgical field into multiple learning models, thereby executing calculations by each learning model. The information processing device 200 derives an integrated recognition result for the surgical field image from the calculation results of the multiple learning models and outputs information based on the derived recognition result.
[0022] Figure 2 is a block diagram illustrating the internal configuration of the information processing device 200. The information processing device 200 is a dedicated or general-purpose computer comprising a control unit 201, a storage unit 202, an operation unit 203, an input unit 204, a first arithmetic unit 205, a second arithmetic unit 206, an output unit 207, a communication unit 208, and the like. The information processing device 200 may be a computer installed in the operating room or a computer installed outside the operating room. The information processing device 200 may be a server installed in the hospital where laparoscopic surgery is performed or a server installed outside the hospital. The information processing device 200 is not limited to a single computer, but may be a computer system consisting of multiple computers and peripheral devices. The information processing device 200 may be a virtual machine virtually constructed by software.
[0023] The control unit 201 includes, for example, a CPU, ROM, and RAM. The ROM in the control unit 201 stores control programs that control the operation of each hardware component of the information processing device 200. The CPU in the control unit 201 executes the control programs stored in the ROM and various computer programs stored in the memory unit 202 (described later), thereby controlling the operation of each hardware component and making the entire device function as an information processing device according to this application. The RAM in the control unit 201 temporarily stores data used during the execution of calculations.
[0024] In this embodiment, the control unit 201 is configured to include a CPU, ROM, and RAM, but the configuration of the control unit 201 is arbitrary and can be any arithmetic circuit or control circuit that includes a GPU (Graphics Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), quantum processor, volatile or non-volatile memory, etc. Furthermore, the control unit 201 may also have functions such as a clock that outputs date and time information, a timer that measures the elapsed time from the time a measurement start instruction is given to the time a measurement end instruction is given, and a counter that counts numbers.
[0025] The storage unit 202 includes storage devices such as hard disks and flash memory. The storage unit 202 stores computer programs executed by the control unit 201, various data acquired from external sources, and various data generated within the device.
[0026] The computer programs stored in the memory unit 202 include a recognition processing program PG1 that causes the control unit 201 to execute processing for recognizing recognition targets included in the surgical field image, and a display processing program PG2 that causes the control unit 201 to execute processing for displaying information based on the recognition results on the display device 130. Note that the recognition processing program PG1 and the display processing program PG2 do not need to be independent computer programs and may be implemented as a single computer program. These programs are provided, for example, by a non-temporary recording medium M on which the computer programs are recorded in a readable format. The recording medium M is a portable memory such as a CD-ROM, USB memory, or SD (Secure Digital) card. The control unit 201 reads the desired computer program from the recording medium M using a reading device (not shown in the figure) and stores the read computer program in the memory unit 202. Alternatively, the computer programs may be provided by communication. In this case, the control unit 201 downloads the desired computer program via the communication unit 208 and stores the downloaded computer program in the memory unit 202.
[0027] The memory unit 202 includes a first learning model 310 and a second learning model 320. The first learning model 310 is a trained model that has been trained to output information about a first organ contained in a surgical field image in response to the input of a surgical field image. The first organ includes specific organs such as the esophagus, stomach, large intestine, pancreas, spleen, ureter, lung, prostate, uterus, gallbladder, liver, and vas deferens, or non-specific organs such as connective tissue, fat, nerves, blood vessels, muscles, and membranous structures. The first organ is not limited to a specific organ and may be any structure present in the body. Similarly, the second learning model 320 is a trained model that has been trained to output information about a second organ contained in a surgical field image in response to the input of a surgical field image. The second organ includes the specific or non-specific organs described above. The second organ is not limited to a specific organ and may be any structure present in the body. In the following explanation, if it is not necessary to distinguish between specific organs and non-specific organs, they will simply be referred to as organs.
[0028] In Embodiment 1, the second organ to be recognized has similar characteristics to the first organ, and an organ that may be mistaken for the first organ in the recognition process is selected. For example, if the recognition target by the first learning model 310 is loose connective tissue, then nerve tissue (or membrane structure) is selected as the recognition target by the second learning model 320. The combination of the first and second organs is not limited to (a) loose connective tissue and nerve tissue (or membrane structure), but may also be a combination of (b) fat and pancreas, (c) fat to be removed and fat to be preserved, (d) two of bleeding, hemorrhage marks, and blood vessels, (e) two of ureters, arteries, and membrane structures, (f) stomach and intestines, (g) liver and spleen, etc.
[0029] Embodiment 1 describes a case where the first organ is loose connective tissue and the second organ is nerve tissue. That is, the first learning model 310 is trained to output information about the loose connective tissue contained in the surgical field image in response to the input of the surgical field image. Such a first learning model 310 is generated by training it using an appropriate learning algorithm, with the surgical field image obtained by imaging the surgical field and the ground truth data showing the portion of the first organ (the loose connective tissue portion in Embodiment 1) within the surgical field image as training data. The ground truth data showing the portion of the first organ is generated by manual annotation by a specialist such as a physician. The same applies to the second learning model 320.
[0030] The first and second learning models 310 and 320 may be generated internally by the information processing device 200 or by an external server. In the latter case, the information processing device 200 can download the first and second learning models 310 and 320 generated by the external server via communication and store the downloaded first and second learning models 310 and 320 in the storage unit 202. The storage unit 202 stores definition information for the first and second learning models 310 and 320, including information about the layers that the first and second learning models 310 and 320 have, information about the nodes that constitute each layer, and parameters determined by learning, such as weight coefficients and biases between nodes.
[0031] The operation unit 203 is equipped with operating devices such as a keyboard, mouse, touch panel, contactless panel, stylus pen, and microphone for voice input. The operation unit 203 receives operations from the operator or other user and outputs information related to the received operations to the control unit 201. The control unit 201 performs appropriate processing according to the operation information input from the operation unit 203. In this embodiment, the information processing device 200 is configured to include the operation unit 203, but it may also be configured to receive operations through various devices such as an externally connected CCU 110.
[0032] The input unit 204 is equipped with a connection interface for connecting an input device. In this embodiment, the input device connected to the input unit 204 is the CCU 110. The input unit 204 receives image data of the surgical field image acquired by the laparoscope 11 and processed by the CCU 110. The input unit 204 outputs the input image data to the control unit 201. The control unit 201 may also store the image data acquired from the input unit 204 in the storage unit 202.
[0033] In this embodiment, a configuration is described in which image data of the surgical field is acquired from the CCU 110 through the input unit 204. However, image data of the surgical field may be acquired directly from the laparoscope 11, or image data of the surgical field may be acquired from an image processing device (not shown) that is detachably attached to the laparoscope 11. In addition, the information processing device 200 may acquire image data of the surgical field recorded in the recording device 140.
[0034] The first arithmetic unit 205 includes a processor and memory. An example of a processor is a GPU (Graphics Processing Unit), and an example of memory is VRAM (Video RAM). When a surgical field image is input to the first arithmetic unit 205, it performs calculations using the first learning model 310 with its built-in processor and outputs the calculation results to the control unit 201. In addition, in response to instructions from the control unit 201, the first arithmetic unit 205 draws the image to be displayed on the display device 130 onto its built-in memory and outputs it to the display device 130 via the output unit 207, thereby displaying the desired image on the display device 130.
[0035] The second arithmetic unit 206, like the first arithmetic unit 205, is equipped with a processor, memory, and the like. The second arithmetic unit 206 may be equivalent to the first arithmetic unit 205, or it may have lower processing power than the first arithmetic unit 205. When a surgical field image is input to the second arithmetic unit 206, it performs calculations using the second learning model 320 with its built-in processor and outputs the calculation results to the control unit 201. In addition, in response to instructions from the control unit 201, the second arithmetic unit 206 draws the image to be displayed on the display device 130 onto its built-in memory and outputs it to the display device 130 via the output unit 207, thereby displaying the desired image on the display device 130.
[0036] The output unit 207 is equipped with a connection interface for connecting an output device. In this embodiment, the output device connected to the output unit 207 is the display device 130. When the control unit 201 generates information to be communicated to the operator, such as an integrated recognition result derived from the calculation results of the learning models 310 and 320, it outputs the generated information from the output unit 207 to the display device 130, thereby displaying the information on the display device 130. Alternatively, the output unit 207 may output the information to be communicated to the operator, such as voice or sound.
[0037] The communication unit 208 is equipped with a communication interface for sending and receiving various types of data. The communication interface provided by the communication unit 208 is a communication interface compliant with wired or wireless communication standards used in Ethernet (registered trademark) and WiFi (registered trademark). When data to be transmitted is input from the control unit 201, the communication unit 208 transmits the data to the specified destination. In addition, when the communication unit 208 receives data transmitted from an external device, it outputs the received data to the control unit 201.
[0038] Next, we will explain the surgical field image input to the information processing device 200. Figure 3 is a schematic diagram showing an example of a surgical field image. In this embodiment, the surgical field image is an image obtained by imaging the inside of the patient's abdominal cavity with a laparoscope 11. The surgical field image does not need to be the raw image output by the imaging device 11B of the laparoscope 11; it is sufficient if it is an image (frame image) that has been processed by a CCU 110 or the like.
[0039] The surgical field imaged by the laparoscope 11 includes tissues that make up specific organs, blood vessels, nerves, connective tissue between tissues, tissues containing lesions such as tumors, and tissues such as membranes and layers covering the tissues. The surgeon, while understanding the relationships of these anatomical structures, uses surgical instruments such as energy treatment instruments 12 and forceps 13 to dissect the tissues containing the lesions. Figure 3 shows an example of a surgical field image, illustrating a scene where the membrane covering organ 501 is pulled using forceps 13 to apply appropriate tension to the loose connective tissue 502, and the loose connective tissue 502 is being dissected using the energy treatment instrument 12. Loose connective tissue 502 is fibrous connective tissue that fills the spaces between tissues and organs, and refers to tissue with a relatively small amount of fibers (elastic fibers) that make up the tissue. Loose connective tissue 502 is dissected as needed when exposing organ 501 or when resecting lesions. In the example shown in Figure 3, the loose connective tissue 502 runs vertically (in the direction of the white arrow in the figure), and the nerve tissue 503 runs horizontally (in the direction of the black arrow in the figure) so as to intersect with the loose connective tissue 502.
[0040] Generally, loose connective tissue and nerve tissue appearing in surgical field images are often difficult to distinguish visually because both are whitish in color and run in a linear fashion. Loose connective tissue is removed as needed, but if nerves are damaged during the traction or removal process, functional impairment may occur postoperatively. For example, damage to the inferior gastric nerve in colorectal surgery can cause urinary dysfunction. Similarly, damage to the recurrent laryngeal nerve in esophagectomy or lung resection can cause swallowing difficulties. Therefore, providing surgeons with information on the recognition results of loose connective tissue and nerve tissue would be useful for them.
[0041] The information processing device 200 according to Embodiment 1 performs calculations by the first learning model 310 and calculations by the second learning model 320 on the same surgical field image, derives an integrated recognition result for the surgical field image from the results of the two calculations, and outputs information based on the derived recognition result, thereby providing the surgeon with information about loose connective tissue and nerve tissue.
[0042] The configurations of the first learning model 310 and the second learning model 320 will be described below. Figure 4 is a schematic diagram showing an example configuration of the first learning model 310. The first learning model 310 is a learning model for image segmentation and is constructed using a neural network equipped with convolutional layers, such as SegNet. The learning model 310 is not limited to SegNet; it may be constructed using any neural network capable of image segmentation, such as FCN (Fully Convolutional Network), U-Net (U-Shaped Network), or PSPNet (Pyramid Scene Parsing Network). Alternatively, the learning model 310 may be constructed using a neural network for object detection, such as YOLO (You Only Look Once) or SSD (Single Shot Multi-Box Detector), instead of a neural network for image segmentation.
[0043] The calculations performed by the first learning model 310 are carried out in the first calculation unit 205. When a surgical field image is input, the first calculation unit 205 performs calculations according to the definition information of the first learning model 310, which includes the learned parameters.
[0044] The first learning model 310 includes, for example, an encoder 311, a decoder 312, and a softmax layer 313. The encoder 311 is constructed by alternating convolutional layers and pooling layers. The convolutional layers are multilayered into 2 to 3 layers. In the example in Figure 4, the convolutional layers are shown without hatching, while the pooling layers are shown with hatching.
[0045] In a convolutional layer, the input data is convolved with a filter of a predetermined size (e.g., 3x3 or 5x5). Specifically, the input value at the position corresponding to each element of the filter is multiplied by a weight coefficient pre-set for the filter, and a linear sum of these element-wise multiplications is calculated. The output of the convolutional layer is obtained by adding a set bias to the calculated linear sum. The result of the convolutional operation may be transformed by an activation function. For example, ReLU (Rectified Linear Unit) can be used as an activation function. The output of the convolutional layer represents a feature map in which the features of the input data have been extracted.
[0046] In the pooling layer, local statistics are calculated for the feature map output from the convolutional layer, which is a higher layer connected to the input. Specifically, a window of a predetermined size (e.g., 2x2, 3x3) corresponding to the position of the higher layer is set, and local statistics are calculated from the input values within the window. For example, the maximum value can be used as the statistics. The size of the feature map output from the pooling layer is reduced (downsampled) according to the size of the window. The example in Figure 4 shows that in the encoder 311, the calculations in the convolutional layer and the calculations in the pooling layer are sequentially repeated to sequentially downsample a 224x224 pixel input image to 112x112, 56x56, 28x28, ..., 1x1 feature maps.
[0047] The output of the encoder 311 (a 1x1 feature map in the example in Figure 4) is input to the decoder 312. The decoder 312 is constructed by alternating inverse convolutional layers and inverse pooling layers. The inverse convolutional layers are multilayered into 2 to 3 layers. In the example in Figure 4, the inverse convolutional layers are shown without hatching, while the inverse pooling layers are shown with hatching.
[0048] In the deconvolution layer, a deconvolution operation is performed on the input feature map. The deconvolution operation is an operation that reconstructs the feature map before the convolution operation, based on the assumption that the input feature map is the result of a convolution operation using a specific filter. In this operation, when the specific filter is represented by a matrix, the output feature map is generated by calculating the product of the transpose of this matrix and the input feature map. The result of the deconvolution layer operation may also be transformed by an activation function such as ReLU, as described above.
[0049] The inverse pooling layer of the decoder 312 is individually mapped one-to-one with the pooling layer of the encoder 311, and the mapped pairs have substantially the same size. The inverse pooling layer expands (upsamples) the size of the feature map that has been downsampled in the pooling layer of the encoder 311. The example in Figure 4 shows that the decoder 312 sequentially upsamples to 1×1, 7×7, 14×14, ..., 224×224 feature maps by sequentially repeating operations in the convolutional layer and the pooling layer.
[0050] The output of the decoder 312 (a 224x224 feature map in the example of Figure 4) is input to the softmax layer 313. The softmax layer 313 outputs the probability of a label identifying a region at each position (pixel) by applying a softmax function to the input values from the inverse convolutional layer connected to the input side. The first learning model 310 according to Embodiment 1 only needs to output from the softmax layer 313 the probability indicating whether or not each pixel corresponds to loose connective tissue in response to the input of a surgical field image. The calculation result from the first learning model 310 is output to the control unit 201.
[0051] By extracting pixels whose label probability output from the softmax layer 313 is above a threshold (e.g., 60% or more), an image showing the recognition result of the loose connective tissue portion (recognition image) is obtained. The first processing unit 205 may display the recognition result of the first learning model 310 on the display device 130 by drawing the recognition image of the loose connective tissue portion to the built-in memory (VRAM) and outputting it to the display device 130 via the output unit 207. The recognition image is an image of the same size as the surgical field image and is generated as an image in which a specific color is assigned to the pixels recognized as loose connective tissue. The color assigned to the pixels of loose connective tissue is preferably a color that does not exist inside the human body so that it can be distinguished from organs and blood vessels. Colors that do not exist inside the human body are, for example, cool colors (blues) such as blue or light blue. In addition, information representing transparency is added to each pixel that makes up the recognition image, and an opaque value is set for pixels recognized as loose connective tissue, and a transparent value is set for other pixels. When the recognition image generated in this way is overlaid on the surgical field image, the loose connective tissue can be displayed on the surgical field image as a structure with a specific color.
[0052] In the example shown in Figure 4, a 224x224 pixel image is used as the input image to the first learning model 310. However, the size of the input image is not limited to the above and can be set appropriately according to the processing power of the information processing device 200, the size of the surgical field image obtained from the laparoscope 11, etc. Furthermore, the input image to the first learning model 310 does not need to be the entire surgical field image obtained from the laparoscope 11; it may be a partial image generated by cutting out the area of interest from the surgical field image. Since the area of interest, which includes the target of treatment, is often located near the center of the surgical field image, for example, a partial image may be used in which the area near the center of the surgical field image is cut out in a rectangular shape to about half the size of the original. By reducing the size of the image input to the first learning model 310, it is possible to increase the processing speed while improving recognition accuracy.
[0053] Figure 5 is a schematic diagram showing an example configuration of the second learning model 320. The second learning model 320 includes an encoder 321, a decoder 322, and a softmax layer 323, and is configured to output information about the nerve tissue portion contained in the surgical field image in response to the input of the surgical field image. The configuration of the encoder 321, decoder 322, and softmax layer 323 of the second learning model 320 is the same as that of the first learning model 310, so a detailed explanation will be omitted.
[0054] The calculations performed by the second learning model 320 are carried out in the second calculation unit 206. When a surgical field image is input, the second calculation unit 206 performs calculations according to the definition information of the second learning model 320, which includes the learned parameters. In the first embodiment, the second learning model 320 only needs to output a probability from the softmax layer 323 indicating whether or not each pixel corresponds to nerve tissue in response to the input surgical field image. The calculation results from the second learning model 320 are output to the control unit 201.
[0055] By extracting pixels where the probability of the label output from the softmax layer 323 is above a threshold (e.g., 60% or more), an image showing the recognition result of the nerve tissue portion (recognition image) is obtained. The second processing unit 206 may draw the recognition image of the nerve tissue portion to its built-in memory (VRAM) and output it to the display device 130 via the output unit 207, thereby displaying the recognition result of the second learning model 320 on the display device 130. The structure of the recognition image showing nerve tissue is the same as that of loose connective tissue, but it is preferable that the color assigned to the pixels of nerve tissue is a color that can be distinguished from loose connective tissue (e.g., green or yellow).
[0056] The operation of the information processing device 200 will be described below. Figure 6 is a flowchart showing the processing procedure performed by the information processing device 200 according to Embodiment 1. The control unit 201 of the information processing device 200 reads and executes the recognition processing program PG1 and the display processing program PG2 from the storage unit 202, and performs the following processing procedure. When laparoscopic surgery is started, the surgical field image obtained by imaging the surgical field with the imaging device 11B of the laparoscope 11 is output to the CCU 110 via the universal code 11D as needed. The control unit 201 of the information processing device 200 acquires the frame-by-frame surgical field image output from the CCU 110 from the input unit 204 (step S101). The control unit 201 performs the following processing each time it acquires a frame-by-frame surgical field image.
[0057] The control unit 201 sends frame-by-frame surgical field images acquired through the input unit 204 to the first calculation unit 205 and the second calculation unit 206, and also gives the first calculation unit 205 and the second calculation unit 206 an instruction to start calculations (step S102).
[0058] When the control unit 201 issues an instruction to start the calculation, the first calculation unit 205 executes the calculation performed by the first learning model 310 (step S103). Specifically, the first calculation unit 205 performs calculations by an encoder 311 that generates a feature map from the input surgical field image and sequentially downsamples the generated feature map, calculations by a decoder 312 that sequentially upsamples the feature map input from the encoder 311, and calculations by a softmax layer 313 that identifies each pixel of the feature map finally obtained from the decoder 312. The first calculation unit 205 outputs the calculation results from the learning model 310 to the control unit 201 (step S104).
[0059] When the control unit 201 issues an instruction to start the calculation, the second calculation unit 206 executes the calculation performed by the second learning model 320 (step S105). Specifically, the second calculation unit 206 generates a feature map from the input surgical field image, performs calculations using the encoder 321 to sequentially downsample the generated feature map, performs calculations using the decoder 322 to sequentially upsample the feature map input from the encoder 321, and performs calculations using the softmax layer 323 to identify each pixel of the feature map finally obtained from the decoder 322. The second calculation unit 206 outputs the calculation results from the learning model 320 to the control unit 201 (step S106).
[0060] In the flowchart of Figure 6, for convenience, the procedure is shown as performing calculations by the first calculation unit 205 first, followed by calculations by the second calculation unit 206. However, it is preferable that the calculations by the first calculation unit 205 and the second calculation unit 206 be performed simultaneously.
[0061] The control unit 201 derives an integrated recognition result for the surgical field image based on the calculation results from the first learning model 310 and the calculation results from the second learning model 320. Specifically, the control unit 201 performs the following processes.
[0062] The control unit 201 refers to the calculation results from the first learning model 310 and performs the recognition process for loose connective tissue (step S107). The control unit 201 can recognize loose connective tissue included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 313 of the first learning model 310 is above a threshold (e.g., 60% or more).
[0063] The control unit 201 refers to the calculation results from the second learning model 320 and performs the neural tissue recognition process (step S108). The control unit 201 can recognize the neural tissue included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 323 of the second learning model 320 is above a threshold (e.g., 60% or more).
[0064] The control unit 201 determines whether the recognition result for loose connective tissue and the recognition result for nerve tissue overlap (step S109). In this step, it checks whether a specific structure included in the surgical field image is recognized as loose connective tissue on one hand and as nerve tissue on the other. Specifically, the control unit 201 determines that the recognition results overlap if a single pixel in the surgical field image is recognized as loose connective tissue on one hand and as nerve tissue on the other. Alternatively, the control unit 201 may compare the area in the surgical field image recognized as loose connective tissue with the area in the surgical field image recognized as nerve tissue and determine whether the recognition results overlap. For example, if the overlap of the two is greater than a predetermined percentage (e.g., 40% or more) in terms of area ratio, it may be determined that the recognition results overlap, and if it is less than the predetermined percentage, it may be determined that the recognition results do not overlap.
[0065] If the control unit 201 determines that there are no duplicate recognition results (S109: NO), it outputs the recognition results for loose connective tissue and nerve tissue (step S110). Specifically, the control unit 201 instructs the first calculation unit 205 to superimpose the recognized image of loose connective tissue onto the surgical field image, and instructs the second calculation unit 206 to superimpose the recognized image of nerve tissue onto the surgical field image. The first calculation unit 205 and the second calculation unit 206 draw the recognized images of loose connective tissue and nerve tissue onto their built-in VRAMs in response to instructions from the control unit 201, and output them to the display device 130 via the output unit 207, thereby displaying the recognized images of loose connective tissue and nerve tissue superimposed on the surgical field image.
[0066] Figure 7 is a schematic diagram showing an example of the display of the recognition results in Embodiment 1. In the display example in Figure 7, for the sake of drawing creation, the parts recognized as loose connective tissue 502 are shown as hatched areas, and the parts recognized as nerve tissue 503 are shown as separate hatched areas. In reality, pixels recognized as loose connective tissue 502 are displayed in, for example, a blue color, and pixels recognized as nerve tissue 503 are displayed in, for example, a green color. By viewing the recognition results displayed on the display device 130, the operator can distinguish between loose connective tissue 502 and nerve tissue 503, and for example, while being aware of the presence of nerve tissue 503 that should not be damaged, the operator can peel off the loose connective tissue 502 using the energy treatment instrument 12.
[0067] In this embodiment, the system is configured to display both loose connective tissue recognized based on the calculation results of the first calculation unit 205 and nerve tissue recognized based on the calculation results of the second calculation unit 206. However, it is also possible to configure the system to display only one of them. Furthermore, the tissue to be displayed may be selected by the operator, or switched by the operator's operation.
[0068] In step S109 of the flowchart shown in Figure 6, if it is determined that the recognition results are duplicates (S109: YES), the control unit 201 outputs warning information indicating that similar structures have been recognized (step S111).
[0069] Figure 8 is a schematic diagram showing an example of warning information display in Embodiment 1. Figure 8 shows an example of warning information display when a structure 504 included in the surgical field image is recognized as loose connective tissue based on the calculation result of the first calculation unit 205 and as nerve tissue based on the calculation result of the second calculation unit 206. The first learning model 310 used by the first calculation unit 205 is trained to output information about loose connective tissue in response to the input of a surgical field image, and the second learning model 320 used by the second calculation unit 206 is trained to output information about nerve tissue in response to the input of a surgical field image. However, since loose connective tissue and nerve tissue have similar external characteristics, if recognition processing is performed independently using the two learning models 310 and 320, there is a possibility that the recognition results will overlap. If the recognition results overlap, presenting only one of the recognition results to the surgeon may lead to the misidentification of what is actually loose connective tissue as nerve tissue, or vice versa. If the recognition results are duplicated, the control unit 201 can prompt the operator to confirm by displaying warning information as shown in Figure 8.
[0070] In the example shown in Figure 8, the control unit 201 is configured to display warning text information superimposed on the surgical field image. However, the warning text information may be displayed outside the display area of the surgical field image, or on another display device (not shown). Instead of displaying warning text information, the control unit 201 may display a warning graphic or provide a warning by outputting audio or sound.
[0071] (Extreme Variation 1-1) The control unit 201 of the information processing device 200 may stop outputting information based on the recognition results if the recognition results of loose connective tissue by the first learning model 310 and the recognition results of neural tissue by the second learning model 320 overlap.
[0072] (Variations 1-2) The control unit 201 of the information processing device 200 may, if the recognition result of loose connective tissue by the first learning model 310 and the recognition result of nerve tissue by the second learning model 320 overlap, select the recognition result with the higher confidence level and output information based on the selected recognition result. The confidence level of the recognition result by the first learning model 310 is calculated based on the probability output from the softmax layer 313. For example, the control unit 201 can calculate the confidence level by finding the average of the probability values for each pixel recognized as loose connective tissue. The same applies to the confidence level of the recognition result by the second learning model 320. For example, if the structure 504 shown in Figure 8 is recognized as loose connective tissue with a 95% confidence level by the first learning model 310, and the same structure 504 is recognized as nerve tissue with a 62% confidence level by the second learning model 320, the control unit 201 can present the operator with the recognition result that this structure 504 is loose connective tissue.
[0073] (Variations 1-3) If the recognition result of loose connective tissue by the first learning model 310 and the recognition result of nerve tissue by the second learning model 320 overlap, the control unit 201 of the information processing device 200 may derive a recognition result according to the confidence level and display the recognition result in a display manner according to the confidence level. Figure 9 is a schematic diagram showing an example of displaying the recognition result according to the confidence level. For example, if the first learning model 310 recognizes a structure 505 included in the surgical field image as loose connective tissue with a confidence level of 95%, but the second learning model 320 does not recognize the same structure 505 as nerve tissue, the control unit 201 may color this structure 505 in a blue color (black in the diagram) and present it to the surgeon. Similarly, if the second learning model 320 recognizes structure 506 in the surgical field image as nerve tissue with 90% confidence, and the first learning model 310 recognizes the same structure 506 as loose connective tissue, the control unit 201 presents this structure 506 to the surgeon, colored, for example, a greenish color (white in the diagram). On the other hand, if the first learning model 310 recognizes structure 507 in the surgical field image as loose connective tissue with 60% confidence, and the second learning model 320 recognizes the same structure 507 as nerve tissue with 60% confidence, the control unit 201 presents this structure 507 to the surgeon, colored, for example, an intermediate color between blue and green (gray in the diagram). Instead of changing the color according to the confidence level, a configuration that changes saturation, transparency, etc., may be adopted.
[0074] (Variations 1-4) If both structures recognized as loose connective tissue and structures recognized as nerve tissue are present in the surgical field image, the control unit 201 of the information processing device 200 may derive information such as the appropriate positional relationship between the two structures, the distance to feature points, the distance to other structures, and the area of other structures from the relationship between the two structures.
[0075] As described above, in Embodiment 1, an integrated recognition result for organs included in the surgical field can be obtained based on the calculation results of the first learning model 310 and the calculation results of the second learning model 320. The information processing device 200 issues a warning or stops outputting information if the recognition results for similar structures are duplicated, thereby avoiding presenting the surgeon with results that may be incorrect.
[0076] (Embodiment 2) Embodiment 2 describes a configuration that combines organ recognition and event recognition to derive an integrated recognition result.
[0077] The information processing device 200 according to Embodiment 2 includes a first learning model 310 for recognizing organs and a third learning model 330 for recognizing events. The organs recognized by the first learning model 310 are not limited to loose connective tissue, but may be any pre-defined organs. The events recognized by the third learning model 330 are events such as bleeding, injury, and pulsation. The other components of the information processing device 200 are the same as those in Embodiment 1, so their description is omitted.
[0078] Figure 10 is a schematic diagram showing an example configuration of the third learning model 330. The third learning model 330 comprises an encoder 331, a decoder 332, and a softmax layer 333, and is configured to output information about events occurring within a surgical field image in response to an input of a surgical field image. The information about events output by the third learning model 330 includes information about events such as bleeding, injury (burn marks from the energy treatment instrument 12), and pulsation. The third learning model 330 is not limited to a learning model for image segmentation or object detection, but may also be a learning model based on CNN (Convolutional Neural Networks), RNN (Recurrent Neural Networks), LSTM (Long Short Term Memory), GAN (Generative Adversarial Network), etc.
[0079] The calculations performed by the third learning model 330 are carried out in the second calculation unit 206. When a surgical field image is input, the second calculation unit 206 performs calculations according to the definition information of the third learning model 330, which includes the learned parameters. The third learning model 330 only needs to output a probability from the softmax layer 333 indicating whether or not an event has occurred for the input surgical field image. The calculation results from the third learning model 330 are output to the control unit 201. The control unit 201 determines that an event has occurred in the surgical field image if the probability of the label output from the softmax layer 333 is above a threshold (for example, 60% or more). The control unit 201 may determine whether or not an event has occurred on a pixel-by-pixel basis in the surgical field image, or it may determine whether or not an event has occurred on a unit basis in the surgical field image.
[0080] Figure 11 is a flowchart showing the processing steps performed by the information processing device 200 according to Embodiment 2. The information processing device 200 performs steps S201 to S206 in the same manner as in Embodiment 1 each time it acquires a surgical field image. The control unit 201 of the information processing device 200 acquires the calculation results from the first learning model 310 and the calculation results from the third learning model 330, and derives an integrated recognition result for the surgical field image based on these calculation results. Specifically, the control unit 201 performs the following processing.
[0081] The control unit 201 refers to the calculation results from the first learning model 310 and performs organ recognition processing (step S207). The control unit 201 can recognize organs included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 313 of the first learning model 310 is above a threshold (e.g., 60% or more).
[0082] The control unit 201 refers to the calculation results from the third learning model 330 and performs event recognition processing (step S208). The control unit 201 can determine whether an event has occurred for each pixel by extracting pixels where the probability of the label output from the softmax layer 333 of the third learning model 330 is above a threshold (for example, 60% or more).
[0083] The control unit 201 determines whether or not an event has occurred in the organ recognized in step S207 (step S209). The control unit 201 compares the pixel recognized as an organ in step S207 with the pixel recognized as having had an event in step S208, and if they match, it determines that an event has occurred in the recognized organ.
[0084] If the control unit 201 determines that no event has occurred in the recognized organ (S209: NO), it terminates the processing according to this flowchart. The control unit 201 may also display the recognition results of organs individually without associating them with events, or display the recognition results of events individually without associating them with organs.
[0085] If the control unit 201 determines that an event has occurred in a recognized organ (S209: YES), it outputs information about the organ where the event occurred (step S210). The control unit 201 may, for example, display the name of the organ where the event occurred as text information superimposed on the surgical field image. Instead of superimposing the information on the surgical field image, the name of the organ where the event occurred may be displayed outside the surgical field image, or it may be output by sound or voice.
[0086] Figure 12 is a schematic diagram showing an example of the display of the recognition result in Embodiment 2. The display example in Figure 12 shows an example in which text information indicating that bleeding has occurred from the surface of the stomach is superimposed on the surgical field image. The control unit 201 may display information about organs damaged by the energy treatment instrument 12, etc., not limited to organs where bleeding has occurred, or it may display information about organs where pulsation is occurring. When displaying information about organs such as blood vessels where pulsation is occurring, for example, the organ may be displayed flashing in synchronization with the pulsation. The synchronization does not necessarily have to perfectly match the pulsation, and a periodic display close to the pulsation may also be used.
[0087] When the control unit 201 detects bleeding from a specific organ (e.g., a vital blood vessel), it may acquire data on the patient's vital signs (pulse, blood pressure, respiration, body temperature), which are constantly detected by sensors not shown in the diagram, and display the acquired data on the display device 130. Furthermore, when the control unit 201 detects bleeding from a specific organ (e.g., a vital blood vessel), it may notify an external device via the communication unit 208. The external device to which the notification is received may be a terminal carried by an anesthesiologist, or it may be an in-hospital server that centrally manages events within the hospital.
[0088] Furthermore, if the control unit 201 detects bleeding from an organ, it may change the threshold used for organ recognition or stop organ recognition altogether. In addition, if the control unit 201 detects bleeding from an organ, it may switch to an improved learning model (not shown) for bleeding and continue organ recognition using this learning model, as organ recognition may become difficult in such cases.
[0089] Furthermore, if there is a risk of anemia due to bleeding, the control unit 201 may automatically estimate the amount or rate of bleeding and suggest a blood transfusion. The control unit 201 can estimate the amount of bleeding by calculating the bleeding area on the image, and can estimate the rate of bleeding by calculating the change in the bleeding area over time.
[0090] As described above, in Embodiment 2, an integrated recognition result obtained by combining the learning model 310 for organ recognition and the learning model 320 for event recognition can be presented to the operator.
[0091] (Embodiment 3) Embodiment 3 describes a configuration that combines organ recognition and device recognition to derive an integrated recognition result.
[0092] The information processing device 200 according to Embodiment 3 includes a first learning model 310 for recognizing organs and a fourth learning model 340 for recognizing devices. The organs recognized by the first learning model 310 are not limited to loose connective tissue, but may be any pre-defined organs. The devices recognized by the fourth learning model 340 are surgical instruments used during surgery, such as energy treatment instruments 12 and forceps 13. The other components of the information processing device 200 are the same as those in Embodiment 1, so their description is omitted.
[0093] Figure 13 is a schematic diagram showing an example configuration of the fourth learning model 340. The fourth learning model 340 comprises an encoder 341, a decoder 342, and a softmax layer 343, and is configured to output information about devices included in a surgical field image in response to an input of a surgical field image. The device information output by the fourth learning model 340 is information about surgical instruments used during surgery, such as energy treatment instruments 12 and forceps 13. The fourth learning model 340 is not limited to a learning model for image segmentation or object detection, but may also be a learning model based on CNN, RNN, LSTM, GAN, etc. Multiple fourth learning models 340 may be provided depending on the type of device.
[0094] The calculations performed by the fourth learning model 340 are carried out in the second calculation unit 206. When a surgical field image is input, the second calculation unit 206 performs calculations according to the definition information of the fourth learning model 340, which includes the learned parameters. The fourth learning model 340 only needs to output a probability from the softmax layer 343 indicating whether or not each pixel corresponds to a specific device in response to the input surgical field image. The calculation results from the fourth learning model 340 are output to the control unit 201. If the probability of the label output from the softmax layer 343 is above a threshold (for example, 60% or more), the control unit 201 determines that a specific device included in the surgical field image has been recognized.
[0095] Figure 14 is a flowchart showing the processing steps performed by the information processing device 200 according to Embodiment 3. The information processing device 200 performs steps S301 to S306 in the same manner as in Embodiment 1 each time it acquires a surgical field image. The control unit 201 of the information processing device 200 acquires the calculation results from the first learning model 310 and the calculation results from the fourth learning model 340, and derives an integrated recognition result for the surgical field image based on these calculation results. Specifically, the control unit 201 performs the following processing. It is assumed that the storage unit 202 stores information on organs and devices recognized in the most recent (for example, one frame prior) surgical field image.
[0096] The control unit 201 refers to the calculation results from the first learning model 310 and performs organ recognition processing (step S307). The control unit 201 can recognize organs included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 313 of the first learning model 310 is above a threshold (e.g., 60% or more).
[0097] The control unit 201 refers to the calculation results from the fourth learning model 340 and performs device recognition processing (step S308). The control unit 201 can recognize devices included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 343 of the fourth learning model 340 is above a threshold (e.g., 60% or more).
[0098] The control unit 201 determines whether the device is moving on the organ recognized in step S307 (step S309). The control unit 201 reads the organ and device information recognized by the most recent surgical field image from the storage unit 202, compares the read organ and device information with the newly recognized organ and device information, and detects a change in relative position to determine whether the device is moving on the organ.
[0099] If the control unit 201 determines that the device is not moving on the organ (S309: NO), it terminates the processing according to this flowchart. The control unit 201 may display the organ recognition result independently of the device, and may also display the device recognition result independently of the organ.
[0100] If the control unit 201 determines that the device is moving on an organ (S309: YES), it generates display data with a modified display mode for the organ and outputs it to the display device 130 via the output unit 207 (step S310). The control unit 201 may change the display mode by changing the display color, saturation, and transparency of the organ, or by making the organ blink. The display device 130 displays the organ with the modified display mode. When a device is moving on an organ, there is a possibility that the information processing device 200 may misidentify the organ, so changing the display mode prompts the surgeon to make a judgment. The control unit 201 may also instruct the first calculation unit 205 to change the display mode, and the display mode may be changed by the processing of the first calculation unit 205.
[0101] (Variation 3-1) The control unit 201 of the information processing device 200 may stop the organ recognition process if it determines that the device is moving (not stopped) on the organ. If the organ recognition process is stopped, the control unit 201 may continue the device recognition process and resume the organ recognition process when it determines that the device has stopped.
[0102] (Variation 3-2) The control unit 201 of the information processing device 200 may stop outputting the organ recognition result if it determines that the device is moving (not stopped) on the organ. In this case, the calculations by the first learning model 310 and the fourth learning model 340, the organ recognition process based on the calculation result of the first learning model 310, and the device recognition process based on the calculation result of the fourth learning model 340 are performed continuously, and the display of the recognition image showing the organ recognition result is stopped. The control unit 201 should resume output processing when it determines that the device has stopped.
[0103] (Modified example 3-3) The control unit 201 of the information processing device 200 may derive information about the organ being processed by the device based on the organ recognition process and the device recognition process. The control unit 201 can derive information about the organ being processed by the device by comparing the pixels recognized as organs in step S307 with the pixels recognized as devices in step S308. The control unit 201 may output the information about the organ being processed by the device and display it, for example, on the display device 130.
[0104] (Modification 3-4) The control unit 201 of the information processing device 200 may acquire dimensional information of the recognized device and derive dimensional information of the recognized organ based on the acquired dimensional information of the device.
[0105] Figure 15 is a flowchart showing the procedure for deriving dimensional information. The control unit 201 acquires the dimensional information of the recognized device (step S321). The device's dimensional information may be pre-stored in the storage unit 202 of the information processing device 200, or it may be stored in an external device. In the former case, the control unit 201 can acquire the dimensional information by reading the desired information from the storage unit 202, and in the latter case, the control unit 201 can acquire the dimensional information by accessing the external device. Note that the dimensional information does not need to be the dimensions of the entire device, but may be the dimensions of a part of the device (for example, the cutting edge).
[0106] The control unit 201 calculates the ratio between the dimensions of the image information of the device portion indicated by the acquired dimensional information and the dimensions of the recognized organ on the image (step S322).
[0107] The control unit 201 calculates the dimensions of the organ based on the device dimensional information acquired in step S321 and the dimensional ratio calculated in step S322 (step S323). The control unit 201 outputs the calculated organ dimensional information and may display it on, for example, the display device 130.
[0108] (Modifications 3-5) The control unit 201 of the information processing device 200 may derive information about organs damaged by a device based on organ recognition processing and device recognition processing. For example, if the control unit 201 recognizes a device on an organ and determines that a part of the organ has discolored, it determines that organ damage caused by a device has been detected. Devices on organs are recognized by a procedure similar to the procedure shown in the flowchart of Figure 14. Discoloration of organs is recognized by changes in pixel values over time. When the control unit 201 detects organ damage caused by a device, it outputs information to that effect and displays it, for example, on the display device 130.
[0109] (Variations 3-6) The control unit 201 of the information processing device 200 may derive information indicating whether the device being used on an organ is appropriate, based on organ recognition processing and device recognition processing. The storage unit 202 of the information processing device 200 shall have a definition table that defines the relationship between the type of organ and the devices that can be used (or should not be used) for each organ. This definition table defines, for example, that sharp forceps should not be used on the intestine. The control unit 201 recognizes organs and devices from the surgical field image and determines whether the device being used on the organ is appropriate by referring to the definition table described above. If it is determined that the device is inappropriate (for example, if sharp forceps are being used on the intestine), the control unit 201 outputs information indicating that the wrong device is being used and displays it, for example, on the display device 130. The control unit 201 may also issue a warning by voice or warning sound.
[0110] (Variations 3-7) When the control unit 201 of the information processing device 200 recognizes a device, it may output operation support information for that device and display it, for example, on the display device 130. The device operation support information is the device's user manual and can be stored in the storage unit 202 of the information processing device 200 or in an external device.
[0111] As described above, in Embodiment 3, an integrated recognition result obtained by combining the learning model 310 for organ recognition and the learning model 340 for device recognition can be presented to the surgeon.
[0112] (Embodiment 4) Embodiment 4 describes a configuration that combines organ recognition and scene recognition to derive an integrated recognition result.
[0113] The information processing device 200 according to Embodiment 4 includes a first learning model 310 for recognizing organs and a fifth learning model 350 for recognizing scenes. The organs recognized by the first learning model 310 are not limited to loose connective tissue, but may be any pre-defined organs. The first learning model 310 may have separate models for each type of organ to accommodate various organs. The scenes recognized by the fifth learning model 350 are, for example, characteristic scenes that show characteristic scenes of surgery. The other configurations of the information processing device 200 are the same as in Embodiment 1, so their description is omitted.
[0114] Figure 16 is a schematic diagram showing an example configuration of the fifth learning model 350. The fifth learning model 350 comprises an input layer 351, an intermediate layer 352, and an output layer 353, and is configured to output information about the scene shown in the surgical field image in response to the input of a surgical field image. The information about the scene output by the fifth learning model 350 includes the probability that the scene contains specific organs such as blood vessels, important nerves, and specific organs (such as ureters and spleen), the probability that the scene involves characteristic surgical procedures such as vascular dissection and lymph node dissection, and the probability that the scene involves characteristic procedures (such as vascular clips and automatic anastomoses) performed using specific surgical devices (such as vascular clips and automatic anastomosers). The fifth learning model 350 is constructed, for example, by a CNN. Alternatively, the fifth learning model 350 may be a learning model constructed by an RNN, LSTM, GAN, etc., or it may be a learning model for image segmentation or object detection.
[0115] The calculations performed by the fifth learning model 350 are carried out in the second calculation unit 206. When a surgical field image is input, the second calculation unit 206 performs calculations according to the definition information of the fifth learning model 350, which includes the learned parameters. The fifth learning model 350 outputs the probability that the input surgical field image corresponds to a specific scene from each node constituting the output layer 353. The calculation results from the fifth learning model 350 are output to the control unit 201. The control unit 201 performs scene recognition by selecting the scene with the highest probability from among the probabilities of each scene output from the output layer 353.
[0116] Figure 17 is a flowchart showing the processing steps performed by the information processing device 200 according to Embodiment 4. The control unit 201 of the information processing device 200 acquires frame-by-frame surgical field images output from the CCU 110 via the input unit 204 (step S401). The control unit 201 performs the following processing each time it acquires frame-by-frame surgical field images.
[0117] The control unit 201 sends frame-by-frame surgical field images acquired through the input unit 204 to the first calculation unit 205 and the second calculation unit 206, and also gives an instruction to the second calculation unit 206 to start calculations (step S402).
[0118] When the control unit 201 issues an instruction to start the calculation, the second calculation unit 206 executes the calculation performed by the fifth learning model 350 (step S403). Specifically, the first calculation unit 205 executes the calculations in the input layer 351, the hidden layer 352, and the output layer 353 that constitute the fifth learning model 350, and outputs the probability of corresponding to a specific scene from each node of the output layer 353. The second calculation unit 206 outputs the calculation results from the fifth learning model 350 to the control unit 201 (step S404).
[0119] The control unit 201 performs scene recognition processing based on the calculation results from the second calculation unit 206 (step S405). That is, the control unit 201 identifies the scene represented by the current surgical field image by selecting the scene with the highest probability among the probabilities of each scene output from the output layer 353.
[0120] The control unit 201 selects a learning model for organ recognition according to the identified scene (step S406). For example, if the scene recognized in step S405 includes the ureter, the control unit 201 selects a learning model for ureter recognition. Also, if the scene recognized in step S405 is a scene of lymph node dissection of the upper edge of the pancreas during gastric cancer surgery, the control unit 201 selects a learning model for lymph node recognition, a learning model for pancreas recognition, a learning model for stomach recognition, etc. Also, if the scene recognized in step S405 is a scene of ligation using a vascular clip, the control unit 201 selects a learning model for vascular recognition, a learning model for surgical devices, etc. The control unit 201 can select a learning model for organ recognition according to the identified scene, not limited to the ureter, lymph node, pancreas, or stomach. Hereinafter, the learning model for organ recognition selected in step S406 will be referred to as the first learning model 310. The control unit 201, along with the information of the selected first learning model 310, issues an instruction to the first arithmetic unit 205 to start the calculation.
[0121] When the control unit 201 issues an instruction to start the calculation, the first calculation unit 205 executes the calculation performed by the first learning model 310 (step S407). Specifically, the first calculation unit 205 performs calculations using an encoder 311 that generates a feature map from the input surgical field image and sequentially downsamples the generated feature map, calculations using a decoder 312 that sequentially upsamples the feature map input from the encoder 311, and calculations using a softmax layer 313 that identifies each pixel of the feature map finally obtained from the decoder 312. The first calculation unit 205 outputs the calculation results from the learning model 310 to the control unit 201 (step S408).
[0122] The control unit 201 refers to the calculation results from the first learning model 310 and performs organ recognition processing (step S409). The control unit 201 can recognize organs included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 313 of the first learning model 310 is above a threshold (e.g., 60% or more).
[0123] The control unit 201 outputs the organ recognition result (step S410). Specifically, the control unit 201 instructs the first calculation unit 205 to superimpose the recognized organ image onto the surgical field image. The first calculation unit 205, in response to the instruction from the control unit 201, draws the recognized organ image into its built-in VRAM and outputs it to the display device 130 via the output unit 207, thereby superimposing the recognized organ image onto the surgical field image. Alternatively, a learning model that recognizes when a specific scene has ended may be used to instruct the termination of calculations in the first learning model or the start of a different learning model.
[0124] (Variation 4-1) If the information processing device 200 is configured to perform recognition processing for a specific organ (i.e., if it has only one first learning model 310), it may be configured not to perform organ recognition processing until it recognizes a scene that includes this specific organ.
[0125] Figure 18 is a flowchart showing the processing procedure in modified example 4-1. The control unit 201 of the information processing device 200 performs scene recognition processing each time a surgical field image is acquired, in the same procedure as shown in Figure 17 (steps S421 to S425).
[0126] The control unit 201 determines whether or not a specific scene has been recognized through scene recognition processing (step S426). The control unit 201 only needs to determine whether or not the scene is one of the pre-set scenes, such as a scene including a ureter, a lymph node dissection scene, or a scene of ligation using a vascular clip. If it determines that a specific scene has not been recognized (S426: NO), the control unit 201 terminates the processing according to this flowchart. On the other hand, if it determines that a specific scene has been recognized (S426: YES), the control unit 201 gives an instruction to the first calculation unit 205 to start calculation.
[0127] When the control unit 201 issues an instruction to start calculations, the first calculation unit 205 executes calculations using the first learning model 310 (step S427) and outputs the calculation results from the learning model 310 to the control unit 201 (step S428).
[0128] The control unit 201 refers to the calculation results from the first learning model 310 and performs organ recognition processing (step S429). The control unit 201 can recognize organs included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 313 of the first learning model 310 is above a threshold (e.g., 60% or more).
[0129] The control unit 201 outputs the organ recognition result (step S430). Specifically, the control unit 201 instructs the first calculation unit 205 to superimpose the recognized organ image onto the surgical field image. The first calculation unit 205, in response to the instruction from the control unit 201, draws the recognized organ image into its built-in VRAM and outputs it to the display device 130 via the output unit 207, thereby superimposing the recognized organ image onto the surgical field image.
[0130] (Modification 4-2) The control unit 201 of the information processing device 200 may acquire prior information according to the recognition result of the scene recognition process and perform organ recognition processing by referring to the acquired prior information. Figure 19 is a conceptual diagram showing an example of a prior information table. Prior information is registered in the prior information table according to the surgical scene. In gastric cancer surgery, when lymph node dissection of the upper edge of the pancreas is performed, the pancreas is located below the lymph nodes. For example, regarding the scene of lymph node dissection of the upper edge of the pancreas in gastric cancer surgery, the prior information table registers the prior information that the pancreas is located below the lymph nodes. The prior information table is not limited to the information shown in Figure 19, but various types of prior information according to various scenes are registered. The prior information table may be prepared in the storage unit 202 of the information processing device 200, or it may be prepared in an external device.
[0131] Figure 20 is a flowchart showing the processing procedure in modified example 4-2. The control unit 201 of the information processing device 200 performs scene recognition processing each time a surgical field image is acquired, in the same procedure as shown in Figure 17 (steps S441 to S445).
[0132] The control unit 201 accesses the storage unit 202 or an external device and acquires prior information according to the recognized scene (step S446). After acquiring the prior information, the control unit 201 gives an instruction to the first arithmetic unit 205 to start the calculation.
[0133] When the control unit 201 issues an instruction to start calculations, the first calculation unit 205 executes calculations using the first learning model 310 (step S447) and outputs the calculation results from the first learning model 310 to the control unit 201 (step S448).
[0134] The control unit 201 refers to the calculation results from the first learning model 310 and performs organ recognition processing (step S449). The control unit 201 can recognize organs included in the surgical field image by extracting pixels in which the probability of the label output from the softmax layer 313 of the first learning model 310 is above a threshold (e.g., 60% or more).
[0135] The control unit 201 determines whether the organ recognition result is consistent with the prior information (step S450). For example, in a gastric cancer surgery, if prior information is obtained that the pancreas is located below the lymph nodes, but the control unit 201 recognizes the pancreas as being located above the lymph nodes, the control unit 201 can determine that the organ recognition result is inconsistent with the prior information.
[0136] If the control unit 201 determines that the organ recognition result does not match the prior information (S450: NO), it outputs warning information (step S451). The control unit 201 outputs text information from the output unit 207 indicating that the organ recognition result does not match the prior information, and superimposes it within the display area of the surgical field image. The control unit 201 may also display the warning text information outside the display area of the surgical field image, or it may display the warning text information on another display device (not shown). Instead of displaying the warning text information, the control unit 201 may display a warning graphic or provide a warning by outputting sound or audio.
[0137] In this embodiment, a warning is output if the organ recognition result does not match the prior information. However, since there is a possibility of misrecognition when the two do not match, the output processing of the organ recognition result may be stopped. In this case, the calculations by the first learning model 310 and the fifth learning model 350, the organ recognition processing based on the calculation result of the first learning model 310, and the scene recognition processing based on the calculation result of the fifth learning model 350 are performed continuously, but the display of the recognized image showing the organ recognition result is stopped. Alternatively, the control unit 201 may be configured to stop the organ recognition processing instead of stopping the output processing.
[0138] If the organ recognition process determines that it is consistent with the prior information (S450: YES), the control unit 201 outputs the organ recognition result (step S452). Specifically, the control unit 201 instructs the first calculation unit 205 to superimpose the recognized organ image onto the surgical field image. The first calculation unit 205, in response to the instruction from the control unit 201, draws the recognized organ image into its built-in VRAM and outputs it to the display device 130 via the output unit 207, thereby superimposing the recognized organ image onto the surgical field image.
[0139] In the flowchart of Figure 20, the system is configured to refer to prior information after performing organ recognition processing, but prior information may also be referred to when performing organ recognition processing. For example, if prior information is obtained that the pancreas is located below the lymph nodes, the control unit 201 may generate mask information to exclude the upper part of the lymph nodes from the recognition target, and give a start operation instruction to the first calculation unit 205 along with the generated mask information. The first calculation unit 205 can then apply the mask to the surgical field image and perform calculations using the first learning model 310 based on the partial image outside the masked area.
[0140] (Modification 4-3) The control unit 201 of the information processing device 200 may change the threshold used for organ recognition depending on the recognized scene. Figure 21 is a flowchart of the processing procedure in modified example 4-3. The control unit 201 of the information processing device 200 performs scene recognition processing each time a surgical field image is acquired, in the same procedure as shown in Figure 17 (steps S461 to S465).
[0141] The control unit 201 determines whether or not a specific scene has been recognized through scene recognition processing (step S466). The control unit 201 only needs to determine whether or not the scene is one of the pre-set scenes, such as a scene including a ureter or a lymph node dissection scene.
[0142] If the control unit 201 determines that a specific scene has been recognized (S466: YES), it sets the threshold used for organ recognition to a relatively low first threshold (< second threshold) (step S467). In other words, the control unit 201 sets the threshold so that the organ to be recognized is more easily detected. After setting the threshold, the control unit 201 gives a calculation start instruction to the first calculation unit 205.
[0143] On the other hand, if it is determined that a specific scene is not being recognized (S466: NO), the control unit 201 sets the threshold used for organ recognition to a relatively higher second threshold (> first threshold) (step S468). In other words, the control unit 201 sets the threshold so that the organ to be recognized is less likely to be detected. After setting the threshold, the control unit 201 gives a calculation start instruction to the second calculation unit 206.
[0144] When the control unit 201 issues an instruction to start calculations, the first calculation unit 205 executes calculations using the first learning model 310 (step S469) and outputs the calculation results from the learning model 310 to the control unit 201 (step S470).
[0145] The control unit 201 refers to the calculation results from the first learning model 310 and performs organ recognition processing (step S471). The control unit 201 compares the probability of the label output from the softmax layer 313 of the first learning model 310 with the threshold (first threshold or second threshold) set in step S467 or step S468, and can recognize organs included in the surgical field image by extracting pixels that are equal to or greater than the threshold.
[0146] The control unit 201 outputs the organ recognition result (step S472). Specifically, the control unit 201 instructs the first calculation unit 205 to superimpose the recognized organ image onto the surgical field image. The first calculation unit 205, in response to the instruction from the control unit 201, draws the recognized organ image into its built-in VRAM and outputs it to the display device 130 via the output unit 207, thereby superimposing the recognized organ image onto the surgical field image.
[0147] (Modification 4-4) The control unit 201 of the information processing device 200 may change the threshold for the period until the target organ is recognized and the period after the target organ is recognized. Figure 22 is a flowchart of the processing procedure in modified example 4-4. The control unit 201 of the information processing device 200 performs scene recognition processing and organ recognition processing each time a surgical field image is acquired, in the same procedure as shown in Figure 17 (steps S481 to S488). A third threshold (<first threshold) is pre-set for the threshold used in the organ recognition processing.
[0148] The control unit 201 determines whether or not an organ has been recognized by the organ recognition process in step S488 (step S489). For example, if the number of pixels in which the probability of the label output from the softmax layer 313 of the first learning model 310 is determined to be above the third threshold (e.g., 30% or more) is greater than or equal to a predetermined number, the control unit 201 can determine that an organ has been recognized from the surgical field image.
[0149] If the control unit 201 determines that it has not recognized an organ (S489: NO), it sets a third threshold (step S490). That is, the control unit 201 maintains a preset threshold. By maintaining the threshold at a relatively low value until an organ is recognized, the organ becomes easier to detect, and the information processing device 200 can function as a sensor for organ detection.
[0150] If the control unit 201 determines that an organ has been recognized (S489: YES), it sets a first threshold that is higher than the third threshold (step S491). In other words, once the system begins to recognize an organ, the control unit 201 changes the threshold used for organ recognition from the third threshold to the first threshold (> third threshold), thereby improving the accuracy of organ recognition.
[0151] The control unit 201 outputs the organ recognition result (step S492). Specifically, the control unit 201 instructs the first calculation unit 205 to superimpose the recognized organ image onto the surgical field image. The first calculation unit 205, in response to the instruction from the control unit 201, draws the recognized organ image into its built-in VRAM and outputs it to the display device 130 via the output unit 207, thereby superimposing the recognized organ image onto the surgical field image.
[0152] (Modifications 4-5) The control unit 201 of the information processing device 200 may derive the proper names of recognized organs from the calculation results of the first learning model 310 based on the information of the scene recognized by scene recognition. Figure 23 is a conceptual diagram showing an example of a proper name table. The proper name table registers the proper names of organs in association with the surgical scene. For example, the inferior mesenteric artery and inferior mesenteric plexus often appear in the scene of sigmoid colon cancer surgery. In the proper name table, the inferior mesenteric artery and inferior mesenteric plexus are registered as proper names of organs in relation to the scene of sigmoid colon cancer surgery. The proper name table may be prepared in the storage unit 202 of the information processing device 200, or it may be prepared in an external device.
[0153] The control unit 201 can access the memory unit 202 or an external device according to the scene recognized by scene recognition and read the proper names of organs registered in the proper name table, thereby presenting the proper name of the organ identified by organ recognition to the surgeon. Figure 24 is a schematic diagram showing an example of organ name display. When the control unit 201 recognizes a sigmoid colon cancer surgery scene by scene recognition and recognizes a blood vessel using the first learning model 310 for blood vessel recognition, it estimates that the blood vessel is the inferior mesenteric artery by referring to the proper name table. The control unit 201 can then display the text information that the proper name of the recognized blood vessel is the inferior mesenteric artery on the display device 130. Alternatively, the control unit 201 may display the proper name on the display device 130 only when it receives instructions from the surgeon through the operation unit 203, etc.
[0154] Furthermore, the control unit 201 recognizes the sigmoid colon cancer surgery scene through scene recognition, and when it recognizes a nerve using the first learning model 310 for nerve recognition, it can refer to the proper nomenclature table to estimate that the nerve is the inferior mesenteric plexus and display it on the display device 130.
[0155] (Variations 4-6) The control unit 201 of the information processing device 200 may derive structural information based on the scene information recognized by scene recognition and the organ information recognized by organ recognition. In the modified example 4-6, the structural information derived includes information on organs not recognized by organ recognition, and information on lesions such as cancer and tumors.
[0156] The control unit 201 derives information about structures by referring to a structure table. Figure 25 is a conceptual diagram showing an example of a structure table. The structure table registers information about known structures for each scenario. The structure information is, for example, textbook information about organs, including information such as the organ's name, location, and direction of course. For example, by referring to the structure table shown in Figure 25, if the control unit 201 recognizes the portal vein during gastric surgery, it can present the surgeon with information that the left gastric vein branches off from the portal vein. Similarly, if the control unit 201 recognizes the right gastric artery during gastric surgery, it can present the surgeon with information that the base of the right gastric artery is shaped like the Japanese character for "person".
[0157] A structure table may be prepared for each patient. The patient-specific structure table registers information about the lesion area obtained in advance using other medical images or examination methods. For example, information about the lesion area obtained in advance using CT (Computed Tomography) images, MRI (Magnetic Resonance Imaging) images, ultrasound images, optical coherence tomography images, angiography images, etc., is registered. Using these medical images, information about lesion areas inside organs that do not appear in the observation image (surgical field image) of the laparoscope 11 can be obtained. When the control unit 201 of the information processing device 200 recognizes a scene or organ, it can refer to the structure table and present information about lesion areas that do not appear in the surgical field image.
[0158] Figure 26 is a schematic diagram showing an example of how a structure is displayed. Figure 26A shows an example of how the right gastric artery is displayed when it is recognized during gastric surgery. The control unit 201 recognizes the right gastric artery through scene recognition or organ recognition, refers to the structure table, and reads out information that the base of the right gastric artery is shaped like the Japanese character for "person". Based on the information read from the structure table, the control unit 201 can display on the display device 130 the textual information that the base of the right gastric artery is shaped like the Japanese character for "person" and predict the course of blood vessels that have not yet been confirmed.
[0159] Figure 26B shows an example of displaying a lesion that is not visible in the surgical field image. If information on lesions obtained in advance for a specific patient is registered in the structure table, the control unit 201, upon recognizing the scene and organ, refers to the patient-specific structure table to read information on lesions that are not visible in the surgical field image. Based on the information read from the structure table, the control unit 201 can display images of lesions not visible in the surgical field image or text information indicating the presence of lesions inside organs on the display device 130.
[0160] The control unit 201 may display the above structure as a three-dimensional image. The control unit 201 can display the structure as a three-dimensional image by reconstructing multiple tomographic images obtained in advance by CT or MRI using techniques such as surface rendering or volume rendering. The control unit 201 may display the three-dimensional image as an object in augmented reality (AR) by superimposing it on the surgical field image, or it may display it as a virtual reality (VR) object separately from the surgical field image. The objects in augmented reality or virtual reality may be presented to the surgeon through a head-mounted display not shown in the figure. The information processing device 200 according to this embodiment can recognize and display structures such as organs that appear in the surgical field image using the first learning model 310, and can display structures that do not appear in the surgical field image using AR or VR technology. The surgeon can visually perceive the structures of organs that appear in the surgical field image and the internal structures that do not appear in the surgical field image, and can easily grasp the overall picture.
[0161] (Modification 4-7) The control unit 201 of the information processing device 200 may predict possible events based on the scene information recognized by scene recognition and the organ information recognized by organ recognition. A case table, which collects cases that have occurred in past surgeries, is used to predict events. Figure 27 is a conceptual diagram showing an example of a case table. The case table registers information on cases that occurred during surgery, associated with the scene recognized by scene recognition and the organ recognized by organ recognition. The table in Figure 27 shows an example in which many cases of bleeding from the inferior mesenteric artery in a sigmoid colon cancer surgery scene are registered.
[0162] Figure 28 is a schematic diagram showing an example of event display. When the control unit 201 recognizes the scene of sigmoid colon cancer surgery through scene recognition and recognizes the inferior mesenteric artery through organ recognition, it can refer to the case table to understand that bleeding has occurred frequently in the past, and can display this information as text on the display device 130. In addition, the control unit 201 may also display special cases that have occurred in the past or cases that should be notified to the surgeon as prior information on the display device 130, not limited to cases that have occurred frequently in the past.
[0163] As described above, in Embodiment 4, based on the calculation results obtained from the first learning model 310 for organ recognition and the calculation results obtained from the fifth learning model 350 for scene recognition, an integrated recognition result for the surgical field image can be derived, and information based on the derived recognition result can be provided to the surgeon.
[0164] (Embodiment 5) Embodiment 5 describes a configuration that combines event recognition and scene recognition to derive an integrated recognition result.
[0165] The information processing device 200 according to Embodiment 5 derives information on characteristic scenes in an event according to the recognition result of the event. As a specific example, when the information processing device 200 recognizes organ damage as an event, a configuration in which the scene of damage occurrence and the scene of damage repair are derived as characteristic scenes will be described. The information processing device 200 includes a third learning model 330 that has been trained to output information on organ damage in response to the input of a surgical field image, and a fifth learning model 350 that has been trained to output information on the scene of damage occurrence and the scene of damage repair in response to the input of a surgical field image. Each time the information processing device 200 acquires a surgical field image, the first arithmetic unit 205 performs calculations by the third learning model 330, and the second arithmetic unit 206 performs calculations by the fifth learning model 350. The information processing device 200 also temporarily records the surgical field image (video) input from the input unit 204 in the storage unit 202.
[0166] Figure 29 is a flowchart showing the processing steps performed by the information processing device 200 according to Embodiment 5. Each time a surgical field image is acquired, the control unit 201 refers to the calculation results by the third learning model 330 and performs recognition processing for events (organ damage), and determines whether or not there is organ damage based on the recognition result. If it is determined that there is organ damage, the control unit 201 performs the following processing.
[0167] Each time a surgical field image is acquired, the control unit 201 performs scene recognition processing by referring to the calculation results of the fifth learning model 350 (step S501), and determines whether or not the scene where the injury occurred has been recognized based on the result of the scene recognition processing (step S502). The control unit 201 compares the scene of the previous frame with the scene of the current frame, and if it switches from a scene where the organ is not damaged to a scene where the organ is damaged, it determines that the scene where the injury occurred has been recognized. If it determines that the scene where the injury occurred has not been recognized (S502: NO), the control unit 201 proceeds to step S505.
[0168] If the control unit 201 determines that it has recognized the scene of injury (S502: YES), it extracts a partial video containing the scene of injury (step S503). The control unit 201 extracts a partial video containing the scene of injury by specifying, for example, a time before the time of injury (e.g., a few seconds before) as the start point of the partial video and the time of injury as the end point of the partial video, using the surgical field image (video) temporarily recorded in the memory unit 202. Alternatively, the control unit 201 may further record the surgical field image (video) and extract a partial video containing the scene of injury by specifying a time before the time of injury as the start point of the partial video and a time after the time of injury (e.g., a few seconds after) as the end point of the partial video.
[0169] The control unit 201 stores the extracted portion of the video in the storage unit 202 (step S504). The control unit 201 extracts the portion of the video between the start point and end point specified in step S503 from the recorded video and stores it separately in the storage unit 202 as a video file.
[0170] The control unit 201 determines whether or not it has recognized a scene of damage repair based on the result of the scene recognition process in step S501 (step S505). The control unit 201 compares the scene of the previous frame with the scene of the current frame and determines that it has recognized a scene of damage repair if it has switched from a scene in which the organ is damaged to a scene in which the organ is not damaged. If it determines that it has not recognized a scene of damage repair (S505: NO), the control unit 201 terminates the process according to this flowchart.
[0171] If the control unit 201 determines that it has recognized a repair scene of damage (S505: YES), it extracts a partial video including the repair scene (step S506). The control unit 201 extracts a partial video including the repair scene by specifying, for example, a point in time before the repair (e.g., a few seconds before) as the start point of the partial video and the repair time as the end point of the partial video, using the surgical field image (video) temporarily recorded in the memory unit 202. Alternatively, the control unit 201 may further record the surgical field image (video) and extract a partial video including the repair scene by specifying a point in time before the repair as the start point of the partial video and a point in time after the repair (e.g., a few seconds after) as the end point of the partial video.
[0172] The control unit 201 stores the extracted portion of the video in the storage unit 202 (step S507). The control unit 201 extracts the portion of the video between the start point and end point specified in step S506 from the recorded video and stores it separately in the storage unit 202 as a video file.
[0173] The control unit 201 registers information about the recognized scene in the scene record table each time it recognizes a scene. The scene record table is stored in the storage unit 202. Figure 30 is a conceptual diagram showing an example of a scene record table. The scene record table stores, for example, the date and time the scene was recognized, the name that identifies the recognized scene, and the video file of the partial image extracted when the scene was recognized, in association with each other. In the example in Figure 30, the video file showing the scene of damage occurrence and the video file showing the scene of damage repair recognized from the start to the end of surgery are registered in the scene record table in association with the date and time.
[0174] The control unit 201 may display information about scenes registered in the scene recording table on the display device 130. The control unit 201 may also display the scene information on the display device 130 in a table format so that, for example, a surgeon or other person can select any scene. Alternatively, the control unit 201 may place objects (UI) such as thumbnails and icons representing each scene on the display screen and accept scene selections from the surgeon or other person. When the control unit 201 accepts a scene selection from the surgeon or other person, it reads the video file of the corresponding scene from the storage unit 202 and plays the read video file. The played video file is displayed on the display device 130.
[0175] Alternatively, the control unit 201 may be configured to play a video file of a scene it has selected. The control unit 201 may refer to the scene recording table and, if the scene in which damage occurred is registered but the scene in which the damage was repaired is not registered, it may play the scene in which the damage occurred to inform the operator that no repair was performed.
[0176] In this embodiment, the system recognizes the scene of damage occurrence and the scene of repair, extracts partial video footage from each recognized scene, and stores it as a video file in the storage unit 202. However, the scenes to be recognized are not limited to the scene of damage occurrence and the scene of repair. The control unit 201 can recognize characteristic scenes in various events that may occur during surgery, extract partial video footage of the recognized characteristic scenes, and store it as a video file in the storage unit 202.
[0177] For example, the control unit 201 may, upon recognizing a bleeding event, recognize the scene of bleeding onset and the scene of hemostasis, extract partial video footage from each recognized scene, and store it as a video file in the storage unit 202. Alternatively, when the control unit 201 recognizes a bleeding event, it may recognize the scene in which an artificial object such as gauze enters the surgical field image and the scene in which it exits the surgical field image, extract partial video footage from each recognized scene, and store it as a video file in the storage unit 202. Since the introduction of gauze into the patient's body is not limited to bleeding, the control unit 201 may, regardless of the recognition of a bleeding event, extract partial video footage from scenes in which gauze enters the surgical field image and scenes in which gauze exits the surgical field image during surgery, and store it as a video file in the storage unit 202. Furthermore, the control unit 201 may recognize scenes in which artificial objects such as hemostatic clips, restraint bands, and sutures, not limited to gauze, enter the surgical field image and scenes in which they enter the surgical field image, and store video files corresponding to the partial images of these scenes in the storage unit 202.
[0178] As described above, in Embodiment 5, since partial video footage (video files) showing characteristic scenes of events that occurred during surgery are stored in the storage unit 202, it is possible to easily review the events.
[0179] (Embodiment 6) The control unit 201 of the information processing device 200 may cause both the first calculation unit 205 and the second calculation unit 206 to perform calculations using the first learning model 310 for the same surgical field image, and then evaluate the calculation results from the first calculation unit 205 and the calculation results from the second calculation unit 206 to derive error information from either the first calculation unit 205 or the second calculation unit 206.
[0180] Figure 31 is a flowchart showing the processing steps performed by the information processing device 200 according to Embodiment 6. When the control unit 201 acquires a surgical field image (step S601), it causes both the first calculation unit 205 and the second calculation unit 206 to perform calculations using the first learning model 310 (step S602).
[0181] The control unit 201 obtains the calculation result of the first calculation unit 205 and the calculation result of the second calculation unit 206 (step S603) and evaluates the calculation result (step S604). The control unit 201 determines whether the calculation result of the first calculation unit 205 and the calculation result of the second calculation unit 206 differ by a predetermined amount or more (step S605). If the difference is less than the predetermined amount (S605: NO), the control unit 201 terminates the process according to this flowchart. If it is determined that the difference is greater than the predetermined amount (S605: YES), the control unit 201 determines that an error has been detected (step S606) and outputs warning information (step S607). The warning information may be displayed on the display device 130 or output by sound or voice.
[0182] (Embodiment 7) Embodiment 7 describes a configuration that estimates the usage state of a surgical instrument by combining recognition results from multiple learning models.
[0183] The information processing device 200 according to Embodiment 7 includes a first learning model 310 for recognizing organs, a fourth learning model 340 for recognizing devices, and a fifth learning model 350 for recognizing scenes.
[0184] The first learning model 310 is a learning model that, in response to input of a surgical field image, outputs a probability indicating whether or not each pixel constituting the surgical field image corresponds to the organ to be recognized. The control unit 201 of the information processing device 200 acquires calculation results from the first learning model 310 in real time and analyzes them in a time series to estimate the effect of the surgical instrument on the organ. For example, the control unit 201 calculates the number of pixels recognized as organs, changes in area, or the amount of pixels recognized as organs that have disappeared by the first learning model 310, and if it determines that a predetermined amount of pixels have disappeared, it estimates that the surgical instrument has moved onto the organ.
[0185] The fourth learning model 340 is a learning model that is trained to output information about devices appearing in a surgical field image in response to the input of a surgical field image. In Embodiment 7, the fourth learning model 340 is trained to include information about the open / closed state of the forceps 13 as information about the device.
[0186] The fifth learning model 350 is a learning model that, when a surgical field image is input, is trained to output information about the scene shown in the surgical field. In Embodiment 7, the fifth learning model 350 is trained to include information about the scene in which an organ is grasped as the information about the scene.
[0187] Figure 32 is an explanatory diagram illustrating the estimation method in Embodiment 7. The surgical field image shown in Figure 32 shows a scene in which connective tissue is being removed between tissue ORG1, which constitutes an organ, and tissue ORG2, which contains a lesion such as a malignant tumor. At this time, the surgeon grasps tissue ORG2, which contains the lesion, with forceps 13 and spreads it in an appropriate direction to expose the connective tissue that exists between tissue ORG2, which contains the lesion, and tissue ORG1, which is to be preserved. The surgeon then removes the exposed connective tissue using an energy treatment instrument 12, thereby separating tissue ORG2, which contains the lesion, from tissue ORG1, which is to be preserved.
[0188] The control unit 201 of the information processing device 200 inputs a surgical field image, as shown in Figure 32, into the fourth learning model 340 and obtains the calculation result, recognizing that the forceps 13 are closed. However, based solely on the calculation result of the fourth learning model 340, the control unit 201 cannot determine whether the forceps 13 are closed while grasping an organ or without grasping an organ.
[0189] Therefore, in Embodiment 7, the usage state of the forceps 13 is estimated by further combining the calculation results of the first learning model 310 and the fifth learning model 350. That is, if the control unit 201 recognizes from the calculation result of the fourth learning model 340 that the forceps 13 is in a closed state, from the calculation result of the first learning model 310 that the forceps 13 is on an organ, and from the calculation result of the fifth learning model 350 that it is a situation in which the organ is being grasped, then the control unit 201 can recognize that the forceps 13 is closed while grasping the organ.
[0190] Figure 33 is a flowchart showing the estimation procedure in Embodiment 7. The control unit 201 of the information processing device 200 acquires frame-by-frame surgical field images output from the CCU 110 via the input unit 204 (step S701). The control unit 201 performs the following processing each time it acquires frame-by-frame surgical field images.
[0191] The control unit 201 obtains the calculation result from the fourth learning model 340 (step S702) and determines whether the forceps 13 is closed or not (step S703). The calculation by the fourth learning model 340 is performed, for example, by the second calculation unit 206. The control unit 201 only needs to obtain the calculation result from the fourth learning model 340 from the second calculation unit 206. If it is determined that the forceps 13 is not closed (S703: NO), the control unit 201 terminates the process according to this flowchart.
[0192] If the control unit 201 determines that the forceps 13 is closed (S703: YES), it obtains the calculation result from the first learning model 310 (step S704) and determines whether the forceps 13 is present on the organ or not (step S705). The calculation by the first learning model 310 is performed, for example, by the first calculation unit 205. The control unit 201 only needs to obtain the calculation result from the first learning model 310 from the first calculation unit 205. If the control unit 201 determines that the forceps 13 is not present on the organ (S705: NO), it terminates the processing according to this flowchart.
[0193] If the control unit 201 determines that the forceps 13 is on an organ (S705: YES), it obtains the calculation result from the fifth learning model 350 (step S706) and determines whether or not the organ is being grasped (step S707). The calculation by the fifth learning model 350 is performed by the second calculation unit 206, for example, while the calculation by the fourth learning model 340 is not being performed. The control unit 201 only needs to obtain the calculation result from the fifth learning model 350 from the second calculation unit 206. If the control unit 201 determines that the organ is not being grasped (S707: NO), it terminates the processing according to this flowchart.
[0194] If the control unit determines that an organ is being grasped (step S707: YES), it estimates that the forceps have grasped the organ (step S708). The estimation result may be displayed on the surgical field image. For example, when the organ grasp by the forceps 13 is completed, the color of the forceps 13 may be increased as an indicator.
[0195] In this embodiment, a configuration for estimating the state in which an organ is grasped by forceps 13 has been described. However, by combining the recognition results of multiple learning models, it is possible to estimate not only grasping by forceps 13 but also procedures using various surgical devices (grasping, cutting, dissection, etc.).
[0196] (Embodiment 8) Embodiment 8 describes a configuration in which the optimal learning model is selected according to the input surgical field image.
[0197] The information processing device 200 according to Embodiment 8 includes multiple learning models that recognize the same recognition target. As an example, a configuration that recognizes the same organ using a first learning model 310 and a second learning model 320 will be described. The organs recognized by the first learning model 310 and the second learning model 320 are not limited to loose connective tissue or nerve tissue, but can be any pre-set organ.
[0198] In one example, the first learning model 310 and the second learning model 320 are constructed using different neural networks. For example, the first learning model 310 is constructed using SegNet, and the second learning model 320 is constructed using U-Net. The combination of neural networks used to construct the first learning model 310 and the second learning model 320 is not limited to the above; any neural network can be used.
[0199] Alternatively, the first learning model 310 and the second learning model 320 may be learning models with different internal configurations. For example, the first learning model 310 and the second learning model 320 may be learning models constructed using the same neural network, but with different types and numbers of layers, number of nodes, and node connections.
[0200] Furthermore, the first learning model 310 and the second learning model 320 may be learning models trained using different training data. For example, the first learning model 310 may be a learning model trained using training data that includes ground truth data annotated by a first expert, and the second learning model 320 may be a learning model trained using training data that includes ground truth data annotated by a second expert different from the first expert. Alternatively, the first learning model 310 may be a learning model trained using training data that includes surgical field images taken at one medical institution and annotation data (ground truth data) for those surgical field images, and the second learning model 320 may be a learning model trained using training data that includes surgical field images taken at another medical institution and annotation data (ground truth data) for those surgical field images.
[0201] When a surgical field image is input to the information processing device 200, the first calculation unit 205 performs calculations based on the first learning model 310, and the second calculation unit 206 performs calculations based on the second learning model 320. The control unit 201 of the information processing device 200 analyzes the calculation results from the first calculation unit 205 and the calculation results from the second calculation unit 206, and selects the optimal learning model for organ recognition (in this embodiment, either the first learning model 310 or the second learning model) based on the analysis results.
[0202] Figure 34 is an explanatory diagram illustrating the analysis method for the calculation results. Each learning model for recognizing organs outputs a probability (confidence level) as a calculation result indicating whether or not each pixel corresponds to the target organ. When the number of pixels is aggregated for each confidence level, a distribution like that shown in Figures 34A to 34C is obtained. In the graphs shown in Figures 34A to 34C, the horizontal axis represents the confidence level, and the vertical axis represents the number of pixels (proportion of the entire image). Ideally, each pixel is classified as having a confidence level of 1 (100% probability of being an organ) or a confidence level of 0 (0% probability of being an organ). Therefore, when the distribution of confidence levels is examined based on the calculation results obtained from an ideal learning model, a polarized distribution like that shown in Figure 34A is obtained.
[0203] When the control unit 201 of the information processing device 200 obtains calculation results from the first learning model 310 and the second learning model 320, it aggregates the number of pixels for each confidence level and selects the learning model that has a distribution close to the ideal distribution. For example, if the distribution obtained from the calculation results of the first learning model 310 is the distribution shown in Figure 34B, and the distribution obtained from the calculation results of the second learning model 320 is the distribution shown in Figure 34C, the control unit 201 selects the second learning model 320 because the latter is closer to the ideal distribution.
[0204] The control unit 201 determines whether the distribution is close to an ideal distribution by evaluating each distribution using, for example, evaluation coefficients such that the evaluation value increases as the confidence level approaches 1 or 0. Figure 35 shows an example of an evaluation coefficient table. Such an evaluation coefficient table is pre-prepared in the storage unit 202. In the example in Figure 35, the evaluation coefficients are set to take higher values as the confidence level approaches 1 or 0.
[0205] When the control unit 201 obtains the aggregated number of pixels for each confidence level, it calculates a score indicating the quality of the distribution by multiplying it by an evaluation coefficient. Figure 36 shows an example of the score calculation results. Figures 36A to 36C show the results of calculating the score for each distribution shown in Figures 34A to 34C. The score calculated from the ideal distribution is the highest. When the score is calculated for the distribution obtained from the calculation results of the first learning model 310, the total score is 84, and when the score is calculated for the distribution obtained from the calculation results of the second learning model 320, the total score is 188. In other words, the second learning model 320 has a higher score than the first learning model 310, so the control unit 201 selects the second learning model 320 as the appropriate learning model.
[0206] Figure 37 is a flowchart showing the processing steps performed by the information processing device 200 according to Embodiment 8. When the control unit 201 acquires a surgical field image (step S801), it causes the first calculation unit 205 to perform calculations by the first learning model 310 (step S802), and obtains the calculation results by the first learning model 310 (step S803). The control unit 201 aggregates the number of pixels for each confidence level for the first learning model 310 (step S804), and calculates a distribution score (first score) by multiplying each by an evaluation coefficient (step S805).
[0207] Similarly, the control unit 201 causes the second calculation unit 206 to perform calculations using the second learning model 320 on the surgical field image acquired in step S801 (step S806), and obtains the calculation results from the second learning model 320 (step S807). The control unit 201 aggregates the number of pixels for each confidence level for the second learning model 320 (step S808), and calculates a distribution score (second score) by multiplying each by an evaluation coefficient (step S809).
[0208] In this flowchart, for convenience, the calculations for the first learning model 310 (S802-S805) are performed first, followed by the calculations for the second learning model 320 (S806-S809). However, these steps may be performed in any order, or concurrently.
[0209] The control unit 201 compares the first score and the second score and determines whether the first score is greater than or equal to the second score (step S810).
[0210] If the control unit 201 determines that the first score is equal to or greater than the second score (S810: YES), it selects the first learning model 310 as the appropriate learning model (step S811). Thereafter, the control unit 201 performs organ recognition processing using the selected first learning model 310.
[0211] If the control unit 201 determines that the first score is less than the second score (S810: NO), it selects the second learning model 320 as the appropriate learning model (step S812). Thereafter, the control unit 201 performs organ recognition processing using the selected second learning model 320.
[0212] As described above, in Embodiment 8, a more appropriate learning model can be selected and the organ recognition process can be executed.
[0213] The information processing device 200 may perform organ recognition processing using the calculation results of the first learning model 310 in the foreground and perform calculations using the second learning model 320 in the background. The control unit 201 may periodically evaluate the first learning model 310 and the second learning model 320 and switch the learning model used for organ recognition according to the evaluation results. Furthermore, the control unit 201 may evaluate the first learning model 310 and the second learning model 320 at the timing instructed by the operator or the like, and switch the learning model used for organ recognition according to the evaluation results. In addition, by combining this with the scene recognition described in Embodiment 4, the first learning model 310 and the second learning model 320 may be evaluated when a feature scene is recognized, and the learning model used for organ recognition may be switched according to the evaluation results.
[0214] In Embodiment 8, an example of application to the learning models 310 and 320 for organ recognition was described. However, for the learning model 330 for event recognition, the learning model 340 for device recognition, and the learning model 350 for scene recognition, default and optional models may also be prepared, and the model used for recognition may be switched by evaluating the calculation results of these models.
[0215] In Embodiment 8, a method using evaluation coefficients was described as an evaluation method for the first learning model 310 and the second learning model 320. However, evaluation is not limited to methods using evaluation coefficients; various statistical indicators can be used. For example, the control unit 201 may calculate the variance and standard deviation of the distribution and determine that the distribution is polarized if the variance and standard deviation are high. Alternatively, the control unit 201 may evaluate the calculation results of each model by taking the value of 100 minus the percentage of pixels as the value of the vertical axis of the graph and calculating the kurtosis and skewness of the graph. Furthermore, the control unit 201 may evaluate the calculation results of each model using the mode, percentiles, etc.
[0216] (Embodiment 9) Embodiment 9 describes a configuration in which organ recognition is performed in the foreground and bleeding event recognition is performed in the background.
[0217] Figure 38 is a sequence diagram showing an example of processing performed by the information processing device 200 according to Embodiment 9. The information processing device 200 receives image data of the surgical field image output from the CCU 110 as it is received (video input) through the input unit 204. The input surgical field image is continuously recorded in, for example, the storage unit 202. The control unit 201 loads the surgical field image into memory frame by frame and instructs the first arithmetic unit 205 and the second arithmetic unit 206 to perform calculations. If the first arithmetic unit 205 has the processing capability to process surgical field images at 30 FPS, the control unit 201 only needs to instruct the first arithmetic unit 205 to perform calculations each time the surgical field image is loaded into memory at a frame rate of 30 FPS. The first arithmetic unit 205 performs calculations (inference processing) based on the first learning model 310 in response to instructions from the control unit 201 and performs processing to draw the recognized organ image based on the calculation result onto the built-in VRAM. The recognized organ image drawn on the VRAM is output through the output unit 207 and displayed on the display device 130. The information processing device 200 according to Embodiment 9 can continuously display recognized images of organs included in the surgical field image by performing organ recognition (inference) processing and drawing processing by the first arithmetic unit 205 in the foreground.
[0218] The second processing unit 206 performs recognition processing for events that do not need to be constantly monitored in the background. Embodiment 9 describes a configuration for recognizing bleeding events as an example, but a configuration for recognizing organ damage events or scenes may be used instead of bleeding events. The computing power of the second processing unit 206 may be lower than that of the first processing unit 205. For example, if the second processing unit 206 has the computing power to process surgical field images at 6 FPS, as shown in Figure 38, the second processing unit 206 only needs to process surgical field images input at 30 FPS once every five times.
[0219] Furthermore, if the second arithmetic unit 206 has higher processing power, it may perform recognition processing for multiple events in the background. For example, if the second arithmetic unit 206 has the same processing power as the first arithmetic unit 205, it is possible to perform recognition processing for up to five types of events, and events such as bleeding and organ damage can be performed sequentially.
[0220] Figure 39 is a flowchart showing the procedure of processing performed by the first arithmetic unit 205. The first arithmetic unit 205 determines whether the current state is in inference mode (step S901). If it is in inference mode (S901: YES), the first arithmetic unit 205 obtains the latest frame loaded into memory by the control unit 201 (step S902) and performs inference processing (step S903). In the inference processing, the first arithmetic unit 205 performs calculations according to the first learning model 310 for organ recognition. By performing calculations using the first learning model 310, an organ recognition image is obtained. The first arithmetic unit 205 performs drawing processing to draw the organ recognition image on the built-in VRAM (step S904). The organ recognition image drawn on VRAM is output to the display device 130 via the output unit 207 and displayed on the display device 130. After completing the drawing processing, the first arithmetic unit 205 returns the process to step S901.
[0221] If the current state is not inference mode (S901:NO), the first arithmetic unit 205 determines whether the current state is playback mode or not (step S905). If the current state is not playback mode (S905:NO), the first arithmetic unit 205 returns to step S901.
[0222] If the current state is playback mode (S905: YES), the first calculation unit 205 acquires the bleeding log created by the second calculation unit 206 (step S906) and acquires a specified frame from the saved partial video (step S907). Here, the partial video is a video of the surgical field having frames for a time range from the start of bleeding to the end of bleeding. Alternatively, the partial video may include the start of bleeding and have frames for a predetermined time range from before the start of bleeding to a time when a predetermined time range has elapsed. The second calculation unit 206 performs a drawing process to draw the partial video on the built-in VRAM (step S908). The partial video drawn on the VRAM is output to the display device 130 via the output unit 207 and displayed on the display device 130. After completing the drawing process, the first calculation unit 205 returns the process to step S905.
[0223] Figure 40 is a flowchart showing the procedure of processing performed by the second arithmetic unit 206. The second arithmetic unit 206 acquires the latest frame loaded into memory by the control unit 201 (step S921) and performs inference processing (step S922). In the inference processing, the second arithmetic unit 206 performs calculations according to the third learning model 330 for event recognition.
[0224] The second calculation unit 206 performs bleeding determination processing and logging processing based on the calculation results of the third learning model 330 (step S923). Based on the calculation results of the third learning model 330, the second calculation unit 206 determines that the start of bleeding has been recognized if the number of pixels recognized as bleeding (or the proportion within the surgical field image) exceeds a threshold. Furthermore, after recognizing the start of bleeding, the second calculation unit 206 determines that the end of bleeding has been recognized if the number of pixels recognized as bleeding (or the proportion within the surgical field image) falls below the threshold. When the second calculation unit 206 recognizes the start of bleeding or the end of bleeding, it records the start of bleeding or the end of bleeding in the log.
[0225] The second calculation unit 206 determines whether the sets for bleeding start and bleeding end are available (step S924). If the sets are not available (S924: NO), the second calculation unit 206 returns to step S921. If the sets are available, the second calculation unit 206 saves the partial image (step S925). That is, the second calculation unit 206 only needs to store the image of the surgical field having frames for the time range from the bleeding start to the bleeding end in the storage unit 202. Alternatively, the second calculation unit 206 may store the image of the surgical field as a partial image in the storage unit 202, including the bleeding start time and having frames from before the bleeding started to a predetermined time range.
[0226] As described above, in Embodiment 9, organ recognition is performed in the foreground and event recognition is performed in the background. If an event such as bleeding is recognized in the second calculation unit 206, the first calculation unit 205 can perform the event drawing process recorded as a log instead of the organ recognition process and drawing process.
[0227] Furthermore, the control unit 201 may continuously monitor the load on the first calculation unit 205 and the second calculation unit 206, and, depending on the load on each calculation unit, execute the calculations that the second calculation unit 206 should perform in the first calculation unit 205.
[0228] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims, not in the sense described above, and all modifications within the sense and scope equivalent to the claims are intended.
[0229] For example, the information processing device 200 according to Embodiments 1 to 8 is configured to derive an integrated recognition result from the calculation results of two types of learning models, but it may also be configured to derive an integrated recognition result from the calculation results of three or more types of learning models. For example, by combining Embodiment 1 and Embodiment 2, the information processing device 200 may derive an integrated recognition result based on the calculation results of two types of learning models 310 and 320 for organ recognition and the calculation result of the learning model 330 for event recognition. Alternatively, by combining Embodiment 2 and Embodiment 3, the information processing device 200 may derive an integrated recognition result based on the calculation result of the learning model 310 for organ recognition, the calculation result of the learning model 330 for event recognition, and the calculation result of the learning model 340 for device recognition. The information processing device 200 is not limited to these combinations, and may also derive an integrated recognition result by appropriately combining Embodiments 1 to 8.
[0230] Furthermore, in this embodiment, the first arithmetic unit 205 performs calculations using one learning model, and the second arithmetic unit 206 performs calculations using the other learning model. However, the hardware that performs calculations using the learning models is not limited to the first arithmetic unit 205 and the second arithmetic unit 206. For example, the control unit 201 may perform calculations using two or more types of learning models, or a virtual machine may be prepared and the calculations using two or more types of learning models may be performed on the virtual machine.
[0231] Furthermore, although the order of processing is defined by a flowchart in this embodiment, some processing may be executed concurrently, and the order of processing may be changed.
[0232] The matters described in each embodiment can be combined with each other. Furthermore, the independent and dependent claims described in the claims can be combined with each other in any combination, regardless of the form of reference. In addition, the claims use a form in which claims referencing two or more other claims (multi-claim form), but are not limited to this. A form in which multi-claims referencing at least one multi-claim (multi-multi-claim) may also be used. [Explanation of symbols]
[0233] 10 Trocca 11 Laparoscopy 12 Energy Treatment Tools 13 forceps 110 Camera Control Unit (CCU) 120 Light source device 130 Display device 140 Recording device 200 Information Processing Devices 201 Control Unit 202 Storage section 203 Operation section 204 Input section 205 1st calculation section 206 2nd calculation section 207 Output section 208 Communications Department 310 First Learning Model 320 Second Learning Model PG1 Recognition Processing Program PG2 Display Processing Program
Claims
1. A first arithmetic unit that performs calculations using a first learning model that has been trained to output information for identifying organs included in a surgical field image in response to the input of a surgical field image, A second calculation unit performs calculations using a second learning model that has been trained to output information about the scene shown in the surgical field image in response to the input of the surgical field image. A derivation unit that derives information of structures related to the organ recognized from the calculation results of the first calculation unit and the scene recognized from the calculation results of the second calculation unit, An output unit that outputs the information of the derived structure and An information processing device equipped with the following features.
2. The derivation unit derives information about a structure related to an organ recognized from the calculation result of the first calculation unit and a scene recognized from the calculation result of the second calculation unit by referring to a structure table that stores information about a structure in association with information about an organ and information about a scene. The information processing apparatus according to claim 1.
3. The information of the structure includes the anatomical name of the structure. The information processing apparatus according to claim 2.
4. The information of the structure includes the anatomical position or direction of travel of the structure. The information processing apparatus according to claim 2.
5. The structure table stores information about the lesion area for each patient as information about the structure, The derivation unit derives patient-specific lesion information based on the organ information recognized from the calculation results of the first calculation unit and the scene information recognized from the calculation results of the second calculation unit. The information processing apparatus according to claim 2.
6. A first learning model, trained to output information for identifying organs included in a surgical field image in response to the input of the surgical field image, performs calculations. A second learning model, trained to output information about the scene shown in the surgical field image in response to the input of the surgical field image, performs calculations. Information on structures related to the organs recognized by the calculation results of the first learning model and the scenes recognized by the calculation results of the second learning model is derived. Output the information of the derived structure. An information processing method in which processing is performed by a computer.
7. A first learning model, trained to output information for identifying organs included in a surgical field image in response to the input of the surgical field image, performs calculations. A second learning model, trained to output information about the scene shown in the surgical field image in response to the input of the surgical field image, performs calculations. Information on structures related to the organs recognized by the calculation results of the first learning model and the scenes recognized by the calculation results of the second learning model is derived. Output the information of the derived structure. A computer program that causes a computer to perform a process.