Information processing system, endoscope system, method of operating the information processing system and program

The information processing system addresses the challenge of unclear surgical scenes in endoscope images by using trained models to enhance scene detection and display accuracy, ensuring stable support information display.

JP2026111585APending Publication Date: 2026-07-06OLYMPUS CORPORATION(JP) +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
OLYMPUS CORPORATION(JP)
Filing Date
2024-12-24
Publication Date
2026-07-06

AI Technical Summary

Technical Problem

The surgical scene in endoscope images is often not clearly distinguishable, leading to instability in determining scene changes, which affects the accuracy of support information display.

Method used

An information processing system using trained models to detect surgical scene information and support information, where the detection rate of consistent scene information exceeds a threshold, allowing for accurate scene selection and superimposed display on the endoscope image.

Benefits of technology

Improves the accuracy of surgical scene estimation and enhances the clarity of support information display by ensuring consistent scene detection and appropriate model selection based on detection rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026111585000001_ABST
    Figure 2026111585000001_ABST
Patent Text Reader

Abstract

Providing information processing systems that more reliably recognize surgical scenes from endoscopic images. [Solution] The information processing system 5 includes a processor 10 that performs display processing for the display 7. The processor 10 detects surgical scene information for the endoscopic image based on a first trained model 21. The processor 10 also stores the detected surgical scene information in a recording unit 30, and selects a surgical scene if the detection rate of surgical scene information indicating the same surgical scene is equal to or greater than a first threshold or exceeds a first threshold. The processor 10 also determines a second trained model 22 according to the selected surgical scene. The processor 10 also detects support information for the endoscopic image based on the determined second trained model 22, and performs processing to superimpose the detected support information onto the endoscopic image and display it on the display 7.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system, an endoscope system, an operation method of an information processing system, a program, and the like.

Background Art

[0002] A method of displaying support information such as diagnosis based on an endoscope image is known. Patent Document 1 discloses a method of identifying the change of a surgical scene using a change trigger of a surgical scene.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The surgical scene cannot always be clearly distinguished from the endoscope image. Therefore, if it is immediately determined that the surgical scene has changed because the endoscope image has changed, the stability of the determination of the surgical scene may decrease.

Means for Solving the Problems

[0005] One aspect of the present disclosure relates to an information processing system that includes a processor for display processing on a display, wherein the processor detects surgical scene information on an endoscope image based on a first trained model that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among a plurality of surgical scenes, based on first training data to which a first annotation, which is an annotation of a surgical scene, has been attached to a training image, the detected surgical scene information is stored in a recording unit, the detection rate of the same surgical scene information is equal to or greater than a first threshold or exceeds the first threshold, the surgical scene is selected, a second trained model is determined according to the selected surgical scene based on second training data to which a second annotation, which is an annotation of support information, has been attached to the training image, the support information is detected on the endoscope image based on the determined second trained model, and the detected support information is superimposed on the endoscope image and displayed on the display.

[0006] Other aspects of this disclosure relate to an endoscope system including the information processing system described above and an endoscope.

[0007] Another aspect of the present disclosure relates to a method for operating an information processing system that performs display processing on a display, comprising: detecting surgical scene information in an endoscope image based on a first trained model that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among a plurality of surgical scenes, based on first training data to which a first annotation, which is an annotation of a surgical scene, has been attached to a training image; accumulating the detected surgical scene information in a recording unit, and selecting a surgical scene when the detection rate of the same surgical scene information is equal to or greater than a first threshold or exceeds the first threshold; determining a second trained model that has been trained to output support information based on second training data to which a second annotation, which is an annotation of support information, has been attached to the training image, according to the selected surgical scene; detecting the support information in the endoscope image based on the determined second trained model; and displaying the detected support information superimposed on the endoscope image on the display.

[0008] Other aspects of this disclosure relate to a program that causes a computer to perform the following steps: detecting surgical scene information in an endoscope image based on a first trained model that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among a plurality of surgical scenes, based on first training data to which a first annotation, which is an annotation of a surgical scene, has been attached to a training image; accumulating the detected surgical scene information in a recording unit, and selecting a surgical scene when the detection rate of the same surgical scene information is equal to or greater than a first threshold or exceeds the first threshold; determining a second trained model that has been trained to output support information based on second training data to which a second annotation, which is an annotation of support information, has been attached to the training image, according to the selected surgical scene; detecting the support information in the endoscope image based on the determined second trained model; and displaying the detected support information superimposed on the endoscope image on a display. [Brief explanation of the drawing]

[0009] [Figure 1] A block diagram illustrating an example configuration of an endoscope system, including an information processing system. [Figure 2] A diagram illustrating the hardware involved in inference processing based on the first and second trained models. [Figure 3] A diagram illustrating the input and output data included in the first training data. [Figure 4] A diagram illustrating the input and output data included in the second training data set. [Figure 5] A flowchart illustrating an example of the processing method according to this embodiment. [Figure 6] A diagram illustrating an example of the relationship between surgical scenes and the first threshold. [Figure 7] A flowchart illustrating an example of the process involved in updating surgical scene information. [Figure 8] A diagram illustrating an example of accumulating surgical scene information. [Figure 9]A diagram for explaining the update of surgical scene information. [Figure 10] Another diagram for explaining the update of surgical scene information. [Figure 11] (A) is a diagram conceptually explaining the accuracy of surgical scene information. (B) is a diagram explaining an example of the accumulation of surgical scene information regarding the effect of the method of this embodiment. [Figure 12] A flowchart for explaining an example of the process related to the accumulation of surgical scene information. [Figure 13] A flowchart for explaining another example of the process related to the accumulation of surgical scene information. [Figure 14] A diagram for explaining an example of the inference by the first learned model. [Figure 15] A diagram for explaining an example of the relationship between the surgical scene and the second threshold. [Figure 16] (A) is a diagram explaining an example when the highest confidence level of the surgical scene does not exceed the second threshold. (B) is a diagram explaining another example when the highest confidence level of the surgical scene does not exceed the second threshold. [Figure 17] A flowchart for explaining an example of the process for determining support information. [Figure 18] A block diagram for explaining an example of a system for creating the first learned model. [Figure 19] A flowchart for explaining an example of the process related to the creation of the first teacher data. [Figure 20] (A)(B) are diagrams for explaining examples of displays related to surgical scene information. [Figure 21] (A)(B) are diagrams for explaining another examples of displays related to surgical scene information.

MODE FOR CARRYING OUT THE INVENTION

[0010] Hereinafter, preferred embodiments of the present disclosure will be described in detail. Note that the embodiments described below do not unduly limit the content described in the claims, and not all of the configurations described in the embodiments are essential constituent elements.

[0011] FIG. 1 is a block diagram for explaining a configuration example of the endoscope system 1 of the present embodiment. The endoscope system 1 includes an endoscope 3 and an information processing system 5.

[0012] The endoscope 3 is, for example, a rigid endoscope in which most of the insertion portion is rigid. Since the rigid endoscope is well known, detailed illustration of the configuration is omitted. Hereinafter, laparoscopic cholecystectomy (hereinafter abbreviated as lapacole) as an endoscopic surgery using the endoscope 3 as a rigid endoscope will be exemplified and described, but the method of the present embodiment is not precluded from being applied to other endoscopic surgeries or surgeries and examinations using a flexible endoscope.

[0013] The information processing system 5 includes a processor 10, a memory 20, and a recording unit 30. In the memory 20, a first learned model 21 and a second learned model 22 are stored.

[0014] The processor 10 of the present embodiment is configured by the following hardware. The hardware can include at least one of a circuit for processing digital signals and a circuit for processing analog signals. For example, the hardware can be composed of one or more circuit devices mounted on a circuit board or one or more circuit elements. The one or more circuit devices are, for example, ICs or the like. The one or more circuit elements are, for example, resistors, capacitors, or the like.

[0015] Furthermore, as shown in Figure 1, the information processing system 5 of this embodiment includes a memory 20 and a processor 10 that operates based on the information stored in the memory 20. The information includes, for example, programs and various types of data. The programs include, for example, a first trained model 21, a second trained model 22, and programs related to processing described later in Figure 5. The processor 10 can be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), etc. The memory 20 may be a semiconductor memory such as SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory), a register, a magnetic storage device such as a hard disk drive, or an optical storage device such as an optical disk drive. For example, the memory 20 stores instructions that can be read by a computer, and processing is realized when these instructions are executed by the processor 10. The instructions here may be instructions from an instruction set that constitutes a program, or instructions that instruct the hardware circuit of the processor 10 to operate. The memory 20 is also called a storage device.

[0016] Information processing system 5 is composed of information processing devices such as personal computers, servers, or processing devices dedicated to medical systems. In this case, the processor and memory included in the information processing device correspond to the processor 10 and memory 20 of information processing system 5. Alternatively, information processing system 5 may be a cloud system in which multiple information processing devices are connected by a network. In this case, the processor and memory included in the information processing devices constituting the cloud system correspond to the processor 10 and memory 20 of information processing system 5.

[0017] The processor 10 in this embodiment performs display processing on the display 7. For example, the tip of the endoscope 3 includes an imager (not shown, hereafter simply referred to as "imager"), and the processor 10 receives an image signal from the imager via a cable, input interface, etc. (not shown). The image signal received by the processor 10 is processed by an amplifier, A / D converter, etc. (not shown) to create an endoscopic image. The endoscopic image in this embodiment is a still image acquired from the imager at regular intervals, with the unit being a first hour. For convenience, the collection of still images captured by the imager is referred to as an endoscopic video. Furthermore, as shown in Figure 1, by connecting the information processing system 5 and the display 7, the imaged endoscopic image is output to the display 7 via an output unit 14, which will be described later. This allows, for example, in endoscopic surgery, the user to observe the biological tissue in the body cavity while performing procedures on the biological tissue with instruments. In this embodiment, the user refers to, for example, the surgeon handling the instruments, the scopist operating the endoscope 3, and all persons involved in the procedure. In this embodiment, support information may also be output to the display 7 along with the endoscopic image, details of which will be described later.

[0018] Surgical scene information is stored in the recording unit 30 by a method described later. Surgical scene information is data that identifies a surgical scene and is stored in the recording unit 30 as memory. Details of the surgical scene information will be described later. The recording unit 30 can be implemented using semiconductor memory or the like, similar to the memory 20 mentioned above. The processor 10 then performs the processing described later based on the surgical scene information stored in the recording unit 30. In this embodiment, the memory 20 and the recording unit 30 are shown separately, but the information processing system 5 may be configured so that a part of the memory 20 is used as the recording unit 30 and the surgical scene information is stored in the memory 20.

[0019] The first pre-trained model 21 is trained using the first training data, and the second pre-trained model 22 is trained using the second training data. In this embodiment, "training" specifically refers to "machine learning," or more precisely, "deep learning," but hereafter it will be referred to as "machine learning" or simply "training." At least a part of the first pre-trained model 21 in this embodiment includes a neural network. The same applies to the second pre-trained model 22. Although not shown in the diagram, it has an input layer into which data is input, an intermediate layer that performs calculations based on the output from the input layer, and an output layer that outputs data based on the output from the intermediate layer. Nodes in a given layer are connected to nodes in adjacent layers, and each connection is assigned a weighting coefficient. Each node multiplies the output of the preceding node by the weighting coefficient and calculates the sum of the multiplication results. Furthermore, each node adds a bias to the sum and applies an activation function to the sum result to obtain the output of that node. By sequentially executing this process from the input layer to the output layer, the output of the neural network is obtained. Various activation functions are known, such as the sigmoid function and the ReLU function, and these can be widely applied in this embodiment.

[0020] The neural network in this embodiment is more specifically a Convolutional Neural Network (CNN). Although not shown in the diagram, the CNN includes convolutional layers and pooling layers that perform convolutional operations. The convolutional layers perform filtering. The pooling layers perform pooling operations that reduce the size in the vertical and horizontal directions. More specifically, the CNN obtains its output by performing operations in a fully connected layer after multiple operations in the convolutional and pooling layers. A fully connected layer is a layer that performs operations when all nodes of the previous layer are mapped to the nodes of a given layer. Furthermore, as a more specific method for the CNN, known models such as EFFICIENTNET can be adopted.

[0021] Although not shown in Figure 1, the information processing system 5 of this embodiment includes hardware for the processor 10 to read the first trained model 21 or the second trained model 22, which includes the neural network described above, from the memory 20 and perform inference as described later. More specifically, as shown in Figure 2, the information processing system 5 further includes an input unit 12 and an output unit 14. The input unit 12 is, for example, an image interface that receives endoscopic images from an imager. The output unit 14 is an interface that transmits the inference results based on the first trained model 21 or the second trained model 22 to an external device. In this embodiment, the external device is specifically, for example, a display 7.

[0022] Furthermore, while Figures 1 and 2 illustrate one processor 10 and one memory 20, and show that a first trained model 21 and a second trained model 22 are stored in the memory 20, the configuration of the information processing system 5 in this embodiment is not limited to this. For example, the information processing system 5 may be configured to separately include a first module containing memory and a processor for performing inference based on the first trained model 21, and a second module containing memory and a processor for performing inference based on the second trained model 21. The first and second modules may then cooperate to perform the processing described later. The information processing system 5 may also further include processors that control the first and second modules, respectively. In the following explanation, to simplify the description, the main entity performing the processing in the information processing system 5 is simply referred to as the processor 10. Also, while Figures 1 and 2 illustrate that one first trained model 21 and one second trained model 22 are stored in the memory 20, in the information processing system 5 of this embodiment, multiple first trained models 21 may be stored in the memory 20. Similarly, in the information processing system 5 of this embodiment, multiple second trained models 22 may be stored in the memory 20. For example, by preparing a first trained model 21 according to the type of surgery, the method of this embodiment can be applied to multiple types of surgery with a single endoscopy system 1. Furthermore, by preparing a trained model 22 for each surgical scene, the method of this embodiment can be suitably applied to surgeries involving multiple surgical scenes while suppressing the excessive size of the network included in the second trained model 22.

[0023] The machine learning in this embodiment is supervised learning. In supervised learning, training data is a dataset that associates input data with correct labels, and is also called training data. More specifically, for example, the first trained model 21 is generated by performing machine learning on the first trained model 121 using the first training data. The input data for the first training data is endoscopic images used as training images. These training images are annotated with surgical scenes. In other words, the user manually annotates the images by looking at the endoscopic images used as training images and judging the surgical scenes based on tacit knowledge. That is, the first training data is a dataset in which the training images, which are the input data, are marked with correct labels indicating which surgical scene each training image belongs to. As a result, for example, when applying the method of this embodiment to a surgery that includes K surgical scenes from the 1st to the Kth surgical scenes, as shown in Figure 3, by inputting endoscopic images used as training images, the confidence level for the 1st to the Kth surgical scenes is output from each node. Although not shown in the diagram, the output layer of the neural network included in the first training model 121 has K nodes. The first node contains information representing the confidence that the class corresponding to the input endoscopic image belongs to class 1. The second to K nodes are similar, with each node containing information representing the confidence that the input endoscopic image belongs to class 2 to class K. Classes 1 to K correspond to the first to K surgical scenes. Furthermore, if the output layer is a known softmax layer, for example, the K outputs are a set of probability data that sum to 1. In other words, data representing the confidence of the class related to each node is output from each node. Then, by performing machine learning using a predetermined learning algorithm, for example, the weighting coefficients included in the neural network of the first training model 121 are optimized, resulting in the first trained model 21. The predetermined learning algorithm is, for example, a supervised learning algorithm using backpropagation.

[0024] Furthermore, when creating the first training data, the user may select a predetermined endoscopic image from the available endoscopic images and add annotations to it; details will be described later. Also, in this embodiment, the annotations related to the first training data may be conveniently referred to as the first annotation.

[0025] The second trained model 22 is generated by performing machine learning on the second training model 122 using the second training data. The input data for the second training data consists of endoscopic images used as training images. The dataset of this input data and the output data, which includes the training images with correct labels for supporting information, constitutes the second training data. Examples of supporting information include the names and locations of lesions present in the training images used as input data.

[0026] Furthermore, for example, positional and morphological information of the target area in a surgical scene may be used as supporting information. Positional and morphological information refers to information about position and information about shape. For example, as shown in A10 of Figure 4, an endoscopic image related to Calot triangulation in laparotomy is used as training image input data in a surgical scene. In laparotomy, the Rubiere groove, cystic duct, lower edge of S4, and common bile duct are used as landmarks during the surgery. In other words, in this embodiment, the Rubiere groove, cystic duct, lower edge of S4, and common bile duct are the target areas. The lower edge of S4 refers to the lower edge of the medial side of the left lobe of the liver. However, due to certain circumstances, it may be difficult for the user to clearly grasp this positional and morphological information from the endoscopic image. These circumstances include, for example, the ambiguity of the boundary of the end portion of the Rubiere groove, the cystic duct being covered with fat, the ambiguity of the boundary of the lower edge of S4, and the common bile duct being covered by the liver. In other words, the training image shown in A10 is an endoscopic image in which the positional and morphological information of the Rubiere groove, cystic duct, lower edge of S4, and common bile duct cannot be directly identified.

[0027] As shown in Figure 4, A20, the output data is an endoscopic image annotated by attaching the tags shown in A21, A22, A23, and A24 to the training image A10. The tag shown in A21 indicates the Rubiere groove, the tag shown in A22 indicates the cystic duct, the tag shown in A23 indicates the lower edge of S4, and the tag shown in A24 indicates the common bile duct. In other words, the second training model 122 in Figure 4 is machine-trained using a dataset of an endoscopic image in Calot triangulation of laparoscopy and an image in which the positional shape information of the Rubiere groove, cystic duct, lower edge of S4, and common bile duct is detected as supporting information and superimposed on the endoscopic image as second training data. In other words, when the positional shape method is used as supporting information, the neural network included in the second trained model 22 is configured such that the endoscopic image is input to the input layer and detection information indicating the positional shape of the target region is output from the output layer.

[0028] Annotation of the second training data is performed, for example, by a user proficient in laparoscopic cholangiocarcinoma looking at the training image shown in A10, identifying the Rubiere sulcus, cystic duct, lower edge of S4, and common bile duct based on tacit knowledge, and tagging each of them. In other words, the output data for the second training data is, more specifically, a set of map data with flagged pixels in the tagged regions shown in A21, map data with flagged pixels in the tagged regions shown in A22, map data with flagged pixels in the tagged regions shown in A23, map data with flagged pixels in the tagged regions shown in A24, and training image data.

[0029] While annotation may be performed by an experienced user on all training images, the system may also be configured to automatically annotate training images via tracking for frames following the image frame corresponding to the training image annotated by the user (hereinafter simply referred to as "frame"). A detailed method is disclosed, for example, in International Publication No. 2020 / 110278. By performing machine learning using the second training data thus created, the weighting coefficients included in the neural network of the second training model 122 are optimized, resulting in the second trained model 22. In this embodiment, the annotation related to the second training data may be conveniently referred to as the second annotation.

[0030] An example of the processing method of this embodiment will be explained using the flowchart in Figure 5. The processing in Figure 5 is repeated, for example, by a timer interrupt. To make the method of this embodiment easier to understand, the processing in Figure 5 is assumed to be performed in response to the timing when surgical scene information is accumulated in the recording unit 30. For example, if the aforementioned first time is 0.2 seconds, then 5 pieces of surgical scene information are accumulated per second, and the processing in Figure 5 is performed 5 times.

[0031] In Figure 5, the processor 10 recognizes the surgical scene (step S100). For example, the processor 10 reads the first pre-trained model 21, uses the acquired endoscopic image as input data, and infers which surgical scene the acquired endoscopic image belongs to. In other words, step S100 can be described as detecting the surgical scene using the first pre-trained model 21.

[0032] Subsequently, the processor 10 updates the surgical scene information based on the surgical scene recognized in step S100 (step S200). Updating the surgical scene information in step S200 means deciding whether to change the surgical scene information and changing or maintaining the surgical scene information, details of which will be described later. Note that if the number of surgical scene information entries in the recording unit 30 has not reached the upper limit, the processor 10 may omit steps S200 and beyond.

[0033] Subsequently, the processor 10 selects a surgical scene based on the surgical scene information updated in step S200 (step S300). As will be described in detail later, in this embodiment, the surgical scene recognized in step S100 does not necessarily match the surgical scene selected in step S300.

[0034] Subsequently, the processor 10 determines the support information based on the surgical scene selected in step S300 (step S400). For example, as mentioned above, if the memory 20 contains multiple second trained models 22, the processor 10 selects the second trained model 22 corresponding to the surgical scene selected in step S300. Step S400 may also include a process to decide not to display the support information, details of which will be described later.

[0035] While users can decide how to classify surgical scenes as they see fit, one example is to associate surgical scenes with the phases of the surgery being performed. Figure 6 shows examples of commonly classified laparoscopic cholecystic phases (hereinafter sometimes simply referred to as "phases"). Phase 0 represents surgical scenes classified as "other" among the laparoscopic cholecystic phases. Phase 1 represents surgical scenes classified as "preparation" among the laparoscopic cholecystic phases. Phase 2 represents surgical scenes classified as "Calot's triangle dissection" among the laparoscopic cholecystic phases. Phase 3 represents surgical scenes classified as "clipping and dissection" among the laparoscopic cholecystic phases. Phase 4 represents surgical scenes classified as "gallbladder dissection" among the laparoscopic cholecystic phases. Phase 5 represents surgical scenes classified as "gallbladder bagging and gallbladder retrieval" among the laparoscopic cholecystic phases. Phase 6 represents surgical scenes classified as "washing and coagulation" among the laparoscopic cholecystic phases.

[0036] Furthermore, in subsequent illustrations, "Phase 0" may be abbreviated simply as "P0," "Phase 1" as "P1," "Phase 2" as "P2," and "Phase 3" as "P3." Similarly, in subsequent illustrations and explanations, "Phase 4" may be abbreviated simply as "P4," "Phase 5" as "P5," and "Phase 6" as "P6." Also, "P0" in Figures 8 and later may indicate either Phase 0 as a surgical scene or information about the surgical scene corresponding to Phase 0. The same applies to "P1" through "P6."

[0037] Figure 6 also shows that a first threshold is set for each procedural phase. The first threshold for Phase 0 is 0.84, for Phase 1 it is 0.72, for Phase 2 it is 0.72, for Phase 3 it is 0.36, for Phase 4 it is 0.68, for Phase 5 it is 0.40, and for Phase 6 it is 0.52.

[0038] A more detailed example of the process in step S200 will be explained using the flowchart in Figure 7. The processor 10 stores the surgical scene information in the recording unit 30 based on the result of step S100 (step S210). The processor 10 then determines whether or not there is any new surgical scene information that exceeds the first threshold (step S270). More specifically, in step S270, the processor 10 determines whether or not the detection rate of the most frequent surgical scene information among the detection rates of the surgical scene information recorded in the recording unit 30 exceeds the first threshold. Therefore, for example, even if new surgical scene information is detected, if the detection rate of the detected new surgical scene information is not the most frequent, the processor 10 determines NO in step S270 for that new surgical scene information.

[0039] If the processor 10 determines that there is new surgical scene information that exceeds the first threshold (YES in step S270), it updates the information with the new surgical scene information (step S280) and terminates the flow. On the other hand, if the processor 10 determines that there is no new surgical scene information that exceeds the first threshold (NO in step S270), it maintains the current surgical scene (step S290) and terminates the flow. In other words, if the detection rate of each surgical scene information recorded in the recording unit 30 is all below the first threshold, the processor 10 determines NO in step S270.

[0040] In this embodiment, step S270 is indicated as a determination of whether the detection rate of surgical scenes is equal to or greater than the first threshold. For example, if the detection rate of surgical scenes is 0.8 and the first threshold is 0.8, the processor 10 determines YES in step S270. However, the processing in step S270 is not limited to this. For example, the processor 10 may determine YES in step S270 if the detection rate of surgical scenes exceeds the first threshold, and NO in step S270 if the detection rate of surgical scenes is equal to or less than the first threshold. The user can determine the determination criteria as appropriate. In this case, for example, if the detection rate of surgical scenes is 0.8 and the first threshold is 0.8, the processor 10 determines NO in step S270. The same applies to steps S220 and S620, which will be described later.

[0041] The process in step S200 is conceptually explained. For example, when surgery begins, the endoscopic system 1 is activated and the imager starts acquiring endoscopic images. Then, the process shown in Figure 5 is started at a predetermined timing. The processor 10 then infers that the acquired endoscopic image in step S100 in the first frame is a surgical scene corresponding to phase 0. As a result, the processor 10 stores the surgical scene information corresponding to phase 1 in the recording unit 30.

[0042] In the following explanation, it will be assumed that the recording unit 30 stores 10 surgical scene information entries. However, this is merely an example for the sake of explanation, and the number of surgical scene information entries stored in the recording unit 30 is not limited to 10. In Figure 8, the set of 10 boxes arranged horizontally on the page conceptually represents the number of surgical scene information entries that can be stored in the recording unit 30. Furthermore, the leftmost box represents the oldest surgical scene information entry stored at the time of the event, and the entries to the right of the page represent entries stored at more recent times. In other words, the surgical scene information arranged horizontally on the page is arranged chronologically, and the rightmost surgical scene information is the most recent surgical scene information acquired by the processor 10 at the time of the event. Also, for the sake of explanation, the oldest surgical scene information stored in the recording unit 30 will be referred to as the "1st surgical scene information entry." Similarly, for example, the newest surgical scene information stored in the recording unit 30 will be referred to as the "10th surgical scene information entry." Similarly, among the surgical scene information stored in the recording unit 30, the second oldest surgical scene information is referred to as the "second oldest surgical scene information," and among the surgical scene information stored in the recording unit 30, the second newest surgical scene information is referred to as the "ninth oldest surgical scene information."

[0043] Subsequently, in the second frame, the processor 10 infers through step S100 that the acquired endoscopic image is a surgical scene corresponding to phase 1. As a result, the processor 10 further stores surgical scene information corresponding to phase 1 in the recording unit 30. Similarly, in the third, fourth, fifth, sixth, seventh, eighth, ninth, and tenth frames, the processor 10 infers through step S100 that the acquired endoscopic image is a surgical scene corresponding to phase 1. As a result, the maximum number of 10 pieces of surgical scene information are stored in the recording unit 30.

[0044] Subsequently, in the 11th frame, step S100 is performed, and the processor 10 infers that the acquired endoscopic image is a surgical scene corresponding to phase 1. In this case, the processor 10 discards the first surgical scene information, i.e., the surgical scene information accumulated in the first frame, shifts the 2nd to 10th surgical scene information to the 1st to 9th surgical scene information, and newly accumulates the surgical scene information corresponding to phase 1 acquired in the 11th frame as the 10th surgical scene information. Subsequently, in the 12th frame, step S100 is performed, and the processor 10 infers that the acquired endoscopic image is a surgical scene corresponding to phase 1. In this case, the processor 10 discards the first surgical scene information, i.e., the surgical scene information accumulated in the second frame, and newly accumulates the surgical scene information corresponding to phase 1 acquired in the 12th frame as the 10th surgical scene information. More generally, if V(n,k) is the kth surgical scene information (k is one of 10 integers from 0 to 9) stored in the recording unit 30 in the nth frame (where n is an integer greater than or equal to 10), then it can be said that the processor 10 performs the process of setting V(n,k) = V(n-1,k+1).

[0045] The process in step S270 will be conceptually explained. As a premise, as shown in Figure 9, in the (N-8)th frame, surgical scene information corresponding to phase 1 is stored in the recording unit 30, and in step S300, the surgical scene is determined to be phase 1.

[0046] In the (N-7)th frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to phase 2. As a result, as shown in B1, the surgical scene information corresponding to phase 2 is stored as the 10th surgical scene information. Since the total number of boxes corresponding to the recording unit 30 is 10, and the number of boxes storing the surgical scene information for phase 2 is 1, the detection rate of the surgical scene information for phase 2 is 0.1, which is not the mode. Therefore, the processor 10 determines NO in step S270 and maintains the current surgical scene as phase 1 in step S290. As a result, the processor 10 selects phase 1 as the surgical scene in step S300.

[0047] In the subsequent (N-6)th frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to phase 2. As a result, as shown in B2 of Figure 9, the surgical scene information corresponding to phase 2 is accumulated as the 9th and 10th surgical scene information. Therefore, the detection rate of surgical scenes corresponding to phase 2 is 0.2, which is not the mode, so the processor 10 determines NO in step S270 and maintains the current surgical scene as phase 1 in step S290. Consequently, the processor 10 selects phase 1 as the surgical scene in step S300.

[0048] In the subsequent (N-5)th frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to phase 2. As a result, as shown in B3 of Figure 9, the surgical scene information corresponding to phase 2 is accumulated as the 8th, 9th, and 10th surgical scene information. Therefore, the detection rate of surgical scenes corresponding to phase 2 is 0.3, which is not the mode, so the processor 10 determines NO in step S270 and maintains the current surgical scene as phase 1 in step S290. Consequently, the processor 10 selects phase 1 as the surgical scene in step S300.

[0049] In the subsequent (N-4)th frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to phase 2. As a result, as shown in B4 of Figure 9, the surgical scene information corresponding to phase 2 is accumulated as the 7th, 8th, 9th, and 10th surgical scene information. Therefore, the proportion of surgical scene information for phase 2 in the (N-4)th frame is 0.4, and the detection rate of surgical scenes corresponding to phase 2 is not the mode. Thus, the processor 10 determines NO in step S270 and maintains the current surgical scene as phase 1 in step S290. Consequently, the processor 10 selects phase 1 as the surgical scene in step S300.

[0050] In the subsequent (N-3)th frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to phase 2. As a result, as shown in B5 of Figure 10, the surgical scene information corresponding to phase 2 is accumulated as the 6th, 7th, 8th, 9th, and 10th surgical scene information. Consequently, the detection rate of the phase 2 surgical scene information in the (N-3)th frame is 0.5, which is the mode along with the surgical scene information corresponding to phase 1. However, since the detection rate of the phase 2 surgical scene information, 0.5, is smaller than the first threshold corresponding to phase 2, which is 0.72, the processor 10 determines NO in step S270 and maintains the current surgical scene as phase 1 in step S290. As a result, the processor 10 selects phase 1 as the surgical scene in step S300.

[0051] In the subsequent (N-2)th frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to phase 2. As a result, as shown in B6 of Figure 10, the surgical scene information corresponding to phase 2 is accumulated as the 5th, 6th, 7th, 8th, 9th, and 10th surgical scene information. Consequently, the proportion of surgical scene information corresponding to phase 2 in the (N-2)th frame is 0.6, which is the mode. However, since the detection rate of phase 2 surgical scene information, 0.6, is smaller than the first threshold corresponding to phase 2, which is 0.72, the processor 10 determines NO in step S270 and maintains the current surgical scene as phase 1 in step S290. As a result, the processor 10 selects phase 1 as the surgical scene in step S300.

[0052] In the subsequent (N-1)th frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to phase 2. As a result, as shown in B7 of Figure 10, the surgical scene information corresponding to phase 2 is accumulated as the 4th, 5th, 6th, 7th, 8th, 9th, and 10th surgical scene information. Consequently, the proportion of surgical scene information for phase 2 in the (N-1)th frame is 0.7, which is the mode. However, since the detection rate of surgical scene information for phase 2, 0.7, is smaller than the first threshold corresponding to phase 2, which is 0.72, the processor 10 determines NO in step S270 and maintains the current surgical scene as phase 1 in step S290. As a result, the processor 10 selects phase 1 as the surgical scene in step S300.

[0053] In the subsequent Nth frame, the processor 10 infers that the endoscopic image acquired in step S100 corresponds to the surgical scene information corresponding to Phase 2. As a result, as shown in B8 of Figure 10, the surgical scene information corresponding to Phase 2 is accumulated as the 3rd, 4th, 5th, 6th, 7th, 8th, 9th, and 10th surgical scene information. Consequently, the proportion of Phase 2 surgical scene information in the Nth frame is 0.8, which is the mode. Furthermore, since the detection rate of Phase 2 surgical scene information, 0.8, is greater than the first threshold corresponding to Phase 2, which is 0.72, the processor 10 determines YES in step S270 and updates the current surgical scene to Phase 2 in step S280. Thus, the first threshold has technical significance as a criterion for determining whether the first trained model 21 is confidently judging the surgical scene information.

[0054] Subsequently, in step S300, the processor 10 selects phase 2 as the surgical scene. Then, in step S400, the processor 10 decides to output support information based on inference by the trained model 22 corresponding to phase 2. Finally, in step S500, the processor 10 outputs the endoscopic image acquired in step S100 and the support information determined in step S400 to the display 7. As a result, the endoscopic image and support information are displayed superimposed on the display 7.

[0055] As described above, the information processing system 5 of this embodiment includes a processor 10 that performs display processing on the display 7. The processor 10 detects surgical scene information on the endoscopic image based on a first trained model 21 that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among multiple surgical scenes, based on first training data to which first annotations, which are annotations of surgical scenes, have been attached to the training image. The processor 10 also stores the detected surgical scene information in the recording unit 30 and selects the surgical scene if the detection rate of surgical scene information indicating the same surgical scene is equal to or greater than a first threshold or exceeds a first threshold. The processor 10 also determines a second trained model 22 that has been trained to output support information, based on second training data to which second annotations, which are annotations of support information, have been attached to the training image, according to the selected surgical scene. The processor 10 also detects support information on the endoscopic image based on the determined second trained model 22 and displays the detected support information superimposed on the endoscopic image on the display 7.

[0056] As described above, the information processing system 5 of this embodiment includes the processor 10, and can detect support information based on the second trained model 22 associated with the surgical scene, and can superimpose the detected support information onto the endoscopic image and display it on the display 7. Furthermore, because the information processing system 5 of this embodiment includes the processor 10, it can store surgical scene information detected based on the first trained model 21 for the endoscopic image in the recording unit 30. In addition, because the information processing system 5 of this embodiment includes the processor 10, it can improve the accuracy of surgical scene estimation by selecting a surgical scene associated with the surgical scene information when the detection rate of identical surgical scene information stored in the recording unit 30 is equal to or exceeds a first threshold.

[0057] As mentioned above, in surgeries where procedures are classified based on the content of the procedure, a method is known to associate the procedure with the surgical scene and predict the surgical scene from endoscopic images. However, it is not always possible to uniquely grasp the procedure from endoscopic images. In other words, the boundaries between adjacent procedures in surgery are not always clear. Conceptually, as shown in Figure 11(A), it is thought that the certainty of each surgical scene may change smoothly over time. For example, as shown in C1 of Figure 11(A), the certainty of the surgical scene corresponding to phase 2 is high between timing t1 and timing t2. Therefore, if the endoscopic images acquired during the period shown in C1 are used as input data for the first trained model 21, it is expected that the output will have a very high degree of confidence that it is a surgical scene corresponding to phase 2. Similarly, as shown in C3 of Figure 11(A), the probability of the surgical scene corresponding to phase 3 is high between timing t3 and timing t4. Therefore, if the endoscopic images acquired during the period shown in C3 are used as input data for the first trained model 21, it is expected that the output will have a very high degree of confidence that the surgical scene corresponds to phase 3.

[0058] On the other hand, as shown in C2, the accuracy of both Phase 2 and Phase 3 is low between timing t2 and timing t3. Therefore, if endoscopic images acquired during the period shown in C2 are used as input data, the confidence level for the surgical scene corresponding to Phase 2 and the confidence level for the surgical scene corresponding to Phase 3 will not be very high as output.

[0059] For example, although not shown in the diagram, an endoscopic image showing the gallbladder infundibulum being grasped with grasping forceps has a high degree of confidence that it is Phase 2 (Calot triangulation), and an endoscopic image showing the cystic duct being clipped with clippers has a high degree of confidence that it is Phase 3 (clipping, cutting). However, an endoscopic image showing the grasping forceps independently may be acquired because it was captured at a point in the middle of the Calot triangulation, or it may be acquired because it was captured after the Calot triangulation has been completed. It is difficult to accurately determine whether it is Phase 2 or Phase 3 from such an endoscopic image. If the surgical scene with the highest confidence level as an output result is mechanically adopted, then in the period shown in C2, there is a possibility that Phase 2 and Phase 3 will be mixed in the output result of the trained model 21. As a result, the second trained model 22, which is suitable for the surgical scene, may not be selected, and the support information may not be displayed accurately. In the conventional method, for example, as shown in C4 of Figure 11(B), when a surgical scene corresponding to Phase 3 is detected, the information processing system 5 immediately operates to display support information corresponding to Phase 3.

[0060] In this respect, by applying the method of this embodiment, surgical scene information is stored in the recording unit 30, and when the number of stored identical surgical scene information exceeds a first threshold, the second trained model 22 corresponding to the surgical scene corresponding to the surgical scene information that exceeds the first threshold is read out, thereby enabling the display of more appropriate support information. For example, as shown in C4 of Figure 11(B), at the timing when a surgical scene corresponding to phase 3 is detected, the mode of the surgical scene information stored in the recording unit 30 is the surgical scene information corresponding to phase 2, so the surgical scene corresponding to phase 2 is selected.

[0061] Furthermore, when a user creates first training data based on endoscopic images from Phase 2, for example, they may use endoscopic images acquired during the period shown in C1 as input data and perform surgical scene annotation on those endoscopic images. Similarly, when a user creates first training data based on endoscopic images from Phase 3, for example, they may use endoscopic images acquired during the period shown in C3 as input data and perform surgical scene annotation on those endoscopic images. By doing so, it is possible to create first training data based on endoscopic images that best match the characteristics of each surgical scene, thereby improving the inference accuracy of the first trained model 21.

[0062] Furthermore, although not shown in the diagram, an endoscopic image primarily showing the cystic duct and common bile duct is considered to belong to Phase 1 (preparation) if it is intended to provide an overview of the entire Calot triangle, but it is considered to belong to Phase 2 (Calot triangle development) if it is intended to indicate the start of Calot triangle development. Therefore, it is difficult to determine whether an endoscopic image primarily showing the cystic duct and common bile duct is in Phase 1 or Phase 2. In such cases, applying the method of this embodiment is useful.

[0063] Furthermore, the method of this embodiment may be implemented as an endoscope system 1. That is, the endoscope system 1 of this embodiment includes the information processing system 5 described above and the endoscope 3. By doing so, the same effects as described above can be obtained.

[0064] Furthermore, the method of this embodiment may be implemented as an operation method for the information processing system 5. In other words, this embodiment relates to an operation method for the information processing system 5 that performs display processing on the display 7. The operation method of the information processing system 5 includes the step of detecting surgical scene information for an endoscopic image based on a first trained model 21 that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among multiple surgical scenes, based on first training data to which a second annotation, which is an annotation of a surgical scene, has been attached to a training image. The operation method of the information processing system 5 of this embodiment further includes the step of accumulating the detected surgical scene information in the recording unit 30 and selecting the surgical scene when the detection rate of surgical scene information indicating the same surgical scene is equal to or greater than a first threshold or exceeds a first threshold. The operation method of the information processing system 5 of this embodiment further includes the step of determining a second trained model 22 that has been trained to output support information, based on second training data to which a second annotation, which is an annotation of support information, has been attached to a training image, according to the selected surgical scene. Furthermore, the operation method of the information processing system 5 in this embodiment further includes the steps of detecting support information for the endoscopic image based on the determined second trained model 22, and displaying the detected support information superimposed on the endoscopic image on the display 7. By doing so, the same effects as described above can be obtained.

[0065] Furthermore, the method of this embodiment may be implemented as a program. Specifically, the program of this embodiment causes the computer to perform the step of detecting surgical scene information on an endoscopic image based on a first trained model 21 that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among multiple surgical scenes, based on first training data to which a first annotation, which is an annotation of a surgical scene, has been attached to the training image. The program of this embodiment also causes the computer to perform the step of storing the detected surgical scene information in a recording unit 30 and selecting the surgical scene if the detection rate of surgical scene information indicating the same surgical scene is equal to or greater than a first threshold or exceeds a first threshold. Furthermore, the program of this embodiment causes the computer to perform the step of determining a second trained model 22 that has been trained to output support information, based on second training data to which a second annotation, which is an annotation of support information, has been attached to the training image, according to the selected surgical scene. Furthermore, the program of this embodiment causes the computer to perform the steps of detecting support information for the endoscopic image and displaying the detected support information superimposed on the endoscopic image on the display 7, based on the determined second trained model 22. By doing so, the same effects as described above can be obtained.

[0066] Furthermore, the processor 10 may determine a second trained model 22, which is trained to output support information based on a second training data to which a second annotation, representing positional and morphological information of the target region in the selected surgical scene, is attached to the training image, according to the selected surgical scene. In this way, positional and morphological information can be displayed with high accuracy as support information.

[0067] Furthermore, the processor 10 may detect surgical scene information based on the first trained model 21 at predetermined intervals and use a predetermined number of surgical scene information obtained at the most recent predetermined number of timings, including the current timing, to determine the detection ratio of each surgical scene. The processor 10 may also select a surgical scene corresponding to the mode of the calculated detection ratio of each surgical scene if the mode is equal to or exceeds the first threshold. In this way, an information processing system 5 can be constructed in which surgical scene information obtained at the current timing is not immediately selected, and surgical scene information is selected using the mode and the first threshold.

[0068] Furthermore, if the detection rate of each of the requested surgical scenes is below or less than the first threshold, the processor 10 does not need to update the selection result of the surgical scene. This prevents the surgical scene from being changed when the determination of the surgical scene is ambiguous.

[0069] Furthermore, the first threshold may be set differently for each surgical scene. For example, when applying the method of this embodiment to a surgery that includes a first surgical scene and a second surgical scene, a first first threshold may be set for the first surgical scene, and a second first threshold different from the first first threshold may be set for the second surgical scene. For example, as mentioned above, it may be difficult to determine from the endoscopic image that it is the first surgical scene, but it is not so difficult to determine from the endoscopic image that it is the second surgical scene, in which case the second first threshold may be set lower than the first first threshold. For example, if the detection rate of the first surgical scene becomes the mode, the processor 10 may perform a process in step S270 to compare the detection rate of the first surgical scene with the first first threshold. Similarly, if the detection rate of the second surgical scene becomes the mode, the processor 10 may perform a process in step S270 to compare the detection rate of the second surgical scene with the second first threshold. In this way, the processor 10 selects the first surgical scene when the detection rate of first surgical scene information indicating the first surgical scene is equal to or greater than the first first threshold, or when it exceeds the first first threshold, and selects the second surgical scene when the detection rate of second surgical scene information indicating the second surgical scene is equal to or greater than the second first threshold, or when it exceeds the second first first threshold. In this way, an information processing system 5 can be constructed in which the first threshold is appropriately set according to the type of surgical scene. For example, in surgical scenes that are easy to detect from endoscopic images, setting the first threshold low allows the timing of updating the surgical scene information to change the surgical scene information to be earlier.

[0070] The method of this embodiment is not limited to the above, and various modifications can be made, such as by adding other features. For example, step S210 in Figure 7 may be as shown in the example processing shown in the flowchart of Figure 12. In Figure 12, the processor 10 determines whether the confidence level of the highest surgical scene is equal to or greater than the second threshold (step S220). If the confidence level of the highest surgical scene is equal to or greater than the second threshold (YES in step S220), the processor 10 stores the surgical scene information relating to the surgical scene with the highest confidence level in the recording unit 30 (step S230). On the other hand, if the confidence level of the highest surgical scene is not equal to or greater than the second threshold (NO in step S220), the processor 10 stores information indicating that the detection of surgical scene information is invalid in the recording unit 30 (step S240).

[0071] For example, step S210 in Figure 7 may be performed as shown in the flowchart in Figure 13. Although the process similar to that in Figure 12 will not be explained, in Figure 13, if the confidence level of the highest surgical scene is not above the second threshold (NO in step S220), the processor 10 stores the surgical scene information of the previous frame in the recording unit 30 (step S250).

[0072] More specifically, for example, in step S100, the endoscopic image was input as input data to the first trained model 21, and inference as shown in Figure 14 was performed. In Figure 14, the inference results output are a confidence score of 0.02 for phase 0, 0.07 for phase 1, 0.83 for phase 2, 0.03 for phase 3, 0.02 for phase 4, 0.02 for phase 5, and 0.01 for phase 6.

[0073] Here, the second threshold is set to a constant value of 0.80 regardless of the phase (i.e., surgical scene). In this case, the processor 10 determines YES in step S220 because the confidence level for phase 2 is 0.83, which is 0.80 or higher, and in step S230, it stores the surgical scene information related to phase 2 in the recording unit 30. In other words, the second threshold has technical significance as a criterion for determining whether the surgical scene information, as an inference result based on the first trained model 21, is worth storing in the recording unit 30.

[0074] In this way, in the information processing system 5 of this embodiment, the processor 10 stores surgical scene information in the recording unit 30 when the confidence level of a surgical scene detected based on the first trained model 21 is equal to or greater than the second threshold, or when it exceeds the second threshold. In this way, a mechanism can be constructed in which surgical scenes are selected based on the accumulation of surgical scene information with a high confidence level.

[0075] Furthermore, for example, the value related to the second threshold may be set differently for each surgical scene. For example, when applying the method of this embodiment to a surgery that includes a first surgical scene and a second surgical scene, a first second threshold may be set for the first surgical scene, and a second second threshold different from the first second threshold may be set for the second surgical scene. In this case, if the confidence level of the first surgical scene is the highest by inference based on the first trained model 21, and that confidence level is equal to or greater than the first second threshold, the processor 10 should determine YES in step S220. Similarly, if the confidence level of the second surgical scene is the highest by inference based on the first trained model 21, and that confidence level is equal to or greater than the second second threshold, the processor 10 should determine YES in step S220.

[0076] More specifically, Figure 15 shows examples of setting second thresholds for each of the procedure phases 0 to 6 described in Figure 6. In Figure 15, the second threshold for phase 0 is 0.79, for phase 1 it is 0.82, for phase 2 it is 0.85, for phase 3 it is 0.78, for phase 4 it is 0.82, for phase 5 it is 0.80, and for phase 6 it is 0.85.

[0077] For example, let's assume that step S220 was performed in the Mth frame, using the inference results mentioned above in Figure 14 and the second threshold shown in Figure 15. In this case, the surgical scene with the highest confidence is Phase 2, but the confidence level for Phase 2 in Figure 14 is 0.83, which is less than the second threshold of 0.85 shown in Figure 15. Therefore, the processor 10 determines NO in step S220.

[0078] In this embodiment of the information processing system 5, the processor 10 stores first surgical scene information in the recording unit 30 when the confidence level of the first surgical scene detected based on the first trained model 21 is equal to or greater than the first second threshold, or when it exceeds the first second threshold. The processor 10 also stores second surgical scene information in the recording unit 30 when the confidence level of the second surgical scene detected based on the first trained model 21 is equal to or greater than the second second threshold, or when it exceeds the second second threshold. In this way, an information processing system 5 can be constructed that determines the final surgical scene information while taking into account the difficulty of judging the surgical scene. For example, in situations where judging the surgical scene is easy, it is convenient to update the surgical scene information more quickly by lowering the second threshold of the surgical scene in question.

[0079] The case where the processor 10 determines NO in step S220 will be explained in more detail. For example, let's assume that in the (M-1) frame, surgical scene information corresponding to phase 1 has been stored in the recording unit 30. Since the recording unit 30 has 10 pieces of surgical scene information corresponding to phase 1 stored, in the (M-1) frame, the processor 10 selects phase 1 in step S300. For example, if step S210 is performed using the processing example shown in Figure 12, the processor 10 determines NO in step S220 and proceeds to step S240. As a result, as shown in Figure 16(A), the oldest surgical scene information (surgical scene information corresponding to phase 1) is deleted, and information indicating invalid surgical scene detection (shown as "P7" in Figure 16 for convenience), shown as D1, is newly stored. Note that the information indicating invalid surgical scene detection is not the surgical scene information itself, so it is not subject to step S270. For example, even if 10 pieces of information indicating invalid surgical scene detection are stored in the recording unit 30, the processor 10 will determine NO in step S270 and maintain the surgical scene information that was updated before the information indicating invalid surgical scene detection was stored in the recording unit 30.

[0080] Also, for example, when step S210 is performed according to the processing example shown in FIG. 13, if the processor 10 determines NO in step S220, it performs step S250. As a result, as shown in FIG. 16(B), the oldest surgical scene information (surgical scene information corresponding to phase 1) is deleted. Also, as shown in D2, the surgical scene information (surgical scene information corresponding to phase 1) of the (M-1)th frame, which is the previous frame, is accumulated. More generally, when the confidence level in the s-th surgical scene when step S100 is performed in the n-th frame is C(n, s) and the second threshold value in the s-th surgical scene is T(s), the processor 10 sets V(n, 9) = S(s) when the relationship C(n, s) ≥ T(s) is satisfied. V(n, 9) is the 10th surgical scene information in the recording unit 30 in the n-th frame. S(s) is the surgical scene information corresponding to the s-th surgical scene. On the other hand, the processor 10 sets V(n, 9) = V(n, 8) when the relationship C(n, s) < T(s) is satisfied. V(n, 9) is the 9th surgical scene information in the recording unit 30 in the n-th frame.

[0081] Thus, in the information processing system 5 of the present embodiment, when the confidence level of the surgical scene information detected based on the first learned model 21 is less than or equal to the second threshold value or less than the second threshold value, the processor 10 causes the recording unit 30 to accumulate information indicating invalid detection of the surgical scene information or accumulates the same surgical scene information as the previous frame as the surgical scene information of the current frame in the recording unit 30. By doing so, in a situation where it is difficult to determine the surgical scene from the endoscopic image, an information processing system 5 can be constructed that does not change the current surgical scene.

[0082] As mentioned above, in the information processing system 5 of this embodiment, multiple second trained models 22 may be included in the memory 20 depending on the surgical scene. However, it is not necessary to prepare a second trained model 22 that corresponds to all surgical scenes; a second trained model 22 that corresponds to some surgical scenes may be included in the memory 20. In this case, for example, step S400 in Figure 5 may be more specifically as shown in the example processing shown in the flowchart in Figure 17.

[0083] In Figure 17, the processor 10 performs a process to determine the surgical scene selected in step S300 (step S410). If a surgical scene related to phase 1 or phase 2 is selected in step S300, the processor 10 performs a process to display support information based on the second trained model 22. More specifically, for example, if a surgical scene related to phase 1 is selected in step S300, the processor 10 reads the second trained model 22 corresponding to the surgical scene related to phase 1 from memory 20 and performs inference processing based on the second trained model 22. Similarly, if a surgical scene related to phase 2 is selected in step S300, the processor 10 reads the second trained model 22 corresponding to the surgical scene related to phase 2 from memory 20 and performs inference processing based on the second trained model 22.

[0084] On the other hand, if the processor 10 selects a surgical scene related to phase 0, phase 3, phase 4, phase 5, or phase 6 in step S300, it terminates the flow. In other words, if the processor 10 selects a surgical scene related to phase 0, phase 3, phase 4, phase 5, or phase 6 in step S300, it does not perform the process of reading the second trained model 22, and therefore it is decided not to display the support information in step S400. Thus, in the information processing system 5 of this embodiment, the processor 10 decides whether or not to display the support information detected by the second trained model 22 depending on which of the multiple surgical scenes the selected surgical scene is. In this way, an information processing system 5 can be constructed that displays support information for surgical scenes where it is appropriate to display support information.

[0085] As mentioned above, the annotation of the first training data (first annotation) is performed manually by the user to create the first training data, but it is also possible to make the first training data automatically created. For example, as shown in Figure 18, the first training data is created by the first training data creation device 200, and the learning device 100 learns based on the created first training data and the first training model 121 (not shown in Figure 18) to generate the first trained model 21. The information processing system 5 then performs inference using the generated first trained model 21.

[0086] In Figure 18, the first training data creation device 200 includes a processor 210 and a memory 220. The processor 210 can be implemented with hardware similar to that of the processor 10 described in Figure 1, and the memory 220 can be implemented with semiconductor memory or the like, similar to the memory 20 described in Figure 1. The third trained model 223 is stored in the memory 220.

[0087] The third trained model 223 is trained using the third training data. The third training data is similar to the first training data in that it uses endoscopic images as input data and annotates surgical scenes onto those endoscopic images, but the details will be described later. The processor 210 then functions as the first training data creation unit 211 by performing the processing shown in the flowchart in Figure 19, for example.

[0088] In Figure 19, the processor 210 inputs the endoscopic image to the third pre-trained model 223 (step S610). More specifically, for example, the endoscopic image is input to the processor 210 via an input unit (not shown), the processor 210 reads the third pre-trained model 223, and performs inference based on the third pre-trained model 223. The processor 210 then determines whether the confidence level of the output data is above the third threshold (step S620). If the processor 210 determines that the confidence level of the output data is above the third threshold (YES in step S620), it includes the dataset of the input endoscopic image and output data in the first training data (step S630) and terminates the flow. On the other hand, if the processor 210 determines that the confidence level of the output data is not above the third threshold (NO in step S620), it terminates the flow. In other words, output data with a confidence level below the third threshold is not included in the first training data.

[0089] In this embodiment, the information processing system 5 is trained to output surgical scene information based on first training data with first annotations when a surgical scene with a confidence level of 3 or higher or exceeding the 3rd threshold is output as a result of inputting training images. In this way, the first training data can be acquired automatically.

[0090] Alternatively, the third training data may be created by acquiring endoscopic images from the endoscopic video every two hours, and having the user annotate the acquired endoscopic images with surgical scenes. The second hour is longer than the first hour mentioned above. In other words, the third trained model 223 is trained with a smaller dataset than the first trained model 21.

[0091] In this embodiment, the information processing system 5 is based on endoscopic images extracted every first hour from endoscopic video, which is a collection of endoscopic images. The first training data is created as a dataset of input endoscopic images and output surgical scenes when endoscopic images are input to the third trained model 223 and the confidence level of the surgical scene output by the third trained model 223 is equal to or greater than the third threshold. The third trained model 223 is created every second hour, which is longer than the first hour, based on the third training data, which has the first annotations attached to the endoscopic images extracted from the endoscopic video. In this way, the first training data can be automatically acquired while reducing the user's burden compared to when the user performs the annotation related to the first training data.

[0092] Furthermore, although not shown in the diagrams, the first training data may be created by using a predetermined image recognition model instead of the third trained model 223. The predetermined image recognition model recognizes a predetermined object in the endoscopic image which is the input data, and outputs a surgical scene associated with the recognized predetermined object as output data. The predetermined image recognition model does not necessarily have to be a deep learning model. For example, in the case of laparoscopic endo

[0093] Furthermore, the information processing system 5 may also display predetermined information relating to the surgical scene in addition to the aforementioned support information. This predetermined information relating to the surgical scene may be, for example, information about the number of surgical scene information stored in the recording unit 30, information about the surgical scene selected in step S300, or information about the transition process of the surgical scene. The transition process of the surgical scene refers to, for example, the process of transitioning from the first surgical scene to the second surgical scene. In other words, in the information processing system 5 of this embodiment, the processor 10 displays at least one of the number of surgical scene information stored in the recording unit 30, the selected surgical scene information, and the transition process of the surgical scene on the display 7. In this way, the user can grasp the information that can be derived from the surgical scene information stored in the recording unit 30.

[0094] Furthermore, for example, the processor 10 may display a graph on the display 7 showing the number of each surgical scene indicated by the surgical scene information stored in the recording unit 30. In this way, the user can visually grasp the number or proportion of surgical scene information stored in the recording unit 30. Specifically, for example, an example display like the one shown in Figure 20(A) is displayed on the display 7. In Figure 20(A), the number of surgical scenes stored in the recording unit 30 is displayed as a histogram as shown in E1. In other words, the processor 10 displays a histogram on the display 7 showing the number of each surgical scene indicated by the surgical scene information stored in the recording unit 30. In this way, the user can visually grasp the variability of the surgical scene information stored in the recording unit 30.

[0095] Alternatively, an example display as shown in Figure 20(B) may be displayed on the display 7. In Figure 20(A), the number of surgical scenes stored in the recording unit 30 is displayed as a bar graph as shown in E2. In other words, the processor 10 displays the number of each surgical scene indicated by the surgical scene information stored in the recording unit 30 as a bar graph on the display 7. In this way, the user can visually grasp the proportion of each surgical scene information stored in the recording unit 30. The processor 10 may also display the number of each surgical scene indicated by the surgical scene information stored in the recording unit 30 as a pie chart on the display 7.

[0096] Furthermore, for example, the processor 10 may display the surgical scene itself, as shown in E20, instead of the graph display shown in E10 in Figure 21(A). The graph display shown in E10 is a graph showing the breakdown of surgical scenes stored in the recording unit 30, similar to the graph display shown in E1 in Figure 20(A). The graph display shown in E10 indicates that the detection rate of surgical scene information related to Phase 2 exceeds the first threshold, and therefore E20 indicates that a surgical scene related to Phase 2 has been selected.

[0097] Furthermore, if, for example, multiple types of surgical scene information are stored in the recording unit 30, the processor 10 may display information indicating that the surgical scene is transitioning. For example, suppose that at the first timing, the detection rate of first surgical scene information corresponding to Phase 1, which is the first surgical scene, was stored in the recording unit 30 to the extent that it exceeded the first threshold, and therefore the first surgical scene was selected at step S300 in Figure 5. Then, at the subsequent second timing, suppose that second surgical scene information corresponding to Phase 2, which is the second surgical scene, continues to be stored in the recording unit 30, but the detection rate of second surgical scene information does not exceed the first threshold. E30 in Figure 21(B) shows the number of detected surgical scene information at the second timing as a histogram graph. In this case, the processor 10 may display, for example, the information shown at E40 in Figure 21(B). The display at E40 includes the surgical scene information shown at E41, the object shown at E42, and the surgical scene information shown at E43. More specifically, to indicate a situation where the surgical scene is transitioning from the first surgical scene to the second surgical scene, E41 represents the first surgical scene, E43 represents the second surgical scene, and E42 represents an object consisting of a one-way arrow. The object shown in E42 is not limited to a one-way arrow. For example, if the information for the first surgical scene and the information for the second surgical scene are detected in approximately the same proportion, the object shown in E42 may be a bidirectional arrow. In this way, in the information processing system 5 of this embodiment, when the first surgical scene information based on the first surgical scene and the second surgical scene information based on the second surgical scene are stored in the recording unit 30, the processor 10 displays information on the display 7 indicating that the surgical scene is transitioning from the first surgical scene to the second surgical scene, using the first surgical scene information, the second surgical scene information, and a predetermined object. In this way, the user can visually grasp that the surgical scene is transitioning from the first surgical scene to the second surgical scene.

[0098] Although this embodiment has been described in detail above, it will be readily apparent to those skilled in the art that many modifications are possible without substantially departing from the novelty and effects of this disclosure. Therefore, all such modifications are included within the scope of this disclosure. For example, any term that appears at least once in the specification or drawings alongside a broader or synonymous term may be replaced with that different term anywhere in the specification or drawings. Furthermore, all combinations of this embodiment and its modifications are also included within the scope of this disclosure. Also included are information processing systems, endoscope systems, methods for operating information processing systems, and programs. The configuration and operation of the device are not limited to those described in this embodiment, and various modifications are possible. [Explanation of symbols]

[0099] 1…Endoscope system, 3…Endoscope, 5…Information processing system, 7…Display, 10…Processor, 12…Input unit, 14…Output unit, 20…Memory, 21…First trained model, 22…Second trained model, 30…Recording unit, 100…Learning device, 121…First training model, 122…Second training model, 200…First training data creation device, 210…Processor, 211…First training data creation unit, 220…Memory, 223…Third trained model, t1, t2, t3, t4…Timing

Claims

1. Includes a processor that performs display processing for the display, The aforementioned processor, Based on a first trained model that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among multiple surgical scenes, based on first training data in which first annotations, which are annotations of surgical scenes, are attached to training images, the surgical scene information is detected on the endoscopic image. The detected surgical scene information is stored in the recording unit, and when the detection rate of surgical scene information representing the same surgical scene is equal to or greater than a first threshold, or when it exceeds the first threshold, the surgical scene is selected. A second trained model, which is trained to output the support information based on second training data to which support information annotations are attached to the training images, is determined according to the selected surgical scene. Based on the determined second trained model, the support information is detected for the endoscopic image. An information processing system characterized by superimposing the detected support information onto the endoscopic image and displaying it on the display.

2. In the information processing system of claim 1, The aforementioned processor, An information processing system characterized by determining a second trained model, which is trained to output the support information based on the second training data to which the second annotation, which is positional and morphological information of the target region in the selected surgical scene, is attached to the training image, according to the selected surgical scene.

3. In the information processing system of claim 1, The aforementioned processor, An information processing system characterized by determining whether or not to display the support information detected by the second trained model, depending on which of the multiple surgical scenes the selected surgical scene is.

4. In the information processing system of claim 1, The aforementioned processor, If the detection rate of the first surgical scene information indicating the first surgical scene is equal to or greater than the first threshold, or if it exceeds the first threshold, the first surgical scene is selected. An information processing system characterized in that the second surgical scene is selected when the detection rate of the second surgical scene information indicating the second surgical scene is equal to or greater than a second first threshold, or when it exceeds the second first threshold.

5. In the information processing system of claim 1, The aforementioned processor, At predetermined intervals, the surgical scene information based on the first trained model is detected. Using the predetermined number of surgical scene information obtained at the most recent predetermined number of timings, including the current timing, the detection ratio of each surgical scene is determined. An information processing system characterized in that, if the mode of the detection ratio of each of the obtained surgical scenes is equal to or greater than the first threshold, or exceeds the first threshold, the surgical scene corresponding to the mode is selected.

6. In the information processing system of claim 5, The aforementioned processor, An information processing system characterized in that, if each of the detection ratios of the obtained surgical scenes is less than or equal to the first threshold, the selection result of the surgical scene is not updated.

7. In the information processing system of claim 1, The aforementioned processor, An information processing system characterized in that, when the confidence level of the surgical scene detected based on the first trained model is equal to or greater than a second threshold, or when it exceeds the second threshold, the surgical scene information is stored in the recording unit.

8. In the information processing system of claim 7, The aforementioned processor, An information processing system characterized in that, if the confidence level of the surgical scene information detected based on the first trained model is less than or equal to the second threshold, information indicating invalid detection of the surgical scene information is stored in the recording unit, or the same surgical scene information as in the previous frame is stored in the recording unit as the surgical scene information for the current frame.

9. In the information processing system of claim 1, The aforementioned processor, If the confidence level of the first surgical scene detected based on the first trained model is equal to or greater than the first second threshold, or if it exceeds the first second threshold, the first surgical scene information indicating the first surgical scene is stored in the recording unit. An information processing system characterized in that, when the confidence level of the second surgical scene detected based on the first trained model is equal to or greater than a second second threshold, or when it exceeds the second second threshold, second surgical scene information indicating the second surgical scene is stored in the recording unit.

10. In the information processing system of claim 1, The first pre-trained model described above is: An information processing system characterized in that, when a surgical scene having a confidence level of a third threshold or higher, or a confidence level exceeding the third threshold, is output as a result of inputting the aforementioned training images, the system is trained to output the surgical scene information based on the first training data to which the first annotation has been attached.

11. In the information processing system according to claim 10, The aforementioned training images are This image is based on the endoscopic images extracted every hour from the endoscopic video, which is a collection of endoscopic images. The first training data is: When the endoscopic image is input to the third trained model, if the confidence level of the surgical scene output from the third trained model is equal to or greater than the third threshold, or exceeds the third threshold, a dataset is created consisting of the input endoscopic image and the output surgical scene. The third pre-trained model described above is: An information processing system characterized in that, at intervals of two hours longer than the first hour, the endoscopic images extracted from the endoscopic video are created based on third training data to which the first annotation has been applied.

12. In the information processing system according to claim 10, The first training data is: An information processing system characterized in that, when an image recognition model detects a predetermined object representing the characteristics of the surgical scene in the endoscopic image, the system creates a dataset with the input endoscopic image and the surgical scene relating to the detected predetermined object as output data.

13. In the information processing system of claim 1, The aforementioned processor, An information processing system characterized by displaying on the display the number of surgical scene information stored in the recording unit, the selected surgical scene information, and at least one of the transition processes of the surgical scene.

14. In the information processing system described in claim 13, The aforementioned processor, An information processing system characterized by displaying a graph on the display showing the number of each surgical scene indicated by the surgical scene information stored in the recording unit.

15. In the information processing system described in claim 14, The aforementioned processor, An information processing system characterized by displaying a histogram of the number of each surgical scene indicated by the surgical scene information stored in the recording unit on the display.

16. In the information processing system described in claim 14, The aforementioned processor, An information processing system characterized by displaying the number of each surgical scene indicated by the surgical scene information stored in the recording unit as a bar graph on the display.

17. In the information processing system of claim 13, The aforementioned processor, An information processing system characterized in that, when first surgical scene information based on a first surgical scene and second surgical scene information based on a second surgical scene are stored in the recording unit, the system displays on the display information indicating that the surgical scene has transitioned from the first surgical scene to the second surgical scene, using the first surgical scene information, the second surgical scene information, and a predetermined object.

18. The information processing system of claim 1, Endoscope and, An endoscopic system characterized by including the following.

19. A method for operating an information processing system that performs display processing on a display, The steps include detecting surgical scene information in an endoscopic image based on a first trained model that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among a plurality of surgical scenes, based on first training data in which first annotations, which are annotations of surgical scenes, are attached to the training image, and The detected surgical scene information is stored in the recording unit, and when the detection rate of the same surgical scene information is equal to or greater than a first threshold, or when it exceeds the first threshold, the surgical scene is selected. The steps include determining a second trained model, which is trained to output the support information based on second training data to which support information annotations are attached to the training images, according to the selected surgical scene, The steps include detecting the support information for the endoscopic image based on the determined second trained model, The steps include: superimposing the detected support information onto the endoscopic image and displaying it on the display; A method for operating an information processing system, characterized by including the following:

20. The steps include detecting surgical scene information in an endoscopic image based on a first trained model that has been trained to output surgical scene information indicating the surgical scene to which the training image belongs among a plurality of surgical scenes, based on first training data in which first annotations, which are annotations of surgical scenes, are attached to the training image, and The detected surgical scene information is stored in the recording unit, and when the detection rate of the same surgical scene information is equal to or greater than a first threshold, or when it exceeds the first threshold, the surgical scene is selected. The steps include determining a second trained model, which is trained to output the support information based on second training data to which support information annotations are attached to the training images, according to the selected surgical scene, The steps include detecting the support information for the endoscopic image based on the determined second trained model, The steps include: superimposing the detected support information onto the endoscopic image and displaying it on a display; A program characterized by causing a computer to execute something.

Citation Information

Patent Citations

  • Automated provision of real-time custom procedural surgical guidance

    US9788907B1