Surgery support device, surgery support method, and surgery support program

The surgical support device addresses complications in surgery by estimating and outputting surgical insights, improving learning and reducing errors through image recognition and thought process estimation.

JP7811416B1Active Publication Date: 2026-02-05DIREAVA CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2025064616
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2026-02-05
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Complications during surgery, particularly in esophageal cancer surgery, are common due to misunderstandings, lack of attention, and communication errors among surgical teams, making it difficult for students to learn from experienced surgeons and hindering effective shared decision-making.

Method used

A surgical support device that performs image recognition on surgical images, estimates situation judgment, decision-making, and thought processes using user queries and reference information, and outputs these insights to assist learners.

Benefits of technology

The device provides insight into the thoughts and decision-making processes of skilled physicians during surgery, enhancing learning and reducing communication errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811416000001_ABST
    Figure 0007811416000001_ABST
Patent Text Reader

Abstract

It is impossible to know what an experienced surgeon is thinking during surgery. [Solution] A surgical assistance device according to one embodiment of the present invention comprises an image acquisition unit that acquires surgical images, an information acquisition unit that acquires queries and reference information from a user, a thought estimation unit that performs image recognition on the surgical images and uses the queries and reference information from the user to estimate at least one of situational judgment, decision-making, and the thought process leading to the situational judgment or decision-making in relation to the surgical images, and an output unit that changes the output format depending on the results of the image recognition, the queries and reference information from the user, and at least one of the situational judgment, decision-making, and thought process, and outputs at least one of the situational judgment, decision-making, and thought process to the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a surgery assistance device, a surgery assistance method, and a surgery assistance program. [Background technology]

[0002] Complications during surgery remain common, particularly after esophageal cancer surgery, where there is a 2% mortality rate and a complication rate of over 50%. Complication rates vary depending on the surgeon's experience and the facility, with more experienced surgeons and larger facilities achieving better outcomes, while less experienced surgeons tend to experience more complications. Causes of complications include misunderstandings during surgery, lack of attention, and communication errors. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2020-194439 Summary of the Invention [Problem to be solved by the invention]

[0004] In surgical education, it is important to watch a surgery in real time and learn "what the surgeon is thinking and doing," "what do they see," "why does the surgeon do this," and "what will they do next?" In other words, it is important to learn how surgeons judge situations and make decisions during surgery. However, because surgeons are focused on the operation, it is difficult for students to speak up, and as a result, students are unable to learn effectively. Furthermore, even during surgery, there are many situations where shared decision-making within the same team is required, but communication errors can lead to complications.

[0005] Therefore, the present invention aims to understand the thoughts of skilled doctors during surgery. [Means for solving the problem]

[0006] A surgical support device according to one embodiment of the present invention comprises an image acquisition unit that acquires surgical images, an information acquisition unit that acquires queries and reference information from a user, a thought estimation unit that performs image recognition on the surgical images and uses the query and reference information from the user to estimate at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making in relation to the surgical images, and an output unit that changes the output form depending on the results of the image recognition, the query and reference information from the user, and at least one of the situation judgment, the decision-making, and the thought process, and outputs at least one of the situation judgment, the decision-making, and the thought process to the user. [Effects of the Invention]

[0007] The present invention allows insight into the thoughts of skilled physicians during surgery. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram for explaining an overview according to one embodiment of the present invention. [Figure 2] 1 is a diagram showing an overall configuration according to an embodiment of the present invention; [Figure 3] 1 is a functional block diagram of a surgery assistance device according to an embodiment of the present invention. [Figure 4] 1 is a flowchart of a process for estimating situation judgment, decision-making, and thought processes according to an embodiment of the present invention. [Figure 5] FIG. 1 is a diagram for explaining estimation of situational judgment, decision-making, and thought processes according to one embodiment of the present invention. [Figure 6] FIG. 1 is a diagram for explaining estimation of situational judgment, decision-making, and thought processes according to one embodiment of the present invention. [Figure 7] FIG. 1 is a diagram for explaining image recognition of a surgical image according to an embodiment of the present invention. [Figure 8] FIG. 1 is a diagram for explaining image recognition of a surgical image according to an embodiment of the present invention. [Figure 9]FIG. 10 is a diagram illustrating the creation of a prompt according to an embodiment of the present invention. [Figure 10] FIG. 1 is a diagram illustrating classification of queries and answers according to an embodiment of the present invention. [Figure 11] FIG. 1 is a diagram illustrating classification of queries and answers according to an embodiment of the present invention. [Figure 12] 1 is an example of an interactive display according to one embodiment of the present invention. [Figure 13] 1 is an example of an interactive display according to one embodiment of the present invention. [Figure 14] 1 is a hardware configuration diagram of a surgery assistance device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0010] <Terminology> In this specification, a "surgical image" refers to an image taken of a patient undergoing surgery, and may be a video or a still image. In this specification, "surgery" may refer to any type of surgery, such as surgery (digestive, vascular, breast, thyroid), cardiovascular surgery, respiratory surgery, neurosurgery, orthopedic surgery, plastic surgery, urology, obstetrics and gynecology, ophthalmology, otolaryngology, oral surgery, pediatric surgery, transplant surgery, psychiatry, trauma surgery, etc., including, for example, the following (1) and (2): (1) Procedural operations according to the surgical procedure (grasping, incising, and dissecting organs and other objects in surgical operations, determining the extent of resection for malignant tumors, ensuring safe margins, dissection of nearby lymph nodes, resection of digestive organs (gastrectomy, intestinal resection, liver resection, gallbladder resection, etc.), reconstructive procedures (anastomosis, bypass surgery), use of staplers (automatic suturing instruments), clipping of aneurysms and artificial vascular replacement in vascular surgery, reduction and fixation of fractures (plates, screws, intramedullary nails) in orthopedic surgery, replacement of artificial joints, removal of brain tumors in neurosurgery, clipping of cerebral aneurysms, submucosal dissection in endoscopic treatment, etc.) (2) Suturing and ligation operations (clipping, ligation, hemostasis using bipolar and monopolar instruments, use of hemostatic agents (gelatin sponge, fibrin glue), suturing tissues such as the digestive tract and blood vessels, use of surgical staplers, insertion of drains) In this specification, an "expert" is a medical professional who meets certain conditions (for example, a doctor such as a surgeon, a medical intern, a medical student, a nurse, a nursing student, a clinical engineer, etc.). For example, an expert is someone who has been working in the medical field for a threshold amount of time or has performed a threshold amount of surgeries.

[0011] <Summary> FIG. 1 is a diagram illustrating an overview of one embodiment of the present invention. In this invention, a surgery support device 10 (described in detail later) performs image recognition on a surgical image and estimates at least one of a situation judgment, decision-making, and thought process (specifically, a thought process leading to a situation judgment and a thought process leading to a decision) of an expert (e.g., an experienced doctor) regarding the surgical image using a query and reference information from a user. The surgery support device 10 changes the output format depending on the result of the image recognition, the query and reference information from the user, and at least one of the situation judgment, decision-making, and thought process, and outputs at least one of the situation judgment, decision-making, and thought process to the user.

[0012] For example, the surgery support device 10 performs image recognition on a surgical image to recognize the type and position of the part (e.g., anatomical structures such as organs and nerves), the type and position of the instrument, the presence or absence of bleeding, and the surgical procedure. The surgery support device 10 inputs the image recognition results (e.g., the part recognition result is the right recurrent laryngeal nerve, the instrument recognition result is bipolar forceps, the bleeding recognition result is no bleeding, and the surgical procedure recognition result is the right recurrent laryngeal nerve lymph node), a query from the user (e.g., "Please tell me the situation in this scene and the future plan"), and reference information (e.g., new surgeon, no surgical experience, surgical procedure (esophagus)) into a generation AI (artificial intelligence) that generates text, and outputs the situation assessment, decision-making, and thought process of an experienced surgeon (e.g., "The lymph node located outside the right recurrent laryngeal nerve is being dissected with a bipolar forceps. Please perform the dissection carefully because there is a risk of paralysis due to recurrent laryngeal nerve damage."), and presents it to the user.

[0013] <System configuration> 2 is a diagram showing the overall configuration according to one embodiment of the present invention. The surgery support system 1 can include a surgery support device 10, an input device 11, an output device 12, and an imaging device 20. Each of these will be described below.

[0014] <<Surgical support equipment>> The surgery support device 10 is a device that estimates at least one of the situation judgment, decision-making, and thought process (specifically, the thought process leading to the situation judgment and the thought process leading to the decision-making) of an expert (e.g., an experienced doctor) in relation to surgical images. When the surgery support device 10 is operating during surgery, the surgical images being captured by the imaging device 20 are used, and when the surgery support device 10 is operating outside of surgery, recorded surgical images captured by the imaging device 20 in the past are used. The surgery support device 10 is made up of one or more computers.

[0015] The surgery support device 10 may be a computer (local server) installed in the operating room, or a computer installed outside the operating room (for example, a cloud server; a terminal installed in the operating room accesses the cloud server).

[0016] <<Input Device>> The input device 11 is a microphone, keyboard, mouse, etc., used to give various instructions to the surgery assistance device 10 .

[0017] <<Output device>> The output device 12 is a monitor, display, etc. that displays images and text acquired from the surgery assistance device 10, a speaker, etc. that plays back audio acquired from the surgery assistance device 10, etc.

[0018] <<Imaging device>> The imaging device 20 is a device (for example, an endoscopic camera) that captures images of a patient undergoing surgery. Images captured by the imaging device 20 are called surgical images (which may be moving images or still images).

[0019] At least two of the surgery assistance device 10, the input device 11, the output device 12, and the imaging device 20 may be implemented in one device.

[0020] <Function block> 3 is a functional block diagram of a surgery support device 10 according to one embodiment of the present invention. The surgery support device 10 can include an image acquisition unit 101, an information acquisition unit 102, a knowledge database processing unit 103, a knowledge database storage unit 104, a thought estimation unit 105, an output unit 106, an interaction processing unit 107, and a data storage unit 108. By executing a program, the surgery support device 10 can function as the image acquisition unit 101, the information acquisition unit 102, the knowledge database processing unit 103, the thought estimation unit 105, the output unit 106, and the interaction processing unit 107. Each of these units will be described below.

[0021] The image acquisition unit 101 acquires surgical images.

[0022] [[Surgery Images]] Here, the surgical images will be described. The image acquisition unit 101 may acquire in real time the surgical images being captured by the imaging device 20, or may acquire surgical images captured in the past and stored in the data storage unit 108 or the like.

[0023] The information acquisition unit 102 acquires a query (e.g., a question) and reference information from a user (i.e., a person who wants to know the situation judgment, decision-making, and thought process of an expert). For example, the information acquisition unit 102 acquires a query and reference information input by the user via the input device 11 or the like.

[0024] The knowledge database processing unit 103 acquires knowledge data from a knowledge database (for example, medical books, reports by experienced doctors, etc.) stored in the knowledge database storage unit 104.

[0025] The knowledge database storage unit 104 stores a knowledge database.

[0026] The thought estimation unit 105 performs image recognition on the surgical image acquired by the image acquisition unit 101, and estimates at least one of the expert's "situation judgment," "decision-making," and "thought process leading to the situation judgment or decision-making" regarding the surgical image, using the query and reference information from the user acquired by the information acquisition unit 102. Specifically, the thought estimation unit 105 makes the estimation using the visual language described in FIGS.

[0027] The output unit 106 changes the output format depending on the result of image recognition by the thought estimation unit 105, the query and reference information from the user acquired by the information acquisition unit 102, and at least one of "situation judgment", "decision making", and "thought process" estimated by the thought estimation unit 105, and outputs at least one of "situation judgment", "decision making", and "thought process" estimated by the thought estimation unit 105 to the user. Note that the output unit 106 may output at least one of "situation judgment", "decision making", and "thought process" estimated by the thought estimation unit 105 together with the result of image recognition by the thought estimation unit 105.

[0028] The interaction processing unit 107 changes the output format of the information output by the output unit 106 in response to a user operation (an instruction input by the user via the input device 11 or the like).

[0029] The data storage unit 108 stores various types of data.

[0030] Furthermore, the surgical support device 10 may further include a learning unit that learns a language model (also called a language generation model, hereinafter referred to as LM), an auxiliary information addition determination unit that determines whether to add auxiliary information to the output, and an emergency state determination unit that determines whether the surgical image is in an emergency state.

[0031] <Processing method> FIG. 4 is a flowchart of a process for estimating situation judgment, decision making, and thought processes according to one embodiment of the present invention.

[0032] In step 1 (S1), the information acquisition unit 102 of the surgery support device 10 acquires a query and reference information from a user. In this way, by acquiring the reference information, the output form of the natural language (i.e., the situation judgment, decision-making, and thought process of an expert such as an experienced doctor) output in S8 can be changed according to the acquired reference information.

[0033] In step 2 (S2), the knowledge database processing unit 103 of the surgery support device 10 acquires knowledge data from the knowledge database stored in the knowledge database storage unit 104. In this way, by referring to knowledge data (for example, medical books, reports by experienced doctors, etc.), it is possible to estimate the situation judgment, decision-making, and thought process of an expert such as an experienced doctor.

[0034] In step 3 (S3), the image acquisition unit 101 of the surgery assistance device 10 acquires a surgery image.

[0035] In step 4 (S4), the thought estimation unit 105 of the surgery assistance device 10 generates a feature amount of the query and a feature amount of the reference information by using the query and the reference information from the user acquired in S1.

[0036] In step 5 (S5), the thought estimation unit 105 of the surgery assistance device 10 performs image recognition on the surgery image acquired in S3.

[0037] In step 6 (S6), the thought estimation unit 105 of the surgery assistance device 10 generates image features using the surgical image acquired in S3 or the result of the image recognition in S5.

[0038] The thought estimation unit 105 may generate image features using a machine learning model that has undergone expression learning (in which arbitrary information is associated with an image).The thought estimation unit 105 may also generate image features using the results of image recognition specialized for a specific task (for example, recognition of a body part (such as an organ), recognition of an instrument, recognition of bleeding, recognition of a process, etc.).

[0039] In step 7 (S7), the thought estimation unit 105 of the surgery support device 10 estimates at least one of the situation judgment, decision-making, and thought process of the expert regarding the surgical image. The thought estimation unit 105 may use information on the thought process in image recognition (specifically, annotation data by a doctor or the like is used to generate an image recognition model), information on the thought process output by the generation AI, information on the thought process acquired by the generation AI repeatedly generating answers (specifically, image information obtained from the result of image recognition, language information previously output by the LM (language model), and information in the knowledge database are used as input to the LM (language model)), or information on the thought process linked by knowledge data in the knowledge database (specifically, knowledge data stored in the knowledge database is acquired by RAG and the knowledge data is input to the LM (language model)).

[0040] In step 8 (S8), the output unit 106 of the surgery assistance device 10 outputs at least one of the situation judgment, decision-making, and thought process estimated in S7 by changing the output format depending on the result of image recognition, the query and reference information from the user, and at least one of the situation judgment, decision-making, and thought process. The output unit 106 can also output the result of image recognition.

[0041] In step 9 (S9), the interaction processing unit 107 of the surgery assistance device 10 changes the output format of the information output by the output unit 106 in response to a user operation (an instruction input by the user via the input device 11 or the like).

[0042] In step 10 (S10), the query and reference information input by the user, the image recognition results, the situation judgment, the decision-making, and the thought process information are recorded in the data storage unit 108 of the surgery support device 10.

[0043] The processing of each step in FIG. 4 will be described in detail below.

[0044] [S1 (Query / reference information acquisition)] The information acquisition unit 102 of the surgery support device 10 acquires a query and reference information from a user. For example, the voice of the query is acquired via a headset with a microphone worn by the user (surgeon). For example, the text of the query is acquired via a keyboard. If the query is voice, any voice recognition (specifically, voice recognition processing using the frequency of the voice waveform as a feature, for example, GMM-HMM or deep learning) is performed to convert the voice into text. Voiceprint authentication may also be performed to acquire role information (which may be a surgeon, assistant, nurse, or individual), and the role information may be added to the text output by voice recognition.

[0045] [[Query]] Here, we will explain the query. A query is a prompt (for example, a question) to the generation AI. Note that a query may include specific fixed words or sentences.

[0046] [[Reference information]] Here, we will explain the reference information. The reference information includes at least one of information about the user (i.e., a person who wants to know the situation judgment, decision-making, and thought process of an expert), information about the surgery, and information about the patient. For example, the reference information includes the user's medical history, surgical history, age, and other profile information, pre-surgery information such as the surgical procedure, and patient chart information.

[0047] [S2 (Acquisition of knowledge data)] The knowledge database processing unit 103 of the surgery assistance device 10 acquires knowledge data from the knowledge database stored in the knowledge database storage unit 104. The knowledge database processing unit 103 can acquire knowledge data from the knowledge database stored in the knowledge database storage unit 104 using a Retrieval-Augmented Generation (RAG) method (vector search, graph search).

[0048] [[Knowledge Database]] The knowledge database will now be described. The knowledge database stores data on medical textbooks, surgery reports, and annotations by doctors. Specifically, the data on medical textbooks, surgery reports, and annotations by doctors is stored in a hierarchical structure for each surgical procedure. Therefore, only knowledge data corresponding to the reference information can be acquired (for example, if the pre-operative information is the surgical procedure for the esophagus, only knowledge data related to the esophageal surgery can be acquired). Note that the knowledge data is not limited to text, and may be an image.

[0049] The above-mentioned reference information may be stored in the knowledge database storage unit 104 , and the information acquisition unit 102 may acquire the reference information stored in the knowledge database storage unit 104 .

[0050] [S3 (Acquisition of surgical images)] The image acquisition unit 101 of the surgery assistance device 10 acquires surgery images.

[0051] [S4 (Generation of query and reference information features)] The thought estimation unit 105 of the surgery assistance device 10 uses the query and reference information from the user to generate feature quantities of the query and feature quantities of the reference information.

[0052] The processing of S4 will be explained with reference to the generation of query and reference information features in Figure 5. Using a Tokenizer, the text of the query and reference information is broken down into tokens, and the tokens are input to the Embedding Layer of the Language Model, which outputs Embedding (features of the query and features of the reference information (embedded representation)). Also, using a Tokenizer, the text of knowledge data in the knowledge database is broken down into tokens, and the tokens are input to the Embedding Layer of the Language Model, which outputs Embedding (features of the knowledge data (embedded representation)).

[0053] The thought estimation unit 105 may generate the embedding after explicitly adding a flag indicating that it is specific reference information or a parameter that controls the generated text.

[0054] In the case of specific reference information, the thought estimation unit 105 may add a prompt to change the text (i.e., sentences about situation judgment, decision-making, and thought processes). For example, surgical experience may be defined as information to be input in advance as reference information, and if the surgical experience is 0-3 years, a brief and easy-to-understand explanation may be added to the prompt. Note that a user interface for inputting surgical experience may also be provided.

[0055] In the case of specific reference information, the thought estimation unit 105 may assign a special token. For example, surgical experience may be defined as reference information to be input in advance, and if the surgical experience is 0-3 years, the Tokenizer assigns a special token (for example, a beginner token) to the token of the reference information. When a beginner token is assigned, the text generated by the LM (language model) (i.e., sentences on situational assessment, decision-making, and thought processes) is designed to be concise and easy to understand (i.e., learned during learning of the LM (language model)). Note that a user interface for inputting surgical experience may be provided.

[0056] In the case of specific reference information, the thought estimation unit 105 may change the output control parameters (parameters that control the generated text). For example, surgical experience may be defined as reference information to be input in advance, and if the surgical experience is 0-3 years, the output control parameters of the LM (language model) may be changed to change the length or variety of the generated text (i.e., sentences about situational judgment, decision-making, and thought processes). A user interface for inputting surgical experience may also be provided. For example, changing the temperature of the LM (language model) output or the distribution of attention during processing. For example, changing the number of tokens in the output of the LM (Language Model). For example, changing the parameters of the Top K samples of the output tokens of the LM (Language Model).

[0057] [S5 (Image Recognition)] The thought estimation unit 105 of the surgery support device 10 performs image recognition on the surgical image. For example, the thought estimation unit 105 uses a semantic segmentation technique to recognize the pixel-level positions (x-coordinates and y-coordinates) on the surgical image where a specific part (such as an organ), instrument, or bleeding is detected. In addition, for example, the thought estimation unit 105 uses a classification technique to recognize the surgical steps from the surgical image.

[0058] Figure 7 is a diagram for explaining image recognition of a surgical image according to one embodiment of the present invention. As shown in [Recognition of parts (organs, etc.), instruments, and bleeding] in Figure 7, the position and likelihood on the surgical image where a part (organ, etc.) is detected, the position and likelihood on the surgical image where an instrument is detected, and the position and likelihood on the surgical image where bleeding is detected are recognized. As shown in [Recognition of surgical steps] in Figure 7, the surgical steps are recognized.

[0059] The thought estimating unit 105 can also perform image recognition and recognize the incision line, the margin for the incision line, and the incision direction as shown in Fig. 8. As shown in Fig. 8, the thought estimating unit 105 can recognize the incision line (the location to be incised), the margin for the incision line, the incision direction (the direction to be incised), and the pulling direction from the surgical image (specifically, perform regression estimation in the x and y directions for the incision line).

[0060] [S6 (Image feature generation)] The thought estimation unit 105 uses the surgical image or the result of image recognition to generate image features (that is, features that serve as the basis for at least one of situation judgment, decision-making, and thought process).

[0061] The processing of S6 will be described with reference to the generation of features of the surgical image in Fig. 5. Note that the visual language model used in the present invention is not limited to Fig. 5, and may be Example 1 or Example 2 in Fig. 6, as long as it is a module that can handle visual information and language information.

[0062] [[Generating Embeddings from Surgical Images]] The surgical image is input to a representation-trained Vision Encoder (which associates arbitrary information with the image), which outputs the features of the surgical image. Next, the features of the surgical image are input to a Projector, which outputs Embedding (features of the surgical image (embedded representation)) that will be input to a Language Model (LM).

[0063] [[When generating embeddings from image recognition results]] First, as described above, image recognition of surgical images is performed using the Image Recognition Model. Image recognition specialized for a specific task (e.g., recognition of body parts (organs, etc.), recognition of instruments, recognition of bleeding, recognition of procedures, etc.) is performed. Next, prompts are created based on the results of image recognition, and the prompts are broken down into tokens using a Tokenizer. Next, the tokens are input into the Embedding Layer of the Language Model (LM), and Embedding (features (embedded representations) of the surgical image) is output.

[0064] For example, if an image recognition model that detects nerves recognizes a nerve as a result of image recognition, a prompt saying "A nerve is present on the right side" is created, and a token is created using the Tokenizer. For example, if an image recognition model that recognizes surgical procedures recognizes a dissection procedure as a result of image recognition, a prompt saying "The dissection procedure is being performed" is created, and a token is created using the Tokenizer. The prompt is created by applying the information from the image recognition results to a predetermined sentence (template sentence) such as "B is present on side A."

[0065] Note that when multiple objects are detected or multiple image recognition models are used, multiple prompts may be input to the Tokenizer to create tokens. Also, in Example 1 of Figures 5 and 6, instead of Prompt + Tokenizer, a Projector may be used to create an Embedding to be input to the Language Model (LM) (in this case, intermediate features of the output backbone of the image recognition model are used). Also, the output results of the image recognition model may be used as input to the Vision Encoder.

[0066] [S7 (Thought estimation processing)] The thought estimation unit 105 estimates the situation judgment, decision-making, and thought process of the expert regarding the surgical image. Specifically, the thought estimation unit 105 inputs embeddings (embedded expressions) into a language model (LM) to generate text (i.e., sentences of situation judgment, decision-making, and thought process) (for example, a decoder-only transformer is used).

[0067] The process of S7 will be described with reference to the thought estimation in FIG.

[0068] [[Annotation]] In the present invention, when training a language model (LM), the situational judgment, decision-making, and thought process of a doctor or the like are used as annotation data for the training data so that the language model (LM) outputs the situational judgment, decision-making, and thought process of a doctor or the like. Specifically, the annotation data used when training the language model (LM) is categorized, and learning is performed by providing text of Object (visible object), Action (behavior of visible object, etc.), Situation (current situation), Prediction (next predicted situation), Plan (next action to be taken), and Reason (reason that led to Situation, Prediction, and Plan). By creating annotation text so that the same words appear between Object, Action, Situation, Prediction, Plan, and Reason, each category is linked and learning becomes easier. For example, an image recognition model or vision encoder capable of generating embeddings that contribute to annotation data (i.e., categorized annotation data) such as those shown below is used. Image recognition models that recognize body parts (organs, etc.) and instruments generate features for outputting objects. The image recognition model that recognizes the process generates features to output the action and situation. The image recognition model that recognizes bleeding generates features to output the situation and prediction. The Vision Encoder generates features for all of these categories.

[0069] An example of annotation is shown below. Object: Right recurrent nerve, bipolar Action: Bipolar delamination Situation: The area around the right recurrent laryngeal nerve is dissected for right recurrent laryngeal nerve dissection. Prediction: No bleeding and no problems with the peeling process Plan: Right recurrent nerve dissection Reason: There is no bleeding and the surgery is proceeding smoothly, so we will continue with the right recurrent nerve dissection. However, please be careful as there is a risk of paralysis due to recurrent nerve damage.

[0070] In the above annotation example, the surgical procedure is for the esophagus. In this case, if the following method can be used for image recognition specialized for a specific task, the language model (LM) will be able to output answers to Object, Action, and Situation with high accuracy, and will be able to output answers to Prediction, Plan, and Reason after comprehensively considering these conditions. The Vision Encoder generates arbitrary image information that includes all of these and inputs it into the language model (LM). Image recognition techniques: area recognition (recurrent laryngeal nerve), instrument recognition (bipolar), action recognition (dissection, instrument action), surgical procedure recognition (nerve dissection), bleeding recognition (presence or absence of bleeding)

[0071] For example, the image recognition model is an object detection model (which may be object detection that estimates a bounding box of a detection target from an image, or segmentation that estimates the position of a detection target from an image at the pixel level (the following will explain the case of segmentation)). In the case of segmentation, an image is input into the image recognition model (object detection model), and the position of the detection target is estimated from the image at the pixel level.

[0072] [[Create Prompt]] This section explains how to create a prompt. For example, if a nerve is detected as a result of image recognition, a prompt saying "A nerve is present at the bottom center of the screen" is created. The prompt can be expressed as a sentence from a template with various patterns, such as "A nerve is visible near the bottom center of the screen." The sentence of the prompt is created according to the coordinate position and likelihood of the object obtained as a result of image recognition.

[0073] As shown in Example 1 in Figure 9, a prompt can be created based on the approximate location of the detected object. First, the surgical image is divided into two vertically and three horizontally, creating six separate regions. Next, based on the results of image recognition, it is determined which of the six divided regions the segmentation class (for example, the right recurrent laryngeal nerve class) belongs to.Then, it is determined that the region (for example, the bottom center) that contains the most segmentation class (for example, the right recurrent laryngeal nerve class) is the region where the object (in this example, the right recurrent laryngeal nerve) was detected. Based on the result of determining which region it belongs to, a prompt (for example, "The right recurrent laryngeal nerve is shown near the bottom center of the screen") is created.

[0074] As shown in [Example 2] in Figure 9, a prompt can be created based on the relative positional relationship between the objects to be detected. First, randomly select a detected object (or select an important organ). For example, the right recurrent laryngeal nerve is selected. Next, the coordinates of the location of the selected object (in this example, the right recurrent laryngeal nerve) and the coordinates of the location of another detected object (e.g., bipolar forceps) are used to calculate the closest distance between the two objects and the direction from one object to the other. A threshold is set for the distance between two such objects, and if the distance between the two objects is less than the threshold, the objects are determined to be close to each other. When objects are close together, the system generates a prompt (e.g., "The bipolar forceps are visible to the right of the right recurrent laryngeal nerve") based on the relative position (i.e., distance and direction) between the two objects.

[0075] For example, a prompt can be created based on the likelihood of an object detected. The prompt can be created based on the maximum likelihood (or the minimum or average) of the detected object. For example, the likelihood (0-1) can be divided into high (1.0-0.66), medium (0.66-0.33), and low (0.33-0) and the prompt can be changed accordingly. High: Nerves are present Center: Nerve may be present Low: Nerve is likely not present

[0076] For example, a prompt can be created based on the likelihood of multiple surgical steps in surgical step recognition. By applying Softmax to predefined surgical step classes (for example, surgical step A (likelihood: 0.38), surgical step B (likelihood: 0.4), surgical step C (likelihood: 0.12), surgical step D (likelihood: 0.1)), the likelihood of each class is output so that the sum of the likelihoods of each class is 1.0. Normally, "step B," with the maximum likelihood, becomes the surgical step recognition result, but since the difference in likelihood between surgical step A and surgical step B is very small, a prompt with the sentence "either step A or step B is being performed" is created.

[0077] It is also possible to create a combination of the above prompts (e.g., "A bipolar forceps is visible near the right recurrent laryngeal nerve, located near the bottom center of the screen. There is no bleeding present.").

[0078] As mentioned above, in Example 1 of Figures 5 and 6, a Projector can be used instead of a Prompt+Tokenizer to create an Embedding that will be input to the Language Model (LM) (in this case, the intermediate features of the output of the backbone of the image recognition model are used (input to the Projector)). Because the image recognition results are interpreted using a Projector and Language Model (LM) that have been trained based on the image recognition result information and its caption, this method has greater expressive power than the method of applying a template prompt.

[0079] Below is a specific example of creating a pronto from the results of image recognition.

[0080] For example, if the tumor invades the blood vessels more than expected, more extensive resection and / or revascularization may be required. (1) Using image recognition results to visualize tumor infiltration into blood vessels (2) Detect the positions (coordinates) of tumors and blood vessels and calculate their proximity (3) Using the information from (1) and (2) above, create an embedding that will be input to the language model (LM).

[0081] For example, when faced with unexpected serious complications such as massive bleeding, tissue fragility, or damage to surrounding organs, the most appropriate repair method can be selected. (1) Using image recognition results, we visualize the occurrence of complications such as severe bleeding, tissue fragility, and damage to surrounding organs. (2) Calculate the amount of bleeding (area of ​​bleeding detection result (segmentation)), the amount and magnitude of tissue damage (area of ​​tissue damage range (segmentation) and estimation of damage amount (regression)), and recognition of damage to surrounding organs of the organ being treated (presence or absence of surrounding organs other than the organ being treated). (3) Using the information from (1) and (2) above, create an embedding that will be input to the language model (LM).

[0082] [Utilization of reference information and knowledge data] In one embodiment of the present invention, the content to be output is changed according to the judgment of the situation using at least one of reference information (for example, advance information on surgery, medical record information) and knowledge data.

[0083] [[Training the Model]] Here, the learning of the visual language model used in the present invention will be explained. The image recognition model, Vision Encoder, and Language Model (LM) in Fig. 5 are each trained independently. 1. Train an image recognition model for a specific task using supervised learning. 2. Train the Vision Encoder using contrastive image and text training, such as SigLIP or CLIP. 3. Train a language model (LM). Like Llama, pre-training is performed using a token prediction task, followed by additional training using a QA task, and reinforcement learning using RLHF (Reinforcement Learning from Human Feedback). 4. The Vision Encoder is frozen so that it does not learn, and the image is input to the Vision Encoder. The resulting image features are input to the Projector to obtain the embedding as the output. The text to be annotated to the image is input to the Tokenizer and the embedding layer of the language model (LM) to obtain the embedding of the true image caption. Loss is calculated between the embedding output by the Projector and the true embedding, and learning is performed. 5. Training the entire visual language model An image is input into the Image Recognition Model to obtain image recognition results for a specific task. The image recognition results are then matched to a predefined prompt, and the prompt is passed through the Tokenizer to create tokens. The image is input into the Vision Encoder to obtain image features. Embeddings are then created from the image features via the Projector. The query and reference text are input into the Tokenizer to create tokens. Embeddings are then created from the tokens of the query and reference text and the tokens created from the image recognition results via the Embedding Layer of the Language Model (LM). The embeddings of the query and reference text, the embedding of the image recognition results, and the embedding of the Vision Encoder are combined and used as input to the Language Model (LM). Only the Projector and Language Model (LM) are trained. In other words, the Image Recognition Model and Vision Encoder are trained in a frozen state.

[0084] Note that instead of using a generative AI (language model (LM)) that generates text as described above, an image generation model or a grounding model that solves image recognition tasks using arbitrary language and image information as input may be used. In other words, by using these models, the thought estimation unit 105 can output not language but arbitrary generated images or the results of image recognition, rather than inputting a prompt and outputting linguistic information.

[0085] For example, by using an image generation model in the thought estimation process, it is possible to generate images such as an ideal development view or an incision line.

[0086] For example, by configuring the system to use a Visual Grounding model for thought estimation processing, it is possible to perform image recognition tasks using instructions in any language. In the configuration shown in Figure 5, a specific image recognition model was placed before the language model (LM), but in this configuration, a Visual Grounding model is added after the language model (LM), and language and images are input to the Visual Grounding model, which then interprets the language and images and performs image recognition (for example, displaying an incision line).

[0087] [S8 (Output processing)] The output unit 106 changes the output format depending on the results of image recognition, the query and reference information from the user, and at least one of the situation judgment, decision-making, and thought process, and outputs at least one of the estimated situation judgment, decision-making, and thought process (for example, visually displaying it on a screen such as a monitor or display, or outputting it as audio to a headset, etc.).

[0088] The output unit 106 can classify the category of the query or the answer to the query (for example, determine whether it is a category related to an object (including parts of an organ, instruments, gauze, etc.)) and change the output format. The output unit 106 can change the display format of the sentences output by the language model (LM model) according to the results of image recognition. For example, the output unit 106 may avoid outputting sentences at a position on the surgical image that overlaps with an object included in the query or answer, or at the position of the outline of the object on the surgical image, or conversely, may output sentences at a position near the object. The output unit 106 can change the display format of the image recognition results depending on the content of the query or answer. For example, when the query or answer is related to an object, the output unit 106 may light up or blink the object on the surgical image, or display a sentence.

[0089] A method for classifying the above-mentioned query or answer category (for example, determining whether the category is related to an object (including parts of an organ, instruments, gauze, etc.)) will be described with reference to FIGS. 10 and 11.

[0090] FIG. 10 is a diagram illustrating classification of queries and answers according to an embodiment of the present invention. As shown in FIG. 10, the similarity between an embedding (vector) of a text collection (query database) of queries related to an object is calculated. A common method for calculating the similarity between texts is to calculate the distance between the texts by converting the texts into embeddings (vectors) and calculating the similarity. Using this method, if the similarity between a text collection (query database) of queries related to an object and a user's query is equal to or greater than a threshold, the query is determined to be related to the object, and the display format can be changed.

[0091] FIG. 11 is a diagram illustrating the classification of queries and answers according to an embodiment of the present invention. As shown in FIG. 11, the language model (LM) is trained to output tokens and the like that determine the category of a query or answer. As described above, when training the language model (LM), the annotation data used in training the language model (LM) is categorized, and learning is performed by providing text for Object (visible object), Action (behavior of visible object, etc.), Situation (current situation), Prediction (predicted next situation), Plan (action to be taken next), and Reason (reason that led to the Situation, Prediction, and Plan). By utilizing this annotation data, when outputting answers according to each category during training of the language model (LM), special tokens or specific words (prefixes) are assigned and output. In other words, when the language model (LM) outputs an answer, a flag indicating the category can be assigned to the answer. The output unit 106 refers to the special token or the like assigned to the answer, and can change the display format if the answer is assigned a token ([OBJ]) indicating an object.

[0092] When a method for determining whether a query or answer is related to such an object determines that the query or answer is related to "where to cut," the output unit 106 can display the location or direction of the cut (incision line, margin relative to the incision line, incision direction).

[0093] The output unit 106 may change the sentence (for example, shorten or lengthen the sentence) depending on whether the surgery assistance device 10 is operating during surgery (i.e., assumed to be operating for surgeons) or when the surgery assistance device 10 is operating outside of surgery (i.e., assumed to be operating for student education). The reference information may indicate whether the surgery assistance device 10 is operating during surgery or outside of surgery.

[0094] The output format is the surgeon's situational judgment, decision-making, and thought process during surgery, as shown below. An example is shown below. These situational judgment, decision-making, and thought processes are used as annotation data for the language model (LM) training data. 1. Situational assessment during incision and visual field development Determine appropriate port location and incision (optimize incision area considering potential adhesions. Adjust incision line based on tumor location and surrounding anatomy) - Procedures and areas of the surgical field to be unfolded - Unfolding method (if there are strong adhesions, choose between sharp or blunt dissection. Decide which tissue to pull with forceps, in which direction, and with what strength) 2. Situational assessment during main operations Identification of anatomical structures (confirmation and estimation of the location and type of exposed blood vessels, nerves, and organs, or those not yet exposed due to being covered by fat or membranes. What to do if abnormal anatomy is found, such as mutated blood vessels. Distinguishing and determining when there are similar anatomical structures. Changing strategies if the extent of cancer infiltration is wider than expected. Evaluation of blood flow and tumor location using ICG fluorescence, ultrasound, or endoscopy. Whether or not to perform cholangiography.) Resection decision-making (whether to dissect certain tissues or lymph nodes, and how far to dissect them; whether the resection or lymph node dissection was complete; whether normal tissue should be preserved or removed; whether additional resection is necessary; whether important nerves (e.g., autonomic nerves in rectal cancer surgery) should be preserved or removed; what is the procedure (sharp or blunt dissection; stapled or hand-sewn anastomosis)?) 3. Decision-making in the event of bleeding or injury - What to do in case of bleeding (Identifying the source of bleeding. Deciding how to stop the bleeding (can the bleeding be stopped by direct pressure? Should I clamp, clip, suture, or use an electric scalpel?). Deciding what to do in case of massive bleeding (should I quickly block off the main blood vessels and prioritize bleeding control? Should I use an autologous blood collection device (cell saver)?) 4. Decision-making on reconstruction and anastomosis -Selection of anastomosis method (end-to-end anastomosis, end-to-side anastomosis, or side-to-side anastomosis; hand-sewn anastomosis or stapled anastomosis; whether to evaluate blood flow using ICG fluorescence and expand the resection area if blood flow at the anastomosis site is poor) Anastomotic leak risk assessment (whether to re-anastomose if there are signs of poor blood flow, switch the reconstruction route, place a prophylactic drain, or consider a bypass or workaround) 5. Decision-making for unexpected situations - Strategy changes based on intraoperative findings (if unexpected metastasis such as peritoneal dissemination is discovered during surgery, should the radical surgery be continued, not resected, or changed to palliative surgery? What to do if another disease is discovered? Equipment problems and unforeseen circumstances (for example, if a surgical support robot or laparoscope cannot be used, should the procedure be switched to laparotomy or continued with equipment replacement? Should an X-ray be taken when gauze runs out?) 6. Thought process through surgery The thought process leading to situational assessment and decision-making (for example, recognizing the anatomy during surgery and noting that "the branching of blood vessels is abnormal," making an inference that "it looks like it's going to get even more curved further ahead," deciding that "if we continue this way, there will be a lot of bleeding," and taking action such as "changing the dissection line," a series of processes from vision → inference → judgment → action)

[0095] [Other output formats] Detailed labeling and enhanced interpretation In one embodiment of the present invention, the AI's confidence score (e.g., ambiguity when a recognized object belongs to multiple categories) is displayed for each object, allowing the user to modify it. Support for time series data In one embodiment of the present invention, the movement of an object is tracked over a series of surgical images (eg, frames of a video) and a description of the changes is displayed. Integration with generative AI In one embodiment of the present invention, related images and stories are automatically generated based on recognized objects and scenes, enabling their use in augmented reality (AR) and VR environments. Audio support In one embodiment of the present invention, surgical image descriptions are read aloud to improve accessibility for the visually impaired. Adding background knowledge In one embodiment of the present invention, information from Wikipedia and scientific papers related to recognized objects is provided to provide a deeper understanding of the context.

[0096] [S9 (Interaction Processing)] The interaction processing unit 107 changes the output format of information output by the output unit 106 in response to user operations (instructions input by the user via the input device 11, etc.). For example, the interaction processing unit 107 generates explanations in multiple languages, such as Japanese, English, and French, and makes them switchable with a click (multilingual translations are provided via prompts). For example, the interaction processing unit 107 reads out an explanation of a surgical image by clicking a button.

[0097] [S10 (Data recording)] The data storage unit 108 records the query and reference information entered by the user, the image recognition results, situation judgment, decision-making, and thought process information. The recorded information is used in a series of processes from the next time the user enters a query to the output of the image recognition results and answer. For example, it may be used as reference information or for image recognition that requires time-series processing.

[0098] <Interactive display> In one embodiment of the invention, the interaction processing unit 107 of the surgical support device 10 highlights and displays the detected object (e.g., a part of an organ, an instrument, etc.) or text obtained as a result of the image recognition in response to the user's operation on the image recognition results and text (i.e., sentences of situational judgment, decision-making, and thought process) output by the output unit 106.

[0099] For example, as shown in [Object Description] in Fig. 12, when the user clicks on a specific area of ​​the image on the screen (for example, a body part (organ, etc.), an object such as an instrument), the corresponding description is displayed in a highlighted manner. Specifically, when the position information of the object has been acquired by image recognition, and the user clicks on the area of ​​the object on the image on the screen (for example, clicking on the area of ​​the right recurrent laryngeal nerve), a prompt such as "Please explain the right recurrent laryngeal nerve" is input to the language model (LM), and the output result (for example, "It is the right recurrent laryngeal nerve. It is part of the vagus nerve and is a nerve that goes around the subclavian artery in the thoracic cavity") is displayed.

[0100] For example, an explanation of an object in the current situation may be given, as in [Explanation of the object in the current situation] in Fig. 12. A prompt such as "Please explain the current situation in the right recurrent laryngeal nerve" is input to the language model (LM), and the output result (for example, "The right recurrent laryngeal nerve. Risky behavior was observed when the surrounding area was dissected, so there is a risk of paralysis") is displayed.

[0101] The interaction processing unit 107 of the surgery support device 10 may prioritize outputting an "explanation of the current situation" during surgery, and prioritize outputting a "general explanation" when playing back recorded surgical images.

[0102] For example, as shown in FIG. 13, when a user clicks on a specific portion of text on the screen, the corresponding area of ​​the image is highlighted. Specifically, when position information of an object is obtained by image recognition, if the user clicks on a part of the text on the screen (such as an organ) or an object such as an instrument (for example, clicking on the "right recurrent laryngeal nerve"), the result of image recognition is used to display the position of that object on the image. For example, a transparent mask may be displayed to highlight an area such as the right recurrent laryngeal nerve, or the outline may be colored, or these transparent masks or outlines may be made to blink. The text of the clicked object and the highlighted display of the object in the image may be displayed in a corresponding color (for example, displayed in the same color).

[0103] <text style> In one embodiment of the present invention, the thought estimation unit 105 of the surgery support device 10 changes the style of text (i.e., sentences describing situational judgment, decision-making, and thought processes) to provide at least one of a general explanation, a specialized (medical, engineering, etc.) explanation, and an emotional (artistic, psychological, etc.) explanation. Specifically, the thought estimation unit 105 can generate multiple interpretations for the same surgical image, such as a general explanation (e.g., performing lymph node dissection around the right recurrent laryngeal nerve), a specialized explanation (e.g., performing bipolar dissection of the lymph nodes outside the right recurrent laryngeal nerve. Please perform the dissection carefully), and an emotional explanation (e.g., be very careful during surgery because nerve damage can cause complications). A user interface is provided for inputting a desired explanation method, and a prompt describing the desired explanation is input into a language model (LM) to generate an answer. Template prompts are defined in advance to change the input to the language model (LM). Specifically, when training a language model (LM), a style conversion module that performs training to solve the task of changing the style of text is attached after the output of the language model (LM).

[0104] For example, if a user wants a professional explanation, the system adds information about the explanation method to the user's query "What are you doing?" and changes the prompt to "What are you doing? Please explain professionally." For example, if a user wants an emotional explanation, the system adds information about the explanation method to the user's query "What are you doing?" and changes the prompt to "What are you doing? Please also tell me how the surgeon feels during the operation."

[0105] <Adding supplementary information to answers> In one embodiment of the present invention, auxiliary information such as disclaimers is added before and after the answer (leaving the final decision to the surgeon). When using AI, the answer (output) may make a plausible mistake such as hallucination, which can sometimes be fatal. In this way, by adding auxiliary information such as disclaimers before and after the answer, the surgeon can use the answer only as a reference.

[0106] For example, the following auxiliary information is available: (1) When a procedure is irreversible (cannot be undone), provide supplementary information to the answer. (2) Provide at least one answer with a score above the threshold and include the score as supporting information. (3) Provide supplementary information instructing participants to ask the same question (query) repeatedly but in different ways.

[0107] Specifically, the surgery support device 10 determines whether a surgical operation is irreversible based on the surgical image, query, and reference information, and if it is determined to be an irreversible surgical operation, it assigns auxiliary information to the answer. The surgery support device 10 also calculates a score (confidence score) for each answer and assigns the score as auxiliary information to at least one answer whose score is equal to or greater than a threshold. The surgery support device 10 also assigns auxiliary information to any answer, instructing the patient to repeatedly ask the same question in different wording.

[0108] This section explains how to provide supplemental information to an answer when an irreversible procedure is being performed. For example, the question is "What should I do next?" (1) Suggestions such as "remove the object" cannot be reversed, and if the notification is incorrect, it could be fatal, so if such an inference is made, no notification will be made. (2) Because a suggestion such as "remove the object" cannot be reversed and would be fatal if it is an incorrect notification, if such an inference is made, a disclaimer sentence is generated at the beginning or end of the suggestion to notify the surgeon. This allows the surgeon to make a final confirmation before resection.

[0109] For example, when performing lymph node dissection around the right recurrent nerve during robotic esophageal cancer surgery, resection of the right recurrent nerve itself must be avoided, and only the area surrounding it must be resected. Once a surgical "resection" is performed, there is no turning back, and irreversible complications may result. Therefore, if a response includes a suggestion for a procedure such as "resection," be sure to inform the patient to check the surrounding anatomical structures, that there is no turning back, and that the final decision is left to the surgeon.

[0110] [Determining whether supplementary information is required] Here, a description will be given of a determination as to whether or not auxiliary information should be added. In cases where it is necessary to add auxiliary information, it is determined whether or not the instruction is particularly related to an irreversible procedure (hereinafter, either method a or b may be used). a. Use of knowledge databases - Pre-register words and sentences related to irreversible procedures in the knowledge database Search the generated sentences against words and sentences of irreversible procedures in the knowledge database If the search returns a hit, the operation is determined to be an irreversible procedure. b. Other (variations) 1.Generate a language model. Obtain the score of words related to irreversible procedures from the attention map of the document, and determine whether the procedure is irreversible based on the score. 2. When training the language model (LM), special tokens are used to generate the output of irreversible manipulation operations. <information>Train it to assign 3. The generated sentences are judged using a language processing model to determine whether they refer to irreversible procedures.

[0111] [Generating answers with additional information] If it is determined that the answer should be given auxiliary information, the surgery support device 10 adds the auxiliary information and regenerates the answer sentence. For example, if the sentence generated by the language model (LM) is "Please cut the outside of the nerve," the surgery support device 10 adds auxiliary information as follows and presents the regenerated sentence (hereinafter, any one of methods a, b, and c may be used). a. Prepare a template statement and add it to the beginning or end of the generated document. Please cut the outside of the nerve. However, please use appropriate judgment as the procedure described in the answer is dangerous. Please use my answer as a reference and proceed with caution. Make sure to cut outside the nerve. b. Prompt the language model (LM) to regenerate the sentence by adding auxiliary information. For example, the prompt: "Please add auxiliary information to the sentence 'Please cut the outside of the nerve' and regenerate the sentence" is given to the language model (LM). - Cut outside the nerve. However, be sure to check the location of the nerve. The surgeon will make the cut. After carefully checking the location of the nerve, the surgeon should cut outside the nerve at his discretion. c. Provide auxiliary information as a sound Information audio Beep d. Use a trained language model (LM) to generate auxiliary sentences A language model (LM) is used that is trained to output auxiliary sentences when it is determined that an input related to an irreversible manipulative operation has been made based on the method for determining irreversible manipulative operations. In this case, the process of determining whether to add auxiliary information and adding the auxiliary information is included in the process of generating an answer.

[0112] This section explains the case where at least one answer with a score above a threshold is given (for example, the answer is read aloud) and the score is attached as auxiliary information (for example, the score is displayed). The questioner cannot know the degree of confidence in the answer. By notifying the confidence as auxiliary information, the questioner can understand the degree of confidence in the answer. In a right recurrent nerve perineal lymph node dissection during robotic esophageal cancer surgery, there are situations where the outside of the right recurrent nerve should be incised and situations where the inside should be incised. Depending on the confidence level of the answer, the surgeon can interpret the answer and proceed with the surgery. Also, even if the answer is incorrect, if the confidence level is displayed, the surgeon can interpret the incorrect answer and make the correct decision.

[0113] [Output multiple answer candidates and their likelihoods] 1. The output of the language model (LM) is calculated for each output (one token (one word)) using the softmax between all tokens in the dictionary held by the language model (LM) and one token in the output of the language model (LM). In other words, the likelihood is calculated for all words contained in the dictionary at this time. 2. To generate multiple sentences, the top K tokens are selected from the tokens with the highest likelihood to the dictionary tokens calculated by Softmax. 3. The language model (LM) infers the next token based on the token it has inferred. Based on the top K tokens obtained, it predicts the top K of the next tokens. By repeating this process to generate documents, it is possible to generate multiple sentences. 4. The language model (LM) calculates the likelihood of a sentence by averaging the log-likelihood of all output tokens.

[0114] [Generating answers with additional information] The surgery support device 10 adds auxiliary information to regenerate the answer sentence, and presents the regenerated sentence with likelihood information added as follows (hereinafter, any of methods a, b, and c may be used). a. Prepare a template sentence and add likelihood information to the beginning or end of the generated document. Cut outside the nerve, but make sure to check the nerve's location. Score "***". b. A prompt is given to the language model (LM) to regenerate the sentence by adding likelihood information. For example, the prompt: "Please regenerate the sentence by adding a confidence level of "***" to the sentence "Please cut the outside of the nerve." is given to the language model (LM). Cut outside the nerve, but make sure to check the location of the nerve. The confidence level for this answer is "***". c. The language model (LM) is prompted to regenerate a sentence that takes into account the likelihood information. For example, the prompt: "Please cut the outside of the nerve" is given a confidence level of "***" for the nerve detection (image recognition result) and the sentence is regenerated. Cut the outside of the nerve. However, the likelihood of detecting the nerve is low at "***". The accuracy of the answer may be low.

[0115] This section explains the case where auxiliary information is added to instruct the questioner to repeatedly ask the same question (query) using different expressions. When the question or answer is related to the target procedure, the questioner, such as a surgeon, is instructed to ask the same question using different expressions. This allows the surgeon to confirm that the answers are consistent. (For example, in response to the question "What would you like to do next?", the answer is "Perform a gastrectomy. Just to be sure, please ask the same question using different expressions.")

[0116] [Generating answers with additional information] If the question or answer is determined to be related to the target action, the surgery support device 10 regenerates the answer text to encourage the surgeon to ask the question again. The regenerated text is presented by adding a text encouraging the surgeon to ask the question again as follows (hereinafter, either method a or b may be used): a. Prepare a template text and add a text encouraging the surgeon to ask the question again to the beginning or end of the generated document. Perform a gastrectomy, but ask the same question in a different way. · Perform a gastrectomy, but ask the same question in different phrases to ensure consistency in your answers. b. A prompt is given to the language model (LM) to regenerate the sentence by adding a sentence that encourages the question to be asked again. For example, a prompt is given to the language model (LM) to regenerate the sentence by adding a sentence that asks the same question again in a different way to the sentence "Please perform a gastrectomy."

[0117] Examples of auxiliary information are shown below. (1) Explanation of points to keep in mind when performing the procedure (e.g., "Pay careful attention to the location of the nerve" when performing perineural resection) (2) Explanation of the adverse events that may occur if the target procedure is performed incorrectly (e.g., "If the resection site is incorrect, nerve paralysis may occur" when performing perineural resection). (3) An explanation that the surgeon bears the final responsibility (e.g., "The use of this system in this surgery is intended to supplement existing standard procedures, and the final diagnosis and treatment decisions are left to the surgeon's discretion.") (4) An explanation that there are certain limitations to the accuracy and safety of the system (e.g., "There is a possibility that the algorithm and measurement accuracy of this system may contain certain errors, and the surgeon should use it as an auxiliary means after fully understanding these limitations"). (5) Explanation that the system will be used in conjunction with standard procedures and will not be completely dependent on the system (e.g., "The use of this system in this surgery will be for the purpose of supplementing existing standard procedures"). (6) Proper explanation will be given to patients and their families, and consent will be obtained (e.g., "The use of this system will be fully explained to patients or their families in advance, and consent will be obtained before use"). (7) Explanation of compliance with applicable laws, regulations, and ethical guidelines (e.g., "The use of this system will be in compliance with applicable laws, regulations, and ethical guidelines, and patient safety will be the top priority.")

[0118] <Changes in how we respond in emergencies> In one embodiment of the present invention, when a question (query) requires urgent attention, the response method is changed so that the surgeon or other surgeon can immediately understand and recognize it. During surgery, the surgeon's hands are occupied, and especially in an emergency, they cannot take their eyes off the patient, so they do not have time to read or listen to the long generated text during the operation. In this way, by changing the response method in an emergency, medical personnel can immediately understand and recognize the situation and how to deal with it in an emergency.

[0119] Specifically, the surgery assistance device 10 determines whether or not an emergency situation exists based on the surgery image, query, and reference information, and if it is determined that an emergency situation exists, changes the way of replying.

[0120] The method of determining an emergency will be explained. 1. Add special tokens to the annotation data to indicate danger. To add special tokens to the training data, for example, add special tokens to the annotation text for data that contains sentences such as "it's dangerous, so do it urgently." <warning>Specifically, by calculating the similarity between the annotation data and a group of sentences that represent dangerous or emergency situations, a special token can be assigned to the annotation data. 2. Special Tokens <warning>The annotated data is used to train a Vision Encoder and a Language Model (LM). 3. LM (Language Model) further special tokens <warning>If a rating is given, the system will learn additional tasks such as shortening the description, limiting answers to keywords only, and adding proper nouns. 4. The Vision Encoder and LM (Language Model) trained through the above process are used to obtain natural language output. By learning in 2, the Vision Encoder and LM (Language Model) add special tokens to the output when danger occurs. <warning>Furthermore, by learning step 3, the LM (language model) can change the linguistic information it outputs in dangerous situations. 5. LM (Language Model) special tokens <warning>When this command is output, the output format of the user interface is changed (specifically, using voice volume or beeping).

[0121] Below are some examples of how to change the way you respond. (1) Summarize your answer in short sentences (e.g., "There is bleeding from the left mesenteric artery. Use a clip to stop bleeding." "There is a perforation, so proceed to laparotomy.") (2) Answer using keywords only (e.g., "bleeding, suction"). (3) Do not give multiple instructions at once, but clearly communicate each instruction one by one (e.g., "Remove this blood vessel, then stop the bleeding and insert a drain," rather than "Ligate the blood vessel. If OK, insert the drain next.") (4) Add proper nouns (e.g., be sure to include specific anatomical parts, etc.) (5) Use voice dynamics (e.g., use a calm tone for normal instructions and a louder voice for emergency instructions). (6) Direct instructions by name (e.g., including the name of the surgeon or assistant)

[0122] Furthermore, by using a classifier that recognizes anatomical structures and surgical procedures, answers that cannot be reverted can be automatically deleted. For example, after a classifier recognizes that a gastrectomy has already been completed, if the answer given to the question "What will you do next?" is "Perform a gastrectomy," the answer will not be output (because gastrectomy is not usually performed twice).

[0123] Below are some examples of emergency situations and suggestions. Bleeding: For massive bleeding or major vascular injury, pressure hemostasis, hemostasis with electric scalpels, clips, and sutures, use of vascular sealing devices, arrangement of emergency blood transfusion (red blood cells, fresh frozen plasma, platelets), immediate occlusion of blood flow with forceps, vascular anastomosis or bypass surgery, preparation of artificial vascular grafts, etc. Circulation: In the case of cardiac arrest, hypotension, hypertension, or air embolism, cardiac massage (CPR) is initiated, adrenaline is administered intravenously, oxygen is administered, and artificial ventilation is performed. Ensuring circulating blood volume, identifying bleeding sites, and controlling hemostasis. Fluid transfusions, vasopressors (noradrenaline, phenylephrine), and anesthetic adjustments are also performed. Breathing: For airway obstruction, pneumothorax, and hypoxemia, secure the airway (bag-mask ventilation, re-intubation), administer steroids (in case of laryngeal edema), immediately perform chest drainage, needle deaeration, administer oxygen, adjust artificial ventilation, and reconfirm tracheal intubation. Neurology: For intraoperative awakening and cerebral infarction, additional anesthetics should be administered, BIS monitoring (electroencephalography) should be performed, prompt postoperative CT / MRI blood pressure management should be performed, and anticoagulant therapy should be considered. Infection: In case of intraoperative sepsis or intestinal perforation, wound cleaning and antibiotic administration are performed. In case of septic shock, noradrenaline is administered, and the perforation is repaired by suturing. Intraperitoneal cleaning and broad-spectrum antibiotic administration are performed. Anesthesia: For malignant hyperthermia and anaphylaxis, dantrolene administration, cooling treatment (ice-cold infusion, ice packs), intravenous adrenaline, steroids, and antihistamines. Instruments: In the event of a surgical equipment failure or robot trouble, prepare spare instruments, interrupt the surgery and select an alternative procedure, switch to manual operation, or proceed to open surgery.

[0124] <Query example> This includes the locations of objects that appear during surgery (organs, instruments, gauze, bleeding, liquids such as bile and pancreatic juice, sutures), actions taken by the surgeon, assistants, and nurses during surgery (resection, coagulation, hemostasis, etc.), progress of the surgery (lymph node dissection in progress, completion of gastrectomy, etc.), reasons for events and actions (reason the surgeon performed the resection, etc.), suggested actions and points to note (the next incision line to dissect, points to note for the next action), evaluation of the technique, etc. Other items include those listed below.

[0125] (1) Questions about anatomy and physiology (What is the detailed anatomy of the area to be operated on? What are the blood vessels and nerves? What is the function of this organ? What important structures are nearby (blood vessels, nerves, lymph nodes, etc.)? What is the boundary between the area that should be removed and the area that should be left intact? How much bleeding is physiologically acceptable? Is the color or hardness of a particular tissue normal or pathological? etc.) (2) Questions about surgical techniques (how to hold the forceps correctly, what is the appropriate amount of force to use, where and how to make the incision for safety, to what layer should the incision be made, what are the appropriate spacing and strength of the sutures, what is the appropriate method of tying the suture (buried suture, running suture, simple ligation, etc.), what is the appropriate method for dissection, what is the direction and angle of excision, how can bleeding be minimized, how should tissue be held to prevent damage, which device (electric scalpel, ultrasonic scalpel, etc.) should be used, which forceps is appropriate for grasping tissue, etc.) (3) Questions about how to use surgical equipment and tools (What is the official name of this tool? What is the correct way to hold and use this tool? When should a specific tool be used? What is the appropriate wattage setting for this electric scalpel? What are the different types of sutures and their characteristics? Which type of needle should be selected (round, square, inverted triangular, etc.)? How should drains be inserted and what type is appropriate? How should tissue adhesives and hemostatic agents be used and what are their applications?) (4) Questions about anesthesia and general management (What type of anesthesia and its effect on surgery? How should changes in blood pressure and pulse rate be handled during surgery? At what level of blood loss does a blood transfusion become necessary? How should body temperature be managed during surgery? What should be done if there are problems with airway management? How should pain be managed after surgery? How should body fluids be managed during surgery (fluid infusion and electrolyte adjustment)?) (5) Questions about intraoperative complications and troubleshooting (How should you stop bleeding if an artery is injured? What is the best way to deal with venous injury? What should you do if unexpected adhesions form? What should you do if an instrument malfunctions? What should you do if a tumor is ruptured during surgery? What should you do if the patient's blood pressure drops suddenly during surgery? What should you do if a sudden pneumoperitoneum (air leak during laparoscopic surgery) occurs? What should you do if an instrument is dropped or sterilization is compromised?) (6) Questions regarding postoperative care (What are the risks of postoperative complications (SSI, intestinal obstruction, suture leakage, etc.)? What is the appropriate timing for extubation? What is the optimal method for wound management? What are the criteria for resuming eating after surgery? How should postoperative drainage be managed? What are the discharge criteria and frequency of follow-up? Is postoperative antibiotic administration necessary?) (7) Questions about the patient's condition and surgical indications (Why was this patient a surgical indication? Were there no other options? What surgical options were available? What points should not be overlooked in the preoperative evaluation? If the patient had complications, how should the surgical procedure be changed?) (8) Questions about the selection of surgical procedures and their rationale (Why was this procedure chosen (open vs. laparoscopic vs. robotic surgery)? What other procedures were available for similar cases? What is the evidence (success rate and complication rate) for this procedure? Why has the procedure changed (modified surgery or new approach)? Why do procedures differ depending on the institution or surgeon even for the same disease?) (9) Questions about medical safety and ethics (What should be explained in the preoperative informed consent? If an unexpected situation occurs during surgery, to what extent should the family be informed? What should be done if a medical error occurs? What should be done if the patient requests DNR (Do Not Resuscitate)? What ethical considerations should be taken into account when a medical intern intervenes during surgery for educational purposes?) (10) Questions about teamwork and communication during surgery (How can I accurately understand the surgeon's intentions? How can I ask questions at the right time during surgery? How can I smoothly cooperate with nurses? How can I cooperate with anesthesiologists? How can I manage stress during surgery? What are the tips for maintaining concentration during surgery?)

[0126] <Input other than surgical images> In one embodiment of the present invention, in addition to surgical images, other information may also be acquired and input into the visual language model.

[0127] (1) Visual characteristics specific to surgery Anatomical structures and surgical procedures: Changes in the position and shape of anatomical structures such as the recurrent laryngeal nerve, pancreas, stomach, and colon, as well as progress of surgical procedures such as "lymph node no. 6 dissection" and their completion status. Operating room environment: Medical staff wearing sterile blue or green surgical gowns, gloves, and masks Surgical instruments: types of instruments used in laparoscopic, thoracoscopic, and robotic-assisted surgery, such as electric scalpels, bipolar forceps, monopolar forceps, Potts scissors, suction tubes, and sealing devices, as well as the position and operation of the instruments, changes in frequency of instrument use, timing of use, reactions to instrument operation (sparks, power, smoke, amount of aspirated fluid discharged, quality, and color), instrument troubles, suture needles, etc. Patient: Changes in the patient's skin or face, changes in vital signs before, during, or after surgery Blood: Quantity and quality of blood, gauze, etc. Feedback information for robotic surgery: range of motion, mechanical information, arm position (2) Context information Marks and logos: Hospital names and medical institution logos Healthcare professionals: doctors and nurses (3) Metadata Photo data: EXIF ​​information of the image, etc. File names and associated data: image file names, etc. (4) Comparison with external information Similarity Image Search: Inference from Surgical Images on the Internet (5) Patient and surgical information Patient's gender, age, stage of cancer, blood tests, imaging tests such as CT, MRI, and endoscopy, whether or not preoperative treatment (chemotherapy, radiation therapy, etc.) was performed, medical history, past medical history, lifestyle history, family history, height, weight, BMI, blood type, history of allergies, surgical procedure, surgical instruments to be used, operating room, scheduled time of surgery, attending physician, whether or not a blood transfusion was performed, whether or not a drain was inserted, abnormalities or complications during surgery, type of surgical support robot, console operation time, use of 3D navigation, application of real-time image guidance, differences from planned surgery, omissions and risks in preoperative preparation, etc.

[0128] <Effects> In one embodiment of the present invention, it is possible to know how experienced doctors and other medical professionals make situational judgments and decisions during surgery, and also to know the thought process that leads to situational judgments and decision-making.

[0129] <Hardware configuration> 14 is a hardware configuration diagram of a surgery support device 10 according to one embodiment of the present invention. The surgery support device 10 can include a control unit 1001, a main memory unit 1002, an auxiliary memory unit 1003, an input unit 1004, an output unit 1005, and an interface unit 1006. Each of these will be described below.

[0130] The control unit 1001 is a processor (for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc.) that executes various programs installed in the auxiliary storage unit 1003.

[0131] The main memory unit 1002 includes a non-volatile memory (Read Only Memory (ROM)) and a volatile memory (Random Access Memory (RAM)). The ROM stores various programs, data, etc. required for the control unit 1001 to execute various programs installed in the auxiliary memory unit 1003. The RAM provides a working area into which the various programs installed in the auxiliary memory unit 1003 are expanded when executed by the control unit 1001.

[0132] The auxiliary storage unit 1003 is an auxiliary storage device that stores various programs and information used when the various programs are executed.

[0133] The input unit 1004 is an input device through which the operator of the surgery assistance apparatus 10 inputs various instructions to the surgery assistance apparatus 10 .

[0134] The output unit 1005 is an output device that outputs the internal state of the surgery assistance device 10 and the like.

[0135] The interface unit 1006 is a communication device for connecting to a network and communicating with other devices.

[0136] The above-disclosed embodiments include, for example, the following aspects. (Appendix 1) an image acquisition unit for acquiring a surgical image; an information acquisition unit that acquires queries and reference information from users; a thought estimation unit that performs image recognition on the surgical image and estimates at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image using the query from the user and the reference information; an output unit that changes an output format in accordance with the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputs at least one of the situation judgment, the decision-making, and the thought process to the user; A surgical assistance device comprising: (Appendix 2) The surgical support device according to claim 1, wherein the thought estimation unit generates, from the surgical image, a feature that serves as a basis for at least one of the situation judgment, the decision-making, and the thought process. (Appendix 3) The surgical support device described in Appendix 1 is characterized in that the thought estimation unit generates features that form the basis of at least one of the situation judgment, the decision-making, and the thought process from the surgical image by using one or more image recognition models, such as a Vision Encoder that has learned embedded representations of images and text, and an image recognition model that performs given surgical image recognition tasks of organ detection, instrument detection, surgical step recognition, action recognition, and bleeding detection. (Appendix 4) The surgical support device according to any one of appendices 1 to 3, wherein the thought estimation unit inputs the features of the surgical image and the features of the image recognition result, as well as a query from the user and the reference information, into a language generation model, and outputs at least one of the situation judgment, the decision-making, and the thought process in natural language. (Appendix 5) 5. The surgical support device according to claim 4, wherein the thought estimation unit is trained using categorized annotation data. (Appendix 6) 4. The surgical support device according to claim 1, wherein the output unit outputs at least one of the situation judgment, the decision-making, and the thought process together with the result of the image recognition. (Appendix 7) The surgical support device of claim 6, wherein if the user's query is related to a specific thing or a specific event, the output unit illustrates the image recognition results on the image. (Appendix 8) When the user's query is a query about a detected object obtained as a result of the image recognition, 8. The surgical support device of claim 7, wherein the output unit displays the detected object in the surgical image or text in an emphasized manner. (Appendix 9) If the user's query is a query regarding an incision location, 8. The surgical support device according to claim 7, wherein the output unit displays the incision position in the surgical image in an emphasized manner. (Appendix 10) The surgical support device according to any one of appendices 1 to 3, wherein the output unit changes the display position of at least one of the situation judgment, the decision-making, and the thought process based on the results of the image recognition. (Appendix 11) The surgical support device of claim 10, wherein the output unit displays at least one of the situation judgment, the decision-making, and the thought process so as not to overlap with the detected object obtained as a result of the image recognition. (Appendix 12) The surgical support device of claim 10, wherein the output unit displays at least one of the situation assessment, the decision-making, and the thought process near a detected object obtained as a result of the image recognition. (Appendix 13) 7. The surgical assistance device of claim 6, further comprising an interaction processing unit that highlights and displays a detected object or text obtained as a result of the image recognition in accordance with the user's operation on the image recognition result and text output by the output unit. (Appendix 14) The surgical assistance device of Appendix 4, wherein the thought estimation unit regenerates a query according to the result of the image recognition, the query from the user and the reference information, and at least one of the situation judgment, the decision-making, and the thought process, or changes the style of the natural language to be output by using a language model trained to explicitly change the output, or by changing output parameters of the language model. (Appendix 15) The surgical support device described in Appendix 14, wherein the thought estimation unit changes the style of natural language depending on the state of the reference information, which relates to a specific thing or a specific event, and the state of the thing or event. (Appendix 16) If the reference information indicates a proficiency level of the user, The surgical support device of claim 15, wherein the output unit outputs text that is concise or detailed depending on the user's level of proficiency. (Appendix 17) If the reference information indicates whether the surgery assistance device is operating during surgery or outside of surgery, The surgical support device of claim 15, wherein the output unit outputs text that is concise or detailed depending on the operating status. (Appendix 18) The surgical support device described in Appendix 14, wherein the thought estimation unit changes the style of the text to provide at least one of a general explanation, a technical explanation, and an emotional explanation. (Appendix 19) A surgical assistance device as described in any of appendices 1 to 3, further comprising a learning unit that applies the results of the image recognition to template sentences, converts them into input for a language generation model, and learns the language generation model. (Appendix 20) an auxiliary information addition determination unit that determines whether or not to add auxiliary information to the output in response to the user query and / or the output of the thought estimation unit; 4. The surgery support device according to claim 1, wherein an output of the thought estimation unit is changed depending on a determination result of the auxiliary information addition determination unit. (Appendix 21) The surgical support device of claim 20, wherein the auxiliary information addition determination unit determines to add auxiliary information to the response when the user's query and / or the output of the thought estimation unit is related to an irreversible procedural operation. (Appendix 22) 21. The surgical support device according to claim 20, wherein the thought estimation unit adds, when adding auxiliary information, a likelihood of an image recognition result by the thought estimation unit or a likelihood of generation of linguistic information as auxiliary information. (Appendix 23) 21. The surgical support device according to claim 20, wherein the thought estimation unit, when adding auxiliary information, adds auxiliary information instructing the answerer to repeatedly ask the same question in different phrases. (Appendix 24) an emergency state determination unit that determines whether the surgical image used for the input is an emergency state by using information obtained in the process of the user's query and / or the thought estimation unit; A surgical support device as described in any one of appendices 1 to 3, wherein the output of the thought estimation unit and the display format of the output unit are changed depending on the judgment result of the auxiliary information addition judgment unit, depending on the judgment result of the emergency state. (Appendix 25) A surgical support device as described in Appendix 24, wherein when the emergency state determination unit determines that an emergency state exists, the thought estimation unit outputs a proper noun or keyword and outputs a short sentence that does not include multiple instructions in one sentence. (Appendix 26) A surgery support device as described in Appendix 24, wherein, when the emergency state determination unit determines that an emergency state exists, the thought estimation unit outputs the names of the members participating in the surgery when instructed. (Appendix 27) 25. The surgical support device according to claim 24, wherein when the emergency state determination unit determines that an emergency state exists, the output unit outputs sound with varying volume. (Appendix 28) For surgical support devices, acquiring surgical images; Obtaining queries and references from users; performing image recognition on the surgical image, and using the query from the user and the reference information, estimating at least one of a situation judgment, a decision, and a thought process leading to the situation judgment or the decision regarding the surgical image; changing an output format according to the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputting at least one of the situation judgment, the decision-making, and the thought process to the user; A program that executes the following. (Appendix 29) A method performed by a surgical assistance device, comprising: acquiring a surgical image; receiving queries and references from users; performing image recognition on the surgical image and using the query from the user and the reference information to estimate at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image; a step of changing an output format according to the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputting at least one of the situation judgment, the decision-making, and the thought process to the user; A method comprising:

[0137] Although the examples of the present invention have been described in detail above, the present invention is not limited to the specific embodiments described above, and various modifications and variations are possible within the scope of the gist of the present invention as set forth in the claims. [Explanation of symbols]

[0138] 1. Surgical support system 10 Surgical support equipment 11 Input Devices 12 Output Devices 20 Imaging device 101 Image acquisition unit 102 Information acquisition department 103 Knowledge database processing unit 104 Knowledge database storage unit 105 Thought estimation part 106 Output section 107 Interaction Processing Unit 108 Data storage unit 1001 control section 1002 Main memory 1003 Auxiliary storage unit 1004 Input section 1005 Output section 1006 Interface section< / warning> < / warning> < / warning> < / warning> < / warning> < / information>

Claims

1. an image acquisition unit for acquiring a surgical image; an information acquisition unit that acquires queries and reference information from users; a thought estimation unit that performs image recognition on the surgical image and estimates at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image using the query from the user and the reference information; an auxiliary information addition determination unit that determines to add auxiliary information to the output of the thought estimation unit when at least one of the query from the user and the output of the thought estimation unit is related to an irreversible procedural operation; an output unit that changes an output format in accordance with the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputs at least one of the situation judgment, the decision-making, and the thought process to the user; Equipped with A surgical assistance device, wherein the reference information relates to a specific thing or a specific event, and the thought estimation unit changes the natural language style of the output of the thought estimation unit depending on the state of the thing or the event.

2. An image acquisition unit for acquiring surgical images; an information acquisition unit that acquires queries and reference information from users; a thought estimation unit that performs image recognition on the surgical image and estimates at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image using the query from the user and the reference information; an emergency state determination unit that determines whether the surgical image is an emergency state using at least one of a query from the user and information obtained during the processing of the thought estimation unit; an output unit that changes an output format in accordance with the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputs at least one of the situation judgment, the decision-making, and the thought process to the user; Equipped with an output of the thought estimation unit or an output form of the output unit is changed according to a result of the determination of an emergency state; A surgical assistance device, wherein the reference information relates to a specific thing or a specific event, and the thought estimation unit changes the natural language style of the output of the thought estimation unit depending on the state of the thing or the event.

3. The surgery support device according to claim 1 , wherein the thought estimation unit generates, from the surgical image, a feature quantity that serves as a basis for at least one of the situation judgment, the decision-making, and the thought process.

4. 3. The surgical support device according to claim 1, wherein the thought estimation unit generates features that form the basis of at least one of the situation judgment, the decision-making, and the thought process from the surgical image by using one or more image recognition models that perform given surgical image recognition tasks, such as a Vision Encoder that has learned embedded representations of images and text, organ detection, instrument detection, surgical step recognition, action recognition, and bleeding detection.

5. 3. The surgical support device according to claim 1, wherein the thought estimation unit inputs the features of the surgical image and the features of the image recognition result, as well as a query from the user and the reference information, into a language generation model, and outputs at least one of the situation judgment, the decision-making, and the thought process in natural language.

6. The surgery support device according to claim 5 , wherein the thought estimation unit is trained using categorized annotation data.

7. The surgery assistance device according to claim 1 , wherein the output unit outputs at least one of the situation judgment, the decision-making, and the thought process together with the result of the image recognition.

8. The surgery assistance device according to claim 7 , wherein when the query from the user is related to a specific thing or a specific event, the output unit displays the result of the image recognition on an image.

9. When the query from the user is a query about a detected object obtained as a result of the image recognition, The surgery assistance device according to claim 8 , wherein the output unit displays the detected object in the surgical image or text in an emphasized manner.

10. If the query from the user is a query regarding an incision position, The surgery assistance device according to claim 8 , wherein the output unit displays the incision position in the surgery image in an emphasized manner.

11. The surgery support device according to claim 1 or 2, wherein the output unit changes a display position of at least one of the situation judgment, the decision-making, and the thought process based on the result of the image recognition.

12. The surgery assistance device according to claim 11 , wherein the output unit displays at least one of the situation judgment, the decision-making, and the thought process so as not to overlap with a detected object obtained as a result of the image recognition.

13. The surgery assistance device according to claim 11 , wherein the output unit displays at least one of the situation assessment, the decision-making, and the thought process near a detected object obtained as a result of the image recognition.

14. 8. The surgical assistance device according to claim 7, further comprising an interaction processing unit that highlights and displays a detected object or the text obtained by the result of the image recognition in accordance with the result of the image recognition output by the output unit and an operation by the user on the text.

15. 6. The surgical assistance device according to claim 5, wherein the thought estimation unit regenerates a query to change the output, or uses a language generation model trained to explicitly change the output, or changes output parameters of a language generation model, in accordance with the result of the image recognition, the query from the user and the reference information, and at least one of the situation judgment, the decision-making, and the thought process, thereby changing the style of the natural language to be output.

16. If the reference information indicates a proficiency level of the user, The surgery assistance device according to claim 1 , wherein the output unit outputs a simple or detailed sentence depending on the user's level of proficiency.

17. If the reference information indicates whether the surgery assistance device is operating during surgery or outside of surgery, The surgery support device according to claim 1 or 2, wherein the output unit outputs a simple or detailed text depending on the operation status.

18. The surgery support device according to claim 1 or 2, wherein the thought estimation unit changes the style of the text to provide at least one of a general explanation, a technical explanation, and an emotional explanation.

19. The surgery assistance device according to claim 1 or 2, further comprising a learning unit that applies the result of the image recognition to a template sentence, converts the result into an input for a language generation model, and learns the language generation model.

20. the auxiliary information addition determination unit determines whether to add auxiliary information to the output of the thought estimation unit in response to at least one of a query from the user and an output of the thought estimation unit, The surgery support device according to claim 1 , wherein an output of the thought estimation unit is changed depending on a determination result of the auxiliary information addition determination unit.

21. The surgery support device according to claim 20 , wherein the thought estimation unit adds, when adding auxiliary information, a likelihood of the result of the image recognition by the thought estimation unit or a likelihood of generation of linguistic information as auxiliary information.

22. The surgery support device according to claim 20 , wherein the thought estimation unit, when adding the auxiliary information, adds auxiliary information to the output of the thought estimation unit instructing the patient to repeatedly ask the same question in different expressions.

23. The surgical support device of claim 2, wherein when the emergency state determination unit determines that an emergency state exists, the thought estimation unit outputs a proper noun or keyword and outputs a short sentence that does not include multiple instructions in one sentence.

24. The surgery support device according to claim 2 , wherein, when the emergency state determination unit determines that an emergency state exists, the thought estimation unit outputs names of members participating in the surgery when instructed.

25. The surgery assistance device according to claim 2 , wherein when the emergency state determination unit determines that an emergency state exists, the output unit outputs a voice with varying volume.

26. For surgical support devices, acquiring surgical images; Obtaining queries and references from users; performing image recognition on the surgical image, and using the query from the user and the reference information, estimating at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image; determining, when at least one of the query from the user and the result of the estimation is related to an irreversible procedural operation, to add auxiliary information to the result of the estimation; changing an output format according to the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputting at least one of the situation judgment, the decision-making, and the thought process to the user; Execute The reference information relates to a specific thing or a specific event, and the program changes the natural language style of the estimation result depending on the state of the thing or the event.

27. ​​A surgical assistance device, acquiring surgical images; Obtaining queries and references from users; performing image recognition on the surgical image, and using the query from the user and the reference information, estimating at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image; determining whether the surgical image is in an emergency state using at least one of a query from the user and information obtained during the estimation process; changing an output format according to the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputting at least one of the situation judgment, the decision-making, and the thought process to the user; Execute The result of the estimation or the output form is changed depending on the result of the determination of the emergency state; The reference information relates to a specific thing or a specific event, and the program changes the natural language style of the estimation result depending on the state of the thing or the event.

28. A method performed by a surgical assistance device, comprising: acquiring a surgical image; receiving queries and references from users; performing image recognition on the surgical image and using the query from the user and the reference information to estimate at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image; determining, when at least one of the query from the user and the result of the estimation is related to an irreversible procedural operation, to add auxiliary information to the result of the estimation; a step of changing an output format according to the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputting at least one of the situation judgment, the decision-making, and the thought process to the user; Including, The method, wherein the reference information relates to a specific thing or a specific event, and the natural language style of the estimation result is changed depending on the state of the thing or the event.

29. A method executed by a surgical assistance device, comprising: acquiring a surgical image; receiving queries and references from users; performing image recognition on the surgical image and using the query from the user and the reference information to estimate at least one of a situation judgment, a decision-making, and a thought process leading to the situation judgment or the decision-making regarding the surgical image; determining whether the surgical image is in an emergency state using at least one of a query from the user and information obtained during the estimation process; a step of changing an output format according to the result of the image recognition, the query from the user, the reference information, and at least one of the situation judgment, the decision-making, and the thought process, and outputting at least one of the situation judgment, the decision-making, and the thought process to the user; Including, The result of the estimation or the output form is changed depending on the result of the determination of the emergency state; The method, wherein the reference information relates to a specific thing or a specific event, and the natural language style of the estimation result is changed depending on the state of the thing or the event.

Citation Information

Patent Citations

  • Surgical robot system, computer readable storage medium, and electronic device

    CN118121312A

  • Surgical Recognition System

    JP2020532347A

  • Phamaceutical composition for preventing or treating Coronavirus infectious disease comprising Coronavirus G-quadruplex ligands as an active ingredient

    KR1020220037944A

  • Tool tracking during surgical procedures

    US20140341424A1

  • System and method for determining, adjusting, and managing resection margin about a subject tissue

    US20210196423A1