Method and apparatus for visual multi-modality based surgical stage recognition
The method and apparatus improve surgical stage recognition by employing visual multi-modality techniques to extract and fuse kinematics-based indices, training an AI model to accurately recognize surgical stages, addressing limitations of previous methods.
Patent Information
- Application Number
- JP2025529931
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2023-09-22
- Publication Date
- 2025-11-14
AI Technical Summary
Existing methods for recognizing surgical stages are limited in their ability to account for interactions between surgical instruments, organs, and other factors such as camera cleaning and bleeding control, leading to inaccurate surgical stage recognition.
A method and apparatus that utilize visual multi-modality by extracting visual kinematics-based indices from surgical videos, fusing feature data using a fusion module, and training an AI model to recognize surgical stages based on these data.
Enhances the accuracy of surgical stage recognition by integrating multiple visual modalities, allowing for more precise analysis of surgical progress and actions.
Smart Images

Figure 2025537349000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a method and apparatus for surgical stage recognition, and more particularly to a method and apparatus for surgical stage recognition based on visual multi-modality. [Background technology]
[0002] Accurate recognition and analysis of surgical stages can optimize the progress of a surgery by providing efficient communication and accurate situational assessment among surgical parties. Accurate recognition of surgical stages can also be useful in post-operative patient monitoring and in classifying common surgical procedures to provide educational materials.
[0003] However, recognizing surgical stages is a difficult task that involves interactions between surgical instruments, organs included in the surgical field, camera cleaning, bleeding control, etc. Previously, research has been conducted on techniques for automatically recognizing surgical stages by analyzing surgical images, but these techniques have had limitations in that they cannot take into account all of the interactions described above related to surgical stages. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure has been made in view of the above circumstances, and an object of the present disclosure is to provide a method and apparatus for recognizing surgical stages based on visual multi-modality.
[0005] The problems that the present disclosure aims to solve are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below. [Means for solving the problem]
[0006] To solve the above technical problem, a method for recognizing surgical stages based on visual multiple modalities, performed by an apparatus according to the present disclosure, includes the steps of: extracting a plurality of visual kinematics-based indices based on a surgical video consisting of a plurality of frames corresponding to a plurality of surgical stages; acquiring first feature data for the surgical video and second feature data for the plurality of visual kinematics-based indices; acquiring fused third feature data by applying a fusion module trained to fuse data to the first feature data and the second feature data; and training a first artificial intelligence (AI) model to recognize each of the plurality of surgical stages based on the third feature data.
[0007] In addition, an apparatus according to the present disclosure for solving the above-mentioned technical problems includes a memory storing at least one process for recognizing surgical stages based on multiple visual modalities, and a processor that performs an operation of recognizing the surgical stages by executing the process, wherein the processor extracts a plurality of visual kinematics-based indices based on surgical images consisting of a plurality of frames corresponding to a plurality of surgical stages, obtains first data features for the surgical images, obtains second feature data for the plurality of visual kinematics-based indices, obtains fused third feature data by applying a fusion module trained to fuse data to the first feature data and the second feature data, and trains a first artificial intelligence (AI) model to recognize each of the plurality of surgical stages based on the third feature data.
[0008] In addition, a computer program stored on a computer-readable recording medium for embodying the present disclosure may also be provided.
[0009] In addition, a computer-readable recording medium having a computer program for implementing the present disclosure recorded thereon may also be provided. [Effects of the Invention]
[0010] According to the solution to the aforementioned problem of the present disclosure, a method and apparatus for recognizing surgical stages based on visual multi-modality can be provided.
[0011] According to the solution to the above-mentioned problem of the present disclosure, a method and apparatus can be provided for training an artificial intelligence model that can more accurately recognize surgical stages based on images showing the progress of the surgery and information about the surgical actions.
[0012] The effects of the present disclosure are not limited to those mentioned above, and other effects not mentioned above will be clearly understood by those skilled in the art from the description below. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a schematic diagram of a system for implementing a method for recognizing surgical stages based on visual multiple modalities according to one embodiment of the present disclosure; [Figure 2] FIG. 1 is a block diagram illustrating the configuration of an apparatus for recognizing surgical stages based on visual multiple modalities according to an embodiment of the present disclosure. [Figure 3] 1 is a flowchart illustrating a method for recognizing surgical stages based on visual multiple modalities according to one embodiment of the present disclosure. [Figure 4] FIG. 1 illustrates the overall architecture of a method for surgical stage recognition based on visual multiple modalities. [Figure 5] 10A and 10B are diagrams illustrating a process of extracting feature data from a surgical image to recognize a surgical stage according to an embodiment of the present disclosure. [Figure 6] FIG. 10 is a diagram illustrating a process of extracting third feature data by a fusion module according to an embodiment of the present disclosure. [Figure 7] FIG. 10 is a diagram illustrating a process in which an apparatus according to an embodiment of the present disclosure recognizes a surgical stage using a trained AI model. DETAILED DESCRIPTION OF THE INVENTION
[0014] The same reference numerals refer to the same elements throughout this disclosure. This disclosure does not describe all elements of each embodiment, and general content in the technical field to which this disclosure pertains or overlapping content in the embodiments will be omitted. The terms "unit, module, component, block" used in this specification may be embodied in software or hardware, and depending on the embodiment, multiple "units, modules, components, blocks" may be embodied as one component, or one "unit, module, component, block" may include multiple components.
[0015] Throughout this specification, when a part is said to be "coupled" to another part, this includes not only direct coupling but also indirect coupling, and indirect coupling includes connection via a wireless communication network.
[0016] Furthermore, when a part is described as "comprising" a certain element, this does not mean that it excludes other elements, but that it may further include other elements, unless otherwise specified.
[0017] Throughout this specification, when an element is said to be "on" another element, this includes not only when the element is in contact with the other element, but also when there is another element between the two elements.
[0018] The terms "first," "second," etc. are used to distinguish one component from another, and the components are not limited to the terms described above.
[0019] The singular expression includes the plural expression unless the context clearly indicates otherwise.
[0020] The identification numbers used in each step are for convenience of explanation, and do not dictate the order of the steps; the steps may be performed in a different order than specified unless the context clearly dictates a particular order.
[0021] The working principle and embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0022] The term "device according to the present disclosure" as used herein includes all of the various devices capable of performing computations and providing results to a user. For example, the device according to the present disclosure may include all of a computer, a server device, and a portable terminal, or may take any one of these forms.
[0023] Here, the computer may include, for example, a notebook computer, a desktop computer, a laptop computer, a tablet PC, a slate PC, etc., equipped with a web browser.
[0024] The server device is a server that communicates with external devices and processes information, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.
[0025] The portable terminal may be, for example, a wireless communication device that ensures portability and mobility, and may include any kind of handheld-based wireless communication device such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, smartphone, etc., as well as wearable devices such as watches, rings, bracelets, anklets, necklaces, glasses, contact lenses, or head-mounted devices (HMD).
[0026] For purposes of describing this disclosure, a "user" is a medical professional, and may be, but is not limited to, a doctor, nurse, clinical pathologist, medical imaging specialist, or technician who repairs / controls medical equipment.
[0027] For purposes of describing this disclosure, "surgery" refers to a surgical procedure in which an incision is made in the skin or mucous membrane to treat disease or injury, and "surgical instruments" can refer to any instrument used to perform surgery.
[0028] For purposes of describing this disclosure, "visual multi-modality" may refer to multiple types of data that are visually embodied (e.g., surgical video data and visual kinematics-based indexes, etc.).
[0029] FIG. 1 is a schematic diagram of a system 1000 for implementing a method for surgical stage recognition based on visual multiple modalities according to one embodiment of the present disclosure.
[0030] As shown in FIG. 1 , a system 1000 for implementing a method for recognizing surgical stages based on visual multi-modality may include an apparatus 100, a hospital server 200, a database 300, and an AI model 400.
[0031] 1, the device 100 is shown as being implemented in the form of a single desktop, but is not limited thereto. As described above, the device 100 may refer to various types of devices or a group of devices to which one or more types of devices are connected.
[0032] The device 100, hospital server 200, database 300, and artificial intelligence (AI) model 400 included in the system 1000 can communicate via a network W. Here, the network W can include a wired network and a wireless network. For example, the network can include various networks such as a local area network (LAN), a metropolitan area network (MAN), and a wide area network (WAN).
[0033] The network W may also include the well-known World Wide Web (WWW). However, the network W according to the embodiments of the present disclosure is not limited to the networks listed above, and may also include, at least in part, a well-known wireless data network, a well-known telephone network, or a well-known wired and wireless television network.
[0034] The device 100 can acquire surgical video consisting of multiple frames corresponding to multiple surgical stages via the hospital server 200 and / or database 300. However, this is only one example, and the device 100 can acquire surgical video captured via a camera connected to the device 100 wirelessly / wired.
[0035] The apparatus 100 can extract a plurality of visual kinematics-based indices based on the surgical video. The plurality of visual kinematics-based indices can include movement and interrelationship information of one or more surgical instruments included in the surgical video.
[0036] The apparatus 100 can acquire third feature data by fusing the first feature data for the surgical video and the second feature data for the plurality of visual kinematics-based indices, and can train the AI model 400 to recognize surgical stages based on the third feature data.
[0037] The operations related to this will be specifically described later with reference to the drawings.
[0038] The hospital server 200 (e.g., a cloud server) can capture and store surgical videos of patients, and can transmit the stored surgical videos to the device 100, the database 300, or the AI model 400.
[0039] The hospital server 200 can protect personal information of the parties involved in the surgical video by pseudonymizing or anonymizing the parties involved in the surgical video. In addition, the hospital server can encrypt and store information related to the age, sex, height, weight, and birth status of the patient involved in the surgical video, which is input by the user.
[0040] The database 300 may store various feature data generated by the device 100 and one or more parameters / instructions for utilizing the AI model 400. Although FIG. 1 illustrates a case where the database 300 is implemented outside the device 100, the database 300 may also be implemented as a component of the device 100.
[0041] The AI model 400 is an artificial intelligence model trained to recognize surgical stages from surgical images. The AI model 400 can be trained to recognize surgical stages from a dataset constructed of actual surgical images and associated feature data. The learning method can include, but is not limited to, supervised training / unsupervised training. The detection data output through the AI model 400 can be stored in the database 300 and / or the memory of the device 100.
[0042] FIG. 1 shows a case where the AI model 400 is implemented outside the device 100 (e.g., implemented on a cloud-based platform), but this is not limited to this and the AI model 400 can be implemented as a component of the device 100.
[0043] FIG. 2 is a block diagram illustrating the configuration of an apparatus 100 for visual multi-modality based surgical stage recognition according to one embodiment of the present disclosure.
[0044] 2, the device 100 may include a memory 110, a communication module 120, a display 130, an input module 140, and a processor 150. However, the device 100 is not limited thereto, and the software and hardware configuration of the device 100 may be modified / added / omitted within a range obvious to a person skilled in the art depending on the required operation.
[0045] The memory 110 can store data supporting various functions of the device 100 and at least one process or program for the operation of the processor 150, can store at least one process for recognizing surgical stages based on visual multi-modality according to the present disclosure, can store input / output data (e.g., a full surgical video consisting of multiple frames, one or more visual kinematics-based indexes, etc.), can store multiple application programs (or applications) run by the device, and data and commands for the operation of the device 100. At least some of these application programs can be downloaded from an external server via wireless communication.
[0046] Such memory 110 may include at least one type of storage medium among flash memory type, hard disk type, solid state disk type (SSD type), silicon disk drive type (SDD type), multimedia card micro type, card type memory (e.g., SD or XD memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk.
[0047] The memory 110 may also include a database that is separate from the device but connected via wire or wirelessly, i.e., the database shown in FIG. 1 may be embodied as a component of the memory 110.
[0048] The communication module 120 may include one or more components that enable communication with external devices, and may include, for example, at least one of a broadcast reception module, a wired communication module, a wireless communication module, a short-range communication module, and a location information module.
[0049] The wired communication module may include various wired communication modules such as a local area network (LAN) module, a wide area network (WAN) module, or a value added network (VAN) module, as well as various cable communication modules such as a universal serial bus (USB), a high definition multimedia interface (HDMI), a digital visual interface (DVI), a recommended standard 232 (RS-232), a power line communication module, or a plain old telephone service (POTS).
[0050] In addition to Wi-Fi modules and WiBro (Wireless Broadband) modules, the wireless communication module may further include wireless communication modules that support various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), and 4G, 5G, and 6G.
[0051] The display 130 displays (outputs) information (e.g., surgical images of a patient, surgical stage recognition information corresponding to specific frames constituting the surgical images, surgical skill scores, etc.) processed by the device 100. For example, the display can display execution screen information of an application program (e.g., an application) run by the device 100, or UI (User Interface) or GUI (Graphical User Interface) information based on such execution screen information.
[0052] The input module 140 is for receiving information input from a user, and when information is input via the user input unit, the processor 150 can control the operation of the device 100 in accordance with the input information.
[0053] The input module 140 may include hardware physical keys (e.g., buttons, dome switches, jog wheels, jog switches, etc., located on at least one of the front, back, and side of the device) and software touch keys. For example, the touch keys may be virtual keys, soft keys, or visual keys displayed on the touchscreen display 130 through software processing, or may be touch keys located on a portion other than the touchscreen. Meanwhile, the virtual keys or visual keys may have various forms and be displayed on the touchscreen, and may be, for example, graphics, text, icons, videos, or combinations thereof.
[0054] The processor 150 can control the overall operation and functions of the device 100. Specifically, the processor 150 can be implemented as a memory that stores data for an algorithm or a program that reproduces the algorithm for controlling the operation of components in the device 100, and at least one processor (not shown) that performs the above-mentioned operations using the data stored in the memory. In this case, the memory and the processor can be implemented as separate chips, or the memory and the processor can be implemented as a single chip.
[0055] In addition, the processor 150 can control any one or more of the above-mentioned components in combination to implement various embodiments of the present disclosure on the device 100, as described in Figures 3 to 7 below.
[0056] FIG. 3 is a flowchart illustrating a method for recognizing surgical stages based on visual multiple modalities performed by an apparatus according to one embodiment of the present disclosure.
[0057] The processor 150 of the device 100 can extract a plurality of visual kinematics-based indices based on a surgical video comprising a plurality of frames corresponding to a plurality of surgical stages (S310).
[0058] Here, the plurality of visual kinematics-based indexes may refer to information indicating movement and interrelationship information of one or more surgical instruments included in a surgical image.
[0059] Specifically, the processor 150 may input a surgical video consisting of multiple frames into a second AI model trained to perform a semantic segmentation algorithm to obtain semantic segmentation mask data, and may extract multiple visual kinematics-based indexes from the semantic segmentation mask data.
[0060] Here, a semantic segmentation algorithm refers to an algorithm that classifies all pixels in a video (or multiple frames / images that make up the video) into a predetermined number of classes. The semantic segmentation algorithm can segment / classify / identify one or more body organs and surgical instruments that are the target of surgery from the video (or multiple frames / images that make up the video), and mask the segmented / classified / identified pixel areas.
[0061] Thus, semantic segmentation mask data may refer to data that masks pixel regions classified as body organs and surgical instruments in a video (or multiple frames / images that make up the video).
[0062] The processor 150 may extract a plurality of visual kinematics-based indices from the semantic segmentation mask data corresponding to one or more surgical instruments included in the surgical video.
[0063] Specifically, the processor 150 can extract feature data associated with the motion of one or more surgical instruments using semantic segmentation mask data corresponding to the surgical instruments, and the apparatus can extract multiple visual kinematics-based indices using the extracted feature data associated with the motion of the surgical instruments.
[0064] 4, the processor 150 can acquire a plurality of frames (400-1, 400-2, ... 400-N) (N is a natural number equal to or greater than 1) showing a plurality of surgical steps that constitute a surgical video. Here, the surgical video can be composed of frames showing the entire surgical process, but is not limited to this.
[0065] For example, as shown in Fig. 5, a surgery may be divided into a plurality of steps (e.g., 20 steps), and processor 150 may acquire images captured for each step. The plurality of frames (400-1, 400-2, . . . 400-N) shown in Fig. 4 may represent frames constituting images captured for each step.
[0066] The processor 150 inputs a plurality of frames (400-1, 400-2, ..., 400-N) into a visual kinematics based index extractor 405 to extract a plurality of visual kinematics based indices (λ1, λ2, ..., λ N ) can be obtained. Here, the visual kinematics-based index extractor 405 can include a second AI model 410 trained to perform a semantic segmentation algorithm.
[0067] The processor 150 can input a plurality of frames (400-1, 400-2, ..., 400-N) into the second AI model 410 and obtain semantic segmentation data (420-1, 420-2, ..., 420-N) corresponding to one or more surgical instruments. The processor 150 can calculate a plurality of visual kinematics-based indices (λ1, λ2, ..., λ3) according to the semantic segmentation data (420-1, 420-2, ..., 420-N) corresponding to one or more surgical instruments. N ) can be obtained.
[0068] Visual kinematics-based indices can be categorized into types based on the motion of surgical instruments or the relationships between surgical instruments. Instrument motion can be measured by path length, speed, centroid, velocity, bounding box, and economy of area (EOA).
[0069] The measurement of the movement index (of a surgical instrument) can be embodied as in Equations 1 to 3.
[0070]
number
[0071]
number
[0072]
number
[0073] Here, PL represents the path length in the current time frame (t), and T represents the time range for computing the index. The path length can be composed of a cumulative path length and a partial path length.
[0074] D(x, t) can measure the difference in the x-axis between the previous and current time frames. x and y can represent the center of mass of the object in the frame. The center of mass represents the average position value in the x and y coordinates of the semantic segmentation mask. s is the velocity over the time range T, and v can represent the velocity in the X or Y direction over the time interval △. bw and bh are the width and height of the bounding box, respectively, and W and H are the width and height of the image, respectively. The bounding box can consist of four values: top, left, box width, box height (bx, by, bw, bh).
[0075] The processor 150 may acquire first feature data for the surgical image and second feature data for a plurality of visual kinematics-based indexes (S320).
[0076] Specifically, the processor 150 may input the surgical image and the plurality of visual kinematics-based indexes into a third AI model to obtain first and second feature data, where the third AI model may be configured based on at least one of a convolutional neural network (CNN) model and a long short term memory (LSTM) model.
[0077] The CNN model refers to the structure of a neural network model trained to perform convolutional operations, and the LSTM model refers to the structure of a neural network model designed to enable long-term memory, compensating for the shortcoming of the RNN (recurrent neural network) model, which is that it cannot remember information that is distant from the currently output data.
[0078] Referring to FIG. 4, the processor 150 receives a surgical image (i.e., a plurality of frames constituting the surgical image) (400-1, 400-2, . . . 400-N) and a plurality of visual kinematics-based indices (λ1, λ2, . . . , λ N ) can be input to the third AI model 430 to obtain the first feature data and the second feature data.
[0079] FIG. 4 shows a surgical video (i.e., a plurality of frames constituting the surgical video) (400-1, 400-2, . . . 400-N) and a plurality of visual kinematics-based indices (λ1, λ2, λ...). N ) represents the same case. However, this is only an example, and the surgical image (i.e., a plurality of frames constituting the surgical image) (400-1, 400-2, . . . 400-N) and a plurality of visual kinematics-based indexes (λ1, λ2, λ..., λ N ) can be input into different models.
[0080] For example, the first feature data may include feature data associated with a specific object (e.g., a bodily organ on which surgery is performed or a surgical instrument) in multiple frames constituting the surgical video, and the second feature data may include a movement pattern of the surgical instrument, etc.
[0081] In yet another embodiment of the present disclosure, a surgical skill score for a user of at least one surgical instrument can be calculated based on a motion path and motion pattern of the at least one surgical instrument associated with a plurality of visual kinematics-based indices. The device can utilize a trained module to calculate the surgical skill based on predefined surgical instrument paths and motion patterns. The surgical skill score can determine whether the user of the surgical instrument is a novice, an expert, or an expert.
[0082] The processor 150 may obtain fused third feature data by applying a fusion module trained to fuse data to the first feature data and the second feature data (S330).
[0083] 4, the processor 150 may apply a fusion module 440 to the first feature data and the second feature data to obtain third feature data. The processor 150 may concatenate each feature data and perform a convolution operation on the concatenated feature data (concatenated feature) to obtain third feature data 450 (convolution-based fusion module, Fused Feature).
[0084] 6(a), the processor 150 may concatenate the first feature data and the second feature data (Concatenated Feature). The processor 150 may apply a fusion module to the concatenated first feature data and the second feature data to obtain fused third feature data. Here, the fusion module may be configured based on a multi-layer perceptron (MLP).
[0085] 6(b), the fusion module may, under the control of the processor 150, apply a stop-gradient algorithm to the first feature data and the second feature data to obtain enhanced data for enhancing the interaction between the first feature data and the second feature data, and may, under the control of the processor 150, perform a convolution operation on the enhanced data to obtain third feature data.
[0086] To apply the stop-gradient algorithm to the first feature data and the second feature data, the device can obtain a contrastive loss using Equations 4 to 6. The processor 150 can use the contrastive loss to identify / learn the similarity between the feature data.
[0087]
number
[0088]
number
[0089]
number
[0090] Here, the following formulas (7) and (8) may represent the first feature data and the second feature data, respectively.
[0091]
number
[0092]
number
[0093] And, the following equation (9) having the same dimension and different views can be generated by a projector configured with MLP.
[0094]
number
[0095] a i and b iEach of these may represent different visual feature data, p represents the order of the normal vector (norm), and m1 and m2 may represent the index of the surgical image and visual kinematics basis, respectively.
[0096] The processor 150 may then perform a convolution operation on the enhancement data for enhancing the interaction between the first feature data and the second feature data to obtain third feature data.
[0097] The processor 150 can train the first AI model to recognize each of the plurality of surgical stages based on the third feature data (S340).
[0098] That is, when a specific frame of any surgical image is input, the first AI model can be trained by the device to output information regarding the surgical stage indicated by the specific frame (i.e., information for distinguishing the surgical stage).
[0099] 7, the processor 150 can input a surgical video consisting of frames showing seven surgical stages to the first AI model. When the first frame 610 and the second frame 620 of the surgical video are played / selected, the first AI model can be trained to output calot triangle dissection and gallbladder dissection as the surgical stages corresponding to each frame.
[0100] Meanwhile, the disclosed embodiments may be embodied in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, which, when executed by a processor, generates program modules to perform the operations of the disclosed embodiments. The recording medium may be embodied as a computer-readable recording medium.
[0101] Computer-readable recording media include all types of recording media that store computer-readable instructions, such as ROM (Read Only Memory), RAM (Random Access Memory), magnetic tape, magnetic disk, flash memory, and optical data storage devices.
[0102] The disclosed embodiments have been described above with reference to the accompanying drawings. Those skilled in the art will understand that the present disclosure may be embodied in forms different from the disclosed embodiments without changing the technical concept or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be construed as limiting.
Claims
1. a memory having stored therein at least one process for recognizing surgical stages based on visual multiple modalities; a processor that performs the process to recognize the surgical stage; Including, The processor: extracting a plurality of visual kinematics-based indices based on a surgical video comprising a plurality of frames corresponding to a plurality of surgical stages; obtaining first feature data for the surgical image; and obtaining second feature data for the plurality of visual kinematics-based indexes; applying a fusion module trained to fuse data to the first feature data and the second feature data to obtain fused third feature data; and training a first artificial intelligence (AI) model to recognize each of the plurality of surgical steps based on the third feature data.
2. When extracting the indexes of the plurality of visual kinematic bases, the processor: The surgical video consisting of the plurality of frames is input to a second AI model trained to perform a semantic segmentation algorithm to obtain semantic segmentation mask data; The apparatus of claim 1, wherein the plurality of visual kinematics-based indices are extracted from the semantic segmentation mask data corresponding to one or more surgical instruments included in the surgical video.
3. 3. The apparatus of claim 2, wherein the plurality of visual kinematics-based indices includes movement and interrelationship information of the one or more surgical instruments.
4. When the processor acquires the first feature data and the second feature data, inputting the surgical image and each of the plurality of visual kinematics-based indices into a third AI model to obtain the first feature data and the second feature data; 4. The apparatus of claim 3, wherein the third AI model includes at least one of a transformer, a convolutional neural network (CNN) model, and a long short term memory (LSTM) model.
5. When the processor acquires the third feature data, Concatenating the first feature data and the second feature data; applying the fusion module to the concatenated first feature data and the second feature data to obtain the third feature data; The apparatus of claim 1 , wherein the fusion module comprises a multi-layer perceptron-based fusion module.
6. The fusion module comprises: applying a stop-gradient algorithm to the first feature data and the second feature data to obtain enhancement data for enhancing interaction between the first feature data and the second feature data; The apparatus of claim 1 , wherein the third feature data is obtained by performing a convolution operation on the enhanced data.
7. The processor: The device of claim 1, further comprising: a surgical skill score for a user of at least one surgical instrument based on a path and pattern of movement of the at least one surgical instrument associated with the indexes of the plurality of visual kinematic bases.
8. the first artificial intelligence model trained based on the third feature data, The apparatus according to claim 1 , wherein, when a specific frame of another surgical video is input by the apparatus, information about a surgical stage indicated by the specific frame is output.
9. 1. A method for recognizing surgical stages based on visual multiple modalities performed by a device, comprising: extracting a plurality of visual kinematics-based indices based on a surgical video including a plurality of frames corresponding to a plurality of surgical steps; acquiring first feature data for the surgical image and second feature data for the plurality of visual kinematics-based indexes; obtaining fused third feature data by applying a fusion module trained to fuse data to the first feature data and the second feature data; training a first artificial intelligence (AI) model to recognize each of the plurality of surgical steps based on the third feature data; A method comprising:
10. The step of extracting the plurality of visual kinematics-based indices comprises: inputting the surgical video consisting of the plurality of frames into a second AI model trained to perform a semantic segmentation algorithm to obtain semantic segmentation mask data; extracting the plurality of visual kinematics-based indices from semantic segmentation mask data corresponding to one or more surgical instruments included in the surgical image; 10. The method of claim 9, comprising:
11. The method of claim 10, wherein the plurality of visual kinematics-based indices includes movement and interrelationship information of the one or more surgical instruments.
12. The step of acquiring the first feature data and the second feature data includes: inputting the surgical image and each of the plurality of visual kinematics-based indices into a third AI model to obtain the first feature data and the second feature data; Including, 12. The method of claim 11, wherein the third AI model includes at least one of a transformer, a convolutional neural network (CNN) model, and a long short term memory (LSTM) model.
13. The step of acquiring the third feature data includes: Concatenating the first feature data and the second feature data; applying the fusion module to the concatenated first feature data and the second feature data to obtain the third feature data; Including, The method of claim 9, wherein the fusion module comprises a multi-layer perceptron-based fusion module.
14. The fusion module comprises: applying a stop-gradient algorithm to the first feature data and the second feature data to obtain enhancement data for enhancing interaction between the first feature data and the second feature data; The method of claim 9, wherein the third feature data is obtained by performing a convolution operation on the enhanced data.
15. The method of claim 9, further comprising calculating a surgical skill score for a user of the at least one surgical instrument based on a movement path and movement pattern of the at least one surgical instrument associated with the plurality of visual kinematics-based indexes.
Citation Information
Patent Citations
Surgical operation training program and surgical operation training system
JP2016170362A
Surgical decision support using decision-theoretic models
JP2020537205A
Using model data to generate an enhanced depth map in a computer-assisted surgical system
US20220175473A1
Self-supervised representation learning using bootstrapped latent representations
WO2021245277A1
Surgery details evaluation system, surgery details evaluation method, and computer program
WO2022181714A1