Method and apparatus for tracking anatomical structures in video assisted surgery

Through SLAM and machine learning, the difficulty of anatomical structure tracking and identification in laparoscopic surgery is solved, real-time tracking and alerting of anatomical structures in video-assisted surgery is achieved, and surgical safety and efficiency are improved.

CN120344211APending Publication Date: 2025-07-18GENESIS MEDTECH INTERNATIONAL PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380087049.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-19
Filing Date
2023-11-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In laparoscopic and video-assisted surgery, prior art is difficult to track and identify anatomical structures in real time, especially when camera angles change, resulting in limited vision and potential risk of anatomical damage to the surgeon.

Method used

The camera trajectory is tracked using synchronous positioning and map construction (SLAM) technology, combined with machine learning models to mark areas of interest, and adaptively redraw markers in the video stream, providing visual and acoustic alerts to prevent surgical tools from contacting the anatomical structure.

Benefits of technology

Real-time tracking of anatomical structures at all camera angles is achieved, reducing surgeon mental fatigue, reducing the risk of anatomical structure damage, and improving the safety and efficiency of the surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120344211A_ABST
    Figure CN120344211A_ABST
Patent Text Reader

Abstract

The present disclosure describes an intra-operative anatomical structure tracking and alerting system for improving contextual awareness of a surgeon during laparoscopic and video-assisted procedures. During laparoscopic and video-assisted procedures, a surgeon may mark anatomical structures for systemic tracking (e.g., pain triangles in blood vessels, nerves, bile ducts, hernia repair, and the like) and then continue the procedure. Such indicia structures may become invisible as the mirror body moves. The system uses synchronous localization and mapping (SLAM) to track and restore the precise position of the mirror body. When the camera moves to an angle at which the previously marked structure is re-visible, the system will redraw the mark to indicate successful tracking of the structure. Further, if a surgical tool (e.g., an energy device) is close to the anatomical structure, an alert (e.g., a sound or on-screen visual alert) may be sent to the surgeon to alert attention.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to Related Applications

[0002] This patent application claims priority to U.S. Provisional Patent Application No. 63 / 387,984, filed on December 19, 2022, the entire content of which is incorporated herein by reference. Technical Field

[0003] This application relates to laparoscopic and video - assisted surgery, and more particularly to an intraoperative anatomical structure tracking and alert system for enhancing a surgeon's situational awareness during laparoscopic and video - assisted surgery. Background Art

[0004] During a surgical procedure, it is crucial for a surgeon to accurately identify key anatomical structures (such as safe and anatomical danger zones, major blood vessels, organs, nerves, etc.) at all camera angles. A major challenge in laparoscopic or video - assisted surgery is the limited field - of - view conditions for the surgeon due to the small viewing angle of the laparoscope or other types of cameras. Although a surgeon can successfully identify these structures at most camera angles, continuously tracking them at all angles can lead to significant mental fatigue and may sometimes be infeasible. Additionally, surgical instruments (such as scalpels and energy devices) may inadvertently contact these anatomical structures, causing patient injury. As used herein, an "energy device" may include a hand - held instrument that emits energy, such as high - frequency, ultrasonic, cryogenic, or thermal energy, beams of different spectra, and microwaves for closing blood vessels, cutting tissue, ablating abnormal tissue or tumors, and hemostasis.

[0005] Previously, there have been attempts to use machine learning / deep learning / computer vision to identify key structures in surgical scenarios, such as safe / hazardous areas during laparoscopic surgery, identification of blood vessels during surgery, etc. For example, see Madani, A. et al., "Artificial Intelligence for Intraoperative Guidance: Using Semantic Segmentation to Identify Surgical Anatomy during Laparoscopic Cholecystectomy". Ann Surg. August 1, 2022; 276(2):363-369. doi: 0.1097 / SLA.0000000000004594. These methods identify and highlight the target anatomical structures when they are visible. However, when the camera moves to an angle where these structures are no longer visible, the identification is lost. Additionally, when multiple cameras are used during surgery, the safe / hazardous areas visible on one screen are not easily transferred to other screens / video streams.

[0006] Therefore, there is a need for methods and devices that can clearly indicate safe / hazardous areas in all video streams in real time, even when the camera angles change during the surgical procedure. SUMMARY OF THE INVENTION

[0007] In one aspect, the present invention provides a method for tracking and positioning regions of interest (such as anatomical structures) from all camera angles. This will enable monitoring of potential or actual contact between surgical instruments and anatomical structures. The method includes obtaining a snapshot of the video stream of the surgical procedure. Then, using a suitable marking tool, the regions of interest are marked on the snapshot. Subsequently, when the camera moves, the viewing angle changes, or the camera focal length changes, the markings are adaptively redrawn at the current position of the regions of interest so that the surgeon is always aware of the position of the regions of interest. At any time during the surgery, if a surgical tool (such as a scalpel, energy device, etc.) approaches the markings, an alert is provided. The alert can be in the form of a visual alert, such as a bright flashing light, text appearing on the screen, etc., or a combination thereof. Alternatively, the alert can be in the form of a sound, such as an alarm sound. Additionally, the alert can vary according to the gap between the surgical tool and the region of interest. Thus, as the gap distance decreases, the color of the visual alert may change from an orange shade to a red shade, or the volume of the sound alert may increase. In this way, the surgeon can focus on performing the surgery without having to worry about potentially damaging any anatomical structures during the procedure.

[0008] In another aspect, the present invention provides a device for use during a surgical procedure. The device includes an imaging device (such as a laparoscope camera or an endoscope camera, etc.), which is configured to transmit a video stream. The device of the present invention further includes a display unit for displaying the video stream. The display unit is also capable of displaying a snapshot image of the acquired video stream. To this end, the display unit is configured to display a plurality of windows, such as a separate window for each individual camera source, a separate window for zooming in on the image, a separate window for the video stream snapshot, etc.

[0009] The device of the present invention further includes a marking tool that allows marking of regions of interest. The marking tool can be a keyboard (where marking is done by using arrow keys and the enter key or other predefined keys), a mouse, a screen that supports touch input (if the screen supports it), a stylus for a screen that supports stylus input, a handheld controller (such as a wireless mouse), etc. and combinations thereof.

[0010] Then, the device includes a memory as part of it, which includes a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD-HD), etc. and combinations thereof. The memory contains code instructions related to the method of the present invention described herein. The memory is also used to store a model for identifying image parts. For example, the model can be used to detect regions of interest, anatomical structures, surgical tools, etc. In some specific instances, the model is used to detect the tip of a surgical tool. The model is preferably based on a machine learning (ML) model, an artificial intelligence (AI) model, a computer vision (CV) model, etc. and combinations thereof. In a preferred embodiment, the model is generated and trained through a specific database annotated by a surgeon.

[0011] The device further includes a processor configured to execute the said code. The said code includes instructions for performing at least the following steps: receiving the video stream; acquiring a snapshot of a specific video stream; communicating with the marking tool; receiving an instruction from the user to mark a region of interest on the snapshot using the marking tool; using the model to identify the region of interest on subsequent video streams; rendering subsequent video streams by adaptively redrawing the mark at the location where the region of interest is present; identifying the location of the surgical tool; and estimating the gap between the tip of the surgical tool and the mark.

[0012] Then, the device includes an alarm tool that issues an alarm if the gap between the tip of the surgical tool and the marker is below the threshold estimated by the processor. The "threshold" used herein refers to the distance between the tip of the surgical tool and the marker. This distance can be measured at an appropriate position on the marker. For example, when the marker is a line, the distance to the closest point on the marker to the tip will be considered. Alternatively, the threshold can be the distance from a predetermined point on the marker to the tip. Depending on the nature of the surgical tool used, the threshold may vary. For example, when using a scalpel, the threshold may be 1 micron, while when using an energy device, the threshold may be 10 microns. In other cases, the threshold may mean that the tip of the surgical tool actually overlaps the marker. In some cases, the alarm tool is included in the display unit, and the alarm will appear in the form of text or an image, in a suitable color (such as a flashing red text box). In other cases, the alarm tool takes the form of a sound alarm, which may include a speaker or integrate a sound card into, for example, the display unit. Alarms that combine vision and sound are also within the scope of the present invention.

[0013] The method and device of the present invention provide significant advantages over the prior art because the identified region of interest is marked in one image, and subsequently, all video sources / streams will include the marker at the location where the region of interest exists. This is achieved using Simultaneous Localization and Mapping (SLAM), which maintains a memory of the previously seen regions. Thus, even in sources / streams / frames where the marker structure is not visible, the location of the region of interest is generally known. Subsequently, when the camera moves back to an angle or focal length where the marker structure is visible again, it will be marked and highlighted again to attract the attention of the surgeon.

[0014] In a preferred embodiment, the initial identification and marking of the region of interest are done by a trained person (such as a surgeon, healthcare professional, surgical assistant, etc.) because some anatomical features cannot be recognized by AI / ML / CV models. Thus, any potential errors that may arise from incorrect model recognition can be eliminated. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] These and other features, aspects, and advantages of the present invention will be better understood when the following detailed description is read in conjunction with the accompanying drawings, in which like characters represent like parts throughout the figures:

[0016] Figure 1 is a schematic diagram of the method steps based on an embodiment of the present invention.

[0017] Figure 2A is an exemplary display of a hernia repair surgery.

[0018] Figure 2B is another exemplary display of a laparoscopic cholecystectomy.

[0019] Figure 3 Shows the tip tracking of a surgical tool based on an embodiment of the present invention (the green cross represents the neural network output of the detected monopolar curved scissors tip).

[0020] Figure 4 Is a block diagram representation of a device based on an embodiment of the present invention. Detailed implementation

[0021] The definitions provided herein are intended to facilitate the understanding of certain terms frequently used herein and are not intended to limit the scope of the present disclosure.

[0022] As used in this specification and the appended claims, the singular forms "a", "an" and "the" include embodiments having a plurality of referents unless the context clearly dictates otherwise.

[0023] Unless otherwise noted, all numbers expressing feature sizes, amounts and physical properties used in this specification and the claims are to be understood as being modified in all instances by the term "about". Accordingly, unless indicated to the contrary, the numerical parameters set forth in the foregoing specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by those skilled in the art utilizing the teachings disclosed herein.

[0024] As used in this specification and the appended claims, the term "or" is generally used in its inclusive sense of "and / or" unless the context clearly dictates otherwise.

[0025] Reference Figure 1 , which is a schematic diagram of the method steps based on an embodiment of the present invention, is generally denoted by the numeral 100. The method includes obtaining at least one video stream 102 from a suitable imaging device (such as an endoscope or laparoscope). Then, the method involves obtaining a snapshot of the video stream in step 104. The manner of obtaining a snapshot of the video stream may be well known in the art and may include, for example but not limited to: pressing a key on a keyboard, clicking via a handheld controller, clicking on a specific part of a touchpad or screen, gestures, voice commands, etc. and combinations thereof.

[0026] Then, the method includes marking the region of interest in step 106. As described herein, the marking of the region of interest can be done manually and / or with the assistance of an AI or ML-based model. The marking is preferably performed by a human, such as a surgeon, surgical assistant, nurse, other healthcare practitioner, etc. This is because not all anatomical structures can be well recognized by the model. The marking can be achieved by suitable means, such as mouse clicks on a certain area of the screen representing the region of interest, using a stylus, touch, gestures, etc. and combinations thereof. As an illustration of this step, when the anatomical structure of interest is visible on the laparoscopic monitor, the surgeon keeps the laparoscope stationary and takes a snapshot to obtain a still image. Then, the touchpad or mouse is used to mark the anatomical structure of interest on such an image. The surgeon can pause the surgery to mark the anatomical structure, or an assistant or nurse can draw such marks. Lines or filled shapes can be used to mark blood vessels or nerves, and polygons can be used to mark regions, such as the pain triangle in a hernia repair surgery. The region of interest is used to identify the critical space around the anatomical structure, which is visible to the surgeon during the surgery but should not be contacted by the surgical tools, or the space is narrow and the entry of the surgical tools into this space may cause the risk of accidental contact with the anatomical structure. Therefore, it is advantageous for the surgeon to mark the region of interest in advance so that the region of interest can be avoided during the surgery and the surgery is only performed at the surgical site. It must be noted that the region of interest can include multiple positions around the surgical site.

[0027] Then, the method of the present invention involves identifying the region of interest in all video streams in step 108. The method of the present invention advantageously uses Simultaneous Localization and Mapping (sometimes abbreviated as SLAM in the industry, such as OrbSLAM-2) to track the camera trajectory and construct a map of the surgical scene as the laparoscope moves around the surgical site. Specifically, SLAM is a method of simultaneously constructing a map and localizing the camera in that map, allowing the camera to map an unknown environment. As the camera moves, the 3D structure of the surgical scene is gradually constructed. Using this intra-operative constructed 3D model of the surgical scene, each point in the current laparoscopic frame can be mapped to a 3D point on the 3D model; on the other hand, each 3D point on the 3D model can be projected into the 2D space of the current frame, and if the projected position of the 3D point is outside the frame boundary, it means that the 3D point is not currently visible. To re-identify the anatomical structures (regions of interest) marked by the surgeon, the 3D points of the marked structures are projected back into the current frame to visualize these points within the current frame boundary. Depending on the camera angle, there can be 3 cases:

[0028] 1. The entire region of interest is visible in the current frame and is highlighted;

[0029] 2. Only a partial region of interest is visible, so only this visible part is highlighted in the current frame; additional virtual reality frames can be constructed to overlay and render the invisible part on the current frame;

[0030] 3. The entire region of interest is outside the current frame boundary and is not highlighted in the current frame; additional virtual reality frames can be constructed to overlay and render the invisible part on the current frame.

[0031] In this way, in step 110, the markers are reproduced in the video stream that is always visible for all regions of interest. Thus, the surgeon is always aware of the location of the anatomical structure during the operation.

[0032] Figure 2A is an exemplary display of a hernia repair operation using the method of the present invention, where line 202 divides the safe area 204 and the unsafe area 206. In addition, visual aids such as graphical indicators or emojis (such as smiling faces and crying faces) can be used to indicate the safe and unsafe areas, so that the surgeon can more easily identify them during the operation.

[0033] Figure 2B is another exemplary display of a laparoscopic cholecystectomy performed using the method of an embodiment of the present invention, showing more subtle changes within the safe and unsafe areas divided by line 208. The safe area is further marked with triangles (exemplary polygons) to indicate the very safe area D 210 and the relatively safe area I 212. An appropriate color selection can also be applied for the surgeon to quickly identify. Similarly, the unsafe area is further marked with a first triangle labeled "pain" 214 (possibly using a certain color to represent a mild danger area), a second triangle labeled "destruction" 216 (appropriately colored, such as red for severe danger), and a third triangle labeled F 218.

[0034] Then, in step 112 of the method based on an embodiment of the present invention, the position of the tip of the surgical tool is identified. This can be the tip of a scalpel or an energy device, as an example of a surgical tool. This is preferably achieved using a trained AI / ML / CV model to ensure speed and accuracy. In some cases, a neural network is trained to detect various surgical tools, such as scalpels or energy devices.

[0035] Figure 3 Shows an exemplary surgical tool tip 302.

[0036] Then, in step 114, the gap between the tip of the surgical tool and the marker is estimated. If the gap is below a threshold, or there is an overlap between the position of the tip of the surgical tool and the marker, it indicates a potential contact between the tip of the surgical tool and the anatomical structure. This may be a cause for concern as it can lead to further complications, accidental injuries, etc. An alert will be sent to the surgeon, which can be in the form of a visual alert on the screen (such as a colored flashing light or a pop-up text), an audible alert, or a combination thereof.

[0037] Now refer to Figure 4 , which is a block diagram representation of a device based on an embodiment of the present invention, generally denoted by the numeral 400. The device 400 includes at least one imaging device 402 configured to capture a video including a plurality of frames. The video can be of any part inside the body, including, for example, the gastrointestinal tract or the urinary tract, etc. The device 400 also includes a medical image processing system 404, which includes a memory 406, a display unit 408, a marking tool 410, a processor 412, as well as a communication interface and software for performing the functions specified by the present invention.

[0038] The processor 412 used herein may include one or more processors, controllers, control modules, or other processing devices. The processor 412 can be implemented using a general-purpose or a dedicated processing engine, such as a microprocessor, a controller, a graphics processing unit (GPU), a field-programmable gate array (FPGA), or other control logic. The processor 412 is connected to a bus or any other communication medium to facilitate interaction with other components shown herein or external communication.

[0039] The memory 406 includes non-volatile memory (such as one or more hard disk drives) and volatile memory (such as random access memory (RAM)). Other memory modules may also be included in the memory 406. For example, preferably, random access memory (RAM) or other dynamic memory can be used to store information and instructions to be executed by the processor 412. The hard disk drive, solid-state drive, or other main memory can also be used to store temporary variables or other intermediate information during the execution of instructions by the processor 412. In addition, read-only memory (“ROM”) or other static storage devices can be used to store static information and instructions of the processor 412.

[0040] In addition, the memory 406 may also include one or more various forms of information storage mechanisms, which may include, for example, media drives and storage unit interfaces. The media drive may include a drive or other mechanism to support fixed or removable storage media. For example, a hard disk drive, a floppy disk drive, a tape drive, an optical disk drive, a CD or DVD drive (R or RW), or other removable or fixed media drives may be provided. Thus, the storage media may include, for example, hard disks, floppy disks, tapes, cartridges, optical disks, CDs or DVDs, or other removable or fixed media that can be read, written to, or accessed by the media drive. As shown in these examples, the storage media may include computer-usable storage media having computer software or data stored thereon.

[0041] In an alternative embodiment, the memory 406 may include other similar tools for allowing a computer program or other instructions or data to be loaded into a suitable computing module. Such tools may include, for example, fixed or removable storage units and interfaces. Examples of such storage units and interfaces may include program cartridges and cartridge interfaces, removable memories (such as flash memory or other removable memory modules) and memory slots, PCMCIA slots and cards, and other fixed or removable storage units and interfaces that allow software and data to be transferred from the storage unit to the computing module.

[0042] The device is deployed physically in the operating room and connected to the console of the surgical system. The device is configured to perform real-time inference of AI / ML / CV models. The device will also record the tracking, SLAM, and detection results generated during the surgery and save such information as a system log to disk. The device can be a stand-alone device that contains a tracking and alert system provided within the physical device envisioned in one embodiment of the present invention. Alternatively, the device can be provided as software and integrated into the existing infrastructure in a suitable environment such as a hospital.

[0043] In cases where components or modules of a technology are implemented, in whole or in part, using software, in one embodiment, these software elements can be implemented to operate with a computing or processing module capable of performing the functionality associated with them. After reading this specification, it will be apparent to those skilled in the relevant art how to implement the technology using other computing modules or architectures. A computing module can represent, for example, computing or processing capabilities found in desktop computers, laptop computers, and notebook computers; handheld computing devices (PDAs, smartphones, cell phones, palmtop computers, etc.); mainframes, supercomputers, workstations, or servers; or any other type of special-purpose or general-purpose computing device (such as desired or appropriate for a given application or environment). A computing module can also represent computing capabilities embedded within or otherwise available to a given device. For example, computing modules can be found in other electronic devices, such as digital cameras, navigation systems, cellular phones, portable computing devices, modems, routers, WAPs, terminals, and other electronic devices that may include some form of processing capability.

[0044] The computing module can also include a communication interface for allowing software and data to be transferred between the computing module and external devices. Examples of communication interfaces include modems, network interfaces (such as Ethernet, network interface cards, WiMedia, IEEE802.XX, or other interfaces), communication ports (such as USB ports, IR ports, RS232 ports, Bluetooth ® interfaces, or other ports), or other such communication interfaces. The software and data transferred through the communication interface are typically carried on signals, which can be electrical, electromagnetic (including optical), or other signals that can be exchanged by a given communication interface. These signals can be provided to the communication interface through wired or wireless communication media. Some examples of channels can include telephone lines, cellular links, RF links, optical links, network interfaces, local or wide area networks, and other wired or wireless communication channels.

[0045] The computing module also includes input / output (I / O) devices (including but not limited to keyboards, displays, pointing devices, etc.), which can be coupled to the device directly or through an intermediate I / O controller.

[0046] The terms "computer program medium" and "computer-usable medium" generally refer to media such as memories, storage units, and media. These and various other forms of computer program media or computer-usable media can participate in carrying one or more sequences of one or more instructions to a processing device for execution. Such instructions embodied on a medium are commonly referred to as "computer program code" or "computer program product" (which may be grouped in the form of computer programs or other groupings). When executed, such instructions can cause a computing module to perform the features or functions of the disclosed technology discussed herein. The computer program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the C programming language or similar programming languages.

[0047] Although various embodiments of the disclosed technology have been described above, it should be understood that they are presented by way of example only and not by way of limitation. Similarly, the various figures may depict exemplary architectures or other configurations of the disclosed technology, and doing so is to assist in understanding the features and functions that may be included in the disclosed technology. The disclosed technology is not limited to the exemplary architectures or configurations shown, but rather various alternative architectures and configurations can be used to achieve the desired features. In fact, it will be apparent to those skilled in the art how to implement alternative functional, logical, or physical partitioning and configurations to achieve the desired features of the technology disclosed herein. Additionally, a large number of different component module names other than those described herein can also be applied to the various partitions. Further, with respect to flowcharts, operation descriptions, and method claims, the order of steps presented herein should not require that various embodiments perform the functions in the same order, unless the context otherwise requires.

[0048] Although the above-disclosed technology is described in terms of various exemplary embodiments and implementations, it should be understood that the various features, aspects, and functions described in one or more of the individual embodiments are not limited in their applicability to the particular embodiments in which they are described, but rather can be applied singly or in various combinations to one or more other embodiments of the disclosed technology, whether or not such embodiments are described and whether or not such features are presented as part of the described embodiments. Accordingly, the breadth and scope of the technology disclosed herein should not be limited by any of the above exemplary embodiments.

[0049] The terms and phrases used in this document, and variations thereof, unless otherwise expressly stated, shall be construed as open-ended and not limiting. By way of example of the foregoing: the term "including" shall be understood to mean "including but not limited to" or similar; the term "example" is used to provide exemplary instances of the item being discussed, rather than an exhaustive or limiting list thereof; the term "a" or "an" shall be understood to mean "at least one", "one or more" or similar; adjectives such as "conventional", "traditional", "normal", "standard", "known" and terms with similar meanings shall not be construed as limiting the item described to items available during a given time period or at a given time, but shall be understood to encompass conventional, traditional, normal or standard techniques available or known at any time, present or future. Similarly, when this document refers to techniques that are obvious or known to a person of ordinary skill in the art, such techniques encompass techniques that are obvious or known to a skilled person at any time, present or future.

[0050] There are broad words and phrases such as "one or more", "at least", "but not limited to" or other similar phrases which in some instances should not be construed as meaning that the narrower case is required or intended in the absence of such broad phrases. The use of the term "module" does not mean that the components or functions described as or claimed to be part of a module are all configured in a common package. In fact, any or all of the various components of a module, whether control logic or otherwise, may be combined in a single package or maintained separately and may further be distributed in multiple groupings or packages or across multiple locations.

[0051] In addition, the various embodiments set forth herein are described in terms of exemplary block diagrams, flowcharts and other illustrations. Those of ordinary skill in the art will appreciate, upon reading this document, that the illustrated embodiments and their various alternatives can be implemented without being limited to the illustrated examples. For example, the block diagrams and their accompanying description should not be construed as mandating a particular architecture or configuration.

Claims

1. A method, comprising: Obtaining at least one video stream of a surgical procedure; Obtaining a snapshot of the video stream; Marking a region of interest on the snapshot; Identifying the presence or absence of the region of interest in subsequent video streams; And Drawing a mark at any position where the region of interest exists in the subsequent video stream to inform the surgeon of the position of the region of interest.

2. The method according to claim 1, wherein the marking is performed by a human.

3. The method according to claim 1, further comprising identifying the position of the tip of a surgical tool in subsequent video streams.

4. The method according to claim 3, further comprising providing an alert when the gap between the tip of the surgical tool and the mark is below a threshold.

5. The method according to claim 3, wherein the position of the tip of the surgical tool is identified by an artificial intelligence method.

6. The method according to claim 1, wherein the mark is a line or a filled shape.

7. The method according to claim 4, wherein the alert is at least one of a visual alert, an audible alert, and a combination thereof.

8. The method according to claim 1, wherein the video stream is used to construct a three-dimensional model using a simultaneous localization and mapping method.

9. The method according to claim 1, wherein the region of interest includes at least one anatomical structure or the space surrounding the anatomical structure.

10. An apparatus for use during a surgical procedure, the apparatus comprising: An imaging device configured to transmit a video stream; A display unit for displaying the video stream; A marking tool; A memory in which a model is stored; And A processor configured to: Receive the video stream; Obtain a snapshot of a specific video stream; Communicate with the marking tool; Receive an instruction from a user to draw a mark at the region of interest on the snapshot; Use the model to identify the region of interest on subsequent video streams by drawing the mark; And Relabel the region of interest on the subsequent video streams.

11. The apparatus according to claim 10, wherein the subsequent video stream is rendered using a three-dimensional model employing a simultaneous localization and mapping method.

12. The apparatus according to claim 10, wherein the model is configured to identify the tip of a surgical tool.

13. The apparatus according to claim 12, wherein the processor is further configured to identify the gap between the tip of the surgical tool and the mark at the region of interest.

14. The apparatus according to claim 13, wherein the processor is configured to provide an alert based on the gap being less than a predefined threshold.

15. The apparatus according to claim 14, wherein the alert is at least one of a visual alert, an audible alert, and a combination thereof.