Systems and methods for image mapping and fusion during surgical procedures
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-12
- Publication Date
- 2026-08-11
Smart Images

Figure CN115484858B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to apparatus, systems, and methods for mapping and fusing images from multiple sources, and more specifically to the fusion of images from multiple sources during a surgical procedure. Background Technology
[0002] Robotic surgical systems, such as remote surgical systems, are used to perform minimally invasive surgical procedures, offering many benefits compared to traditional open surgical techniques, including reduced pain, shorter hospital stays, faster return to normal activities, minimized scarring, reduced recovery time, and less damage to tissues.
[0003] A robotic surgical system may have multiple robotic arms that move attached instruments or tools, such as image-capturing devices, suture devices, electrosurgical instruments, etc., in response to a surgeon viewing images captured by an image-capturing device at the surgical site. During the robotic surgical procedure, each tool is inserted into the patient's body through an opening (natural opening or incision) and positioned to manipulate tissue at the surgical site. The openings are arranged around the patient's body, allowing surgical instruments to be used collaboratively to perform the robotic surgical procedure, and the image-capturing device to visualize the surgical site.
[0004] During robotic surgery, accurately understanding and controlling the position of tools within the surgical site is crucial. Therefore, there is a ongoing need for systems and methods for mapping and fusing images of the surgical site from multiple image sources during robotic surgery. Summary of the Invention
[0005] This disclosure relates to apparatus, systems, and methods for fusing images from multiple sources during a surgical procedure. According to various aspects of this disclosure, a system for object recognition in endoscopic images is proposed. The system includes a light source, a first imaging device, a second imaging device, and an imaging device control unit. The light source is configured to provide light within the surgical site.
[0006] In one aspect of this disclosure, a system for mapping and fusing endoscopic images includes: a light source configured to provide light within a surgical site; a first imaging device configured to acquire an image from the surgical site; a second imaging device configured to acquire an image from the surgical site; a display; and an imaging device control unit configured to control the first imaging device and the second imaging device. The light source is configured to generate a first light including an infrared (IR) band and a second light configured to generate a visible band. The control unit includes a processor and a memory storing instructions. When executed by the processor, the instructions cause the system to: capture a first image of an object within a surgical site using the first imaging device, the image including the first light radiating from the object and a first reference point; capture a second image of the object using the second imaging device, the image including the second light radiating from the object and a second reference point; compare a first position of the first reference point in the first image with a second position of the second reference point in the second image; determine the relative orientation of the first imaging device and the second imaging device based on the comparison; generate an enhanced image that fuses the first image and the second image based on the determined relative orientation; and display the enhanced image on the display.
[0007] In one aspect of this disclosure, the light source can be configured to generate a first light including an infrared (IR) band and a second light configured to generate a visible band.
[0008] In another aspect of this disclosure, the first reference point may include structured light, and the second reference point may include structured light.
[0009] In another aspect of this disclosure, generating the enhanced image may further include determining the viewpoint of the virtual imaging device. The generation of the enhanced image may also be based on the viewpoint of the virtual imaging device.
[0010] In another aspect of this disclosure, generating the enhanced view may further include: determining a first optical path distortion of the first imaging device and a second optical path distortion of the second imaging device; and processing the first image based on the first optical path distortion to match the second optical path distortion.
[0011] In another aspect of this disclosure, when executed, the instruction may also enable the system to perform tracking of the object based on the first reference point and the second reference point.
[0012] In another aspect of this disclosure, the first reference point and the second reference point may include an identifier, texture, dot pattern, and / or a unique identifier.
[0013] In another aspect of this disclosure, the first image may include first distance information for each pixel of the first image. The second image may include second distance information for each pixel of the second image. The relative orientation of the first imaging device may also be based on the first distance information and the second distance information.
[0014] In another aspect of this disclosure, the generation of the enhanced image also includes a portion of the first image that includes the object to represent the object in the enhanced image, and the remainder of the enhanced image includes a fusion of the first image and the second image.
[0015] According to various aspects of this disclosure, a method for mapping and fusing endoscopic images is proposed. The method includes capturing a first image of an object within a surgical site by a first imaging device; and capturing a second image of the object by a second imaging device. The first image includes first light radiated from the object and a first reference point. The second image includes second light radiated from the object and a second reference point. The method further includes: comparing a first position of the first reference point in the first image with a second position of the second reference point in the second image; determining a relative orientation between the first imaging device and the second imaging device based on the comparison; generating an enhanced image fusing the first image and the second image based on the determined relative orientation; and displaying the enhanced image on a display.
[0016] In another aspect of this disclosure, the first light may include an infrared (IR) band, and the second light includes a visible band.
[0017] In another aspect of this disclosure, the first reference point may include structured light, and the second reference point may include structured light.
[0018] In another aspect of this disclosure, generating the enhanced image may further include determining the viewpoint of the virtual imaging device. The generation of the enhanced image may also be based on the viewpoint of the virtual imaging device.
[0019] In another aspect of this disclosure, generating the enhanced view may further include: determining a first optical path distortion of the first imaging device and a second optical path distortion of the second imaging device; and processing the first image based on the first optical path distortion to match the second optical path distortion.
[0020] In another aspect of this disclosure, the method may further include performing tracking of the object based on the first reference point and the second reference point.
[0021] In one aspect of this disclosure, the first reference point and the second reference point may include an identifier, texture, dot pattern, and / or a unique identifier.
[0022] According to various aspects of this disclosure, the first image may include first distance information for each pixel of the first image.
[0023] The second image may include second distance information for each pixel of the second image. The relative orientation of the first imaging device may also be based on the first distance information and the second distance information.
[0024] In another aspect of this disclosure, generating the enhanced image may further include using only the portion of the first image that includes the object to represent the object, and the remaining portion of the enhanced image includes a fusion of the first image and the second image.
[0025] In another aspect of this disclosure, the first imaging device and the second imaging device may include a stereoscopic imaging device.
[0026] According to various aspects of this disclosure, a non-transitory storage medium is proposed that stores a program that causes a computer to execute a computer-implemented method for mapping and fusing endoscopic images. The method includes capturing a first image of an object within a surgical site by a first imaging device; and capturing a second image of the object by a second imaging device. The first image includes first light radiated from the object and a first reference point. The second image includes second light radiated from the object and a second reference point. The method further includes: comparing a first position of the first reference point in the first image with a second position of the second reference point in the second image; determining a relative orientation between the first imaging device and the second imaging device based on the comparison; generating an enhanced image fusing the first image and the second image based on the determined relative orientation; and displaying the enhanced image on a display.
[0027] According to various aspects of this disclosure, a system for mapping and constructing a 3D model includes: a display; a light source configured to provide light within a surgical site; a first imaging device configured to acquire an image from the surgical site; a second imaging device configured to acquire an image from the surgical site; and an imaging device control unit configured to control the first imaging device and the second imaging device. The control unit includes a processor and a memory storing instructions that, when executed by the processor, cause the system to: capture an image of an object within the surgical site by the first imaging device; and capture a second image of the object by the second imaging device; segment the first image to extract a known reference object; determine a first relative position of the first imaging device based on the extracted known reference object; segment the second image to extract the known reference object; determine a second relative position of the second imaging device based on the extracted known reference object; and construct a 3D model of the surgical site based on the determined first and second relative positions.
[0028] In one aspect of this disclosure, the light source can be configured to generate infrared (IR) bands and / or visible bands.
[0029] In another aspect of this disclosure, the first imaging device may include a first view of the surgical site, and the second imaging device may include a second view of the surgical site, which is different from the first view.
[0030] In another aspect of this disclosure, the system may also include a second light source that emits structured light onto the object within the surgical site.
[0031] In another aspect of this disclosure, when executed, the instructions may also cause the processor, when constructing the 3D model, to: determine the location of the virtual imaging device; generate a virtual viewpoint of the 3D model based on the location of the virtual imaging device; and display the virtual viewpoint of the 3D model on a display.
[0032] In another aspect of this disclosure, the virtual perspective may include stereoscopic images.
[0033] In another aspect of this disclosure, the first imaging device may include a 2D imaging device, and the second imaging device may include a stereoscopic imaging device.
[0034] According to various aspects of this disclosure, a method for mapping and constructing a 3D model includes: capturing an image of an object within a surgical site by a first imaging device; capturing a second image of the object by a second imaging device; segmenting the first image to extract a known reference object; determining a first relative position of the first imaging device based on the extracted known reference object; segmenting the second image to extract the known reference object; determining a second relative position of the second imaging device based on the extracted known reference object; and constructing a 3D model of the surgical site based on the determined first relative position and the determined second relative position.
[0035] In one aspect of this disclosure, the method may further include generating at least one of an infrared (IR) band or a visible band from a first light source.
[0036] In another aspect of this disclosure, the first imaging device may include a first view of the surgical site, and the second imaging device may include a second view of the surgical site, which is different from the first view.
[0037] In another aspect of this disclosure, the method may further include emitting structured light onto the object within the surgical site.
[0038] In another aspect of this disclosure, constructing the 3D model may further include: determining the position of a virtual imaging device; generating a virtual viewpoint of the 3D model based on the position of the virtual imaging device; and displaying the virtual viewpoint of the 3D model on a display.
[0039] In another aspect of this disclosure, the first imaging device may include a 2D imaging device, and the second imaging device may include a stereoscopic imaging device.
[0040] In another aspect of this disclosure, the virtual perspective may include stereoscopic images.
[0041] Further details and aspects of various embodiments of this disclosure are described in more detail below with reference to the accompanying drawings. Attached Figure Description
[0042] This document describes embodiments of the present disclosure in conjunction with the accompanying drawings, wherein:
[0043] Figure 1 These are schematic diagrams of the user interface and robot system based on this disclosure;
[0044] Figure 2 yes Figure 1 A perspective view of the linkages of a robot system;
[0045] Figure 3 It is a schematic diagram of the surgical site. Figure 1 The tools of the robotic system are inserted into it;
[0046] Figure 4 This is an illustrative configuration of a visualization or imaging system according to an embodiment of this disclosure; and
[0047] Figure 5 This is a flowchart of a method for mapping and fusing endoscopic images according to an exemplary embodiment of the present disclosure;
[0048] Further details and aspects of exemplary embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. Any of the above aspects and embodiments of this disclosure may be combined without departing from the scope of this disclosure. Detailed Implementation
[0049] The embodiments of the currently disclosed devices, systems, and treatment methods are described in detail with reference to the accompanying drawings, wherein in each of the several views, the same reference numerals refer to the same or corresponding elements. As used herein, the term "distal" refers to the portion of the structure farther from the user, while the term "proximal" refers to the portion of the structure closer to the user. The term "clinician" refers to a physician, nurse, or other care provider and may include support personnel.
[0050] This disclosure is applicable to situations involving the capture of images of surgical sites. Endoscopic systems are provided as examples, but it will be understood that such descriptions are exemplary and do not limit the scope of this disclosure or its applicability to other systems and procedures.
[0051] A straightforward method for visualizing surgical sites in minimally invasive surgery is the use of white light endoscopy, where stereoscopic imaging is particularly desirable in robotic surgery. Near-infrared fluorescence can be used to view function-oriented images, such as indocyanine green dye showing blood perfusion. Unfortunately, these two imaging methods are often present in two separate endoscopes, resulting in images being displayed independently, thus impairing their practicality during surgery. It is desirable to allow these two distinct imaging methods to be displayed together in a single fused display, providing a consistent representation of information about the surgical site.
[0052] Simultaneous structured light projection, whether temporally or geometrically unique, allows common points to be seen simultaneously in two or more endoscopic views. These common spatial points across individual endoscopic views allow 3D dimensional surfaces seen by each endoscope to be rendered in a common coordinate system. With multiple endoscopes knowing these 3D surfaces, a fused dataset can be created and displayed from the most clinically appropriate view.
[0053] A shared 3D data environment can be created by using one or more techniques to obtain the distance between each observed pixel in an image and the image sensor of a camera in a manner known across multiple cameras, allowing common points to be observed across multiple cameras. A means is needed to correlate RGBD images (the RGB values of the color (e.g., red, green, blue) at each pixel and the distance D of the pixel from the camera) to obtain common points. This can be accomplished using structured light projection, which can be observed simultaneously by multiple cameras. This light projection can be performed, for example, in a sequence of points each uniquely displayed, or it can be performed on many points with point sizes and / or shapes used to distinguish them from each other.
[0054] Surfaces observed by multiple cameras are expected to deform as a result of natural processes and surgical manipulations of the site. Therefore, the common points observed by multiple cameras should be continuously updated in a real-time manner, as the surface may deform over a period of time during which a series of points are observed. Thus, either sequential projection needs to be faster than some anticipated deformation rate threshold, or a multiple simultaneous projection point method can be used.
[0055] Because a sufficient number of common 3D points across the observable surgical site can be measured across all cameras, the relative pose of the cameras with respect to each other can be calculated and used to allow image data to be projected onto the normally observed 3D surface. Individual cameras may need to be calibrated to account for optical path distortion, as well as the imager pose relative to that path.
[0056] With all these measurements and calibrations in place, a common data fusion representation of the surgical site is generated. This allows for the generation of a projected view from the desired virtual endoscopic position to provide the surgeon with her desired perspective. This viewing position can be dynamically updated relative to the time-varying data fusion representation, or the data fusion representation can be fixed in time or even recorded and played back, while the camera perspective is manipulated to provide additional insights into the surgical site. It should also be noted that multiple observers can simultaneously view the surgical site from their own chosen perspectives.
[0057] refer to Figure 1 The robotic surgical system 1 according to this disclosure is generally shown as a robotic system 10, a processing unit 30, and a user interface 40. The robotic system 10 typically includes links or arms 12 and a robot base 18. Arms 12 movably support a tool 20 having an end effector 22 configured to act on tissue. Each arm 12 has an end portion 14 supporting the tool 20. Additionally, the end portions 14 of the arms 12 may include imaging devices 16, 16B for imaging the surgical site “S”. The user interface 40 communicates with the robot base 18 via the processing unit 30.
[0058] User interface 40 includes a display device 44 configured to display a three-dimensional image. The display device 44 displays a three-dimensional image of the surgical site “S,” which may include data captured by imaging devices 16, 16B positioned at the distal end 14 of arm 12 and / or data captured by imaging devices positioned around the operating room (e.g., imaging devices positioned within the surgical site “S,” imaging devices positioned adjacent to the patient, imaging devices 56 positioned at the distal end of the imaging link or arm 52). The imaging devices (e.g., imaging devices 16, 16B, 56) may capture visual images, infrared images, ultrasound images, X-ray images, thermal images, and / or any other known real-time images of the surgical site “S.” The imaging devices transmit the captured imaging data to a processing unit 30, which creates a three-dimensional image of the surgical site “S” in real time based on the imaging data and transmits the three-dimensional image to the display device 44 for display. The imaging device 56 is envisioned to be an optical cannula or similar device capable of capturing 2D / 3D images in the visible spectrum, infrared spectrum, or any other spectrum of light, and capable of applying filtering and processing to enhance the captured images / videos.
[0059] The user interface 40 also includes input handles 42 supported on the control arm 43, which allow clinicians to manipulate the robotic system 10 (e.g., the mobile arm 12, the end effector 14 of the arm 12, and / or the tool 20). Each input handle in the input handles 42 communicates with the processing unit 30 to transmit control signals to and receive feedback signals from it. Alternatively or additionally, each input handle in the input handles 42 may include an input device (not shown) that allows the surgeon to manipulate (e.g., clamp, grasp, start, open, close, rotate, advance, slice, etc.) the end effector 22 of the tool 20 supported at the end effector 14 of the arm 12.
[0060] Each input handle in the input handle 42 is movable through a predefined workspace to move the end 14 of the arm 12 into the surgical site “S”. A three-dimensional image on the display device 44 is oriented such that movement of the input handle 42 moves the end 14 of the arm 12, as viewed on the display device 44. It should be understood that the orientation of the three-dimensional image on the display device may be mirrored or rotated relative to a view viewed from above the patient. Additionally, it should be understood that the size of the three-dimensional image on the display device 44 may be scaled to be larger or smaller than the actual structure of the surgical site “S”, thereby allowing clinicians to better visualize the structure within the surgical site “S”. As the input handle 42 moves, the tool 20, and therefore the end effector 22, moves within the surgical site “S”, as detailed below. As detailed herein, movement of the tool 20 may also include movement of the end 14 of the arm 12 supporting the tool 20.
[0061] For a detailed description of the construction and operation of the robotic surgical system 1, please refer to U.S. Patent No. 8,828,023, the entire contents of which are incorporated herein by reference.
[0062] refer to Figure 2 The robot system 10 is configured to support the tool 20 thereon. Figure 1 And selectively relative to the patient's "P" ( Figure 1 The arm 12 includes multiple directional upward-moving tools 20 within a small incision, while simultaneously maintaining the tools 20 within the small incision. The arm 12 includes multiple elongated members or rods 110, 120, 130, 140 that are pivotally connected to each other to provide different degrees of freedom to the arm 12. Specifically, the arm 12 includes a first rod 110, a second rod 120, a third rod 130, and a fourth rod 140.
[0063] The first rod 110 has a first end 110a and a second end 110b. The first end 110a is rotatably coupled to a fixed structure. The fixed structure may be a movable trolley 102 locked in place, a surgical table, a support, an operating room wall, or other structures present in the operating room. A first motor “M1” is operatively coupled to the first end 110a to rotate the first rod 110 about a first axis of rotation A1, which passes through the first end 110a transverse to the longitudinal axis of the first rod 110. The second end 110b of the first rod 110 has a second motor “M2”, which is operatively coupled to the first end 120a of the second rod 120, so that actuation of the motor “M2” affects the rotation of the second rod 120 relative to the first rod 110 about a second axis of rotation A2, which is defined to pass through the second end 110b of the first rod 110 and the first end 120a of the second rod 120. It is foreseeable that the second rotation axis A2 can be transverse to the longitudinal axis of the first rod 110 and the longitudinal axis of the second rod 120.
[0064] The second end 120b of the second rod 120 is operatively connected to the first end 130a of the third rod 130 so that the third rod 130 rotates relative to the second rod 120 about a third rotation axis A3, which passes through the second end 120b of the second rod and the first end 130a of the third rod 130. The third rotation axis A3 is parallel to the second rotation axis A2. The rotation of the second rod 120 about the second rotation axis A2 affects the rotation of the third rod 130 about the third rotation axis A3, such that the first rod 110 and the third rod 130 maintain a substantially parallel relationship to each other.
[0065] The second end 130b of the third rod 130 is operatively connected to the first end 140a of the fourth rod 140. The fourth rod 140 rotates relative to the third rod 130 about a fourth rotation axis A4, which passes through the second end 130b of the third rod 130 and the first end 140a of the fourth rod 140.
[0066] For further reference Figure 3 The fourth rod 140 may be in the form of a track, which supports the slider 142. The slider 142 may slide along an axis parallel to the longitudinal axis of the fourth rod 140 and support the tool 20.
[0067] During a surgical procedure, the robotic system 10 receives input commands from the user interface 40 to move the tool 20, thereby moving the end effector 22 to manipulate and / or act on tissue within the surgical site “S”. Specifically, the levers 110, 120, 130, and 140 of the robotic arm 12 rotate relative to each other, and the slider 142 translates in response to the input commands to position and orient the tool 20 within the surgical site “S”. To control the robotic arm 12, the robotic system 10 calculates the desired tool posture of the tool 20 based on the input commands, captures the tool posture of the tool 20, and manipulates the robotic arm 12 to move the tool 20 to the desired tool posture. Based on the desired tool posture, the robotic system 10 calculates the desired arm posture of the robotic arm 12 to achieve the desired tool posture. Then, the robotic arm 12 responds to the user interface 40 ( Figure 1 The input is captured to determine which of the joysticks 110, 120, 130, and 140 to manipulate in order to achieve the desired arm posture, and thus achieve the desired tool posture of the tool 20 within the surgical site “S”.
[0068] To determine the arm posture of robotic arm 12, robotic system 10 uses an imaging device or endoscope 200 positioned within the surgical site "S" to capture the position and orientation, or tool posture, of tool 20 within the surgical site "S". As detailed below, endoscope 200 is described as capturing tool posture within the surgical site; however, it is contemplated that imaging devices may be used, and each imaging device may include one or more lenses to capture two-dimensional or three-dimensional images.
[0069] Endoscope 200 can be fixed within the surgical site “S” and manipulated by a clinician in the operating room, or it can be attached to another robotic arm 12 to allow manipulation of the position and orientation of endoscope 200 during the surgical procedure. Robotic system 10 uses known techniques to visually capture the tool pose of tool 20 within the surgical site “S” using endoscope 200. Tool 20 may include markings to aid in capturing tool pose, which may include (but are not limited to) different colors, different markings, different shapes, or combinations thereof. The tool pose of tool 20 is captured relative to endoscope 200 in a camera frame, and this tool pose can be transferred to the frame of the surgical site “S”, the frame of tool 20, the frame of robotic arm 12, or any other desired reference frame. It is anticipated that transferring the tool pose of tool 20 to a fixed frame may be advantageous.
[0070] Based on the tool posture of tool 20, robot system 10 can calculate the arm posture of robotic arm 12 using the known kinematic characteristics of robotic arm 12, starting from the tool posture of tool 20 and proceeding towards the first link 110. By calculating the arm posture of robotic arm 12 based on the tool posture of tool 20, the solution for the required tool posture to move tool 20 within the surgical site "S" takes into account any deformation of robotic arm 12 or tool 20 under load. Furthermore, by calculating the arm posture based on the tool posture, it is not necessary to know the position of the fixed structure (e.g., the movable trolley 102) or the first link 110 of arm 12. Figure 2 The connection is independent of where it is attached, and there is no need to determine a solution for moving tool 20 to the desired tool posture. When calculating the solution, the robotic system 10 considers any possible collisions between arm 12 and other arms 12, clinicians in the operating room, patients, or other structures within the operating room. Furthermore, by calculating the tool posture and / or arm posture within a common frame (e.g., the camera frame of a single endoscope), the possible collisions of arm 12 can be estimated by simultaneously calculating the posture of the tool and / or arm using the kinematic properties of each arm (e.g., arm 12) to calculate the position of the rod (e.g., rod 110).
[0071] It is conceivable that the robotic system 10 can be used to simultaneously capture the tool poses of multiple tools 20 via the endoscope 200. By capturing the tool poses of multiple tools 20, the interaction between the tools 20 and their end effectors 22 can be controlled with high precision. This high-precision control can be used to perform automated tasks; for example, suturing tissue. It is foreseeable that by using a single endoscope 200 to capture the tool poses of multiple tools 20, the speed and accuracy of automated tasks can be improved by reducing the need to switch the high-precision tool poses from one camera frame to another during the duration of the automated task.
[0072] It is envisioned that more than one camera and / or endoscope 200 can be used to simultaneously capture the tool posture of the tool 20 within the surgical site “S”. It should be understood that when using multiple cameras, it may be beneficial to convert the position and orientation of the tool 20 to frames other than those defined by one of the cameras.
[0073] It is conceivable that determining the arm posture based on the captured tool posture allows for determining the position of the movable trolley 102 supporting each arm 12 based on the captured tool posture and the kinematic characteristics of the arm 12. After the surgical procedure is completed, the efficiency of the surgical procedure can be determined, and the position of the movable trolley 102 can be recorded. By comparing the position of the movable trolley 102 during the surgical procedure with a high efficiency level, guidance or recommended positions for the movable trolley 102 for a given procedure can be provided to improve the efficiency of future surgical procedures. Improving the efficiency of surgical procedures can reduce costs, surgical time, and recovery time, while also improving surgical outcomes.
[0074] refer to Figure 4 This illustrates a schematic configuration of an endoscope system, which can be Figure 1 The imaging devices 16 and 16B may be different types of systems (e.g., visualization systems, etc.). According to this disclosure, the system includes an imaging device 410, a light source 420, a video system 430, and a display device 440. The light source 420 is configured to provide light to the surgical site via an optical fiber guide 422 through the imaging device 410. The distal end 414 of the imaging device 410 includes an objective lens 436 for receiving or capturing images at the surgical site. The objective lens 436 transfers or transmits the images to an image sensor 432. The images are then transmitted to the video system 430 for processing. The video system 430 includes an imaging device controller 450 for controlling the endoscope and processing the images. The imaging device controller 450 includes a processor 452 connected to a computer-readable storage medium or memory 454, which may be a volatile type of memory such as RAM, or a non-volatile type of memory such as flash memory, disk media, or other types of memory. In various implementations, processor 452 may be another type of processor, such as, but not limited to, digital signal processor, microprocessor, ASIC, graphics processing unit (GPU), field programmable gate array (FPGA), or central processing unit (CPU).
[0075] In various embodiments, memory 454 may be random access memory, read-only memory, disk storage, solid-state memory, optical disk storage, and / or another type of memory. In various embodiments, memory 454 may be separable from imaging device controller 450 and may communicate with processor 452 via a communication bus on a circuit board and / or via a communication cable such as a serial ATA cable or other type of cable. Memory 454 contains computer-readable instructions executable by processor 452 to operate imaging device controller 450. In various embodiments, imaging device controller 450 may include network interface 540 for communication with other computers or servers.
[0076] refer to Figure 5 The flowchart includes various blocks described in an ordered order. However, those skilled in the art will understand that one or more blocks of the flowchart may be performed, repeated, and / or omitted in a different order without departing from the scope of this disclosure. The following description of the flowchart relates to various actions or tasks performed by one or more video systems 430, but those skilled in the art will understand that the video system 430 is exemplary. In various embodiments, the disclosed operations may be performed by another component, device, or system. In various embodiments, the video system 430 or other component / device performs actions or tasks via one or more software applications executing on a processor. In various embodiments, at least some operations may be implemented by firmware, programmable logic devices, and / or hardware circuitry systems. Other implementations are also contemplated within the scope of this disclosure.
[0077] Now for reference Figure 5 This illustrates the operations used for mapping and fusing endoscopic images. In various embodiments, this can be performed by the robotic system 10 described above herein. Figure 5 The operation. In various implementations, it can be performed by another type of system and / or during another type of program. Figure 5 The following description will refer to a robotic system, but it should be understood that such description is exemplary and does not limit the scope of this disclosure or its applicability to other systems and programs.
[0078] Initially, in step 502, a first image of the surgical site is captured via objective lens 436 and transferred to image sensor 432 of the first imaging device 16 of the robotic system 10. In step 504, a second image of the surgical site is captured via objective lens 436 and transferred to image sensor 432 of the second imaging device 16B of the robotic system 10.
[0079] The image may include a first light (e.g., infrared) and a second light (e.g., visible light). For example, two light sources may be present to illuminate the surgical site for the robotic system 10. One light source may be broad-spectrum white light, whose wavelength will be blocked so that the wavelength does not exceed the visible range of approximately 740 nm. The other light source may be pure near-infrared, typically located anywhere between approximately 780 nm and 850 nm. It is envisioned that the first and second light sources can be used simultaneously or in any order.
[0080] As used herein, the term "image" can include still images or moving images (e.g., video). A first image includes a first light source (e.g., infrared). A second image includes a second light source (e.g., visible light). For example, two light sources may be present to illuminate the surgical site for the robotic system 10. One light source may be broad-spectrum white light, whose wavelength will be blocked so that it does not exceed the visible range of approximately 740 nm. The other light source may be pure near-infrared, typically anywhere between approximately 780 nm and 850 nm. It is contemplated that the first and second light sources can be used simultaneously or in any order. In the system, the image sensor 432 of the imaging devices 116, 16B may include a CMOS sensor.
[0081] In the system, when fluorescence-based imaging based on indocyanine green (ICG) is required, the system may include a mode that allows both visible and infrared (IR) illumination to be turned on simultaneously while significantly reducing the illumination intensity of the visible component. ICG-based imaging uses near-infrared light to add contrast to tissue imaging during surgical procedures.
[0082] In the system, the captured images are transmitted to video system 430 for processing. For example, during an endoscopic procedure, a surgeon may use electrosurgical instruments to cut tissue. When the first and second images are captured, they may include objects such as tissue and / or instruments.
[0083] The object may also include reference points configured to assist in the alignment of multiple endoscopic images. Structured light can be used to project these reference points onto the object. Structured light is the process of projecting a known pattern (typically a grid or horizontal bars) onto a scene. The way these patterns deform upon impact with a surface allows vision systems such as robotic system 10 to calculate depth and surface information of objects in the scene. Invisible (or imperceptible) structured light uses structured light without interfering with other computer vision tasks where the projected pattern would be confusing. Exemplary methods include alternating between two completely opposite patterns using IR light or extremely high frame rates. In various embodiments, structured light may include time-sequential and / or geometrically unique structured light projection.
[0084] The IR structured light reference point on the object can be visible in this mode, and is preferred when the tissue shows signs of perfusion. In various embodiments, the IR structured light reference point can be darker relative to the perfusion. The system can retune the IR wavelength of the IR structured light reference point. For example, the ICG IR light can be at 785 nm, while the IR structured light can be above 850 nm. In this way, the structured light IR light will not stimulate the ICG, and vice versa. In various embodiments, the imaging devices 16, 16B and / or the robotic system 10 may include multiple IR sources.
[0085] For example, a reference point may be located on the axis of an instrument or on an organ. A reference point may include a geometry, or, for example, a mark, QR code, texture, dot pattern, and / or unique identifier.
[0086] CMOS imagers used in endoscopes are sensitive to IR and typically have filters that block the reception of this IR to prevent image skew by light invisible to the human eye. Since there is generally no light present inside the human body, all illumination must be added by the endoscopic system; for example, this illumination is not natural light and contains IR. The same ICG capability that can be built into the endoscope can be utilized if the required IR wavelength of the structured light is tuned to the same wavelength used to activate the indocyanine green dye (ICG) that can be used to observe perfusion during surgery. For example, this wavelength can be in the range of approximately 785 nm. The video system 430 accesses the first and second images for further processing.
[0087] In step 506, the video system 430 compares the first position of the first reference point in the first image with the second position of the second reference point in the second image. For example, due to the different relative orientations between the first imaging device and the second imaging device, an object in the first image may be in a slightly different position than the same object in the second image.
[0088] In step 508, the video system 430 determines the relative pose of the first imaging device and the second imaging device based on comparison. For example, the video system 430 may determine, based on comparison, that the first imaging device is a few centimeters to the left of the second imaging device. The first image may include distance information for each pixel of the first and second images. The relative pose of the first imaging device may be based on the distance information in the first image. The relative pose of the second imaging device may be based on the distance information in the second image.
[0089] In step 510, the video system 430 generates an enhanced image that fuses the first and second images based on the determined relative pose. The visible light image may include color information, and the IR image may include a grayscale image. When the two images are fused, some form of blending can be performed. Some regions represented by the data fusion may include only information from a single imaging device. For example, in the enhanced image, an organ may include only color information and not grayscale information, but the remainder of the image will be a mixture of color and grayscale information.
[0090] For example, if depth-enabled images are to be fused into a surgical site representation from multiple cameras with widely varying imaging modalities—i.e., visible light produces full color in one imaging modality while NIR produces grayscale in another—adaptations must be made to blend these. Much like the common variations in transparency-based rendering of NIR information over visible light, some form of appropriate blending is necessary. There may exist areas in the data fusion representation that will only have information from the single camera projected onto them, so that they will not have aspects of blended information. Still, when viewing camera points in those areas, the transition from blending to a single source can be rendered in an appropriate manner to mitigate cognitive dissonance.
[0091] When generating an enhanced image, the video system 430 can determine a first optical path distortion of the first imaging device and a second optical path distortion of the second imaging device. The video system 430 can then process the first image based on the first optical path distortion to match the second optical path distortion.
[0092] The video system 430 can generate a virtual image perspective. The video system 430 can determine the virtual imaging device perspective based on the clinician's input, and the generated enhanced image can be based on the virtual imaging device perspective. For example, a projected view from the desired virtual endoscope position can be generated to provide the clinician with the desired perspective.
[0093] In step 512, the video system 430 displays an enhanced image on a monitor for operator viewing. In various embodiments, the video system 430 may perform tracking of the detected object based on a first reference point and a second reference point. For example, stereoscopic viewing can also be provided by creating two synthetic cameras instead of two synthetic cameras separated by the desired stereo base distance. The resulting two images can be displayed on a typical stereo monitor using standard techniques.
[0094] The flowchart includes various blocks described in an ordered sequence. However, those skilled in the art will understand that one or more blocks of the flowchart may be performed, repeated, and / or omitted in a different order without departing from the scope of this disclosure. The following description of the flowchart relates to various actions or tasks performed by one or more video systems 430, but those skilled in the art will understand that the video system 430 is exemplary. In various embodiments, the disclosed operations may be performed by another component, device, or system. In various embodiments, the video system 430 or other component / device performs actions or tasks via one or more software applications executing on a processor. In various embodiments, at least some operations may be implemented by firmware, programmable logic devices, and / or hardware circuitry systems. Other implementations are also contemplated within the scope of this disclosure.
[0095] In various implementation schemes, the robot system 10 described above can perform the actions. Figure 5 The operation may be performed by another type of system and / or during another type of process. The following description will refer to robotic systems, but it should be understood that such description is exemplary and does not limit the scope of this disclosure or its applicability to other systems and programs.
[0096] Initially, in step 602, a first image of the surgical site is captured via objective lens 436 and transferred to image sensor 432 of the first imaging device 16 of the robotic system 10. In step 604, a second image of the surgical site is captured via objective lens 436 and transferred to image sensor 432 of the second imaging device 16B of the robotic system 10. For example, the first imaging device may be a wide-field 2D imaging device located on / in a cannula. The second imaging device may be, for example, a stereoscopic imaging device, such as an endoscope. The first and second imaging devices may have different perspectives from each other. It is contemplated that the first and second imaging devices may be any combination of imaging devices (e.g., two 2D devices, a 2D device and a stereoscopic device, two stereoscopic devices, etc.). It is also contemplated that more than two imaging devices may be used for image capture.
[0097] Next, at step 606, the method segments the first and second images to extract known references (e.g., surgical instruments or imaging devices). Image segmentation is the process of dividing an image into multiple segments (e.g., a set of pixels called image objects) and assigning labels to each pixel in the image such that pixels with the same label share certain characteristics. Image segmentation can be used for object detection in images, for example, to extract known references (e.g., structured light, organs, or surgical instruments) in images of surgical sites.
[0098] In various embodiments, the first image may include a first location of a first imaging device, and the second image may include the location of a second imaging device. Location information may be based on, for example, a location sensor (e.g., GPS, RFID) or manual input of the location of the imaging device. The first imaging device has a first viewpoint of the surgical site, and the second imaging device may have a second viewpoint of the surgical site, which differs from the first viewpoint. In various embodiments, a known reference object may include structured light projected onto the surgical site. A second light source may be used to project structured light onto the surgical site. The structured light may be captured in the first and second images and may be used, for example, for comparing the locations of the images.
[0099] Next, in step 608, the method determines the relative position of each imaging device in the imaging device based on known reference objects (e.g., a first relative position and a second relative position, respectively).
[0100] Next, in step 610, the method constructs a 3D model based on the first relative position determined by the first imaging device and the second relative position determined by the second imaging device.
[0101] When constructing the 3D model, the method may further include determining the virtual imaging device viewpoint, thereby generating a virtual viewpoint of the 3D model based on the virtual imaging device. For example, a virtual viewpoint of the 3D model from a desired virtual endoscopic position can be generated to provide the surgeon with their desired viewpoint. This viewing position may be dynamically updated relative to a time-varying 3D model representation, or the 3D model representation may be fixed in time or even recorded and played back, while the virtual imaging device viewpoint is manipulated to provide additional surgical site insight. It should also be noted that multiple observers may simultaneously view the surgical site from their own selected viewpoints. It should also be noted that stereoscopic viewing can also be provided by creating two virtual imaging device positions instead of two virtual imaging device positions separated by a desired stereoscopic base. The resulting two images can be displayed on a typical stereoscopic monitor using standard techniques. Next, at step 612, the method displays the stereoscopic virtual imaging device viewpoint.
[0102] The embodiments disclosed herein are examples of this disclosure and may be embodied in various forms. For example, while some embodiments herein are described as separate embodiments, each of these embodiments may be combined with one or more of the other embodiments herein. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but rather as the basis of the claims and as a representative basis for teaching those skilled in the art to apply the contents of this disclosure in virtually any properly detailed description. Throughout the description of the figures, the same reference numerals may refer to similar or equivalent elements.
[0103] The phrases “in one embodiment,” “in an embodiment,” “in some embodiments,” or “in other embodiments” may each refer to one or more of the same or different embodiments according to this disclosure. A phrase in the form “A or B” means “(A), (B), or (A and B).” A phrase in the form “at least one of A, B, or C” means “(A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C).” The term “clinician” may refer to a clinician performing medical procedures or any medical professional, such as a doctor, nurse, technician, medical assistant, etc.
[0104] The system described herein may also utilize one or more controllers to receive various types of information and transform the received information to generate output. The controller may include any type of computing device, computing circuitry, or any type of processor or processing circuitry capable of executing a series of instructions stored in memory. The controller may include multiple processors and / or a multi-core central processing unit (CPU), and may include any type of processor, such as a microprocessor, digital signal processor, microcontroller, programmable logic device (PLD), field-programmable gate array (FPGA), etc. The controller may also include memory to store data and / or instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more methods and / or algorithms.
[0105] Any of the methods, programs, algorithms, or code described herein can be translated into or expressed in a programming language or computer program. As used herein, the terms "programming language" and "computer program" each include any language used to specify instructions for a computer, and include (but are not limited to) the following languages and their derivatives: assembler, Basic, batch files, BCPL, C, C+, C++, Delphi, Fortran, Java, JavaScript, machine code, operating system command languages, Pascal, Perl, PL1, scripting languages, visual Basic, meta-languages that specify their own programs, and all first, second, third, fourth, fifth, or higher generation computer languages. Databases and other data schemas, as well as any other meta-languages, are also included. No distinction is made between interpreted, compiled languages, or languages that use both compilation and interpretation methods. No distinction is made between compiled and source versions of a program. Therefore, references to programs in which programming languages may exist in more than one state (such as source, compiled, target, or linked) are references to any and all such states. References to programs may cover actual instructions and / or the intent of those instructions.
[0106] Any of the methods, programs, algorithms, or code described herein may be contained on one or more machine-readable media or memories. The term "memory" may include a means for providing (e.g., storing and / or transmitting) information in a machine-readable form, such as a processor, computer, or digital processing device. For example, memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, or any other volatile or non-volatile memory storage device. The code or instructions contained thereon may be represented by carrier signals, infrared signals, digital signals, and other similar signals.
[0107] It should be understood that the foregoing description is merely illustrative of this disclosure. Various alternatives and modifications can be devised by those skilled in the art without departing from this disclosure. Therefore, this disclosure is intended to encompass all such alternatives, modifications, and variations. The embodiments described with reference to the accompanying drawings are merely illustrative of certain examples of this disclosure. Other elements, steps, methods, and techniques that are not materially different from those described in the foregoing and / or appended claims are also intended to be included within the scope of this disclosure.
Claims
1. A system for mapping and fusing endoscopic images, the system comprising: monitor; A light source configured to provide light within a surgical site, the light source being configured to generate a first light including an infrared (IR) band and a second light configured to generate a visible band; A first imaging device, configured to acquire images from the surgical site; A second imaging device, configured to acquire images from the surgical site; and An imaging device control unit, configured to control the first imaging device and the second imaging device, the control unit comprising: processor; and A memory, wherein instructions are stored thereon, which, when executed by the processor, cause the system to: A first image of an object within a surgical site, captured by the first imaging device, the first image including the first light radiated from the object and a first reference point; A second image of the object is captured by the second imaging device, the second image including the second light radiated from the object, and a second reference point; The first position of the first reference point in the first image is compared with the second position of the second reference point in the second image; The relative attitudes of the first imaging device and the second imaging device are determined based on the comparison. An enhanced image is generated based on the determined relative pose, fusing the first image and the second image. The first image is an infrared (IR) image including grayscale information, and the second image is a visible light image including color information. When generating the enhanced image, the portion of the second image containing the object is used to represent the object in the enhanced image, and the portion of the enhanced image containing the object includes only the color information. The remaining portion of the enhanced image includes a mixture of the color information from the second image and the grayscale information from the first image. The enhanced image is displayed on the monitor.
2. The system of claim 1, wherein the first reference point comprises structured light, and wherein the second reference point comprises structured light.
3. The system of claim 2, wherein generating the enhanced image further comprises: Determine the viewpoint of the virtual imaging device, and The enhanced image is also generated based on the viewpoint of the virtual imaging device.
4. The system of claim 1, wherein generating the enhanced image further comprises: Determine the first optical path distortion of the first imaging device and the second optical path distortion of the second imaging device; as well as The first image is processed based on the first optical path distortion to match the second optical path distortion.
5. The system of claim 1, wherein the instructions, when executed, further cause the system to perform tracking of the object based on the first reference point and the second reference point.
6. The system of claim 1, wherein the first reference point and the second reference point include identifiers.
7. The system of claim 1, wherein the first reference point and the second reference point comprise at least one of a QR code, a texture, a dot pattern, or a unique identifier.
8. The system of claim 1, wherein the first image includes first distance information for each pixel of the first image. The second image includes second distance information for each pixel of the second image, and The relative attitude of the first imaging device is also based on the first distance information and the second distance information.
9. A method for mapping and fusing endoscopic images, the method comprising: A first image of an object within a surgical site, captured by a first imaging device, the first image including first light radiating from the object and a first reference point; A second image of the object is captured by a second imaging device, the second image including a second reference point and second light radiated from the object, and the second reference point; The first position of the first reference point in the first image is compared with the second position of the second reference point in the second image; The relative attitudes of the first imaging device and the second imaging device are determined based on the comparison. An enhanced image is generated based on the determined relative pose, which is an infrared (IR) image including grayscale information and a visible light image including color information. When the enhanced image is generated, the portion of the second image that includes the object is used to represent the object in the enhanced image, and the portion of the enhanced image that includes the object only includes the color information. The remaining portion of the enhanced image includes a mixture of the color information of the second image and the grayscale information of the first image. as well as The enhanced image is displayed on the monitor. The first image and the second image are static images.
10. The method of claim 9, wherein the first light comprises an infrared (IR) band and the second light comprises a visible band.
11. The method of claim 9, wherein the first reference point comprises structured light, and wherein the second reference point comprises structured light.
12. The method of claim 11, wherein generating the enhanced image further comprises: Determine the viewpoint of the virtual imaging device, and The enhanced image is also generated based on the viewpoint of the virtual imaging device.
13. The method of claim 9, wherein generating the enhanced image further comprises: Determine the first optical path distortion of the first imaging device and the second optical path distortion of the second imaging device; as well as The first image is processed based on the first optical path distortion to match the second optical path distortion.
14. The method of claim 9, wherein the method further comprises performing tracking of the object based on the first reference point and the second reference point.
15. The method of claim 9, wherein the first reference point and the second reference point include identifiers.
16. The method of claim 9, wherein the first reference point and the second reference point comprise at least one of a QR code, a texture, a dot pattern, or a unique identifier.
17. The method of claim 9, wherein the first image includes first distance information for each pixel of the first image. The second image includes second distance information for each pixel of the second image, and The relative attitude of the first imaging device is also based on the first distance information and the second distance information.
18. The method of claim 9, wherein the first imaging device and the second imaging device comprise a stereoscopic imaging device.
19. A non-transitory storage medium storing a program that causes a computer to execute a computer-implemented method for mapping and fusing endoscopic images, the computer-implemented method comprising: A first image of an object within a surgical site, captured by a first imaging device, the first image including first light radiating from the object and a first reference point; A second image of the object is captured by a second imaging device, the second image including second light radiated from the object and a second reference point; The first position of the first reference point in the first image is compared with the second position of the second reference point in the second image; The relative attitudes of the first imaging device and the second imaging device are determined based on the comparison. An enhanced image is generated based on the determined relative pose, which is an infrared (IR) image including grayscale information and a visible light image including color information. When the enhanced image is generated, the portion of the second image that includes the object is used to represent the object in the enhanced image, and the portion of the enhanced image that includes the object only includes the color information. The remaining portion of the enhanced image includes a mixture of the color information of the second image and the grayscale information of the first image. as well as The enhanced image is displayed on the monitor.
Citation Information
Patent Citations
Medical workstation
US8828023B2
Wavelength diverse scintillation reduction
CN103460223A
An image capture unit in a surgical instrument
CN103889353A
Surgical HUB spatial awareness to determine devices in operating theater
US20190201104A1