Endoscope master-slave motion control method and surgical robot system
By integrating a multi-field imaging unit onto the endoscope to generate synthetic scene images, the problem of unclear image field of view during remote operation is solved, improving the accuracy of surgical operations and the operating experience, and enhancing the integration of the surgical robot system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2026-03-24
AI Technical Summary
In minimally invasive surgery, existing technologies struggle to provide a clear visual field during remote operation, resulting in a poor operator experience and impacting surgical precision and efficiency.
By integrating first and second imaging units on the endoscope, images from different fields of view are captured, and a composite scene image is generated and displayed on a display device, including actual and virtual images of the end-effectors, thereby improving the operator's field of vision and operational experience at the surgical site.
It improves the precision of surgical procedures and the operator's remote operation experience, provides a clear visual field, and enhances the integration and operational precision of the surgical robot system.
Smart Images

Figure CN115517615B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of medical devices, and in particular to a method for master-slave motion control of an endoscope and a surgical robot system. BACKGROUND
[0002] In modern minimally invasive surgery, a surgical robot system with an endoscope is needed for surgical operation. The surgical robot system for teleoperation includes a master operator for operation and an endoscope for shooting images inside a cavity under the control of the master operator. In actual scenarios, the endoscope is arranged to include an execution arm that accepts control of the master operator, and the operator controls the execution arm to extend into the cavity by teleoperation of the master operator, and then shoots images by an imaging unit arranged on the endoscope. Sometimes, in order to improve the integration of the surgical robot system, an end effector is also integrated on the endoscope to take into account image acquisition and surgical operation.
[0003] The surgical robot has high requirements for operation precision and human-computer interaction experience. During teleoperation, the operator needs to be provided with images corresponding to operation instructions to help the operator perform surgical operation according to the intracavity condition in a state that the field of view is not affected, and to improve the operation experience of the operator. SUMMARY
[0004] In some embodiments, the present disclosure provides a method for master-slave motion control of an endoscope. The endoscope includes an execution arm, a main body arranged at the end of the execution arm, a first imaging unit, a second imaging unit, and an end effector extending from the distal end of the main body. The method can include: determining a current pose of the master operator; determining a target pose of the end of the execution arm based on the current pose of the master operator and a pose relationship between the master operator and the end of the execution arm; generating a driving instruction for driving the end of the execution arm based on the target pose; obtaining a first image from the first imaging unit; obtaining a second image from the second imaging unit, wherein the fields of view of the first image and the second image are different and include an image of the end effector; generating a composite scene image to remove an actual image of the end effector based on the first image and the second image; and generating a virtual image of the end effector in the composite scene image.
[0005] In some embodiments, the present disclosure provides a robotic system, comprising: a master operator comprising a mechanical arm, a handle disposed on the mechanical arm, and at least one master operator sensor disposed at at least one joint of the mechanical arm, the at least one master operator sensor being configured to obtain joint information of the at least one joint; an endoscope comprising an execution arm, a main body disposed at a distal end of the execution arm, a first imaging unit, a second imaging unit, and an end instrument extending from a distal end of the main body; at least one driving device configured to drive the execution arm; at least one driving device sensor coupled to the at least one driving device and configured to obtain state information of the at least one driving device; a control device configured to be connected with the master operator, the at least one driving device, the at least one driving device sensor, and the endoscope, and perform a method according to any one of some embodiments of the present disclosure; and a display device configured to display an image based on an instruction output by the control device.
[0006] In some embodiments, the present disclosure provides a computer device, comprising: a memory configured to store at least one instruction; and a processor coupled to the memory and configured to execute the at least one instruction to perform a method according to any one of some embodiments of the present disclosure.
[0007] In some embodiments, the present disclosure provides a computer-readable storage medium configured to store at least one instruction, the at least one instruction being configured to cause a computer to perform a method according to any one of some embodiments of the present disclosure when executed by the computer. BRIEF DESCRIPTION OF DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings needed in the description of the embodiments of the present disclosure. The drawings in the following description only show some embodiments of the present disclosure, and other embodiments can be obtained by those skilled in the art without creative labor on the basis of the contents of the embodiments of the present disclosure and these drawings.
[0009] Figure 1 A structural block diagram of a robotic system according to some embodiments of the present disclosure is shown;
[0010] Figure 2 A structural schematic diagram of an endoscope according to some embodiments of the present disclosure is shown;
[0011] Figure 3 A structural schematic diagram of an endoscope located in a body lumen according to some embodiments of the present disclosure is shown;
[0012] Figure 4 A schematic diagram of the relative positional relationship between the first imaging unit, the second imaging unit, and the end instrument according to some embodiments of the present disclosure is shown;
[0013] Figure 5A schematic block diagram showing a control device according to some embodiments of the present disclosure;
[0014] Figure 6 A flowchart showing an endoscopic master-slave motion control method according to some embodiments of the present disclosure;
[0015] Figure 7 (a), Figure 7 (b) shows a coordinate system diagram of a robot system according to some embodiments of the present disclosure, wherein Figure 7 (a) is a coordinate system diagram in master-slave motion mapping, Figure 7 (b) is a coordinate system diagram of an endoscope;
[0016] Figure 8 A schematic diagram showing a master operator according to some embodiments of the present disclosure;
[0017] Figure 9 A flowchart showing a method of displaying a scene image based on a display mode instruction according to some embodiments of the present disclosure;
[0018] Figure 10 A schematic diagram showing multi-scene display on a display device according to some embodiments of the present disclosure;
[0019] Figure 11 A flowchart showing a method of generating a synthetic scene image based on a first image and a second image according to some embodiments of the present disclosure;
[0020] Figure 12 A flowchart showing a method of generating a three-dimensional synthetic scene image based on a first image and a second image according to some embodiments of the present disclosure;
[0021] Figure 13 A flowchart showing a method of generating a depth map based on an optical flow field and a pose of an imaging unit according to some embodiments of the present disclosure;
[0022] Figure 14 A flowchart showing a method of generating a three-dimensional actual scene image based on a first image and / or a second image according to some embodiments of the present disclosure;
[0023] Figure 15 A flowchart showing a method of generating a virtual image of an end effector in a synthetic scene image according to some embodiments of the present disclosure;
[0024] Figure 16 A schematic block diagram showing a computer device according to some embodiments of the present disclosure;
[0025] Figure 17 A schematic diagram showing a robot system according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0026] To make the technical problems solved by this disclosure, the technical solutions adopted, and the technical effects achieved clearer, the technical solutions of the embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely exemplary embodiments of this disclosure, and not all embodiments.
[0027] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this disclosure, it should be noted that unless otherwise expressly specified and limited, the terms "installed," "connected," "coupled," and "coupled" should be interpreted broadly. For example, they can refer to fixed connections or detachable connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.
[0028] In this disclosure, the end closer to the operator (e.g., a doctor) is defined as the proximal end, proximal or rear end, or rear part, and the end closer to the surgical patient is defined as the distal end, distal or anterior end, or front part. It is understood that embodiments of this disclosure can be used in medical devices or surgical robots, as well as other non-medical devices. In this disclosure, a reference coordinate system can be understood as a coordinate system capable of describing the pose of an object. Depending on the actual positioning requirements, the reference coordinate system can be selected with the origin of a virtual reference object or the origin of a physical reference object as its origin. In some embodiments, the reference coordinate system can be a world coordinate system, or a coordinate system in the space where a point on the main manipulator, the end effector arm, the endoscope body, the end effector, the imaging unit, or the cavity is located, or the operator's own perception coordinate system, etc.
[0029] In this disclosure, an object can be understood as an object or target that needs to be positioned, such as an actuator arm or the end effector of an actuator arm, or a point on a cavity. The pose of the actuator arm or a part thereof (e.g., the end effector) can refer to the pose of the coordinate system defined by the actuator arm, a part of the actuator arm, or a rigidly extended portion of the actuator arm (e.g., the body or end effector of an endoscope) relative to a reference coordinate system.
[0030] Figure 1A structural block diagram of a robot system 100 according to some embodiments of the present disclosure is shown. In some embodiments, such as Figure 1 As shown, the robot system 100 may include a master control carriage 110, a slave carriage 130, and a control device 120. The control device 120 can communicate with the master control carriage 110 and the slave carriage 130, for example, via cable or wireless connection, to achieve communication between them. The master control carriage 110 functions as the operating and interactive end of the robot system 100, and may include a master manipulator for remote operation by the operator and a display device for displaying images of the operating area. The slave carriage 130 functions as the working end of the robot system 100, and includes an endoscope for capturing images and performing tasks. The control device 120 enables master-slave mapping between the master manipulator in the master control carriage 110 and the endoscope in the slave carriage 130, allowing the master manipulator to control the movement of the endoscope. In some embodiments, the distal portion of the endoscope is configured to enter the operating area via a cavity through a cannula, sheath, etc., to capture images of the target area in the scene and generate two-dimensional or three-dimensional scene images, which are then displayed on a display device. The endoscope may include a distal end-effector, which may be a surgical tool such as a clamp, hemostatic device, or drug delivery device. In some embodiments, the robotic system 100 may be configured to process the captured scene images to generate various images, such as an actual scene image including the end-effector and / or a composite scene image including a virtual image of the end-effector, and selectively display them on a display device based on instructions (e.g., display mode instructions). By controlling the robotic system 100, the end-effector can be operated to perform surgical operations on the object to be operated on (e.g., pathological tissue, etc.) at the surgical site, either by contact or non-contact, while having a field of vision of the operating area and surgical site. The cannula or sheath may be fixed to a human or animal body to form an opening (e.g., an incision or natural opening), the cavity may be the trachea, esophagus, vagina, intestine, etc., the operating area may be the area where the surgical operation is performed, and the scene may be the cavity or the operating area. Those skilled in the art will understand that the master control carriage 110 and the driven carriage 130 may adopt other structures or forms, such as bases, supports, or buildings. The master control carriage 110 and the driven carriage 130 may also be integrated into the same device.
[0031] Figure 2 A schematic diagram of the structure of an endoscope 200 according to some embodiments of the present disclosure is shown. Figure 3 This diagram illustrates the structure of an endoscope 200 according to some embodiments of the present disclosure located within a cavity A in a body (e.g., in a human or animal body). Figure 2 and Figure 3As shown, the endoscope 200 can include an execution arm 210 and an endoscope body 221. In some embodiments, the execution arm 210 can be a continuum capable of controllable bending, can include a segment capable of bending in at least one degree of freedom at a distal end, and the endoscope body 221 can be disposed at the distal end of the execution arm 210. The execution arm 210 can be disposed on the slave trolley 130, and there can be a master-slave motion mapping relationship between the pose of the master operator and the pose of the distal end of the execution arm. In some embodiments, the execution arm can change the steering of the endoscope body 221 based on the operation instruction issued by the operator to avoid important organs in the body or to adapt to the complex bending of the cavity A, so as to facilitate the endoscope body 221 to feed into the cavity A to reach the operation area or to retreat from the cavity A. In some embodiments, the execution arm can adjust the pose of the endoscope body 221 based on the operation instruction, so as to facilitate the imaging unit (for example, Figure 2 and Figure 3 the first imaging unit 230 and the second imaging unit 240 as shown, Figure 4 the first imaging unit 430 and the second imaging unit 440 as shown, or Figure 17 the first imaging unit 1761 and the second imaging unit 1762 as shown) to shoot the operation area, and to align the end effector (for example, Figure 2 and Figure 3 the end effector 260 as shown, Figure 4 the end effector 460 as shown, or Figure 17 the end effector 1765 as shown) to the surgical site in the operation area.
[0032] The endoscope 200 can further include the first imaging unit 230, the second imaging unit 240, and the end effector 260. In some embodiments, the endoscope 200 can further include at least one illumination unit 250. As Figure 2 shown, the endoscope body 221 can be generally cylindrical, and the cross-sectional shape can be circular or elliptical, etc. to meet different functional needs.
[0033] The first imaging unit 230 can be used to capture a first image, and the second imaging unit 240 can be used to capture a second image. In some embodiments, the first imaging unit 230 and the second imaging unit 240 can be, for example, a CCD camera, each including a set of image sensors and image lenses. The image lenses can be disposed at the distal end of the image sensors and aligned with at least a corresponding image sensor, thereby facilitating the image sensors to capture target areas in the scene through the image lenses. In some embodiments, the image lenses can include multiple convex lenses and concave lenses, which are distributed to form an optical imaging system. For example, the distal surface of the image lens can be a curved convex lens, such as a spherical lens, an ellipsoidal lens, a conical lens, a frustum lens, etc. The image lens can include at least one convex surface to increase the field of view that can be captured.
[0034] The illumination unit 250 provides illumination to facilitate imaging by the first imaging unit 230 and the second imaging unit 240. For example... Figure 2 As shown, in some embodiments, three illumination units 250 are arranged along the circumferential edge of the endoscope 200, located between two imaging units or between an imaging unit and the end instrument 260, but are not limited thereto. The number and arrangement of the illumination units 250 can be changed according to actual needs. For example, there may be two illumination units 250, located on the left and right sides of the endoscope 200. Alternatively, to further increase the illumination intensity, the number of illumination units 250 may be greater than three. In some embodiments, the cross-section of the illumination unit 250 may be crescent-shaped, thereby making full use of the space on the endoscope 200, which helps to achieve miniaturization of the endoscope and increase the illumination field of view, but is not limited thereto, the illumination unit 250 may also be other shapes. In some embodiments, the illumination unit 250 may include a light source and one or more optical fibers coupled to the light source, and an illumination channel for arranging the optical fibers may be formed inside the endoscope body 221. In some embodiments, the light source of the illumination unit 250 may be, for example, an LED light source.
[0035] In some embodiments, the distal instrument 260 may be configured to extend distally from the endoscope body 221 to perform surgical procedures, as described later. In some embodiments, the endoscope 200 may further include a ranging unit (not shown) for measuring the distance between the endoscope 200 and the surgical site. The ranging unit may be, for example, a ranging sensor such as a laser rangefinder. By providing the ranging unit, the distance between the endoscope 200 (e.g., the distal end face of the endoscope body 221) and the surgical site can be determined, thereby enabling the further determination of the distance between the distal instrument 260 and the surgical site.
[0036] Figure 4A schematic diagram illustrating the relative positional relationship between a first imaging unit 430, a second imaging unit 440, and an end effector 460 according to some embodiments of the present disclosure is shown. Figure 4 As shown, in some embodiments, the first imaging unit 430 may be configured to be located on one side of the endoscope body 421 relative to the end instrument 460 and have a first field of view corresponding to the orientation of its own optical axis L1.
[0037] It should be understood that the first field of view of the first imaging unit 430 is formed in a roughly conical shape centered on the optical axis L1. Figure 4 The cross-section of the first field of view is schematically shown, and the plane containing this cross-section is perpendicular to the optical axis L1. Furthermore, it should be understood that the planes containing the first imaging unit 430 and the second imaging unit 440 (described below), the distal surface of the end effector 460, the first field of view, and the cross-section of the second field of view (described below) are not located on the same plane. This is for ease of description. Figure 4 The plane containing the first imaging unit 430 and the second imaging unit 440 described below, the distal surface of the end effector 460, the first field of view, and the cross section of the second field of view described below are shown on the same plane.
[0038] The first field of view includes field of view 431 and field of view 432, where field of view 431 is the portion of the first field of view not obstructed by the end effector 460, and field of view 432 is the portion of the first field of view obstructed by the end effector 460. For example... Figure 4 As shown, the fields of view 431 and 432 cover the entire cross-section of the operating region B (or cavity A). The first imaging unit 430 can capture a first image of the operating region B (or cavity A) under the first field of view, the first image including an image of the end effector 460.
[0039] The second imaging unit 440 can be configured to be located on the opposite side of the endoscope body 421 relative to the end-effector 460, and has a second field of view corresponding to its own optical axis L2. Similar to the first imaging unit 430, the second field of view of the second imaging unit 440 is formed in a generally conical shape centered on the optical axis L2. Figure 4 A cross-section of the second field of view is schematically shown, the plane of which is perpendicular to the optical axis L2. The second field of view includes field of view 441 and field of view 442, where field of view 441 is the portion of the second field of view not obstructed by the end effector 460, and field of view 442 is the portion of the second field of view obstructed by the end effector 460. Figure 4 As shown, the fields of view 441 and 442 cover the entire cross-section of the operating region B (or cavity A). The second imaging unit 440 can capture a second image of the operating region B (or cavity A) under the second field of view. This second image has a different field of view from the first image and includes an image of the end effector 460.
[0040] In some embodiments, the optical axis L1 of the first imaging unit 430 and the optical axis L2 of the second imaging unit 440 may be parallel to the axis L0 of the endoscope body 421, respectively, and the axis L3 of the endoscope 460 may be parallel to the axis L0 of the endoscope body 421 and offset from the line connecting the first imaging unit 430 and the second imaging unit 440. For example, the first imaging unit 430, the second imaging unit 440, and the endoscope 460 may be configured such that the optical axis L1 of the first imaging unit 430, the optical axis L2 of the second imaging unit 440, and the axis L3 of the endoscope 460 are perpendicular to the distal surface of the endoscope body 421, the first imaging unit 430 and the second imaging unit 440 are symmetrically distributed with respect to the endoscope 460, and the axis L3 of the endoscope 460 is located below the line connecting the first imaging unit 430 and the second imaging unit 440. By configuring the first imaging unit 430 and the second imaging unit 440 with their optical axes parallel and symmetrically distributed on both sides of the end effector 460, the first image captured by the first imaging unit 430 and the second image captured by the second imaging unit 440 can be made symmetrical to each other, which helps to process the first image and the second image and can improve the image generation quality and image processing speed of the robot system.
[0041] Those skilled in the art will understand that although the first and second images are described as examples in this disclosure for ease of explanation, the embodiments of this disclosure can be applied to processing sequences of first and second images to form continuous video frame processing and display. Therefore, the capturing, processing, and display of sequences of first and second images fall within the scope of this disclosure and the protection scope of the claims of this disclosure.
[0042] In this disclosure, the end device (e.g., Figure 2 and Figure 3 The end effector 260 shown Figure 4 The end effector 460 shown Figure 17 The distal end instrument 1765 shown may include surgical tools such as hemostatic devices (e.g., electrocoagulation hemostatic devices), clamp devices, and drug delivery devices to meet different surgical needs. The following description uses an electrocoagulation hemostatic device as the distal end instrument.
[0043] In some embodiments, the distal end device can be configured to be fixedly connected proximally to the distal end of the endoscope body, thereby allowing the position of the distal end device to be changed by adjusting the position of the endoscope body, thus aligning the distal end device with the surgical site in the operating area. In some embodiments, the distal end device can be a bipolar electrocoagulation hemostasis device. For example, the distal end device may include at least one first electrode, at least one second electrode, and an insulating body, wherein the at least one first electrode and at least one second electrode are alternately disposed on the circumferentially outer side of the insulating body, and at least a portion of the first electrode and at least a portion of the second electrode are exposed. When a high-frequency current is applied, the at least partially exposed first electrode and the at least partially exposed second electrode form a circuit for electrocoagulation hemostasis.
[0044] Figure 5 A structural block diagram of a control device 500 according to some embodiments of the present disclosure is shown. The control device 500 is used to control an endoscope based on an endoscope master-slave motion control method. Figure 5 As shown, the control device 500 may include a pose determination module 510, a drive module 520, an image processing module 530, a virtual image generation module 540, a display signal generation module 550, and a scene output module 560. The pose determination module 510 can be used to determine the pose of the endoscope 200. In some embodiments, the pose determination module 510 can determine the target pose of the end effector 210 based on the current pose of the master manipulator, thereby determining the pose of the endoscope body 221 and the first imaging unit 230, the second imaging unit 240, and the end effector 260 disposed on the endoscope body 221. The drive module 520 is used to generate drive commands to drive the actuator 210. In some embodiments, the drive module 520 can generate drive commands for driving the end effector 210 based on the target pose of the end effector 210. The image processing module 530 can be configured to receive images from the first imaging unit (e.g., ...). Figure 2 The first imaging unit 230 shown Figure 4 The first imaging unit 430 shown or Figure 17 The first imaging unit 1761 shown receives the first image and receives it from the second imaging unit (e.g., Figure 2 The second imaging unit 240 shown Figure 4 The second imaging unit 440 shown or Figure 17 The second imaging unit 1762 shown receives the second image and generates a synthetic scene image and an actual scene image based on the first and second images. In some embodiments, the image processing module 530 can generate a scene image with depth information based on the pose changes of the first imaging unit 230 and the second imaging unit 240 on the endoscope body 221, as described later. The virtual image generation module 540 can be used to generate an end-effector (e.g., ...) in the synthetic scene image.Figure 2 and Figure 3 The end effector 260 shown Figure 4 The end effector 460 shown Figure 17 A virtual image of the end effector 1765 is shown. The display signal generation module 550 generates a display signal based on a display mode command. In some embodiments, the display mode command is used to switch the image display mode based on operator actions (e.g., controlling endoscope movement, controlling the end effector's operation, or selecting a display mode). The scene output module 560 can, based on the display signal, switch between outputting a synthetic scene image with a virtual image or a real scene image to the display device, or simultaneously output both a synthetic scene image with a virtual image and a real scene image to the display device. It should be understood that the control devices of this disclosure include, but are not limited to, the structures described above; any control device capable of controlling a robot system is within the scope of this disclosure.
[0045] In some embodiments, the display device of the robot system 100 can display images based on instructions output by the control device. Imaging unit (e.g., Figure 2 The first imaging unit 230 and the second imaging unit 240 shown are... Figure 4 The first imaging unit 430 and the second imaging unit 440 shown are either Figure 17 The first imaging unit 1761 and the second imaging unit 1762 shown can be used to acquire images of the operating area and transmit the acquired images to a control device (e.g., Figure 1 The control device 120 shown or Figure 5 The control device 500 shown. The image is processed by the image processing module in the control device (e.g., Figure 5After processing by the image processing module 530, the image is displayed on the display device. The operator can perceive the pose change of the end effector relative to the reference coordinate system in real time through the image on the display device (e.g., an actual scene image or a synthetic scene image of the cavity wall). For example, the displacement direction in the image may differ from the position change direction of the end effector relative to the reference coordinate system, and the rotation direction in the image may also differ from the attitude change direction of the end effector relative to the reference coordinate system. The pose of the master manipulator relative to the reference coordinate system is the pose actually perceived by the operator. The pose change perceived by the operator through remote operation of the master manipulator conforms to a specific pose relationship with the pose change in the image perceived by the operator on the display device. By remotely operating the master manipulator, the pose change of the master manipulator is converted into the pose change of the end effector based on the pose relationship, thereby realizing the pose control of the endoscope body at the end of the manipulator. When the operator grips the handle of the master controller and moves it to operate the actuator arm, based on the principle of intuitive operation, the changes in posture or position in the image perceived by the operator are equal to or proportional to the changes in posture of the master controller perceived by the operator. This helps to improve the operator's teleoperation experience and teleoperation accuracy.
[0046] In this disclosure, during remote operation, the endoscope is controlled by the master controller to move to the desired position and orientation according to the operator's wishes, and intracavitary images corresponding to the operation instructions and field of view requirements are provided to the operator.
[0047] Some embodiments of this disclosure provide an endoscope master-slave motion control method. Figure 6 A flowchart of an endoscope master-slave motion control method 600 according to some embodiments of the present disclosure is shown. In some embodiments, some or all of the steps in method 600 may be performed by a robotic system (e.g., Figure 1 The robot system 100 shown Figure 17 The control device (e.g., for the robot system 1700 shown) Figure 5 The control device 500 shown or Figure 17 The method 600 is executed by the control device 1770 shown. The control device may include a computing device. The method 600 may be implemented by software, firmware, and / or hardware. In some embodiments, the method 600 may be implemented as computer-readable instructions. These instructions may be executed by a general-purpose processor or a special-purpose processor (e.g., a dedicated processor). Figure 17 The control device 1770 shown reads and executes these instructions. In some embodiments, these instructions may be stored on a computer-readable medium.
[0048] Figure 7 (a) Figure 7 (b) A schematic diagram of the coordinate system of a robot system according to some embodiments of the present disclosure is shown, wherein Figure 7(a) is a schematic diagram of the coordinate system in the master-slave motion mapping. Figure 7 (b) is a schematic diagram of the coordinate system of the endoscope. Figure 7 (a) and Figure 7 (b) defines the coordinate systems as follows: the actuator base coordinate system {Tb}, with its origin located at the actuator base or the sheath outlet. Aligned with the extension line of the base or the axial direction of the sheath sleeve. Direction such as Figure 7 As shown in (a), the coordinate system of the end effector arm is {Tt}, with the origin located at the end effector arm. Consistent with the axial direction of the end, Direction such as Figure 7 As shown in (a). The reference coordinate system {w} can be the coordinate system of the space where the actuator arm of the master manipulator or endoscope is located, such as the actuator arm base coordinate system {Tb}, or the world coordinate system, such as... Figure 7 As shown in (a). In some embodiments, the operator's tactile sensation can be used as a reference; when the operator is seated in front of the main control panel, the tactile sensation is upward. Direction, the perceived forward direction is... Direction. The main operator's base coordinate system is {CombX}, and the coordinate axis directions are as follows: Figure 7 As shown in (a). The main controller's handle coordinate system {H}, with coordinate axes oriented as follows. Figure 7 As shown in (a), the imaging unit coordinate system {lens} has its origin at the center of the imaging unit, and the optical axis is oriented as follows. Direction, after the field of vision is straightened, the upper part is Direction. In some embodiments, such as Figure 7 As shown in (b), the imaging unit coordinate system {lens} may include a first imaging unit coordinate system {lens1} and a second imaging unit coordinate system {lens2}, wherein the optical axis direction of the first imaging unit is... Direction, after the field of vision is straightened, the upper part is The direction of the optical axis of the second imaging unit is... Direction, after the field of vision is straightened, the upper part is Direction. The monitor coordinate system {Screen} has its origin at the center of the monitor, and the direction perpendicular to the screen image inwards is... Positive direction, the top of the screen is Direction. In some embodiments, the display coordinate system {Screen} may be consistent with the definition of the field of view direction of the imaging unit coordinate system {lens}, for example, consistent with the definition of the field of view direction of the first imaging unit coordinate system {lens1} or the second imaging unit coordinate system {lens2}. Those skilled in the art will understand that the pose of the imaging unit coordinate system {lens} changes with the movement of the actuator arm's end effector, and the field of view displayed on the display also changes accordingly. The correspondence between the operator's perceived direction and the movement direction of the actuator arm's end effector also changes, but the operator's perceived coordinate system (e.g., the reference coordinate system {w}) and the imaging unit coordinate system {lens} have a predetermined correspondence. For example, when the operator moves forward (e.g., along...), the perceived direction changes with the movement direction of the actuator arm's end effector. When the main manipulator is pushed (in the direction of view), since the display coordinate system {Screen} and the imaging unit coordinate system {lens} have the same definition for the field of view direction, the end of the actuator arm can be controlled along the direction of view. Directional movement. Similarly, there is a specific correspondence in the x and y directions. This provides the operator with an intuitive operational experience.
[0049] The following is based on Figure 7 (a) and Figure 7 (b) The endoscope master-slave motion control method 600 is described using the coordinate system shown in the diagram as an example. However, those skilled in the art will understand that the endoscope master-slave motion control method 600 can be implemented using other coordinate system definitions.
[0050] refer to Figure 6 In step 601, the current pose of the master operator can be determined. The current pose includes the current position and the current orientation. In some embodiments, the current pose of the master operator can be the pose relative to the master operator's base coordinate system {CombX} or reference coordinate system {w}. For example, the pose of the master operator is the pose of the coordinate system {H} defined by the master operator's handle or a portion thereof relative to the master operator's base coordinate system {CombX} (e.g., the coordinate system defined by the support or base on which the master operator is located, or the world coordinate system). In some embodiments, determining the current position of the master operator includes determining the current position of the master operator's handle relative to the master operator's base coordinate system {CombX}, and determining the current orientation of the master operator includes determining the current orientation of the master operator's handle relative to the master operator's base coordinate system {CombX}. In some embodiments, step 601 can be performed by a control device 500 (e.g., pose determination module 520).
[0051] In some embodiments, the current pose of the master operator can be determined based on coordinate transformation. For example, the current pose of the handle can be determined based on the transformation relationship between the coordinate system {H} of the master operator's handle and the base coordinate system {CombX} of the master operator. Typically, the base coordinate system {CombX} of the master operator can be set on the bracket or base on which the master operator is located, and the base coordinate system {CombX} of the master operator remains unchanged during teleoperation. The base coordinate system {CombX} of the master operator can be the same as the reference coordinate system {w} or have a predetermined transformation relationship.
[0052] In some embodiments, the current pose of the master manipulator can be determined based on a master manipulator sensor. In some embodiments, current joint information of at least one joint of the master manipulator is received, and the current pose of the master manipulator is determined based on the current joint information of at least one joint. For example, the current pose of the master manipulator is determined based on the current joint information of at least one joint obtained by the master manipulator sensor. The master manipulator sensor is disposed at at least one joint position of the master manipulator. For example, the master manipulator includes at least one joint, and at least one master manipulator sensor is disposed at at at least one joint. The current pose of the master manipulator is calculated based on the joint information (position or angle) of the corresponding joint obtained by the master manipulator sensor. For example, the current position and current orientation of the master manipulator are calculated based on a forward kinematics algorithm.
[0053] In some embodiments, the master actuator includes at least one attitude joint for controlling the attitude of the handle. Determining the current attitude of the handle of the master actuator includes: obtaining joint information of at least one attitude joint, and determining the current attitude of the master actuator based on the joint information of at least one attitude joint. Figure 8 A schematic diagram of a master operator 800 according to some embodiments of the present disclosure is shown. The master operator 800 may be mounted on a main control carriage, for example... Figure 1 The main control console 110 is shown. (As shown...) Figure 8As shown, the main manipulator 800 includes a multi-degree-of-freedom robotic arm 801, which includes position joints and attitude joints. The attitude joints adjust the attitude of the main manipulator 800, controlling it to achieve a target attitude through one or more attitude joints. The position joints adjust the position of the main manipulator 800, controlling it to achieve a target position through one or more position joints. Main manipulator sensors are located at the attitude and position joints of the robotic arm 801 to acquire joint information (position or angle) corresponding to the attitude and position joints. In some embodiments, the pose of the main manipulator 800 can be represented by a set of joint information (e.g., a one-dimensional matrix composed of this joint information). Based on the acquired joint information, the current pose of the handle 802 of the main manipulator 800 relative to the main manipulator base coordinate system {CombX} can be determined. For example, the main manipulator 800 may include seven joints arranged sequentially from proximal to distal. The proximal end of the main manipulator 800 may be the end closer to the main control carriage (e.g., the end connected to the main control carriage), and the distal end of the main manipulator 800 may be the end farther from the main control carriage (e.g., the end where the handle 802 is located). Joints 5, 6, and 7 are attitude joints used to adjust the attitude of the handle 802 of the main manipulator 800. Based on joint information (such as angles) acquired by the main manipulator sensors of the attitude joints and a forward kinematics algorithm, the current attitude of the main manipulator 800 is calculated. Joints 1, 2, and 3 are position joints used to adjust the position of the handle 802 of the main manipulator 800. Based on joint information (such as position) acquired by the main manipulator sensors of the position joints and a forward kinematics algorithm, the current position of the main manipulator 800 is calculated.
[0054] Continue to refer to Figure 6 In step 603, the target pose of the end effector can be determined based on the current pose of the master manipulator and the pose relationship between the master manipulator and the end effector of the actuator arm. For example, a master-slave mapping relationship can be established between the master manipulator and the end effector of the actuator arm, and the pose of the end effector of the actuator arm can be controlled by remotely operating the master manipulator.
[0055] In some embodiments, the pose relationship between the master manipulator and the end effector of the actuator arm may include the relationship between the pose change of the master manipulator and the pose change of the end effector of the actuator arm, such as being equal or proportional. In some embodiments, the display coordinate system {Screen} may be consistent with the definition of the field of view direction of the imaging unit coordinate system {lens} to obtain an intuitive control experience. The pose change of the imaging unit coordinate system {lens} has a predetermined relationship with the pose change of the master manipulator relative to the reference coordinate system {w}, such as the change in position or orientation being the same or proportional. The imaging unit coordinate system {lens} and the end effector coordinate system {Tt} have a predetermined transformation relationship, such as... Figure 7As shown in (b), the pose change in the end-effector coordinate system {Tt} can therefore be calculated based on the pose change in the imaging unit coordinate system {lens}. Thus, the pose relationship can include the relationship between the pose change of the end-effector relative to the current end-effector coordinate system {Tt} and the pose change of the master manipulator relative to the reference coordinate system {w}, for example, the changes in position or orientation are the same or proportional. The reference coordinate system {w} includes the coordinate system of the space where the master manipulator or the end-effector is located, or the world coordinate system. In some embodiments, step 603 can be performed by the control device 500 (e.g., the pose determination module 520).
[0056] Determining the target pose of the end effector includes: determining the previous pose of the master manipulator, determining the initial pose of the end effector, and determining the target pose of the end effector based on the previous and current poses of the master manipulator and the initial pose of the end effector. The previous and current poses of the master manipulator can be the poses of the master manipulator's handle relative to the master manipulator's base coordinate system {CombX} or reference coordinate system {w}. Based on the previous and current poses of the master manipulator relative to the reference coordinate system {w}, the pose change of the master manipulator relative to the reference coordinate system {w} can be calculated, thereby obtaining the pose change of the end effector relative to the current end effector coordinate system {Tt}. The initial and target poses of the end effector can be the poses of the end effector relative to the end effector's base coordinate system {Tb}. The target pose of the actuator arm's end effector relative to the actuator arm base coordinate system {Tb} can be determined based on the initial pose of the end effector relative to the actuator arm base coordinate system {Tb}, the pose change of the end effector relative to the current end effector coordinate system {Tt}, and the transformation relationship between the current end effector coordinate system {Tt} and the actuator arm base coordinate system {Tb}. The transformation relationship between the current end effector coordinate system {Tt} and the actuator arm base coordinate system {Tb} can be determined based on the initial pose of the end effector relative to the actuator arm base coordinate system {Tb}.
[0057] The pose of the actuator arm's end effector can include the pose of the actuator arm's end effector coordinate system {Tt} relative to the actuator arm's base coordinate system {Tb}. The actuator arm's base coordinate system {Tb} can be the coordinate system of the base on which the actuator arm is mounted, the coordinate system of the sheath through which the actuator arm's end effector passes (e.g., the coordinate system of the sheath exit), the coordinate system of the actuator arm's remote center of motion (RCM), etc. For example, the actuator arm's base coordinate system {Tb} can be set at the sheath exit position, and during teleoperation, the actuator arm's base coordinate system {Tb} can remain unchanged.
[0058] In some embodiments, previous joint information of at least one joint of the master manipulator can be received, and the previous pose of the master manipulator can be determined based on the previous joint information of at least one joint. For example, the previous pose and current pose of the master manipulator's handle can be determined based on joint information of the master manipulator read from the master manipulator's sensors at a previous time and at the current time. The position change of the master manipulator's handle can be determined based on the previous position and current position of the handle relative to the master manipulator's base coordinate system {CombX}. The attitude change of the master manipulator's handle can be determined based on the previous attitude and current attitude of the handle relative to the master manipulator's base coordinate system {CombX}.
[0059] In some embodiments, multiple control loops can be executed. In each control loop, the current pose of the master manipulator obtained in the previous control loop is determined as the previous pose of the master manipulator in the current control loop, and the target pose of the end effector of the actuator arm obtained in the previous control loop is determined as the starting pose of the end effector of the actuator arm in the current control loop. For example, for the first control loop, the initial pose of the master manipulator (e.g., the zero position of the master manipulator) can be used as the previous pose of the master manipulator in the first control loop. Similarly, the initial pose of the end effector of the actuator arm (e.g., the zero position of the actuator arm) can be used as the starting pose of the end effector of the actuator arm in the first control loop. In some embodiments, a control loop can correspond to the time interval between two frames of images acquired by the imaging unit of the endoscope.
[0060] In some embodiments, the pose change of the master manipulator can be determined based on its previous and current poses. The pose change of the end effector can be determined based on the pose change of the master manipulator and the pose relationship between the master manipulator and the end effector of the actuator arm. The target pose of the end effector can be determined based on its initial pose and the pose change of its end effector.
[0061] Positional relationships can include both positional relationships and attitude relationships. The positional relationship between the master manipulator and the end effector of the actuator arm can include the relationship between the positional change of the master manipulator and the positional change of the end effector of the actuator arm, such as being equal or proportional. The attitudeal relationship between the master manipulator and the end effector of the actuator arm can include the relationship between the attitude change of the master manipulator and the attitude change of the end effector of the actuator arm, such as being equal or proportional.
[0062] In some embodiments, method 600 further includes: determining the current position of the handle of the master operator relative to a reference coordinate system, determining the previous position of the handle relative to the reference coordinate system, determining the starting position of the end effector of the actuator relative to the actuator base coordinate system, and determining the target position of the end effector of the actuator relative to the actuator base coordinate system based on the previous and current positions of the handle relative to the reference coordinate system, the transformation relationship between the current end effector coordinate system and the actuator base coordinate system, and the starting position of the end effector of the actuator relative to the actuator base coordinate system. The transformation relationship between the current end effector coordinate system {Tt} and the actuator base coordinate system {Tb} can be determined based on the starting pose of the end effector of the actuator relative to the actuator base coordinate system {Tb}. For example, in a control loop, the previous position of the master operator in the current control loop can be determined based on the current pose of the master operator obtained in the previous control loop, or the previous position of the master operator can be determined based on the joint information of the master operator corresponding to the previous time read by the master operator sensor, and the current position of the master operator can be determined based on the joint information of the master operator corresponding to the current time read by the master operator sensor. The position change of the master manipulator is determined based on the previous and current positions of the handle relative to the reference coordinate system. The starting position of the end effector in the current control cycle can be determined based on the target pose of the end effector obtained from the previous control cycle. The position change of the end effector is determined based on the position change of the master manipulator and the pose relationship between the master manipulator and the end effector. The target position of the end effector is determined based on its starting position and the position change.
[0063] In some embodiments, method 600 further includes: determining the current orientation of the handle of the master operator relative to a reference coordinate system; determining the previous orientation of the handle relative to the reference coordinate system; determining the initial orientation of the end effector of the actuator relative to the actuator arm base coordinate system; and determining the target orientation of the end effector of the actuator arm relative to the actuator arm base coordinate system based on the previous and current orientations of the handle relative to the reference coordinate system, the transformation relationship between the current end effector coordinate system and the actuator arm base coordinate system, and the initial orientation of the end effector of the actuator arm relative to the actuator arm base coordinate system. The transformation relationship between the current end effector coordinate system {Tt} and the actuator arm base coordinate system {Tb} can be determined based on the initial orientation of the end effector of the actuator arm relative to the actuator arm base coordinate system {Tb}. For example, in a control loop, the previous orientation of the master operator in the current control loop can be determined based on the current orientation of the master operator obtained in the previous control loop, or the previous orientation of the master operator can be determined based on the joint information of the master operator corresponding to the previous time read by the master operator sensor, and the current orientation of the master operator can be determined based on the joint information of the master operator corresponding to the current time read by the master operator sensor. The attitude change of the master manipulator is determined based on the previous and current attitudes of the handle relative to the reference coordinate system. The initial attitude of the end effector in the current control cycle can be determined based on the target pose of the end effector obtained from the previous control cycle. The attitude change of the end effector is determined based on the attitude change of the master manipulator and the pose relationship between the master manipulator and the end effector. The target attitude of the end effector is determined based on its initial attitude and the attitude change.
[0064] In some embodiments, the endoscope has an imaging unit on its main body. The imaging unit coordinate system {lens} and the end-effector coordinate system {Tt} have a predetermined transformation relationship. The display coordinate system {Screen} and the imaging unit coordinate system {lens} have the same definition for the field of view direction. For example, as... Figure 2As shown, the endoscope 200 includes an endoscope body 221 and an actuator arm 210. An imaging unit is disposed on the endoscope body 221, comprising a first imaging unit 230 and a second imaging unit 240. The coordinate system {lens1} of the first imaging unit or the coordinate system {lens2} of the second imaging unit has a predetermined transformation relationship with the end-effector coordinate system {Tt} of the actuator arm. A control device (e.g., control device 500) can generate a synthetic scene image and / or a real scene image of the environment surrounding the endoscope in the first image coordinate system {img1} or the second image coordinate system {img2} based on a first image from the first imaging unit and a second image from the second imaging unit. The first image coordinate system {img1} has a predetermined transformation relationship with the first imaging unit coordinate system {lens1}, and the second image coordinate system {img2} has a predetermined transformation relationship with the second imaging unit coordinate system {lens2}, as detailed later. The display coordinate system {Screen} can coincide with either the first imaging unit coordinate system {lens1} or the second imaging unit coordinate system {lens2}, maintaining a consistent field of view. The pose change of the image on the display relative to the display coordinate system {Screen} and the pose change of the end effector relative to the execution arm base coordinate system {Tb} maintain the same amount of change but opposite direction. Thus, when the operator grips the handle of the main controller, the pose change of the image on the display perceived by the operator maintains a preset transformation relationship with the pose change of the handle of the main controller perceived by the operator.
[0065] In some embodiments, when the display shows images from different imaging units, different transformation relationships between the imaging unit coordinate system and the end effector coordinate system can be used. For example, when the display shows a first image from the first imaging unit, a predetermined transformation relationship between the first imaging unit coordinate system {lens1} and the end effector coordinate system {Tt} can be used in the control.
[0066] Continue to refer to Figure 6In step 605, a drive command for driving the end effector of the actuator arm can be generated based on the target pose. For example, a drive signal for the actuator arm can be calculated based on the target pose of the end effector relative to the actuator arm base coordinate system {Tb}. In some embodiments, the control device can send a drive signal to at least one drive device based on the target pose of the end effector of the actuator arm to control the movement of the end effector of the actuator arm. In some embodiments, the control device can determine the drive signal of at least one drive device controlling the movement of the actuator arm based on the target pose of the end effector of the actuator arm using an inverse kinematics numerical iterative algorithm of the actuator arm kinematic model. It should be understood that the kinematic model can be a mathematical model representing the kinematic relationship between the joint space and the task space of the actuator arm. For example, the kinematic model can be established by methods such as the Denavit-Hartenberg (DH) parameter method and the exponential product representation method. In some embodiments, the target pose of the end effector of the actuator arm is the target pose of the end effector of the actuator arm in the reference coordinate system. In some embodiments, step 605 can be executed by the control device 500 (e.g., drive module 520).
[0067] Continue to refer to Figure 6 In step 607, a first image is obtained from the first imaging unit. In some embodiments, the first imaging unit is configured to be located on one side of the endoscope body relative to the end-effector. As the endoscope moves within the scene, the first imaging unit continuously captures the first image in a first field of view, and the control device 500 (e.g., image processing module 530) can receive the first image from the first imaging unit. In some embodiments, the end-effector is located within the first field of view of the first imaging unit, and the first image includes an image of the end-effector captured from one side of the body.
[0068] Continue to refer to Figure 6 In step 609, a second image is obtained from the second imaging unit, wherein the fields of view of the first image and the second image are different and include an image of the endoscope. In some embodiments, the second imaging unit is configured to be located on the opposite side of the endoscope body relative to the endoscope. As the endoscope moves within the scene, the second imaging unit continuously captures second images in a second field of view different from the first field of view of the first imaging unit, and the control device 500 (e.g., image processing module 530) can receive the second images from the second imaging unit. In some embodiments, the endoscope is located within the second field of view of the second imaging unit, and the second image includes an image of the endoscope captured from the opposite side of the body.
[0069] Continue to refer to Figure 6In step 611, a synthetic scene image is generated based on the first image and the second image to remove the actual image of the end effector. In this disclosure, due to the occlusion by the end effector, neither the first imaging unit nor the second imaging unit can capture the entire scene. In some embodiments, the control device 500 (e.g., image processing module 530) can use computer vision processing to fill in the portion of the other image occluded by the end effector using either the first image or the second image, thereby generating a two-dimensional or three-dimensional synthetic scene image with the end effector removed. For example, an exemplary method for generating a synthetic scene image based on the first image and the second image may include, as shown below... Figure 12 The method 1200 is shown. In some embodiments, the computer vision processing may include a feature point detection algorithm that can extract feature points from a first image and a second image for matching, thereby achieving two-dimensional stitching of the first image and the second image. In some embodiments, the computer vision processing may include an image sequence optical flow reconstruction algorithm that can determine the depth of a pixel in scene space based on the optical flow of the pixel in the image, thereby reconstructing the scene in three dimensions. By generating a synthetic scene image, a more complete scene image, at least partially unobstructed by the end effector, can be displayed on a display device, helping the operator to observe the cavity and operating area without obstruction.
[0070] Continue to refer to Figure 6 In step 613, a virtual image of the end effector is generated in the synthesized scene image. For example, an exemplary method for generating a virtual image of the end effector in the synthesized scene image may include, for instance,... Figure 17 The method 1700 is shown. In some embodiments, the control device 500 (e.g., virtual image generation module 540) can generate a virtual image of the end-effector at the position corresponding to the end-effector in the synthesized scene image using a real-time rendering method. By generating a virtual image of the end-effector in the synthesized scene image, the actual position and size of the end-effector can be indicated to the operator without obstructing the operator's view of the scene, thereby avoiding collisions with the walls of the cavity or operating area due to the inability to see the end-effector during operation.
[0071] In some embodiments, method 600 may further include switching scene modes based on display mode instructions. Figure 9 A flowchart illustrating a method 900 for displaying scene images based on display mode instructions according to some embodiments of the present disclosure is provided. In some embodiments, some or all steps of method 900 may be performed by a robot system (e.g., Figure 1 The robot system 100 shown Figure 17 The control device (e.g., for the robot system 1700 shown) Figure 5 The control device 500 shown or Figure 17The method 900 is executed by the control device 1770 shown. The control device may include a computing device. The method 900 may be implemented by software, firmware, and / or hardware. In some embodiments, the method 900 may be implemented as computer-readable instructions. These instructions may be executed by a general-purpose processor or a special-purpose processor (e.g., a dedicated processor). Figure 17 The control device 1770 shown reads and executes these instructions. In some embodiments, these instructions may be stored on a computer-readable medium.
[0072] refer to Figure 9 In step 901, a real-scene image is generated based on the first image and / or the second image to display the actual image of the end effector. In some embodiments, the first image captured by the first imaging unit or the second image captured by the second image unit can be used as a two-dimensional real-scene image, which includes the actual image of the end effector. In some embodiments, the control device 500 (e.g., image processing module 530) can generate a three-dimensional real-scene image based on the first image or the second image using a computer vision algorithm. For example, an exemplary method for generating a three-dimensional real-scene image based on the first image or the second image may include, for example, Figure 16 Method 1600. In some embodiments, the computer vision algorithm may include an image sequence optical flow reconstruction algorithm, which can determine the depth of a pixel in the scene space based on the optical flow of a pixel in a first image or a second image, thereby reconstructing the actual scene in three dimensions. In some embodiments, two actual scene images with different fields of view may be generated simultaneously based on the first image and the second image for side-by-side display on a display device.
[0073] Continue to refer to Figure 9 In step 903, in response to a display mode instruction, a composite scene image and / or an actual scene image with a virtual image of the end effector are displayed. The display mode instruction may include, for example, at least one of a drive instruction, an end effector operation instruction, and a display mode selection instruction.
[0074] A drive command is used to drive the movement of the actuator arm to control the movement of the endoscope within the cavity. In some embodiments, method 900 further includes: generating a first display signal in response to the drive command to display at least a synthetic scene image with virtual images. In some embodiments, the drive command for driving the movement of the actuator arm may include a feed command, a retraction command, or a turning command for the endoscope. For example, the drive command may be determined based on the operator's operation on the master manipulator. The display signal generation module 550 of the control device 500 may generate a first display signal in response to the drive command. The scene output module 560 may output a synthetic scene image with at least virtual images of the end instruments to the display device for display in response to the first display signal. The drive module 520 may control the feed, retraction, or turning of the endoscope based on the drive command. By displaying at least a synthetic scene image with the end instruments removed, the operator can avoid operating the endoscope feed or turning with an incomplete field of vision, thereby avoiding unnecessary surgical risks. On the other hand, displaying at least a synthetic scene image with the end instruments removed when controlling the retraction of the endoscope from the operating area and cavity facilitates the operator's observation of the surgical effect or control of the safe withdrawal of the endoscope. In some embodiments, the target pose of the endoscope's end-effector can be determined based on the master-slave motion mapping relationship between the pose of the master manipulator and the pose of the end-effector, as well as the current pose of the master manipulator. Then, a drive command is generated based on the target pose of the end-effector, which may be, for example, a drive signal related to the target pose. In some embodiments, the method for determining the target pose of the end-effector can be implemented similarly to step 603 in method 600, and the method for generating the drive command can be implemented similarly to step 605 in method 600.
[0075] End-effector operation commands can be used to control the operation of end-effectors. In some embodiments, method 900 may further include generating a second display signal in response to the end-effector operation command to at least display an image of the actual scene. In some embodiments, the end-effector operation command may include an activation command for the end-effector, which indicates the commencement of end-effector operation. In the case of an electrocoagulation hemostasis device, the activation command may be, for example, turning on a power source. In some embodiments, the scene output module 560 of the control device 500 may, in response to the end-effector operation command, display an image of the actual scene or switch the display screen from a composite scene image of a virtual image with the end-effector to the actual scene image, thereby allowing the operator to perform surgical procedures while observing the end-effector, which helps improve operational accuracy. In some embodiments, in response to the operator completing or temporarily suspending the end-effector operation, such as by disconnecting the power source, the scene output module 560 of the control device 500 may display a composite scene image of a virtual image with the end-effector or switch the display screen from the actual scene image to the composite scene image of a virtual image with the end-effector, thereby allowing the operator to confirm the effectiveness of the surgical procedure with a full field of view. In some embodiments, end-device operation commands may have higher priority in display mode control than other drive commands other than the back command and the automatic exit command described later. This allows the operator to more intuitively see the actual situation inside the body when the end-device is triggered to begin operation.
[0076] Those skilled in the art will understand that, in this disclosure, display mode control priority refers to the higher priority of a display mode instruction when multiple display mode instructions exist simultaneously. In some embodiments, the operator can control the endoscope's feed or directional movement via feed or directional commands while issuing an end-device operation command, thereby achieving a first compound operation of controlling the endoscope's movement while activating the end-device for work. For example, when the end-device is an electrocoagulation hemostasis device, the operator can control the endoscope's feed while activating the electrocoagulation hemostasis device to achieve slight contact and compression of the tissue by the electrocoagulation hemostasis device. By prioritizing the display of actual scene images based on end-device operation commands, it can be ensured that the operator performs surgical procedures based on the actual situation inside the body. Similarly, for other compound operations involving multiple operations, the display mode can also be controlled based on the display mode control priority of various operations or operation commands.
[0077] In some embodiments, method 900 may further include controlling the endoscope to move away from the operating area based on a retraction instruction in the driving commands; and generating a first display signal to at least display a synthetic scene image with virtual images in response to the endoscope moving away from the operating area beyond a threshold. The endoscope moving away from the operating area beyond the threshold may include determining whether the distance the endoscope retracts exceeds the threshold, or determining whether the cumulative value of the position change of the master operator corresponding to the retraction instruction exceeds the threshold, the position change corresponding to the retraction instruction in the driving commands. For example, as... Figure 7 As shown in (a), the position change of the main controller corresponding to the back command can be the negative of the main controller handle coordinate system (H) in the reference coordinate system {w} (e.g., the world coordinate system). The amount of positional change in direction. For example, in multiple control cycles, a retraction command proportional to the amount of positional change of the main manipulator can be determined based on the operator's operation on the main manipulator. The drive module 520 of the control device 500 can control the endoscope to move away from the operating area based on the retraction command, and the control device 500 can obtain the amount of retraction positional change of the main manipulator in each control cycle (e.g., store it in memory) and accumulate these positional changes. When the accumulated value of the positional change exceeds a predetermined threshold, the display signal generation module 550 sends a first display signal to the scene output module 560 to display at least a synthetic scene image with virtual images, so that the operator can observe it and avoid damage to internal organs or cavities during the retraction process. This display operation for the retraction command can have a higher priority in display mode control than other drive commands (e.g., feed commands or turning commands, etc.) and end-device operation commands, except for the automatic withdrawal command described later. In this way, the display mode can be automatically adjusted in a timely manner when the operator intends to retract the endoscope, so that the operator can observe the internal condition. In some embodiments, the operator can simultaneously issue a terminal instrument operation command and control the endoscope to retract via a retraction command, achieving a second compound operation of controlling the endoscope's movement while the terminal instrument is activated for work. For example, when the terminal instrument is a clamping device, the operator can control the endoscope to retract while activating the clamping device to hold tissue, achieving slight traction and dissection of the tissue by the clamping device. By displaying a composite scene image based on the terminal instrument operation command when the endoscope is less than a threshold away from the operating area, and displaying a virtual image with the terminal instrument based primarily on the retraction command when the endoscope is more than a threshold away from the operating area, the operator can be automatically provided with the actual or complete field of view inside the body for observation. In some embodiments, method 1100 may further include terminating the terminal instrument operation command in response to the endoscope being more than a threshold away from the operating area. This prevents the terminal instrument from damaging internal organs or cavities during the retraction of the endoscope.
[0078] In some embodiments, the drive commands controlling the movement of the endoscope may further include an automatic withdrawal command, which can be used to control the endoscope to automatically withdraw from the body. In some embodiments, method 900 may further include controlling the endoscope to withdraw from the body based on the automatic withdrawal command, and generating a first display signal in response to the automatic withdrawal command to at least display a synthetic scene image with virtual images. The automatic withdrawal command allows the operator to quickly withdraw the endoscope. Similarly, the automatic withdrawal command may have higher priority than other drive commands (e.g., feed commands, steering commands, or retraction commands, etc.) and end-effector manipulation commands in display mode control. Furthermore, the automatic withdrawal command may automatically terminate other drive commands and end-effector manipulation commands that are being executed, and may disable other drive commands, end-effector manipulation commands, or display mode selection commands. Alternatively, the automatic withdrawal command may also allow the triggering of a display mode selection command during execution, so that the operator can switch display modes for better observation of the body.
[0079] Display mode selection instructions can be used to manually trigger display mode switching, for example, by manual input from an operator. In some embodiments, display mode selection instructions may include at least one of a composite scene display instruction, a real scene display instruction, and a multi-scene display instruction. Specifically, the composite scene display instruction displays a composite scene image of a virtual image with an end-effector, the real scene display instruction displays a real scene image, and the multi-scene display instruction simultaneously displays both the composite scene image of the virtual image with an end-effector and the real scene image. For example, a multi-scene display instruction may display at least a portion of the real scene image in a first window of the display device and at least a portion of the composite scene image of the virtual image with an end-effector in a second window of the display device.
[0080] In some embodiments, the display mode selection command may have a higher priority than other display mode commands, such as drive commands and end-effector operation commands, in terms of display mode control. Because the display mode selection command requires operator intervention and expresses the operator's direct display needs, it can have a higher priority during operation.
[0081] Figure 10 This diagram illustrates multi-scene display on a display device 1000 according to some embodiments of the present disclosure. For example... Figure 10As shown, in some embodiments, the display device 1000 may include a first window 1010 and a second window 1020, with the first window 1010 surrounding the second window 1020 from the outside, forming a so-called picture-in-picture display. In some embodiments, at least a portion of the actual scene image may be displayed in the first window 1010, and at least a portion of a composite scene image of a virtual image with an endoscope may be displayed in the second window 1020. For example, the scene output module 950 of the control device 900 may, in response to a multi-scene display command, simultaneously output the actual scene image and the composite scene image of a virtual image with an endoscope to the display device 1000. The display device 1000 may display a portion of the actual scene image in the first window 1010, and display, for example, the surgical site in the composite scene image of a virtual image with an endoscope in the second window 1020. This display method allows the operator to simultaneously see the environment around the endoscope and the surgical site in front of the endoscope, improving operational accuracy while suppressing discomfort caused to the operator by repeatedly switching images. It should be understood that the ways to present two scene images simultaneously on a display device include, but are not limited to, the methods described above. For example, the display device 1000 may also display the first window 1010 and the second window 1020 side by side in a split-screen manner.
[0082] This disclosure provides some embodiments of a method for generating a composite scene image based on a first image and a second image. Figure 11 A flowchart illustrating a method 1100 for generating a synthetic scene image based on a first image and a second image according to some embodiments of the present disclosure is shown. In some embodiments, some or all of the steps in method 1100 may be performed by a robotic system (e.g., Figure 1 The robot system 100 shown Figure 17 The control device (e.g., of the robot system 1700 shown) Figure 5 The control device 500 shown or Figure 17 The method 1100 is executed by the control device 1770 shown. The control device may include a computing device. The method 1100 may be implemented by software, firmware, and / or hardware. In some embodiments, the method 1100 may be implemented as computer-readable instructions. These instructions may be executed by a general-purpose processor or a special-purpose processor (e.g., a dedicated processor). Figure 17 The control device 1770 shown reads and executes these instructions. In some embodiments, these instructions may be stored on a computer-readable medium.
[0083] refer to Figure 11 In step 1101, a supplementary image is determined based on the first image or the second image. The supplementary image includes the portion of the second image or the first image that is obscured by the end-effector. In some embodiments, the supplementary image may be determined based on the first image, and this supplementary image includes the portion of the second image that is obscured by the end-effector. For example, such as...Figure 4 As shown, the first imaging unit 430 can capture a first image within a first field of view (e.g., the sum of fields of view 431 and 432). The first image includes a first environmental image (e.g., an image of the cavity wall) located within field of view 431 and an image of the end-effector 460 located within field of view 432. The first environmental image may include an image located within field of view 431' (a portion of field of view 431), which is a supplementary image used to synthesize with a second image captured by the second imaging unit 440, corresponding to the portion of the second image obscured by the end-effector 460. Similarly, in some embodiments, a supplementary image may be determined based on the second image, which includes the portion of the first image obscured by the end-effector. For example, the second imaging unit 440 can capture a second image within a second field of view (e.g., the sum of fields of view 441 and 442). The second image includes a second environmental image (e.g., an image of the cavity wall) located within field of view 441 and an image of the end-effector 460 located within field of view 442. The second environmental image may include an image located within the field of view 441' (a portion of the field of view 441), which is a supplementary image used to synthesize with the first image captured by the first imaging unit 430, corresponding to the portion of the first image that is obscured by the end effector 460.
[0084] In some embodiments, the position of the supplementary image in the first image or the position of the supplementary image in the second image can be determined based on the spatial positional relationship between the first imaging unit 430, the second imaging unit 440 and the end effector 460, thereby separating the supplementary image from the first image or the second image. The supplementary image can be used to stitch together with the second environmental image in the second image or the first environmental image in the first image to generate a stitched image.
[0085] Continue to refer to Figure 11 In step 1103, a first environmental image or a second environmental image is determined based on the first image or the second image. The first environmental image and the second environmental image do not include the image of the end-device. In some embodiments, the image of the end-device 460 can be removed from the first image or the second image based on the difference between the environmental image and the image of the end-device 460 to obtain the first environmental image in the first image or the second environmental image in the second image. For example, the image of the end-device 460 can be removed from the first image or the second image based on color features, boundary features, texture features, or spatial relationship features to generate the first environmental image or the second environmental image.
[0086] Continue to refer to Figure 11In step 1105, the first environmental image or the second environmental image and the supplementary image are stitched together to generate a stitched image. The following explanation uses the stitching of the first environmental image and the supplementary image as an example to illustrate the generation of the stitched image. It should be understood that the stitched image can also be generated by stitching the second environmental image and the supplementary image.
[0087] In some embodiments, feature points can be extracted from the first environment image and the supplementary image using a feature point detection algorithm. The feature point detection algorithm can be any one of the following: Harris (corner detection), SIFT (Scale Invariant Feature Transform), SURF (Speeded-Up Robust Features), and ORB (Oriented Fast and Rotated Brief). For example, feature points can be extracted from the edges of the first environment image and the supplementary image using a feature point detection algorithm, and a feature point database can be established based on the feature point data structure. The feature point data structure can include the feature point's position coordinates, scale, orientation, and feature vector, etc.
[0088] In some embodiments, a feature matching algorithm can be used to perform feature matching on feature points of the edges of the first environment image and the supplementary image, thereby determining the correlation between the edges of the first environment image and the edges of the supplementary image. The feature matching algorithm can be any one of the following: brute-force matching algorithm, cross-matching algorithm, KNN (k-nearest neighbor classification) matching algorithm, and RANSAC (Random Sample Consensus) matching algorithm.
[0089] In some embodiments, a registration image can be generated based on a first environmental image and / or a supplementary image. For example, the transformation relationship between the first image coordinate system {img1} and the second image coordinate system {img2} can be determined based on the spatial positional relationship between the first imaging unit 430 and the second imaging unit 440 (e.g., the transformation relationship between the first imaging unit coordinate system {lens1} and the second imaging unit coordinate system {lens2}), and the supplementary image in the second image coordinate system {img2} is transformed into an image in the first image coordinate system {img1} based on this transformation relationship, thereby generating a registration image for image fusion with the first environmental image. In some embodiments, the first environmental image and the supplementary image can also be transformed into images in the reference coordinate system {w} based on the transformation relationship between the first image coordinate system {img1}, the second image coordinate system {img2}, and the reference coordinate system {w} (e.g., the coordinate system of the end of the endoscope).
[0090] In some embodiments, a stitched image can be generated by stitching together a first environment image and a registration image. For example, the edges of the first environment image and the registration image can be aligned and stitched together based on successfully matched feature points in the first environment image and the supplementary image to generate a stitched image. In some embodiments, the generated stitched image can be a two-dimensional stitched image. In some embodiments, the two-dimensional stitched image can serve as a two-dimensional composite scene image.
[0091] In some embodiments, method 1100 may further include processing the first environment image, the second environment image, the supplementary image, or the stitched image to generate a three-dimensional synthetic scene image. Figure 12 A flowchart illustrating a method 1200 for generating a three-dimensional synthetic scene image based on a first image and a second image according to some embodiments of the present disclosure is provided. In some embodiments, some or all of the steps in method 1200 may be performed by a robotic system (e.g., Figure 1 The robot system 100 shown Figure 17 The control device (e.g., for the robot system 1700 shown) Figure 5 The control device 500 shown or Figure 17 The method 1200 is executed by the control device 1770 shown. The control device may include a computing device. The method 1200 may be implemented by software, firmware, and / or hardware. In some embodiments, the method 1200 may be implemented as computer-readable instructions. These instructions may be executed by a general-purpose processor or a special-purpose processor (e.g., a dedicated processor). Figure 17 The control device 1770 shown reads and executes these instructions. In some embodiments, these instructions may be stored on a computer-readable medium.
[0092] refer to Figure 12 In step 1201, for at least one of the first environmental image, the second environmental image, the supplementary image, or the stitched image, the optical flow field of the image is determined based on the image and the previous frame image (two consecutive frames). The optical flow field includes the optical flow of multiple pixels in the image. The following explanation uses the first environmental image in the first image as an example.
[0093] In some embodiments, as the endoscope (e.g., endoscope 420) moves within the cavity, a first imaging unit (e.g., first imaging unit 430) captures images of the cavity environment with a continuously changing field of view (corresponding to the direction of the optical axis), resulting in a plurality of first images arranged in a frame sequence. A first environment image can be determined by removing images of the end-effector (e.g., end-effector 460) from the first images. The method for determining the first environment image can be implemented similarly to step 1103 in method 1100.
[0094] Pixels in the first environmental image correspond to object points in the environment. In the sequence of the first environmental images, pixels move between adjacent frames (e.g., the previous frame and the current frame of the image), generating optical flow. This optical flow is a two-dimensional vector describing the positional changes of the pixels, corresponding to the three-dimensional motion vector of the object points in the environment, and is the projection of the three-dimensional motion vector of the object points onto the image plane. In some embodiments, the optical flow of pixels in the first environmental image can be calculated using the previous and current frames of the first environmental image. In some embodiments, by calculating the optical flow of multiple pixels in the first environmental image, the optical flow field of the first environmental image can be obtained. The optical flow field is the instantaneous velocity field generated by the movement of pixels in the first environmental image on the image plane, including the instantaneous motion vector information of the pixels, such as the direction and speed of the pixel's movement.
[0095] Continue to refer to Figure 12 In step 1203, a depth map of the image is generated based on the optical flow field of the image and the pose of the imaging unit corresponding to the image. This depth map includes the depth of object points corresponding to multiple pixels. In some embodiments, the depth value of the object point in the cavity environment corresponding to the pixel can be determined based on the optical flow of the pixel in the optical flow field of the first environment image and the pose change of the first imaging unit, thereby generating a depth map of the first environment image based on the depth of the object point. For example, an exemplary method for generating a depth map of an image based on the optical flow field and the pose of the imaging unit may include, as shown below... Figure 13 Method 1300.
[0096] Figure 13 A flowchart illustrating a method 1300 for generating a depth map based on the pose of an optical flow field and an imaging unit, according to some embodiments of the present disclosure, is provided. In some embodiments, some or all of the steps in method 1300 may be performed by a robotic system (e.g., Figure 1 The robot system 100 shown Figure 17 The control device (e.g., for the robot system 1700 shown) Figure 5 The control device 500 shown or Figure 17 The method 1300 is executed by the control device 1770 shown. The control device may include a computing device. The method 1300 may be implemented by software, firmware, and / or hardware. In some embodiments, the method 1300 may be implemented as computer-readable instructions. These instructions may be executed by a general-purpose processor or a special-purpose processor (e.g., a dedicated processor). Figure 17 The control device 1770 shown reads and executes these instructions. In some embodiments, these instructions may be stored on a computer-readable medium.
[0097] refer to Figure 13In step 1301, the focus of the optical flow field is determined based on the optical flow field of the image. For example, when generating the optical flow field based on a first environmental image, during endoscope advance or retraction, the optical flows of multiple pixels in the first environmental image are not parallel to each other, and the extensions of the optical flow vectors converge at the focus of the optical flow field, which is a fixed point in the optical flow field. In some embodiments, the optical flow vectors in the optical flow field may have a corresponding relationship with the focus, and each optical flow vector may converge to a different focus. In some embodiments, such as during endoscope advance, the focus of the optical flow field may include an expansion focus (FOE), which is the convergence point of the optical flow vectors extending in the opposite direction. In some embodiments, such as during endoscope retraction, the focus of the optical flow field may include a contraction focus (FOC), which is the convergence point of the optical flow vectors extending in the forward direction.
[0098] Continue to refer to Figure 13 In step 1303, based on the focal point of the optical flow field, the distances between multiple pixels and the focal point are determined. For example, in the first environmental image, the distances between multiple pixels and the focal point can be determined in the first image coordinate system.
[0099] Continue to refer to Figure 13 In step 1305, based on the optical flow field of the image, the velocities of multiple pixels in the optical flow field are determined. The velocity of the optical flow of a pixel can be the ratio between the distance the pixel moves in the optical flow field (the length of the optical flow) and the time interval between two consecutive frames. In some embodiments, in the first environmental image, the distances that multiple pixels move in the optical flow field can be determined in the first image coordinate system. In some embodiments, the time interval between two consecutive frames (the time of each frame) can be, for example, 1 / 60 of a second, but is not limited thereto and can be appropriately adjusted according to imaging requirements.
[0100] Continue to refer to Figure 13In step 1307, the velocity of the imaging unit is determined based on the pose of the imaging unit corresponding to the image. In some embodiments, the pose of the imaging unit can be determined based on the target pose of the end of the actuator arm and the pose relationship between the end of the actuator arm and the imaging unit. For example, the pose of the first imaging unit can be determined based on the target pose of the end of the actuator arm and the pose transformation relationship between the first imaging unit coordinate system {lens1} and the end of the actuator arm coordinate system {Tt}. In some embodiments, the method for determining the target pose of the end of the actuator arm can be implemented similarly to step 603 in method 600. In some embodiments, the previous pose and current pose of the first imaging unit can be determined based on the initial pose and target pose of the end of the actuator arm in a control loop. The previous pose of the first imaging unit corresponds to the previous frame of the image and can be obtained based on the initial pose of the end of the actuator arm in a control loop, and the current pose of the first imaging unit corresponds to the current frame of the image and can be obtained based on the target pose of the end of the actuator arm in the same control loop. Based on the previous pose and current pose of the first imaging unit, the distance the first imaging unit moves can be determined, and thus the velocity of the first imaging unit can be determined based on the distance the first imaging unit moves and the time interval between two consecutive frames.
[0101] Continue to refer to Figure 13 In step 1309, a depth map of the image is determined based on the distances between multiple pixels and the focal point, the velocities of the multiple pixels in the optical flow field, and the velocity of the imaging unit. In some embodiments, in the optical flow field generated from the first environment image, the depth value (depth information) of an object point can be determined based on the distance between a pixel and the focal point, the velocity of that pixel in the optical flow field, and the velocity of the first imaging unit. The depth value of the object point can be the distance between the object point and the image plane of the first imaging unit. By calculating the depth value of the object point for each pixel in the first environment image, a depth map of the first environment image can be obtained.
[0102] In some embodiments, the distances of multiple pixels moving within the optical flow field can be determined based on the optical flow field of the image, and the distances of the imaging unit moving can be determined based on the pose of the imaging unit corresponding to the image. Thus, the depth map of the image is determined based on the distances between the multiple pixels and the focal point, the distances of the multiple pixels moving within the optical flow field, and the distances of the imaging unit moving.
[0103] Continue to refer to Figure 12In step 1205, based on the depth map of the image and the pixel coordinates of multiple pixels, the spatial coordinates of the object points corresponding to the multiple pixels are determined. In some embodiments, the spatial coordinates of the object points can be determined based on the pixel coordinates of the pixels in the first environment image in the first image coordinate system and the depth value of the object points corresponding to the pixels, thereby realizing the transformation of the two-dimensional pixels in the first environment image to three-dimensional coordinates.
[0104] Continue to refer to Figure 12 In step 1207, color information of multiple pixels is obtained based on the image. In some embodiments, a color feature extraction algorithm can be used to extract the color information of pixels in the first environmental image. The color feature extraction algorithm can be any of the following methods: color histogram, color set, color moment, color aggregation vector, etc.
[0105] Continue to refer to Figure 12 In step 1209, point cloud fusion is performed on the image based on the color information of multiple pixels and the spatial coordinates of object points to generate a three-dimensional point cloud. In some embodiments, the spatial coordinates of object points can be transformed based on the intrinsic parameter matrix of the first imaging unit, and point cloud fusion is performed on the first environmental image based on the color information of the pixels to generate a three-dimensional point cloud, which includes the three-dimensional spatial coordinates and color information of the object points. In some embodiments, the intrinsic parameters of the first imaging unit can be known or obtained through calibration.
[0106] In some embodiments, method 1200 can be used to process a first environment image or a second environment image and a supplementary image to generate a three-dimensional point cloud. For example, in some embodiments, feature extraction and stereo matching can be performed on the three-dimensional point clouds of the first environment image and the supplementary image to stitch the first environment image and the supplementary image together, thereby generating a three-dimensional stitched image. In some embodiments, method 1200 can also be used to process a registration image generated based on the supplementary image to generate a three-dimensional point cloud, and feature extraction and stereo matching can be performed on the three-dimensional point clouds of the first environment image and the registration image to stitch together and generate a three-dimensional stitched image. It should be understood that a three-dimensional stitched image can also be generated by stitching together a second environment image in a second image and its corresponding supplementary image.
[0107] In some embodiments, a two-dimensional or three-dimensional stitched image can be used as a composite scene image. In some embodiments, a two-dimensional composite scene image can also be processed to generate a three-dimensional composite scene image. For example, method 1200 can be used to process the stitched image generated in method 1100 to achieve the conversion of the composite scene image from two-dimensional to three-dimensional.
[0108] In some embodiments, method 1100 may further include generating a three-dimensional real-scene image based on at least one of a first image or a second image. Figure 14 A flowchart illustrating a method 1400 for generating a three-dimensional real-world scene image based on a first image and / or a second image according to some embodiments of the present disclosure is provided. In some embodiments, some or all steps of method 1400 may be performed by a robotic system (e.g., Figure 1 The robot system 100 shown Figure 17 The control device (e.g., for the robot system 1700 shown) Figure 5 The control device 500 shown or Figure 17 The method 1400 is executed by the control device 1770 shown. The control device may include a computing device. The method 1400 may be implemented by software, firmware, and / or hardware. In some embodiments, the method 1400 may be implemented as computer-readable instructions. These instructions may be executed by a general-purpose processor or a special-purpose processor (e.g., a dedicated processor). Figure 17 The control device 1770 shown reads and executes these instructions. In some embodiments, these instructions may be stored on a computer-readable medium.
[0109] refer to Figure 14 In step 1401, for either the first or second image, the optical flow field of the image is determined based on the image and its previous frame. The optical flow field includes the optical flow of multiple pixels in the image. In some embodiments, step 1401 can be implemented similarly to step 1201 in method 1200.
[0110] Continue to refer to Figure 14 In step 1403, a depth map of the image is generated based on the optical flow field of the image and the pose of the imaging unit corresponding to the image. The depth map includes the depth of object points corresponding to multiple pixels. In some embodiments, step 1403 can be implemented similarly to step 1203 in method 1200.
[0111] Continue to refer to Figure 14 In step 1405, based on the depth map of the image and the pixel coordinates of multiple pixels, the object point spatial coordinates corresponding to the multiple pixels are determined. In some embodiments, step 1405 can be implemented similarly to step 1205 in method 1200.
[0112] Continue to refer to Figure 14 In step 1407, color information of multiple pixels is obtained based on the image. In some embodiments, step 1407 can be implemented similarly to step 1207 in method 1200.
[0113] Continue to refer to Figure 14In step 1409, point cloud fusion is performed on the image based on the color information of multiple pixels and the spatial coordinates of object points to generate a three-dimensional point cloud. In some embodiments, step 1409 can be implemented similarly to step 1209 in method 1200.
[0114] This disclosure provides some embodiments of a method for generating a virtual image of an end-effector in a synthetic scene image. Figure 15 A flowchart illustrating a method 1500 for generating a virtual image of an end effector in a synthetic scene image according to some embodiments of the present disclosure is provided. In some embodiments, some or all of the steps in method 1500 may be performed by a robotic system (e.g., Figure 1 The robot system 100 shown Figure 17 The control device (e.g., for the robot system 1700 shown) Figure 5 The control device 500 shown or Figure 17 The method 1500 is executed by the control device 1770 shown. The control device may include a computing device. The method 1500 may be implemented by software, firmware, and / or hardware. In some embodiments, the method 1500 may be implemented as computer-readable instructions. These instructions may be executed by a general-purpose processor or a special-purpose processor (e.g., a dedicated processor). Figure 17 The control device 1770 shown reads and executes these instructions. In some embodiments, these instructions may be stored on a computer-readable medium.
[0115] refer to Figure 15 In step 1501, the position and size of the end effector in the synthesized scene image are determined. In some embodiments, the position and size of the end effector in the synthesized scene image can be determined based on the inherent parameters of the end effector, which may include the positional parameters of the end effector on the subject (e.g., the relative positional relationship with the first imaging unit and the second imaging unit), orientation parameters, and size parameters. For example, the inherent parameters of the end effector may be known or obtained through calibration. In some embodiments, the position and size of the end effector in the synthesized scene image may also be determined based on the edges of the first environment image, the second environment image, or the supplementary image.
[0116] Continue to refer to Figure 15 In step 1503, a virtual image of the end effector is generated in the composite scene image. In some embodiments, the virtual image of the end effector can be generated in the composite scene image through real-time rendering. For example, the virtual image of the end effector can be generated for each frame of the composite scene image. In some embodiments, the virtual image of the end effector may include outlines and / or transparent entities to indicate the end effector. This allows the position and size of the end effector to be shown without obstructing the operator's view.
[0117] In some embodiments, method 6000 may further include generating a first virtual ruler for indicating distance along the axial direction of a virtual image of the end-effector in the synthesized scene image. This first virtual ruler may be generated along the contour of the virtual image of the end-effector to indicate the length of the end-effector, helping to improve operator precision. In some embodiments, the method for generating the first virtual ruler may be implemented similarly to step 1503 of method 1500.
[0118] In some embodiments, method 6000 may further include determining the distance between the distal end of the instrument and the surgical site, and updating a first virtual ruler to a second virtual ruler based on the distance between the distal end of the instrument and the surgical site. For example, the distance between the distal end of the instrument and the surgical site can be measured by a ranging unit on the instrument, and a second virtual ruler can be generated based on that distance. The second virtual ruler may include information indicated by the first virtual ruler and distance information between the distal end of the instrument and the surgical site. By updating the first virtual ruler to the second virtual ruler, the length of the distal instrument and the distance between the distal instrument and the surgical site can be shown simultaneously, which helps to further improve the operator's operational accuracy. In some embodiments, the method for generating the second virtual ruler can be implemented similarly to step 1503 in method 1500.
[0119] In some embodiments of this disclosure, a computer device is also provided, including a memory and a processor. The memory may be used to store at least one instruction, and the processor is coupled to the memory for executing the at least one instruction to perform some or all of the steps in the method of this disclosure, such as... Figure 6 , Figure 9 , Figures 11-15 Some or all of the steps in the method disclosed herein.
[0120] Figure 16 A schematic block diagram of a computer device 1600 according to some embodiments of the present disclosure is shown. See also Figure 16 The computer device 1600 may include a central processing unit (CPU) 1601, a system memory 1604 including random access memory (RAM) 1602 and read-only memory (ROM) 1603, and a system bus 1605 connecting the various components. The computer device 1600 may also include an input / output system and a mass storage device 1607 for storing the operating system 1613, application programs 1614, and other program modules 1615. The input / output devices include an input / output control unit 1610 mainly composed of a display 1608 and input devices 1609.
[0121] Mass storage device 1607 is connected to central processing unit 1601 via a mass storage control device (not shown) connected to system bus 1605. Mass storage device 1607 or computer-readable media provides non-volatile storage for computer devices. Mass storage device 1607 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drives.
[0122] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, flash memory or other solid-state storage technologies, CD-ROM, or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The aforementioned system memories and mass storage devices can be collectively referred to as memory.
[0123] Computer device 1600 can be connected to network 1612 via network interface unit 1611 connected to system bus 1605.
[0124] The system memory 1604 or mass storage device 1607 is also used to store one or more instructions. The central processing unit 1601 implements all or part of the steps of the methods in some embodiments of this disclosure by executing the one or more instructions.
[0125] In some embodiments of this disclosure, a computer-readable storage medium is also provided, storing at least one instruction that is executed by a processor to cause a computer to perform some or all of the steps in the methods of some embodiments of this disclosure, such as... Figure 6 , Figure 9 , Figures 11-15 Some or all of the steps in the disclosed method. Examples of computer-readable storage media include memory for computer programs (instructions), such as read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage devices.
[0126] Figure 17A schematic diagram of a robot system 1700 according to some embodiments of the present disclosure is shown. In some embodiments of the present disclosure, see [link to schematic diagram]. Figure 17 The robot system 1700 includes: a master manipulator 1710, a control device 1770, a drive device 1780, a slave tool 1720, and a display device 1790. The master manipulator 1710 includes a robotic arm, a handle mounted on the robotic arm, and at least one master manipulator sensor mounted at at least one joint on the robotic arm. The at least one master manipulator sensor is used to obtain joint information of at least one joint. In some embodiments, the master manipulator 1710 includes a six-degree-of-freedom robotic arm, with a master manipulator sensor mounted at each joint to generate joint information (such as joint angle data). In some embodiments, the master manipulator sensor employs a potentiometer and / or an encoder. The slave tool 1720 is equipped with an endoscope 1730. In some embodiments, the slave tool 1720 includes a motion arm 1740, and the endoscope 1730 may be mounted at the distal end of the motion arm 1740. In some embodiments, the endoscope 1730 includes an actuator arm 1750 and an endoscope body 1760 disposed at the end of the actuator arm 1750. A first imaging unit 1761, a second imaging unit 1762, and a distal end-effector 1765 are disposed on the endoscope body 1760. The first imaging unit 1761 is used to capture a first image, the second imaging unit 1762 is used to capture a second image, and the distal end-effector 1765 is used to perform surgical procedures. A display device 1790 is used to display the images output by the endoscope 1730. A control device 1770 is configured to connect to the actuator arm 1740 and the drive device 1780 to control the movement of the endoscope 1730, and is communicatively connected to the endoscope 1730 to process the images output by the endoscope 1730. The control device 1770 is used to perform some or all of the steps in the methods of some embodiments of this disclosure, such as... Figure 6 , Figure 9 , Figures 11-15 Some or all of the steps in the method disclosed herein.
[0127] Note that the above are merely exemplary embodiments and technical principles of this disclosure. Those skilled in the art will understand that this disclosure is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this disclosure. Therefore, although this disclosure has been described in detail through the above embodiments, this disclosure is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this disclosure, the scope of which is determined by the scope of the appended claims.
Claims
1. A robot system, characterized in that, include: The master manipulator includes a robotic arm, a handle disposed on the robotic arm, and at least one master manipulator sensor disposed at at least one joint on the robotic arm, wherein the at least one master manipulator sensor is used to obtain joint information of the at least one joint; An endoscope includes an actuating arm, a body disposed at the end of the actuating arm, a first imaging unit, a second imaging unit, and an end-effector extending from the distal end of the body, wherein the first imaging unit is used to capture a first image, the second imaging unit is used to capture a second image, the first image and the second image have different fields of view and include an image of the end-effector; At least one drive device for driving the actuator arm; At least one drive device sensor, coupled to the at least one drive device and used to obtain state information of the at least one drive device; A display device for displaying images; and A control device is configured to communicate with the main manipulator, the at least one drive device, the at least one drive device sensor, the endoscope, and the display device. The control device is configured to determine the current pose of the main manipulator, determine the target pose of the end effector based on the current pose of the main manipulator and the pose relationship between the main manipulator and the end effector of the actuator, generate a drive command for driving the end effector based on the target pose, obtain a first image from the first imaging unit, obtain a second image from the second imaging unit, generate a composite scene image based on the first image and the second image to remove the actual image of the end effector, and generate a virtual image of the end effector in the composite scene image. The control device is also configured to: Based on the first image and / or the second image, generate an actual scene image to display the actual image of the end effector; and In response to a display mode command, the synthetic scene image with the virtual image and / or the actual scene image are displayed; The display mode instructions include the drive instructions and the end-device operation instructions for controlling the operation of the end-device. The drive instructions include a feed instruction for controlling the endoscope's advance or a steering instruction for controlling the endoscope's rotation. The control device is further configured to: In response to the feed command or the steering command, a first display signal is generated to at least display the synthetic scene image with the virtual image; and In response to an end-device operation command, a second display signal is generated to at least display the actual scene image; The drive command also includes a retraction command for controlling the retraction of the endoscope, and the control device is further configured to: Based on the retraction command, the endoscope is controlled to move away from the operating area; and In response to the endoscope moving away from the operating area beyond a threshold, the first display signal is generated to at least display the synthetic scene image with the virtual image; The display mode control priority of the display mode instruction includes: The display mode control priority of the back movement command is higher than the display mode control priority of the end-effector operation command; and The display mode control priority of the end-effector operation command is higher than that of the feed command and the steering command.
2. The robot system according to claim 1, characterized in that, The display mode instruction also includes a display mode selection instruction for selecting a display mode, and the control device is further configured to generate a display signal corresponding to the selected display mode in response to the display mode selection instruction; The display mode control priority of the display mode instruction also includes: The display mode selection instruction has a higher display mode control priority than the drive instruction and the end-device operation instruction.
3. The robot system according to claim 1, characterized in that, The control device is also configured to: Determine whether the distance the endoscope retracts exceeds a threshold; or Determine whether the cumulative value of the position change of the main operator corresponding to the back instruction exceeds a threshold.
4. The robot system according to claim 1, characterized in that, The drive command also includes an automatic exit command, and the control device is further configured to: Based on the aforementioned automatic exit command, the endoscope is controlled to exit the body; and In response to the automatic exit command, the first display signal is generated to at least display the synthetic scene image with the virtual image; The display mode control priority of the display mode instruction also includes: The display mode control priority of the automatic exit command is higher than the display mode control priority of the end-device operation command.
5. The robot system according to claim 2, characterized in that, The display mode selection instruction includes at least one of the following: A composite scene display instruction is used to display the composite scene image containing the virtual image; The actual scene display command is used to display the actual scene image; or A multi-scene display instruction is used to display at least a portion of the actual scene image in a first window and at least a portion of the composite scene image with the virtual image in a second window.
6. The robot system according to claim 1, characterized in that, The pose relationship includes at least one of the following: The position change at the end of the actuator arm is equal to or proportional to the position change of the master manipulator; and / or The change in attitude at the end of the actuator arm is equal to or proportional to the change in attitude of the master manipulator.
7. The robot system according to claim 1, characterized in that, The control device is also configured to: Determine the current pose of the main operator's handle relative to the reference coordinate system; Determine the previous pose of the handle relative to the reference coordinate system; Determine the initial pose of the end effector relative to the actuator base coordinate system; as well as Based on the previous and current poses of the handle relative to the reference coordinate system, the transformation relationship between the current end coordinate system of the actuator arm and the base coordinate system of the actuator arm, and the initial pose of the end of the actuator arm relative to the base coordinate system of the actuator arm, the target pose of the end of the actuator arm relative to the base coordinate system of the actuator arm is determined.
8. The robot system according to claim 1, characterized in that, The control device is also configured to: Execute multiple control loops, where in each control loop: The current pose of the master operator obtained in the previous control cycle is determined as the previous pose of the master operator in the current control cycle; as well as The target pose of the end of the actuator arm obtained in the previous control cycle is determined as the starting pose of the end of the actuator arm in the current control cycle.
9. The robot system according to claim 1, characterized in that, The control device is also configured to: Based on the first image or the second image, a supplementary image is determined, the supplementary image including the portion of the second image or the first image that is obscured by the end device; Based on the first image or the second image, a first environmental image or a second environmental image is determined, wherein the first environmental image and the second environmental image do not include the image of the end device; as well as The first environmental image or the second environmental image and the supplementary image are stitched together to generate a stitched image.
10. The robot system according to claim 9, characterized in that, The control device is further configured to: for at least one of the first image, the second image, the first environment image, the second environment image, the supplementary image, or the stitched image, Based on the image and the previous frame of the image, the optical flow field of the image is determined, and the optical flow field includes the optical flow of multiple pixels in the image; Based on the optical flow field of the image and the pose of the imaging unit corresponding to the image, a depth map of the image is generated, and the depth map includes the depth of the object points corresponding to the plurality of pixels; Based on the depth map of the image and the pixel coordinates of the plurality of pixels, determine the object point spatial coordinates of the object point corresponding to the plurality of pixels; Based on the image, obtain the color information of the plurality of pixels; as well as The image is fused using the color information of the multiple pixels and the spatial coordinates of the object points to generate a three-dimensional point cloud.
11. The robot system according to claim 10, characterized in that, The control device is also configured to: Based on the optical flow field of the image, determine the focal point of the optical flow field; Based on the focal point of the optical flow field, the distance between the plurality of pixels and the focal point is determined; Based on the optical flow field of the image, the velocity of the plurality of pixels in the optical flow field is determined; The velocity of the imaging unit is determined based on the pose of the imaging unit corresponding to the image. as well as The depth map of the image is determined based on the distance between the plurality of pixels and the focal point, the velocity of the plurality of pixels in the optical flow field, and the velocity of the imaging unit.
12. The robot system according to claim 10 or 11, characterized in that, The control device is also configured to: The pose of the imaging unit is determined based on the target pose of the end of the actuator arm and the pose relationship between the end of the actuator arm and the imaging unit.
13. The robot system according to claim 1, characterized in that, The control device is also configured to: Determine the position and size of the end effector in the synthesized scene image; and A virtual image of the end effector is generated in the synthesized scene image.
14. The robot system according to claim 13, characterized in that, The virtual image of the end effector includes outlines and / or transparent entities to show the end effector; and / or In the synthesized scene image, a first virtual ruler for indicating distance is generated along the axial direction of the virtual image of the end effector.
15. The robot system according to claim 14, characterized in that, The control device is also configured to: Determine the distance between the endoscope and the surgical site; and Based on the distance between the endoscope and the surgical site, the first virtual ruler is updated to the second virtual ruler.
16. A computer device, comprising: Memory, used to store at least one instruction; as well as A processor, coupled to the memory, is configured to perform the following steps: Determine the current pose of the main manipulator of the surgical robot system; Based on the current pose of the main manipulator and the pose relationship between the main manipulator and the end of the endoscope's actuator arm, the target pose of the end of the actuator arm is determined. Based on the target pose, drive commands are generated for driving the endoscope of the surgical robot system. The endoscope includes the actuator arm, a main body disposed at the end of the actuator arm, a first imaging unit, a second imaging unit, and an end instrument extending from the distal end of the main body. The drive commands are used to drive the end of the actuator arm. Obtain a first image from the first imaging unit; A second image is obtained from the second imaging unit, wherein the fields of view of the first image and the second image are different and include an image of the endoscope's end instrument; Based on the first image and the second image, a synthetic scene image is generated to remove the actual image of the end effector; A virtual image of the end effector is generated in the synthesized scene image; Based on the first image and / or the second image, generate an actual scene image to display the actual image of the end effector; and In response to a display mode command, the synthetic scene image with the virtual image and / or the actual scene image are displayed; The display mode command includes the drive command and the end-device operation command for controlling the operation of the end-device. The drive command includes a feed command for controlling the endoscope feed or a steering command for controlling the endoscope rotation. The display of the synthetic scene image with the virtual image and / or the actual scene image in response to the display mode command further includes: In response to the feed command or the steering command, a first display signal is generated to at least display the synthetic scene image with the virtual image; and In response to an end-device operation command, a second display signal is generated to at least display the actual scene image; The driving command further includes a retraction command for controlling the retraction of the endoscope, and the display of the synthetic scene image with the virtual image and / or the actual scene image in response to the display mode command further includes: Based on the retraction command, the endoscope is controlled to move away from the operating area; and In response to the endoscope moving away from the operating area beyond a threshold, the first display signal is generated to at least display the synthetic scene image with the virtual image; The display mode control priority of the display mode instruction includes: The display mode control priority of the back movement command is higher than the display mode control priority of the end-effector operation command; and The display mode control priority of the end-effector operation command is higher than that of the feed command and the steering command.
17. A computer-readable storage medium for storing at least one instruction, which, when executed by a computer, causes the computer to perform the following steps: Determine the current pose of the main manipulator of the surgical robot system; Based on the current pose of the main manipulator and the pose relationship between the main manipulator and the end of the endoscope's actuator arm, the target pose of the end of the actuator arm is determined. Based on the target pose, drive commands are generated for driving the endoscope of the surgical robot system. The endoscope includes the actuator arm, a main body disposed at the end of the actuator arm, a first imaging unit, a second imaging unit, and an end instrument extending from the distal end of the main body. The drive commands are used to drive the end of the actuator arm. Obtain a first image from the first imaging unit; A second image is obtained from the second imaging unit, wherein the fields of view of the first image and the second image are different and include an image of the endoscope's end instrument; Based on the first image and the second image, a synthetic scene image is generated to remove the actual image of the end effector; A virtual image of the end effector is generated in the synthesized scene image; Based on the first image and / or the second image, generate an actual scene image to display the actual image of the end effector; and In response to a display mode command, the synthetic scene image with the virtual image and / or the actual scene image are displayed; The display mode command includes the drive command and the end-device operation command for controlling the operation of the end-device. The drive command includes a feed command for controlling the endoscope feed or a steering command for controlling the endoscope rotation. The display of the synthetic scene image with the virtual image and / or the actual scene image in response to the display mode command further includes: In response to the feed command or the steering command, a first display signal is generated to at least display the synthetic scene image with the virtual image; and In response to an end-device operation command, a second display signal is generated to at least display the actual scene image; The driving command further includes a retraction command for controlling the retraction of the endoscope, and the display of the synthetic scene image with the virtual image and / or the actual scene image in response to the display mode command further includes: Based on the retraction command, the endoscope is controlled to move away from the operating area; and In response to the endoscope moving away from the operating area beyond a threshold, the first display signal is generated to at least display the synthetic scene image with the virtual image; The display mode control priority of the display mode instruction includes: The display mode control priority of the back movement command is higher than the display mode control priority of the end-effector operation command; and The display mode control priority of the end-effector operation command is higher than that of the feed command and the steering command.
Citation Information
Patent Citations
Endoscope with dual image sensors
CN113453606A
Master-slave motion control method, robot system, equipment and storage medium
CN113876434A
Tool localization system with image enhancement and method of operation thereof
US20150170381A1