A sleeve pose recognition method and system based on binocular vision

By using a binocular vision-based cannula pose recognition method, and utilizing a binocular camera system to recognize and compensate for eye movements, the high complexity and insufficient accuracy caused by reliance on external markers in existing technologies are solved. This achieves efficient and accurate cannula pose recognition and eye movement compensation, thereby improving the safety and stability of robot-assisted surgery.

CN120182360BActive Publication Date: 2025-11-21BEIJING XIANWEI MEDICAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510661100.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-11-21
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

In existing technologies, robot-assisted intraocular surgery systems rely heavily on external markers when identifying scleral cannulas, resulting in high complexity of surgical scenarios and insufficient real-time performance and estimation accuracy, making it difficult to meet the requirements of high timeliness and high stability.

Method used

A binocular vision-based cannula pose recognition method is adopted. By building a binocular camera system, binocular images are acquired, cannula and cannula hole are identified, disparity map is generated and converted into depth map, and boundary point cloud is used to determine the position and orientation of cannula, reducing the dependence on external markers.

Benefits of technology

It improves the accuracy and speed of cannula pose recognition, can actively compensate for eye movements, enhances the safety and stability of surgery, and simplifies the operation complexity of aligning instruments with the scleral cannula.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182360B_ABST
    Figure CN120182360B_ABST
Patent Text Reader

Abstract

The application relates to the field of medical devices and provides a cannula pose recognition method and system based on binocular vision. First, a binocular camera system covering the shooting range of eyeballs is built, and a group of binocular images are collected, and a cannula and a cannula hole in the two images are respectively recognized. Then, a parallax map is generated based on the cannula and the cannula hole in the two images, the parallax map is converted into a depth map in combination with the internal parameter of the binocular camera system, and the boundary point cloud of the cannula hole is obtained. Finally, the cannula position and the fitting plane of the cannula hole are determined according to the boundary point cloud, and the normal vector is taken as the cannula pose. The application can reduce the operation complexity of instrument alignment to the cannula in robot-assisted surgery, does not need to additionally set a marker, effectively improves the precision and speed of cannula pose recognition, directly recognizes eyeball movement through cannula pose change, helps to enhance the compensation effect of eyeball movement, and improves the safety of robot-assisted surgery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical devices, specifically relating to a method and system for cannula pose recognition based on binocular vision. Background Technology

[0002] Intraocular surgery is a highly delicate procedure primarily used in the treatment of ophthalmic diseases such as retinal repair and vitrectomy. Due to the intricate and fragile nature of the intraocular tissues, the precision and stability required for the procedure are extremely high, thus demanding a high level of skill from the surgeon. In recent years, the development of robot-assisted surgical systems has enabled them to offer high-precision and stable motion control, helping to reduce the difficulty of surgical procedures and improve safety and success rates.

[0003] In intraocular surgery, surgeons typically use specialized instruments to fix a scleral cannula to the sclera of the patient's eyeball before inserting surgical instruments into the eye. While robot-assisted surgical systems can effectively reduce the difficulty of intraocular procedures, they introduce new challenges. For example, the robot-assisted system requires precise alignment with the scleral cannula fixed to the surface of the eyeball to smoothly insert surgical instruments; this step is complex and time-consuming, limiting surgical efficiency.

[0004] Furthermore, during surgery, patients may experience slight displacement of their eyeballs due to spontaneous breathing, minor muscle movements, or other uncontrolled physiological factors, resulting in changes in the relative position between the surgical instruments and the eyeball. If such displacements are not compensated for in real time, they may affect surgical precision or even cause tissue damage. Therefore, to ensure surgical safety, real-time compensation for eyeball movements is necessary.

[0005] Existing methods for compensating for eye movement are mainly divided into two categories: passive compensation and active compensation. Passive compensation methods typically reduce relative movement between the eye and the instrument through physical fixation. For example, studies have used methods such as fixing the patient's head and eyeball, or directly mounting the surgical robot to the patient's head to suppress relative movement. However, while these methods can alleviate displacement between the instrument and the eye to some extent, they have several limitations. On the one hand, physical fixation may cause patient discomfort; on the other hand, due to the compliance of the periocular soft tissues, passive fixation cannot completely restrict the subtle movements of the eyeball within the orbit, thus affecting operational stability.

[0006] In contrast, active compensation methods use external sensors to sense eye movements in real time and feed them back to the robotic system, driving the robot to make dynamic adjustments to achieve real-time compensation for eye displacement. Existing active compensation technologies mainly rely on visual sensing systems to identify and track scleral cannulas or markers. For example, some literature proposes using a binocular miniature camera and an inertial measurement unit (IMU) to track ArUco markers on both sides of the scleral cannulas to estimate the cannulas pose; another method uses a YOLO network to perform target detection on the cannulas image and feeds the detected area into an SC6D network to perform pose estimation in monocular RGB images. These methods have good active response capabilities and can adapt to complex eye movement scenarios. However, due to their high dependence on external markers, they increase the complexity of the surgical scenario, and still have shortcomings in real-time performance and estimation accuracy, making it difficult to meet the requirements of high-timeliness and high-stability intraocular surgical applications. Summary of the Invention

[0007] This invention provides a binocular vision-based method and system for cannula pose recognition, which solves the problems of existing technologies that rely heavily on external markers, resulting in high complexity of surgical scenarios and insufficient real-time performance and estimation accuracy.

[0008] To address the aforementioned technical problems, the present invention discloses the following technical solutions:

[0009] One aspect of the present invention provides a binocular vision-based cannula pose recognition method for determining the pose of the cannula on the eyeball, comprising:

[0010] Build a binocular camera system whose shooting range covers the eyeball;

[0011] Acquire a set of binocular images captured by a binocular camera system, wherein the binocular images include a first image and a second image;

[0012] Identify the sleeve and sleeve hole in the two images respectively;

[0013] Generate a disparity map based on the sleeve and sleeve hole in the two images;

[0014] By combining the intrinsic parameters of the binocular camera system, the disparity map is converted into a depth map, and the boundary point cloud of the casing hole is obtained.

[0015] The sleeve position and the fitting plane of the sleeve hole are determined based on the boundary point cloud, and the normal vector of the fitting plane is used as the sleeve attitude.

[0016] Optionally, the binocular camera system with a shooting range covering the eyeball includes:

[0017] Configure the binocular camera according to the preset application scenario;

[0018] The intrinsic and extrinsic parameters of the stereo camera are calibrated using a preset calibration method and a calibration board.

[0019] Optionally, after performing the step of acquiring a set of binocular images captured by the binocular camera system, the method further includes:

[0020] Epipolar correction is performed on the first and second images to ensure that the preset corresponding points in the two images are on the same horizontal line.

[0021] Optionally, identifying the sleeve and sleeve hole in the two images respectively includes:

[0022] A pre-trained target detection network was used to detect the sleeve region in the first and second images, respectively.

[0023] The cannula region is cropped out from the first image as a first sub-image, and the cannula region is cropped out from the second image as a second sub-image;

[0024] A pre-trained semantic segmentation network was used to identify the sleeve and sleeve hole in the first and second sub-images, respectively.

[0025] Optionally, generating a disparity map based on the sleeve and sleeve aperture in the two images includes:

[0026] The first and second sub-images are input into a pre-trained disparity estimation network, and the identified sleeves and sleeve holes are used as mask regions to output a disparity map of the mask regions, with the first sub-image as the reference image.

[0027] Optionally, the intrinsic parameters of the binocular camera system are used to convert the disparity map into a depth map and obtain the boundary point cloud of the casing hole, including:

[0028] The correlation between parallax and depth was obtained based on the intrinsic parameters of a binocular camera system.

[0029] The disparity map is converted into a depth map using the aforementioned correlation.

[0030] Map each pixel in the depth map to a 3D spatial point, as a sleeve point cloud;

[0031] The pixels corresponding to the edge of the sleeve hole in the first sub-image are determined, and the three-dimensional spatial points corresponding to the pixels are extracted from the sleeve point cloud as the boundary point cloud of the sleeve hole.

[0032] Optionally, determining the casing position and the fitting plane of the casing hole based on the boundary point cloud, and using the normal vector of the fitting plane as the casing attitude, includes:

[0033] Calculate the average coordinates of the point cloud at the casing hole boundary to determine the spatial position of the casing.

[0034] The least squares method is used to generate a fitting plane for the casing hole based on the boundary point cloud;

[0035] The normal vector of the fitted plane is obtained as the sleeve orientation.

[0036] Optionally, the method further includes:

[0037] A simulated binocular camera system is established under a simulation environment based on the intrinsic and extrinsic parameters of the binocular camera system.

[0038] An eye simulation model in the simulation environment is constructed based on the size and relative position of the eyeball and the cannula.

[0039] In a simulated environment, a simulated binocular camera system is used to simulate and photograph an eye simulation model to obtain simulated binocular images.

[0040] Optionally, the method further includes:

[0041] Acquire multiple sets of real binocular images of the eyeballs captured by the binocular camera system in a real environment, and multiple sets of simulated binocular images of the eyeball simulation model captured in a simulated environment;

[0042] The sleeves and sleeve holes in the real and simulated binocular images are labeled, and a training dataset is constructed.

[0043] The YOLOv11 network was trained using the training dataset to obtain an object detection network, which was used to detect the sleeve region in the image.

[0044] The Segformer network was trained using the training dataset to obtain a semantic segmentation network, which was used to segment the sleeve and sleeve hole in the image.

[0045] The LightStereo network is trained using the training dataset to obtain a disparity estimation network, which is used to generate disparity maps.

[0046] Another aspect of the present invention discloses a binocular vision-based cannula pose recognition system for determining the pose of the cannula on the eyeball, the system being applied to the binocular vision-based cannula pose recognition method described in the foregoing aspect.

[0047] This invention discloses a binocular vision-based method and system for scleral cannula pose recognition. First, a binocular camera system covering the eyeball is constructed, and a set of binocular images are captured by the system. The cannula and cannula aperture are identified in each of the two images. Then, a disparity map is generated based on the cannula and cannula aperture in the two images. This disparity map is then converted into a depth map using the intrinsic parameters of the binocular camera system, obtaining the boundary point cloud of the cannula aperture. Finally, the cannula position and the fitting plane of the cannula aperture are determined based on the boundary point cloud, and its normal vector is used as the cannula pose. This invention reduces the operational complexity of aligning instruments with the scleral cannula in robot-assisted surgery and can be used for active eye movement compensation. Compared with existing methods, it eliminates the need for additional markers and effectively improves the accuracy and speed of cannula pose recognition. By directly identifying eye movements through cannula pose change information, it comprehensively compensates for eye movements caused by various factors, helping to enhance the compensation effect of eye movements and improve the safety of robot-assisted surgery.

[0048] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify essential or necessary features of this disclosure, nor is it intended to limit the scope of this disclosure. Attached Figure Description

[0049] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0050] Figure 1 This is a flowchart illustrating a binocular vision-based cannula pose recognition method disclosed in an embodiment of the present invention.

[0051] Figure 2 This is one implementation disclosed in an embodiment of the present invention. Figure 1 A flowchart illustrating step S100;

[0052] Figure 3 This is one implementation disclosed in an embodiment of the present invention. Figure 1 A flowchart illustrating step S300;

[0053] Figure 4 This is one implementation disclosed in an embodiment of the present invention. Figure 1 A flowchart of step S500;

[0054] Figure 5 This is one implementation disclosed in an embodiment of the present invention. Figure 1 A flowchart illustrating step S600;

[0055] Figure 6This is a schematic diagram of a process for obtaining simulated binocular images according to an embodiment of the present invention;

[0056] Figure 7 This is a schematic diagram of a network training process disclosed in an embodiment of the present invention. Detailed Implementation

[0057] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0058] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0059] Figure 1 This is a flowchart illustrating a binocular vision-based cannula pose recognition method disclosed in an embodiment of the present invention, used to determine the pose of the cannula on the eyeball. Figure 1 As shown, the method includes the following steps:

[0060] Step S100: Construct a binocular camera system with a shooting range covering the eyeball.

[0061] The embodiments of this invention first require the construction of a vision system with spatial depth perception capabilities for subsequent pose estimation. The binocular camera system constructed must at least meet the following conditions:

[0062] (1) The field of view must cover the entire eyeball and its surrounding structures.

[0063] (2) It has sufficient resolution and frame rate to support the accuracy requirements of subsequent image processing.

[0064] (3) It can stably acquire synchronized left and right image pairs, namely the first image and the second image proposed in the subsequent embodiments.

[0065] In one embodiment of the present invention, such as Figure 2 As shown, step S100 can be implemented using the following sub-steps:

[0066] Step S101: Set up the binocular camera according to the preset application scenario.

[0067] Based on the intended application, determine the specific requirements of the shooting environment, such as whether the subject contains structures like eyeballs or cannulas; select appropriate baseline distance and focal length based on field of view size and pixel density; avoid overexposure or reflection interference, and add auxiliary lighting if necessary; and ensure that the optical axes of the two cameras are as parallel as possible and their center heights are consistent, etc.

[0068] Step S102: Using a preset calibration method, calibrate the intrinsic and extrinsic parameters of the binocular camera based on the calibration board.

[0069] In a specific embodiment of the present invention, the Zhang Zhengyou calibration method can be used to calibrate the intrinsic and extrinsic parameters of the camera. The intrinsic parameters of the camera include the intrinsic parameter matrices of the left and right cameras, distortion coefficients, etc.; the extrinsic parameters of the camera include the rotation matrix and translation vector between the left and right cameras, etc.

[0070] Step S200: Acquire a set of binocular images captured by the binocular camera system.

[0071] A binocular image consists of a first image and a second image, which are images captured by the left and right cameras, respectively. The first image is the left image, and the second image is the right image.

[0072] In one embodiment of the present invention, after performing step S200, it is also necessary to perform epipolar correction on the first image and the second image so that the preset corresponding points in the two images are located on the same horizontal line.

[0073] For example, after obtaining the intrinsic and extrinsic parameters of the two cameras, the correction transformation matrix of the left and right cameras is first calculated using OpenCV. Then, a mapping table of the left and right images is generated. Finally, the original image is stretched or compressed according to the mapping table to complete the epipolar correction.

[0074] In the corrected left and right images, each set of corresponding points is on the same horizontal line, which facilitates subsequent disparity calculation.

[0075] Step S300: Identify the sleeve and sleeve hole in the two images respectively.

[0076] In one embodiment of the present invention, such as Figure 3 As shown, step S300 can be implemented by the following sub-steps:

[0077] Step S301: Use a pre-trained target detection network to detect the sleeve region in the first image and the second image respectively.

[0078] The left image (first image) and right image (second image) after epipolar correction are input into the pre-trained object detection network YOLOv11. This object detection network has the ability to detect sleeve regions under complex lighting and different imaging conditions; the specific training process will be described in subsequent embodiments. The YOLOv11 network performs feature extraction and object recognition on the image, outputting a set of bounding boxes and corresponding class labels and confidence scores. To improve the reliability of object detection, a confidence threshold is preset; only bounding boxes with scores higher than this threshold are retained, and the largest bounding box or the bounding box located at the center of the image is preferentially selected as the finally detected sleeve region.

[0079] Step S302: Cropping out the sleeve region from the first image as a first sub-image, and cropping out the sleeve region from the second image as a second sub-image.

[0080] Based on the cannula region detected by the YOLOv11 network, this region is cropped from the original image to obtain a first sub-image corresponding to the cannula region in the first image and a second sub-image corresponding to the cannula region in the second image. To ensure uniformity of subsequent network input, these two sub-images are standardized to a fixed size of 192×160. This image cropping effectively reduces the processing area, improving the efficiency of subsequent neural network inference, and also reduces background interference for subsequent segmentation and matching tasks.

[0081] Step S303: Use a pre-trained semantic segmentation network to identify the sleeve and sleeve hole in the first sub-image and the second sub-image respectively.

[0082] The cropped first and second sub-images are input into the pre-trained semantic segmentation network Segformer. This network employs an encoder-decoder structure and possesses strong global context modeling capabilities, enabling fine segmentation of different structural regions in the image. The training phase covers various morphological and positional variations of cannulas and cannula holes, ensuring the network achieves good segmentation results for images under different environments. The specific training process will be described in subsequent embodiments. The output of the Segformer network is a pixel-level segmentation map of the same size as the input image, with each pixel assigned to a specific category, including background, cannula, and cannula hole.

[0083] The semantic segmentation results can be used to generate a mask map, retaining only the regions in the image that belong to the sleeve or sleeve hole, providing a precise region of interest for the subsequent disparity map calculation process.

[0084] Step S400: Generate a disparity map based on the sleeve and sleeve hole in the two images.

[0085] In one embodiment of the present invention, step S400 can be implemented in the following manner:

[0086] The first and second sub-images are input into a pre-trained disparity estimation network, and the identified sleeves and sleeve holes are used as mask regions to output a disparity map of the mask regions, which uses the first sub-image as the reference image.

[0087] After semantic segmentation of the cannula region and cannula hole region in the first and second sub-images, a mask image is obtained for each sub-image. For example, in the mask image, pixels marked "1" correspond to "cannula", pixels marked "2" correspond to "cannula hole", and pixels marked "0" correspond to "background". This mask image will serve as a region constraint for pixel matching in the subsequent disparity estimation process, ensuring that the network only performs calculations within the region of interest, reducing interference from irrelevant regions, and improving computational efficiency and accuracy.

[0088] The first and second sub-images, after object detection and cropping, are input together into the pre-trained disparity estimation network LightStereo. LightStereo is a lightweight end-to-end stereo matching neural network capable of generating high-resolution, high-precision disparity maps. The network input consists of the first and second sub-images and a mask image, enabling disparity search and matching only within the mask region. Finally, the network outputs a disparity map with the first sub-image as the reference viewpoint, representing the horizontal displacement of each pixel within the mask region between the first and second sub-images. The output disparity map is a two-dimensional matrix, where the value of each pixel represents the disparity value of that point between the two images, which can be further used for subsequent depth estimation and point cloud reconstruction.

[0089] Step S500: Combine the intrinsic parameters of the binocular camera system to convert the disparity map into a depth map and obtain the boundary point cloud of the casing hole.

[0090] In one embodiment of the present invention, such as Figure 4 As shown, step S500 can be completed using the following sub-steps:

[0091] Step S501: Obtain the correlation between disparity and depth based on the intrinsic parameters of the binocular camera system.

[0092] Based on the known intrinsic parameters of the binocular camera system, a mathematical mapping relationship between disparity and depth is established. This mapping relationship can be encoded into a function to achieve real-time conversion from disparity to depth.

[0093] Step S502: Use correlation to convert the disparity map into a depth map.

[0094] A point-by-point transformation is performed on each pixel in the disparity map. During the transformation process, to ensure the accuracy of the result, the precision unit of disparity must be considered. For example, median filtering can be used to smooth out abnormal depth values. The transformation result is a depth map of the same size as the disparity map, where each pixel value represents the depth (Z-coordinate) of its corresponding point in three-dimensional space.

[0095] Step S503: Map each pixel in the depth map to a 3D spatial point to form a sleeve point cloud.

[0096] Each pixel in the depth map is restored to three-dimensional spatial coordinates (X,Y,Z) through back projection, resulting in a set of three-dimensional coordinates that constitute a complete point cloud of the casing area.

[0097] Step S504: Determine the pixel points corresponding to the edge of the sleeve hole in the first sub-image, and extract the three-dimensional spatial points corresponding to the pixel points in the sleeve point cloud as the boundary point cloud of the sleeve hole.

[0098] The region marked as the cannula opening in the semantic segmentation result is extracted, and then the Canny edge detection operator is used on this region map to extract the edge pixel positions. The Canny operator has good edge detection performance and outputs a set of two-dimensional pixel coordinates. Subsequently, combined with the position of these two-dimensional coordinates in the depth map, the corresponding three-dimensional spatial points are found to complete the mapping from the two-dimensional edge to the three-dimensional boundary point cloud, thus constructing the boundary point cloud of the cannula opening.

[0099] Step S600: Determine the sleeve position and the fitting plane of the sleeve hole based on the boundary point cloud, and use the normal vector of the fitting plane as the sleeve attitude.

[0100] In one embodiment of the present invention, such as Figure 5 As shown, step S600 can be implemented using the following sub-steps:

[0101] Step S601: Calculate the average coordinates of the point cloud at the casing hole boundary as the spatial position of the casing.

[0102] Statistical analysis was performed on the extracted point cloud of the casing hole edge, and the spatial average coordinates of all points were calculated. The obtained average coordinates were used to represent the spatial position of the casing.

[0103] Step S602: Generate the fitting plane of the casing hole based on the boundary point cloud using the least squares method.

[0104] Using the least squares method or other methods, an optimal plane is fitted based on all boundary point clouds to describe the approximate distribution of all edge points of the casing hole.

[0105] Step S603: Obtain the normal vector of the fitting plane as the sleeve orientation.

[0106] This normal vector is the spatial unit vector of the casing orifice orientation, which can represent the orientation of the casing.

[0107] In one embodiment of the present invention, such as Figure 6 As shown, after performing step S100 to build a binocular camera system with a shooting range covering the eyeball, the following steps are also included:

[0108] Step S021: Establish a simulated binocular camera system in the simulation environment based on the intrinsic and extrinsic parameters of the binocular camera system.

[0109] A simulated stereo camera system is built in 3D simulation software based on the intrinsic parameters (including focal length, principal point position, distortion parameters, etc.) and extrinsic parameters (rotation matrix and translation vector between the left and right cameras, etc.) of the stereo camera. All parameters of the simulated camera must strictly correspond to the calibration results of the real camera to ensure that the simulated image has consistent geometric characteristics with the real image. This allows the simulated image obtained in the simulation environment to be used for neural network training and successfully transferred to the real scene.

[0110] Step S022: Construct an eye simulation model in the simulation environment based on the size and relative position of the eyeball and the cannula.

[0111] A three-dimensional simulation model is constructed based on the actual eyeball model and the dimensions of the cannula structure (such as radius, aperture, height, insertion depth, etc.) and their relative spatial relationships. The modeling process can be completed using 3D modeling software and then imported into the simulation environment. It is necessary to ensure that the geometric accuracy of the model is consistent with the real object in order to generate image data with training value.

[0112] Step S023: In a simulation environment, a simulated binocular camera system is used to simulate and photograph the eye simulation model to obtain a simulated binocular image.

[0113] Using a pre-built simulated binocular camera system, the simulated model is sampled and photographed from different angles and under different lighting conditions, generating a large number of left and right image pairs (simulated binocular images). The image pairs need to cover various changes in position, angle, and distance to enhance the network's adaptability to complex real-world scenes.

[0114] In one embodiment of the present invention, such as Figure 7 As shown, the following sub-steps can be used to train the object detection network, semantic segmentation network, and disparity estimation network:

[0115] Step S024: Acquire multiple sets of real binocular images of the eyeballs captured by the binocular camera system in a real environment, and multiple sets of simulated binocular images of the eyeball simulation model captured in a simulated environment.

[0116] The system collects real image pairs obtained by capturing images of the real eyeball and the cannula using a real binocular camera, as well as a large number of simulated images generated by the simulation system.

[0117] Step S025: Label the sleeves and sleeve holes in the real stereo images and simulated stereo images, and construct the training dataset.

[0118] The sleeves and sleeve holes in all images (real and simulated images) are labeled. The labeling work can be completed by combining semi-automatic tools and manual correction. A dataset is constructed based on the labeled images. In one embodiment of this invention, the dataset is divided into a training dataset and a validation dataset, wherein the training dataset is used for model training and the validation dataset is used to verify the accuracy of the model.

[0119] Step S026: Train the YOLOv11 network using the training dataset to obtain an object detection network, which is used to detect the sleeve region in the image.

[0120] The YOLOv11 network is trained using a training dataset. This network structure can quickly detect sleeve regions in an image based on image features and output bounding box coordinates. In other embodiments disclosed in this invention, the training network is not limited to the YOLOv11 network; other networks can also be used for training. During training, data augmentation strategies (such as rotation, brightness variation, cropping, etc.) can be employed to improve network robustness, and the loss function can be optimized to ensure detection accuracy.

[0121] Step S027: Train the Segformer network using the training dataset to obtain a semantic segmentation network, which is used to segment the sleeve and sleeve hole in the image.

[0122] Segformer is a high-precision, lightweight Transformer-structured semantic segmentation network. After training on a training dataset, it can extract two fine-grained category regions from an image: the entire sleeve and the sleeve hole, providing support for subsequent mask disparity calculations. In other embodiments disclosed in this invention, the training network is not limited to Segformer; other networks can also be used for training.

[0123] Step S028: Train the LightStereo network using the training dataset to obtain a disparity estimation network, which is used to generate a disparity map.

[0124] The LightStereo network is trained under supervised supervision using simulated stereo images from the training dataset, enabling it to estimate disparity maps based on image pairs and segmented mask regions. The training objective is to minimize the error between the predicted disparity and the true disparity. In other embodiments disclosed in this invention, the training network is not limited to the LightStereo network; other networks can also be used for training. Since the depth information of each pixel in the simulated image is computable, the trained model exhibits high accuracy and generalization ability.

[0125] This invention also discloses a binocular vision-based cannula pose recognition system for determining the pose of the cannula on the eyeball. This system is applied to the binocular vision-based multi-stage cannula pose recognition method disclosed in the foregoing embodiments.

[0126] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A binocular vision-based method for cannula pose recognition, used to determine the pose of the cannula on the eyeball, characterized in that, include: Build a binocular camera system whose shooting range covers the eyeball; Acquire a set of binocular images captured by a binocular camera system, wherein the binocular images include a first image and a second image; Identify the sleeve and sleeve hole in the two images respectively, including: A pre-trained target detection network was used to detect the sleeve region in the first and second images, respectively. The cannula region is cropped out from the first image as a first sub-image, and the cannula region is cropped out from the second image as a second sub-image; A pre-trained semantic segmentation network was used to identify the sleeve and sleeve hole in the first and second sub-images, respectively. Generate a disparity map based on the sleeve and sleeve hole in the two images; By combining the intrinsic parameters of the binocular camera system, the disparity map is converted into a depth map, and the boundary point cloud of the casing hole is obtained. The sleeve position and the fitting plane of the sleeve hole are determined based on the boundary point cloud, and the normal vector of the fitting plane is used as the sleeve attitude.

2. The sleeve pose recognition method according to claim 1, characterized in that, The binocular camera system with a shooting range covering the eyeball includes: Configure the binocular camera according to the preset application scenario; The intrinsic and extrinsic parameters of the stereo camera are calibrated using a preset calibration method and a calibration board.

3. The sleeve pose recognition method according to claim 1, characterized in that, After performing the step of acquiring a set of binocular images captured by the binocular camera system, the method further includes: Epipolar correction is performed on the first and second images to ensure that the preset corresponding points in the two images are on the same horizontal line.

4. The sleeve pose recognition method according to claim 1, characterized in that, The process of generating a disparity map based on the sleeve and sleeve hole in two images includes: The first and second sub-images are input into a pre-trained disparity estimation network, and the identified sleeves and sleeve holes are used as mask regions to output a disparity map of the mask regions, with the first sub-image as the reference image.

5. The sleeve pose recognition method according to claim 4, characterized in that, The intrinsic parameters of the binocular camera system are used to convert the disparity map into a depth map and obtain the boundary point cloud of the casing hole, including: The correlation between parallax and depth was obtained based on the intrinsic parameters of a binocular camera system. The disparity map is converted into a depth map using the aforementioned correlation. Map each pixel in the depth map to a 3D spatial point, as a sleeve point cloud; The pixels corresponding to the edge of the sleeve hole in the first sub-image are determined, and the three-dimensional spatial points corresponding to the pixels are extracted from the sleeve point cloud as the boundary point cloud of the sleeve hole.

6. The sleeve pose recognition method according to claim 5, characterized in that, The step of determining the sleeve position and the fitting plane of the sleeve hole based on the boundary point cloud, and using the normal vector of the fitting plane as the sleeve attitude, includes: Calculate the average coordinates of the point cloud at the casing hole boundary to determine the spatial position of the casing. The least squares method is used to generate a fitting plane for the casing hole based on the boundary point cloud; The normal vector of the fitted plane is obtained as the sleeve orientation.

7. The sleeve pose recognition method according to claim 1, characterized in that, The method further includes: A simulated binocular camera system is established under a simulation environment based on the intrinsic and extrinsic parameters of the binocular camera system. An eye simulation model in the simulation environment is constructed based on the size and relative position of the eyeball and the cannula. In a simulated environment, a simulated binocular camera system is used to simulate and photograph an eye simulation model to obtain simulated binocular images.

8. The sleeve pose recognition method according to claim 7, characterized in that, The method further includes: Acquire multiple sets of real binocular images of the eyeballs captured by the binocular camera system in a real environment, and multiple sets of simulated binocular images of the eyeball simulation model captured in a simulated environment; The sleeves and sleeve holes in the real and simulated binocular images are labeled, and a training dataset is constructed. The YOLOv11 network was trained using the training dataset to obtain an object detection network, which was used to detect the sleeve region in the image. The Segformer network was trained using the training dataset to obtain a semantic segmentation network, which was used to segment the sleeve and sleeve hole in the image. The LightStereo network is trained using the training dataset to obtain a disparity estimation network, which is used to generate disparity maps.

Citation Information

Patent Citations

  • Dynamic traffic information acquisition method and system based on binocular stereoscopic vision

    CN118262511A