Cannula pose recognition method and system based on binocular vision
Through the cannula posture recognition method based on binocular vision, the posture of the cannula on the eye is identified, and the problem of insufficient complexity and accuracy caused by relying on external markers in the prior art is solved, efficient and accurate eye movement compensation is achieved, and surgical safety is improved.
Patent Information
- Application Number
- CN202510661100.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The high reliance on external markers in the prior art leads to high complexity of surgical scenarios and insufficient in real-time and estimation accuracy, making it difficult to meet the application needs of high-time and high-stability intraocular surgery.
Using a cannula position recognition method based on binocular vision, a binocular camera system covering the eyeball is built, binocular images are collected, the casing and casing holes are identified, the parallax map is generated, and the boundary point cloud of the casing holes is converted into a depth map, and the casing posture is determined by fitting the plane.
The operation complexity of the instrument alignment of the scleral cannula in robot-assisted surgery is reduced, and active eye movement compensation is achieved without additional markers, the accuracy and speed of cannula position recognition is improved, the eye movement compensation effect is enhanced, and the safety of robot-assisted surgery is improved.
Smart Images

Figure CN120182360A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical devices, and particularly relates to a method and system for identifying the pose of a cannula based on binocular vision. Background Art
[0002] Intraocular surgery is a highly delicate type of surgery, mainly applied in the treatment processes of ophthalmic diseases such as retinal repair and vitrectomy. Due to the fine and fragile intraocular tissue structure, extremely high requirements are imposed on the precision and stability of the operation, thus posing extremely high requirements on the technical level of surgeons. In recent years, with the development of robot-assisted surgical systems, their high-precision and high-stability motion control capabilities help reduce the surgical operation difficulty for doctors and improve the safety and success rate of surgeries.
[0003] In intraocular surgery, doctors usually need to first use special instruments to fix the scleral cannula to the sclera of the patient's eyeball, and then insert the surgical instrument into the eye through the cannula for operation. Although robot-assisted surgical systems can effectively reduce the difficulty of doctors' operations during intraocular procedures, they introduce new challenges. For example, the robot-assisted system needs to accurately align with the scleral cannula fixed on the surface of the eyeball to smoothly insert the surgical instrument into the eye. This step is complex and time-consuming, limiting the surgical efficiency.
[0004] In addition, during the surgery, the patient may have small displacements of the eyeball due to factors such as spontaneous breathing, slight muscle movement, or other uncontrolled physiological factors, resulting in changes in the relative position between the surgical instrument and the eyeball. If such displacements are not compensated in real time, it may affect the surgical precision and even cause tissue damage. Therefore, to ensure the safety of the surgery, real-time compensation for eyeball movement is required.
[0005] Existing eyeball movement compensation methods are mainly divided into two categories: passive compensation and active compensation. Among them, passive compensation methods usually reduce the relative movement between the eyeball and the instrument through physical fixation. For example, some studies have used methods such as fixing the patient's head, fixing the eyeball, or directly mounting the surgical robot on the patient's head to achieve the suppression of relative movement. However, although these methods can alleviate the displacement problem between the instrument and the eyeball to a certain extent, they have many limitations. On the one hand, physical fixation may cause discomfort to the patient; on the other hand, due to the compliance of the periorbital soft tissues, passive fixation cannot completely restrict the small movements of the eyeball within the orbit, thus affecting the operation stability.
[0006] In contrast, the active compensation method uses an external sensor to sense eye movements in real time and feeds back to the robotic system, driving the robot to make dynamic adjustments to achieve real-time compensation for eye displacement. Existing active compensation technologies mainly rely on vision sensing systems to identify and track scleral sleeves or markers. For example, some literature proposes integrating a binocular micro camera with an inertial measurement unit (IMU) to track ArUco markers on both sides of the scleral sleeve to estimate the sleeve pose; another method uses the YOLO network to perform object detection on the sleeve image and sends the detected region to the SC6D network to perform pose estimation under monocular RGB images. Such methods have good active response capabilities and can adapt to complex eye movement scenarios. However, due to their high dependence on external markers, they increase the complexity of the surgical scenario and still have deficiencies in terms of real-time performance and estimation accuracy, making it difficult to meet the requirements of high-timeliness and high-stability intraocular surgery applications. Summary of the Invention
[0007] The present invention provides a method and system for identifying the pose of a sleeve based on binocular vision to solve the problems in the prior art that highly rely on external markers, resulting in a relatively high complexity of the surgical scenario and deficiencies in terms of real-time performance and estimation accuracy.
[0008] To solve the above technical problems, the embodiments of the present invention disclose the following technical solutions: One aspect of the present invention provides a method for identifying the pose of a sleeve based on binocular vision, which is used to determine the pose of the sleeve on the eye and includes: Construct a binocular camera system with a shooting range covering the eye; Collect a set of binocular images captured by the binocular camera system, where the binocular images include a first image and a second image; Identify the sleeve and the sleeve hole in each of the two images respectively; Generate a disparity map based on the sleeve and the sleeve hole in the two images; Convert the disparity map into a depth map in combination with the internal parameters of the binocular camera system and obtain the boundary point cloud of the sleeve hole; Determine the sleeve position and the fitting plane of the sleeve hole according to the boundary point cloud, and use the normal vector of the fitting plane as the sleeve pose.
[0009] Optionally, the construction of the binocular camera system with a shooting range covering the eye includes: Set the binocular camera according to a preset application scenario; Use a preset calibration method to calibrate the internal and external parameters of the binocular camera based on a calibration board.
[0010] Optionally, after the step of collecting a set of binocular images captured by the binocular camera system, the method further includes: Perform epipolar rectification on the first image and the second image so that the preset corresponding points in the two images are on the same horizontal line.
[0011] Optionally, the separately identifying the casing and the casing hole in the two images includes: Using a pre-trained object detection network to detect the casing regions in the first image and the second image respectively; Cropping out the casing region in the first image as the first sub-image, and cropping out the casing region in the second image as the second sub-image; Using a pre-trained semantic segmentation network to identify the casing and the casing hole in the first sub-image and the second sub-image respectively.
[0012] Optionally, the generating a disparity map based on the casing and the casing hole in the two images includes: Inputting the first sub-image and the second sub-image into a pre-trained disparity estimation network, and using the identified casing and casing hole as the mask regions to output the disparity map of the mask regions, where the disparity map takes the first sub-image as the reference image.
[0013] Optionally, the combining the internal parameters of the binocular camera system to convert the disparity map into a depth map and obtaining the boundary point cloud of the casing hole includes: Obtaining the correlation relationship between disparity and depth based on the internal parameters of the binocular camera system; Using the correlation relationship to convert the disparity map into a depth map; Mapping each pixel point in the depth map to a three-dimensional space point as the casing point cloud; Determining the pixel points corresponding to the edges of the casing hole in the first sub-image, and extracting the three-dimensional space points corresponding to the pixel points from the casing point cloud as the boundary point cloud of the casing hole.
[0014] Optionally, the determining the casing position and the fitting plane of the casing hole according to the boundary point cloud and taking the normal vector of the fitting plane as the casing attitude includes: Calculating the average coordinates of the boundary point cloud of the casing hole as the spatial position of the casing; Using the least squares method to generate the fitting plane of the casing hole based on the boundary point cloud; Obtaining the normal vector of the fitting plane as the casing attitude.
[0015] Optionally, the method further includes: Establishing a simulated binocular camera system in the simulation environment according to the internal parameters and external parameters of the binocular camera system; Constructing an eyeball simulation model in the simulation environment based on the sizes and relative positions of the eyeball and the casing; Performing simulated shooting on the eyeball simulation model using the simulated binocular camera system in the simulation environment to obtain simulated binocular images.
[0016] Optionally, the method further includes: Obtaining multiple sets of real binocular images of the eyeball captured by the binocular camera system in a real environment, and multiple sets of simulated binocular images of the eyeball simulation model captured in a simulation environment; Labeling the cannulas and cannula holes in the real binocular images and the simulated binocular images, and constructing a training dataset; Training the YOLOv11 network using the training dataset to obtain an object detection network for detecting the cannula area in the image; Training the Segformer network using the training dataset to obtain a semantic segmentation network for segmenting the cannulas and cannula holes in the image; Training the LightStereo network using the training dataset to obtain a disparity estimation network for generating a disparity map.
[0017] Another aspect of the present invention discloses a cannula pose recognition system based on binocular vision for determining the pose of the cannula on the eyeball. The system is applied to the cannula pose recognition method based on binocular vision described in the foregoing aspect.
[0018] A cannula pose recognition method and system based on binocular vision disclosed by the present invention. First, a binocular camera system with a shooting range covering the eyeball is built, a set of binocular images captured by the binocular camera system is collected, and the cannulas and cannula holes in the two images are respectively recognized. Then, a disparity map is generated based on the cannulas and cannula holes in the two images, and the disparity map is converted into a depth map by combining the internal parameters of the binocular camera system to obtain the boundary point cloud of the cannula hole. Finally, the cannula position and the fitting plane of the cannula hole are determined according to the boundary point cloud, and its normal vector is used as the cannula pose. The present invention can reduce the operation complexity of aligning the instrument with the scleral cannula in robot-assisted surgery and can be used for active eyeball movement compensation. Compared with the existing methods, there is no need to additionally set markers, and the accuracy and speed of cannula pose recognition are effectively improved. The eyeball movement is directly recognized through the cannula pose change information, and the eyeball movement caused by various factors is comprehensively compensated, which helps to enhance the compensation effect of the eyeball movement and improve the safety of robot-assisted surgery.
[0019] The invention content part is provided to introduce the selection of concepts in a simplified form, which will be further described in the specific implementation manners below. The invention content part is not intended to identify the important features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the exemplary embodiments of the present disclosure in more detail with reference to the accompanying drawings, in which, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0021] Figure 1 Schematic flowchart of a method for identifying the pose of a casing based on binocular vision disclosed in an embodiment of the present invention; Figure 2 Disclosed in an embodiment of the present invention for realizing Figure 1 Schematic flowchart of step S100 in Figure 3 Disclosed in an embodiment of the present invention for realizing Figure 1 Schematic flowchart of step S300 in Figure 4 Disclosed in an embodiment of the present invention for realizing Figure 1 Schematic flowchart of step S500 in Figure 5 Disclosed in an embodiment of the present invention for realizing Figure 1 Schematic flowchart of step S600 in Figure 6 Schematic flowchart of a method for obtaining simulated binocular images disclosed in an embodiment of the present invention; Figure 7 Schematic flowchart of a process for realizing network training disclosed in an embodiment of the present invention. Detailed implementation manners
[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0023] As used herein, the term "including" and its variants mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an exemplary embodiment" and "an embodiment" mean "at least one exemplary embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may be other explicit and implicit definitions below.
[0024] Figure 1 Schematic flowchart of a method for identifying the pose of a casing based on binocular vision disclosed in an embodiment of the present invention, used to determine the pose of the casing on the eyeball. As Figure 1As shown, the method includes the following steps: Step S100: Build a binocular camera system with a shooting range covering the eyeball.
[0025] In the embodiment of the present invention, a vision system with spatial depth perception ability needs to be constructed first for subsequent pose estimation. The built binocular camera system should at least meet the following conditions: (1) The field of view needs to cover the entire eyeball and its surrounding structures.
[0026] (2) Have sufficient resolution and frame rate to support the accuracy requirements of subsequent image processing.
[0027] (3) Be able to stably obtain synchronized left and right image pairs, that is, the first image and the second image proposed in subsequent embodiments.
[0028] In an embodiment disclosed in the present invention, as Figure 2 shown, the following sub-steps can be used to implement step S100: Step S101: Set the binocular camera according to the preset application scenario.
[0029] According to the preset application purpose, determine the specific requirements of the shooting environment. For example, whether the shooting object includes structures such as the eyeball and cannula; select appropriate baseline distance and focal length according to the field of view size and pixel density; avoid overexposure or reflection interference, and add auxiliary lighting if necessary; and ensure that the optical axes of the two cameras are as parallel as possible and the center heights are the same.
[0030] Step S102: Use a preset calibration method to calibrate the internal and external parameters of the binocular camera based on a calibration board.
[0031] In a specific embodiment disclosed in the present invention, the Zhang-Zhengyou calibration method can be used to complete the calibration of the internal and external parameters of the camera. Among them, the internal parameters of the camera include the internal parameter matrices, distortion coefficients, etc. of the left and right cameras; the external parameters of the camera include the rotation matrix, translation vector, etc. between the left and right cameras.
[0032] Step S200: Collect a group of binocular images captured by the binocular camera system.
[0033] The binocular images include the first image and the second image, that is, the images captured by the left and right cameras. Among them, the first image is the left image and the second image is the right image.
[0034] In an embodiment disclosed in the present invention, after executing step S200, the first image and the second image also need to be epipolar corrected so that the preset corresponding points in the two images are on the same horizontal line.
[0035] For example, after obtaining the internal and external parameters of two cameras, first, use OpenCV to calculate the rectification transformation matrices of the left and right cameras, then, generate the mapping tables for the left and right images, and finally, stretch or compress the original images according to the mapping tables to complete epipolar rectification.
[0036] For the rectified left and right images, each pair of corresponding points lies on the same horizontal line, facilitating subsequent disparity calculation.
[0037] Step S300: Identify the casing and the casing hole in the two images respectively.
[0038] In an embodiment disclosed by the present invention, as Figure 3 shown, step S300 can be implemented by the following sub-steps: Step S301: Use a pre-trained object detection network to detect the casing regions in the first image and the second image respectively.
[0039] Input the rectified left image (the first image) and the right image (the second image) into the pre-trained object detection network YOLOv11 respectively. This object detection network has the ability to detect the casing region under complex lighting and different imaging conditions. The specific training process will be described in the subsequent embodiments. The YOLOv11 network extracts features and performs object recognition on the images, and outputs a set of bounding boxes and corresponding class labels and confidence scores. To improve the reliability of object detection, a confidence threshold is preset, and only the bounding boxes with scores higher than this threshold will be retained, and the largest bounding box or the bounding box located at the center of the image is preferentially selected as the finally detected casing region.
[0040] Step S302: Crop the casing region in the first image as the first sub-image, and crop the casing region in the second image as the second sub-image.
[0041] According to the casing region detected by the YOLOv11 network, crop this region in the original image to obtain the first sub-image corresponding to the casing region in the first image and the second sub-image corresponding to the casing region in the second image respectively. To adapt to the unity of subsequent network input, these two sub-images are standardized to a fixed size of 192×160. By performing cropping processing on the images, on the one hand, the processing area is effectively reduced, improving the inference efficiency of the subsequent neural network, and on the other hand, the interference of the background to the subsequent segmentation and matching tasks can also be reduced.
[0042] Step S303: Use a pre-trained semantic segmentation network to identify the casing and the casing hole in the first sub-image and the second sub-image respectively.
[0043] The first sub-image and the second sub-image obtained by cropping are respectively input into the pre-trained semantic segmentation network Segformer. This network adopts an encoder-decoder structure and has strong global context modeling ability, which can finely segment different structural regions in the image. The training stage covers various morphological and positional changes of the casing and the casing hole, ensuring that the network has good segmentation effects on images in different environments. The specific training process will be described in the subsequent embodiments. The output of the Segformer network is a pixel-level segmentation map with the same size as the input image, and each pixel is assigned to a specific category, including background, casing, and casing hole.
[0044] The semantic segmentation result can be used to generate a mask map, which only retains the regions in the image that belong to the casing or the casing hole, providing an accurate region of interest for the subsequent disparity map calculation process.
[0045] Step S400: Generate a disparity map based on the casing and the casing hole in the two images.
[0046] In an embodiment disclosed by the present invention, step S400 can be implemented in the following manner: Input the first sub-image and the second sub-image into the pre-trained disparity estimation network, and use the identified casing and casing hole as the mask regions to output the disparity map of the mask regions, with the first sub-image as the reference image.
[0047] After completing the semantic segmentation of the casing region and the casing hole region in the first sub-image and the second sub-image, the corresponding mask maps in each sub-image are obtained. For example, in the mask map, the pixels marked as "1" correspond to "casing", the pixels marked as "2" correspond to "casing hole", and the pixels marked as "0" correspond to "background". This mask map will be used as the region constraint for pixel matching in the subsequent disparity estimation process, ensuring that the network only performs calculations within the region of interest, reducing interference from irrelevant regions, and improving the calculation efficiency and accuracy.
[0048] Input the first sub-image and the second sub-image that have undergone object detection and cropping processing into the pre-trained disparity estimation network LightStereo. LightStereo is a lightweight end-to-end stereo matching neural network that can generate high-resolution and high-precision disparity maps. The network inputs are the first sub-image, the second sub-image, and the mask map, realizing disparity search and matching only within the mask regions. Finally, the network outputs a disparity map with the first sub-image as the reference perspective, that is, the horizontal displacement amount of each pixel within the mask region between the first sub-image and the second sub-image. The output disparity map is a two-dimensional matrix, and the value of each pixel represents the disparity value of this point between the binocular images, which can be further used for subsequent depth estimation and point cloud reconstruction.
[0049] Step S500: Convert the disparity map into a depth map in combination with the internal parameters of the binocular camera system, and obtain the boundary point cloud of the casing hole.
[0050] In an embodiment disclosed in the present invention, as Figure 4 shown, the following sub-steps can be adopted to complete Step S500: Step S501: Obtain the correlation between disparity and depth based on the internal parameters of the binocular camera system.
[0051] Based on the known internal parameter values of the binocular camera system, establish a mathematical mapping relationship between disparity and depth, and this mapping relationship can be encoded as a function to achieve real-time conversion from disparity to depth.
[0052] Step S502: Convert the disparity map into a depth map using the correlation.
[0053] Perform point-by-point conversion on each pixel in the disparity map. During the conversion process, to ensure the accuracy of the result, the accuracy unit of the disparity needs to be considered. For example, median filtering can be used to smooth abnormal depth values. The conversion result is a depth map with the same size as the disparity map, and each pixel value in the image represents the depth (Z-direction coordinate) of its corresponding point in the three-dimensional space.
[0054] Step S503: Map each pixel point in the depth map to a three-dimensional space point as the casing point cloud.
[0055] Restore each pixel point in the depth map to three-dimensional space coordinates (X, Y, Z) through back-projection, and obtain a set of three-dimensional coordinate sets to form a complete casing area point cloud.
[0056] Step S504: Determine the pixel points corresponding to the edges of the casing hole in the first sub-image, and extract the three-dimensional space points corresponding to the pixel points from the casing point cloud as the boundary point cloud of the casing hole.
[0057] Extract the area marked as the casing hole in the semantic segmentation result, and then use the Canny edge detection operator on this area map to extract the edge pixel positions. The Canny operator has good edge detection performance and outputs a set of two-dimensional pixel coordinates. Subsequently, combine the positions of these two-dimensional coordinates in the depth map to find their corresponding three-dimensional space points, complete the mapping from two-dimensional edges to three-dimensional boundary point clouds, and form the boundary point cloud of the casing hole opening.
[0058] Step S600: Determine the position of the casing and the fitting plane of the casing hole according to the boundary point cloud, and use the normal vector of the fitting plane as the casing attitude.
[0059] In an embodiment disclosed in the present invention, as Figure 5 shown, the following sub-steps can be adopted to implement Step S600: Step S601: Calculate the average coordinates of the boundary point cloud of the casing hole as the spatial position of the casing.
[0060] Perform statistical analysis on the extracted edge point cloud of the casing hole, calculate the spatial average coordinates of all its points, and use the obtained average coordinates to represent the spatial position of the casing.
[0061] Step S602: Generate a fitting plane of the casing hole based on the boundary point cloud using the least squares method.
[0062] Adopt the least squares method or other methods to fit an optimal plane based on all the boundary point clouds to describe the plane where all the edge points of the casing hole are approximately distributed.
[0063] Step S603: Obtain the normal vector of the fitting plane as the casing attitude.
[0064] This normal vector is the spatial unit vector of the orientation of the casing hole opening and can represent the pose of the casing.
[0065] In an embodiment disclosed by the present invention, as Figure 6 shown, after performing step S100 of setting up a binocular camera system with a shooting range covering the eyeball, the following steps are further included: Step S021: Establish a simulated binocular camera system in the simulation environment according to the internal and external parameters of the binocular camera system.
[0066] Build a simulated binocular camera system in a 3D simulation software according to the internal parameters (including focal length, principal point position, distortion parameters, etc.) and external parameters (rotation matrix and translation vector between the left and right cameras, etc.) of the binocular camera. All parameters of the simulated camera need to strictly correspond to the calibration results of the real camera to ensure that the simulated images and real images have consistent geometric characteristics, so that the simulated images obtained in the simulation environment can be used for neural network training and successfully transferred to the real scene.
[0067] Step S022: Construct an eyeball simulation model in the simulation environment based on the sizes and relative positions of the eyeball and the casing.
[0068] Construct a 3D simulation model according to the actual sizes of the eyeball model and the casing structure (such as radius, aperture, height, insertion depth, etc.) and their relative spatial relationships. The modeling process can be completed using 3D modeling software and imported into the simulation environment. It is necessary to ensure that the geometric accuracy of the model is consistent with the real object to generate image data with training value.
[0069] Step S023: Use the simulated binocular camera system to perform simulated shooting on the eyeball simulation model in the simulation environment to obtain simulated binocular images.
[0070] Use the established simulated binocular camera system to sample and photograph the above-mentioned simulation model from different perspectives and under different lighting conditions, generating a large number of left and right image pairs (simulated binocular images). The image pairs need to cover various changes such as positions, angles, distances, etc., to enhance the network's adaptability to actual complex scenarios.
[0071] In an embodiment disclosed by the present invention, as Figure 7 shown, the following sub-steps can be adopted to implement the training of the object detection network, semantic segmentation network, and disparity estimation network: Step S024: Obtain multiple groups of real binocular images of the eyeball taken by the binocular camera system in the real environment, and multiple groups of simulated binocular images of the eyeball simulation model taken in the simulation environment.
[0072] Respectively collect the real image pairs obtained by the real binocular camera photographing the real eyeball and the cannula, and a large number of simulated images generated in the simulation system.
[0073] Step S025: Label the cannula and cannula holes in the real binocular images and simulated binocular images, and construct a training dataset.
[0074] Label the cannula and cannula holes in all images (real images and simulated images). The labeling work can be completed in combination with semi-automatic tools and manual correction. Based on the labeled images, a dataset is constructed. In an embodiment disclosed by the present invention, the dataset is divided into a training dataset and a validation dataset, where the training dataset is used for model training, and the validation dataset is used to verify the accuracy of the model.
[0075] Step S026: Use the training dataset to train the YOLOv11 network to obtain an object detection network for detecting the cannula area in the image.
[0076] Use the training dataset to train the YOLOv11 network. This network structure can quickly detect the cannula area in the image according to the image features and output the bounding box coordinates. In other embodiments disclosed by the present invention, the training network is not limited to the YOLOv11 network, and other networks can also be used for training. During the training process, data augmentation strategies (such as rotation, brightness change, cropping, etc.) can be adopted to improve the network's robustness, and the loss function can be optimized to ensure the detection accuracy.
[0077] Step S027: Use the training dataset to train the Segformer network to obtain a semantic segmentation network for segmenting the cannula and cannula holes in the image.
[0078] Segformer is a high-precision and lightweight semantic segmentation network based on the Transformer architecture. After being trained with a training dataset, it can extract two fine-grained category regions, namely the overall casing and the casing holes, from the image, providing support for subsequent masked disparity calculation. In other embodiments disclosed in the present invention, the training network is not limited to the Segformer network, and other networks can also be used for training.
[0079] Step S028: Train the LightStereo network with the training dataset to obtain a disparity estimation network for generating a disparity map.
[0080] Supervise and train the LightStereo network based on the simulated binocular images in the training dataset so that it can estimate the disparity map based on the image pairs and the segmented mask regions. The training objective is to minimize the error between the predicted disparity and the true disparity. In other embodiments disclosed in the present invention, the training network is not limited to the LightStereo network, and other networks can also be used for training. Since the depth information of each pixel in the simulated image is computable, the trained model has high accuracy and generalization ability.
[0081] An embodiment of the present invention also discloses a binocular vision-based casing pose recognition system for determining the pose of the casing on the eyeball. This system is applied to the binocular vision-based multi-stage casing pose recognition method disclosed in the foregoing embodiments.
[0082] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technologies in the market, or to enable other ordinary skilled persons in the technical field to understand the embodiments disclosed herein.
Claims
1. A binocular vision-based sleeve pose recognition method for determining the pose of the sleeve on the eyeball, characterized in that, Including: Construct a binocular camera system with a shooting range covering the eyeball; Collect a set of binocular images captured by the binocular camera system, where the binocular images include a first image and a second image; Identify the cannula and the cannula hole in the two images respectively; Generate a disparity map based on the cannula and the cannula hole in the two images; Combine the internal parameters of the binocular camera system to convert the disparity map into a depth map, and obtain the boundary point cloud of the cannula hole; Determine the position of the cannula and the fitting plane of the cannula hole according to the boundary point cloud, and use the normal vector of the fitting plane as the cannula attitude.
2. The sleeve pose recognition method according to claim 1, characterized in that, The constructing of the binocular camera system with a shooting range covering the eyeball includes: Set the binocular camera according to the preset application scenario; Use the preset calibration method to calibrate the internal and external parameters of the binocular camera based on the calibration board.
3. The sleeve pose recognition method according to claim 1, characterized in that, After performing the step of collecting a set of binocular images captured by the binocular camera system, the method further includes: Perform epipolar correction on the first image and the second image so that the preset corresponding points in the two images are on the same horizontal line.
4. The sleeve pose recognition method according to any one of claims 1 to 3, characterized in that, The respectively identifying the cannula and the cannula hole in the two images includes: Use the pre-trained object detection network to detect the cannula region in the first image and the second image respectively; Crop the cannula region in the first image as the first sub-image, and crop the cannula region in the second image as the second sub-image; Use the pre-trained semantic segmentation network to identify the cannula and the cannula hole in the first sub-image and the second sub-image respectively.
5. The sleeve pose recognition method according to claim 4, characterized in that, The generating of the disparity map based on the cannula and the cannula hole in the two images includes: Input the first sub-image and the second sub-image into the pre-trained disparity estimation network, and use the identified cannula and cannula hole as the mask region to output the disparity map of the mask region, where the disparity map is based on the first sub-image as the reference image.
6. The sleeve pose recognition method according to claim 5, characterized in that, The combining of the internal parameters of the binocular camera system to convert the disparity map into a depth map and obtaining the boundary point cloud of the cannula hole includes: Obtain the correlation between disparity and depth based on the internal parameters of the binocular camera system; Use the correlation to convert the disparity map into a depth map; Map each pixel point in the depth map to a three-dimensional space point as the cannula point cloud; Determine the pixel points corresponding to the edge of the cannula hole in the first sub-image, and extract the three-dimensional space points corresponding to the pixel points from the cannula point cloud as the boundary point cloud of the cannula hole.
7. The sleeve pose recognition method according to claim 6, characterized in that, The determining of the position of the cannula and the fitting plane of the cannula hole according to the boundary point cloud and using the normal vector of the fitting plane as the cannula attitude includes: Calculate the average coordinates of the boundary point cloud of the cannula hole as the spatial position of the cannula; Use the least squares method to generate the fitting plane of the cannula hole based on the boundary point cloud; Obtain the normal vector of the fitting plane as the cannula attitude.
8. The sleeve pose recognition method according to claim 1, characterized in that, The method further includes: Establish a simulated binocular camera system in the simulation environment according to the internal and external parameters of the binocular camera system; Construct an eyeball simulation model in the simulation environment based on the size and relative position of the eyeball and the cannula; Use the simulated binocular camera system to simulate the shooting of the eyeball simulation model in the simulation environment to obtain simulated binocular images.
9. The casing pose recognition method according to claim 8, wherein, The method further includes: Obtain multiple sets of real binocular images of the eyeball captured by the binocular camera system in a real environment, and multiple sets of simulated binocular images of the eyeball simulation model captured in a simulation environment; Annotate the cannulas and cannula holes in the real binocular images and simulated binocular images, and construct a training dataset; Use the training dataset to train the YOLOv11 network to obtain an object detection network for detecting the cannula region in the image; Use the training dataset to train the Segformer network to obtain a semantic segmentation network for segmenting the cannulas and cannula holes in the image; Use the training dataset to train the LightStereo network to obtain a disparity estimation network for generating a disparity map.
10. A binocular vision-based casing pose recognition system for determining the pose of the casing on the eyeball, wherein, The system is applied to the binocular vision-based cannula pose recognition method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Tree fruit three-dimensional pose identification method and system based on single two-dimensional image
CN115810188A
Dynamic traffic information acquisition method and system based on binocular stereoscopic vision
CN118262511A
Method for positioning poses of screw holes based on three-dimensional point cloud
CN118628569A
Binocular electronic hard tube endoscope
CN211066498U
Eye surgery surgical system and computer implemented method for providing the position of at least one trocar point
US20210059857A1
Cited By
Lightweight human body detection and distance measurement method and system suitable for stage lamp
CN121074953A