Monocular Visual Pose Estimation Method and System Based on EPnP Algorithm

Through the combination of EPnP algorithm and a monocular camera, the precise position estimation and real-time tracking of the robotic arm under a single Marker marker are achieved, which solves the problems of large error accumulation and calculation volume in optical navigation positioning and tracking technology, and improves the accuracy and efficiency of orthopedic surgical robots.

CN119090958BActive Publication Date: 2025-07-08SUZHOU ZOEZEN ROBOT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411019817.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-06-18
Filing Date
2024-07-29
Publication Date
2025-07-08
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

The existing optical navigation positioning and tracking technology has problems such as the large number of Marker markers, large cumulative errors, large calculation amounts and high cost in orthopedic surgical robots, which affects the surgical path planning and robot accuracy.

Method used

The monocular visual pose estimation method based on the EPnP algorithm is used to combine the EPnP non-iteration algorithm with the monocular camera. By tracking the patient's movement process by the robotic arm under a Marker marker, the iterative optimization algorithm is used to obtain the precise position and adjust the end position of the robotic arm in real time.

Benefits of technology

It reduces registration error, reduces calculation amount, reduces cost, and improves registration accuracy and calculation efficiency, adapts to a dynamically changing medical environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119090958B_ABST
    Figure CN119090958B_ABST
Patent Text Reader

Abstract

The present invention provides a monocular vision pose estimation method and system based on the EPnP algorithm, which relates to the technical field of visual positioning. The method includes obtaining the internal parameters of a monocular camera; arranging spatial reference markers to determine the world coordinates and pixel coordinates of feature points; representing the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establishing a camera perspective projection model to form a system of linear equations, and determining the coordinates of the virtual control points in the camera coordinate system; introducing a parallel perspective projection model to obtain an initial estimate and determining the coordinates of the control points in the camera coordinate system to calculate the initial value; taking the reduction of the coordinate system distance difference as the optimization objective, using the initial value, and solving through an iterative optimization algorithm to obtain the accurate pose; fixing the monocular camera at the end of the robotic arm and recording the reference pose; when movement occurs, the camera acquires the current pose, and in combination with the reference pose and the current pose of the end of the robotic arm, calculates in real time the target pose to which the end of the robotic arm should move.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual positioning, and particularly to a monocular vision pose estimation method and system based on the EPnP algorithm. Background Art

[0002] The development of digital medical technology has promoted the progress of orthopedic surgical robot navigation technology. Most of the existing orthopedic surgical robot navigation systems monitor the moving body position of patients during surgery by means of NDI optical navigation positioning and tracking technology. First, multiple Marker markers are used to represent the spatial pose information of surgical instruments, patient pain points, and positioning scales. Then, NDI unifies the above Marker markers into the same coordinate system, and determines the starting and ending points of the surgical path planning, the positioning position of the robot, and the real-time tracking of the moving situation of the patient during surgery by means of the spatial data provided by the coordinate system. Therefore, the accuracy of the spatial pose information of each Marker marker directly affects the surgical path planning and the accuracy of the robot. The more Marker markers there are, the greater the accumulated error and the greater the computational amount. For this reason, the present invention proposes a monocular vision pose estimation method based on the EPnP algorithm, which combines the EPnP non-iterative algorithm with a monocular camera. Under the condition of only one Marker marker, the manipulator can track the movement process of the patient through the monocular camera;

[0003] The existing optical navigation positioning and tracking technology uses NDI to real-time track and display the positions of multiple Marker markers, which has problems such as a large number of Marker markers, large accumulated errors, large computational amounts, and high costs. The present invention proposes a monocular vision pose estimation method based on the EPnP algorithm, which combines the EPnP non-iterative algorithm with a monocular camera. Under the condition of only one Marker marker, it solves the problems of error accumulation and large computational amount caused by the need to match multiple Markers for intraoperative positioning by the NDI camera itself. This method reduces the registration error, improves the registration accuracy, reduces the computational amount, and reduces the cost. The application of the present invention can at least solve some problems of the existing technology. Summary of the Invention

[0004] The embodiments of the present invention provide a monocular vision pose estimation method and system based on the EPnP algorithm, which can at least solve some problems in the existing technology.

[0005] In the first aspect of the embodiments of the present invention,

[0006] A monocular vision pose estimation method based on the EPnP algorithm is provided, including:

[0007] Obtain the internal parameters of the monocular camera through camera calibration; determine the world coordinates of the feature points on the spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; represent the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establish a camera perspective projection model, form a linear equation system regarding the camera coordinates of the virtual control points, and determine the camera coordinates corresponding to the virtual control points in the camera coordinate system; wherein, the monocular camera is fixed at the end of the robotic arm, and the optical axis of the monocular camera coincides with the center point of the end of the robotic arm;

[0008] Introduce a parallel perspective projection model, obtain an initial estimate of the camera pose through an iterative parallel perspective solution process, determine the control point coordinates in the camera coordinate system based on the initial estimate of the camera pose, and calculate the initial values of the homogeneous barycentric coordinate coefficients of the control points; taking the reduction of the distance difference between the control points in the camera coordinate system and the world coordinate system as the optimization objective, use the initial values of the homogeneous barycentric coordinate coefficients of the control points, and solve through an iterative optimization algorithm to obtain the accurate pose of the spatial reference marker relative to the camera;

[0009] Take the initial pose of the spatial reference marker relative to the end of the robotic arm as the reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculate the current pose of the end of the robotic arm, combine the reference pose and the current pose of the end of the robotic arm, calculate the target pose of the end of the robotic arm in real time, and move the robotic arm according to the target position.

[0010] In an alternative embodiment,

[0011] Obtain the internal parameters of the monocular camera through camera calibration; determine the world coordinates of the feature points on the spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; represent the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establish a camera perspective projection model, form a linear equation system regarding the camera coordinates of the virtual control points, and determine the camera coordinates corresponding to the virtual control points in the camera coordinate system including:

[0012] Calibrate the monocular camera using the Zhang Zhengyou calibration method to obtain the internal parameter matrix of the monocular camera, including the camera focal length and the optical center coordinates;

[0013] Take the checkerboard within the camera's field of view as the planar reference marker, detect the pixel coordinates of the corner points of the planar reference marker through a feature extraction algorithm, select n of the corner points as feature points, and calculate the world coordinates of the feature points in the checkerboard coordinate system corresponding to the checkerboard based on the size of the checkerboard and the arrangement of the corner points;

[0014] In the chessboard coordinate system, select 4 non-coplanar points as virtual control points, establish the transformation relationship from the chessboard coordinate system to the virtual control point coordinate system corresponding to the virtual control points, and represent the world coordinates of the feature points through coordinate transformation and homogeneous barycentric coordinates to obtain the coordinate representation of the feature points in the virtual control point coordinate system;

[0015] Based on the internal parameter matrix, combine the world coordinates of the feature points and the coordinate representation of the feature points in the virtual control point coordinate system to establish a camera imaging perspective projection model equation set;

[0016] In the camera imaging perspective projection model equation set, introduce the unknown coordinate representation of the virtual control points in the camera coordinate system, transform the camera imaging perspective projection model equation set into a linear equation set of the control points in the camera coordinate system, and solve the linear equation set through the singular value decomposition numerical optimization method to obtain the coordinates of the 4 virtual control points in the camera coordinate system.

[0017] In an optional embodiment,

[0018] It further includes:

[0019] The internal parameter matrix of the monocular camera has the following formula:

[0020]

[0021] Where, K represents the internal parameter matrix, f x represents the focal length of the camera in the x-axis direction, f y represents the focal length of the camera in the y-axis direction, u0 represents the abscissa of the optical center of the camera, and v0 represents the ordinate of the optical center of the camera;

[0022] The camera imaging perspective projection model equation set has the following formula:

[0023]

[0024] Where, j represents the virtual control point index, i represents the feature point index, α ij represents the homogeneous barycentric coordinate of the i-th feature point in the virtual control point coordinate system, x j c represents the x component of the j-th virtual control point in the camera coordinate system, y j c represents the y component of the j-th virtual control point in the camera coordinate system, z j c represents the z component of the j-th virtual control point in the camera coordinate system, u i represents the abscissa of the i-th feature point on the image plane, v iDenote the ordinate of the \(i\)-th feature point on the image plane.

[0025] In an alternative embodiment,

[0026] Introduce a parallel perspective projection model. Through an iterative parallel perspective solution process, obtain an initial estimate of the camera pose. Based on the initial estimate of the camera pose, determine the coordinates of the control points in the camera coordinate system, and calculate the initial values of the homogeneous barycentric coordinate coefficients of the control points; taking reducing the distance difference between the control points in the camera coordinate system and the world coordinate system as the optimization objective, use the initial values of the homogeneous barycentric coordinate coefficients of the control points, and solve through an iterative optimization algorithm to obtain the accurate pose of the spatial reference marker relative to the camera, including:

[0027] By establishing a parallel perspective projection model, reconstruct the system of equations of the camera imaging perspective projection model into a linear system of equations regarding the control point coordinates and the homogeneous barycentric coordinate coefficients;

[0028] By iteratively solving the system of equations of the parallel perspective projection model, obtain an initial estimate of the camera pose, and based on the initial estimate, calculate the initial coordinates of the control points in the camera coordinate system. Based on the initial coordinates of the control points in the camera coordinate system, solve the homogeneous barycentric coordinate system of equations by the least squares method to obtain the initial values of the homogeneous barycentric coordinate coefficients of the control points;

[0029] Based on the camera pose and the initial values of the homogeneous barycentric coordinate coefficients of the control points, construct an objective function based on the distance difference between the control points in the camera coordinate system and the world coordinate system, minimize the objective function to obtain the optimal camera pose estimate, and obtain the accurate pose of the spatial reference marker relative to the camera.

[0030] In an alternative embodiment,

[0031] The objective function includes:

[0032] Its formula is as follows:

[0033]

[0034] where \(\beta\) represents the optimization objective, \(j\), \(l\) represent the virtual control point indices, \(C\) j c represents the coordinates of the \(j\)-th virtual control point in the camera coordinate system, \(C\) l c represents the coordinates of the \(l\)-th virtual control point in the camera coordinate system, \(C\) j w represents the coordinates of the \(j\)-th virtual control point in the world coordinate system, \(C\) l w represents the coordinates of the \(l\)-th virtual control point in the world coordinate system.

[0035] In an alternative embodiment,

[0036] The initial pose of the spatial reference marker relative to the end of the robotic arm is used as the reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculates the current pose of the end of the robotic arm, combines the reference pose and the current pose of the end of the robotic arm, and calculates the target pose of the end of the robotic arm in real time. Moving the robotic arm according to the target position includes:

[0037] Place a spatial reference marker within the workspace corresponding to the end of the robotic arm. The spatial reference marker consists of multiple control points with known coordinates;

[0038] Based on the monocular camera completely capturing the spatial reference marker as a reference, when the robotic arm is in the initial position, record the initial joint angles of the robotic arm, calculate the initial pose of the spatial reference marker relative to the monocular camera, and use the initial pose as the reference pose;

[0039] The monocular camera captures the image corresponding to the spatial reference marker in real time and calculates the current pose of the spatial reference marker relative to the monocular camera;

[0040] According to the reference pose and the current pose, calculate the pose change amount of the spatial reference marker. Calculate the joint angle change amount based on the pose change amount and the initial joint angles, and compensate the joint angle change amount into the current joint angles of the robotic arm to obtain the target joint angles of the end of the robotic arm. Move the robotic arm based on the target joint angles.

[0041] In an alternative embodiment,

[0042] The monocular camera capturing the image corresponding to the spatial reference marker in real time includes:

[0043] Use a 4x4 homogeneous transformation matrix to represent the relative pose between coordinate systems. The homogeneous transformation matrix means the pose of coordinate system j relative to coordinate system i.

[0044] Initially, the relative position of the end of the robotic arm and the spatial reference marker is the reference, and the formula is as follows:

[0045]

[0046] where, represents the constant homogeneous transformation matrix between the camera coordinate system and the target pose coordinate system, Represents the initial homogeneous transformation matrix between the camera coordinate system and the manipulator base coordinate system, Represents the initial homogeneous transformation matrix between the manipulator base coordinate system and the target pose coordinate system;

[0047] After the spatial reference marker moves, the relative pose between the spatial reference marker and the end of the manipulator remains unchanged, satisfying the following formula:

[0048]

[0049] Where, Represents the homogeneous transformation matrix of the camera coordinate system relative to the manipulator base coordinate system, Represents the homogeneous transformation matrix of the target pose coordinate system relative to the manipulator base coordinate system;

[0050] At the same time, the pose of the end of the manipulator in the base coordinate system of the manipulator satisfies the following formula:

[0051]

[0052] Where, Represents the homogeneous transformation matrix of the manipulator end coordinate system relative to the manipulator base coordinate system, Represents the homogeneous transformation matrix of the transformation matrix of the camera coordinate system relative to the manipulator end coordinate system;

[0053] When the spatial reference marker moves to a new position, based on the base coordinate system, the new pose formula is as follows:

[0054]

[0055] Where, Represents the homogeneous transformation matrix of the new position of the target pose relative to the camera coordinate system, Represents the homogeneous transformation matrix of the new position of the target pose relative to the manipulator base coordinate system, Represents the homogeneous transformation matrix of the new position of the target pose relative to the camera coordinate system.

[0056] In the second aspect of the embodiments of the present invention,

[0057] A monocular vision pose estimation system based on the EPnP algorithm is provided, including:

[0058] A first unit for obtaining the internal parameters of a monocular camera through camera calibration; determining the world coordinates of feature points on a spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; representing the world coordinates of the feature points as a linear combination of four non-coplanar virtual control points, establishing a camera perspective projection model, forming a system of linear equations regarding the camera coordinates of the virtual control points, and determining the camera coordinates corresponding to the virtual control points in the camera coordinate system; wherein the monocular camera is fixed at the end of a robotic arm, and the optical axis of the monocular camera coincides with the center point of the end of the robotic arm;

[0059] A second unit for introducing a parallel perspective projection model, obtaining an initial estimate of the camera pose through an iterative parallel perspective solution process, determining the control point coordinates in the camera coordinate system based on the initial estimate of the camera pose, and calculating the initial values of the homogeneous barycentric coordinate coefficients of the control points; taking the reduction of the distance difference between the control points in the camera coordinate system and the world coordinate system as an optimization objective, and using the initial values of the homogeneous barycentric coordinate coefficients of the control points, solving through an iterative optimization algorithm to obtain the accurate pose of the spatial reference marker relative to the camera;

[0060] A third unit for using the initial pose of the spatial reference marker relative to the end of the robotic arm as a reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculates the current pose of the end of the robotic arm, combines the reference pose and the current pose of the end of the robotic arm, calculates the target pose of the end of the robotic arm in real time, and moves the robotic arm according to the target position.

[0061] In a third aspect of the embodiments of the present invention,

[0062] There is provided an electronic device, comprising:

[0063] A processor;

[0064] A memory for storing instructions executable by the processor;

[0065] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0066] In a fourth aspect of the embodiments of the present invention,

[0067] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0068] In the embodiments of the present invention, an iterative optimization algorithm can be used to obtain the accurate pose of the spatial reference marker relative to the camera, improving the accuracy of pose estimation. By introducing the parallel perspective projection model and the iterative parallel perspective solving process, an initial estimate of the camera pose can be quickly obtained without an initial pose, providing an effective starting point for further iterative optimization, not only improving the convergence speed of the algorithm but also enhancing the accuracy of the final pose estimation. By fixing the monocular camera at the end of the robotic arm and capturing the new pose of the spatial reference marker in real time, the method can calculate and adjust the target pose of the end of the robotic arm in real time to adapt to the position change of the spatial reference marker. The real-time pose tracking and adjustment mechanism provides strong support for the precise operation and automated operation of the robotic arm. Since it does not depend on the specific shape or size of the spatial reference marker and only requires identifiable feature points to be arranged within the camera's field of view, it has high flexibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a schematic flow chart of the monocular vision pose estimation method based on the EPnP algorithm in the embodiments of the present invention;

[0070] Figure 2 is a schematic structural diagram of the monocular vision pose estimation system based on the EPnP algorithm in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0072] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0073] Figure 1 is a schematic flow chart of the monocular vision pose estimation method based on the EPnP algorithm in the embodiments of the present invention, as Figure 1 shown, the method includes:

[0074] S101. Obtain the internal parameters of the monocular camera through camera calibration; determine the world coordinates of the feature points on the spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; represent the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establish a camera perspective projection model, form a linear equation system regarding the camera coordinates of the virtual control points, and determine the camera coordinates corresponding to the virtual control points in the camera coordinate system; wherein, the monocular camera is fixed at the end of the robotic arm, and the optical axis of the monocular camera coincides with the center point of the end of the robotic arm;

[0075] S102. Introduce a parallel perspective projection model, obtain an initial estimate of the camera pose through an iterative parallel perspective solution process, determine the control point coordinates in the camera coordinate system based on the initial estimate of the camera pose, and calculate the initial values of the homogeneous barycentric coordinate coefficients of the control points; taking the reduction of the distance difference between the control points in the camera coordinate system and the world coordinate system as the optimization objective, use the initial values of the homogeneous barycentric coordinate coefficients of the control points, and solve through an iterative optimization algorithm to obtain the accurate pose of the spatial reference marker relative to the camera;

[0076] S103. Take the initial pose of the spatial reference marker relative to the end of the robotic arm as the reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculate the current pose of the end of the robotic arm, combine the reference pose and the current pose of the end of the robotic arm, calculate the target pose of the end of the robotic arm in real time, and move the robotic arm according to the target position.

[0077] In an alternative embodiment,

[0078] Obtain the internal parameters of the monocular camera through camera calibration; determine the world coordinates of the feature points on the spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; represent the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establish a camera perspective projection model, and form a linear equation system regarding the camera coordinates of the virtual control points, and determining the camera coordinates corresponding to the virtual control points in the camera coordinate system includes:

[0079] Calibrate the monocular camera using the Zhang Zhengyou calibration method to obtain the internal parameter matrix of the monocular camera, including the camera focal length and the optical center coordinates;

[0080] Take the checkerboard within the camera's field of view as the planar reference marker, detect the pixel coordinates of the corner points of the planar reference marker through a feature extraction algorithm, select n of these corner points as feature points, and calculate the world coordinates of the feature points in the checkerboard coordinate system corresponding to the checkerboard based on the size of the checkerboard and the arrangement of the corner points;

[0081] In the chessboard coordinate system, select 4 non-coplanar points as virtual control points, establish the transformation relationship from the chessboard coordinate system to the virtual control point coordinate system corresponding to the virtual control points, and represent the world coordinates of the feature points through coordinate transformation and homogeneous barycentric coordinates to obtain the coordinate representation of the feature points in the virtual control point coordinate system;

[0082] Based on the internal parameter matrix, combined with the world coordinates of the feature points and the coordinate representation of the feature points in the virtual control point coordinate system, establish a system of equations for the camera imaging perspective projection model;

[0083] In the system of equations for the camera imaging perspective projection model, introduce the unknown coordinate representation of the virtual control points in the camera coordinate system, transform the system of equations for the camera imaging perspective projection model into a system of linear equations for the control points in the camera coordinate system, and solve the system of linear equations through the singular value decomposition numerical optimization method to obtain the coordinates of the 4 virtual control points in the camera coordinate system.

[0084] In an alternative embodiment,

[0085] It further includes:

[0086] The internal parameter matrix of the monocular camera has the following formula:

[0087]

[0088] where K represents the internal parameter matrix, f x represents the focal length of the camera in the x-axis direction, f y represents the focal length of the camera in the y-axis direction, u0 represents the abscissa of the optical center of the camera, and v0 represents the ordinate of the optical center of the camera;

[0089] The system of equations for the camera imaging perspective projection model has the following formula:

[0090]

[0091] where j represents the virtual control point index, i represents the feature point index, α ij represents the homogeneous barycentric coordinate of the i-th feature point in the virtual control point coordinate system, x j c represents the x component of the j-th virtual control point in the camera coordinate system, y j c represents the y component of the j-th virtual control point in the camera coordinate system, z j c represents the z component of the j-th virtual control point in the camera coordinate system, u i represents the abscissa of the i-th feature point on the image plane, v iDenote the ordinate of the \(i\)-th feature point on the image plane.

[0092] In an alternative embodiment,

[0093] Introduce a parallel perspective projection model, obtain an initial estimate of the camera pose through an iterative parallel perspective solution process, determine the coordinates of control points in the camera coordinate system based on the initial estimate of the camera pose, and calculate the initial values of the homogeneous barycentric coordinate coefficients of the control points; with the goal of minimizing the distance difference between the control points in the camera coordinate system and the world coordinate system, use the initial values of the homogeneous barycentric coordinate coefficients of the control points, and solve through an iterative optimization algorithm to obtain the accurate pose of the spatial reference marker relative to the camera, including:

[0094] By establishing a parallel perspective projection model, reconstruct the system of equations of the camera imaging perspective projection model into a linear system of equations regarding the control point coordinates and the homogeneous barycentric coordinate coefficients;

[0095] By iteratively solving the system of equations of the parallel perspective projection model, obtain an initial estimate of the camera pose, and based on the initial estimate, calculate the initial coordinates of the control points in the camera coordinate system. Based on the initial coordinates of the control points in the camera coordinate system, solve the homogeneous barycentric coordinate system of equations by the least squares method to obtain the initial values of the homogeneous barycentric coordinate coefficients of the control points;

[0096] Based on the camera pose and the initial values of the homogeneous barycentric coordinate coefficients of the control points, construct an objective function based on the distance difference between the control points in the camera coordinate system and the world coordinate system, minimize the objective function to obtain the optimal camera pose estimate, and obtain the accurate pose of the spatial reference marker relative to the camera.

[0097] In an alternative embodiment,

[0098] The objective function includes:

[0099] Its formula is as follows:

[0100]

[0101] Where \(\beta\) represents the optimization objective, \(l\), \(j\) represent the virtual control point indices, \(C\) l c represents the coordinates of the \(l\)-th virtual control point in the camera coordinate system, \(C\) j c represents the coordinates of the \(j\)-th virtual control point in the camera coordinate system, \(C\) l w represents the coordinates of the \(l\)-th virtual control point in the world coordinate system, \(C\) j w represents the coordinates of the \(j\)-th virtual control point in the world coordinate system.

[0102] In an alternative embodiment,

[0103] The initial pose of the spatial reference marker relative to the end of the robotic arm is used as the reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculates the current pose of the end of the robotic arm, combines the reference pose and the current pose of the end of the robotic arm, and calculates the target pose of the end of the robotic arm in real time. Moving the robotic arm according to the target position includes:

[0104] Place a spatial reference marker within the working space corresponding to the end of the robotic arm, and the spatial reference marker consists of multiple control points with known coordinates;

[0105] Based on the monocular camera completely capturing the spatial reference marker as a reference, when the robotic arm is in the initial position, record the initial joint angles of the robotic arm, calculate the initial pose of the spatial reference marker relative to the monocular camera, and use the initial pose as the reference pose;

[0106] The monocular camera captures the image corresponding to the spatial reference marker in real time and calculates the current pose of the spatial reference marker relative to the monocular camera;

[0107] According to the reference pose and the current pose, calculate the pose change amount of the spatial reference marker, calculate the joint angle change amount according to the pose change amount and the initial joint angles, and compensate the joint angle change amount into the current joint angles of the robotic arm to obtain the target joint angles of the end of the robotic arm. Based on the target joint angles, move the robotic arm.

[0108] In an alternative embodiment,

[0109] The monocular camera capturing the image corresponding to the spatial reference marker in real time includes:

[0110] Using a 4x4 homogeneous transformation matrix to represent the relative pose between coordinate systems, and the homogeneous transformation matrix means the pose of coordinate system j relative to coordinate system i.

[0111] Initially, the relative position of the end of the robotic arm and the spatial reference marker is a reference, and the formula is as follows:

[0112]

[0113] Wherein, represents the constant homogeneous transformation matrix between the camera coordinate system and the target pose coordinate system, Represents the initial homogeneous transformation matrix between the camera coordinate system and the manipulator base coordinate system, Represents the initial homogeneous transformation matrix between the manipulator base coordinate system and the target pose coordinate system;

[0114] After the spatial reference marker moves, the relative pose between the spatial reference marker and the end of the manipulator remains unchanged, satisfying the following formula:

[0115]

[0116] Where, Represents the homogeneous transformation matrix of the camera coordinate system relative to the manipulator base coordinate system, Represents the homogeneous transformation matrix of the target pose coordinate system relative to the manipulator base coordinate system;

[0117] At the same time, the pose of the end of the manipulator in the base coordinate system of the manipulator satisfies the following formula:

[0118]

[0119] Where, Represents the homogeneous transformation matrix of the manipulator end coordinate system relative to the manipulator base coordinate system, Represents the homogeneous transformation matrix of the transformation matrix of the camera coordinate system relative to the manipulator end coordinate system;

[0120] The spatial reference marker moves to a new position. Based on the base coordinate system, the new pose formula is as follows:

[0121]

[0122] Where, Represents the homogeneous transformation matrix of the new position of the target pose relative to the camera coordinate system, Represents the homogeneous transformation matrix of the new position of the target pose relative to the manipulator base coordinate system, Represents the homogeneous transformation matrix of the new position of the target pose relative to the camera coordinate system.

[0123] In this embodiment, by representing the world coordinates of feature points as a linear combination of non-coplanar virtual control points and establishing an accurate camera perspective projection model, the method can obtain the accurate pose of the spatial reference marker relative to the camera through an iterative optimization algorithm, improving the accuracy of pose estimation; by introducing a parallel perspective projection model and an iterative parallel perspective solution process, the method can quickly obtain an initial estimate of the camera pose without an initial pose, providing an effective starting point for further iterative optimization, not only improving the convergence speed of the algorithm but also the accuracy of the final pose estimation; by fixing a monocular camera at the end of the robotic arm and capturing the new pose of the spatial reference marker in real time, the method can calculate and adjust the target pose at the end of the robotic arm in real time to adapt to the position change of the spatial reference marker. The real-time pose tracking and adjustment mechanism provides strong support for the accurate operation and automated operation of the robotic arm; the method does not depend on the specific shape or size of the spatial reference marker, only requires identifiable feature points to be arranged within the camera's field of view, and has high flexibility; since the method only requires a monocular camera as a visual sensor and realizes pose estimation and pose tracking through software algorithms, compared with systems that require multiple sensors or high-cost hardware, the method can reduce the hardware cost and complexity of the system while ensuring high precision.

[0124] A specific embodiment of this application is as follows:

[0125] The robotic arm tracking the patient's movement mainly includes two key steps. The first is that the monocular camera at the end of the robotic arm obtains the position and pose of the target marker. When the target moves, the displacement and direction of the target movement can be calculated. The second is that the robotic arm moves the same distance in the same direction according to the displacement and direction of the target movement, keeping the relative position between the end of the robotic arm and the target unchanged, thereby realizing tracking the patient's movement.

[0126] First, the monocular camera obtains the pose information of the target marker;

[0127] The monocular estimating the target pose estimates the position and pose of the camera according to a set of spatial point and image point pair information. A spatial point in the target coordinate system {O w -X w Y w Z w} has coordinates in the camera coordinate system {O c -X c Y c Z c} as and pixel coordinates p in the image plane coordinate system {O f -uv} are i (u i , vi )。After the camera is calibrated, the parameters of the camera internal parameter matrix can be obtained: focal length f, optical center coordinates (u0, v0), and physical length of a single pixel (d x , d y ), let f x = f / d x 、f u = f / d y , and construct the internal parameter matrix of the camera:

[0128]

[0129] According to the camera perspective projection model, the coordinates of a spatial point in the camera coordinate system and the projection coordinates p i (u i , v i ) have the following relationship:

[0130]

[0131] In the formula, λ i represents the depth factor, and K represents the internal parameter matrix of the camera;

[0132] The fiducial point P i 's coordinates in the camera coordinate system and its coordinates in the world coordinate system satisfy the following relationship:

[0133]

[0134] In the formula, R ∈ SO(3), representing the rotation relationship between the world coordinate system and the camera coordinate system, and t ∈ R 3 represents the translation relationship between the world coordinate system and the camera coordinate system.

[0135] Through the above relationship model, when the camera internal parameter K is known, through n pairs of matching image coordinates {p i , i = 1, 2,..., n} and their corresponding world coordinate system coordinates to calculate the relative pose R and t between the camera coordinate system and the world coordinate system; for a point in the world coordinate system, it can be expressed as a linear combination of 4 non-coplanar virtual control points :

[0136]

[0137] Among them, α ij is the homogeneous barycentric coordinate of the feature point , satisfying And theoretically, non-coplanar virtual control points It can be arbitrarily selected;

[0138] In particular, for the same fiducial point P i the homogeneous barycentric coordinate relationship with the control points remains unchanged under different coordinate systems, that is, the homogeneous barycentric coordinates of the fiducial point P in the camera coordinate system i and the control points also satisfy the relationship:

[0139]

[0140] Therefore, the camera perspective projection model can be modified to:

[0141]

[0142] After arranging this formula, a system of equations about can be obtained:

[0143]

[0144] For n fiducial points, we can obtain 2n equations:

[0145] M 2n×12 x 12×1 = 0 (8)

[0146] In the formula the solution x of the system of equations lies in the null space of the coefficient matrix M and can be expressed as:

[0147]

[0148] In the formula, v i is the eigenvector corresponding to the zero eigenvalue of M T When the coefficient β i is determined, the coordinates of the 4 virtual control points in the camera coordinate system are also determined, and thus the coordinates of the fiducial points in the camera coordinate system can be calculated After obtaining the coordinates of the fiducial points in the world coordinate system and the camera coordinate system, the pose of the camera can be calculated according to the solution method of the absolute orientation problem;

[0149] After solving for the value of the coefficient β i the accuracy of the solution can be further improved by Gauss - Newton optimization. The goal of Gauss - Newton optimization is to reduce the distance difference between the control points in the camera coordinate system and the world coordinate system, and its objective function is:

[0150]

[0151] The EPnP algorithm can achieve good results when the number of 3D-2D points n > 6. However, when the number of 3D-2D points is small, such as n = 4 or 5, the EPnP+GN algorithm is unstable. The factor causing the instability of the algorithm is the coefficient β obtained by the EPnP algorithm i whose initial value sometimes seriously deviates from the correct value, resulting in the subsequent Gauss-Newton optimization being unable to converge correctly. Therefore, the EPnP algorithm is improved by introducing a parallel perspective projection model. Through it, an initial estimate R0 and t0 of the pose are obtained. According to the relationship the coordinates of the control points in the camera coordinate system can be obtained and then Since:

[0152]

[0153] Therefore, after knowing V and x, the coefficient β = [β1, β2, β3, β4] can be obtained T whose initial value is β0 = V T V) -1 V T x;

[0154] The parallel perspective projection model is used to approximate the pose solution under the perspective projection model through an iterative parallel perspective solution process. It can be regarded as a first-order approximation of the perspective projection. Let the camera pose R = [i, j, k] T and t = [t x , t y , t z T , then the parallel perspective projection equation is expressed as:

[0155]

[0156] where u0 = t x / t z , v0 = t y / t z , define:

[0157]

[0158] Then I p and J p can be linearly solved from Equation (12), and then by taking the modulus of I p and J p two solutions of t z are obtained. Taking the mean of the two solutions as the estimate of t z , the three elements of the translation vector can be obtained as follows:

[0159] ​

[0160] According to the orthogonality of R, we have:

[0161]

[0162] If we let [·] × represent the skew-symmetric matrix corresponding to the three-dimensional vector, then the above equation gives k as:

[0163]

[0164] Substituting it into Equation (12) gives the corresponding i and j;

[0165] After obtaining i and j, an initial estimate R0 and t0 of the camera pose can be obtained. According to the relationship the coordinates of the control point in the camera coordinate system can be obtained Furthermore, we can obtain Furthermore, optimizing the Gauss-Newton optimization objective function can optimize the accurate pose R and t:

[0166]

[0167] Second, the robotic arm tracks the moving target marker;

[0168] The camera is fixed at the end of the robotic arm to track the Marker (goal). During this process, the camera records the position of the target marker. When the target marker moves, the end of the robotic arm tracks the target marker to keep the relative pose between the end and the target unchanged;

[0169] Now assume that the pose of the marker recorded by the camera (cam) in space is The relative position reference between the end of the robotic arm and the target at the initial time is After moving the target marker, to keep the relative pose between the end and the target unchanged, the pose of the camera in the robotic arm base coordinate system (base) should satisfy:

[0170]

[0171] The pose of the end in the base coordinate system (base) should satisfy:

[0172]

[0173] When the target marker moves to a new position the new pose of the end in the base coordinate system (base) should satisfy:

[0174]

[0175] In the above formula,

[0176] (1) is the pose of the current end of the robotic arm in the base coordinate system, which is known;

[0177] (2) is the pose of the end marker relative to the end of the robotic arm, which is a constant and can be obtained by hand-eye calibration and is known;

[0178] (3) is the pose of the two markers in the camera coordinate system, which is known;

[0179] (4) is the relative position reference recorded initially, which is known;

[0180] Therefore, according to Equation (20), when the target moves, the homogeneous matrix that the robotic arm needs to move to track the target can be calculated.

[0181] In this embodiment, by using the EPnP algorithm and its improvement, combined with the Gauss-Newton optimization method, the method can accurately estimate the pose of the target marker in the case of a small number of 3D-2D corresponding points; the method is particularly suitable for practical application scenarios with a limited number of landmark points; the introduced parallel perspective projection model provides a relatively accurate initial pose estimate for the EPnP algorithm, making the subsequent optimization process more robust and greatly improving the stability and accuracy of the algorithm when the number of landmark points is small; the method captures the pose change of the target marker in real time and calculates the distance and direction that the robotic arm needs to move to keep the relative position with the target unchanged; the real-time response ability enables the method to be applicable to dynamically changing scenarios, such as instrument tracking in medical surgery or sample operation in the laboratory; by automatically capturing the position and pose of the target marker and calculating the corresponding robotic arm movement, it reduces the complexity of the operation and the dependence on professional operators, improving the safety and efficiency of the operation; the method is not limited by the shape and size of specific markers, and as long as the pose of the marker can be accurately captured by a monocular camera, precise tracking can be achieved, which has flexibility and enables the method to be widely applied to a variety of different occasions and tasks; the embodiment demonstrates a solution that combines advanced vision algorithms and mechanical control technologies, providing a new technical route for fields such as robotic arm control and machine vision recognition, and having good promotion and application prospects.

[0182] Figure 2 is a schematic structural diagram of the monocular vision pose estimation system based on the EPnP algorithm according to an embodiment of the present invention, as Figure 2 shown, the system includes:

[0183] A first unit is configured to obtain the internal parameters of a monocular camera through camera calibration; determine the world coordinates of feature points on a spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; represent the world coordinates of the feature points as a linear combination of four non-coplanar virtual control points, establish a camera perspective projection model, form a system of linear equations regarding the camera coordinates of the virtual control points, and determine the camera coordinates of the virtual control points corresponding to the camera coordinate system; wherein, the monocular camera is fixed at the end of a robotic arm, and the optical axis of the monocular camera coincides with the central point of the end of the robotic arm.

[0184] A second unit is configured to introduce a parallel perspective projection model, obtain an initial estimate of the camera pose through an iterative parallel perspective solving process, determine the coordinates of control points in the camera coordinate system based on the initial estimate of the camera pose, and calculate the initial values of the homogeneous barycentric coordinate coefficients of the control points; taking the reduction of the distance difference between the control points in the camera coordinate system and the world coordinate system as an optimization objective, and using the initial values of the homogeneous barycentric coordinate coefficients of the control points, solve through an iterative optimization algorithm to obtain the accurate pose of the spatial reference marker relative to the camera.

[0185] A third unit is configured to use the initial pose of the spatial reference marker relative to the end of the robotic arm as a reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculates the current pose of the end of the robotic arm, combines the reference pose and the current pose of the end of the robotic arm, calculates the target pose of the end of the robotic arm in real time, and moves the robotic arm according to the target position.

[0186] In a third aspect of the embodiments of the present invention,

[0187] There is provided an electronic device, including:

[0188] A processor;

[0189] A memory for storing instructions executable by the processor;

[0190] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0191] In a fourth aspect of the embodiments of the present invention,

[0192] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0193] The present invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A monocular vision pose estimation method based on the EPnP algorithm, characterized in that Including: Obtaining the internal parameters of a monocular camera through camera calibration; Determining the world coordinates of the feature points on the spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; expressing the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establishing a camera perspective projection model, forming a linear equation system regarding the camera coordinates of the virtual control points, and determining the camera coordinates corresponding to the virtual control points in the camera coordinate system; wherein, the monocular camera is fixed at the end of the robotic arm, and the optical axis of the monocular camera coincides with the center point of the end of the robotic arm; Introducing a parallel perspective projection model, obtaining an initial estimate of the camera pose through an iterative parallel perspective solution process, determining the control point coordinates in the camera coordinate system based on the initial estimate of the camera pose, and calculating the initial values of the homogeneous barycentric coordinate coefficients of the control points; taking reducing the distance difference between the control points in the camera coordinate system and the world coordinate system as the optimization objective, and using the initial values of the homogeneous barycentric coordinate coefficients of the control points, solving through an iterative optimization algorithm to obtain the accurate pose of the spatial reference marker relative to the camera; Taking the initial pose of the spatial reference marker relative to the end of the robotic arm as the reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculates the current pose of the end of the robotic arm, combines the reference pose and the current pose of the end of the robotic arm, calculates the target pose of the end of the robotic arm in real time, and moves the robotic arm according to the target pose; The monocular camera captures the image corresponding to the spatial reference marker in real time, including: Using a 4x4 homogeneous transformation matrix to represent the relative pose between coordinate systems, the homogeneous transformation matrix means the pose of coordinate system j relative to coordinate system i ; Initially, the relative position between the end of the robotic arm and the spatial reference marker is a reference, and the formula is as follows: ; Among them, represents the constant homogeneous transformation matrix between the camera coordinate system and the target pose coordinate system, represents the initial homogeneous transformation matrix between the camera coordinate system and the robot base coordinate system, represents the initial homogeneous transformation matrix between the robot base coordinate system and the target pose coordinate system; After the spatial reference marker moves, the relative pose between the spatial reference marker and the end of the robotic arm remains unchanged, satisfying the following formula: ; Among them, represents the homogeneous transformation matrix of the camera coordinate system relative to the manipulator base coordinate system, represents the homogeneous transformation matrix of the target pose coordinate system relative to the manipulator base coordinate system; Meanwhile, the pose of the end of the robotic arm in the base coordinate system of the robotic arm satisfies the following formula: ; Among them, represents the homogeneous transformation matrix of the coordinate system at the end of the robotic arm relative to the base coordinate system of the robotic arm, represents the homogeneous transformation matrix of the transformation matrix of the camera coordinate system relative to the coordinate system at the end of the robotic arm; When the spatial reference marker moves to a new position, based on the base coordinate system, the new pose formula is as follows: ; Among them, represents the homogeneous transformation matrix of the new position of the target pose relative to the camera coordinate system, represents the homogeneous transformation matrix of the new position of the target pose relative to the robot arm base coordinate system, represents the homogeneous transformation matrix of the new position of the target pose relative to the camera coordinate system.

2. The method according to claim 1, wherein Obtaining the internal parameters of a monocular camera through camera calibration; determining the world coordinates of the feature points on the spatial reference marker arranged within the field of view of the monocular camera, and the pixel coordinates of the feature points on the camera image; expressing the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establishing a camera perspective projection model, forming a linear equation system regarding the camera coordinates of the virtual control points, and determining the camera coordinates corresponding to the virtual control points in the camera coordinate system, including: Calibrating the monocular camera using the Zhang Zhengyou calibration method to obtain the internal parameter matrix of the monocular camera, including the camera focal length and the optical center coordinates; Taking the checkerboard within the camera's field of view as a planar reference marker, detecting the pixel coordinates of the corner points of the planar reference marker through a feature extraction algorithm, selecting n of the corner points as feature points, and calculating the world coordinates of the feature points in the checkerboard coordinate system corresponding to the checkerboard based on the size of the checkerboard and the arrangement of the corner points; In the chessboard coordinate system, select 4 non-coplanar points as virtual control points, establish the transformation relationship from the chessboard coordinate system to the virtual control point coordinate system corresponding to the virtual control points, and represent the world coordinates of the feature points through coordinate transformation and homogeneous barycentric coordinates to obtain the coordinate representation of the feature points in the virtual control point coordinate system; Based on the internal parameter matrix, combine the world coordinates of the feature points and the coordinate representation of the feature points in the virtual control point coordinate system to establish a camera imaging perspective projection model equation set; In the camera imaging perspective projection model equation set, introduce the unknown coordinate representation of the virtual control points in the camera coordinate system, transform the camera imaging perspective projection model equation set into a linear equation set of the control points in the camera coordinate system, and solve the linear equation set through the singular value decomposition numerical optimization method to obtain the coordinates of the 4 virtual control points in the camera coordinate system.

3. The method according to claim 2, characterized in that It also includes: The internal parameter matrix of the monocular camera, and its formula is as follows: ; Among them, K represents the intrinsic parameter matrix, f x represents the focal length of the camera in the x axis direction, f y represents the focal length of the camera in the y axis direction, u 0 represents the abscissa of the optical center of the camera, v 0 represents the ordinate of the optical center of the camera; The camera imaging perspective projection model equation set, and its formula is as follows: ; Among them, j represents the virtual control point index, i represents the feature point index, α ij represents the i th homogeneous barycentric coordinate of the feature point in the virtual control point coordinate system, x j c represents the j th x component of the virtual control point in the camera coordinate system, y j c represents the j th y component of the virtual control point in the camera coordinate system, z j c represents the j th z component of the virtual control point in the camera coordinate system, u i represents the i th abscissa of the feature point on the image plane, v i represents the i th ordinate of the feature point on the image plane.

4. The method according to claim 1, wherein Introduce a parallel perspective projection model, obtain an initial estimate of the camera pose through an iterative parallel perspective solution process, determine the control point coordinates in the camera coordinate system based on the initial estimate of the camera pose, and calculate the initial values of the control point homogeneous barycentric coordinate coefficients; Taking the reduction of the distance difference between the control points in the camera coordinate system and the world coordinate system as the optimization goal, use the initial values of the control point homogeneous barycentric coordinate coefficients, and solve through an iterative optimization algorithm to obtain the precise pose of the spatial reference marker relative to the camera, including: By establishing a parallel perspective projection model, reconstruct the equation set of the camera imaging perspective projection model into a linear equation set about the control point coordinates and the homogeneous barycentric coordinate coefficients; By iteratively solving the equation set of the parallel perspective projection model, obtain an initial estimate of the camera pose, and calculate the initial coordinates of the control points in the camera coordinate system according to the initial estimate. Based on the initial coordinates of the control points in the camera coordinate system, solve the homogeneous barycentric coordinate equation set through the least squares method to obtain the initial values of the control point homogeneous barycentric coordinate coefficients; Based on the camera pose and the initial values of the control point homogeneous barycentric coordinate coefficients, construct an objective function based on the distance difference between the control points in the camera coordinate system and the world coordinate system, minimize the objective function to obtain the optimal camera pose estimate, and obtain the precise pose of the spatial reference marker relative to the camera.

5. The method according to claim 4, wherein The objective function includes: Its formula is as follows: ; Among them, β represents the optimization objective, j , l represents the virtual control point index, C j c represents the coordinates of the j th virtual control point in the camera coordinate system, C l c represents the coordinates of the l th virtual control point in the camera coordinate system, C j w represents the coordinates of the j th virtual control point in the world coordinate system, C l w represents the coordinates of the l th virtual control point in the world coordinate system.

6. The method according to claim 1, wherein Take the initial pose of the spatial reference marker relative to the end of the robotic arm as the reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculates the current pose of the end of the robotic arm, combines the reference pose and the current pose of the end of the robotic arm, and calculates the target pose of the end of the robotic arm in real time. Move the robotic arm according to the target position, including: Place a spatial reference marker in the working space corresponding to the end of the robotic arm, and the spatial reference marker consists of multiple control points with clear coordinates; Based on the complete capture of the spatial reference marker by the monocular camera as a benchmark, when the robotic arm is in the initial position, record the initial joint angles of the robotic arm, calculate the initial pose of the spatial reference marker relative to the monocular camera, and use the initial pose as the reference pose; The monocular camera captures the image corresponding to the spatial reference marker in real time and calculates the current pose of the spatial reference marker relative to the monocular camera; According to the reference pose and the current pose, calculate the pose change amount of the spatial reference marker, calculate the joint angle change amount according to the pose change amount and the initial joint angles, and compensate the joint angle change amount into the current joint angles of the robotic arm to obtain the target joint angles at the end of the robotic arm. Based on the target joint angles, move the robotic arm.

7. A monocular vision pose estimation system based on the EPnP algorithm, which is used to implement the method described in any one of the preceding claims 1-6, characterized in that, Comprising: A first unit for obtaining the internal parameters of the monocular camera through camera calibration; Determine the world coordinates of the feature points on the spatial reference marker arranged within the field of view of the monocular camera and the pixel coordinates of the feature points in the camera image; represent the world coordinates of the feature points as a linear combination of 4 non-coplanar virtual control points, establish a camera perspective projection model, form a linear equation system regarding the camera coordinates of the virtual control points, and determine the camera coordinates of the virtual control points corresponding to the camera coordinate system; wherein, the monocular camera is fixed at the end of the robotic arm, and the optical axis of the monocular camera coincides with the center point at the end of the robotic arm; A second unit for introducing a parallel perspective projection model, obtaining an initial estimate of the camera pose through an iterative parallel perspective solution process, determining the control point coordinates in the camera coordinate system based on the initial estimate of the camera pose, and calculating the initial values of the control point homogeneous barycentric coordinate coefficients; taking reducing the distance difference between the control points in the camera coordinate system and the world coordinate system as the optimization goal, using the initial values of the control point homogeneous barycentric coordinate coefficients, and solving through an iterative optimization algorithm to obtain the precise pose of the spatial reference marker relative to the camera; A third unit for using the initial pose of the spatial reference marker relative to the end of the robotic arm as the reference pose; when the spatial reference marker moves, the camera captures the new pose of the spatial reference marker in real time, calculates the current pose of the end of the robotic arm, combines the reference pose and the current pose of the end of the robotic arm, calculates the target pose of the end of the robotic arm in real time, and moves the robotic arm according to the target pose.

8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Robot 3D visual guidance grabbing method

    CN118143929A