Image visual servo bolt alignment method based on real-time depth estimation
Through the image visual servo method based on real-time depth estimation, the bolt depth is calculated using the YOLOv7 model and direct linear transformation, the accuracy and cost problems of bolt alignment in unstructured scenarios are solved, and efficient and accurate alignment of the robot in contact network maintenance is achieved.
Patent Information
- Application Number
- CN202510429800.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art is difficult to accurately obtain the depth information of the bolt in unstructured scenarios, resulting in high cost of visual servo hardware and insufficient control accuracy.
The image vision servo method based on real-time depth estimation is used to detect the bolt corner points through the YOLOv7 deep learning model, and the external parameter matrix of the bolt coordinate system to the camera coordinate system is calculated by combining direct linear transformation and singular value decomposition. The virtual vision servo iterative optimization is used to obtain accurate depth values, and finally the robotic arm alignment bolts are controlled using image vision servo.
It realizes accurate alignment of bolts by robots in unstructured scenarios, reduces hardware costs, improves control accuracy and efficiency, and is suitable for contact network maintenance of urban rail transit and high-speed rail networks.
Smart Images

Figure CN120347516A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and particularly to an image visual servo bolt alignment method based on real-time depth estimation. Background Art
[0002] With the rapid development of urban rail transit and high-speed rail networks, the maintenance and inspection of catenaries have become increasingly crucial. The catenary is an important part of power supply, and its safety and stability directly affect the normal operation of railway transportation. In the maintenance of catenaries, bolt tightening is of great significance. As an important connecting part of the catenary structure, the tightening state of bolts is related to the overall safety of the catenary.
[0003] Traditional manual inspection methods are not only inefficient but also have many potential safety hazards. Using robots for bolt tightening can significantly improve the operation efficiency, ensure higher tightening quality, and promote the automation and intelligence development of catenary maintenance. An important prerequisite for a robot to complete the bolt tightening operation is to ensure that the robot accurately aligns with the bolt. However, due to factors such as geographical environment and manual installation, there are significant differences in the catenary structures even on the same line. This poses a huge challenge to the robot's recognition and positioning of bolts.
[0004] To solve the above problems, a technical solution of using visual servo control to align the robot with the bolt has been proposed. Visual servo is a closed-loop control process that uses a visual sensor to perceive external information and guide the robot to complete a specific task. Its core is to combine the visual system with the robot control system, continuously adjust the robot's behavior through visual feedback, minimize the error between the current visual features and the desired visual features, and thus achieve the desired control goal. Using visual servo technology can significantly improve the accuracy, flexibility, and intelligence of the robot in completing operations in unstructured scenarios. According to different visual feedback information, visual servo can be divided into two categories: position-based visual servo (PBVS) and image-based visual servo (IBVS). Compared with PBVS, IBVS has less dependence on the camera calibration accuracy and does not require prior geometric knowledge of the target object. Two important prerequisites for using IBVS are to obtain accurate image feature point coordinates and corresponding depth information. In particular, the depth information will significantly affect the convergence speed of the servo control process, and inaccurate depth values may even lead to servo failure. However, it is very difficult to obtain accurate depth values. At present, deep learning models or advanced visual devices are usually adopted, but this will significantly increase the hardware cost. Summary of the Invention
[0005] To solve the problems that it is difficult to obtain accurate depth values in the prior art and the hardware cost of visual servo is relatively high, the present invention proposes an image visual servo bolt alignment method based on real-time depth estimation, which can enable the robotic arm to accurately align bolts in an unstructured scene and minimize the hardware cost as much as possible to solve the above problems.
[0006] This application discloses an image visual servo bolt alignment method based on real-time depth estimation, which is characterized by the following steps:
[0007] S1. Obtain the RGB image of the component through the camera, and use the YOLOv7 deep learning model for image processing to obtain the pixel coordinates of the bolt corner points;
[0008] S2. Based on the pixel coordinates of the bolt corner points and the actual size of the bolt, use the direct linear transformation algorithm to calculate the homography matrix, which describes the conversion relationship from the bolt coordinate system to the camera coordinate system and the conversion relationship from the camera coordinate system to the pixel coordinate system;
[0009] S3. Combine the homography matrix with the camera internal parameter matrix to calculate the truncated extrinsic parameter matrix of the camera. The camera internal parameter matrix describes the conversion relationship from the camera coordinate system to the pixel coordinate system;
[0010] S4. Through rotation matrix orthonormalization and singular value decomposition reconstruction, restore the complete extrinsic parameter matrix from the bolt coordinate system to the camera coordinate system. The extrinsic parameter matrix describes the conversion relationship from the bolt coordinate system to the camera coordinate system;
[0011] S5. Use virtual visual servo to iteratively optimize the complete extrinsic parameter matrix;
[0012] S6. Combine the optimized complete extrinsic parameter matrix with the bolt size to calculate the depth values of each corner point of the bolt relative to the camera;
[0013] S7. Use the obtained bolt corner point coordinates and depth values, and use the image-based visual servo method to control the robotic arm to align with the bolt.
[0014] Preferably, the image processing in S1 includes object detection, semantic segmentation, and line fitting operations.
[0015] Preferably, the calculation formula of the homography matrix in S2 is as follows:
[0016] The mapping relationship between the pixel coordinates of the bolt corner points and the corner point coordinates in the bolt coordinate system is:
[0017]
[0018] where are the corner point coordinates in the bolt coordinate system, are the corner coordinates in the pixel coordinate system, and H is the homography matrix, which is specifically expressed as follows:
[0019]
[0020] Based on the correspondence between the pixel coordinates of the four sets of bolt corner points and the two-dimensional homogeneous coordinates in the bolt coordinate system, the homography matrix H can be calculated and obtained.
[0021] Preferably, the step S3 includes the following steps:
[0022] The homography matrix is split into a 3×4 camera intrinsic matrix and a 4×3 truncated extrinsic matrix, and the formula is as follows:
[0023]
[0024] Among them, K is the camera intrinsic matrix, E is the truncated extrinsic matrix with the original third column deleted, f x , f y , u0, and v0 are the four intrinsic parameters of K. The complete extrinsic matrix [R, T] describes the transformation relationship from the bolt coordinate system O b -X b Y b Z b to the camera coordinate system O c -X c Y c Z c where R is the rotation matrix and T is the displacement vector, and the expression is:
[0025] Expand and solve Equation (4) to obtain the elements in the truncated extrinsic matrix E:
[0026]
[0027] Process the obtained truncated extrinsic matrix E by multiplying it by a scaling factor s:
[0028]
[0029] The magnitude of the scaling factor s is the reciprocal of the geometric mean of the lengths of the vectors [R1, R4, R7] T and [R2, R5, R8] T , and the sign is determined according to the sign of T Z .
[0030] Preferably, the step S4 includes the following steps:
[0031] The complete displacement vector T has been obtained through the truncated extrinsic matrix E, and the missing column vector of the rotation matrix R is restored through the cross product operation:
[0032]
[0033] The rotation matrix that is strictly unitary orthogonal is obtained by performing singular value decomposition and reconstruction on R:
[0034] R = USV T ; (8)
[0035] where U and V are 3×3 orthogonal matrices, S is a 3×3 non - negative diagonal matrix, and the reconstructed rotation matrix R is the product of matrices U and V;
[0036] Thus, the complete external parameter matrix Q = [R, T] is obtained, and then the complete external parameter matrix is further processed:
[0037]
[0038] where multiplying Q on the left by F means that the external parameter matrix Q rotates 180° around the X - axis of the true camera coordinate system, and the obtained is the complete external parameter matrix of the bolt coordinate system O b -X b Y b Z b relative to the camera coordinate system O c -X c Y c Z c .
[0039] Preferably, the S5 includes the following steps:
[0040] The homogeneous coordinates of the corner point in the bolt coordinate system O b -X b Y b Z b are b P = b P X , b P Y , 0, 1] T , combined with the complete external parameter matrix to obtain the homogeneous coordinates of the corner point in the camera coordinate system O c -X c Y c Z c as c P = c P X , c P Y , c P Z , 1] T , and the calculation formula is:
[0041]
[0042] Among them, c P Z is the depth value of the bolt corner point relative to the camera coordinate system O c -X c Y c Z c ;
[0043] Suppose there is a bolt corner point P i , and its homogeneous coordinate in the bolt coordinate system O b -X b Y b Z b is The pixel coordinate obtained by image detection is [u i , v i T , and according to the principle of the perspective model, the pixel coordinate [u i , v i T is projected onto the image coordinate system O-xy as follows:
[0044]
[0045] Using the complete extrinsic parameter matrix to perform forward projection to obtain the projection point of P i on the image coordinate system O-xy as Obtained from the principle of the perspective model:
[0046]
[0047] Set the desired feature of the virtual visual servo as Set the current feature as Set the servo error as Expand Equation (10) to obtain the depth value of the feature point as:
[0048]
[0049] The calculation formula for the image Jacobian matrix of the feature point is as follows:
[0050]
[0051] Add up all the of each corner point to obtain the virtual visual servo error e v , and then add up all the of each corner point to obtain the image Jacobian matrix L v of the virtual visual servo;
[0052] Perform visual servo control through the following formula:
[0053]
[0054] where λ v is the virtual visual servo gain, is the Moore-Penrose pseudoinverse of L v , and v vc is the spatial velocity of the virtual camera;
[0055] Use T to update the complete extrinsic parameter matrix The calculation formula is as follows:
[0056]
[0057] where k is the number of iterative optimizations, and the updated complete extrinsic parameter matrix is used for forward projection in the next iteration;
[0058] Define an error When err k+1 - err k < 10 -8 the virtual visual servo iterative optimization ends.
[0059] Preferably, the S6 includes the following steps:
[0060] Combine b P i = b P Xi , b P Yi , 0, 1] T with Equation (10) to obtain the accurate depth value of the bolt corner point P i c P Z .
[0061] Preferably, the S7 includes the following steps:
[0062] Obtain the corner pixel coordinates [u i , v i T from image detection, and convert it to the image coordinate system [x i , y i T through Equation (11), and substitute [x i , y i T and the corresponding depth value c P Zi into Equation (14) to obtain the image Jacobian matrix L i corresponding to each corner point;
[0063] Set the desired feature of the image-based visual servo to The current feature is set to s i = [x i , y i T , and the error is set to Sum up the e of each point i to obtain the servo error e, and sum up the L of each point i to obtain the image Jacobian matrix L, and obtain the camera speed v through the following formula c :
[0064] v c = -λL + e; (17)
[0065] where λ is the image visual servo gain, and L + is the Moore-Penrose pseudoinverse of L;
[0066] Use v c to perform motion control on the robotic arm, and repeat the execution of S1 to S7;
[0067] Define a total error of the image-based visual servo When Err < 0.00005, the image-based visual servo converges and the bolt alignment is completed.
[0068] Advantages of the present invention:
[0069] (1) The present invention only relies on accurate RGB data and does not adopt learning-based or adaptive technologies, so there are no excessive requirements for visual devices and computer computing power.
[0070] (2) The depth estimation algorithm proposed by the present invention does not use any relevant models of robotic arm kinematics and dynamics, which improves the generality of the algorithm. Finally, using the obtained bolt corner coordinates and depth values combined with IBVS can control the robot to accurately align the bolt. Description of the drawings
[0071] Figure 1 is the flowchart of the image visual servo bolt alignment method based on real-time depth estimation according to the embodiment of the present invention;
[0072] Figure 2 is the experimental scenario simulation diagram according to the embodiment of the present invention;
[0073] Figure 3 is the RGB image captured by the camera according to the embodiment of the present invention;
[0074] Figure 4 is the result diagram of the bolt corner points extracted by the image algorithm according to the embodiment of the present invention;
[0075] Figure 5 Bolt coordinate system diagram of an embodiment of the present invention;
[0076] Figure 6 Optical perspective projection model diagram of the camera of an embodiment of the present invention;
[0077] Figure 7 Corner depth estimation curve graph of an embodiment of the present invention;
[0078] Figure 8 Total servo error curve graph of an embodiment of the present invention. Detailed implementation manners
[0079] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the following takes examples with reference to the attached drawings and further elaborates on the present application in detail.
[0080] The present application discloses an image vision servo bolt alignment method based on real-time depth estimation, and the process is as Figure 1 shown.
[0081] In this embodiment, in the catenary environment built in the laboratory, taking an M12 bolt installed on the inclined cross-arm component as an example, as Figure 2 shown, the method steps proposed in this embodiment are described in detail. This embodiment uses a UR16e collaborative robot, and the vision device is an Intel Realsense D435 camera. The resolution of the camera is set to 1920×1080, and the frame rate is 30fps. It is installed at the end of the robotic arm in the eye in hand manner. The movement of the robotic arm is controlled using the official robot control and data exchange protocol RTDE (Real-Time Data Exchange). The whole process runs under the Ubuntu 20.04 LTS operating system, and the software architecture is based on the visual servo platform (VISP) and the robot operating system (ROS). The program algorithm is designed using the C++ language and the Python language. VISP is an open-source visual servo platform developed and designed by the IRISA-Inria team. The specific steps are as follows:
[0082] S1. Obtain the RGB image of the component through the camera, and use the YOLOv7 deep learning model for image processing to obtain the pixel coordinates of the bolt corner points.
[0083] Use the D435 camera to obtain a frame of RGB image of the whole component, as Figure 3 shown. Select the lowermost M12 bolt as the target bolt. Use the YOLOv7 deep learning model for image processing. After operations such as target detection, semantic segmentation, and straight line fitting, finally the pixel coordinates of the bolt corner points can be extracted. The extracted bolt corner points are asFigure 4 as shown
[0084] S2. Based on the pixel coordinates of the bolt corner points and the actual size of the bolt, the homography matrix is calculated using the direct linear transformation algorithm. The homography matrix describes the conversion relationship from the bolt coordinate system to the camera coordinate system and the conversion relationship from the camera coordinate system to the pixel coordinate system.
[0085] Since the selected bolt model is M12 and the distance between opposite sides of the hexagon on its surface is 18 mm, as Figure 5 shown. Assume that a bolt coordinate system O Figure 5 -X b -X b Y b Z b is established at the center of the bolt plane. Now, [X, Y, 1] T is used to represent a certain corner point in the bolt coordinate system. Originally, three-dimensional coordinates should be used, but since the Z coordinate on the bolt plane is all 0, the corner point is rewritten in the two-dimensional homogeneous form with Z implicitly zero. The coordinate of this bolt corner point corresponding to the pixel coordinate system O′-uv is [u, v, 1] Figure 6 , and the mapping relationship between the two is unique and can be represented by the homography matrix. This mapping relationship can be obtained by the direct linear transformation (DLT) algorithm. The mapping relationship between the pixel coordinates of the bolt corner points and the corner point coordinates in the bolt coordinate system is as follows: T where,
[0086]
[0087] where, is the corner point coordinate in the bolt coordinate system, is the corner point coordinate in the corresponding pixel coordinate system, and H is the homography matrix, which is expressed as follows:
[0088]
[0089] Based on the corresponding relationship between the pixel coordinates of four groups of bolt corner points and the two-dimensional implicit homogeneous coordinates in the bolt coordinate system, the homography matrix H can be calculated.
[0090] Since the number of point features usually selected for image-based visual servo is 4, so select Figure 4 points 1, 2, 4, 5 in Figure 5 and points 1, 2, 3, 4 in
[0091]
[0092] Substitute these four groups of corresponding points into Equation (1) to calculate the homography matrix H. Expand all the matrices to get the following equation: Use the Gaussian elimination method to solve Equation (3) and normalize the element h9 in H to 1. Finally, the required homography matrix H is obtained.
[0093] S3. Combine the homography matrix and the camera intrinsic matrix to calculate the truncated extrinsic matrix of the camera. The camera intrinsic matrix describes the conversion relationship from the camera coordinate system to the pixel coordinate system.
[0094] The homography matrix is split into a 3×4 camera intrinsic matrix and a 4×3 truncated extrinsic matrix, and the formula is as follows:
[0095]
[0096] Figure 6 is the optical perspective projection model used by the vision device. The homography matrix H corresponds to the projection transformation, which describes the conversion relationship from the bolt coordinate system O b -X b Y b Z b to the pixel coordinate system O′-uv. K is the camera intrinsic matrix, which describes the conversion relationship from the camera coordinate system O c -X c Y c Z c to the pixel coordinate system O′-uv. The four intrinsic parameters f x 、f y 、u0, v0 in K can be obtained in advance by camera intrinsic calibration. E is the truncated extrinsic matrix with the original third column deleted. The reason for truncation is that two-dimensional homogeneous points with Z implicitly zero are used to calculate the matrix H. The complete extrinsic matrix [R, T] describes the conversion relationship from the bolt coordinate system O b -X b Y b Z b to the camera coordinate system O c -X c Y c Z c where R is the rotation matrix and T is the displacement vector, and the expression is:
[0097] Expand and solve Equation (4) to obtain the elements in the truncated extrinsic matrix E:
[0098]
[0099] Since the column vectors of the rotation matrix are all unit vectors, the obtained truncated extrinsic matrix E needs to be multiplied by a scaling factor s for processing:
[0100]
[0101] The magnitude of the scaling factor s is the vectors [R1, R4, R7] T and [R2, R5, R8] TReciprocal of the geometric mean of the lengths, and the sign is determined according to T Z The sign of Z is determined by the sign of T. If T Z > 0, the sign of the scaling factor s is negative because Figure 6 the imaging plane of Figure 6 is a virtual plane, and the real imaging plane is behind the camera lens. In the real camera coordinate system, Z C and Y C will also be inverted. Therefore, the correct sign of T Z must remain negative.
[0102] In this way, the truncated extrinsic parameter matrix E is obtained.
[0103] S4. Through the orthonormalization of the rotation matrix and the reconstruction by singular value decomposition, the complete extrinsic parameter matrix from the bolt coordinate system to the camera coordinate system is restored. The extrinsic parameter matrix describes the conversion relationship from the bolt coordinate system to the camera coordinate system.
[0104] The complete displacement vector T = [T X , T Y , T Z has been obtained through the truncated extrinsic parameter matrix E T , but the rotation matrix R is not complete yet. According to the unit orthonormal relationship between the column vectors of the rotation matrix, the missing third column vector of R is restored by taking the cross product of the first two column vectors. The formula is as follows:
[0105]
[0106] The complete rotation matrix R can be obtained through the above formula, but it is not strictly unit orthonormal. Therefore, singular value decomposition reconstruction is still required:
[0107] R = USV T ; (8)
[0108] where U and V are 3×3 orthogonal matrices, and S is a 3×3 non - negative diagonal matrix. The reconstructed rotation matrix R is the product of matrices U and V.
[0109] Thus, the complete extrinsic parameter matrix Q = [R, T] is obtained. However, since the real imaging plane is behind the camera lens, the image formed will be upside - down. Therefore, Y C and Z C in the real camera coordinate system will be inverted compared to Figure 6 . And in this embodiment, the subsequent vision servo based on the image uses the Figure 6 camera coordinate system. Therefore, Q needs to be further processed:
[0110]
[0111] Among them, multiplying Q on the left by F means that the extrinsic parameter matrix Q rotates 180° around the X - axis of the real camera coordinate system, and the obtained is the bolt coordinate system O b -X b Y b Z b with respect to the camera coordinate system O c -X c Y c Z c of the complete extrinsic parameter matrix.
[0112] S5. Use virtual visual servo (VVS) to iteratively optimize the complete extrinsic parameter matrix.
[0113] If the actual size of the bolt is known in advance, the homogeneous coordinates of the corner points in the bolt coordinate system O b -X b Y b Z b are b P = b P X , b P Y , 0, 1] T . Combining with the complete extrinsic parameter matrix to obtain the homogeneous coordinates of the corner points in the camera coordinate system O c -X c Y c Z c are c P = c P X , c P Y , c P Z , 1] T , and the calculation formula is:
[0114]
[0115] where c P Z is the depth value of the bolt corner point with respect to the camera coordinate system O c -X c Y c Z c . This depth will be used for subsequent image-based visual servo to control the robot motion. However, the obtained by the DLT algorithm is inaccurate because the DLT algorithm cannot take into account the non-linear distortion. The inaccurate means that the estimated corner point depth c P Z is inaccurate. Therefore, virtual visual servo needs to be used to optimize the iterative matrix
[0116] Virtual visual servoing regards the pose matrix optimization problem as a visual servoing task based on 2D images. There is an error between the bolt corner points obtained by image detection and the corresponding corner points obtained by forward projection through the matrix. Define a virtual camera and use image-based visual servoing to minimize this error. When the error is minimized, the pose of the virtual camera at this time is the pose of the real camera.
[0117] Suppose there is a bolt corner point P I , in the bolt coordinate system O b -X B Y B Z b The homogeneous coordinate is The pixel coordinates obtained by image detection are [u i , v i T , According to the principle of the perspective model, the pixel coordinates [u i , v i T are projected onto Figure 6 the image coordinate system O-xy as
[0118]
[0119] Then use the complete extrinsic parameter matrix to perform forward projection to obtain the projection point of P i on the image coordinate system O-xy as Obtained from the principle of the perspective model:
[0120]
[0121] Set the expected feature of virtual visual servoing as Set the current feature as Set the servo error as Using virtual visual servoing also requires obtaining the depth value of the feature point, set as Expand Equation (10) to obtain the depth value of the feature point as:
[0122]
[0123] The calculation formula for the image Jacobian matrix of the feature point is as follows:
[0124]
[0125] Because there are four feature points, the of all four corner points are all added up to obtain the virtual visual servo error e v , and then the of the four corner points are added up to obtain the image Jacobian matrix L of virtual visual servoingv ;
[0126] Visual servo control is performed through the following formula:
[0127]
[0128] where λ v is the artificially set virtual visual servo gain, is the Moore-Penrose pseudoinverse of L v , and v vc is the spatial velocity of the virtual camera.
[0129] Since v vc ∈se(3), its exponential map corresponds to the displacement transformation matrix T∈SE(3), which represents the displacement of the virtual camera per second. Use T to update the complete extrinsic parameter matrix The calculation formula is as follows:
[0130]
[0131] where k is the number of iterative optimizations, and the updated complete extrinsic parameter matrix is used for forward projection in the next iteration;
[0132] Define an error When err k+1 - err k < 10 -8 the virtual visual servo iterative optimization ends. At this time, is accurate.
[0133] S6. Combine the optimized complete extrinsic parameter matrix with the bolt size to calculate the depth values of each corner point of the bolt relative to the camera. After obtaining the accurate , and knowing the actual size of the bolt, combine b P i = b P Xi , b P Yi , 0, 1] T Combine with Equation (10) to obtain the accurate depth value i of the bolt corner point P c P Z .
[0134] S7. Use the obtained bolt corner point coordinates and depth values to control the robotic arm to align with the bolt using the image-based visual servo (IBVS) method.
[0135] The corner point pixel coordinates [u i , v i are obtained by image detectionT , it is transformed into the image coordinate system [x i , y i through Equation (11). T , [x i , y i is substituted into Equation (14) together with the corresponding depth value T P c to obtain the image Jacobian matrix L corresponding to each corner point. Zi . i
[0136] Set the desired feature of image-based visual servo to Set the current feature to s i = [x i , y i . T Set the error to Stack the e i , L i of the four corner points respectively to obtain the servo error e and the image Jacobian matrix L. The camera velocity v is obtained through the following formula c :
[0137] v c = -λL + e; (17)
[0138] where λ is the image visual servo gain set manually, and L + is the Moore-Penrose pseudoinverse of L.
[0139] Use v c to control the motion of the robotic arm, and repeat the execution of S1 to S7.
[0140] Define the total error of image-based visual servo When Err < 0.00005, the image-based visual servo converges, and at this moment, the bolt alignment is just completed.
[0141] The depth estimation curves of the four corner points of the bolt in a certain control process of IBVS are as Figure 7 shown, and the total error curve of the servo is as Figure 8 shown. The depth before IBVS servo is set to 40 cm, and the set desired servo convergence depth is 20 cm, that is, when the camera plane is 20 cm away from the bolt plane, the bolt is just aligned. To verify the effectiveness, 10 verifications were carried out in total. The results show that the average convergence time of using IBVS to control the robot to align the bolt is 8.723 s, and the depth estimation accuracy is always kept within 2 cm, which can meet the requirements of automated operations and improve the maintenance work efficiency.
[0142] In summary, the present application aims to achieve the rapid alignment of a target bolt by a robot in an unstructured scenario. To achieve closed-loop control, visual feedback information is introduced, and the image-based visual servo method (IBVS) is used to control the movement of the robotic arm. The prerequisite for applying this method is to obtain accurate bolt corner coordinates and the depth value of the corner relative to the camera. To obtain accurate bolt corner coordinates in an unstructured scenario, deep learning methods such as the YOLOv7 model must be used. To obtain the accurate depth value of the bolt corner, the present application proposes a real-time depth estimation algorithm for coplanar feature points based on the principle of homography transformation. This algorithm only relies on accurate RGB data and does not adopt learning-based or adaptive methods, so it does not have high requirements for visual devices and computer computing power. Moreover, the depth estimation algorithm does not use any models related to the kinematics and dynamics of the robotic arm, improving the generality of the algorithm. Finally, by using the obtained bolt corner coordinates and depth value and combining with IBVS, the robot can be controlled to accurately align with the bolt.
[0143] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. An image vision servo bolt alignment method based on real-time depth estimation, characterized in that It includes the following steps: S1. Obtain the RGB image of the component through a camera, perform image processing using the YOLOv7 deep learning model, and obtain the pixel coordinates of the bolt corner points; S2. Based on the pixel coordinates of the bolt corner points and the actual size of the bolt, calculate the homography matrix using the direct linear transformation algorithm; S3. Combine the homography matrix and the camera internal parameter matrix to calculate the truncated external parameter matrix of the camera; S4. Reconstruct through rotation matrix orthonormalization and singular value decomposition to restore the complete external parameter matrix from the bolt coordinate system to the camera coordinate system; S5. Use virtual visual servo to iteratively optimize the complete external parameter matrix; S6. Combine the optimized complete external parameter matrix with the bolt size to calculate the depth values of each corner point of the bolt relative to the camera; S7. Use the obtained bolt corner point coordinates and depth values, and use the image-based visual servo method to control the robotic arm to align with the bolt.
2. The method for image vision servo bolt alignment based on real-time depth estimation according to claim 1, wherein, The image processing in S1 includes object detection, semantic segmentation, and line fitting operations.
3. The method for image vision servo bolt alignment based on real-time depth estimation according to claim 2, wherein The calculation formula of the homography matrix in S2 is as follows: The mapping relationship between the pixel coordinates of the bolt corner points and the corner point coordinates in the bolt coordinate system is: Among them, is the corner point coordinate in the bolt coordinate system, is the corner point coordinate in the pixel coordinate system, and H is the homography matrix, which is specifically expressed as follows: Based on the correspondence between the pixel coordinates of four groups of bolt corner points and the two-dimensional homogeneous coordinates in the bolt coordinate system, the homography matrix H can be calculated.
4. The method for image vision servo bolt alignment based on real-time depth estimation according to claim 3, wherein S3 includes the following steps: Split the homography matrix into a 3×4 camera internal parameter matrix and a 4×3 truncated external parameter matrix, and the formula is as follows: Among them, K is the internal parameter matrix of the camera, E is the truncated external parameter matrix with the original third column removed, f x , f y , u0, and v0 are the four internal parameter values of K. The complete external parameter matrix [R, T] describes the transformation relationship from the bolt coordinate system O b -X b Y b Z b to the camera coordinate system O c -X c Y c Z c . Among them, R is the rotation matrix and T is the displacement vector. The expression is as follows: Expand and solve Equation (4) to obtain each element in the truncated external parameter matrix E: Multiply the obtained truncated external parameter matrix E by a scaling factor s for processing: The magnitude of the scaling factor s is the geometric mean reciprocal of the lengths of the vectors [R1, R4, R7] T and [R2, R5, R8] T and the sign is determined according to the sign of T Z .
5. The method for image vision servo bolt alignment based on real-time depth estimation according to claim 4, characterized in that, S4 includes the following steps: The complete displacement vector T has been obtained through the truncated external parameter matrix E, and the missing column vector of the rotation matrix R is restored through cross product operation: Perform singular value decomposition reconstruction on R to obtain a strictly unitary orthogonal rotation matrix: R = USV T ; (8) Among them, U and V are 3×3 orthogonal matrices, S is a 3×3 non-negative diagonal matrix, and the reconstructed rotation matrix R is the product of matrices U and V; Thus, the complete external parameter matrix Q = [R, T] is obtained, and then the complete external parameter matrix is further processed: Among them, Q multiplied by F means that the external parameter matrix Q rotates 180° around the X-axis of the true camera coordinate system, and the obtained is the bolt coordinate system O b -X b Y b Z b relative to the camera coordinate system O c -X c Y c Z c complete external parameter matrix.
6. The method for image vision servo bolt alignment based on real-time depth estimation according to claim 5, wherein S5 includes the following steps: The homogeneous coordinates of the corner point in the bolt coordinate system O b -X b Y b Z b are b P = b P X , b P Y , 0, 1] T . Combining with the complete external parameter matrix we get the homogeneous coordinates of the corner point in the camera coordinate system O c -X c Y c Z c as c P = c P X , c P Y , c P Z , 1] T . The calculation formula is: Among them, c P Z is the depth value of the bolt corner point relative to the camera coordinate system O c -X c Y c Z c ; Assume there is a bolt corner point P i , in the bolt coordinate system O b -X b Y b Z b The homogeneous coordinate is The pixel coordinate obtained by image detection is [u i , v i T . According to the principle of the perspective model, the pixel coordinate [u i , v i T is projected onto the image coordinate system O-xy as follows: Use the complete external parameter matrix Perform forward projection to obtain P i The projection point on the image coordinate system O-xy is Obtained from the principle of the perspective model: Set the desired feature of virtual visual servoing as Set the current feature as Set the servo error as Expanding Equation (10) gives the depth value of the feature point as: The calculation formula of the image Jacobian matrix of the feature points is as follows: Sum up all of to obtain the virtual visual servo error e v . Then, sum up all of to obtain the image Jacobian matrix L of the virtual visual servo v ; Perform visual servo control through the following formula: where λ v is the virtual visual servo gain, is the Moore-Penrose pseudoinverse of L v , and v vc is the spatial velocity of the virtual camera; Update the complete external parameter matrix using T The calculation formula is as follows: where k is the number of iterative optimizations, and the updated complete extrinsic parameter matrix is used for forward projection in the next iteration; Define an error When err k+1 -err k <10 -8 The virtual visual servo iterative optimization ends.
7. The method for image vision servo bolt alignment based on real-time depth estimation according to claim 6, wherein S6 includes the following steps: Let b P i = b P Xi , b P Yi , 0, 1] T Combining with Equation (10) gives the accurate depth value of the bolt corner point P i c P Z . 8. The method for image vision servo bolt alignment based on real-time depth estimation according to claim 7, wherein S7 includes the following steps: The corner pixel coordinates [u i , v i T are obtained by image detection and converted to the image coordinate system [x i , y i T through Equation (11). Then, [x i , y i T and the corresponding depth value c P Zi are substituted into Equation (14) to obtain the image Jacobian matrix L i corresponding to each corner point; Set the desired feature of image-based visual servoing as Set the current feature as s i =[x i , y i T , and set the error as Sum up the e i at each point to obtain the servo error e, sum up the L i at each point to obtain the image Jacobian matrix L, and obtain the camera velocity v through the following formula c :[[]] v c = -λL + e; (17) where λ is the image visual servo gain, and L + is the Moore-Penrose pseudoinverse of L; Use v c Perform motion control on the robotic arm and repeat the execution of S1 to S7; Define the total error of image-based visual servoing When Err < 0.00005, the image-based visual servoing converges and the bolt alignment is completed.
Citation Information
Patent Citations
Robot visual servo operation method based on finite time control
CN116872216A
Visual servo positioning method for bolt group of overhead line system of electrified railway
CN118505805A
Rapid alignment method based on relative position relation of part bolts
CN118650619A
Online position correction method for contact network maintenance mechanical arm based on SAC algorithm
CN119260726A
Bolt 6D pose detection and visual servo method based on 3D template matching
CN119515783A