This invention discloses a bolt
pose estimation method based on
monocular RGB images, applicable to industrial production and intelligent manufacturing. The method involves inputting a
monocular RGB image or video frame into a pre-defined
feature extraction network to obtain a high-dimensional feature map of the image or video frame to be detected. This high-dimensional feature map is then input into a convolutional
pose estimation network to predict the bolt's 2D keypoint
confidence map and affinity map. The prediction results are used to search for the optimal match among all potential symmetrical poses of the bolt, and the network loss is calculated to update the network parameters. A
heuristic algorithm is used to search within the prediction results of the
confidence map and affinity map to obtain the 2D coordinates of all bolt keypoints. Finally, a
direct linear transformation algorithm is used to establish the connection between the 2D and 3D keypoints of the bolt, calculating the 6D
pose of the bolt.