This disclosure provides a method, apparatus,
system, device, medium, and program product for hand
pose reconstruction. The method includes: acquiring an
image sequence simultaneously captured from at least two angles; determining the
hand region in each frame of the
image sequence; wherein the
hand region of at least a portion of the images in the
image sequence is determined based on the hand regions of adjacent images; identifying hand
pose parameters in the hand regions; and constructing a three-dimensional hand model based on the hand
pose parameters. This disclosure, based on multiple views, avoids keypoint drift and shape loss caused by
occlusion of the real hand during the three-dimensional hand model construction process, and has the advantages of low computational load and short
inference time. Furthermore, since the
hand region of some images in the image sequence is determined based on the hand regions of their adjacent images, target recognition is not required for each frame, thereby reducing the frequency of target recognition and further improving the speed and efficiency of three-dimensional hand model construction.