Mediape-based hand space 3D key point calculation method
By establishing mathematical relationships and camera projection constraints, combined with a least-squares optimization algorithm, the transformation process of Mediapipe hand key point detection is simplified, solving the problems of redundant degrees of freedom and error accumulation in existing technologies, and achieving efficient 3D displacement control to meet the application requirements of virtual reality and augmented reality.
Patent Information
- Application Number
- CN202510899455.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-30
AI Technical Summary
The existing Mediapipe hand key point detection method cannot achieve 3D displacement control at the same time, and has problems of redundant degrees of freedom and error accumulation.
By establishing the mathematical relationship between the 3D coordinates of the hand key points and the 3D coordinates, introducing the camera projection constraint, and combining the least squares optimization algorithm for iterative solution, the transformation process is simplified, the degree of freedom redundancy is reduced, and error transmission is avoided.
It achieves efficient and stable 3D key point solution in hand space, supports 3D displacement control in virtual reality and augmented reality, and improves the naturalness, smoothness and precision of operation.
Smart Images

Figure CN120726698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and image processing, and in particular to a method for calculating 3D key points in hand space based on Mediapipe. Background Art
[0002] The core goal of 3D hand keypoint detection technology is to develop advanced methods and tools to accurately calculate and analyze gestures in holographic 3D interactive visualization terminal systems, thereby enabling efficient spatial human-computer interaction. 3D hand keypoint detection is a core research direction in computer vision and is widely used in fields such as virtual reality (VR), augmented reality (AR), human-computer interaction (HCI), and industrial automation, providing technical support for contactless interaction and enhanced spatial freedom.
[0003] Among them, the Mediapipe-based hand 3D key point detection method is a complex technical means, which aims to achieve high-precision real-time capture and solution of hand posture. The Mediapipe framework has become an industry benchmark technology due to its high efficiency, real-time performance and high-precision 3D key point solution capabilities. The hand key points defined by Mediapipe are as follows: Figure 1 As shown in the figure, placing your hand in front of the camera, the Mediapipe HandLandmarksDetection module outputs the hand's chirality (Handedness), keypoint 2D coordinates (Landmarks), and keypoint 3D coordinates (WorldLandmarks). All the output 3D coordinates of the hand keypoints are based on the WRIST keypoint numbered 0 in the figure as the geometric center origin.
[0004] However, existing technologies have certain limitations in practical applications. For example, when using MediaPipe to perform hand interaction with virtual objects, either 3D coordinates are used for in-place rotation control, but not translation control; or 2D coordinates are used for planar displacement control, but not depth control, which is not a complete 3D displacement control. Therefore, existing MediaPipe hand keypoint detection suffers from the problem of not being able to output spatial 3D. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned shortcomings and provide a method for solving 3D key points in hand space based on Mediapipe, which is simple to calculate, reduces redundant degrees of freedom, has tighter constraints, avoids error accumulation caused by error propagation in step-by-step optimization, and has high solution stability.
[0006] To achieve the above object, the specific solutions of the present invention are as follows:
[0007] A first aspect of the present invention provides a method for calculating 3D key points in hand space based on Mediapipe, which may specifically include the following steps:
[0008] S100, establishing a mathematical relationship between the 3D coordinates of the hand key point space and the 3D coordinates of the hand key point;
[0009] S200, introducing a camera projection constraint relationship and establishing a correlation model between the 2D coordinates of the hand key points and the 3D coordinates of the hand key points in space;
[0010] S300 , based on the association model and the mathematical relationship, and in combination with the 2D coordinates and 3D coordinate data of the hand key points, solve the 3D coordinates of the hand key points in space.
[0011] The present invention further provides that the mathematical relationship between the 3D coordinates of the hand key point space and the 3D coordinates of the hand key point is expressed by the following formula:
[0012] SpaceLandmarks[i]=R1 -1 *(WorldLandmarks[i]-t1)
[0013] Among them, i represents the i-th hand key point, SpaceLandmarks[i] is the spatial 3D coordinate of the hand key point, WorldLandmarks[i] is the 3D coordinate of the hand key point, R1 and t1 are rotation and translation matrices.
[0014] The present invention further provides that the association model between the 2D coordinates of the hand key points and the 3D coordinates of the hand key points space is expressed by the following formula:
[0015] Landmarks[i]=Intrinsic×g(SpaceLandmarks[i])
[0016] Among them, Landmarks[i] is the 2D coordinate of the key point of the hand, the g(x) function represents the distortion transformation applied to x, and Intrinsic is the intrinsic parameter matrix of the camera.
[0017] The present invention further comprises: solving the spatial 3D coordinates of the hand key points based on the association model and the mathematical relationship, combining the 2D coordinates of the hand key points and the 3D coordinate data of the hand key points, specifically including:
[0018] Select at least three 2D coordinates and 3D coordinates of key hand points, and use the least squares optimization algorithm to iteratively optimize and solve the rotation and translation matrices R1 and t1 to obtain stable rotation and translation matrices R1 and t1;
[0019] According to the stable rotation and translation matrices R1 and t1, the corresponding 3D coordinates of the hand key points are obtained.
[0020] The present invention further comprises the following iterative steps of the least squares optimization algorithm:
[0021] S301, initializing the rotation and translation matrices R1 and t1;
[0022] S302, calculating the spatial 3D coordinates of the hand key points according to the current rotation and translation matrices R1 and t1 using the mathematical relationship;
[0023] S303, calculating the projected 2D coordinates of the hand key points according to the association model, combined with the camera intrinsic parameter matrix and distortion parameters;
[0024] S304, calculating the error between the projected 2D coordinates of the hand key points and the actually obtained 2D coordinates of the hand key points;
[0025] S305, updating the rotation and translation matrices R1 and t1 according to the error;
[0026] S306 , repeating steps S302 to S305 until the error meets a preset threshold or reaches a maximum number of iterations.
[0027] The present invention further provides that the objective function of the least squares optimization algorithm is:
[0028]
[0029] Among them, N≥3 and N≤21.
[0030] The present invention further includes establishing a mathematical relationship between the spatial 3D coordinates of the hand key points and the 3D coordinates of the hand key points, specifically comprising:
[0031] Output the 2D coordinates and 3D coordinates of 21 hand key points through the Mediapipe algorithm;
[0032] The generation process of the 3D coordinates of the hand key points is reversely analyzed; based on the analysis results, the transformation process is simplified to a single rotation and translation matrix operation, thereby obtaining the mathematical relationship between the spatial 3D coordinates of the hand key points and the 3D coordinates of the hand key points.
[0033] A second aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the above-mentioned solution method.
[0034] A third aspect of the present invention provides a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the steps of the solution method described above are implemented.
[0035] A fourth aspect of the present invention provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the solution method described above are implemented.
[0036] The beneficial effects of the present invention are as follows: by simplifying the complex transformation process and combining it with the camera projection transformation relationship, the present invention utilizes a least-squares optimization algorithm for iterative optimization and solution, effectively and stably calculating the spatial 3D coordinates of the hand key points from the 3D coordinates of the hand key points. This method is computationally simple, reduces redundant degrees of freedom, provides tighter constraints, avoids error accumulation caused by error propagation during step-by-step optimization, and provides a highly stable solution. This method addresses the existing MediaPipe hand key point detection's inability to output spatial 3D coordinates, meeting the needs of practical application scenarios such as virtual reality, augmented reality, and gesture recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a diagram of the hand key point numbers and names defined by the existing Mediapipe;
[0038] Figure 2 It is a schematic flow diagram of the present invention; DETAILED DESCRIPTION
[0039] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the scope of implementation of the present invention is not limited thereto.
[0040] like Figures 1 to 2 As shown in FIG, the embodiment of the present invention is a method for calculating the 3D key points of the hand space based on Mediapipe. Mediapipe is a machine learning-based framework for detecting the key points of the hand, such as Figure 1 As shown in Figure 1, the output includes the 2D and 3D coordinates of 21 hand keypoints. 2D coordinates are typically based on pixel positions in the image plane, while 3D coordinates reflect the relative position of the hand in space. The hand image is captured by a camera, and the algorithm uses a deep learning model to predict the depth information of the keypoints, generating 3D coordinates. These 3D coordinates are typically relative to the center point of the hand and are called the 3D coordinates of the hand keypoints.
[0041] like Figure 2 As shown, the method for calculating hand space 3D key points based on Mediapipe in this embodiment may specifically include the following steps:
[0042] Step S100: Establishing a mathematical relationship between the 3D coordinates of the hand key points in space and the 3D coordinates of the hand key points.
[0043] Specifically, the 2D coordinates and 3D coordinates of the 21 hand key points in each frame of the image identified by the Mediapipe hand key point detection module are obtained; these key points cover the palm, finger joints, and other parts. The 2D coordinates of the hand key points are the key point position information on the image plane; the 3D coordinates of the hand key points are the three-dimensional position information after preliminary processing in the Mediapipe internal coordinate system. The generation process of the 3D coordinates of the hand key points is reversely analyzed; based on the analysis results, the transformation process is simplified to a single rotation and translation matrix operation, thereby obtaining the mathematical relationship between the spatial 3D coordinates of the hand key points and the 3D coordinates of the hand key points.
[0044] Specifically, through reverse analysis, it can be seen that the original 3D coordinates of the hand key points are usually obtained by rotating and translating the 3D coordinates of the hand key point space to obtain the intermediate transition 3D coordinates (TransLandmarks), and then subtracting the transition 3D coordinates numbered by WRIST from all the transition 3D coordinates.
[0045] The original hand keypoint space coordinates (SpaceLandmarks) need to undergo two steps of transformation:
[0046] Rotation and translation transformation: transform to intermediate transition 3D coordinates (TransLandmarks) through matrix R and translation vector t;
[0047] Wrist decentralization: Taking the wrist point (WRIST) as the origin, all coordinates are subtracted from TransLandmarks[0] to obtain the decentralized coordinates (WorldLandmarks);
[0048] The transformation formula is as follows:
[0049] TransLandmarks[i]=R0×SpaceLandmarks[i]+t0
[0050] WorldLandmarks[i]=TransLandmarks[i]-TransLandmarks[0]
[0051] That is:
[0052] WorldLandmarks[i]=(R0×SpaceLandmarks[i]+t0)-(R0×
[0053] SpaceLandmarks[0]+t0)
[0054] Among them, i represents the i-th key point, R0 and t0 are the original rotation and translation matrices, SpaceLandmarks is the 3D coordinate of the hand key point to be determined, and WorldLandmarks is the original 3D coordinate of the hand key point.
[0055] Without loss of generality, the "rotation-translation transformation" process and the "transition 3D coordinates minus the WRIST number" process can be combined into a "new rotation-translation transformation" process, because a new R1 and t1 can always be found to achieve the combination of the two processes. Therefore, the above transformation process formula can be simplified to:
[0056] WorldLandmarks[i]=R1×SpaceLandmarks[i]+t1
[0057] Where i represents the i-th keypoint, R1 and t1 are the rotation and translation matrices to be calculated, and SpaceLandmarks[i] is the 3D spatial coordinate of the hand keypoint to be calculated. For all keypoints in each frame, the rotation and translation matrices R1 and t1 are shared because they are a unified transformation for all keypoints in the current frame, without involving transformations between different frames.
[0058] If the R1 and t1 rotation and translation matrices are obtained, the calculation formula for the 3D coordinates of the hand key points is:
[0059] SpaceLandmarks[i]=R1 -1 *(WorldLandmarks[i]-t1) (1)
[0060] Among them, i represents the i-th hand key point, SpaceLandmarks[i] is the spatial 3D coordinate of the hand key point, WorldLandmarks[i] is the 3D coordinate of the hand key point, R1 and t1 are rotation and translation matrices.
[0061] That is, according to formula (1), the generation process of the original 3D coordinates of the hand key points is simplified into a new rotation and translation transformation process; thereby, the redundancy of the degrees of freedom can be reduced while ensuring that the degrees of freedom remain unchanged, the constraints are tighter, and the error accumulation caused by the error transmission of the step-by-step optimization is avoided.
[0062] The derivation of this formula ensures a clear mapping relationship between the 3D coordinates of the hand key points in space and the 3D coordinates of the hand key points.
[0063] Step S200: Introduce camera projection constraint relationship and establish an association model between the 2D coordinates of the hand key points and the 3D coordinates of the hand key points in space.
[0064] Specifically, the camera's intrinsic matrix (Intrinsic) and distortion parameters (Distortion) are calibrated in advance using a camera calibration tool. Then, the 3D coordinates of the hand key points in space and the 2D coordinates of the hand key points satisfy the camera projection transformation relationship. The formula is as follows:
[0065] Landmarks[i]=Intrinsic×g(SpaceLandmarks[i]) (2)
[0066] Among them, Landmarks[i] is the 2D coordinate of the key point of the hand, the g(x) function represents the distortion transformation applied to x, and Intrinsic is the intrinsic parameter matrix of the camera;
[0067] Therefore, the camera projection transformation relationship between the 3D coordinates of the hand key points and the 2D coordinates of the hand key points can be established through formula (2). This is applicable to complex imaging scenarios such as fisheye lenses and wide-angle lenses, and can achieve high-precision spatial coordinate solution using a low-cost RGB camera.
[0068] Step S300: Based on the association model and the mathematical relationship, the 2D coordinates of the hand key points and the 3D coordinate data of the hand key points are combined to solve the 3D coordinates of the hand key points in space.
[0069] Specifically, by combining formulas (1) and (2), it can be seen that at least three 2D coordinates and 3D coordinates of the hand key points are required to solve the rotation and translation matrices R1 and t1, as well as the corresponding spatial 3D coordinates. Therefore, at least three 2D coordinates and 3D coordinates of the hand key points are selected, and the least squares optimization algorithm is used to iteratively optimize and solve the rotation and translation matrices R1 and t1 to obtain stable rotation and translation matrices R1 and t1; based on the stable rotation and translation matrices R1 and t1, the corresponding spatial 3D coordinates of the hand key points are obtained. By introducing the camera projection constraint relationship, a correlation model between the 2D coordinates of the hand key points and the spatial 3D coordinates of the hand key points is established, and combined with the least squares optimization algorithm for iterative solution, the accuracy and stability of the spatial 3D coordinate solution of the hand key points can be significantly improved.
[0070] Furthermore, parameter sharing is implemented: for the same frame, all hand keypoints share the same rotation and translation matrices R1 and t1, enforcing consistency in the rigid body motion of the hand skeleton. This parameter sharing reduces the number of variables, the number of iterations, and improves computational efficiency.
[0071] Furthermore, the iterative steps of the least squares optimization algorithm of this embodiment may specifically include:
[0072] Step S301, initializing the rotation and translation matrices R1 and t1;
[0073] Step S302: Calculate the 3D coordinates of the hand key points using the mathematical relationship according to the current rotation and translation matrices R1 and t1;
[0074] Step S303: Calculate the projected 2D coordinates of the hand key points according to the association model, combined with the camera intrinsic parameter matrix and distortion parameters;
[0075] Step S304: Calculate the error between the projected 2D coordinates of the hand key points and the actually acquired 2D coordinates of the hand key points;
[0076] Step S305: Update the rotation and translation matrices R1 and t1 according to the error;
[0077] Step S306: Repeat steps S302 to S305 until the error meets a preset threshold or the maximum number of iterations is reached. This results in stable rotation and translation matrices R1 and t1. The corresponding 3D coordinate data of the hand key points is then obtained using the stable rotation and translation matrices R1 and t1. A least squares optimization algorithm is used for iterative optimization to achieve a more stable solution.
[0078] Specifically, during the iterative optimization process of the least squares optimization algorithm, the objective function is:
[0079]
[0080] Among them, N≥3 and N≤21.
[0081] In practical applications, by mapping the spatial 3D coordinates of the hand key points to the holographic three-dimensional interactive visualization terminal system, accurate solution results of the hand posture are generated to support spatial human-computer interaction of gesture operations.
[0082] In order to better enable relevant personnel in this technical field to fully understand and implement the present invention, the specific implementation principle of the present invention is further supplemented below with reference to a specific application scenario.
[0083] First, in the virtual reality scene, the user captures gesture operations in real time through the camera. The image data collected by the camera is processed by the Mediapipe algorithm to generate the 2D coordinates and 3D coordinates of the hand key points. At this time, the 2D coordinates of the hand key points directly correspond to the position information in the image plane, while the 3D coordinates of the hand key points are the three-dimensional position information preliminarily calculated in the Mediapipe internal coordinate system. In order to ensure the accurate solution of the 3D coordinates of the hand key points, it is necessary to establish a rotation and translation matrix to describe the mathematical relationship between the 3D coordinates of the hand key points and the 3D coordinates of the hand key points. Through the formula WorldLandmarks[i]=R1×SpaceLandmarks[i]+t1, where R1 and t1 are the rotation and translation matrices to be calculated, combined with the inverse analysis of the 3D coordinate generation process of the hand key points output by Mediapipe, it is derived that SpaceLandmarks[i]=R1 -1 The core of this formula is to convert the 3D coordinates in the Mediapipe internal coordinate system into the 3D coordinates in the real world, thereby providing accurate location information for subsequent gesture operations.
[0084] Next, to establish a correlation model between the 2D coordinates of the hand keypoints and their spatial 3D coordinates, a camera projection constraint relationship is introduced. By using a camera calibration tool to obtain the camera intrinsic parameter matrix and distortion parameters, a camera projection transformation relationship is constructed between the spatial 3D coordinates of the hand keypoints and their 2D coordinates. The formula Landmarks[i] = Intrinsic × g(SpaceLandmarks[i]) is used to describe this transformation relationship, where Landmarks[i] is the 2D coordinate of the hand keypoint, the g(x) function represents the distortion transformation applied to x, and Intrinsic is the camera intrinsic parameter matrix. This formula allows the spatial 3D coordinates of the hand keypoints to be mapped to the image plane, thereby achieving a direct correlation between the 2D and 3D coordinates. The core function of this correlation model is to provide the necessary constraints for the optimization algorithm to ensure the accuracy of the final solution.
[0085] Then, the least squares optimization algorithm is combined to perform an iterative solution to obtain a stable rotation and translation matrix. At least three 2D coordinates and 3D coordinates of the key points of the hand are selected, and the objective function is constructed by combining mathematical relationships and association models. The formula of the objective function is:
[0086]
[0087] Where N ≥ 3 and N ≤ 21. After initializing R1 and t1 in the rotation and translation matrices, the 3D coordinates of the hand key points are calculated using mathematical relationships, and the projected 2D coordinates of the hand key points are calculated by combining the camera intrinsic parameter matrix and distortion parameters. By comparing the error between the projected 2D coordinates of the hand key points and the actual 2D coordinates of the hand key points, the rotation and translation matrices R1 and t1 are updated based on the error. Repeat the above steps until the error meets the preset threshold or the maximum number of iterations is reached, thus obtaining a stable rotation and translation matrix. Finally, based on the stable rotation and translation matrix, the 3D coordinates of the hand key points are calculated to complete the accurate solution of the 3D coordinates of the hand key points.
[0088] After the 3D coordinates of the hand key points are calculated, they can be mapped to the holographic three-dimensional interactive visualization terminal system to generate the hand posture solution results. In the virtual reality scene, users can use gestures to perform operations such as grabbing, moving, and rotating virtual objects. For example, when a user extends a finger and approaches a virtual object, the 3D coordinates of the hand key points in space can accurately reflect the relative position relationship between the finger and the virtual object. By updating the 3D coordinates of the hand key points in space in real time, the system can accurately judge the user's intention and perform corresponding operations. In addition, the high-precision calculation of the 3D coordinates of the hand key points in space makes gesture operations more natural and smooth, solving the problem that the existing technology cannot support translation control and depth control at the same time.
[0089] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the above-mentioned solution method.
[0090] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned solution method, resolving the issue with existing MediaPipe hand keypoint detection's inability to output spatial 3D coordinates. Compared to the prior art, the computer-readable storage medium provided in this embodiment offers the same beneficial effects as the MediaPipe-based method for solving spatial 3D hand keypoints provided in the aforementioned embodiment, and is not further elaborated upon here.
[0091] An embodiment of the present invention further provides a computer program product, comprising a computer program that, when executed by a processor, implements the steps of the solution method described above. The computer program product provided in this embodiment can solve the problem that the existing MediaPipe hand key point detection cannot output spatial 3D coordinates. Compared with the prior art, the beneficial effects of the computer program product provided in this embodiment are the same as the beneficial effects of the MediaPipe-based hand spatial 3D key point solution method provided in the above embodiment, and will not be repeated here.
[0092] The above description is only a preferred embodiment of the present invention. Therefore, any equivalent changes or modifications made according to the structure, characteristics and principles described in the scope of the patent application of the present invention are included in the protection scope of the patent application of the present invention.
Claims
1. A method for calculating 3D key points in hand space based on Mediapipe, characterized in that: The steps include: S100, establishing a mathematical relationship between the 3D coordinates of the hand key point space and the 3D coordinates of the hand key point; S200, introducing a camera projection constraint relationship and establishing a correlation model between the 2D coordinates of the hand key points and the 3D coordinates of the hand key points in space; S300 , based on the association model and the mathematical relationship, and in combination with the 2D coordinates and 3D coordinate data of the hand key points, solve the 3D coordinates of the hand key points in space.
2. The method for calculating 3D key points in hand space based on Mediapipe according to claim 1, characterized in that: The mathematical relationship between the 3D coordinates of the hand key point space and the 3D coordinates of the hand key point is expressed by the following formula: SpaceLandmarks[i]=R1 -1 *(WorldLandmarks[i]-t1) Among them, i represents the i-th hand key point, SpaceLandmarks[i] is the spatial 3D coordinate of the hand key point, WorldLandmarks[i] is the 3D coordinate of the hand key point, R1 and t1 are rotation and translation matrices.
3. The method for calculating 3D key points in hand space based on Mediapipe according to claim 2, characterized in that: The association model between the 2D coordinates of the hand key points and the 3D coordinates of the hand key points space is expressed by the following formula: Landmarks[i]=Intrinsic×g(SpaceLandmarks[i]) Among them, Landmarks[i] is the 2D coordinate of the key point of the hand, the g(x) function represents the distortion transformation applied to x, and Intrinsic is the intrinsic parameter matrix of the camera.
4. The method for calculating 3D key points in hand space based on Mediapipe according to claim 3, characterized in that: The method of solving the spatial 3D coordinates of the hand key points based on the association model and the mathematical relationship and combining the 2D coordinates of the hand key points and the 3D coordinate data of the hand key points specifically includes: Select at least three 2D coordinates and 3D coordinates of key hand points, and use the least squares optimization algorithm to iteratively optimize and solve the rotation and translation matrices R1 and t1 to obtain stable rotation and translation matrices R1 and t1; According to the stable rotation and translation matrices R1 and t1, the corresponding 3D coordinates of the hand key points are obtained.
5. The method for calculating 3D key points in hand space based on Mediapipe according to claim 4, characterized in that: The iterative steps of the least squares optimization algorithm include: S301, initializing the rotation and translation matrices R1 and t1; S302, calculating the spatial 3D coordinates of the hand key points according to the current rotation and translation matrices R1 and t1 using the mathematical relationship; S303, calculating the projected 2D coordinates of the hand key points according to the association model, combined with the camera intrinsic parameter matrix and distortion parameters; S304, calculating the error between the projected 2D coordinates of the hand key points and the actually obtained 2D coordinates of the hand key points; S305, updating the rotation and translation matrices R1 and t1 according to the error; S306 , repeating steps S302 to S305 until the error meets a preset threshold or reaches a maximum number of iterations.
6. The method for calculating 3D key points in hand space based on Mediapipe according to claim 4, characterized in that: The objective function of the least squares optimization algorithm is: Among them, N≥3 and N≤21.
7. The method for calculating 3D key points in hand space based on Mediapipe according to claim 2, characterized in that: The establishing of the mathematical relationship between the spatial 3D coordinates of the hand key points and the 3D coordinates of the hand key points specifically includes: Output the 2D coordinates and 3D coordinates of 21 hand key points through the Mediapipe algorithm; The generation process of the 3D coordinates of the hand key points is reversely analyzed; based on the analysis results, the transformation process is simplified to a single rotation and translation matrix operation, thereby obtaining the mathematical relationship between the spatial 3D coordinates of the hand key points and the 3D coordinates of the hand key points.
8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the solution method according to any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the solution method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the solution method according to any one of claims 1 to 7 are implemented.