A method and device for estimating hand posture

Through the combination of left and right cameras and inertial sensors, the 3D coordinates of the hand joint node are estimated, which solves the problems of complex calculations, poor real-time performance and limited interaction range in the prior art, and achieves more efficient and accurate hand posture estimation.

CN113536931BActive Publication Date: 2025-05-13HISENSE VISUAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110665272.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-16
Publication Date
2025-05-13
Estimated Expiration
2041-06-16

AI Technical Summary

Technical Problem

The existing 3D spatial pose estimation method of hand joint nodes is complex, real-time and accurate in calculations, and the use of depth cameras increases device power consumption and limits interaction range.

Method used

The hand images are collected separately by left and right cameras, combined with the motion data collected by the inertial sensor, and the position information of the AR device is determined. Based on the 2D coordinates of the hand joint node extracted from the hand image and the pose information of the AR device, the 3D coordinates of the hand joint node are estimated.

Benefits of technology

The hardware load of 3D coordinate estimation of hand joint nodes is reduced, the accuracy of estimation is improved, the dependence on AR device processing performance is reduced, and the interaction range is expanded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113536931B_ABST
    Figure CN113536931B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of augmented reality, and provides a method and device for estimating a hand posture, which obtains a first hand image captured by a left camera, a second hand image captured by a right camera, and motion data of an AR device captured by an inertial sensor; determines the posture information of the AR device according to the first hand image, the second hand image and the motion data; obtains a first 2D coordinate set and a second 2D coordinate set from 2D coordinates of hand joint points respectively extracted from the first hand image and the second hand image, and estimates the 3D coordinates of the hand joint points in combination with the posture information of the AR device. The method saves the calculation amount of 3D coordinates, reduces the dependence on the processing performance of the AR device, and improves the accuracy of 3D coordinate estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of augmented reality (AR) technology, and in particular to a hand posture estimation method and device. Background Art

[0002] With the development of artificial intelligence technology, technologies such as augmented reality (AR) and virtual reality (VR) have been widely used, allowing people to immerse themselves in an environment where virtual information and real environment are integrated. The emergence of augmented reality products (such as AR glasses) can provide people with information that cannot be quickly obtained in the real world, help people improve work efficiency, and provide many interesting life experiences.

[0003] Hand joint pose estimation is an important technology for human-computer interaction, which can get rid of the traditional way of inputting information through mouse or keyboard. By accurately estimating the 3D spatial pose of hand joints, it is possible to use simple gestures to provide instructions to the machine to quickly transmit information, thereby improving the efficiency and experience of human-computer interaction.

[0004] At present, the 3D spatial posture estimation methods of hand joints include: first, based on a monocular hand image, a deep learning model is used to directly estimate the 3D coordinate information of the hand joints, but the calculation is complex and heavily dependent on the processing performance of hardware such as the GPU, APU, NPU, and TPU in the AR device, and the real-time performance and accuracy are poor; second, based on a monocular hand image, a deep learning model is used to first estimate the 2D coordinate information of the hand joints, and at the same time, the 3D coordinate information of the hand joints is calculated in conjunction with the hand depth image collected by the depth camera, but the use of the depth camera not only increases the power consumption of the AR device and shortens the standby time, but is also limited by the field of view of the depth camera (the field of view angle is less than or equal to 55°), and the interaction range is narrow; third, the 3D coordinate information of the hand is detected with the help of a third-party device, such as wearing special gloves for the user's hands, which affects the flexibility of the user's hand activities and reduces the user experience. Summary of the invention

[0005] The embodiments of the present application provide a hand posture estimation method and device for reducing the hardware load of an AR device in estimating the 3D coordinates of hand joints and improving the accuracy of 3D coordinate estimation.

[0006] In a first aspect, an embodiment of the present application provides a hand posture estimation method, which is applied to an AR device, including:

[0007] Acquire a first hand image captured by a left camera, a second hand image captured by a right camera, and motion data captured by an inertial sensor;

[0008] Determine the position information of the AR device according to the first hand image, the second hand image and the motion data;

[0009] Extracting a first 2D coordinate set of hand joint points from the first hand image, and extracting a second 2D coordinate set of hand joint points from the second hand image;

[0010] Estimate the 3D coordinates of the hand joint points according to the first 2D coordinate set, the second 2D coordinate set, and the pose information of the AR device.

[0011] In a second aspect, an embodiment of the present application provides an augmented reality (AR) device, including a left camera, a right camera, an inertial sensor, a memory, and a processor;

[0012] The left camera is connected to the processor and is configured to capture a first hand image;

[0013] The right camera is connected to the processor and is configured to capture a first hand image;

[0014] The inertial sensor is connected to the processor and configured to collect motion data of the AR device;

[0015] The memory is connected to the processor and configured to store computer program instructions;

[0016] The processor is configured to perform the following operations according to the computer program instructions:

[0017] Acquire a first hand image captured by the left camera, a second hand image captured by the right camera, and motion data captured by the inertial sensor;

[0018] Determine the position information of the AR device according to the first hand image, the second hand image and the motion data;

[0019] Extracting a first 2D coordinate set of hand joint points from the first hand image, and extracting a second 2D coordinate set of hand joint points from the second hand image;

[0020] Estimate the 3D coordinates of the hand joint points according to the first 2D coordinate set, the second 2D coordinate set, and the pose information of the AR device.

[0021] In a third aspect, the present application provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the hand posture estimation method provided in an embodiment of the present application.

[0022] In the above-mentioned embodiments of the present application, hand images are respectively collected by left and right cameras, and the posture information of the AR device is determined in combination with the motion data collected by the inertial sensor. The 3D coordinates of the hand joint points are estimated based on the 2D coordinates of the hand joint points extracted from the hand images and the posture information of the AR device. On the one hand, the 3D coordinates are estimated based on the 2D coordinates of the hand joint points and the posture information of the AR device, which saves the amount of calculation, reduces the processing load of the hardware, and reduces the dependence on the processing performance of the AR device. On the other hand, the 3D coordinates of the hand joint points are estimated by using the motion data collected by the inertial sensor and the hand images collected by the left and right cameras, which improves the accuracy of the estimation compared to estimating the 3D coordinates of the hand joint points using a monocular hand image. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0024] Figure 1 The embodiment of the present application provides an AR device structure diagram;

[0025] Figure 2 The flowchart of the hand posture estimation method provided by the embodiment of the present application is exemplarily shown;

[0026] Figure 3 The schematic diagram of hand joint points provided by the embodiment of the present application is exemplarily shown;

[0027] Figure 4 The following is a schematic diagram showing the matching of joint points in hand images captured by left and right cameras provided in an embodiment of the present application;

[0028] Figure 5 The schematic diagram of determining the positions of left and right cameras provided in an embodiment of the present application is exemplarily shown;

[0029] Figure 6 The schematic diagram of the hand posture 3D coordinate estimation principle provided by the embodiment of the present application is exemplarily shown;

[0030] Figure 7 The complete hand posture estimation method flow chart provided by the embodiment of the present application is exemplarily shown;

[0031] Figure 8 The functional structure diagram of the AR device provided in the embodiment of the present application is exemplified.

[0032] Fig. 9The hardware structure diagram of the AR device provided in the embodiment of the present application is exemplified. DETAILED DESCRIPTION

[0033] In order to make the purpose, implementation mode and advantages of the present application clearer, the exemplary implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0034] Based on the exemplary embodiments described in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the claims attached to this application. In addition, although the disclosure in this application is introduced according to one or several exemplary examples, it should be understood that each aspect of the disclosure can also constitute a complete implementation method separately.

[0035] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0036] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, for example, they can be implemented in an order other than those given in the diagrams or descriptions of the embodiments of this application.

[0037] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to those components expressly listed but may include other components not expressly listed or inherent to such products or devices.

[0038] The term "module" as used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0039] The present application provides a hand posture estimation method and device. The device may be an AR, VR or other interactive device.

[0040] Taking AR glasses as an example, Figure 1 The structure diagram of the AR device provided in the embodiment of the present application is exemplarily shown. Figure 1As shown, the AR device includes a left display lens 101 and a right display lens 102. When a user wears the AR glasses, the human eye can view video images through the left and right display lenses. A left camera 103 is provided on the left display lens 101, and a right camera 104 is provided on the right display lens 102. The left and right cameras can respectively capture hand images during the interaction process.

[0041] Optionally, in order to increase the field of view angle of the left and right cameras and increase the range of motion of the user's hands, the left and right cameras are fisheye lenses.

[0042] like Figure 1 Not shown, the AR device also includes an inertial measurement unit (IMU). Generally, an IMU contains three single-axis accelerometers and three single-axis gyroscopes. The accelerometer is used to measure the acceleration of the AR device, and the gyroscope is used to measure the angular velocity of the AR device. By measuring the angular velocity and acceleration of the AR device in three-dimensional space, the position information of the AR device can be solved in combination with the visual images collected by the left and right cameras. Generally speaking, the IMU should be installed at the center of gravity of the object being measured, that is, the center of gravity of the AR glasses.

[0043] It should be noted that Figure 1 The camera 105 in the middle of the AR device is a common RGB camera, which is not used in the solution provided in this application and will not be introduced in detail.

[0044] based on Figure 1 The AR device shown in the figure, the present application implements a method for estimating hand posture. Figure 2 As shown, the method can be implemented by an AR device with interactive functions, and mainly includes the following steps:

[0045] S201: Acquire a first hand image captured by a left camera, a second hand image captured by a right camera, and motion data of an AR device captured by an inertial sensor.

[0046] In this step, the AR device acquires the first hand image and the second hand image simultaneously captured by the left and right cameras, and grayscales the acquired first hand image and the second hand image. The motion data of the AR device acquired by the IMU is also acquired. Since the IMU is a device for measuring the three-axis (X, Y, Z axis) attitude angle (angular velocity) and acceleration of an object, the acquired motion data includes the acceleration and angular velocity of the AR device, and is combined with the first hand image and the second hand image to solve the posture information of the AR device.

[0047] S202: Determine the position and posture information of the AR device according to the first hand image, the second hand image, and the motion data.

[0048] In S202, since the angular velocity and acceleration of the AR device measured by the IMU have obvious drift, the error of posture estimation is large when the acquired motion data is integrated multiple times, while the visual data collected by the left and right cameras have no drift and are rich in texture information. In this way, the posture information of the AR device can be determined by fusing the motion data collected by the IMU and the hand images collected by the left and right cameras, and the visual data can be used to optimize the posture error estimated by the IMU, thereby improving the estimation accuracy.

[0049] In the specific implementation, feature points are extracted from the first hand image and the second hand image, and the extracted feature points and the motion data (angular velocity and acceleration) collected by the IMU are input into the Kalman filter to estimate the posture information of the AR device. Optionally, the filter includes an extended Kalman filter (Extended Kalman Filter, EKF), a multi-state Kalman filter (Multi-State Constraint Kalman Filter, MSCKF), etc. The specific estimation process has been introduced in the prior art. Since this part is not the focus of this application, it will not be described in detail here.

[0050] In the embodiment of the present application, the posture information includes position information and rotation angle information. Since the IMU is generally set at the center of gravity of the AR device, the posture information estimated in S202 is the overall posture information of the AR device.

[0051] Assume that the position information of the AR device is recorded as P0(x, y, z), where x, y, and z respectively represent the coordinate values ​​of the AR device on the X, Y, and Z axes of the head coordinate system, and the rotation angle information is recorded as R0, which represents the deflection angle of the AR device in the head coordinate system.

[0052] S203: extracting a first 2D coordinate set of hand joint points from the first hand image, and extracting a second 2D coordinate set of hand joint points from the second hand image.

[0053] In this step, the first hand image is input into the trained deep learning model to obtain the first 2D coordinate set of the hand joints, and the second hand image is input into the trained deep learning model to obtain the second 2D coordinate set of the hand joints, wherein the number of 2D coordinates of the hand joints in the first 2D coordinate set and the second 2D coordinate set is the same, and the number of hand joints is determined by the selected deep learning model. For example, when the deep learning model is the interHand2.6m model, the model will output the 2D coordinates of 21 joints in the order of the index of the hand joints, such as Figure 3 shown.

[0054] The embodiments of the present application do not impose restrictive requirements on the network structure used by the deep learning model, including but not limited to convolutional neural networks (CNN) and recurrent neural networks (RNN).

[0055] In an optional implementation, in the first 2D coordinate set and the second 2D coordinate set output by the deep learning model, the 2D coordinates of the hand joint points may contain erroneous data, affecting the estimation accuracy of the 3D coordinates. Therefore, the 2D coordinates of the corresponding hand joint points in the first 2D coordinate set and the second 2D coordinate set may be matched with feature points, and the 2D coordinates of the hand joint points that do not match in the first 2D coordinate set and the second 2D coordinate set are eliminated.

[0056] In the embodiment of the present application, each sub-joint point in the hand joint point has corresponding 2D coordinates in the first 2D coordinate set and the second 2D coordinate set. To distinguish the description, the coordinates of the sub-joint point in the first 2D coordinate set are marked as the first 2D sub-coordinates, and the coordinates of the sub-joint point in the second 2D coordinate set are marked as the second 2D sub-coordinates. The pixel point corresponding to the first 2D sub-coordinate in the first image is the first center pixel point, and N pixels adjacent to the first center pixel point are selected. The pixel point corresponding to the second 2D sub-coordinate in the second image is the second center pixel point, and N pixels adjacent to the second center pixel point are selected, where N is an integer greater than or equal to 1. Optionally, N=256.

[0057] In a specific implementation, the pixel values ​​of the first central pixel point and the N pixel points adjacent to the first central pixel point are compared to generate a first descriptor, and the pixel values ​​of the second central pixel point and the N pixel points adjacent to the second central pixel point are compared to generate a second descriptor.

[0058] Taking the generation of the first descriptor as an example, the pixel values ​​of the first central pixel and the adjacent 256 pixels are compared. If the pixel value of the first central pixel is less than the pixel value of the adjacent pixel, it is set to 1, otherwise it is set to 0, thereby obtaining a 256-bit 0 / 1 sequence, which is used as the first descriptor.

[0059] After generating the first descriptor and the second descriptor, determine the Hamming distance between the first descriptor and the second descriptor. The larger the Hamming distance, the lower the matching degree. If the Hamming distance is greater than the preset distance threshold, indicating that the first 2D sub-coordinate and the second 2D sub-coordinate do not match, the first 2D sub-coordinate corresponding to the first descriptor and the second sub-2D coordinate corresponding to the second descriptor are eliminated.

[0060] For example, taking the sub-joint with index (number) 12 as an example, Figure 4As shown, (a) is the first hand image captured by the left camera, and (b) is the second hand image captured by the right camera. The Hamming distance between the first description generated by the sub-joint point 12 in the first hand image and the second descriptor generated by the sub-joint point 12 in the second hand image is calculated. When the Hamming distance is greater than a preset threshold, the 2D coordinates of the sub-joint point 12 are removed from the first 2D coordinate set, and the 2D coordinates of the sub-joint point 12 are removed from the second 2D coordinate set.

[0061] S204: Estimate the 3D coordinates of the hand joints according to the first 2D coordinate set, the second 2D coordinate set, and the posture information of the AR device.

[0062] In this step, what is determined in S202 is the posture information of the AR device, and the first 2D coordinate set is the 2D coordinates of the hand joint points extracted from the first hand image captured by the left camera, and the second 2D coordinate set is the 2D coordinates of the hand joint points extracted from the second hand image captured by the right camera. In order to estimate the 3D coordinates of the hand joint points, it is necessary to determine the posture information of the left and right cameras based on the posture information of the AR device.

[0063] In an embodiment of the present application, the deviation of the posture information of the left and right cameras relative to the posture information of the AR device is measurable. Combined with the determined posture information of the AR device, the posture information of the left and right cameras can be determined respectively.

[0064] In a specific implementation, the posture information of the AR device includes the position information P0 and the rotation angle information R0 of the AR device. Figure 5 As shown, according to the position information P0 of the AR device and the position deviation L1 of the left camera relative to the AR device, the position information P1 of the left camera is determined (P1=P0-L1), and according to the rotation angle information R0 of the AR device and the angle deviation R01 of the left camera relative to the AR device, the rotation angle information R1 of the left camera is determined (R1=R0*R01), and the posture information (P1, R1) of the left camera is obtained; according to the position information P0 of the AR device and the position deviation L2 of the right camera relative to the AR device, the position information P2 of the right camera is determined (P2=P0+L2), and according to the rotation angle information R0 of the AR device and the angle deviation R02 of the right camera relative to the AR device, the rotation angle information R2 of the right camera is determined (R2=R0*R02), and the posture information (P2, R2) of the right camera is obtained.

[0065] In S204, after obtaining the posture information of the left and right cameras, the 3D coordinates of the hand joints are estimated according to the first 2D coordinate set, the second 2D coordinate set, the posture information of the left camera, and the posture information of the right camera.

[0066] When estimating the 3D coordinates of the hand joints, triangulation can be used. Figure 6 As shown, O1 is the optical center of the left camera, O2 is the optical center of the right camera, I1 is the first hand image, I2 is the second hand image, p1 is the joint point on the first hand image, p2 is the joint point in the second hand image, and t is the transformation matrix of the second hand image relative to the first hand image (or the transformation matrix of the first hand image relative to the second hand image). Theoretically, O1p1 and O2p2 will intersect at a point P, where P is the position of the joint point in three-dimensional space. However, due to the influence of noise, O1p1 and O2p2 often cannot intersect. The least squares method can be used to solve the point P' that is closest to P, thereby obtaining the 3D coordinates of the joint point.

[0067] The specific calculation process has been encapsulated in the OpenCV function, and the 3D coordinates of the joint points can be directly obtained by calling the function. The specific function is as follows:

[0068] cv::triangulatePoints(PoseL, PoseR, Point2DL, Point2DR, Point3D)

[0069] Among them, PoseL represents the pose information of the left camera, PoseR represents the pose information of the right camera, Point2DL represents the 2D coordinates of the joint point in the first hand image, Point2DR represents the 2D coordinates of the joint point in the second hand image, and Point3D is the solved 3D coordinates.

[0070] In the above embodiment of the present application, the first hand image is collected by the left camera of the AR device, and the second hand image is collected by the right camera. Combined with the angular velocity and acceleration collected by the IMU, the Kalman filter algorithm is used to estimate the posture information of the AR device, which makes full use of the advantages of the camera and the IMU and improves the accuracy of the posture estimation of the AR device; and the posture information of the left and right cameras is determined respectively through the posture information of the AR device. Further, according to the posture information of the left and right cameras, as well as the 2D coordinates of the hand joint points extracted from the first hand image and the 2D coordinates of the hand joint points extracted from the second hand image, the triangulatePoints function is directly called to obtain the 3D coordinates of the hand joint points. On the one hand, compared with estimating the 3D coordinates of the hand joint points directly according to the hand image based on the deep learning model, the calculation complexity is reduced, the calculation amount is saved, thereby reducing the processing load of the hardware in the AR device and reducing the dependence on the processing performance of the AR device; on the other hand, compared with the 2D coordinates, the 3D coordinates of the joint points are estimated by combining the depth camera, which overcomes the limitations of the interactive range and reduces the power consumption of the device; on the other hand, there is no need to use third-party devices, which improves the user experience. In addition, the accuracy of 3D coordinate estimation is improved by discarding mismatched 2D coordinates.

[0071] It should be noted that Figure 2 Taking AR devices as an example, the same applies to other devices with interactive functions (such as VR devices).

[0072] Figure 7 The complete hand posture estimation method flow chart provided by the embodiment of the present application is shown as an example. Figure 7 As shown in the figure, the process mainly includes the following steps:

[0073] S701: Acquire a first hand image captured by a left camera, a second hand image captured by a right camera, and motion data of the AR device captured by an IMU.

[0074] The detailed description of this step is referred to in S201 and will not be repeated here.

[0075] S702: Extract feature points from the first hand image and the second hand image, combine them with motion data, and use a Kalman filter algorithm to fuse and calculate the posture information of the AR device.

[0076] In this step, according to the processing performance of the AR device, loose coupling or tight coupling fusion can be used to calculate the pose information of the AR device. The specific process is shown in S202 and will not be repeated here.

[0077] S703: Determine the posture information of the left camera and the posture information of the right camera respectively according to the posture information of the AR device.

[0078] In this step, the posture information includes position information and rotation angle information. According to the posture information of the AR device and the deviations of the left and right cameras relative to the posture information of the AR device, the posture information of the left and right cameras is determined. The specific process is shown in S204 and will not be repeated here.

[0079] S704: extracting a first 2D coordinate set of hand joint points from the first hand image, and extracting a second 2D coordinate set of hand joint points from the second hand image.

[0080] In this step, a deep learning model is used to extract the 2D coordinates of the joint points in the first hand image and the 2D coordinates of the joint points in the second hand image, respectively. The description of the deep learning model is referred to in S203 and will not be repeated here.

[0081] S705: For each sub-joint point in the hand joint point, generate a first descriptor according to the first hand image, and generate a second descriptor according to the second hand image.

[0082] In this step, the pixel corresponding to the sub-joint point in the first hand image is compared with the pixel values ​​of the N pixels adjacent to the pixel point. If the pixel value of the pixel corresponding to the sub-joint point is less than the pixel value of the adjacent pixel point, it is set to 1, otherwise it is set to 0, thereby obtaining an N-bit 0 / 1 sequence, which is used as the first descriptor. Similarly, the second descriptor is generated. The specific process is shown in S203, which will not be repeated here.

[0083] S706: Determine whether the Hamming distance between the first descriptor and the second descriptor is greater than a preset threshold. If so, execute S707; otherwise, execute S708.

[0084] In this step, the larger the Hamming distance between the first descriptor and the second descriptor, the lower the matching degree of the 2D coordinates of the joint points corresponding to the first descriptor and the second descriptor. In order not to affect the accuracy of the subsequent 3D coordinate estimation, the corresponding 2D coordinates need to be removed from the first 2D coordinate set and the second 2D coordinate set. The specific process is shown in S203 and will not be repeated here.

[0085] S707: Eliminate the 2D coordinates of the sub-joint points corresponding to the first descriptor from the first 2D coordinate set, and eliminate the 2D coordinates of the sub-joint points corresponding to the second descriptor from the second 2D coordinate set.

[0086] The detailed description of this step is referred to S203 and will not be repeated here.

[0087] S708: Estimate the 3D coordinates of the hand joints according to the first 2D coordinate set, the second 2D coordinate set, the posture information of the left camera, and the posture information of the right camera.

[0088] In this step, the 3D coordinates of the hand joint points can be estimated by triangulation. Since the calculation process has been encapsulated in the triangulatePoints function in OpenCV, the 3D coordinates of the hand joint points can be directly estimated by calling the function. The specific process is shown in S204 and will not be repeated here.

[0089] Based on the same technical concept, the embodiment of the present application provides an AR device that can execute the hand posture estimation method process executed by the AR device of the embodiment of the present application and achieve the same technical effect, which will not be repeated here.

[0090] See also Figure 8 , the AR device includes an acquisition module 801, a determination module 802, an extraction model 803, and an estimation module 804:

[0091] An acquisition module 801 is used to acquire a first hand image acquired by a left camera and a second hand image acquired by a right camera, and motion data of an AR device acquired by an inertial sensor;

[0092] A determination module 802 is used to determine the position information of the AR device according to the first hand image, the second hand image and the motion data;

[0093] An extraction module 803, configured to extract a first 2D coordinate set of hand joint points from the first hand image, and to extract a second 2D coordinate set of hand joint points from the second hand image;

[0094] The estimation module 804 is used to estimate the 3D coordinates of the hand joint points according to the first 2D coordinate set, the second 2D coordinate set and the posture information of the AR device.

[0095] Optionally, the estimation module 804 is specifically used for:

[0096] According to the posture information of the AR device, the posture information of the left camera and the posture information of the right camera are determined respectively;

[0097] Estimate the 3D coordinates of the hand joints according to the first 2D coordinate set, the second 2D coordinate set, the posture information of the left camera, and the posture information of the right camera.

[0098] Optionally, the estimation module 804 is specifically used for:

[0099] The position information of the left camera is determined according to the position information of the AR device and the position deviation of the left camera relative to the AR device, and the rotation angle information of the left camera is determined according to the rotation angle information of the AR device and the angle deviation of the left camera relative to the AR device, so as to obtain the position information of the left camera;

[0100] According to the position information of the AR device and the position deviation of the right camera relative to the AR device, the position information of the right camera is determined, and according to the rotation angle information of the AR device and the angle deviation of the right camera relative to the AR device, the rotation angle information of the right camera is determined to obtain the posture information of the right camera.

[0101] Optionally, the device also includes a removal module 805, which is used to perform feature point matching on the 2D coordinates of corresponding hand joint points in the first 2D coordinate set and the second 2D coordinate set, and remove the 2D coordinates of unmatched hand joint points in the first 2D coordinate set and the second 2D coordinate set.

[0102] Optionally, the elimination module 805 is specifically used for:

[0103] For the first 2D sub-coordinate and the second 2D sub-coordinate of each sub-joint in the hand joint, perform the following operations:

[0104] Compare the pixel point corresponding to the first 2D sub-coordinate in the first image with the pixel values ​​of N adjacent pixels to generate a first descriptor, where N is an integer greater than or equal to 1;

[0105] Compare the pixel point corresponding to the second 2D sub-coordinate in the second image with the pixel values ​​of N adjacent pixels to generate a second descriptor;

[0106] Determine the Hamming distance between the first descriptor and the second descriptor, and if the Hamming distance is greater than a preset distance threshold, remove the first 2D sub-coordinate corresponding to the first descriptor and the second sub-2D coordinate corresponding to the second descriptor;

[0107] The first 2D sub-coordinate is the 2D coordinate of the sub-joint point in the first 2D coordinate set, and the second 2D sub-coordinate is the 2D coordinate of the sub-joint point in the second 2D coordinate set.

[0108] Based on the same technical concept, the embodiment of the present application provides an AR device that can execute the hand posture estimation method process executed by the AR device of the embodiment of the present application and achieve the same technical effect, which will not be repeated here.

[0109] See also Fig. 9 The AR includes a left camera 901, a right camera 902, an IMU 903, a memory 904, and a processor 905, wherein the left camera 901, the right camera 902, the IMU 903, the memory 904 and the processor 903 are connected via a bus (in Fig. 9The left camera 901 is configured to collect a first hand image, the right camera 902 is configured to collect a second hand image, the IMU 903 is configured to collect motion data of the AR device, the memory 904 is configured to store computer program instructions, and the processor 905 is configured to execute the embodiments of the present application according to the computer program instructions. Figure 2 The method flow is shown.

[0110] The embodiment of the present application also provides a computer-readable storage medium for storing some instructions, which, when executed, can complete the method of the aforementioned embodiment.

[0111] The embodiments of the present application also provide a computer program product for storing a computer program, wherein the computer program is used to execute the method of the aforementioned embodiment.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0113] For the convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A hand posture estimation method, characterized in that: Applied to augmented reality (AR) devices, including: Acquire a first hand image captured by a left camera and a second hand image captured by a right camera, and motion data of the AR device captured by an inertial sensor; wherein the left camera and the right camera are fisheye lenses, and the first hand image and the second hand image are RGB images; Determine the posture information of the AR device according to the first hand image, the second hand image and the motion data; wherein the posture information includes position information and rotation angle information; Extracting a first 2D coordinate set of hand joint points from the first hand image, and extracting a second 2D coordinate set of hand joint points from the second hand image; Determine the posture information of the left camera according to the position information of the AR device and the posture deviation of the left camera relative to the AR device, and determine the posture information of the right camera according to the position information of the AR device and the posture deviation of the right camera relative to the AR device; wherein each posture deviation includes a position deviation corresponding to the position information and a rotation angle deviation corresponding to the rotation angle information; Estimate the 3D coordinates of the hand joint points according to the first 2D coordinate set, the second 2D coordinate set, the posture information of the left camera, and the posture information of the right camera; The posture information (P1, R1) of the left camera is expressed as follows: P1=P0-L1, R1=R0*R01; The posture information (P2, R2) of the right camera is expressed as: P2=P0-L2, R2=R0*R02; The P0 represents the position information of the AR device, the R0 represents the rotation angle information of the AR device, the L1 represents the position deviation of the left camera relative to the AR device, and the L2 represents the position deviation of the right camera relative to the AR device. The R01 represents the rotation angle deviation of the left camera relative to the AR device, and the R02 represents the rotation angle deviation of the right camera relative to the AR device.

2. The method according to claim 1, characterized in that After extracting a first 2D coordinate set of hand joint points from the first hand image and extracting a second 2D coordinate set of hand joint points from the second hand image, before estimating the 3D coordinates of the hand joint points, the method further includes: The 2D coordinates of the corresponding hand joint points in the first 2D coordinate set and the second 2D coordinate set are matched with each other in feature points, and the 2D coordinates of the hand joint points that do not match each other in the first 2D coordinate set and the second 2D coordinate set are eliminated.

3. The method according to claim 2, characterized in that Performing feature point matching on the first 2D coordinate and the second 2D coordinate, and removing unmatched first 2D coordinates and second 2D coordinates, comprises: For the first 2D sub-coordinate and the second 2D sub-coordinate of each sub-joint point in the hand joint point, perform the following operations: Compare the pixel point corresponding to the first 2D sub-coordinate in the first hand image with the pixel values ​​of N adjacent pixels to generate a first descriptor, where N is an integer greater than or equal to 1; Compare the pixel point corresponding to the second 2D sub-coordinate in the second hand image with the pixel values ​​of N adjacent pixels to generate a second descriptor; Determine a Hamming distance between the first descriptor and the second descriptor, and if the Hamming distance is greater than a preset distance threshold, remove a first 2D sub-coordinate corresponding to the first descriptor and a second sub-2D coordinate corresponding to the second descriptor; The first 2D sub-coordinate is the 2D coordinate of the sub-joint point in the first 2D coordinate set, and the second 2D sub-coordinate is the 2D coordinate of the sub-joint point in the second 2D coordinate set.

4. An augmented reality (AR) device, characterized in that: Includes a left camera, a right camera, an inertial sensor, a memory, and a processor; The left camera is a fisheye lens, connected to the processor, and configured to capture a first hand image, where the first hand image is an RGB image; The right camera is a fisheye lens, connected to the processor, and configured to capture a second hand image, where the second hand image is an RGB image; The inertial sensor is connected to the processor and configured to collect motion data of the AR device; The memory is connected to the processor and configured to store computer program instructions; The processor is configured to perform the following operations according to the computer program instructions: Acquire a first hand image captured by the left camera and a second hand image captured by the right camera, and motion data of the AR device captured by the inertial sensor; Determine the posture information of the AR device according to the first hand image, the second hand image and the motion data; wherein the posture information includes position information and rotation angle information; Extracting a first 2D coordinate set of hand joint points from the first hand image, and extracting a second 2D coordinate set of hand joint points from the second hand image; Determine the posture information of the left camera according to the position information of the AR device and the posture deviation of the left camera relative to the AR device, and determine the posture information of the right camera according to the position information of the AR device and the posture deviation of the right camera relative to the AR device; wherein each posture deviation includes a position deviation corresponding to the position information and a rotation angle deviation corresponding to the rotation angle information; Estimate the 3D coordinates of the hand joint points according to the first 2D coordinate set, the second 2D coordinate set, the posture information of the left camera, and the posture information of the right camera; The posture information (P1, R1) of the left camera is expressed as follows: P1=P0-L1, R1=R0*R01; The posture information (P2, R2) of the right camera is expressed as: P2=P0-L2, R2=R0*R02; The P0 represents the position information of the AR device, the R0 represents the rotation angle information of the AR device, the L1 represents the position deviation of the left camera relative to the AR device, and the L2 represents the position deviation of the right camera relative to the AR device. The R01 represents the rotation angle deviation of the left camera relative to the AR device, and the R02 represents the rotation angle deviation of the right camera relative to the AR device.

5. The AR device according to claim 4, characterized in that: After the processor extracts a first 2D coordinate set of hand joint points from the first hand image and extracts a second 2D coordinate set of hand joint points from the second hand image, and before estimating the 3D coordinates of the hand joint points, the processor is further configured to: The 2D coordinates of the corresponding hand joint points in the first 2D coordinate set and the second 2D coordinate set are matched with each other in feature points, and the 2D coordinates of the hand joint points that do not match each other in the first 2D coordinate set and the second 2D coordinate set are eliminated.

6. The AR device according to claim 5, characterized in that: The processor performs feature point matching on the first 2D coordinate and the second 2D coordinate, and removes the unmatched first 2D coordinate and the second 2D coordinate, and is specifically configured as follows: For the first 2D sub-coordinate and the second 2D sub-coordinate of each sub-joint point in the hand joint point, perform the following operations: Compare the pixel point corresponding to the first 2D sub-coordinate in the first hand image with the pixel values ​​of N adjacent pixels to generate a first descriptor, where N is an integer greater than or equal to 1; Compare the pixel point corresponding to the second 2D sub-coordinate in the second hand image with the pixel values ​​of N adjacent pixels to generate a second descriptor; Determine a Hamming distance between the first descriptor and the second descriptor, and if the Hamming distance is greater than a preset distance threshold, remove a first 2D sub-coordinate corresponding to the first descriptor and a second sub-2D coordinate corresponding to the second descriptor; The first 2D sub-coordinate is the 2D coordinate of the sub-joint point in the first 2D coordinate set, and the second 2D sub-coordinate is the 2D coordinate of the sub-joint point in the second 2D coordinate set.

Citation Information

Patent Citations

  • Binocular three-dimensional imaging method and system

    CN109724537A

  • Gesture interaction method and device based on AR scene, storage medium and communication terminal

    CN110221690A

  • Virtual reality head-mounted display device positioning system based on binocular camera and IMU

    CN111307146A