Estimating a hand pose of a user

The method enhances hand pose estimation in HMIs by generating multidimensional point clouds and creating scene sequences using machine learning, addressing the issue of incomplete data capture for improved gesture recognition.

WO2025181078A1PCT designated stage Publication Date: 2025-09-04GESTIGON GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/055034
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2025-02-25
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing gesture recognition systems for contactless human-machine interfaces (HMIs) face reduced reliability when incomplete measurement data of a hand pose is captured, particularly when only a single sensor is used, leading to inaccurate estimation of free-space gestures.

Method used

A method utilizing machine learning algorithms to generate multidimensional point clouds from sensor data, estimate reference points corresponding to hand anatomical structures, and create a scene sequence to link these clouds, enabling accurate estimation of hand poses even with incomplete data.

Benefits of technology

Improves the estimation of hand poses and associated gestures by enhancing the accuracy of hand pose estimation, particularly in scenarios where partial hand obscuration occurs, allowing for more reliable free-space gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025055034_04092025_PF_FP_ABST
    Figure EP2025055034_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for estimating a hand pose (100, 150) of a user for determining a user input at a human-machine interface, wherein the method comprises: receiving measurement data that are captured using sensors and represent a first hand pose (100) and a temporally subsequent second hand pose (150) of the user; generating a first multi-dimensional point cloud (140), which represents the first hand pose (100), using the measurement data; generating a second multi-dimensional point cloud (170), which represents the second hand pose (150), using the measurement data; estimating first positions of first reference points (130) of the first hand pose (100) in the first multi-dimensional point cloud (140) using an algorithm based on machine learning; supplementing the first multi-dimensional point cloud (140) with the estimated first reference points (130); creating a scene sequence by linking a plurality of in each case first points in the supplemented first multi-dimensional point cloud (140) to a plurality of in each case second points in the second multi-dimensional point cloud (170) using an algorithm; estimating second positions of second reference points of the second hand pose (150) in the second multi-dimensional point cloud (170) using the scene sequence; estimating the second hand pose (150) using the second reference points.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ESTIMATION OF A USER'S HAND POSE

[0002] The present invention relates to a method and a device for estimating a user's hand pose to determine a user input at a human-machine interface. Furthermore, the invention relates to a computer program configured to execute the method.

[0003] While traditional human-machine interfaces (abbreviated to "HMI" or more commonly "HMI") are usually based on a contact-based interaction between a user and a corresponding control element, such as a physical switch or a touch-sensitive surface or display device, HMIs are now known in which a gesture performed by a human user in free space without physical contact with an HMI device is sensorily detected. This sensory detection of such gestures, also known as "free space gestures," can be achieved, for example, using a suitable camera. Mathematical methods, particularly image recognition methods, are then used to draw conclusions from the sensory data regarding the intended user input via the detected gesture.

[0004] Particularly in the context of vehicle-related applications, such MMIs capable of recognizing free-space gestures can be used to provide inputs for operating the vehicle or a subsystem thereof. Furthermore, it is also possible for the free-space gestures to relate to the vehicle's surroundings, such as pointing at an interesting object external to the vehicle that is visible from the vehicle. The recognition of such a free-space gesture serves to provide input to a system intended to provide information related to the object or to prompt the vehicle to react to it.

[0005] Gesture recognition systems are known for implementing such contactless MMIs. These systems use calibrated 2D sensors to detect and track points along the contour of a user's hand, a hand pose, associated with specific features, particularly from above and the side. Based on the detected points, a pointing direction associated with the gesture performed in three-dimensional space can then be estimated. Such systems rely on the recognition of clearly distinguishable features (points, particularly reference points) of the human hand and require fixed and calibrated hardware, preferably with two sensors. When capturing measurement data of a hand pose, especially when only a single sensor is used, it can happen that part of the hand is obscured from the sensor, thus missing important information about the hand pose.Based on such incomplete measurement data with regard to the hand pose, the estimation of the hand pose and a resulting estimated free space gesture can only be carried out with reduced reliability.

[0006] The present invention is based on the object of further improving the estimation of a hand pose, in particular when measurement data for the respective hand pose are missing.

[0007] This object is achieved according to the teaching of the independent claims. Various embodiments and further developments of the invention are the subject of the dependent claims.

[0008] A first aspect of the solution relates to a method, in particular a computer-implemented method, for estimating a user's hand pose for determining a user input at a human-machine interface, wherein the method comprises: (i) receiving sensor-detected measurement data representing a first hand pose and a temporally subsequent second hand pose of the user; (ii) using the measurement data, generating a first multidimensional point cloud representing the first hand pose; (iii) using the measurement data, generating a second multidimensional point cloud representing the second hand pose; (iv) estimating first positions of first reference points, which correspond in particular to an anatomical structure of a hand of the user, of the first hand pose in the first multidimensional point cloud using an algorithm based on machine learning;(vi) Creating a scene sequence by linking a plurality of first points of the supplemented first multidimensional point cloud with a plurality of second points of the second multidimensional point cloud using an algorithm; (vii) Estimating second positions of second reference points, which correspond in particular to the anatomical structure of the user's hand, of the second hand pose in the second multidimensional point cloud using the scene sequence; (viii) Estimating the second hand pose using the second reference points.

[0009] The terms "comprises," "includes," "includes," "has," "has," "with," or any other variation thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a method or apparatus that includes or has a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or that are inherent in such a method or apparatus.

[0010] Furthermore, unless explicitly stated to the contrary, "or" refers to an inclusive "or" rather than an exclusive "or." For example, a condition A or B is satisfied by one of the following conditions: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).

[0011] The terms "a" or "an" as used herein are defined as "one or more." The terms "another" and "another," and any other variations thereof, are defined as "at least one other."

[0012] The term "plurality" as used here is to be understood as meaning "two or more".

[0013] The term “configured” or “set up” to fulfil a specific function (and respective variations thereof) as used here is to be understood to mean that the corresponding device is already in a design or setting in which it can carry out the function or is at least adjustable - i.e. configurable - so that it can carry out the function after being set accordingly. The configuration can be carried out, for example, by appropriately setting parameters of a process sequence or of switches or similar for activating or deactivating functionalities or settings. In particular, the device can have a plurality of predetermined configurations or operating modes, so that the configuration can be carried out by selecting one of these configurations or operating modes.

[0014] The term "point cloud" within the meaning of the invention refers in particular to a set of points in a vector space, in particular three-dimensional space or a two-dimensional subspace thereof, which has an organized or unorganized spatial structure ("cloud"). A point cloud is described by the points it contains, each of which can be recorded in particular by its positions specified by its spatial coordinates. Additional attributes such as geometric normals, color values, temperature values, recording times, measurement accuracies, or other information can be recorded for the points. The term "hand pose" within the meaning of the invention refers in particular to the position of a human hand and the associated fingers of a user at a particular point in time.In contrast to a gesture, especially a free-space gesture, which comprises a movement that may include a plurality of hand poses at different times, the hand pose is therefore static.

[0015] "Reference points" within the meaning of the invention are understood to mean, in particular, points of a, in particular, multidimensional point cloud, in particular a three-dimensional point cloud, which, through their position, in particular their spatial coordinates, in the point cloud, correspond to an anatomical structure of a human hand. In particular, these reference points in the point cloud can represent finger joints of the hand, fingertips, a central point in the palm of the hand, and / or points on the wrist. In particular, using known positions of the reference points mentioned as examples, a hand pose can be estimated using a suitable algorithm, in particular a so-called pose estimator.

[0016] A "TOF camera system" within the meaning of the invention is understood in particular to be a 3D camera system that can measure distances based on a time-of-flight (TOF) method. The measurement principle is based on illuminating the scene to be recorded using a light pulse, and the camera measures the time it takes for the light to reach the object and back for each pixel. Due to the constancy of the speed of light, this required time is directly proportional to the distance. The camera thus provides the distance to the object imaged on it for each pixel. In particular, the camera can have a charge-coupled 3D sensor, also known as a CCD (charge-coupled device).

[0017] The method according to the first aspect makes it possible to estimate a second hand pose, which occurred chronologically after a first hand pose and for which incomplete measurement data was acquired, using a scene sequence. Incomplete measurement data is understood to mean that not all parts of the hand were measured, in particular due to parts of the hand being hidden from the sensor for acquiring the measurement data. The scene sequence comprises a link between a first multi-dimensional point cloud, which represents the first hand pose, and a second multi-dimensional point cloud, which represents the second hand pose. In addition, reference points of the first hand pose for the first multi-dimensional point cloud are estimated during the scene sequence. The second hand pose can then be estimated using the chronologically previous first hand pose and the scene sequence.

[0018] Preferred embodiments of the method are described below, which can be combined with each other as well as with the other aspects described, unless this is expressly excluded or is technically impossible.

[0019] In some embodiments, a geometric line, in particular a straight line, is determined between two first reference points in the first multidimensional point cloud that are adjacent with respect to their respective first positions. The geometric line can represent a bone of a hand that connects the respective first reference points, in particular joints of the hand. By using the geometric line in the scene sequence, the change in the position of this geometric line, in particular of the corresponding bone, from the first hand pose to the second hand pose can also be determined. This can improve the scene sequence.

[0020] In some embodiments, a plurality of geometric lines are determined, wherein each geometric line of the plurality of geometric lines connects two adjacent first reference points in the first multidimensional first point cloud, wherein a skeletal structure for the first hand pose is determined using the plurality of geometric lines, and wherein this skeletal structure is used in creating the scene sequence. Using the skeletal structure makes the scene sequence more accurate. This is because, in addition to the first reference points, the skeletal structure is mapped in the scene sequence, whereby a movement of the skeletal structure from the first hand pose to the second hand pose can be represented in the scene sequence.

[0021] In some embodiments, the method further comprises: (i) using the second multi-dimensional point cloud, estimating and subsequently weighting third reference points of the second hand pose, which correspond in particular to the anatomical structure of the user's hand, at third positions;

[0022] (ii) weighting the second reference points estimated from the scene sequence;

[0023] (iii) comparing the respective weighted second reference points with the respective corresponding weighted third reference points; (iv) determining those preferred reference points whose weighting of the respective comparison has a higher weighting, and wherein these preferred reference points replace the second reference points for determining the second hand pose. This makes it possible to ensure that those reference points from the second reference points and the third reference points with a lower weighting are not taken into account for the estimation of the second hand pose, and thus the reference points with an overall higher weighting, i.e. the preferred reference values, are taken into account for the estimation of the second hand pose. This makes the estimation of the second hand pose more accurate or improved overall.

[0024] In some embodiments, the second reference points are each weighted with a first weighting factor x, and the third reference points are weighted with a second weighting factor (1 - x). This takes different weighting factors into account. On the other hand, the first weighting factor and the second weighting factor are interdependent, such that the first weighting factor is greater than the second weighting factor, and as a result the second reference points can have a higher weighting and can therefore be used preferentially over the third reference points. This can be advantageous because the second reference points are determined using the scene sequence, and therefore the third reference points, which are based on the measurement data for the second hand pose, are not used.This can be particularly advantageous if parts of the hand were not measured during the second hand pose, resulting in incomplete measurement data. Third reference points determined from this may therefore be less accurate or reliable.

[0025] In some embodiments, a minimum distance of the estimated second reference points to the first multidimensional point cloud is estimated using the second positions, wherein only those second reference points whose distance does not exceed a predetermined distance value are used to estimate the second hand pose. This allows the second reference points whose distance is greater than the predetermined distance, i.e., too large, to be excluded from the estimation of the second hand pose. This allows the accuracy of the estimation of the second hand pose to be further improved. In some embodiments, the sensor-detected measurement data represent a free-space gesture of the user, wherein the free-space gesture contains the first hand pose and the chronologically subsequent second hand pose, and wherein the free-space gesture is estimated using the estimated second hand pose, in particular using a gesture classifier.Due to the improved estimation of the second hand pose using the described method, the free-space gesture containing the second hand pose can also be estimated more accurately. The free-space gesture can be used to provide an input for operating the vehicle or a subsystem thereof, particularly in the context of a vehicle-related application.

[0026] In some embodiments, the sensory acquisition of the measurement data is carried out by a TOF camera system. The use of a TOF camera system represents a particularly effective and high-resolution implementation option for the at least one sensor required for sensory acquisition of the measurement data.

[0027] In some embodiments, three-dimensional measurement data is acquired through sensory acquisition of the measurement data, and the generated first multidimensional point cloud and the second multidimensional point cloud are each three-dimensional. This allows for improved estimation of the second hand pose overall, since depth information is also taken into account in a three-dimensional space. This allows, for example, curved fingers in a hand pose to be better detected.

[0028] A second aspect of the solution relates to a device for estimating a hand pose of a user, wherein the device is configured to carry out the method according to one of the preceding claims.

[0029] In some embodiments, the device comprises: (i) a sensor configured to acquire measurement data representing a first hand pose and a temporally subsequent second hand pose of the user; (ii) an estimator configured to: generate a first multi-dimensional point cloud representing the first hand pose using the measurement data; generate a second multi-dimensional point cloud representing the second hand pose using the measurement data; estimate first positions of first reference points of the first hand pose in the first multi-dimensional point cloud using an algorithm based on machine learning; supplement the first multi-dimensional point cloud with the estimated first reference points;to create a scene sequence by linking a plurality of first points of the supplemented first multidimensional point cloud with a plurality of second points of the second multidimensional point cloud using an algorithm; to estimate second positions of second reference points of the second hand pose in the second multidimensional point cloud using the scene sequence; to estimate the second hand pose using the second reference points;

[0030] The device can in particular be a control unit for a vehicle, in particular a motor vehicle.

[0031] A third aspect of the solution relates to a computer program configured to carry out the method according to the first aspect.

[0032] The computer program can, in particular, be stored on a non-volatile data carrier. This is preferably a data carrier in the form of an optical data carrier or a flash memory module. This can be advantageous if the computer program as such is to be handled independently of a processor platform on which the one or more programs are to be executed. In another implementation, the computer program can be present as a file on a data processing unit, in particular on a server, and can be downloadable via a data connection, for example the Internet or a dedicated data connection, such as a proprietary or local network. Furthermore, the computer program can have a plurality of interacting individual program modules.

[0033] The device according to the second aspect of the invention can accordingly comprise a program memory in which the computer program is stored. Alternatively, the device can also be configured to access an external computer program, for example, available on one or more servers or other data processing units, via a communication connection, in particular to exchange data with it that is used during the execution of the method or computer program or that represents outputs of the computer program.

[0034] The features and advantages explained with regard to the first aspect of the solution also apply accordingly to the other aspects described. Further advantages, features, and possible applications emerge from the following description of preferred embodiments in conjunction with the figures.

[0035] This shows

[0036] Figs. 1A and 1B schematically show a camera and a 3D point cloud, respectively, according to an embodiment;

[0037] Fig. 2 schematically shows a flow diagram illustrating an embodiment of a method; and

[0038] Fig. 3 schematically shows a device for determining a hand pose.

[0039] Throughout the figures, the same reference numerals are used for the same or corresponding elements.

[0040] Figs. 1A and 1B schematically show a camera 110 and two 3D point clouds 140, 170, each of a hand pose 100, 150, according to one embodiment. Camera 110 is, in particular, a so-called time-of-flight (TOF) camera, which can also capture depth data, i.e., 3D data.

[0041] By way of example, Figs. 1A and 1B show a 3D spatial section 120 in the form of a cuboid, which can be sensor-captured by the TOF camera 110. A 3D point cloud can be generated from the captured data, which additionally has a corresponding timestamp. Fig. 1A schematically shows a first 3D point cloud 140, and Fig. 1B schematically shows a second 3D point cloud 170, wherein the first 3D point cloud is captured at a different point in time, i.e., has a different timestamp, than the second 3D point cloud.

[0042] According to Fig. 1A, the first 3D point cloud 140 represents a first hand pose 100. The first 3D point cloud 140 shown in Fig. 1A represents a first hand pose 100 with a complete hand with five fingers. In the first hand pose 100, a plurality of first reference points 130 are marked by way of example. The first reference points 130 can, for example, have ends, joints, or other points characterizing a hand. First positions of these first reference points 130 in the first 3D point cloud can be estimated by an algorithm based on machine learning. Therefore, it is possible that a first reference point 130 itself is not part of the 3D point cloud 140, but is located, for example, within the 3D point cloud 140, which, due to the underlying sensory detection, represents a shell of the hand pose 100.

[0043] Furthermore, using the first reference points 130 and / or using a comparable and already stored hand pose, a skeleton or a bone structure of the hand or the associated hand pose can be determined.

[0044] Fig. 1A shows, by way of example, individual first geometric lines 135, in particular straight lines, of such a skeleton relating to the first hand pose 100. Two first reference points 130 are connected by a first geometric line 135, which represents a bone of a hand. By connecting adjacent first reference points 130 with geometric lines, the positions of additional points lying on this line can be determined, and a schematic skeleton of a hand can be determined. These additional points or the skeleton can be used for an improved scene sequence, as explained below in Fig. 2.

[0045] According to Fig. 1B, the second 3D point cloud 170 represents a second hand pose 150. In the second 3D point cloud 170, third reference points 160 are illustrated by way of example at their respective third positions. Two third reference points 160 can be connected by a second geometric line 165.

[0046] The second 3D point cloud 170 shown in Fig. 1B represents a second hand pose 150 in which parts of the hand are missing, in particular the palm and also parts of the fingers adjacent to the palm. By way of example, missing third reference points 180 are also shown in the second hand pose 150.

[0047] In other words, for the second 3D point cloud 170, not all measurements were recorded that would have been required to represent a substantially complete hand or second hand pose 150. This may have occurred because the missing hand parts were obscured from the TOF camera.

[0048] As can be seen in Fig. 1B, the second hand pose 150 is different from the first hand pose 100. Thereafter, there was a movement of the hand between the first hand pose 100 and the second hand pose 150, which changed its pose. The first positions, second positions, and other positions and geometric details in Figs. 1A and 1B can each be described as spatial coordinates in the x, y, and z directions using the schematically illustrated coordinate system.

[0049] Fig. 2 schematically shows a flowchart 200 for illustrating an embodiment of a method for estimating a hand pose 150 of a user for determining a user input at a human-machine interface.

[0050] In a first step S210 of the method, measurement data representing a first hand pose 100 and a temporally subsequent second hand pose 150 of the user are acquired by sensors. The sensory acquisition of the measurement data can be carried out, in particular, by a TOF sensor 110.

[0051] In a further step S220 of the method, a first 3D point cloud 140 is generated using the measurement data, which represents the first hand pose 100.

[0052] In a further step S230 of the method, a second 3D point cloud 170 is generated using the measurement data, which represents the second hand pose 150.

[0053] In a further step S240, first positions of first reference points 130 of the first hand pose 100 in the first 3D point cloud 140 are estimated using an algorithm based on machine learning.

[0054] In a further step S250 of the method, the first 3D point cloud 140 is supplemented with the estimated first reference points 130.

[0055] In a further step S260 of the method, a scene sequence is created by linking a plurality of first points of the supplemented first 3D point cloud 140 with a plurality of second points of the second 3D point cloud 170 using an algorithm.

[0056] Another option for establishing a link between the two point clouds 140, 170 in order to estimate the movement of a scene can be achieved by a so-called "coherent point drift algorithm." In a further step S270 of the method, second positions of second reference points (not shown here) of the second hand pose 150 in the second 3D point cloud 170 are estimated using the scene history.

[0057] In a further step S280, the second hand pose 150 is estimated using the second reference points.

[0058] In Fig. 1B, third reference points 160 are shown for the second hand pose 150. These third reference points 160 arise from the second 3D point cloud 170 according to Fig. 1B. In method step S260, however, the second reference points are determined without using the third reference points 160. However, this can be done within the scope of a further exemplary embodiment in which the third reference points 160 are first estimated, then weighted, and finally compared with weighted corresponding second reference points. The respective reference point from a second reference point and a third reference point 160, which has the higher weighting, can then be used to estimate the second hand pose 150. This allows a more precise estimate of the second hand pose 150, as performed in method step S280.The reference points with the higher weighting are then used accordingly as second reference points in step S280.

[0059] The method can be used, for example, for a gesture control application in a vehicle. The TOF sensor 110 can be mounted in the vehicle's roof lining. The user can perform a gesture or a free-space gesture in an area above the gear lever, which is recorded as measurement data by the TOF sensor. This recording includes the first hand pose 100 and the second hand pose 150. The first hand pose 100 is represented by a generated dense first 3D point cloud 140, whereas the generated second 3D point cloud of the second hand pose 150, as shown in Fig. 1B, is incomplete or incorrect. This can occur because the user rotated or moved the hand between the first hand pose 100 and the second hand pose 150, and the TOF sensor can no longer record measurement data for the entire hand. Therefore, a pose estimation of the second hand pose 150 will be inaccurate because important hand information is missing.Using the method described above, an improved second hand pose 150 is estimated. This improved second hand pose 150, as well as the first hand pose 100, and possibly further hand poses, can be used as input to a trained gesture classifier, which then determines and outputs a gesture designation. Fig. 3 schematically shows a device 300 for determining a first hand pose 100 and a second hand pose 150 according to one embodiment. The device 300 has an estimation device 310 and a sensor 110, in particular a TOF camera, for acquiring measurement data representing a first hand pose 100 and a second hand pose 150 of a user. Other sensor types, in particular sensors based on ultrasound or radar measurement, can also be used instead or in combination therewith.The estimation device 310 can, in particular, be a computer, which can be equipped, in particular, with a processor platform 320 and a program and data memory 330 as well as a data output 340 for outputting output data determined by the computer, in particular in the form of hand poses 100, 150. A computer program consisting of one or more program modules can be stored in the program memory 330 and is configured, when executed on the processor platform 320, to cause the estimation device 310 to execute the method according to the invention, for example, as described with reference to Figure 2.

[0060] While at least one exemplary embodiment has been described above, it should be appreciated that a large number of variations exist. It should also be noted that the described exemplary embodiments are only non-limiting examples and are not intended to limit the scope, applicability, or configuration of the devices and methods described herein. Rather, the foregoing description will provide a guide to implementing at least one exemplary embodiment, with the understanding that various changes in the operation and arrangement of the elements described in an exemplary embodiment may be made without departing from the subject matter defined in the appended claims, as well as their legal equivalents.

[0061] 100, 150 First, second hand pose

[0062] 110 TOF camera

[0063] 120 Captureable 3D area

[0064] 130, 160 First, third reference points

[0065] 140, 170 First, second 3D point cloud

[0066] 135, 165 First, second geometric line

[0067] 180 Missing third reference points

[0068] 200 Flowchart

[0069] S210 Recording measurement data

[0070] S220 Generating a first point cloud

[0071] S230 Generating a second point cloud

[0072] S240 Estimating initial reference points

[0073] S250 Adding the first point cloud

[0074] S260 Creating a Scene Flow

[0075] S270 Estimating second reference points

[0076] S280 Estimating the second hand pose

[0077] 300 device

[0078] 310 Estimating facility

[0079] 320 processor platform

[0080] 330 program and data memory

[0081] 340 Data output

Claims

CLAIMS 1. A method for estimating a hand pose (100, 150) of a user to determine a user input at a human-machine interface, the method comprising: Receiving sensor-detected measurement data representing a first hand pose (100) and a temporally subsequent second hand pose (150) of the user; Using the measurement data, a first multidimensional point cloud (140) is generated, which represents the first hand pose (100); Using the measurement data, a second multidimensional point cloud (170) is generated, which represents the second hand pose (150); Estimating first positions of first reference points (130) of the first hand pose (100) in the first multidimensional point cloud (140) using an algorithm based on machine learning; Supplementing the first multidimensional point cloud (140) with the estimated first reference points (130); Creating a scene sequence by linking a plurality of first points of the supplemented first multi-dimensional point cloud (140) with a plurality of second points of the second multi-dimensional point cloud (170) using an algorithm; Estimating second positions of second reference points of the second hand pose (150) in the second multidimensional point cloud (170) using the scene history; Estimating the second hand pose (150) using the second reference points.

2. Method according to claim 1, wherein a geometric line (135) connecting the two first adjacent reference points (130) is determined between two first reference points (130) in the first multidimensional point cloud (140) which are adjacent with respect to their respective first positions.

3. The method according to claim 2, wherein a plurality of geometric lines (135) are determined, each geometric line (135) of the plurality of geometric Lines each connect two adjacent first reference points (130) in the first multidimensional first point cloud (140), wherein a skeletal structure for the first hand pose (100) is determined using the plurality of geometric lines (135), and wherein this skeletal structure is used in creating the scene sequence.

4. The method according to claim 1 to 3, further comprising: Using the second multidimensional point cloud (170), third reference points (160) at third positions of the second hand pose (150) are estimated and subsequently weighted; the second reference points estimated from the scene sequence are weighted; Comparing the respective weighted second reference points with the respective corresponding weighted third reference points (160); Determining those preferred reference points whose weighting of the respective comparison has a higher weighting, and wherein these preferred reference points replace the second reference points for determining the second hand pose.

5. Method according to one of the preceding claims, wherein the weighting of the second reference points is carried out with a first weighting factor x, and the weighting of the third reference points (160) is carried out with a second weighting factor (1 -x).

6. Method according to one of the preceding claims, wherein using the second positions a minimum distance of the estimated second reference points to the first multi-dimensional point cloud (140) is estimated, wherein only the second reference points are used to estimate the second hand pose (150), the distance of which does not exceed a predetermined distance value.

7. The method according to any one of the preceding claims, wherein the sensor-detected measurement data represent a free-space gesture of the user, wherein the free-space gesture includes the first hand pose (100) and the temporally subsequent second hand pose (150), and wherein the free-space gesture is estimated using the estimated second hand pose (150).

8. Method according to one of the preceding claims, wherein the sensory acquisition of the measurement data is carried out by a TOF camera system.

9. Method according to one of the preceding claims, wherein three-dimensional measurement data are acquired by the sensory acquisition of the measurement data, and the generated first multi-dimensional point cloud (140) and the second multi-dimensional point cloud (170) are each three-dimensional.

10. Device (300) for estimating a hand pose (100, 150) of a user, wherein the device is configured to carry out the method according to one of the preceding claims.

11. Device (300) according to claim 10, comprising: A sensor (110) configured to acquire measurement data representing a first hand pose (100) and a temporally subsequent second hand pose (150) of the user; An estimating facility (310) which is set up: Using the measurement data, to generate a first multidimensional point cloud (140) representing the first hand pose (100); Using the measurement data, to generate a second multidimensional point cloud (170) representing the second hand pose (150); Estimating first positions of first reference points (130) of the first hand pose (100) in the first multidimensional point cloud (140) using an algorithm based on machine learning; supplementing the first multidimensional point cloud (140) with the estimated first reference points (130); Creating a scene sequence by linking a plurality of first points of the supplemented first multidimensional point cloud (140) with a plurality of second points of the second multidimensional point cloud (170) using an algorithm; Estimate second positions of second reference points of the second hand pose (150) in the second multidimensional point cloud (170) using the scene history; Estimate the second hand pose (150) using the second reference points.

12. A computer program configured to carry out the method according to any one of claims 1 to 9

Citation Information

Patent Citations

  • Information processor, control method, and program

    US20200202609A1

  • Machine learning based activity detection utilizing reconstructed 3D arm postures

    US20220051145A1

  • Detection processing device, detection processing method, information processing system

    US20240160294A1

  • Detection processing device, detection processing method, and information processing system

    WO2022196222A1