Method and device for estimating the positions of human key points

A neural network-based method with database-assisted scaling and camera parameter utilization addresses the challenge of obscured key points, enhancing the accuracy of three-dimensional human body point estimation.

WO2026006864A1PCT designated stage Publication Date: 2026-01-08EMOTION3D GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/AT2025/060264
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-06-27
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing methods using depth cameras struggle to accurately measure the three-dimensional positions of key human body points obscured by body limbs, seat belts, or clothing, leading to inaccuracies in determining absolute coordinates.

Method used

A method utilizing a neural network to detect two-dimensional key points in images, followed by a database query to select key points with low variance in average distances, and a position determination unit to calculate absolute coordinates using intrinsic camera parameters, ensuring accurate scaling and transformation.

Benefits of technology

Enables precise estimation of absolute three-dimensional positions of key human body points, overcoming obscuration issues and improving accuracy in vehicle occupant detection systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AT2025060264_08012026_PF_FP_ABST
    Figure AT2025060264_08012026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method and a device for estimating the absolute positions S1..N (abs) of a number N of key points of a person in a three-dimensional space, comprising the steps of: capturing at least one two-dimensional image (2) of the space using at least one image capturing unit (1); receiving the image (1) using a data processing unit (3); detecting, using an image processing unit (4), a number N of two-dimensional key points of the person in the image (2), and estimating the three-dimensional relative positions of the detected key points S1..N (rel) and the estimated distances d(rel).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and device for estimating the positions of human key points

[0002] The invention relates to a method and a device for estimating the positions of key human points in a three-dimensional space.

[0003] Methods for determining the three-dimensional positions of key points of the human body, in particular the positions of the head, eyes, nose, ears, sternum, and various human joints, such as shoulder joints, elbow joints, knee joints, wrist joints, and hip joints, are known from the prior art. Such methods are used particularly in vehicles to detect the positions and body postures of the vehicle occupants.

[0004] To determine the absolute positions of the key points—that is, their three-dimensional positions within the vehicle's coordinate system—image acquisition units in the form of depth cameras or time-of-flight (ToF) cameras are typically used. These cameras use dedicated sensors to measure the distance to the camera for each pixel, the so-called pixel depth. Since the camera's absolute coordinates are calibrated in advance within the vehicle, the absolute coordinates of the key points can be calculated from the measured pixel depths.

[0005] However, the use of depth cameras presents the problem that certain key points of the human body are often obscured, particularly in vehicles. For example, the pixel depths of key points obscured by body limbs, seat belts, or clothing cannot be accurately measured by the depth camera. One object of the invention is to solve this problem and to provide a method and a device for estimating the positions of key points of the human body with greater accuracy.

[0006] These and other problems are solved according to the invention using a method according to claim 1.

[0007] A method according to the invention is used to estimate the absolute three-dimensional positions Si..N (abs) A number N of key points of a person in a three-dimensional space is formed. It comprises the following steps:

[0008] In a first step, an image acquisition unit captures at least a two-dimensional image of the space. This image acquisition unit can be a conventional 2D camera. The image is then received by a data processing unit.

[0009] The data processing unit can be designed as a microcontroller or microcomputer and may include a central processing unit (CPU), volatile semiconductor memory (RAM), non-volatile semiconductor memory (ROM, SSD hard drive), magnetic storage (hard drive) and / or optical storage (CD-ROM), as well as interface units (Ethernet, USB) and the like. The components of such data processing units are generally known to those skilled in the art.

[0010] In the next step, an image processing unit detects the two-dimensional positions of the key points of the human SI ... N from the captured two-dimensional image. For this purpose, known pattern recognition algorithms, particularly from the field of machine learning, can be used. Specifically, a neural network trained to estimate the key positions of one or more people from 2D images can be employed. These key points can be predefined points on the human body, including wrists, arm joints, elbow joints, shoulder joints, knee joints, ankle joints, hip joints, as well as the center of the head, the sternum, and the positions of the eyes, ears, nose, and / or mouth.

[0011] Following or simultaneously with the detection of the two-dimensional key points of the human in the image, an estimation of the relative three-dimensional positions Si..N is performed. (rel) of the detected key points. This estimation can preferably also be performed by a neural network trained to extract the 2-dimensional coordinates and depth vectors (i.e., the distances of the key points from the camera) from the image data. The estimation includes determining the relative distances d (rel) between each pair of identified key points. These estimated key points are therefore located in the camera's coordinate system, but the correct scaling is not yet known.

[0012] Preferably, both the detection of the two-dimensional key points and the estimation of the relative three-dimensional positions of the key points are performed by one and the same neural network, which has been specially trained beforehand to extract the two-dimensional key points of a person from a two-dimensional image and to estimate their respective distances from the camera in three-dimensional space.

[0013] According to the invention, for this purpose, two-dimensional images of people, whose key points are annotated and their 3D distances to each other are provided, can be used as training data for the neural network. However, instead of a neural network, any other method from the field of machine learning and artificial intelligence can also be used.

[0014] The result of the estimation is a list of relative positions of the key points Si..N. (rel) = [xi(rel) , yi (rel) , Zi (rel) ], where i = 1 ... N and N is the number of detected key points. While the estimated 3D positions of the key points are correct in their relations and 3D angles to each other, the absolute 3D coordinates of the key points cannot be derived from this estimate. For example, a tall person farther away from the camera may appear exactly the same in the 2D image as a short person closer to the camera.

[0015] To determine the absolute 3D coordinates of the key points, the next step involves querying a database using the data processing unit to select at least M = 3 key points S'I..M from the key points previously detected in the image. The goal is to select those key points whose absolute distance d (abs)On average, this exhibits the lowest static variance in the overall population. To this end, the database contains the average distances of numerous combinations of key human landmarks, such as the average distance between the eyes, ears, or the distance between the chin and nose. Furthermore, the database includes information on the variance of these values ​​within the population, i.e., information on how much these distances typically vary.

[0016] The result of the query is the absolute distance d. (abs) between those detected key points whose variance in the database is lowest. For example, an average interpupillary distance of d can be obtained from the database. (abs) =7cm can be removed.

[0017] Preferably, the selected key points can be key points from the head area, for example the positions of the eyes, nose, mouth and / or ears of the person, since key points from the head area are rigid and usually have low variance.

[0018] According to the invention, the image processing unit can estimate the gender and / or age of the person in the image and use this information when selecting key points and / or determining the average absolute distance d. (abs) This is taken into account. In order to also consider the gender and / or age of the detected person when selecting the key points, the database can contain the average values ​​of the distances between the key points and the variance of these values ​​for different genders and different age groups of the population.

[0019] The result of the database query is the additional information of the most probable three-dimensional absolute distance d. (abs) two key points detected in the image.

[0020] In a next step, one of the selected key points SI ...M, for example the position of the nose, is used as the origin point [0,0,0] and the M selected key points are scaled so that the distances between any two points have the value of d (abs) agree. The newly scaled and transformed key points S'i...M (rel) form a reference model.

[0021] The reference model is also still in the relative coordinate system, but the exact scaling of the key points 1 ... M has already been carried out. In the next step, the remaining key points Si ... N will be scaled. (rel) scaled to the key points of the reference model.

[0022] In the next step, the distance vectors Vi..M are calculated by a position determination unit. (abs) the image acquisition unit from the selected key points.

[0023] To determine the distance vectors Vi...M (abs) To calculate these parameters, the intrinsic geometric parameters of the image acquisition unit are used. These can include, in particular, the focal length, the principal point, and the distortion coefficients of a camera. For example, if it is known that the camera has a focal length of 50 mm, the vectors Vi...M can be derived from the reference model and the two-dimensional estimates of these key points. (abs)These calculations are performed using well-known programming libraries such as OpenCV and Dlib, which provide corresponding routines, as described in the internet article https: / / leamopencv.com / head-pose-estimation-using-opencv-and-dlib / . In the next step, the positioning unit calculates the absolute three-dimensional positions of the other key points Si..N. (abs) from the relative three-dimensional positions of the key points Si..N (rel) and the now known vectors Vi...M (abs) The key points from the camera. After all key points Si..N (rel) Once the 3D angles are correctly scaled and known, all remaining key points must be calculated using the vectors Vi ...M. (abs) They can be shifted. For example, shifting all points by the intersection of all vectors V is sufficient here. (abs) .

[0024] The calculation results in the absolute three-dimensional positions of the key points Si..N.(abs) known in the vehicle's coordinate system.

[0025] Both the image processing unit and the positioning unit can be provided as separate hardware units, or preferably as software modules in the RAM or ROM of the data processing unit. However, it is also possible for these units to be located externally, for example on a server on the internet, to which the necessary data is transferred.

[0026] According to the invention, the positions of the estimated key points can be transferred into a topological three-dimensional data model, in particular into a graph. The known three-dimensional physiology of the human body can be used to create the topological data model. For example, the identified position of an elbow joint can be represented as a key point in the form of a node in a graph and connected to the positions of the wrist and shoulder joints via edges.

[0027] The invention further relates to a computer-readable storage medium comprising instructions that cause a data processing unit to execute a method according to the invention.

[0028] The invention further relates to a device for estimating the absolute positions of key human points in a three-dimensional space, configured to perform a method according to the invention. The invention further relates to a vehicle comprising a device according to the invention.

[0029] Further features of the invention will become apparent from the claims, the exemplary embodiments and the figures.

[0030] The invention is explained below using an exemplary, non-exclusive embodiment.

[0031] Fig. 1 shows a schematic example of a device according to the invention;

[0032] Fig. 2a shows a schematic representation of a graph during the execution of a method according to the invention;

[0033] Fig. 2b shows a schematic representation of the two-dimensional and absolute three-dimensional positions of the key points during the execution of a method according to the invention.

[0034] Fig. 1 shows a schematic example of a device according to the invention for estimating the absolute positions of human key points Si - SN in a three-dimensional space. The device comprises an image acquisition unit 1, which is configured to acquire a two-dimensional image 2 and provides intrinsic geometric parameters 7, in particular the focal length of the image 2. Furthermore, the device comprises a data processing unit 3 with an image processing unit 4, a database 5, and a position determination unit 6.

[0035] Image processing unit 4 comprises a neural network trained to extract human key points from a two-dimensional image and simultaneously estimate their relative three-dimensional positions. Two-dimensional image data of humans with three-dimensionally annotated key points were used to train this network; that is, the images used for training contain images of humans whose key points are annotated with three-dimensional relative coordinates. Image processing unit 4 is first applied to image 2, so that the human key points Si - SN and their three-dimensional relative positions Si are determined in the two-dimensional image 2. (rel) - SN (rel) The key points are shown in Figure 1. They are represented by the Cartesian coordinates Xi. (rel) , yi (rel) , Zi (rel) They are stored as a data object in the form of a graph.

[0036] Fig. 2a shows a schematic representation of the graphs during the execution of a method according to the invention. On the left side is a graph showing the two-dimensional key points Si - SN of the human in Figure 2.

[0037] The neural network estimates relative three-dimensional coordinates of the key points Si from this. (rel) - SN (rel) as shown in the right-hand image. Since the training data for the neural network is stored as two-dimensional points combined with depth values ​​for these points, the estimated points are also located in the camera coordinate system with the camera origin as the coordinate origin.

[0038] In the next step, the resulting relative key points are transformed into a unit coordinate system. For example, the key points of the head have the following relative coordinates, with the tip of the nose defined as the origin point:

[0039] Key Point Xi (rel) yi (rel) Zi (rel)

[0040] Si< rel Left eye -225.0 170.0 -135.0

[0041] S2< rel > Right eye 225.0 150.0 -135.0

[0042] S3< rel) Tip of nose 0.0 0.0 0.0

[0043] S4 (rel) Sternum 0.0 -250.0 0.0

[0044] The coordinates of the other key points are transformed relative to the origin point, i.e., the tip of the nose. The distances dij indicated in the figure denote the distances between the key points Si and Sj.

[0045] Fig. 2b shows a schematic representation of the two-dimensional and three-dimensional positions of the key points during the execution of a method according to the invention. The relative three-dimensional coordinates Si are determined by the estimation and transformation described above. (rel) - SN (rel)the key points are known; however, information about the absolute coordinates of the key points in three-dimensional space is missing, i.e., their position and scaling in the coordinate system of camera 1.

[0046] To calculate these absolute three-dimensional coordinates, a database 5 is used, which stores the average distances of key human landmarks and their variance in the general population. For example, it might contain information that the interpupillary distance (IPD) in adult males has a probability of 95% between 7.0 cm and 7.2 cm. If the relative three-dimensional coordinates of the two eyes are known from the estimate, these coordinates are then transformed so that their scalar distance has a value of 7.1 cm. If the eyes are not visible in image 2, other key landmarks of the head are used, such as the distance between the nose and the chin. At least M = 3 points must be available for the next step.

[0047] If three of the detected key points are correctly scaled, the information about the distance and orientation to the image acquisition unit 1 is still missing. In Fig. 2b, the distance d taken from database 5 is shown. (abs) The selected key points are shown on the right.

[0048] In the next step, the three selected key points of the head area are transformed into a reference model whose origin point coincides with the camera origin. For example, the key points of the head have the following relative coordinates, with the tip of the nose defined as the origin point:

[0049] Key point Xi yi Zi

[0050] S'i< rel Left eye -22.5 17.0 -13.5

[0051] S'2 (rel) Right eye 22.5 15.0 -13.5

[0052] S'3 (rel)Nose tip 0.0 0.0 0.0 These selected points form a correctly scaled reference model in which the distances between the key points correspond to the actual average distances. Fig. 2b shows on the left the reference model calculated in this way for three exemplary key points Si' (rel) , S2 ,(rel) , S3 ,(rel) .

[0053] In the next step, a position determination unit 6 calculates the absolute vectors Vi. (abs) , V2 (abs) , vs (abs) the image acquisition unit 1 from the estimated two-dimensional key points Si (rel) , S2 (rel) , S3 (rel) and the reference model consisting of S'i (rel) , S'2 (rel) , S'3 (rel)in the coordinate system of the image acquisition unit. For this purpose, the focal length, the principal point of the image, and the distortion parameters of the image acquisition unit 1 are used as geometric parameters 7. Knowing these values, a projection of the 2D points Si, S2, and S3 onto the corresponding 3D points S1 can be performed, as shown in Fig. 2b. (abs) , S2 (abs) and S3( abs) be carried out. Consequently, the absolute coordinates Si (abs) , S2 (abs) and S3( abs) the three selected points in the coordinate system of the image acquisition unit 1, as well as the absolute vectors V (abs) Camera 1 is aware of these key points.

[0054] In the next step, the positioning unit 6 calculates the absolute coordinates of the remaining key points in the coordinate system of the image acquisition unit 1. After all key points Si..N (rel)Once the 3D angles are correctly scaled and known, all remaining key points must be determined using the vectors V. (abs) They can be shifted. For example, shifting all points by the intersection of all vectors V is sufficient here. (abs) .

[0055] However, the invention is not limited to this described embodiment, but also includes further embodiments of the present invention within the scope of the following patent claims.

Claims

Patent claims 1. Procedure for estimating the absolute positions SI..N (abs) a number N of key points of a person in a three-dimensional space, comprising the steps: a. acquisition, by at least one image acquisition unit (1), of at least one two-dimensional image (2) of the space; b. reception, by a data processing unit (3), of the image (1); c. detection, by an image processing unit (4), of a number N of two-dimensional key points of the person in the image (2), and estimation of the three-dimensional relative positions of the detected key points Si..N (rel) as well as the estimated distances d (rel) , characterized in that the following steps are carried out: d. query, by the data processing unit (3), a database (5) to select at least M = 3 detected key points S'I..M and determine the average distance d (abs)the selected key points, e.g. calculation of a reference model using the average distances d (abs) , encompassing the key points S'I...M, and scaling of the key points SI..N (rel) to the key points of the reference model, f. calculation, by a position determination unit (6), the distance vectors Vi..M (abs) The image acquisition unit (1) calculates the absolute three-dimensional positions of the key points SI..N from the two-dimensional key points S'I...M of the reference model, taking into account intrinsic geometric parameters (7) of the image acquisition unit (1), in particular the focal length f and the principal point, as well as g. The position determination unit (6) calculates these positions. (abs) from the relative positions of the key points Si..N (rel) and the vectors Vl...M (abs) .

2. Method according to claim 1, characterized in that the key points are predefined points of the human body, in particular wrists, arm joints, elbow joints, shoulder joints, knee joints, ankle joints, hip joints, as well as the center of the head, sternum, eyes, ears, nose, and / or mouth of a person.

3. Method according to claim 1 or 2, characterized in that for the detection of the two-dimensional key points of the human being in the image (2) and for the estimation of the three-dimensional relative positions of the detected key points SI..N (rel) one and the same neural network is used.

4. Method according to one of claims 1 to 3, characterized in that the selected key points are those key points whose average distances have the lowest statistical variance of all key points detected in the image (2).

5. Method according to claim 4, characterized in that the selected key points are key points from the head area of ​​the human being, for example the positions of the eyes, nose, mouth and / or ears of the human being.

6. Method according to any one of claims 1 to 5, characterized in that the image processing unit estimates the gender and / or age of the person in the image and uses this information when selecting and / or determining the average absolute distance d (abs) considered between the key points 7. Method according to one of claims 1 to 6, characterized in that the key points are stored by the data processing unit (3) in electronically readable form, for example in the form of a table or a graph.

8. Computer-readable storage medium comprising instructions that cause a data processing unit (3) to execute a method according to any one of claims 1 to 7.

9. Device for estimating absolute positions SI ..N (abs) a number N of key points of a person in a three-dimensional space, comprising a. an image acquisition unit (1) configured to acquire at least one two-dimensional image (2) of the space, b. a data processing unit (3) for receiving the image (2), c. an image processing unit (4) i. for detecting a number N of two-dimensional key points of the person in the image (2), and ii. for estimating the relative three-dimensional positions of the detected key points Si..N (rel) as well as the estimated distances d (rel)is designed in that d. the data processing unit (3) is used to query a database (5) to select at least M = 3 detected key points and to determine the average distance d (abs) the selected key points are formed, whereby e. from these average distances d (abs) a reference model consisting of the key positions S'I...M is generated and f. a position determination unit (6) is provided for calculating i. the absolute vectors Vi...M (abs) the image acquisition unit (1 ) is formed from the two-dimensional key points S'I...M of the reference model taking into account intrinsic geometric parameters (7) of the image acquisition unit (1 ), in particular the focal length f and the principal point of the image, and ii. the absolute three-dimensional positions of the key points SI..N (abs) from the relative positions of the key points Si..N (rel)and the values ​​of Vi...M (abs) is trained.

10. Vehicle comprising a device according to claim 9.

Citation Information

Patent Citations

  • A camera-based distance measurement method, control method, device and electronic equipment

    CN114910052B

  • Method and system for multi-view image processing with accurate three-dimensional skeletal reconstruction

    CN117561546A

  • Personalized hrtfs via optical capture

    US20210211825A1

  • Method and device for adjusting or controlling a vehicle component

    US20230120314A1