Computer-implemented method for creating a 3D body joint position of an object

A method for determining 3D body joint positions by estimating 2D positions and using statistical error modeling addresses inefficiencies in existing methods, providing accurate and flexible 3D pose estimation across diverse environments.

DE102024201217A1Pending Publication Date: 2025-08-14ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024201217
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing 3D pose estimation methods for body joints require large amounts of annotated data and are limited to interior recordings, making them inefficient and inflexible for various application scenarios.

Method used

A computer-implemented method that determines 3D body joint positions by first estimating 2D joint positions, manually annotating errors, and using a statistical error model to generate accurate 3D positions through probabilistic modeling, allowing flexible application in different environments.

Benefits of technology

Enables efficient and accurate determination of 3D body joint positions without extensive data sets, enabling applications beyond interior spaces and improving the flexibility and accuracy of 3D pose estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The disclosure includes a computer-implemented method (100) for determining a 3D body joint position (20) of an object (70) comprising the following steps: - generating (102) a first 2D body joint position (12) for at least one body joint point (10) of the object (70) based on a body joint position data set (30) which contains at least one image from a plurality of camera perspectives; - Manually annotating (104) a second 2D body joint position (14) based on the body joint position data set (30) to determine an error deviation of the second 2D body joint position (14) from the generated first 2D body joint position (12); - determining (106) an error model (50) by estimating a statistical error value indicating a deviation (55) of the first 2D body joint position (12) from the second 2D body joint position data (14); - Determining (108) the 3D body joint position (20) for the at least one body joint point (10) on the basis of the first 2D body joint position (12) and the determined error model (50) by applying a statistical method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a computer-implemented method for creating a 3D body joint position of an object. State of the art

[0002] 3D pose estimation involves estimating the positions of 3D body joints of people or animals based on input data, typically available as (RGB) camera images. This 3D pose estimation is typically learned using deep learning models.

[0003] However, training deep learning models for the 3D pose of an object—be it a human or an animal—requires large amounts of annotated or labeled 3D pose data. This data is currently typically captured using a motion capture system with optical markers or a multi-view camera system with synchronized cameras. Due to the size and limitations of these systems, however, these recordings are usually limited to indoor shots of individual people.

[0004] It is therefore the object of the present invention to provide a solution by means of which 3D poses in the form of 3D body joint positions for an object can be determined in an efficient and improved manner. Disclosure of the invention

[0005] This problem is solved by a computer-implemented method for creating a 3D body joint position of an object having the features of the independent claim.

[0006] According to a first aspect, the disclosure relates to a computer-implemented method for creating a 3D body joint position of an object, comprising the following steps: In a first step, a first 2D body joint position is determined for at least one body joint point of the object based on a body joint position data set which contains at least one image from several camera perspectives.

[0007] In a second step, a second 2D body joint position is manually annotated based on the body joint position data set in order to determine an error deviation of the second 2D body joint position from the generated first 2D body joint position.

[0008] In a third step, an error model is determined by estimating a statistical error value that indicates a deviation of the first 2D body joint position from the second 2D body joint position data.

[0009] In a fourth step, the 3D body joint position for the at least one body joint point is determined based on the first 2D body joint position and the determined error model by applying a statistical method.

[0010] A fundamental idea of ​​the present invention is that, based on a 2D pose estimator for at least one determined 2D body joint position of an object using a specific (image) data set containing captured images from multiple camera perspectives of the object, a statistical error value of the 2D pose estimator is determined in order to estimate or determine a 3D body joint position for a body joint point of the object.

[0011] A further aspect of the present invention therefore consists in extending the existing approach for estimating a 2D pose for an object to generate a 3D pose for this object. Since the existing approach to 2D pose estimation is error-prone, it is fundamental for the present invention to model this error - using a few, labeled images. In addition, the determined error measure for estimating the 3D pose points or the 3D body joint points is modeled probabilistically in order to take into account the previously determined error model for the 2D pose or 2D body joint position, respectively, and to be able to provide a confidence estimate for the generated 3D pose or body joint position. The proposed approach of the present invention advantageously allows for the simple modeling of further framework conditions.

[0012] Thus, the present invention allows for the efficient determination of more accurate 3D body joint positions for objects, whether humans or animals, without requiring large amounts of data sets. Furthermore, the determination of these 3D body joint positions is not limited to indoor spaces, but can be flexibly applied to any other location, such as the interior of a vehicle. This makes the present approach flexible for a variety of application scenarios.

[0013] For the purposes of the present invention, a 2D body joint position is defined as information about a position of a joint point of a body joint of an object, such as a human or an animal, in a 2D space.

[0014] Accordingly, a 3D body joint position is defined as information about a position of a joint point of a body joint of an object, such as a human or an animal, in a 3D space.

[0015] A possible embodiment of the method provides that the determination of the 3D body joint position further comprises the following steps: - estimating at least one 3D body joint position for the at least one body joint point based on the first 2D body joint position; - Projecting the at least one estimated 3D body joint position into the image of the camera perspective on the basis of which the first 2D body joint position was formed; - Evaluate the estimated at least one 3D body joint position, how likely the estimated at least one 3D body joint position is based on the previously determined error model.

[0016] One possible embodiment of the method provides for the evaluation step to be performed by using either the most probable 3D body joint position of all estimated 3D body joint positions or a weighted average of all estimated 3D body joint positions. This efficiently selects the best estimate of the object's 3D body joint position.

[0017] One possible embodiment of the method provides for the first 2D body joint position to be generated by applying a machine learning (ML) model. This provides the advantage that the first 2D body joint position can be generated efficiently and easily.

[0018] One possible implementation of the method involves using the Markov Chain Monte Carlo method or the variational inference method as the statistical method. This allows for easy mapping of application-specific scenarios when generating the 3D body joint position.

[0019] One possible embodiment of the method involves determining the error model using a Gaussian mixture model. This offers the advantage of being able to create the error model simply and efficiently.

[0020] One possible embodiment of the method provides for the generation of the 3D body joint position by incorporating at least one of the following pieces of information: temporal information representing a sequence of different camera images; information about a dimension of at least one limb for the respectively determined 2D body joint position and / or determined 3D body joint position; and information about a spatial restriction for the determined 3D body position. In this way, the quality of the determined 3D body joint positions for the selected object can be flexibly and easily adapted to different application scenarios and requirements depending on the available data.

[0021] According to a second aspect, the disclosure relates to a machine-readable data carrier and / or download product comprising the computer program.

[0022] According to a third aspect, the disclosure relates to one or more computers and / or computer instances with the computer program, and / or with the machine-readable data carrier and / or the download product.

[0023] Further measures improving the invention are presented in more detail below together with the description of the preferred embodiments of the invention with reference to figures. Examples of implementation

[0024] It shows: Fig. 1 Schematic flow diagram of the computer-implemented procedure for creating a 3D body joint position of an object.

[0025] Fig. 1 shows a schematic flow diagram of the computer-implemented method 100 for creating a 3D body joint position 20 of an object 70.

[0026] In step 102, a first 2D body joint position 12 is determined for at least one body joint point 10 of the object 70 based on a body joint position data set 30 containing at least one image from multiple camera perspectives. This data set 30 can be a specific data set. The object 70 can be a human or an animal.

[0027] Optionally, the first 2D body joint position 12 can be generated by applying a machine learning (ML) model 40.

[0028] In step 104, a second 2D body joint position 14 is manually annotated or labeled on the basis of the body joint position data set 30 in order to determine an error deviation of the second 2D body joint position 14 from the generated first 2D body joint position 12.

[0029] In a step 106, an error model 50 is determined by estimating a statistical error value that indicates a deviation 55 of the first 2D body joint position 12 from the second 2D body joint position data 14. The estimation of the statistical error value generally refers to an error estimate or an uncertainty estimate for the first 2D body joint position. In step 106, the results from steps 102 and 104 are processed together in a corresponding manner.

[0030] Optionally, the error model 50 can be determined using a Gaussian mixture model.

[0031] In a step 108, the 3D body joint position 20 for the at least one body joint point 10 is generated 108 on the basis of the first 2D body joint position 12 and the determined error model 50 or the determined error measure of the first 2D body joint position 12 by applying a statistical or probabilistic method.

[0032] Optionally, the step 108 of determining the 3D body joint position 20 may further comprise the following steps: - estimating 110 at least one 3D body joint position 20 for the at least one body joint point 10 based on the first 2D body joint position 12; - Projecting 112 the at least one estimated 3D body joint position 20 into the image of the camera perspective on the basis of which the first 2D body joint position 12 was formed; - Evaluate 114 the estimated at least one 3D body joint position 20, how likely the estimated at least one 3D body joint position 20 is based on the previously determined error model 50.

[0033] Optionally, the step 114 of evaluating may be performed by using either the most probable 3D body joint position of all estimated 3D body joint positions 20 or a weighted average of all estimated 3D body joint positions 20.

[0034] Optionally, the generated 3D body joint position 20 can be used to generate training data to train a machine learning (ML) model configured to automatically determine a 3D body joint position of the object 70. The ML model receives 2D images as input data and, after processing them, outputs corresponding 3D images of estimated 3D body joint positions of the object 70.

[0035] The method described above can be implemented using a technical system for generating a 3D body joint position 20. The technical system is configured to use the generated 3D body joint position to create training data, which is used as input data for training an ML model to automatically determine a 3D body joint position of the object 70 using the trained ML model.

[0036] The present invention can be applied in particular to a vehicle having such a technical system in order to control an automated vehicle function based on information about the generated 3D body joint information.

[0037] An important aspect of the present invention therefore consists in applying a confidence / error estimate to the determined first 2D body joint position 12 in order to determine the corresponding 3D body joint position 20 for the at least one body joint point 10, which corresponds to the first 2D body joint position 12. The determined error model 50 or the error estimate for the first 2D body joint position 12 is used to determine how error-prone the first 2D body joint position 12 is. Or, more generally, the determined error model 50 provides a measure of this or how strongly the determined error is distributed among the individual 2D body joint positions 12.

[0038] This information is then used to generate a corresponding 3D body joint position from the 2D body joint position with the corresponding error information using probabilistic modeling methods. This position includes confidence / uncertainty information. This additional confidence / uncertainty information indicates how likely it is that the joint point is actually located at the determined 3D body joint position 20 of the joint point.

[0039] The statistical or probabilistic method can preferably be implemented as a Markov Chain Monte Carlo optimization method or as a variational inference method. However, the present invention is not limited to the application of these statistical methods.

[0040] Optionally, the 3D body joint position 20 can be generated by including at least one of the following information: Inclusion of temporal information representing a sequence of different camera images, information about a dimension of at least one limb for the respectively determined 2D body joint position 12, 14 and / or determined 3D body joint position 20, information about a spatial restriction, for example for a person within a passenger compartment of a vehicle, for the determined 3D body position 20.

[0041] The 3D body joint position 20 generated in this way for the at least one body joint point 10 of the object 70 can then be used, for example, to generate corresponding training data to train a machine learning ML system that automatically estimates or recognizes 3D body joint positions of any object. In addition to humans and animals, semi-automated systems such as robots or robot-like systems can also be considered as objects. These systems have body joints whose 3D body joint positions in space are to be determined with a specified degree of certainty in order to be able to adapt or control their behavior accordingly in a targeted and application-dependent manner.

[0042] In particular, the information about the generated 3D body joint position 20 can be used to detect or predict a state of an object, for example a driver or a person in a vehicle that includes functions for automated driving or autonomous driving.

[0043] If the person being observed is in a vehicle that has implemented at least partially automated vehicle functions or automated vehicle assistance functions, the driving behavior of the vehicle or the autonomously driving system can be flexibly adapted based on the generated or estimated 3D body joint position according to a detected situation in which the person is located.

[0044] For example, a driver's fatigue level or, more generally, a critical driver situation can be detected, predicted or estimated in order to then activate corresponding vehicle functions that put the vehicle into a safe mode or implement a defined safety level of the technical system.

[0045] Further aspects of the present invention are explained below for better understanding: In the current state of the art, a 3D pose or a 3D joint position for a body joint of an object is calculated using a synchronized multi-camera system and labeled or estimated 2D poses using a compensation method across the various camera images. The reprojection error is used as an error measure for optimization. The 3D pose is projected into the image based on the camera information and compared with the 2D pose. This determines the 3D pose that best matches the 2D pose determined by the existing 2D pose approach.

[0046] However, the determination of 2D poses using a 2D pose approach explained above and known in the prior art is subject to errors. Therefore, an important aspect of the present invention is to statistically model this error for each 2D body joint point using (a small amount of) available labeled data. Gaussian mixture models, for example, can be used for this purpose.

[0047] Furthermore, the reprojection error for the adjustment procedure is determined probabilistically, for example, using a Markov Chain Monte Carlo (MCMC) approach, which allows the previously determined error model of the 2D pose estimator to be taken into account. The proposed approach thus selects the 3D pose that probabilistically matches the estimated 2D poses and also allows a confidence estimate of the thus determined 3D pose.

[0048] The computer-implemented method according to the invention thus allows for the seamless integration of automatically generated 2D pose labels and manually annotated 2D poses. The latter can be modeled using a separate error model with low variance. Extensions of the described and inventive approach with respect to a temporal component, i.e., a motion model for each 3D joint point of the 3D pose, are possible using the MCMC approach.

[0049] This also applies to the inclusion of local constraints; for example, no 3D poses should be estimated outside the passenger compartment. Other constraints, such as the length of the respective limbs, can also be modeled, e.g., by uniformly distributing them within a certain, permissible range of values. Hypotheses that lie outside this range can thus be automatically rejected.

[0050] Furthermore, it should be noted that the use of statistical methods, such as Markov Chain Monte Carlo or variational inference, can be designed as follows: A high-dimensional hypothesis space is spanned across all 3D body joint positions. By using the camera parameters assumed to be known (i.e., intrinsic and extrinsic parameters), each 3D body joint position hypothesis can be backprojected into the respective camera image. There, with the help of the previously determined 2D body joint error model, the backprojected 3D body joint position hypothesis can be evaluated for plausibility or assigned an (un)certainty. The final uncertainty value for a hypothesis results from the combination of the evaluations of all backprojections into the respective camera images.The final estimate for the 3D body joint position can be, for example, the most probable 3D body joint position hypothesis or the weighted mean of all 3D body joint positions. Uncertainty modeling can be performed analogously, for example, by considering the variance of the hypotheses.

Claims

[1] Computer-implemented method (100) for determining a 3D body joint position (20) of an object (70) comprising the following steps: - generating (102) a first 2D body joint position (12) for at least one body joint point (10) of the object (70) based on a body joint position data set (30) which contains at least one image from a plurality of camera perspectives; - Manually annotating (104) a second 2D body joint position (14) based on the body joint position data set (30) to determine an error deviation of the second 2D body joint position (14) from the generated first 2D body joint position (12); - determining (106) an error model (50) by estimating a statistical error value indicating a deviation (55) of the first 2D body joint position (12) from the second 2D body joint position data (14); - Determining (108) the 3D body joint position (20) for the at least one body joint point (10) on the basis of the first 2D body joint position (12) and the determined error model (50) by applying a statistical method. [2] The computer-implemented method (100) of claim 1, wherein determining (108) the 3D body joint position (20) further comprises the steps of: - estimating (110) at least one 3D body joint position (20) for the at least one body joint point (10) based on the first 2D body joint position (12); - Projecting (112) the at least one estimated 3D body joint position (20) into the image of the camera perspective on the basis of which the first 2D body joint position (12) was formed; - Evaluating (114) the estimated at least one 3D body joint position (20) as to how likely the estimated at least one 3D body joint position (20) is based on the previously determined error model (50). [3] The computer-implemented method (100) of claim 2, wherein the step (114) of evaluating is performed by using either the most probable 3D body joint position of all estimated 3D body joint positions (20) or a weighted average of all estimated 3D body joint positions (20). [4] Computer-implemented method (100) according to one of the preceding claims, wherein the statistical method is the Markov Chain Monte Carlo method or the variational inference method. [5] Computer-implemented method (100) according to one of the preceding claims, wherein the error model (50) is determined by means of a Gaussian mixture model. [6] Computer-implemented method (100) according to one of the preceding claims, wherein the generation of the 3D body joint position (20) is carried out by including at least one of the following information: inclusion of temporal information which represents a sequence of different camera images, information about a dimension of at least one limb for the respectively determined 2D body joint position (12, 14) and / or determined 3D body joint position (20), information about a spatial restriction for the determined 3D body position (20). [7] Computer-implemented method according to one of the preceding claims, wherein the generated 3D body joint position (20) is used to generate training data to train an ML model which is designed to automatically determine a 3D body joint position of the object 70. [8] Technical system for generating a 3D body joint position (20) according to one of the preceding claims, wherein the technical system is designed to create training data by means of the generated 3D body joint position, which training data are used as input data for training an ML model in order to automatically determine a 3D body joint position of the object 70 by means of the trained ML model. [9] Vehicle comprising a technical system according to claim 7 for controlling an automated vehicle function based on information about the generated 3D body joint information. [10] A computer program containing machine-readable instructions which, when executed on one or more computers and / or computer instances, cause the computer or computer instances to carry out the method according to any one of claims 1 to 6. [11] Machine-readable data carrier and / or download product with the computer program according to claim 10. [12] One or more computers and / or computer instances with the computer program according to claim 10, and / or with the machine-readable data carrier and / or download product according to claim 11.