Method for controlling an occupant function on the basis of a 3D body posture of one or more occupants of a motor vehicle
Patent Information
- Application Number
- PCT/EP2025/056108
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2025-03-06
- Publication Date
- 2025-10-02
AI Technical Summary
Existing methods for determining 3D body posture in motor vehicles are not suitable for real-time occupant function control due to high computational requirements or limited accuracy, making them unsuitable for interior monitoring.
A method involving a depth sensor to record a depth map, identify 2D body postures using an identification model, improve 3D support points based on a body posture model, and control occupant functions using improved 3D postures, with reduced computational intensity.
Enables real-time determination of 3D body postures for controlling vehicle occupant functions, improving accuracy and reducing computational resources.
Smart Images

Figure EP2025056108_02102025_PF_FP_ABST
Abstract
Description
[0001] Method for controlling an occupant function based on a 3D body posture of one or more occupants of a motor vehicle.
[0002] TECHNICAL BACKGROUND
[0003] The invention relates to a method for controlling an occupant function based on a 3D body posture, based on body parameters or based on a 3D body posture, based on body parameters or based on a 3D body posture and on body parameters of one or more occupants of a motor vehicle, a device for carrying out the method, a computer program product, as well as a computer-readable storage medium and a method for training machine learning models.
[0004] STATE OF THE ART
[0005] A variety of methods for determining body posture are known from the state of the art. Typically, these methods determine keypoints from images, and then determine a body posture based on these keypoints. The various methods differ in terms of the type and processing of the input signal, the way the extracted features are parameterized, and the method for inferring the 3D body posture.
[0006] The publications listed below address various subtasks of a method for determining or estimating 3D body poses. For example, in "PifPaf: Composite Fields for Human Pose Estimation," Sven Kreiss, Lorenzo Bertoni, and Alexandre Alahi describe a method based on normal 2D images for identifying or estimating a 2D body pose.
[0007] Zhou et al. in "Objects as Points", http: / / arxiv.org / abs / 1904.07850, describe a method in which a central point of an object and an object frame are determined, whereby object properties are determined based on the object frame and the central point.
[0008] Demirdjian and Varri describe in "Driver Pose Estimation with 3D Time-of-Flight Sensor.", 2009, IEEE Workshop on Computational Intelligence in Vehicles and Vehicular Systems, how, based on depth maps and an articulated model, a 3D body pose is determined using an iterative nearest neighbor method. First, a person must be extracted from the depth map using a Gaussian mixture model trained on an empty vehicle.
[0009] Kim and Kwon, "3D Human Pose Machine with a ToF Sensor using Pre-trained Convolutional Neural Networks.", International Conference on Information and Communication Technology Convergence (2019), and Rodrigues et al., "Top-Down Human Pose Estimation with Depth Images and Domain Adaptation.", 14th International Conference on Computer Vision Theory and Applications (2019), describe a method for determining a 3D body pose from a depth map using a convolutional neural network. The method first removes noise from the depth map as best as possible. An intensity image of the depth map is then input to the convolutional neural network, and the 2D body pose determined by the convolutional neural network is then converted into a 3D body pose using depth information from the depth map.
[0010] Furthermore, deformable 3D mesh models for determining body height and weight are known in the state of the art, see, for example, Pavlakos et al.: "Expressive body capture: 3d hands, face and body from a single image." 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (2019). These models provide a body volume, from which the weight can then be determined using further assumptions. These models primarily use classical optimization methods, in which a large number of model parameters must be optimized to determine body height.
[0011] There are also methods, such as the following reference methods, that use a learning-based approach to reduce runtime. However, they have significant limitations that make them unsuitable for the application of driver interior detection.
[0012] Bashirov et al., "Real-time RGBD-based Extended Body Pose Estimation," 2021 IEEE Winter Conference on Applications of Computer Vision (2021), describes a method in which a skeletal body pose is first determined in an RGBD video using the Azure Kinect Body Tracking SDK. Based on the skeletal pose, a network then determines the pose parameters of a mesh model. Body parameters are not estimated and must be determined separately beforehand.
[0013] For a medical application, Pfitzner et al., "Neural Network-based Visual Body Weight Estimation for Drug Dosage Finding," Conference: Medical Imaging (2016), describes a method in which data from an RGBD sensor is first combined with that from a thermal imaging camera, using the thermal imaging camera data for background subtraction. The resulting segmented point cloud is coded using principal component analysis, and the coded point cloud is processed using a network to estimate the weight of a recorded person. The method requires the recorded person to be in a supine position.
[0014] In Pantanowith et al.: "Estimation of Body Mass Index from photographs using deep Convolutional Neural Networks," Informatics in Medicine Unlocked 26 (2021), a neural network is used to estimate a body mass index, BMI, whereby a person is required to assume a predetermined posture before an image is taken from which the BMI is determined.
[0015] The methods known from the prior art for determining a 3D body posture meet the real-time requirements for a method for controlling an occupant function based on interior monitoring only to a very limited extent or not at all, or require such considerable computing resources that they are not transferable to interior monitoring in a motor vehicle.
[0016] SUMMARY OF THE INVENTION
[0017] The invention is based on the object of providing a method with which a 3D body posture of occupants of a motor vehicle can be determined, if possible in real time, in order to control an occupant function of the motor vehicle based on the determined 3D body posture.
[0018] A further object is to provide a system for carrying out the above method.
[0019] One aspect of the invention relates to a method for controlling an occupant function based on a 3D body posture of one or more occupants of a motor vehicle, comprising recording a depth map using a depth sensor, determining, using an identification model, for each of the occupants detected in the depth map, at least one 2D body posture comprising a plurality of 2D support points, determining one or more 3D body postures based on the depth map and the 2D support points, wherein the 3D body postures comprise 3D support points that lie on a surface defined by the depth map, in particular on the surface of a body of the respective occupant, improving, for each of the 3D body postures, the 3D support points using a body posture model based on the 3D body postures, and controlling the occupant function based on the improved 3D body posture, characterized in thatthat at least one of the improved 3D support points of the improved 3D body posture has a better match with corresponding actual body points of the respective occupant than the respective 3D support point before the improvement.
[0020] For the purposes of the present invention, an "occupant function" is a function of the motor vehicle that affects the occupant in some way. In particular, these can be comfort functions, such as climate control or seat functions; they can also be warning functions, such as issuing a signal in the event of a dangerous or unhealthy posture; or safety functions, such as belt tensioning, an airbag function, or even a braking function or other driving functions.
[0021] For the purposes of the present invention, a "2D posture" comprises a set of 2D support points and connecting axes connecting the 2D support points or a subset of the 2D support points. Various models or specifications exist for a number of support points and connecting axes. For the 2D posture, each support point is assigned a point in a two-dimensional image. When outputting the 2D posture, confidence maps can also be output instead of the pure coordinates in the image. The connecting axes can be output, in particular, as points or points parameterized with a certain order. Alternatively, the connecting axes can also be output as root points, and vectors can be output for the root points to the respective other root points of the respective connecting axes. Any other representation of such connecting axes connecting at least two points is also possible.
[0022] For the purposes of the present invention, a "3D posture" differs from a 2D posture in that the positions of the set of support points and connecting axes are not given in two-dimensional space, but in three-dimensional space.
[0023] For the purposes of the present invention, an "occupant" is a person who is in the motor vehicle.
[0024] For the purposes of the present invention, a "motor vehicle" is a passenger car, a truck or a bus.
[0025] For the purposes of the present invention, an "intensity image" is an image in which an intensity is specified for each pixel. Intensity images can be monochrome or color. In a color intensity image, the intensity is specified for each of the various color components, e.g., for the typical three color channels of the image: red, green, and blue.
[0026] For the purposes of the present invention, a "depth map" is an image, here in particular an image of the interior of a motor vehicle, in which a spatial depth, called depth information, is recorded for each pixel. A depth map can include an intensity image in addition to the depth information.
[0027] Depth maps can be captured using various depth sensors, such as time-of-flight cameras, stereo cameras, or multi-camera systems, with or without active illumination or structured light, or using learning-based methods.
[0028] For the purposes of the present invention, a "depth sensor" comprises one or more sensors with which depth maps can be recorded. As already mentioned above, these can be cameras such as time-of-flight cameras, with or without stereo cameras, multi-camera arrays, with or without active illumination, particularly with structured light. Furthermore, depth maps can also be created with cameras using learning-based methods. Depth sensors, in particular, use infrared light, but other light sources such as laser scanners are also common.
[0029] For the purposes of the present invention, an "identification model" is a processing model by means of which the occupants detected in the intensity image as well as the 2D posture of the occupants are determined and output based on an intensity image detected by the depth sensor.
[0030] For the purposes of the present invention, a "processing model" is a model configured to process input data and output output data. The processing model can be a classic model, created, for example, using classic optimization or analysis methods, or it can be a model trained using a machine learning method, also called a machine learning model.
[0031] For the purposes of the present invention, a "machine learning model" is a processing model, in particular a neural network, that can be trained using supervised or unsupervised learning to process input data and output output data or result data. For the purposes of the present invention, an "input datum" is a datum input into a processing model that is processed by the processing model. For example, the identification model uses the intensity image as input datum.
[0032] For the purposes of the present invention, an "output date" is a date output by a processing model, whereby the processing model can, in particular, output multiple output data items, in particular a result date. In addition to the result date, a processing model can also output one or more intermediate data items.
[0033] For the purposes of the present invention, a "result datum" is a datum output by a processing model, which is calculated by the processing model and typically output by processing the input datum. In a multi-layer model, a final layer of the model, the so-called output layer, outputs the result datum.
[0034] For the purposes of the present invention, a "key point," also called a "base point," is a point on an occupant's body that can be used to characterize the occupant's posture. Typically, this involves points on various joints of an occupant, as well as prominent points, especially prominent points on the face, such as the eyes, nose, mouth, and ears.
[0035] According to the present invention, a "2D vertex" is a vertex that specifies the position of a vertex, also called a keypoint, in two spatial dimensions. 2D vertices are used, particularly in images, to locate the vertices.
[0036] For the purposes of the present invention, a "posture model" is a processing model, in particular a machine learning model, that can be trained or has been trained to determine a 3D posture of an occupant.
[0037] For the purposes of the present invention, a "3D vertex" is a vertex that specifies the position of a vertex, i.e., a key point, in three spatial dimensions. For example, vertices can be specified as 3D vertices in depth maps. Classic, model-based methods demonstrate high accuracy in predicting 3D body postures. A large number of optimization parameters are determined using a classic optimization method. In addition to a sufficiently good initial pose, such optimization methods require large amounts of computing resources or are correspondingly slow, which is why they do not meet the real-time requirements for controlling motor vehicle occupant functions based on interior monitoring.The inventors therefore proposed a multi-stage process in which one or more inaccurate 3D body postures are first determined based on the respective 2D body postures and the respective depth maps. The accuracy of the determined 3D body posture is relatively low; for example, the 3D support points determined for joint points based on the depth map lie on a body surface of the respective occupant, but should be within the occupant's body volume. Subsequently, based on the inaccurate 3D body postures, suitable corrected or improved 3D body postures are determined using a trained posture model. In contrast to classic optimization models, the posture model is far less computationally intensive, which is why it also meets the real-time requirements for controlling occupant functions.
[0038] Preferably, the improvement comprises one or more of the following: correcting a distance of the 3D support points to the depth sensor, wherein the 3D support points lie on a surface of the body of the respective occupant and at least one of the improved 3D support points lies within the body of the respective occupant, and in particular the connecting axes between the corrected 3D support points are also corrected based on the corrected positions of the 3D support points, correcting inconsistent 3D support points so that at least one of the improved 3D support points has a smaller distance to the actual body points of the respective occupant, supplementing unidentified 3D support points, in particular based on symmetry information or context information, so that at least one of the improved 3D support points has a smaller distance to the actual body points of the respective occupant.
[0039] By performing well-defined processing steps during the improvement of the 3D posture in a multi-stage process, for example correcting the depth information or determining improved 3D support points for inconsistent 3D support points or unidentified 3D support points, for example based on symmetry information, the process can meet the real-time requirements for controlling occupant functions, since parameters do not have to be optimized over several iterations, as is the case with an iterative process based on optimization.
[0040] Preferably, an intensity image is input into the identification model as input data, and the identification model outputs the 2D posture comprising a plurality of 2D support points as output data. Determining the 2D posture particularly comprises determining one or more confidence maps that encode the position of the 2D support points and, in particular, connecting axes connecting some of the 2D support points in the depth map. In particular, the output data also includes a vector field that encodes directions in which the connecting axes lie with high probability. The intensity image is particularly recorded by the depth sensor system, or a separate sensor system is used to record the intensity image.
[0041] For the purposes of the present invention, a "confidence map" of a support point or connecting axis indicates the probability, also referred to as certainty or confidence, of the respective support point or connecting axis being located at the location of the confidence map. Depending on a given input datum, the confidence maps can also include depth information.
[0042] Because the identification model outputs both the 2D posture and a confidence value for each of the 2D anchor points, uncertainties in the determination of the 2D anchor points can be taken into account during further processing of the 2D posture. For example, when the improved 3D posture is determined in further processing, they can be incorporated into the determination of the uncertainty of the improved 3D posture. If this uncertainty is used to control the occupant functions, for example, using a suitable filter, such as a Kalman filter, comparatively uncertain values have only a minor influence on the control of the occupant function, making the control more reliable and improved.
[0043] Preferably, the 2D support points approximate or reflect the actual positions of joint points, eyes, nose, ears, mouth, hands, abdominal center and other prominent points of an occupant, especially in the intensity image.
[0044] Preferably, determining the 2D body postures for each occupant comprises restricting the depth map to an occupant frame that captures the 2D support points of the respective 2D body postures and corresponds to the respective occupant.
[0045] In the sense of the present invention, an "occupant frame" is a sub-area of the depth map, whereby for each occupant frame only exactly one set of support points of exactly one occupant is determined, whereby the occupant frames of different occupants can overlap.
[0046] By limiting the output 2D posture to the occupant frame, the pose model only needs to determine the 3D vertices in one part of the depth map at a time. This reduces the likelihood of misinterpreting the vertices of different occupants in the depth map. Preferably, a 3D pose is represented as image coordinates with corresponding depth values from the depth map.
[0047] By viewing a 3D body posture in the image coordinate system with depth, the 3D body posture can be determined particularly easily from the body surface points. The inventors of the present invention have recognized that in many methods for determining a 3D body posture, the support points lie on the surface of a passenger's body, which is precisely determined by the depth information in the depth map in combination with the 2D support points.
[0048] However, especially for the joints of occupants, the support points in the body and not on its surface are relevant, which is why this can be implemented particularly easily and learned by a machine learning model by a relatively simple shift in depth, in particular by shifting a distance from the origin.
[0049] Preferably, determining the 3D body posture comprises converting a parameterization as described above, as image coordinates with depth values, into a parameterization in a Euclidean camera coordinate system whose origin lies in the depth sensor, in particular in the center of the depth sensor. The transformation used is an inverse projection and provides a 3D body posture with 3D support points in a Euclidean coordinate system, in particular with reference to the depth sensor.
[0050] Because the final output of the posture model, the 3D posture, is in camera-centered coordinates, no calibration of extrinsic camera parameters is required for the application of the posture model. A transformation to a global vehicle coordinate system selected for a specific occupant function is therefore only necessary after the 3D posture has been output. Accordingly, the inference and, in particular, the training of the posture model are independent of the quality of extrinsic calibration.
[0051] Preferably, determining the 3D body posture comprises converting the depth map and the 2D body posture into a camera coordinate system by means of an inverse projection, wherein the depth sensor is located at the origin of the camera coordinate system and the 3D support points are specified in the camera coordinate system with the depth sensor at the origin.
[0052] Preferably, the method further comprises determining body parameters of the respective occupant based on the occupant frame and the depth map, wherein the body parameters are input as part of the input data into the body model for determining the 3D body posture. The body parameters particularly comprise one or more of the following parameters: an age parameter, a height parameter, a weight parameter, and / or a gender parameter.
[0053] By determining the body parameters of the occupants, the body parameters can be taken into account, especially when correcting the 3D support points, which can further improve the correction of the 3D support points.
[0054] Preferably, the determination of the body parameters is carried out using a body parameter model, wherein the body parameter model is configured in particular to determine the weight parameter and the size parameter.
[0055] By determining the body parameters using a body parameter model as a regression model instead of a classic optimization method, the real-time requirements of the 3D posture determination process are further met, while the accuracy of determining the 3D posture is further improved, since the body parameters can be used as additional input data for the posture model and contain information about the offset between the body surface points and the 3D posture. Preferably, the body parameter model and the identification model are implemented as a joint convolutional network, wherein the convolutional network determines the size parameter and the weight parameter for each pixel in the depth map, and the confidence maps output by the convolutional network include the size parameter and the weight parameter.
[0056] By implementing the body parameter model and the identification model together in a convolutional network, the required computational resources can be further reduced without losing accuracy and the processing speed can be further increased.
[0057] Preferably, a mean value and a mean deviation are determined for each of the body parameters, wherein the mean value and / or the mean deviation are filtered, in particular using a Kalman filter, and the filtered values, in particular the filtered mean value, are input into the posture model. Because the input data is filtered using a Kalman filter, body parameter values for which a large error was determined have only a minor influence on further processing by the posture model. Furthermore, the variance, in this case the variance of the determined posture, can be used as a measure of the reliability of the determined posture in controlling the occupant function.One aspect of the invention relates to a method for controlling an occupant function based on body parameters of one or more occupants of a motor vehicle, comprising recording a depth map by means of a depth sensor, determining, by means of an identification model, an occupant frame, determining, by means of a body parameter model, body parameters based on the occupant frame and the depth map, controlling the occupant function based on the body parameters.
[0058] The inventors recognized that conventional methods for determining body parameters often involve high computational effort, making them unsuitable for determining body parameters in real time. By determining body parameters using the body parameter model, accuracy can be improved and computational effort reduced, thus meeting the real-time requirements for controlling occupant functions with minimal computational effort.
[0059] Preferably, the method further comprises determining a 3D body posture according to the method for determining a 3D body posture described above, and an input datum of the body parameter model comprises the determined 3D body posture. Preferably, the body parameter model has been trained as a regression model for determining the weight parameter and the height parameter.
[0060] Because the body parameter model was trained as a regression model, a mean and variance can be calculated directly from the machine learning model's output data, particularly using a negative Gaussian log-likelihood loss. If these are input into a suitable filter, such as a Kalman filter, before further processing, the Kalman filter essentially filters out the values with high variance during further processing. In addition to its usability in a filter, the variance also provides a measure of the reliability of an output, which can be used to control occupant function.
[0061] Preferably, the body parameter model implemented as a regression model is implemented as an ensemble, wherein the ensemble comprises a plurality of identical regression models, wherein the regression models of the ensemble were initialized with different model parameters before training, in particular were trained differently or were trained with different data, and the variance of the output data of the regression models of the ensemble is interpreted as the variance of the output data.
[0062] By implementing it as an ensemble, the variance of the output data of the different regression models can be directly interpreted as the variance of the regression, which makes it easy to determine the consistency of the results from the output data.
[0063] Preferably, recording the depth map comprises processing a raw image recorded by the depth sensor and processing the raw image by one or more of the following: rectifying, converting to a target focal length, scaling and tone mapping the intensity image.
[0064] Because the depth maps input into the identification model are always rectified based on a raw image or adapted to a target focal length and / or a specific dynamic range, the identification model can process the depth maps particularly reliably and quickly, since occupants to be identified or their body posture always have a typical dimension or extent in the depth map that is expected by the model.
[0065] Preferably, the method further comprises determining body parameters of the respective occupant based on the occupant frame and the depth map, wherein the body parameters are input as part of the input data into the body model for determining the 3D body posture, and the body parameters in particular comprise one or more of the following parameters: an age parameter, a height parameter, a weight parameter and / or a gender parameter and a set of statistical shape parameters that describe the surface of a 3D model.
[0066] By determining the body parameters of the occupants, the body parameters can be taken into account, especially when correcting the 3D support points, which can further improve the correction of the 3D support points.
[0067] The occupant functions preferably include, in particular, safety functions and comfort functions of the motor vehicle, in particular a control of one or more of the following systems: a belt tensioning system, an airbag system, a comfort reclining system for a passenger.
[0068] Preferably, an output of the 3D posture model comprises mean values and mean deviations, and the mean values and mean deviations are filtered using a Kalman filter before being used to control the occupant function. By filtering the outputs of the posture model using a Kalman filter, output data of the posture model for which a large error has been determined have only a minor influence on controlling the occupant functions. A further aspect of the invention relates to an evaluation module comprising a plurality of machine learning models, in particular an identification model, a posture model, and a body parameter model, wherein the machine learning models have been trained such that they jointly execute the method described above.
[0069] A further aspect of the present invention relates to a method for training the machine learning models of the evaluation module described above.
[0070] A further aspect of the invention relates to a control device for controlling an occupant function according to one of the methods according to one of the claims.
[0071] A further aspect relates to a motor vehicle comprising a control device for controlling an occupant function according to the method described above.
[0072] A further aspect relates to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method described above.
[0073] BRIEF SUMMARY OF THE CHARACTERS
[0074] The invention will be explained in more detail below with reference to the examples shown in the drawings. The drawings show in
[0075] FIG. 1 schematically shows a system for use with the method for controlling an occupant function of a motor vehicle based on a 3D body posture according to an embodiment;
[0076] FIG. 2 schematically shows an apparatus for use with the method according to one embodiment;
[0077] FIG. 3 schematically shows a machine learning model for use with the method according to one embodiment;
[0078] FIG. 4 schematically shows a 2D body posture;
[0079] FIG. 5 schematically shows a method according to an embodiment of the present invention;
[0080] FIG. 6 schematically shows a method according to an embodiment of the present invention; and in
[0081] FIG. 7 schematically shows a method according to an embodiment of the present invention.
[0082] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0083] One embodiment of a motor vehicle 100 includes a depth sensor system 102 and a control device 200 (see FIG. 2). The depth sensor system 102 is communicatively coupled to the control device 200 (for example, with a wired or wireless communication connection). According to FIG. 1, the motor vehicle 100 also includes an interior 106, in which an occupant 104 sits according to this embodiment.
[0084] The depth sensor system 102 is configured to record depth maps of the interior 106 of the motor vehicle 100. To record the depth map 400, the depth sensor system 102 can, for example, use an infrared light source that emits short light pulses and determines a travel time for each pixel of the depth sensor system 102 that a reflected light pulse requires to travel back to the depth sensor system 102. In addition to the travel time, the depth sensor system 102 also determines an intensity of the reflected light pulse. The depth map 400 is determined based on the determined travel time. However, according to alternative embodiments, other known depth sensors can also be used.
[0085] According to the present embodiment, the recorded depth map 400 is passed to the controller 200 for further processing.
[0086] The control device 200 comprises an evaluation module 202, a memory module 204, and a control module 206. The modules can be implemented as either software or hardware modules and are connected to one another via channels 208 (see FIG. 2). The channels 208 are logical data connections between the individual modules. Alternatively, the modules can also be connected to one another via a bus system.
[0087] The storage module 204 stores the depth maps 400 recorded by the depth sensor system 102 and manages the data to be evaluated in the control device 200. The evaluation module 202 is configured to process the depth maps 400 recorded by the depth sensor system 102. In particular, the evaluation module 202 comprises one or more processing models, in particular machine learning models 300. The machine learning models 300 implemented in the evaluation module 202 are, in particular, an identification model 602 and a posture model 606. According to one embodiment of the invention, a body parameter model 604 can also be implemented in the evaluation module 202.
[0088] In accordance with the present invention, the control module 206 is configured to control the control device 200. The control module 206 reads data from the memory module 204 and forwards the data to the evaluation module 202. In addition, the control module 206 forwards the control information for controlling occupant functions to a respective control device 200 for controlling the respective occupant function. The various processing models are, in particular, machine learning models 300. A machine learning model 300 (see FIG. 3) can, in particular, be a neural network with multiple layers. In particular, the machine learning model 300 has an input layer 302, several intermediate layers 304, and an output layer 306. The input layer 302 receives an input datum 308, processes the input datum 308 using the input layer 302, the intermediate layers 304, and the output layer 306, and outputs a result datum 310.Depending on which of the various processing models is involved, the input data and the output data vary. For some machine learning models, intermediate data 312 can also be output.
[0089] The machine learning models 300 can be implemented in particular as regressors, classifiers or semantic segmenters.
[0090] For the purposes of the present invention, a "regressor," also called a regression model, is a processing model, in particular a machine learning model, that performs a regression. If the regressor is a machine learning model, it is trained to perform the regression, in particular through supervised learning or unsupervised learning. For the purposes of the present invention, a "classifier," also called a classification model, is a processing model that assigns a class to an input datum or can be trained to assign a class to the input datum. The result datum can, in particular, be a class assigned to the input datum. The format can, in particular, be a vector, with each entry of the vector corresponding precisely to one of the possible classes to be assigned, and in particular, a "T" entry in the vector indicates the class of the input datum. Alternatively, a class number can be output.As a further alternative, the classifier can also be trained such that the result date is a vector, where each entry in the vector indicates a probability that the respective input date belongs to the class corresponding to the result date entry. Depending on the implementation, the respective format of the result date of classifiers varies, and accordingly, the format of the target date in an annotated dataset for training the classifier also varies.
[0091] For the purposes of the present invention, a "semantic segmenter," also called a semantic segmentation model, is a processing model that assigns an output value to each entry of an input datum. A result datum is also referred to as a segmentation mask or semantic segmentation mask. If the input datum is an image, the semantic segmenter performs an image-to-image mapping, with which each pixel of an input datum is assigned an output value according to a semantics.
[0092] For the purposes of the present invention, "supervised learning" is a learning process in which a machine learning model is trained to perform a desired mapping using an annotated dataset.
[0093] In the sense of the present invention, "unsupervised learning" is a learning process in which a machine learning model is trained without specifying a desired goal, solely on the basis of an unannotated data set, whereby the machine learning model automatically finds or should find certain cluster points in the data set.
[0094] For the purposes of the present invention, an "annotated data set" comprises input data and target data, with each input data item corresponding to an annotation or label, called a target data item, of the target data. The target data is typically generated through complex processing or by manual marking or similar. Based on the target data, a processing model is trained to execute a desired mapping. During training, a result data item is approximated to a target data item over several training rounds using the respective training method. For the purposes of the present invention, a "target data item" is a data item used when training a processing model to execute a processing mapping, to which a result data item output by the processing model, based on the input data item, is to be adapted. The approximation is performed using a target function.For the purposes of the present invention, an "objective function" is in particular a profit function or a loss function that specifies how differences between the result date of the processing model and the target date are evaluated.
[0095] For the purposes of the present invention, a "loss function" is a function that captures differences between the result date and the specified target date. If the result date and the target date are images, for example, the comparison can be made pixel-by-pixel. If the result date and the target date are vectors or tensors, for example, the difference can be made entry-by-entry. The differences can be added absolute values in an L1 loss function. The square sum of the differences is calculated in an L2 loss function. To minimize the loss function, the values of model parameters of the processing model are changed, which can be calculated, for example, using gradient descent and backpropagation. Other possible loss functions include, in particular, a cross-entropy loss, a hinge loss, a logistic loss, a log-likelihood loss, a Gaussian negative log-likelihood loss, or a Kullback-Leibler loss.Especially when training regressors, Gaussian negative log-likelihood loss is often used as a loss function. When using Gaussian negative log-likelihood loss, the machine learning model trained with the loss function is trained to output a variance in addition to an output mean. Alternatively, a standard deviation or the logarithm of the standard deviation can also be output.
[0096] In particular, when training the posture model 606, a backprojection loss function can be used in which a 3D posture is projected using a camera model so that a loss function can be calculated in two-dimensional coordinates.
[0097] For the purposes of the present invention, a "gain function" is an objective function, wherein the gain function, in contrast to the loss function, which captures a difference between the result date and the target date and is minimized during training, captures a match which is maximized during training.
[0098] A method for controlling an occupant function using the above-described motor vehicle 100 is described below with reference to FIGS. 4 to 6. In the described method, in a step 502 (see FIG. 5), a depth map 400 (see FIG. 4) is recorded using the depth sensor system 102 and stored in the memory module 204 for further processing. According to one embodiment of this embodiment, the depth sensor system 102 can also forward the depth map 400 directly to the evaluation module 202 for evaluation; intermediate storage is then unnecessary.
[0099] According to one configuration of this embodiment, recording the depth map 400 comprises a processing step in which the raw image recorded by the depth sensor 102, also called the raw depth map, is processed using conventional image processing methods. In particular, the raw image is rectified, for example, the geometry in the raw image is converted based on a target focal length; alternatively or additionally, tone mapping, also called dynamic compression, can also be applied to the intensity values of the raw image. Rectifying the raw image ensures that, even when using different optics in the depth sensor 102, a rectified image is always input into one of the machine learning models 300, thereby improving the quality of the result data 310 of the respective machine learning model 300.Similarly, converting to a desired target focal length also improves the quality of the result datum 310, since, for example, a distance and, based on the distance, an extension of a body of an occupant 104 can be determined with corresponding precision, thereby improving, for example, the determination of body parameters. The same applies to tone mapping. Because the dynamic range of an image is always normalized according to an expected input, the accuracy and quality of the result datum 310 for all of the machine learning models 300 can also be improved.
[0100] According to a subsequent step 504, an identification model 602 implemented in the evaluation module 202 (see FIG. 6) determines in the depth map 400 for each occupant 104 detected in the depth map 400 a 2D posture, comprising 2D anchor points 402 (see FIG. 4) based on the intensity image of the depth map 400, and a vector field for each connecting axis that indicates in which direction and with which probability the connecting axis points at the respective position of the vector field. According to this embodiment, the result datum 310 of the identification model 602 includes a confidence for each of the anchor points 402 that indicates, for the respective anchor point 402, how high the probability is that the respective anchor point 402 is located there.
[0101] According to one embodiment of this specific embodiment, the identification model 602 can also be configured to determine an occupant frame 406 based on the determined 2D support points 402. The occupant frame 406 encompasses the 2D support points 402 corresponding to an occupant 104. The 2D support points 402 of multiple occupants 104 can overlap on a depth map 400; an occupant frame 406 output by an identification model 602 always only includes support points 402 of one occupant 104 and the corresponding associated connecting axes 404. In particular, according to the embodiment, the identification model 602 outputs a section of the depth map 400 corresponding to a respective occupant frame 406, together with the 2D support points 402 corresponding to the occupant frame 406.
[0102] According to this embodiment, the occupant frame 406 is a so-called bounding box, as determined using common neural networks. Bounding boxes are referred to when they are, in particular, a rectangle enclosing an object.
[0103] According to one refinement of this embodiment, a segmentation mask can also be output instead of the bounding box. In the segmentation mask, each pixel is assigned a value, which is used to identify an occupant 104. For example, all pixels corresponding to a driver are assigned a "1." All pixels corresponding to a passenger are assigned a "2." This continues according to the number of seats in the motor vehicle 100. All pixels that do not belong to an occupant can be assigned a "0." The selected values are, of course, only exemplary. Analogous to the confidence maps for the reference points 402, confidence maps for the various possible occupants 104 in the motor vehicle 100 could also be output.
[0104] Such support points 402 are schematically shown in the depth map 400 in FIG. 4 for occupants 104 of a motor vehicle 100. FIG. 4 schematically shows the intensity image of the depth map 400. According to this embodiment, the support points 402 are precisely the positions of joints of an occupant 104, in particular the shoulder joints, the hip joints, the elbow joints, and the wrists in the depth map 400. Further support points 402 can, in particular, be prominent points on the face, such as the nose, ears, eyes, and mouth.
[0105] In the intensity image, the support points 402 identified by the identification model 602 are already drawn as round circles, whereby the circles in the face have been chosen to be smaller here, particularly for the sake of clarity. The support points 402 are connected to one another by connecting axes 404. Two occupants 104 can be seen in the depth map 400. The identification model 602 has also determined occupant frames 406 for each of the support points 402. The depicted body postures show a slightly forward-bent posture of the upper body, particularly for the left occupant 104. Furthermore, the knees are only poorly visible in the image section, and the feet of the occupants 104 are practically invisible. The extent to which different extremities are detected by a depth sensor 102 depends on the geometry and localization of the depth sensor 102. In particular, multiple depth sensors 102 can be used to detect all extremities.In addition, a different optics can be used in the depth sensor 102 with, for example, a wider or narrower opening angle.
[0106] Other body postures may differ considerably from those shown here, particularly in the position of the extremities. Further forward or backward bending may also be detected, as may rotations. For example, an occupant 104, in particular a front passenger, may have assumed a lying position on their seat. This also applies in particular to occupants 104 in the rear rows of seats. The occupants 104 in the rear rows of seats in the interior 106 of the motor vehicle 100 can be recorded, in particular, using the same depth sensor system 102. According to one embodiment of this embodiment, however, multiple depth sensors 102 can also be installed in the motor vehicle 100, so that occupants 104 in the rear rows of seats can also be reliably detected by a depth sensor system 102 in a depth map 400.According to one embodiment of this embodiment, the identification model 602 outputs only the positions of the support points 402 in the intensity image instead of the confidence maps. The identification model 602 is either trained directly to output the support points 402, or the identification model 602 also calculates confidence maps, and in a final step, the support points 402 are then determined from the confidence maps. According to a further step 506, a 3D body pose is determined from the depth map 400 and the 2D support points 402. For this purpose, the 2D support points 402 are first supplemented by the depth information of the depth map 400, resulting in 3D support points 402. In this case, the 3D support points comprise the 2D support points output by the identification model 602, which now also have a further spatial dimension for the points of the depth map 400, here the depth information of the corresponding point of the intensity image orwhich includes depth information from the depth map 400. The 3D support points 402 form a 3D body posture.
[0107] According to one configuration of this embodiment, the 2D vertices are converted into 3D vertices based on the depth values at the corresponding position in the depth map 400 using the inverse projection matrix of the calibrated depth camera 102 before applying the pose model 606. This step reverses the projection so that the 3D vertices are specified in a Euclidean 3D coordinate system whose origin lies at the camera center of the depth camera 102. The generated 3D vertices form the 3D body posture. Training a machine learning model to correct the 3D vertices in Euclidean coordinates facilitates, for example, learning the inherent symmetry between the left and right halves of the body.
[0108] According to a step 508, the posture model 606, based on the 3D posture, improves the 3D posture and outputs an improved 3D posture to the control module 206. Improving the 3D posture can comprise one or more different improvements. One possible improvement comprises determining improved depth information, wherein the posture model 606, in particular for the 3D support points corresponding to the joint positions of the occupants, corrects the position of the 3D support points, in particular with respect to a depth, such that they no longer lie on a body surface of the respective occupant given by the depth map 400, but rather within the body of the respective occupant. Another alternative improvement comprises correcting inconsistent 3D support points.For example, individual support points 402 in the depth map 400 may be obscured, so that the depth values belong to the obscuring object and not to the surface of the body at the location of the support point 402. Furthermore, in the event of an occupancy, the identification model 602 may output one or more 2D support points that do not correspond at all to the actual position of the respective support point 402 on the occupant. The posture model 606 can be trained to identify such inconsistent support points 402 and, for example, to estimate a better position based on a symmetry of the occupant's body. In particular, it may also be the case that the identification model 602 does not identify certain support points 402 at all; these support points 402 can then also be estimated by the posture model 606. For example, certain support points 402 may lie outside the image captured by the depth sensor 102.Even for such cases, the posture model 606 can estimate vertices 402 whose position is closer to the actual position. This latter improvement is also called supplementing unidentified 3D vertices. For unidentified 3D vertices, vertices 402 are output exclusively with "O" entries. When supplementing unidentified 3D vertices, the "O" entries are replaced by 3D vertices estimated by the posture model 606. Typically, the 3D vertices exhibit a significantly better match to the actual body points than the "O" entries. The simultaneous learning of the described improvements makes it possible to generate a posture model that is robust against all corresponding error sources.
[0109] According to this embodiment, vertices 402 are input into the pose model 606 as input data 308, which significantly reduces the complexity of the processing map to be learned, since no complex image data needs to be processed by the pose model 606, and the pose model 606 thus does not need to learn complex semantics in the image data. This significantly improves real-time processing capability. Furthermore, using vertices 402 as input data 308 also significantly simplifies the generation of synthetic data for training, since complex rendering of image data as input data 308 is not necessary.
[0110] According to one embodiment of the first embodiment, several such 3D body postures can be generated from the 2D support points and the confidence maps and used as an input signal for the posture model 606. Based on the confidence value at the respective 2D support point, a selection method, e.g., random selection from a Gaussian distribution, is controlled, which selects several 2D support points with associated depth values from the environment of the original 2D support point. By inversely projecting the 2D support points with depth values, several corresponding 3D support points are generated, which in turn form several 3D body postures. By selecting the depth values and the 2D support points from the environment of the original support point, the distribution of the depth values and the reliability of the identification model 602 are taken into account. The distribution of several 3D body postures facilitates the learning of the necessary correction by the posture model 606.According to a further refinement of the first embodiment, confidence values output by the identification model 602 can be directly calculated with the 3D support points, for example, by multiplying the confidence values at the 2D support points by the corresponding intermediate or final results of the posture model 606, such as correction values or variances. This can improve the learning and accuracy of the 3D posture and the associated 3D posture uncertainty.
[0111] According to another alternative, the posture model 606 can process the confidence maps directly. The posture model 606 then outputs the 3D posture and its variance as output data.
[0112] According to a step 510, the evaluation module 202 forwards the improved 3D posture to the control module 206, which then controls an occupant function based on the improved 3D posture. The occupant function can include, in particular, safety functions or comfort functions. In particular, the safety functions can include belt tensioner and airbag functions. For example, a 3D posture can indicate that an occupant 104 has assumed a forward-leaning 3D posture. If the control device 200, in particular the control module 206, determines that the vehicle is highly likely to perform a braking maneuver, in particular a severe braking maneuver, then the control device 200 can control the belt tensioner such that the occupant 104 is pulled or forced into an upright posture by the belt tensioner.In the event of an unavoidable accident situation detected by the control module 206, the airbag control can, for example, be adjusted depending on the distance of the head and torso of the occupant 104 from the steering wheel and the weight of the occupant 104, thereby reducing the risk of injury. For example, with the comfort functions, an air conditioning function can be appropriately adjusted, for example, to a seating position. For example, the outlets used for the air conditioning function can be coordinated with the 3D body posture. For example, if an occupant 104 is in a more reclining position, different air outlets of an air conditioning system can be used than if the respective occupant 104 is sitting upright.
[0113] Warning functions include, for example, detecting dangerous body positions, such as reaching backwards or to the glove compartment, or generally detecting that one or both hands are not touching the steering wheel.
[0114] In addition to the 3D body posture determined by the posture model 606, occupant information is also output to the control module 206. The occupant information here includes, in particular, information about a seat, so that an occupant function can be executed that is appropriately tailored to the identified seat in the motor vehicle 100, in particular, for example, the deployment of the airbag or belt tensioner only for the respective occupant 104.
[0115] According to one embodiment of this embodiment, the method may further comprise a step 507 for determining body parameters using a body parameter model 604 (see FIG. 6). The occupant frame 406 determined by the identification model 602 is input into the body parameter model 604 as input data 308. The occupant frame 406 is the sub-region of the depth map 400 corresponding to an occupant, which encompasses the support points 402 of an occupant 104. The occupant frame 406 is output by the identification model 602. Based on the occupant frame 406, the posture model 606 determines the body parameters. The body parameters include, in particular, the height and weight of the respective occupant 104.
[0116] The result datum 310 of the body parameter model 604 here includes the height and weight of the respective occupant 104. In particular, the body parameter model 604 can also output a gender and an age. The body parameters output by the body parameter model 604, in addition to the confidence maps or the 2D support points, are input into the posture model 606 as part of the input datum 308. In particular, the evaluation module 202 combines the 2D support points and the body parameters into an input datum 308 for the posture model 606.
[0117] The inventors have recognized that the quality of the correction of the 3D support points directly depends on body parameters such as the weight and height of the occupants 104, which is why the posture model 606 receives the body parameters height and weight as part of the input data 308 in addition to the 2D support points. Additionally, the gender and age can also be estimated by the body parameter model 604, or the shape parameters of a statistical model, which specify the volume and weight distribution of all segments. These also influence the correction performed by the posture model 606.
[0118] According to one configuration of this embodiment, the result data 310 of the machine learning models 300 can each be filtered through a Kalman filter before further processing. The Kalman filter has the advantage that, in addition to an input measured value, an estimated error of the measured value is also input into the filter. Measured values with a large estimated error are evaluated in the Kalman filter such that they have a smaller influence on the filter output.
[0119] A second embodiment is described below with reference to FIG. 7. The second embodiment differs from the first embodiment in that the step of controlling the occupant function is based on the body parameters. According to a step 702, an image is captured with a depth sensor 102, as in step 502. According to a subsequent step 704, an occupant frame 406 is determined. In contrast to the first embodiment, the identification model 602 according to the second embodiment only determines the occupant frame 406; an output of a 2D body posture comprising 2D support points is omitted in this embodiment. Accordingly, the identification model 602 of the second embodiment is somewhat simpler and only needs to identify the occupants detected in the depth map 400 and output the corresponding occupant frame 406.In particular, in this embodiment, the identification model 602 is designed as a detector that outputs a list of the found occupant frames 406.
[0120] According to a further step 706, the second embodiment comprises determining, by means of the body parameter model 604, body parameters based on the occupant frame 406. The body parameter model 604 corresponds to the body parameter model 604 of the first embodiment.
[0121] According to one aspect of this embodiment, a posture model 606 provides additional input data 308 for the body parameter model 604 to improve accuracy. In this embodiment, for example, the distances between 3D vertices can be used to improve the accuracy of body parameters such as height and weight.
[0122] According to a step 708, the occupant functions are controlled as already described above with reference to the first embodiment, but this time based on the body parameters.
[0123] According to the embodiment described above, as well as its modifications and refinements, the machine learning models 300 can be, in particular, regressors, classifiers, or even semantic segmenters. Accordingly, the respective input data 308, the result data 310, and also the target data during training differ when supervised learning is used.
[0124] If the machine learning models 300 are trained as classifiers, one of the following functions or a combination of cross-entropy loss, hinge loss, exponential loss, square loss, savage loss, tangent loss and Gaussian loss is used as the loss function.
[0125] When training classifiers, the target data used is primarily the classes to be recognized. Depending on the choice of the loss function or output layer 306, a hard or soft assignment to one or more classes can be made.
[0126] If the machine learning models are 300 semantic segmenters, the outputs are in particular segmentation masks.
[0127] According to one embodiment, the depth map 400 comprises the intensity image, and the depth sensor 102 records the depth map 400 and the intensity image simultaneously. According to an alternative, the intensity image can be recorded independently of the depth map 400. In particular, in addition to the depth sensor 102, a camera for recording the intensity image can be provided. Preferably, the camera and the depth sensor 102 are arranged at the same location. According to an alternative, the camera and the depth sensor 102 can be arranged at different locations; in this case, the intensity image and the depth map 400 must be suitably transformed so that they are approximately registered to one another, i.e., the same pixels detect the same objects. The transformation can also be carried out for sub-regions of the depth map 400 and the intensity image.
[0128] According to a further embodiment, the present invention also comprises a device for carrying out the method described above. In particular, a control device 200. The control device 200 can, in particular, be an ECU, an electronic control unit.
[0129] According to a further embodiment, the present invention provides a computer program product comprising instructions that cause a computer to carry out the method described above.
[0130] According to a further embodiment, the present invention provides a method for training the machine learning models 300 described above, i.e., the identification model 602, the body parameter model 604, and the posture model 606.
[0131] According to a further embodiment, the present invention provides a motor vehicle 100 with a device for carrying out the method described above, in particular the control device 200. LIST OF REFERENCE SYMBOLS
[0132] 100 motor vehicles
[0133] 102 Depth sensors
[0134] 104 inmates
[0135] 106 Interior
[0136] 200 control device
[0137] 202 Evaluation module
[0138] 204 memory module
[0139] 206 Control module
[0140] 208 channels
[0141] 300 machine learning model
[0142] 302 Input layer
[0143] 304 Intermediate layer
[0144] 306 Output layer
[0145] 308 Entry date
[0146] 310 Result date
[0147] 312 Interim date
[0148] 400 depth map
[0149] 402 base
[0150] 404 connecting axis
[0151] 406 passenger frame
[0152] 602 Identification model
[0153] 604 Body parameter model
[0154] 606 Posture Model
Claims
CLAIMS 1. A method for controlling an occupant function based on a 3D posture of one or more occupants of a motor vehicle (100), comprising: Recording (502) a depth map (400) by means of a depth sensor (102), Determining (504), by means of an identification model (602), for each of the occupants detected in the depth map (400), at least one 2D body posture comprising a plurality of 2D support points, Determining (506) one or more 3D body postures based on the depth map (400) and the 2D support points, wherein the 3D body postures comprise 3D support points lying on a surface given by the depth map (400), in particular on the surface of a body of the respective occupant, Improving (508), for each of the 3D body postures, the 3D support points using a body posture model (606) based on the 3D body postures, and Controlling (510) the occupant function based on the improved 3D body posture, characterized in that at least one of the improved 3D support points of the improved 3D body posture has a better match with respective corresponding actual body points of the respective occupant than the respective 3D support point (402) before the improvement (508).
2. The method according to claim 1, wherein the improving (508) comprises one or more of the following: Correcting a distance of the 3D support points to the depth sensor (102), wherein the 3D support points lie on a surface of the body of the respective occupant and at least one of the improved 3D support points lies within the body of the respective occupant, and in particular the connecting axes between the corrected 3D support points are also corrected based on the corrected positions of the 3D support points, Correcting inconsistent 3D support points so that at least one of the improved 3D support points has a smaller distance to the actual body points of the respective occupant, Adding unidentified 3D vertices, in particular based on symmetry information or context information, so that at least one of the improved 3D support points better match the actual body points of the respective occupant.
3. The method according to claim 1 or 2, wherein an intensity image is input into the identification model (602) as input data (308) and the identification model (602) outputs the 2D body posture comprising a plurality of 2D support points as output data, wherein determining the 2D body posture comprises, in particular, determining one or more confidence maps that encode the position of the 2D support points and, in particular, connecting axes connecting some of the 2D support points in the depth map (400), in particular the output data also comprises a vector field that encodes directions in which the connecting axes lie with a high probability, wherein the intensity image was recorded in particular by the depth sensor system (102) or a separate sensor system is used to record the intensity image.
4. The method according to any one of claims 1 to 3, wherein the 2D support points approximate actual positions of joint points, eyes, nose, ears, mouth, chin, hands, abdominal center or other prominent points of an occupant.
5. The method according to any one of claims 1 to 4, wherein determining the 2D body postures for each occupant comprises restricting the depth map (400) to an occupant frame (406) enclosing the 2D support points of the respective 2D body postures and corresponding to the respective occupant.
6. The method according to one of claims 1 to 5, wherein determining the 3D body posture comprises converting the depth map (400) and the 2D body posture into a camera coordinate system by means of an inverse projection, the depth sensor (102) is located at the origin of the camera coordinate system, and the 3D support points are given in the camera coordinate system with the depth sensor (102) at the origin.
7. The method according to any one of claims 1 to 6, wherein an output of the 3D body posture comprises respective mean values and mean deviations, and the mean values and mean deviations are filtered by means of a Kalman filter before being used to control the occupant function.
8. The method according to any one of claims 1 to 7, wherein the method further comprises determining (507) body parameters of the respective occupant based on the occupant frame and the depth map, and the body parameters are input as part of the input data into the body model for determining the 3D body posture.
9. The method according to claim 8, wherein the determination of the body parameters is carried out using a body parameter model, wherein the body parameter model is configured in particular to determine the weight parameter and the size parameter.
10. The method according to claim 9, wherein the body parameter model and the identification model are implemented as a joint convolutional network, wherein the convolutional network determines the size parameter and the weight parameter for each pixel in the depth map, and the 2D body pose output by the convolutional network comprises the size parameter and the weight parameter.
11. Method according to one of claims 8 to 10, wherein a mean value and a mean deviation are determined for each of the body parameters, wherein the mean value and the mean deviation are filtered in particular by means of a Kalman filter and the filtered values are input in particular into the body posture model.
12. A method for controlling (510) an occupant function based on body parameters of one or more occupants of a motor vehicle, comprising: Recording (502) a depth map (400) by means of a depth sensor (102), determining (504) an occupant frame by means of an identification model, determining (506) body parameters based on the occupant frame (406) and the depth map (400) by means of a body parameter model, Controlling (510) the occupant function based on the body parameters.
13. The method of claim 12, wherein the method further comprises determining (506) a 3D body posture according to the method of any one of claims 1 to 7, and an input datum (308) of the body parameter model (604) comprises the 3D body posture.
14. The method according to any one of claims 9 to 12, wherein the body parameter model (604) has been trained as a regression model for determining the weight parameter and the size parameter.
15. The method according to claim 14, wherein the body parameter model (604) implemented as a regression model is implemented as an ensemble, the ensemble comprising a plurality of identical regression models, the regression models of the ensemble being initialized with different model parameters before training, and the variance of the output data of the regression models of the ensemble being interpreted as the variance of the output data.
16. The method according to any one of claims 8 to 15, wherein the body parameters comprise in particular one or more of the following parameters: an age parameter, a height parameter, a weight parameter and / or a gender parameter.
17. The method according to any one of claims 1 to 16, wherein the recording (502) of the depth map (400) comprises processing a raw image recorded by the depth sensor (102), and the processing of the raw image comprises one or more of equalization, conversion to a target focal length, scaling, and tone mapping.
18. The method according to one of claims 1 to 17, wherein the occupant functions comprise, in particular, safety functions and comfort functions of the motor vehicle (100), in particular a control of one or more of the following systems: a belt tensioning system, an airbag system, a comfort reclining system for a passenger, a warning system.
19. The method according to any one of claims 1 to 18, wherein the depth map (400) assigns depth information to each point in a captured scene.
20. A method for training the machine learning models (300) to carry out the method according to one of claims 1 to 19.
21. Evaluation module (202) comprising a plurality of machine learning models (300), in particular an identification model (602), a posture model (606) and a body parameter model (604), wherein the machine learning models (300) have been trained such that they jointly execute the method according to one of the preceding claims 1 to 20.
22. Control device (200) for controlling an occupant function according to one of the methods according to one of claims 1 to 19.
23. Motor vehicle (100), comprising a control device (200) for controlling an occupant function according to one of the methods according to one of claims 1 to 19.
24. A computer program product comprising instructions which, when the program is executed by one or more computers, cause them to carry out the method according to any one of the preceding claims 1 to 20.