Method for controlling occupant functions based on 3D body pose of one or more occupants of a motor vehicle
Patent Information
- Application Number
- CN202580017940.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2025-03-06
- Publication Date
- 2026-09-29
AI Technical Summary
[0012]从现有技术中已知的用于确定3D身体姿态的方法仅非常有限地满足直到根本不满足用于基于内部空间监控来控制乘员功能的方法的实时要求,或者为此需要如此显著的计算资源,以至于所述计算资源不能转移到机动车中的内部空间监控上
[0057]发明人已认识到,常见的用于确定身体参数的方法通常与高计算开销相关联,因此所述方法不适合实时确定身体参数。通过借助身体参数模型来确定身体参数,可以改善精度,此外减少了计算开销,因此在仅低的计算开销的情况下也满足用于控制乘员功能的实时要求。
Smart Images

Figure CN122847728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for controlling occupant functions based on 3D body posture, body parameters, or 3D body posture and body parameters of one or more occupants of a motor vehicle; an apparatus for performing the method; a computer program product; a computer-readable storage medium; and a method for training a machine learning model. Background Technology
[0002] Various methods for determining body pose are known in the prior art. These methods typically involve identifying support points, also known as keypoints, from an image and determining the body pose based on these identified support points. Different methods vary in the type and processing of the input signal, the type of parameterization of the extracted features, and the reasoning method for 3D body pose.
[0003] The publications listed below partially address different subtasks of methods for determining or estimating 3D body pose. For example, Sven Kreiss, Lorenzo Bertoni, and Alexandre Alahi describe a method for identifying or determining 2D body pose based on ordinary 2D images in “PifPaf: CompositeFields for Human Pose Estimation”.
[0004] Zhou et al. described a method in “Objects as Points” (http: / / arxiv.orq / abs / 1904.07850) in which the center point and object bounding box of an object are determined, and the object attributes are determined using the object bounding box and the center point.
[0005] In “Driver Pose Estimation with 3D Time-of-Flight Sensor” (2009, IEEE Workshop on Computational Intelligence for Vehicles and Vehicle Systems), Demirdjian and Varri described how to determine 3D body pose based on depth maps and joint models using an iterative nearest neighbor method. First, the person must be extracted from the depth map, for which a Gaussian mixture model trained with the vehicle in an unloaded state is used.
[0006] Kim and Kwon's "3D Human Pose Machine with a ToF Sensor using Pre-trained Convolutional Neural Networks" (International Conference on Convergence of Information and Communication Technologies, 2019) and Rodrigues et al.'s "Top-Down Human Pose Estimation with Depth Images and Domain Adaptation" (14th International Conference on Theory and Applications of Computer Vision, 2019) describe methods for determining 3D body pose from depth maps using convolutional networks. In these methods, interference effects are first removed from the depth map as much as possible. Then, the intensity image of the depth map is input into a convolutional network, where the 2D body pose determined by the convolutional network is subsequently converted into a 3D body pose using the depth information from the depth map.
[0007] Furthermore, deformable 3D mesh models are known in existing technologies for determining height and weight, see, for example, Pavlakos et al., “Expressive body capture: 3D hands, face and body from a single image” (IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2019). These models provide body volume, from which weight can then be determined using other assumptions. These models primarily employ classical optimization methods, where a large number of model parameters must be optimized to determine height.
[0008] There are also methods, such as the following reference method, that follow a learning-based approach to reduce runtime. However, these methods come with significant limitations, making them unsuitable for applications involving driver interior space detection.
[0009] Bashirov et al.'s "Real-time RGBD-based Extended Body Pose Estimation" (IEEE Winter Conference on Applications of Computer Vision 2021) describes a method in which the skeletal body pose is first determined in an RGBD video using the Azure Kinect Body Tracking SDK, and then the network determines the pose parameters of the mesh model based on the skeletal pose. The body parameters are not estimated and must be determined separately beforehand.
[0010] For medical applications, Pfitzner et al.'s paper, "Neural Network-based Visual Body Weight Estimation for Drug Dosage Finding" (Conference: Medical Imaging, 2016), describes a method in which data from an RGBD sensor is first combined with data from a thermal imaging camera, where the thermal imaging data is used for background subtraction. The resulting segmented point cloud is encoded using principal component analysis, and the encoded point cloud is processed by a network to estimate the weight of the person being photographed. This method requires the person being photographed to be in a lying position.
[0011] Pantanowith et al.'s paper, "Estimation of Body Mass Index from photographs using deep Convolutional Neural Networks" (Unlocking Medical Informatics, Vol. 26, 2021), uses neural networks to estimate body mass index (BMI). In this paper, the person being examined must assume a predetermined body posture before the images are taken, and the BMI is determined using the images.
[0012] Methods known from the prior art for determining 3D body poses satisfy, to a very limited extent, the real-time requirements of methods for controlling occupant functions based on interior space monitoring, or require such significant computational resources that these resources cannot be transferred to interior space monitoring in motor vehicles. Summary of the Invention
[0013] The basic objective of this invention is to provide a method that can determine the 3D body posture of a motor vehicle occupant in real time as much as possible, so as to control the occupant functions of the motor vehicle based on the determined 3D body posture.
[0014] Another task is to provide a system for performing the above methods.
[0015] One aspect of the invention relates to a method for controlling occupant functions based on the 3D body pose of one or more occupants of a motor vehicle, comprising: receiving a depth map using a depth sensing mechanism; determining at least one 2D body pose including a plurality of 2D support points for each occupant detected in the depth map using a recognition model; determining one or more 3D body poses based on the depth map and the 2D support points, wherein the 3D body pose includes 3D support points located on a surface given by the depth map, particularly on the surface of the corresponding occupant's body; for each of the 3D body poses, improving the 3D support points based on the 3D body pose using a body pose model; and controlling the occupant functions based on the improved 3D body pose, characterized in that at least one of the improved 3D support points of the improved 3D body pose has better consistency with the respective corresponding actual body points of the corresponding occupant compared to the corresponding 3D support points before the improvement.
[0016] In the context of this invention, "occupant function" refers to a function of a motor vehicle that relates to occupants in some way. This can be, in particular, comfort functions, such as air conditioning or seat functions, as well as warning functions, such as signaling in dangerous or unhealthy body postures, or safety functions, such as seat belt pretensioners or airbag functions, but can also be braking or other driving functions.
[0017] In the context of this invention, a "2D body pose" includes a set of 2D support points and connecting axes that connect the 2D support points or subsets of the 2D support points to each other. Different models or specifications exist for a given number of support points and connecting axes. For a 2D body pose, each support point is assigned exactly one point in the two-dimensional image. However, when outputting the 2D body pose, a confidence map can be output instead of pure coordinates from the image. Connecting axes can be output as points or points parameterized in a certain order; alternatively, connecting axes can also be output as root points and vectors from the root points to corresponding other root points on the corresponding connecting axes. Any other representation of such connecting axes that connect at least two points to each other is also possible.
[0018] In the context of this invention, the difference between "3D body pose" and "2D body pose" is that the positions of the set of support points and connecting axes are given in three-dimensional space, not in two-dimensional space.
[0019] In the context of this invention, "occupant" refers to a person inside a motor vehicle.
[0020] In the context of this invention, "motor vehicle" refers to a passenger car, a freight car, or a public bus.
[0021] In the context of this invention, an "intensity image" is an image that provides the intensity for each pixel. An intensity image can be monochrome or color. In a color intensity image, the intensity is given for different color components, such as the typical three color channels of an image: red, green, and blue.
[0022] In the context of this invention, a "depth map" is an image, specifically an image of the interior space of a motor vehicle, in which spatial depth, referred to as depth information, is detected for each pixel. An intensity image may be included as an adjunct to the depth information.
[0023] Depth maps can be detected using different depth sensing mechanisms, such as time-of-flight cameras, stereo cameras, or combinations of multiple cameras, with or without active illumination or structured light, or using learning-based methods.
[0024] In the context of this invention, a "depth sensing mechanism" includes one or more sensors that can be used to record depth maps. As mentioned above, this can in particular be a camera, such as time-of-flight cameras, with or without a stereo camera, a multi-camera combination, with or without active illumination, especially with structured light. Furthermore, depth maps can also be constructed using cameras with the aid of learning-based methods. In particular, depth sensing mechanisms use infrared light, but other light sources such as laser scanners are also common.
[0025] In the context of this invention, the "recognition model" is a processing model that, with the aid of an intensity image detected by a depth sensing mechanism, determines and outputs the occupant detected in the intensity image and the occupant's 2D body pose.
[0026] In the context of this invention, a "processing model" is a model configured to process input data and output data. A processing model can be, for example, a classical model constructed using classical optimization or analysis methods; similarly, a processing model can also be a model learned using machine learning methods, also known as a machine learning model.
[0027] In the context of this invention, a "machine learning model" is a processing model, especially a neural network, which can be trained by supervised or unsupervised learning to process input data and output data or result data.
[0028] In the context of this invention, "input data" refers to data that is input into and processed by the processing model. For example, a recognition model might use an intensity image as input data.
[0029] In the context of this invention, "output data" refers to data output by the processing model, which in particular can output multiple output data, especially result data. In addition to result data, the processing model can also output one or more intermediate data.
[0030] In the context of this invention, "result data" refers to data output by a processing model, which is typically calculated by the processing model from input data and then output. In a multi-layered model, the last layer, the so-called output layer, outputs the result data.
[0031] In the context of this invention, the "support point," also known as a key point, is a point on the occupant's body by means of which the occupant's body posture can be characterized. Typically, points of different joints of the occupant are used for this purpose, as well as particularly prominent points, especially prominent points on the face, such as the eyes, nose, mouth, and ears.
[0032] According to the present invention, a "2D support point" is a support point that gives the location of a support point (also called a key point) in two spatial dimensions. In particular, 2D support points are used to locate support points in images.
[0033] In the context of this invention, a “body pose model” is a processing model, particularly a machine learning model, that can be trained or has been trained to determine the 3D body pose of an occupant.
[0034] In the context of this invention, a "3D support point" is a support point that provides the location of a support point (i.e., a key point) in three spatial dimensions. For example, in a depth map, a support point can be given as a 3D support point. Classical, model-based methods have high accuracy in predicting 3D body pose. Here, a large number of optimization parameters are determined using classical optimization methods. Besides a sufficiently good initial pose, such optimization methods require significant computational resources or are correspondingly slow, and therefore do not meet the real-time requirements of controlling occupant functions of a motor vehicle based on monitoring the interior space of the vehicle. Therefore, the inventors propose a multi-level method in which one or more imprecise 3D body poses are first determined based on the corresponding 2D body pose and the corresponding depth map. The accuracy of the determined 3D body poses is relatively low; for example, the 3D support points determined for key points are located on the corresponding occupant's body surface through depth map determination, but should be within the occupant's body volume. Subsequently, based on the imprecise 3D body poses, a suitable corrected or improved 3D body pose is determined using the body pose model learned for this purpose. In contrast to the classic optimization model, the body posture model is far less computationally intensive, and therefore also meets the real-time requirements for controlling occupant functions.
[0035] Preferably, the improvement includes one or more of the following: correcting the distance from the 3D support points to the depth sensing mechanism, wherein the 3D support points are located on the surface of the corresponding occupant's body, and at least one of the improved 3D support points is located inside the corresponding occupant's body, and particularly correcting the connecting axis between the corrected 3D support points based on the corrected position of the 3D support points. Correcting inconsistent 3D support points such that at least one of the improved 3D support points has a smaller distance from the actual body point of the corresponding occupant. Supplementing unidentified 3D support points, particularly based on symmetry information or contextual information, such that at least one improved 3D support point has a smaller distance from the actual body point of the corresponding occupant. By performing well-defined processing steps, such as correcting depth information or determining improved 3D support points (e.g., based on symmetry information) during the improvement of 3D body pose in a multi-level method, the method can meet the real-time requirements when controlling occupant functions because it is not necessary to optimize parameters in multiple iterations as in optimization-based iterative methods.
[0036] Preferably, the intensity image is input as input data to the recognition model, and the recognition model outputs a 2D body pose including multiple 2D support points as output data. Determining the 2D body pose specifically includes determining one or more confidence maps in the depth map, which encode the positions of the 2D support points and, in particular, the connecting axes connecting some of the 2D support points to each other. The output data also includes a vector field that encodes the orientation of the connecting axes with a high probability. The intensity image is received, either together with a depth sensing mechanism or using a separate sensing mechanism.
[0037] In the context of this invention, a "confidence map" is provided, indicating the probability (also referred to as security or confidence) with which the corresponding support point or connecting axis lies at the location stated in the confidence map. Depending on the corresponding input data, the confidence map may also include depth information, in particular.
[0038] By identifying not only the 2D body pose but also a confidence value for each 2D support point, the uncertainty in determining the 2D support points can be considered during further processing of the 2D body pose. For example, when determining an improved 3D body pose in further processing, this uncertainty can be incorporated into the determination of the uncertainty of the improved 3D body pose. If this uncertainty is used for occupant control, for example, with a suitable filter such as a Kalman filter, the relatively uncertain value has only a small impact on the control of occupant functions, thus making the control more reliable and better.
[0039] Preferably, the 2D support points are, in particular, the actual locations of the occupant's joints, eyes, nose, ears, mouth, hands, center of abdomen, and other salient points in the intensity image, or representations of these actual locations.
[0040] Preferably, determining the 2D body pose for each occupant includes limiting the depth map to an occupant frame corresponding to the occupant, which surrounds the 2D support points of the corresponding 2D body pose.
[0041] In the context of this invention, a "passenger frame" is a portion of a depth map, wherein for each passenger frame, exactly one set of support points for exactly one passenger is always determined, and passenger frames for different passengers may overlap.
[0042] By recognizing the model and limiting the output 2D body pose to the occupant bounding box, the body pose model only needs to determine the 3D support points separately in a portion of the depth map. This reduces the probability of misinterpreting the support points of different occupants in the depth map.
[0043] Preferably, the 3D body pose is shown as image coordinates with a corresponding depth value from the depth map.
[0044] By viewing 3D body pose in an image coordinate system with depth, 3D body pose can be determined from body surface points in a particularly simple manner. The inventors of this invention have recognized that in many methods for determining 3D body pose, the support points are located on the surface of the occupant's body, and these support points are precisely given by a combination of depth information in a depth map and 2D support points.
[0045] But especially for the occupant's joints, the support points are related to the body rather than its surface, so this can be achieved particularly easily by relatively simple movement in depth, especially by moving the distance from the origin, and can be learned by machine learning models.
[0046] Preferably, determining the 3D body pose includes converting the parameterization of image coordinates with depth values, as described above, into parameterization in a Euclidean camera coordinate system, the origin of which is located within the depth sensing mechanism, particularly at its center. The transformation used is inverse projection, and provides a 3D body pose with 3D support points in the Euclidean coordinate system, particularly relevant to the depth sensing mechanism.
[0047] Since the final output of the body pose model, namely the 3D body pose, is performed in camera-centric coordinates, there is no need to perform external camera parameter calibration for the application of the body pose model. Therefore, conversion to the global vehicle coordinate system selected for a specific occupant function is only required after the 3D body pose is output. Correspondingly, the inference of the body pose model, especially the training, is independent of the quality of external calibration.
[0048] Preferably, determining the 3D body pose includes converting the depth map and the 2D body pose into a camera coordinate system by means of inverse projection, wherein the depth sensing mechanism is located at the origin of the camera coordinate system, and the 3D support point is given in the camera coordinate system having the depth sensing mechanism at the origin.
[0049] Preferably, the method further includes determining the corresponding occupant's body parameters based on the occupant frame and the depth map, wherein the body parameters are input into the body model as part of the input data to determine the 3D body pose.
[0050] The body parameters include, in particular, one or more of the following parameters: age, body type, weight, and / or gender.
[0051] By determining the occupant's body parameters, especially when calibrating 3D support points, the calibration of 3D support points can be further improved.
[0052] Preferably, the determination of the body parameters is performed using a body parameter model, wherein the body parameter model is specifically configured to determine the weight parameter and the body shape parameter.
[0053] By employing a body parameter model as a regression model instead of classical optimization methods to determine body parameters, the real-time requirements are further met in methods for determining 3D body pose, although the accuracy in determining 3D body pose is further improved because the body parameters can be used as additional input data for the body pose model and contain information about the magnitude of the offset between body surface points and the 3D body pose. Preferably, the body parameter model and the recognition model are implemented as a common convolutional network, wherein the convolutional network determines the body shape parameter and the weight parameter for each pixel in the depth map, and the confidence map output by the convolutional network includes the body shape parameter and the weight parameter.
[0054] By implementing the body parameter model and the recognition model together in a convolutional network, the required computational resources can be further reduced without sacrificing accuracy, or the processing speed can be further improved.
[0055] Preferably, an average value and an average deviation are determined for each of the body parameters, wherein the average value and / or average deviation are filtered, in particular, using a Kalman filter, and the filtered values, in particular the filtered average value, are input into the body posture model. By filtering the input data using a Kalman filter, body parameter values determined to have large errors have only a small impact when further processed by the body posture model. Furthermore, the variance, here referring to the variance of the determined body posture, can be used as a measure of the reliability of the determined body posture when controlling occupant functions.
[0056] One aspect of the present invention relates to a method for controlling occupant functions based on body parameters of one or more occupants of a motor vehicle, comprising: receiving a depth map using a depth sensing mechanism; determining an occupant frame using a recognition model; determining body parameters using a body parameter model, based on the occupant frame and the depth map; and controlling the occupant functions based on the body parameters.
[0057] The inventors have recognized that common methods for determining body parameters are often associated with high computational overhead, making them unsuitable for real-time determination. By using body parameter models to determine body parameters, accuracy can be improved, and computational overhead can be reduced, thus meeting the real-time requirements for controlling occupant functions with only low computational overhead.
[0058] Preferably, the method further includes determining a 3D body pose according to the method described above for determining a 3D body pose, and the input data of the body parameter model includes the determined 3D body pose. Preferably, the body parameter model has been trained as a regression model for determining weight parameters and body shape parameters.
[0059] Since the body parameter model has been trained as a regression model, especially with negative Gaussian log-likelihood loss, the mean and variance can be directly calculated from the resulting data of the machine learning model. If these inputs are fed into a suitable filter, such as a Kalman filter, before further processing, values with large variances will be effectively filtered out by the Kalman filter in further processing. In addition to its usability in the filter, variance also provides a measure of the reliability of the output that can be used to control occupant functions.
[0060] Preferably, the body parameter model implemented as a regression model is implemented as an ensemble, wherein the ensemble includes multiple identical regression models, the ensemble regression models are initialized with different model parameters before the training, in particular, they are trained differently or with different data, and the variance of the output data of the ensemble regression models is interpreted as the variance of the output data.
[0061] By implementing integration, the variance of the output data from different regression models can be directly interpreted as the variance of the regression, thus deriving consistency of results from the output data in a simple way.
[0062] Preferably, receiving the depth map includes processing the raw image received by the depth sensing mechanism, and processing the raw image includes one or more of the following: distortion correction, conversion to target focal length, scaling, and tone mapping of the intensity image.
[0063] By inputting depth maps into the recognition model, which are always based on the original image with distortion correction or adapted to the target focal length and / or a defined dynamic range, the recognition model can process depth maps particularly reliably and quickly because the occupant to be identified or their body pose always has a typical size or range of extension expected by the model in the depth map.
[0064] Preferably, the method further includes determining the corresponding occupant's body parameters based on the occupant frame and the depth map, wherein the body parameters are input into the body model as part of the input data to determine the 3D body pose, and the body parameters particularly include one or more of the following parameters: age parameter, body size parameter, weight parameter and / or gender parameter, and a set of statistical shape parameters describing the surface of the 3D model.
[0065] By determining the occupant's body parameters, especially when calibrating 3D support points, the calibration of 3D support points can be further improved.
[0066] Preferably, the occupant functions include, in particular, the safety and comfort functions of the motor vehicle, especially controlling one or more of the following systems: seat belt pretensioner system, airbag system, and a reclining comfort system for the front passenger.
[0067] Preferably, the output of the 3D body pose includes an average value and an average deviation, and the average value and average deviation are filtered by a Kalman filter before being used to control the occupant functions.
[0068] The output of the body posture model is filtered by a Kalman filter, and the output data of the body posture model that is determined to have a large error has only a small impact on the control of occupant functions.
[0069] Another aspect of the present invention relates to an evaluation module comprising multiple machine learning models, particularly a recognition model, a body posture model, and a body parameter model, wherein the machine learning models are trained to jointly perform the above-described method.
[0070] Another aspect of the present invention relates to a method for training a machine learning model for the above-described evaluation module.
[0071] Another aspect of the present invention relates to a control device for controlling occupant functions according to any one of the methods described in any one of the claims.
[0072] On the other hand, it relates to a motor vehicle, including a control device for controlling occupant functions according to the above-described method.
[0073] On the other hand, it relates to a computer program product, including instructions that, when executed by a computer, cause the computer to perform the above-described method. Attached Figure Description
[0074] The invention will now be explained in detail with the aid of examples shown in the accompanying drawings. In the drawings: Figure 1 The system is schematically illustrated for use with a method for controlling occupant functions of a motor vehicle based on 3D body posture according to one embodiment; Figure 2 The illustration schematically shows an apparatus for use with a method according to one embodiment; Figure 3 The illustration schematically shows a machine learning model for use with a method according to one implementation. Figure 4 This schematically illustrates a 2D body pose. Figure 5 This illustration schematically depicts a method according to one embodiment of the present invention; Figure 6 This schematically illustrates a method according to one embodiment of the invention; and Figure 7 The illustration schematically depicts a method according to one embodiment of the present invention. Detailed Implementation
[0075] One embodiment of the motor vehicle 100 includes a depth sensing mechanism 102 and a control device 200 (see...) Figure 2 The depth sensing mechanism 102 is communicatively coupled to the control device 200 (e.g., via a wired or wireless communication connection). According to... Figure 1 The motor vehicle 100 also includes an interior space 106, in which an occupant 104 sits.
[0076] Depth sensing mechanism 102 is configured to record a depth map of the interior space 106 of the vehicle 100. For example, depth sensing mechanism 102 may use an infrared light source to receive the depth map 400. The infrared light source emits short light pulses and determines the flight time required for the reflected light pulse to travel back to depth sensing mechanism 102 for each pixel of the depth sensing mechanism 102. In addition to the flight time, depth sensing mechanism 102 also determines the intensity of the reflected light pulse. The depth map 400 is determined based on the determined flight time. However, according to alternative designs, other known depth sensing mechanisms may also be used.
[0077] According to this embodiment, the recorded depth map 400 is transmitted to the control device 200 for further processing.
[0078] Control device 200 includes evaluation module 202, storage module 204, and control module 206. These modules can be configured as either software or hardware modules and are interconnected via channel 208 (see...). Figure 2 Channel 208 is the logical data connection between the modules. Alternatively, the modules can also be interconnected via a bus system.
[0079] The storage module 204 stores the depth map 400 received by the depth sensing mechanism 102 and manages the data to be evaluated in the control device 200.
[0080] Evaluation module 202 is configured to process the depth map 400 received by depth sensing mechanism 102. Specifically, evaluation module 202 includes one or more processing models, particularly machine learning models 300. The machine learning models 300 implemented in evaluation module 202 are particularly a recognition model 602 and a body pose model 606. According to one embodiment, a body parameter model 604 may also be implemented in evaluation module 202.
[0081] In this invention, the control module 206 is configured to control the control device 200. The control module 206 reads data from the storage module 204 and transmits the data to the evaluation module 202. Furthermore, the control module 206 transmits control information for controlling occupant functions to the corresponding control device 200 for controlling the respective occupant functions.
[0082] The different processing models mentioned, especially machine learning model 300. Machine learning model 300 (see...) Figure 3This can be a neural network, particularly one with multiple layers. Specifically, machine learning model 300 has an input layer 302, multiple intermediate layers 304, and an output layer 306. The input layer 302 receives input data 308, processes the input data 308 using the input layer 302, intermediate layers 304, and output layer 306, and outputs result data 310. The input and result data will vary depending on which of the different processing models described herein applies. For some machine learning models, intermediate data 312 may also be output.
[0083] Machine learning models 300 can be implemented as regressors, classifiers, or semantic segmenters.
[0084] In the context of this invention, a "regressor," also referred to as a regression model, is a processing model that performs regression, particularly a machine learning model. If the regressor is a machine learning model, then the regressor is trained to perform regression, particularly through supervised or unsupervised learning.
[0085] In the context of this invention, a "classifier," also referred to as a classification model, is a processing model that assigns a class to the input data or can be trained to assign a class to the input data. The resulting data can, in particular, be the categories assigned to the input data, wherein the format can, in particular, be a vector, where each entry of the vector corresponds exactly to one of the possible categories to be assigned, and in particular, "1" entries in the vector indicate the category of the input data. Alternatively, category numbers can be output. However, as another alternative, the classifier can also be trained such that the resulting data is a vector, where each entry in the vector gives the probability that the corresponding input data belongs to the category corresponding to the entry in the resulting data. Depending on the implementation, the format of the resulting data of the classifier and the format of the target data in the labeled dataset used to train the classifier may vary.
[0086] In the context of this invention, a "semantic segmenter," also known as a semantic segmentation model, is a processing model that assigns an output value to each entry of the input data. The resulting data is also called a segmentation mask or semantic segmentation mask. If the input data is an image, the semantic segmenter performs an image-to-image mapping, using which it assigns an output value corresponding to the semantic meaning to each pixel of the input data.
[0087] In the context of this invention, "supervised learning" is a learning process in which a machine learning model is trained using a labeled dataset to perform a desired mapping.
[0088] In the context of this invention, "unsupervised learning" is a learning process in which a machine learning model is trained solely on an unlabeled dataset without specifying a desired objective, wherein the machine learning model automatically finds or should find specific clustering points in the dataset.
[0089] In the context of this invention, a "labeled dataset" comprises input data and target data, wherein for each input data point, a corresponding label or marker (referred to as the target data) exists in the target data. The target data is typically generated through laborious processing or by manual labeling or similar methods. A processing model is trained using the target data to perform the desired mapping. During training, the resulting data approximates the target data through multiple rounds of training using appropriate training methods.
[0090] In the context of this invention, "target data" refers to the data used when training a processing model to perform a processing mapping, whereby the output data of the processing model based on the input data should be adapted to said data. Approximation is performed using a target function.
[0091] In the context of this invention, the "objective function" is in particular a profit function or loss function that specifies how to evaluate the difference between the resulting data and the target data of the processing model.
[0092] In the context of this invention, a "loss function" is a function that detects the difference between the resulting data and the predetermined target data. If the resulting data and target data are, for example, images, the comparison can be performed pixel-by-pixel. If the resulting data and target data are, for example, vectors or tensors, the difference can be performed entry-by-entry. In an L1 loss function, the differences can be summed as absolute values. In an L2 loss function, the sum of squares of the differences is formed. To minimize the loss function, the values of the model parameters of the processing model are changed, which can be calculated, for example, by gradient descent and backpropagation. Other loss functions that can be considered include, in particular, cross-entropy loss, hinge loss, logistic loss, log-likelihood loss, Gaussian negative log-likelihood loss, or KL divergence loss.
[0093] Especially when training regressors, Gaussian negative log-likelihood loss is frequently used as the loss function. When using Gaussian negative log-likelihood loss, the machine learning model trained with this loss function is trained to output the variance in addition to the mean of the output. Alternatively, the standard deviation or the logarithm of the standard deviation can also be output.
[0094] Especially when training the body pose model 606, the back projection loss function can be used, in which the 3D body pose is projected with the help of the camera model, so that the loss function can be calculated in two-dimensional coordinates.
[0095] In the context of this invention, a "profit function" is an objective function, wherein, in contrast to a loss function that minimizes the difference between the detected data and the target data during training, the profit function detects consistency, which is maximized during training.
[0096] The following is for reference. Figures 4 to 6 A method for controlling occupant functions when using the aforementioned motor vehicle 100 is described. In the described method, in step 502 (see...) Figure 5 The depth map 400 is received by the depth sensing mechanism 102 (see...) Figure 4 The depth sensing mechanism 102 can also directly transmit the depth map 400 to the evaluation module 202 for evaluation, without the need for intermediate storage.
[0097] According to one embodiment, receiving the depth map 400 includes a processing step in which the original image, also known as the raw depth map, received by the depth sensing mechanism 102 is processed using classical image processing methods. Specifically, distortion correction is performed on the raw image, and the geometry in the raw image is calculated based on the target focal length. Alternatively or supplementarily, tone mapping, also known as dynamic compression, can be applied to the intensity values of the raw image. By correcting the distortion of the raw image, it can be ensured that when using different optics in the depth sensing mechanism 102, the distortion-corrected image is always input into one of the machine learning models 300, thereby improving the quality of the resulting data 310 of the corresponding machine learning model 300. Similarly, conversion to the desired target focal length also improves the quality of the resulting data 310, because, for example, distances can be determined more accurately, and the size range of the occupant 104's body can be determined based on said distances, thereby improving the determination of body parameters, for example. Tone mapping is similar. By always normalizing the dynamic range of an image to correspond to the desired input, the accuracy and quality of the resulting data 310 can be improved again for all the machine learning models 300.
[0098] In the next step 504, the recognition model 602 implemented in the evaluation module 202 (see here) Figure 6 The intensity image based on depth map 400 determines, for each occupant 104 detected in depth map 400, 2D support points 402 (see [reference]). Figure 4The model identifies the 2D body pose and determines a vector field for each connecting axis, which gives the probability of the connecting axis pointing in which direction at the corresponding position in the vector field. According to this embodiment, the result data 310 of the identification model 602 includes a confidence level for each of the support points 402, which gives the probability of the corresponding support point 402 being present there.
[0099] According to one embodiment, the recognition model 602 can also be configured to determine the occupant frame 406 based on the determined 2D support points 402. The occupant frame 406 surrounds the 2D support point 402 corresponding to one occupant 104. Multiple 2D support points 402 of occupants 104 can overlap on the depth map 400, and the occupant frame 406 output by the recognition model 602 always includes only the support point 402 of one occupant 104 and its corresponding associated connecting axis 404. In particular, according to this design, the recognition model 602 outputs a portion of the depth map 400 corresponding to the respective occupant frame 406, along with the 2D support point 402 corresponding to the occupant frame 406.
[0100] According to this implementation, the occupant box 406 is a so-called bounding box, as determined using common neural networks. It is called a bounding box, especially when it involves a rectangle enclosing an object.
[0101] According to one design scheme of this embodiment, a segmentation mask can also be output instead of a bounding box. In the segmentation mask, a value is assigned to each pixel to identify the occupant 104. For example, "1" is assigned to all pixels corresponding to the driver. "2" is assigned to all pixels corresponding to the front passenger. This corresponds to the number of seats in the vehicle 100. "0" can be assigned to all pixels that do not belong to any occupant. The selected values should, of course, be understood exemplarily only. Similar to the confidence map for support point 402, confidence maps can also be output for the different possible occupants 104 in the vehicle 100.
[0102] Such a support point 402 is Figure 4 The occupant 104 of the motor vehicle 100 is schematically shown in depth diagram 400. Figure 4 An intensity image of depth map 400 is schematically shown. According to this embodiment, support points 402 are precisely the locations of the joints of occupant 104, particularly the shoulder, hip, elbow, and wrist joints in depth map 400. Other support points 402 may be prominent points on the face, such as the nose, ears, eyes, and mouth.
[0103] In the intensity image, the support points 402 identified by the recognition model 602 are plotted as circles, where the circles in the face are chosen to be smaller, especially for visual clarity. The support points 402 are connected to each other by connecting shaft 404. Two occupants 104 can be identified in the depth map 400. The recognition model 602 also defines occupant frames 406 for each support point 402. The body posture shown, especially for the left occupant 104, shows a slightly forward-bent upper body posture. Furthermore, in the image portion, the knees are barely visible, and the feet of the occupant 104 are practically not visible at all. The extent to which different limbs are detected by the depth sensing mechanism 102 depends on the geometry and positioning of the depth sensing mechanism 102. In particular, multiple depth sensing mechanisms 102 can also be used to detect all limbs. In addition, another optics in the depth sensing mechanism 102 with, for example, a larger or smaller angle can also be used.
[0104] Other body postures, especially in terms of limb position, can differ significantly from those shown here. Further bending forward or backward, as well as rotation, can also be detected. For example, occupant 104, particularly the front passenger, may assume a reclining position in their seat. This also applies particularly to occupant 104 in the rear seats. Occupant 104 in the rear seats within the interior space 106 of the vehicle 100 can be detected using the same depth sensing mechanism 102. However, according to one embodiment, multiple depth sensing mechanisms 102 can be installed in the vehicle 100, allowing occupant 104 in the rear seats to also be reliably detected in the depth map 400 by the depth sensing mechanisms 102.
[0105] According to one design of this embodiment, the recognition model 602 replaces the confidence map and only outputs the position of the support point 402 in the intensity image. Either the recognition model 602 is directly trained to output the support point 402, or the recognition model 602 also calculates the confidence map, and then the support point 402 is determined from the confidence map in the final step.
[0106] According to another step 506, the 3D body pose is determined from the depth map 400 and the 2D support points 402. To do this, the 2D support points 402 are first supplemented with the depth information from the depth map 400, thereby obtaining the 3D support points 402. In this case, the 3D support points include the 2D support points output by the recognition model 602, which now also have another spatial dimension for points in the depth map 400, here being the depth information of corresponding points in the intensity image or depth information from the depth map 400. The 3D support points 402 form the 3D body pose.
[0107] According to one embodiment, 2D support points are converted to 3D support points before using the body pose model 606, using the inverse projection matrix of a calibrated depth camera 102 and based on depth values at corresponding locations in the depth map 400. This step reverses the projection, thus giving the 3D support points in a Euclidean 3D coordinate system, the origin of which is located at the camera center of the depth camera 102. The resulting 3D support points form the 3D body pose. A machine learning model is trained to correct the 3D support points in Euclidean coordinates, for example, to facilitate learning the inherent symmetry between the left and right halves of the body.
[0108] According to step 508, body posture model 606 improves the 3D body posture based on the 3D body posture and outputs the improved 3D body posture to control module 206. Improving the 3D body posture can include one or more different improvements. One possible improvement includes determining improved depth information, wherein body posture model 606, particularly for 3D support points corresponding to the joint positions of the occupant, corrects the positions of the 3D support points, especially with respect to depth, so that they are no longer located on the body surface of the corresponding occupant given by depth map 400, but rather inside the corresponding occupant's body. Another alternative improvement includes correcting inconsistent 3D support points. For example, individual support points 402 in depth map 400 may be occluded, so that the depth value belongs to the occluded object and not to the body surface at the position of said support point 402. Furthermore, it may be possible that, in the case of occlusion, recognition model 602 outputs one or more 2D support points that do not correspond at all to the actual positions of the corresponding support points 402 in the occupant. Body posture model 606 can be trained to recognize such inconsistent support points 402 and estimate better positions, for example, based on the symmetry of the occupant's body. In particular, it is also possible that the identification model 602 may not have identified the determined support points 402 at all, which can then be estimated by the body pose model 606. For example, the determined support point 402 may be located outside the image detected by the depth sensing mechanism 102. In this case, the body pose model 606 can also estimate support points 402 whose positions are closer to the actual positions. The last improvement mentioned is also called supplementing unidentified 3D support points. For unidentified 3D support points, the output is support point 402 with only "0" entries. When supplementing unidentified 3D support points, the "0" entries are replaced with 3D support points estimated by the body pose model 606. Typically, 3D support points have significantly better consistency with actual body points than "0" entries. Simultaneously learning the described improvement enables the generation of a body pose model that is robust to all corresponding error sources.
[0109] According to this design, the support point 402 is used as input data 308 into the body pose model 606. This significantly reduces the complexity of the processing mapping to be learned, as the complex image data does not need to be processed by the body pose model 606, and therefore the body pose model 606 does not need to learn the complex semantics in the image data. This significantly improves the real-time processing capability. Furthermore, using the support point 402 as input data 308 also significantly simplifies the generation of synthetic data for training, because it eliminates the need to use complex rendering of image data as input data 308.
[0110] According to one design scheme of the first embodiment, multiple 3D body poses can be generated from 2D support points and confidence maps, and used as input signals for body pose model 606. Here, a selection method, such as random selection from a Gaussian distribution, is controlled based on the confidence values at the corresponding 2D support points. This selection method selects multiple 2D support points with corresponding depth values from the region surrounding the original 2D support points. Multiple corresponding 3D support points are generated through inverse projection of the 2D support points with depth values, which in turn form multiple 3D body poses. By selecting depth values and 2D support points from the region surrounding the original support points, the distribution of depth values and the confidence of the recognition model 602 are taken into account. The distribution of multiple 3D body poses facilitates the learning of necessary corrections by the body pose model 606.
[0111] According to another design scheme of the first embodiment, the confidence value output by the recognition model 602 can be directly calculated with the 3D support points, for example, by multiplying the confidence value on the 2D support points with the corresponding intermediate or final result of the body pose model 606, such as a correction value or variance. This can improve the learning and accuracy of 3D body pose and the uncertainty associated with the 3D body pose.
[0112] According to another alternative, the body pose model 606 can directly process the confidence map. The body pose model 606 then outputs the 3D body pose and its variance as output data.
[0113] According to step 510, the evaluation module 202 transmits the improved 3D body posture to the control module 206, which then uses the improved 3D body posture to control occupant functions. Occupant functions here may include, in particular, safety or comfort functions. Specifically, safety functions may include seatbelt pretensioners and airbag functions.
[0114] For example, 3D body posture can indicate that occupant 104 has adopted a forward-leaning 3D body posture. If control device 200, and in particular control module 206, determines that the vehicle is highly likely to perform braking, especially a strong braking operation, control device 200 can control the seat belt pretensioner, causing occupant 104 to be pulled into or forced into an upright body posture by the seat belt pretensioner. In the event of an unavoidable accident detected by control module 206, airbag control can be adapted, for example, based on the distance from occupant 104's head and torso to the steering wheel and based on occupant 104's weight, thereby reducing the risk of injury. In comfort functions, for example, air conditioning functions can be suitably adapted, for example, to the seat position. For example, the vents used for the air conditioning function can be coordinated with the 3D body posture; for example, when occupant 104 is in a more reclined position, different air vents can be used than when the corresponding occupant 104 is sitting upright. Warning features include detecting dangerous body postures, such as reaching back or toward the glove box, or generally detecting one or both hands not touching the steering wheel.
[0115] In addition to the 3D body posture determined by the body posture model 606, occupant information is also output to the control module 206, which in particular includes information about the seat position, so that occupant functions that are coordinated with the seat position identified in the vehicle 100 can be performed, such as triggering the airbag or seat belt pretensioner only for the corresponding occupant 104.
[0116] According to one embodiment of this design, the method may further include methods for using body parameter model 604 (see Figure 6 Step 507 for determining body parameters. The occupant frame 406 determined by the recognition model 602 is input as input data 308 into the body parameter model 604. The occupant frame 406 is a partial region of the depth map 400 corresponding to an occupant, the partial region surrounding the support point 402 of an occupant 104. The occupant frame 406 is output by the recognition model 602. Based on the occupant frame 406, the body parameter model 604 determines the body parameters. Here, the body parameters specifically include the height and weight of the corresponding occupant 104.
[0117] The result data 310 of body parameter model 604 here includes the height and weight of the corresponding occupant 104. In particular, body parameter model 604 can also output gender and age. The body parameters output by body parameter model 604, along with confidence maps or 2D support points, are input into body posture model 606 as part of input data 308. In particular, evaluation module 202 summarizes the 2D support points and body parameters into input data 308 for body posture model 606.
[0118] The inventors have recognized that the correction quality of the 3D support points directly depends on the body parameters of the occupant 104, such as weight and height. Therefore, a body posture model 606 is attached to the 2D support points to receive the body parameters height and weight as part of the input data 308. Additionally, gender and age, or shape parameters of a statistical model giving the volume and weight distribution of all segments, can also be estimated from the body parameter model 604. These also affect the correction performed by the body posture model 606.
[0119] According to one design scheme of this embodiment, the result data 310 of the machine learning model 300 can be filtered by a Kalman filter before further processing. The advantage of the Kalman filter is that, in addition to the input measurement value, the estimation error of the measurement value is also input into the filter. Measurement values with large estimation errors are evaluated in the Kalman filter such that the measurement value has a smaller impact on the filter's output.
[0120] The following is for reference. Figure 7 The second embodiment is described. The difference between the second and first embodiments lies in the step of controlling occupant functions using body parameters. According to step 702, an image is received using the depth sensing mechanism 102 as in step 502. According to the subsequent step 704, occupant frames 406 are determined. Unlike the first embodiment, in the second embodiment, the recognition model 602 only determines the occupant frames 406, omitting the output of 2D body pose including 2D support points. Accordingly, the recognition model 602 of the second embodiment is constructed more simply and only needs to identify the occupants detected in the depth map 400 and output the corresponding occupant frames 406. In particular, in this embodiment, the recognition model 602 is configured as a detector that outputs a list of the found occupant frames 406.
[0121] In another step 706, the second embodiment includes determining body parameters based on the occupant frame 406 using a body parameter model 604. The body parameter model 604 corresponds to the identity parameter model 604 of the first embodiment.
[0122] According to one design of this embodiment, the body posture model 606 provides additional input data 308 for the body parameter model 604 to improve accuracy. In this embodiment, for example, the distance between 3D support points can be used to improve the accuracy of body parameters such as height and weight.
[0123] According to step 708, as described above with reference to the first embodiment, but this time the occupant functions are controlled based on body parameters.
[0124] Based on the above implementation forms, variations, and design schemes, the machine learning model 300 can be a regressor, a classifier, or a semantic segmenter. Correspondingly, when using supervised learning, the corresponding input data 308, result data 310, and target data differ during training.
[0125] If the machine learning model 300 is trained as a classifier, then one or a combination of the following functions should be used as the loss function: cross-entropy loss, hinge loss, exponential loss, squared loss, Savage-Loss, tangent loss, and Gaussian loss.
[0126] When training the classifier, it is particularly important to consider the categories to be identified as the target data. This can be achieved through hard or soft assignment to one or more categories, corresponding to the choice of the loss function or output layer 306.
[0127] If the machine learning model 300 is a semantic segmenter, then the output is specifically the segmentation mask.
[0128] According to one design, the depth map 400 includes an intensity image, and the depth sensing mechanism 102 simultaneously receives both the depth map 400 and the intensity image. According to an alternative, the intensity image can be received independently of the depth map 400. In particular, a camera for receiving the intensity image can be provided attached to the depth sensing mechanism 102. Preferably, the camera and the depth sensing mechanism 102 are located in the same position. According to an alternative, the camera and the depth sensing mechanism 102 can be located in different positions. In this case, the intensity image and the depth map 400 must be appropriately transformed so that they are approximately registered to each other, i.e., the same image points detect the same objects. The transformation can also be performed separately for portions of the depth map 400 and the intensity image.
[0129] According to another embodiment, the invention also includes equipment for performing the above-described methods. In particular, a control device 200. The control device 200 may, in particular, be an ECU, i.e., an electronic control unit.
[0130] According to another embodiment, the present invention provides a computer program product comprising instructions that cause a computer to perform the above-described method.
[0131] According to another embodiment, the present invention provides a method for training the above-mentioned machine learning model 300, namely, recognition model 602, body parameter model 604 and body posture model 606.
[0132] According to another embodiment, the present invention provides a motor vehicle 100 having a device, in particular a control device 200, for performing the above-described methods.
[0133] List of reference numerals
[0134] 100 motor vehicles
[0135] 102 Depth Sensing Mechanism
[0136] 104 crew members
[0137] 106 interior space
[0138] 200 control device
[0139] 202 Assessment Module
[0140] 204 storage module
[0141] 206 control module
[0142] 208 channels
[0143] 300 machine learning models
[0144] 302 Input Layer
[0145] 304 intermediate layer
[0146] 306 Output Layer
[0147] 308 Input Data
[0148] 310 Results Data
[0149] 312 intermediate data
[0150] 400 Depth Map
[0151] 402 support level
[0152] 404 connecting shaft
[0153] 406 crew frame
[0154] 602 recognition model
[0155] 604 Body Parameter Model
[0156] 606 Body Posture Model
Claims
1. A method for controlling occupant functions based on the 3D body pose of one or more occupants of a motor vehicle (100), the method comprising: A depth map (400) is received (502) by means of a depth sensing mechanism (102). Using the recognition model (602), for each occupant detected in the depth map (400), at least one 2D body pose comprising multiple 2D support points is determined (504). Determine (506) one or more 3D body poses based on the depth map (400) and the 2D support points, wherein the 3D body poses include 3D support points located on a surface given by the depth map (400), particularly on the surface of the respective occupant's body. For each of the 3D body poses, the 3D support points are improved (508) based on the 3D body pose using a body pose model (606), and The occupant function is controlled (510) based on the improved 3D body posture, characterized in that, Compared to the corresponding 3D support point (402) before the improvement (508), at least one of the improved 3D support points of the improved 3D body posture has better consistency with the actual body point corresponding to the respective occupant.
2. The method according to claim 1, wherein, The improvement (508) includes one or more of the following: The distance from the 3D support points to the depth sensing mechanism (102) is corrected, wherein the 3D support points are located on the surface of the corresponding occupant's body, and at least one of the improved 3D support points is located inside the corresponding occupant's body, and the connecting axis between the corrected 3D support points is also corrected based, in particular, on the corrected position of the 3D support points. The inconsistent 3D support points are corrected so that at least one of the improved 3D support points has a smaller distance from the corresponding actual body point of the occupant. Supplementing unidentified 3D support points, especially based on symmetry or contextual information, allows at least one of the improved 3D support points to better align with the actual body points of the corresponding occupant.
3. The method according to claim 1 or 2, wherein, The intensity image is input as input data (308) into the recognition model (602), and the recognition model (602) outputs a 2D body pose including the plurality of 2D support points as output data, wherein determining the 2D body pose includes determining one or more confidence maps, the confidence maps encoding the positions of the 2D support points and, in particular, the connecting axes connecting some of the 2D support points in the depth map (400), and the output data also includes a vector field encoding the direction of the connecting axes with a high probability, wherein the intensity image is received together with the depth sensing mechanism (102) or a separate sensing mechanism is used to receive the intensity image.
4. The method according to any one of claims 1 to 3, wherein, The 2D support points approximate the actual positions of the occupant's joints, eyes, nose, ears, mouth, chin, hands, center of abdomen, or other prominent points.
5. The method according to any one of claims 1 to 4, wherein, Determining the 2D body pose for each occupant includes limiting the depth map (400) to an occupant frame (406) that surrounds the 2D support points of the corresponding 2D body pose.
6. The method according to any one of claims 1 to 5, wherein, Determining the 3D body pose involves converting the depth map (400) and the 2D body pose into a camera coordinate system by means of inverse projection, wherein the depth sensing mechanism (102) is located at the origin of the camera coordinate system, and the 3D support point is given in the camera coordinate system with the depth sensing mechanism (102) as the origin.
7. The method according to any one of claims 1 to 6, wherein, The output of the 3D body pose includes an average value and an average deviation, and the average value and average deviation are filtered by a Kalman filter before being used to control the occupant functions.
8. The method according to any one of claims 1 to 7, wherein, The method further includes determining (507) the corresponding occupant's body parameters based on the occupant frame and the depth map, and inputting the body parameters as part of the input data into the body model to determine the 3D body pose.
9. The method according to claim 8, wherein, The determination of the body parameters is performed using a body parameter model, wherein the body parameter model is specifically configured to determine the weight parameter and the body shape parameter.
10. The method according to claim 9, wherein, The body parameter model and the recognition model are implemented as a common convolutional network, wherein the convolutional network determines the body shape parameter and the weight parameter for each pixel in the depth map, and the 2D body pose output by the convolutional network includes the body shape parameter and the weight parameter.
11. The method according to any one of claims 8 to 10, wherein, For each of the body parameters, a mean and a mean deviation are determined, wherein the mean and mean deviation are filtered, in particular, by means of a Kalman filter, and the filtered values are input into the body posture model.
12. A method for controlling (510) occupant functions based on physical parameters of one or more occupants of a motor vehicle, the method comprising: A depth map (400) is received (502) by means of a depth sensing mechanism (102). The occupant frame (504) is determined using a recognition model. Body parameters (506) are determined using a body parameter model, based on the occupant frame (406) and the depth map (400). The occupant functions are controlled (510) based on the said body parameters.
13. The method according to claim 12, wherein, The method further includes determining (506) a 3D body pose by any one of claims 1 to 7, and the input data (308) of the body parameter model (604) includes the 3D body pose.
14. The method according to any one of claims 9 to 12, wherein, The body parameter model (604) has been trained as a regression model to determine the weight parameter and the body shape parameter.
15. The method according to claim 14, wherein, The body parameter model (604) implemented as a regression model is implemented as an ensemble, wherein the ensemble includes multiple identical regression models, the ensemble regression models are initialized with different model parameters prior to the training, and the variance of the output data of the ensemble regression models is interpreted as the variance of the output data.
16. The method according to any one of claims 8 to 15, wherein, The body parameters include, in particular, one or more of the following parameters: age, body type, weight, and / or gender.
17. The method according to any one of claims 1 to 16, wherein, Receiving (502) the depth map (400) includes processing the raw image received by the depth sensing mechanism (102), and processing the raw image includes one or more of distortion correction, conversion to target focal length, scaling, and tone mapping.
18. The method according to any one of claims 1 to 17, wherein, The occupant functions include, in particular, the safety and comfort functions of the motor vehicle (100), and especially control one or more of the following systems: Seatbelt pretensioning system airbag system Comfort reclining system for the co-pilot Warning system.
19. The method according to any one of claims 1 to 18, wherein, The depth map (400) assigns depth information from the detected scene to each point.
20. A method for training a machine learning model (300) to perform the method according to any one of claims 1 to 19.
21. An evaluation module (202) comprising multiple machine learning models (300), particularly a recognition model (602), a body pose model (606), and a body parameter model (604), wherein, The machine learning models (300) are trained to perform the method according to any one of claims 1 to 20.
22. A control device (200) for controlling occupant functions according to any one of claims 1 to 19.
23. A motor vehicle (100) comprising a control device (200) for controlling occupant functions in accordance with any one of claims 1 to 19.
24. A computer program product, comprising instructions that, when executed by one or more computers, cause the computers to perform the method according to any one of claims 1 to 20.