Techniques for generating composite image data for body detection related tasks
By generating synthetic image data, combining 2D skeleton representation and text prompts, downstream AI models are trained to solve the problem of insufficient performance of 3D body pose estimator in the real world and improve its detection capabilities in automated driving and autonomous robot systems.
Patent Information
- Application Number
- CN202510170065.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-15
- Filing Date
- 2025-02-17
- Publication Date
- 2025-08-15
AI Technical Summary
The difficulty in obtaining accurate 3D body posture annotation images in the prior art has led to insufficient performance and robustness of 3D human posture estimators in the real world and insufficient representation of data sets, which limits their application in fields such as automated driving and autonomous robots.
By generating synthetic image data, including 2D skeleton representation, 2D depth map and dense semantic coding, combined with text prompts, downstream AI is trained using conditional image synthesis models to improve its robustness and generalization capabilities, especially when sensor data is limited.
Improves the training efficiency and accuracy of downstream AI in body detection tasks, enhances its performance in automated driving and autonomous robot systems, and is able to handle diverse body postures and environments.
Smart Images

Figure CN120495434A_ABST
Abstract
Description
[0001] The present invention relates to techniques for generating synthetic image data for body detection-related tasks, wherein the synthetic image data can be used to train, validate, and / or test downstream artificial intelligence (AI), in particular downstream neural networks (NN); techniques for training conditional image synthesis models to generate synthetic image data for body detection-related tasks based on sensor data, wherein the synthetic image data can be used to train, validate, and / or test downstream AI, in particular downstream NN; and techniques for training, validating, and / or testing downstream AI, in particular downstream NN, for performing body detection-related tasks based on sensor data. In particular, methods, computing devices, systems, computer program products, and computer-readable storage media are provided. Background Art
[0002] For applications such as automated (especially autonomous) driving, planning the movements of (especially autonomous) robots, operating smart home appliances, and / or controlling surveillance systems, it is crucial not only to recognize the presence of living bodies (e.g., humans or animals), but also to estimate their three-dimensional (3D) pose and / or their movements in the near future. In the context of automated (and / or autonomous) driving, this includes humans, extreme situations such as dangerously maneuvered cars, or near-misses involving pedestrians.
[0003] In principle, trained AI models can be used to estimate 3D body pose from images. However, acquiring images with accurate 3D body pose annotations is often a difficult process. Specialized capture studios are required, and data acquisition is slow. This severely limits the diversity of scene positions, human appearance, and poses that can be represented in such datasets. This means that current evaluations of 3D human pose estimators are performed on datasets that do not represent the full range of situations that occur in the real world. Consequently, the actual real-world performance and robustness of current state-of-the-art 3D human pose estimators are questionable. Summary of the Invention
[0004] In the following, the solution according to the present invention is described with respect to the claimed method and with respect to the claimed computing device. Features, advantages or alternative embodiments herein can be assigned to other claimed objects (e.g., systems, computer programs or computer program products), and vice versa. In other words, features described or claimed in the context of the method can be used to improve claims for the computing device. In this case, the functional features of the method are respectively implemented by the structural elements of the corresponding computing device, and vice versa.
[0005] With respect to the first method aspect, a computer-implemented method for generating synthetic image data for body detection-related tasks based on sensor data is provided, which synthetic image data can be used to train, validate and / or test downstream AI, in particular downstream NN. The method includes the step of receiving visual information related to the body. The visual information includes a two-dimensional (2D) skeletal representation of the body, a 2D projection (in particular, dense) semantic encoding of the body, and a 2D depth map of the body. The method further includes the step of receiving textual prompts related to the appearance of the body and / or environmental information related to the body. The method still further includes the step of generating synthetic image data of the body based on the received textual prompts conditioned by the received visual information. The generation is performed by a conditional image synthesis model.
[0006] By means of a technology including a computer-implemented method, improved robustness and generality of a downstream artificial intelligence (AI, in particular a downstream NN) trained by provided downstream (in particular synthetic) training data is provided. Alternatively or additionally, data augmentation and / or data enrichment (e.g., in view of the diversity of bodies, 3D body poses, appearances, environments and / or positions) is provided, in particular for efficiently training, validating and / or testing the downstream AI (in particular a downstream NN), in particular in the presence of limited (in particular real) data. As a result, the generalization ability of the downstream AI (in particular a downstream NN) can be improved.
[0007] This technology can further enable training, validating, and / or testing downstream AI (particularly downstream NN) for domain conversion tasks (e.g., for converting synthetic images into realistic images).
[0008] Downstream AI (particularly downstream NN) can be configured for object (particularly including body) detection (e.g., 3D) body pose detection, classification and / or semantic segmentation of image data received via sensors (e.g., particularly video cameras, radar sensors, LiDAR sensors, ultrasonic sensors, motion sensors and / or thermal sensors).
[0009] Downstream AI (in particular, downstream NN) can be trained, validated and / or tested by synthetic image data (also referred to as: synthetic images), in particular considering the received visual information related to the body (or any further information related to the body determined based on the visual information) as ground truth.
[0010] The body may comprise a (particularly living) human body and / or a (particularly living) animal body.
[0011] For example, for downstream applications of autonomous driving or driver assistance systems (collectively referred to as automated driving, with levels L1 to L5, and particularly higher levels L3 to L5), the human and / or animal body may include living beings in a traffic situation. Alternatively or additionally, the human and / or animal body may include living beings in a home automation environment.
[0012] In terms of being associated with the same body and / or the same (e.g., 3D) body pose, the fragments of visual information of the body, namely the 2D skeletal representation, the 2D (particularly dense) semantic encoding, and the 2D depth map, can be mutually consistent (and / or include conditions for mutual consistency). Alternatively or additionally, the visual information related to the body can be determined based on the input to the generated body model. The input to the generated body model can include a set of shape parameters and a set of posture parameters related to the body. The output of the generated body model can include a generated representation of the body, in particular a representation including a 3D skeleton (e.g., including major joints and / or bones, for example, including ankle joints, knee joints, and hip joints for specifying the position of lower limbs) and a 3D surface (e.g., as a 3D mesh). The shape parameters can take into account the body proportions, height, and / or weight of the body. Alternatively or additionally, the posture parameters can take into account the 3D positions (and / or joints) of the skeleton of the body (and / or in particular the positions of major joints and / or bones). The posture parameters and the shape parameters can be collectively represented as a 3D body pose.
[0013] The generated body model may account for variations in body shape according to (eg, 3D) body pose.
[0014] The generated body model may include a skinned multi-person linear (SMPL) model, as described by M. Loper et al. in “SMPL: Askinned multi-person linear model,” ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, October 2015, which is incorporated herein by reference.
[0015] Alternatively or additionally, a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body may be determined based on the generated body model. For example, the 2D skeletal representation of the body may include a projection of the 3D position of the skeleton onto the image plane. Alternatively or additionally, the 2D projection (particularly dense) semantic encoding of the body and / or the 2D depth map of the body may include a projection of the 3D surface of the body (e.g., represented by a 3D mesh) onto (particularly the same) image plane.
[0016] The 2D projection encoding of the body may be a semantic encoding, including semantic information for each pixel and / or voxel within the image plane. For example, a pixel may include semantic information (e.g., represented by one or more colors) encoding an anatomical structure (e.g., a limb (e.g., an arm or leg), a head, or a body trunk) that is to be imaged in the (e.g., generated) synthetic image data.
[0017] Alternatively or additionally, the encoding can be dense. Dense encoding can mean that the encoding is performed independently for each pixel and / or voxel. Alternatively or additionally, dense (and / or dense semantic encoding) can refer to the fact that a code (and / or code vector) is provided for each point (and / or pixel and / or voxel) on the surface of the body. This is in contrast to semantic encoding, which only provides codes for certain points (and / or pixels and / or voxels). Figure 7B An example of conventional (especially non-dense) encoding is illustrated in . Figure 7B The grayscale values (and / or shades of gray) in may, for example, provide semantic information only for the joints (e.g., left elbow, right elbow). Alternatively, the semantic encoding may refer only to the limbs (e.g., as large-scale structures).
[0018] A 2D depth map can encode the (e.g., relative) depth of the information being imaged, for each pixel and / or voxel, for example. Bright colors (e.g., white or light gray) can encode information close to the viewer, while dark colors (e.g., dark gray or black) can encode information farther from the viewer. For example, the more anatomical structures are located in the foreground of the composite image data being generated, the brighter the pixels may be.
[0019] The received text prompt (also referred to as text prompt) may include data representing information in a text format (e.g., as raw data and / or in particular a natural language). Receiving may be performed via a digital interface (e.g., a user interface UI, in particular a graphical user interface GUI) and / or a human-machine interface (HMI). Alternatively or additionally, the received text prompt may be received in a sound format to be converted into a processable format as input to the conditional image synthesis model. Further alternatively or additionally, the text prompt may be received from a computer-implemented text prompt generator that is configured to generate a set of different text prompts.
[0020] The received textual prompts relating to the appearance of the body may include information about the appearance of the body; the height of the body; the weight of the body; the proportions of the body (e.g., including the relationship between the dimensions of the body's head and its torso); age; race (particularly in the case of a human body); species (particularly in the case of an animal body), including, for example, a horse, cow, dog and / or cat; biological sex; and / or clothing worn by the (particularly human) body, including, for example, outerwear such as a coat.
[0021] The textual cues related to the appearance of the body may supplement the received visual information with further details about the appearance of the body for use in generating the composite image.
[0022] The text prompt received related to environmental information related to the body may include information about the location (e.g., outside, in an urban area, and / or in a rural surrounding); lighting conditions (e.g., during daytime or during nighttime); the background scene in which the body is to be placed; and / or weather conditions (e.g., sunny, foggy, rainy, and / or snowy).
[0023] The textual hint associated with the environmental information can be used to generate the environment of the body in the synthetic image data (e.g., a realistic version thereof). For example, the textual hint can be "woman on a bridge," or "man walking in a park during a rainy day," or "man wearing a coat on stairs."
[0024] In a further improvement, the received visual information and textual prompts may comprise a plurality of visual information and associated textual prompts, each visual information and associated textual prompt being directed to a (particularly a different) body, so that the generated composite image data will comprise a plurality of bodies.
[0025] Generating synthetic image data of the body by the conditional image synthesis model may involve calculating synthetic images by image processing performed via a digital computer unit.Alternatively or additionally, the generating is not necessarily based on receiving an image by an optical sensor.
[0026] The conditional image synthesis model can be or include a generated AI and / or a generated model. The conditional image synthesis model can be or include a text-to-image model that converts a received text prompt into one or more features of synthesized image data. The synthesized image can be generated while controlling, for example, 3D body pose (as encoded by visual information), appearance, and position (as examples of an environment).
[0027] The conditional image synthesis model can be denoted as upstream AI (and / or upstream model).
[0028] The conditional image synthesis model may include a deep learning (DL) model.
[0029] The conditioning may comprise a combination of a depth map, a (particularly dense) semantic encoding of the body (also referred to as: semantic information) and a 2D skeleton (and / or 2D projection) of the body. Alternatively or additionally, the content of the generated synthetic image data may be controlled by text, in particular in the form of textual cues.
[0030] A downstream AI (also referred to as a downstream model), in particular a downstream NN, may receive the generated synthetic image data and, optionally, the received visual information (and / or its root information, including 3D body pose, and / or pose parameters and shape parameters, and / or any further information related to the body determined based on the visual information) as input. The received visual information (and / or its root information, including 3D body pose, and / or pose parameters and shape parameters, and / or any further information related to the body determined based on the visual information) may encode basic facts associated with the generated synthetic image data for use in performing downstream tasks by the downstream AI, in particular the downstream NN.
[0031] Downstream tasks may include object detection, (e.g., 3D) body pose detection, classification, and / or semantic segmentation of image data. Downstream tasks may enable the detection of obstacles relevant to safe automated (particularly autonomous) driving, planning the movement (and / or operation) of robots and / or household appliances in automated systems, and / or enable the detection of humans and / or animals subject to access control.
[0032] The method may further comprise the step of determining visual information related to the body from the generated body model, the visual information comprising a 2D skeletal representation of the body, a 2D projected (particularly dense) semantic encoding of the body, and a 2D depth map of the body. Determining the visual information may comprise performing a projection of the generated body model onto an image plane.
[0033] A 2D skeleton representation (also referred to as information indicating 2D skeleton positions) may specifically include 2D keypoint information. 2D keypoints may refer to the positions of (e.g., major) joints of a human or artificial body (e.g., a subset of joints representing large-scale joints (and / or motor functions), such as shoulders, hips, elbows, knees, hands, and / or feet). The 2D keypoint information may include 2D projections of the positions of (e.g., major) joints onto the image plane.
[0034] A 2D depth map may encode a 3D body pose based on a (eg simple) 3D surface representation (particularly a 3D mesh). The 2D depth map may include assigning a brightness (and / or color) to each point of the 3D surface representation.
[0035] The 2D keypoints may lie within a volume enclosed by a (eg simple) 3D surface representation (in particular a 3D mesh).
[0036] 2D projection (especially dense) semantic encoding (also called semantic information) can assign information about body parts to each point in the 2D depth map. Semantic information can include assigning a color (and / or brightness) to each point in the 2D depth map. For example, the color can be specific to an anatomical structure, such as an arm or a leg.
[0037] By determining the visual information with the aid of a generated body model, a consistent combination of a 2D skeletal representation of the body, a 2D projective (especially dense) semantic encoding and a 2D depth map is ensured.
[0038] The generated body model may generate a model of the body based on a set of shape parameters and pose parameters. Optionally, the generated body model may include an SMPL model.
[0039] By parameterizing the body according to a set of shape parameters and pose parameters, a variety of (eg, 3D) body poses and body shapes, and / or a variety of visual information related to the body can be generated.
[0040] SMPL includes a learned model of human shape and pose-related shape variations and / or a skinned vertex-based model of natural human pose, as described by M. Loper et al. in "SMPL: A Skinned Multi-Person Linear Model", ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6): 248: 1-248: 16 (October 2015), which is incorporated herein by reference. By using SMPL, it is possible to accurately represent the shape and pose of the human body (and / or 3D body pose). SMPL may include consideration of forward kinematics, linear blending skinning and / or linear regression to obtain 3D keypoints. 3D keypoints may include 3D (particularly coordinates and / or voxel) positions of (e.g., major) joints.
[0041] Alternatively or additionally, for animal bodies, shape and pose models are provided by S. Zuffi et al., “3D Menagerie: Modeling the 3D Shape and Pose of Animals”, http: / / arxiv.org / abs / 1611.07700, [5a], which is incorporated herein by reference. A model is learned for different species (e.g., including quadrupedal mammals).
[0042] The method may further include providing the generated synthetic image data of the body to a downstream AI, particularly a downstream neural network. The downstream AI, particularly the downstream neural network, may be configured to perform body detection-related tasks based on the sensor data. Providing the generated synthetic image data may include providing visual information, a set of shape parameters and pose parameters of a generated body model, the generated body model based on which the visual information is determined, and / or one or more related quantities as ground truth for the body detection-related tasks.
[0043] By providing generated synthetic image data along with basic facts, efficient training, validation and / or testing of downstream AI, especially downstream NN, can be enabled.
[0044] The conditional image synthesis model may include: a generated text-to-image model, which is configured to generate synthetic image data based on a received text prompt; and an image adjustment model, which is configured to encode the received visual information to adjust and / or control (and / or modify) the generated text-to-image model.
[0045] The generated text-to-image model may enable generation and / or synthesis of (particularly realistic) synthetic images that include appearance and / or context information of the textual prompt.
[0046] The generated text-to-image model may include an encoder and a decoder (particularly with skip connections), and optionally one or more intermediate blocks. For example, the generated text-to-image model may include a U-net architecture.
[0047] The image conditioning model may include an encoder, optionally including one or more intermediate blocks, followed by a zero convolution layer. The zero convolution layer may include a (e.g., 1×1) convolution layer in which both weights and biases are initialized to zero.
[0048] The encoder of the image conditioning model can be a (particularly trainable) copy of the encoder of the generated text-to-image model. The output of the optional one or more intermediate blocks and the zero convolutional layer can be fed (e.g., via a cross-attention mechanism) to the decoder of the optional one or more intermediate blocks and the generated text-to-image model.
[0049] The image conditioning model may encode and / or transform the visual information related to the body (i.e. the received 2D skeletal representation, 2D projection (in particular dense) semantic encoding and 2D depth map) so that it can be fed into the text-to-image model (e.g. using a cross-attention mechanism for some layers (in particular decoder layers) of the text-to-image model), thereby enabling the text-to-image model to produce (in particular realistic) synthetic image data that is consistent with the generated body model underlying the visual information. Alternatively or additionally, conditioning the generation of the synthetic image (in particular by the generated text-to-image model) by the encoded (in particular by the image conditioning model) visual information advantageously provides (in particular realistic) synthetic image data comprising the body as well as the ground truth data.
[0050] The image conditioning model may be configured to perform encoding on the received visual information in conjunction with (and / or based on) textual cues related to the appearance of the body and / or environmental information related to the body.
[0051] The generated text-to-image model may start "from scratch", e.g. starting from pure noise as image input (also denoted: corrupted image). Furthermore, the generated text-to-image model may receive textual cues relating to the appearance of the body and / or environmental information related to the body. The image conditioning model may receive visual information comprising a 2D skeletal representation, a 2D projected (in particular dense) semantic encoding and a 2D depth map of the body, e.g. concatenated along the channel dimension. Furthermore, the image conditioning model may receive (in particular the same as received by the generated text-to-image model) image input (e.g. corrupted image and / or pure noise) and / or textual cues relating to the appearance of the body and / or environmental information related to the body.
[0052] The output of the generated text-to-image model (particularly conditioned by the output of the image conditioning model) can include a prediction of noise (and / or pure noise) added to the clean image to create the corrupted image. The (e.g., true and / or clean) image output can be determined (e.g., calculated) based on the predicted noise.
[0053] By combining conditioning on 2D skeletal representations, 2D projected (especially dense) semantic encodings, and 2D depth maps (all of which are included in the received visual information), the photorealistic accuracy of synthesized image data can be significantly improved, and thereby the quality of downstream training.
[0054] The generated text-to-image model may include a diffusion model, in particular a stable diffusion (SD) network. The SD network may include a U-Net architecture with an encoder and a decoder (in particular with skip connections).
[0055] Diffusion models (also known as diffusion probability models and / or score-based generative models) may include machine learning (ML) models and / or generative models configured for image generation, image denoising, restoration and / or super-resolution through a diffusion process, in particular by using a forward process (e.g., adding noise, in particular Gaussian noise, to an image), a reverse process (e.g., predicting noise, in particular Gaussian noise in an image and compensating and / or subtracting accordingly) and a sampling process.
[0056] The diffusion model can be understood as treating the image generation process as a denoising task. Pure noise can be iteratively converted to real images. For this denoising task, a model can be trained, in particular a U-Net. It is possible to work directly in the image space and generate real images directly from pure noise. Alternatively or additionally, the latent diffusion model can perform (and / or conduct) the reverse process in the lower dimensional latent space of a variational autoencoder (VAE). Pure noise can be converted into a latent embedding of an image, which must be decoded to generate a real image. The "latent space image can be converted back to an image in a non-latent (and / or image) space" part can be performed (and / or completed) by the decoder of the VAE. For image generation, there is no need to use the encoder of the VAE. Alternatively or additionally, the encoder of the VAE only needs to be used during the training of the diffusion model.
[0057] SD may use or include a latent diffusion model. SD may involve or include a NN architecture (also referred to as a model) for text-to-image (particularly diffusion) generation. SD aims to learn a diffusion process that generates (e.g., denoised) image datasets from a given probability distribution (and / or the distribution of a given latent embedding dataset). The NN architecture of SD may include a transformer and / or a U-Net having an encoder and decoder with skip connections (e.g., between layers of the encoder and decoder) and / or a cross-attention mechanism (e.g., within layers of the encoder and / or decoder). Alternatively or additionally, the U-Net of SD may include an attention layer.
[0058] The NN architecture of the SD can further include a diffusion process by which the (e.g., noisy and / or corrupted) input image is converted into a latent space image. The latent space image can be converted back into an image in a non-latent (and / or image) space (e.g., converted into a de-noised and / or clean image) via a transformer and / or U-Net.
[0059] SD can advantageously produce high-quality (eg, high-resolution and / or realistic) synthetic images based on textual cues at low computational cost.
[0060] Alternatively or additionally, another type of diffusion model may operate in image space (e.g., as described by the Google Research Brain Team in “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv:2205.11487,
[13] , which is incorporated herein by reference, for which no model weights are published but an open source copy exists). In general, the techniques described herein for generating synthetic image data may also be applicable to other types of diffusion models.
[0061] Alternatively or additionally, the text-to-image model may include DALL·E or one of its successor models (e.g., DALL·E 2), and / or any text-conditional image generation based on CLIP (e.g., as described by A. Ramesh et al. in “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv:2204.06125 [cs.CV],
[14] , which is incorporated herein by reference).
[0062] The image conditioning model may include ControlNet. ControlNet may include an encoder and convolutional layers with a cross-attention mechanism with the generated text-to-image model (particularly the encoder of the SD network).
[0063] The generated models, and in particular the SD network, can be conditioned by predetermined features, in particular by using ControlNet. For the purpose of conditioning, reference is made to L. Zhang et al., “Adding Conditional Control to Text-to-Image Diffusion Models” (2023), available at Computer Vision Foundation and IEEE Xplore, which is incorporated herein by reference.
[0064] The connection between SD-Network and ControlNet can provide particularly computationally efficient regulation, which saves time and (eg, graphics processing unit, GPU) memory.
[0065] The combination of SD-Net and ControlNet can provide a particularly rich variety of (especially realistic) generated synthetic image data, including bodies that are consistent with the generated body models.
[0066] The ControlNet can be initialized using weights from the deep ControlNet by Zhang et al.
[11] . The main difference is that the ControlNet is fine-tuned on data (particularly body detection related tasks) using input (particularly received visual information related to the body, and / or textual cues related to the body's appearance and / or body-related environmental information) according to the techniques described in this paper for generating synthetic image data, to allow for better control over appearance and background than the original ControlNet.
[0067] Regarding the second method aspect, a computer-implemented method for training a conditional image synthesis model to generate synthetic image data for body detection-related tasks based on sensor data is provided, wherein the synthetic image data can be used to train, validate, and / or test downstream AI, particularly downstream NN. The method comprises the step of receiving a synthetic image training dataset. The synthetic image training dataset may include a (particularly 2D) image dataset of a body, which is annotated by a set of 3D body labels applied exogenously to the body when the image dataset is acquired; a (particularly 2D) image dataset of a body annotated by a human expert, wherein the annotations include 3D body pose information; and / or a (particularly 2D) image dataset of a body annotated by 2D keypoint information. The method further comprises the step of training the conditional image synthesis model based on the received synthetic image training dataset. Training the conditional image synthesis model comprises converting the annotations of the synthetic image training dataset into at least a portion of visual information related to the body. The visual information comprises a 2D skeletal representation of the body, a 2D projective (particularly dense) semantic encoding of the body, and a 2D depth map of the body. In particular, the 2D image dataset is regarded as the ground truth for the generated synthetic image data.
[0068] Training of the conditional image synthesis model may further include generating 3D pose information, in particular generating 3D pose information based on 2D keypoint information of annotations of a (in particular 2D) image dataset and / or as an intermediate step before converting the annotations into at least part of visual information related to the body.
[0069] The (particularly 2D) image dataset may comprise real images of the body. Alternatively or additionally, the 2D image dataset may comprise synthetic images of the body.
[0070] A dataset of (especially 2D) images can correspond to the ground truth for the training of conditional image synthesis models.
[0071] Training can be applied to a combination of the generated text-to-image model and the image conditioning model. The image conditioning model (e.g., ControlNet) can be trained specifically using a standard denoising loss function. ControlNet can also take a text prompt as input in order to predict noise, as illustrated in Equation (1).
[0072] By training the conditional image synthesis model, the weights of the nodes (e.g., the hidden layers of NN, especially ControlNet) can be learned.
[0073] The (particularly accurate) 3D body pose may be included by a collection of 3D body markers and / or by annotations (also called labels) performed by human experts.
[0074] The 3D pose information generated from the annotated 2D keypoint information can be represented as pseudo 3D labels.
[0075] The training of the conditional image synthesis model can be monitored by quality control metrics. The quality control metrics can include 3D (e.g., human) pose metrics, particularly mean per-joint position error (mpjpe), Procrustes aligned mean per-joint position error (pa-mpjpe), and / or percentage of correct keypoints (pck). For any of the 3D (e.g., human) pose metrics, a distance threshold can be selected.
[0076] Alternatively or additionally, the quality control metric may include Frechet Inception Distance (FID) and / or an image similarity metric, in particular Learned Perceptual Patch Similarity (LPIPS).
[0077] With respect to a third method aspect, a computer-implemented method for training, validating and / or testing a downstream AI, in particular a downstream NN, for performing body detection related tasks based on sensor data is provided. The method comprises the step of receiving synthetic image data of a body. The synthetic image data of the body may be generated according to the first method aspect. The synthetic image data of the body comprises a generated body model, visual information related to the body and / or related quantities. The visual information may comprise a 2D skeletal representation of the body, a 2D projected (in particular dense) semantic encoding of the body and a 2D depth map of the body. Alternatively or additionally, the related quantities may comprise information about the body derived from the generated body model and / or from the visual information. The method further comprises the step of training the downstream AI, in particular the downstream NN, based on the received synthetic image data of the body. Any of the generated body model, the visual information related to the body and / or the related quantities may be regarded as basic facts.
[0078] The synthetic dataset data provided by this technology can be used for training, benchmarking, validating and / or testing downstream AI (particularly downstream NN). In particular, the use of synthetic image datasets can enable the identification of systematic errors and / or biases of downstream AI (particularly downstream NN).
[0079] By using synthetic image data of the body and ground truth based on visual information, its underlying generated body model and / or information derived therefrom to train, validate and / or test downstream AI, in particular downstream NN, the downstream AI, in particular downstream NN, can be improved in performing body detection related tasks for a wide variety of bodies and environments. Alternatively or additionally, the speed and / or convergence of training downstream AI (in particular downstream NN) can be improved.
[0080] Performing body detection related tasks can be based on the received sensor data. The tasks can include classification, semantic segmentation and / or object detection, particularly body detection.
[0081] The detection of bodies may include in particular the detection of road users, for example pedestrians and / or (in particular wild) animals, such as wild boars or deer.
[0082] The sensor data may be received from a (eg, video) camera, a radar sensor, a LiDAR sensor, an ultrasonic sensor, a motion sensor, and / or a thermal image sensor. The sensor data may alternatively or additionally include a sensed seatbelt position of a driver of the vehicle.
[0083] The method according to the third method aspect may be used for applying downstream AI, in particular downstream NN, to automated (in particular autonomous) driving, planning the movements of a robot, operating household appliances and / or controlling access control systems.
[0084] The application may include video surveillance and / or motion capture (e.g., receiving sensor data therefrom). Alternatively or additionally, the application may include assisted driving, autonomous driving, smart home (and / or household) devices, (e.g., robotic) personal assistants, and / or operating technical equipment (e.g., robots, power tools, and / or manufacturing machines) in a manufacturing environment. Further alternatively or additionally, the application may include security systems, such as for anti-theft, access control, and / or monitoring of the driver of a vehicle.
[0085] Downstream AI (particularly downstream NN) may include receiving sensor data as input in an inference phase and providing task-specific output, particularly including determined classification and / or 3D body pose.
[0086] Regarding a first device aspect, a computing device for generating synthetic image data for use in body detection-related tasks based on sensor data is provided, wherein the synthetic image data can be used to train, validate, and / or test downstream AI, particularly downstream NN. The computing device can be configured to perform any of the steps disclosed in the context of the first method aspect, and / or include any of the features disclosed in the context of the first method aspect.
[0087] Regarding a second device aspect, a computing device for training a conditional image synthesis model to generate synthetic image data for use in body detection-related tasks based on sensor data is provided, wherein the synthetic image data can be used to train, validate, and / or test downstream AI, particularly downstream NN. The computing device can be configured to perform any of the steps disclosed in the context of the second method aspect, and / or include any of the features disclosed in the context of the second method aspect.
[0088] Regarding a third device aspect, a computing device for training, validating, and / or testing a downstream AI, particularly a downstream NN, for performing body detection-related tasks based on sensor data is provided. The computing device may be configured to perform any of the steps disclosed in the context of the third method aspect, and / or include any of the features disclosed in the context of the third method aspect.
[0089] Regarding the system aspect, a system for training, verifying, testing and / or applying downstream AI (particularly downstream NN) is provided. The system comprises: a computing device according to the first device aspect, configured to generate synthetic image data; a computing device according to the second device aspect, configured to use the generated synthetic image data to train a conditional image synthesis model; and a computing device according to the third device aspect, configured to train, verify and / or test the downstream AI (particularly downstream NN). The system further comprises: a downstream AI (particularly downstream NN) configured to perform body-related detection tasks; and at least one sensor and / or image capture device configured to provide sensor data, based on which the body-related detection tasks are performed.
[0090] With respect to a further aspect, a computer program product is provided. The computer program product comprises program elements which, when loaded into a memory of a computing device, cause the computing device to perform the steps of the method according to the first, second and / or third method aspect.
[0091] With respect to a still further aspect, there is provided a computer-readable medium having stored thereon program elements that can be read and executed by a computing device so as to perform the steps of the method according to the first, second and / or third method aspects when the program elements are executed by the computing device. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 Flowchart of a method for generating synthetic image data for body detection related tasks based on sensor data, which synthetic image data can be used to train, validate and / or test downstream AI, especially downstream NN.
[0093] Figure 2 Flowchart of a method for training a conditional image synthesis model to generate synthetic image data for body detection related tasks based on sensor data, which synthetic image data can be used to train, validate and / or test downstream AI, particularly downstream NN.
[0094] Figure 3 Flowchart of a method for training, validating and / or testing downstream AI, particularly downstream NN, for performing body detection related tasks based on sensor data.
[0095] Figure 4 This is an overview of the structure and architecture of a computing device for generating synthetic image data for body detection related tasks based on sensor data, which synthetic image data can be used to train, validate and / or test downstream AI, especially downstream NN.
[0096] Figure 5 This is an overview of the structure and architecture of a computing device for training a conditional image synthesis model to generate synthetic image data for body detection related tasks based on sensor data, which synthetic image data can be used to train, validate and / or test downstream AI, especially downstream NN.
[0097] Figure 6 Provides an overview of the structure and architecture of a computing device for training, validating, and / or testing downstream AI, particularly downstream NN, for performing body detection-related tasks based on sensor data.
[0098] Figures 7A to 7D It is shown that conventional conditioning using only a 2D depth map and a 2D skeletal representation results in the generation of a composite image that is inconsistent with the conditioning and / or inconsistent with the textual prompt "Man in coat in city".
[0099] Figures 8A to 8E We show a first example of combining a 2D skeletal representation, a 2D (especially dense) semantic encoding, and a 2D depth map as conditions with a textual cue “man in coat on stairs”, which leads to consistent and realistic synthetic imaging data.
[0100] Figures 9A to 9E A second example is shown where a 2D skeletal representation, a 2D (especially dense) semantic encoding, and a 2D depth map are combined as conditions with a textual cue “man wearing a coat in the city”, which leads to consistent and realistic synthetic imaging data.
[0101] Figure 10 Shown is an example architecture of a conditional image synthesis model that leverages a combination of stabilized diffusion and ControlNet.
[0102] Figures 11A to 11C Examples of downstream applications are schematically illustrated. DETAILED DESCRIPTION
[0103] Figure 1 A flow chart of a computer-implemented method 100 for generating synthetic image data for body detection related tasks based on sensor data is schematically illustrated, which synthetic image data can be used for training, validating and / or testing downstream AI, in particular downstream NN.
[0104] Method 100 includes a step S102 of receiving visual information related to a body. The visual information includes a two-dimensional (2D) skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body. Method 100 further includes a step S104 of receiving textual cues related to the appearance of the body and / or environmental information related to the body. Method 100 still further includes a step S106 of generating synthetic image data of the body based on the received S104 textual cues conditioned by the received S102 visual information. Generating S106 is performed by a conditional image synthesis model.
[0105] Optionally, the method 100 includes a step S101 of determining visual information related to the body using the generated body model, the visual information including a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body. Determining the visual information S101 may include performing a projection of the generated body model onto an image plane.
[0106] Further optionally, method 100 includes a step S108 of providing the generated S106 synthetic image data of the body to a downstream AI, particularly a downstream neural network. The downstream AI, particularly the downstream neural network, may be configured for body detection-related tasks based on the sensor data. Providing the S106 synthetic image data generated S108 may include providing the visual information, a set of shape parameters and pose parameters of the generated body model, the generated body model based on which the visual information is determined, and / or one or more related quantities as ground truth for the body detection-related tasks.
[0107] Figure 2 A flowchart of a computer-implemented method 200 for training a conditional image synthesis model to generate synthetic image data for body detection-related tasks based on sensor data is schematically illustrated, which synthetic image data can be used to train, validate and / or test downstream AI, in particular downstream NN.
[0108] Method 200 includes a step S202 of receiving a synthetic image training dataset. The synthetic image training dataset may include a (particularly 2D) image dataset of a body, which is annotated by a set of 3D body markers applied exogenously to the body when the image dataset is acquired. Alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated by a human expert, wherein the annotations include 3D body pose information. Further alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated with 2D keypoint information. In the case of the (particularly 2D) image dataset of the body, the training of the conditional image synthesis model includes generating 3D pose information.
[0109] Method 200 further includes a step (S204) of training a conditional image synthesis model based on the synthetic image training dataset received (S202). Training (S204) the conditional image synthesis model includes converting annotations of the synthetic image training dataset into at least a portion of visual information related to the body. The visual information includes a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body. The (particularly 2D) image dataset is used as ground truth for generating synthetic image data.
[0110] Figure 3 A flow chart of a computer-implemented method 300 of training, validating and / or testing a downstream AI, in particular a downstream NN, for performing body detection related tasks based on sensor data is schematically illustrated.
[0111] Method 300 includes a step S302 of receiving synthetic image data of a body. The synthetic image data of the body is generated according to method 100. The synthetic image data of the body includes a generated body model, visual information related to the body, and / or related quantities. The visual information includes a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body. Alternatively or additionally, the related quantities include information about the body derived from the generated body model and / or from the visual information. Method 300 further includes a step S304 of training a downstream AI, particularly a downstream NN, based on the synthetic image data of the body received S302. The generated body model, visual information related to the body, and / or related quantities are regarded as basic facts.
[0112] Figure 4 The architecture of a computing device 400 for generating synthetic image data for body detection related tasks based on sensor data is schematically illustrated, which synthetic image data can be used to train, validate and / or test downstream AI, in particular downstream NN.
[0113] The computing device 400 includes a first input interface 402, which is configured to receive visual information related to the body. The visual information includes a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body. The computing device 400 further includes a second input interface 404, which is configured to receive textual prompts related to at least one of the appearance of the body and / or environmental information related to the body. The computing device 400 still further includes a generation module 406, which includes a conditional image synthesis model. The conditional image synthesis model is configured to generate synthetic image data of the body based on the received textual prompts conditioned by the received visual information.
[0114] Optionally, the computing device 400 includes a determination module 401 configured to determine visual information related to the body using the generated body model, the visual information including a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body. Determining the visual information may include projecting the generated body model onto an image plane.
[0115] Further optionally, computing device 400 includes an output interface 408 configured to provide the generated synthetic image data of the body to a downstream AI, particularly a downstream neural network. The downstream AI, particularly the downstream neural network, may be configured to perform body detection-related tasks based on the sensor data. Providing the generated synthetic image data may include providing visual information, a set of shape parameters and pose parameters of a generated body model, the generated body model based on which the visual information is determined, and / or one or more related quantities as ground truth for the body detection-related tasks.
[0116] Any of the first input interface 402, the second input interface 404, and the optional output interface 408 can be implemented by an input-output interface 410. Alternatively or additionally, the generation module 406 and / or the optional determination module 401 can be implemented by a processing unit. Further alternatively or additionally, the computing device 400 can include at least one memory 414.
[0117] Figure 5The architecture of a computing device 500 for training a conditional image synthesis model to generate synthetic image data for body detection related tasks based on sensor data, which synthetic image data can be used to train, validate and / or test downstream AI, in particular downstream NN.
[0118] The computing device 500 includes an input interface 502 configured to receive a synthetic image training dataset. The synthetic image training dataset may include a (particularly 2D) image dataset of a body, which is annotated by a set of 3D body markers applied exogenously to the body when the image dataset is acquired. Alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated by a human expert, wherein the annotations include 3D body pose information. Further alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated by 2D keypoint information. In particular in the case of a (particularly 2D) image dataset of a body, the training of the conditional image synthesis model includes, for example, generating the image by a 3D pose information generation module ( Figure 5 The computing device 500 further includes a training module 504 configured to train a conditional image synthesis model based on the received synthetic image training dataset. Training the conditional image synthesis model includes converting annotations of the synthetic image training dataset into at least a portion of visual information related to the body. The visual information includes a 2D skeletal representation of the body, a 2D projective (particularly dense) semantic encoding of the body, and a 2D depth map of the body, and the (particularly 2D) image dataset is considered as the ground truth for the generated synthetic image data.
[0119] The input interface 502 may be implemented by an input-output interface 506. Alternatively or additionally, the training module 504 may be implemented by a processing unit 508. Further alternatively or additionally, the computing device 500 may include at least one memory 510.
[0120] Figure 6 The architecture of a computing device 600 for training, validating and / or testing a downstream AI, in particular a downstream NN, for performing body detection related tasks based on sensor data is schematically illustrated.
[0121] The computing device 600 includes an input interface 602, which is configured to receive synthetic image data of a body. The synthetic image data of the body is generated according to method 100. The synthetic image data of the body includes a generated body model, visual information related to the body, and / or related quantities. The visual information includes a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body, and a 2D depth map of the body. Alternatively or additionally, the related quantities include information about the body derived from the generated body model and / or from the visual information. The computing device 600 further includes a training module 604, which is configured to train downstream AI, particularly downstream NN, based on the received synthetic image data of the body. The generated body model, visual information related to the body, and / or related quantities are regarded as basic facts.
[0122] The input interface 602 may be implemented by an input-output interface 606. Alternatively or additionally, the training module 604 may be implemented by a processing unit 608. Further alternatively or additionally, the computing device 600 may include at least one memory 610.
[0123] Any of the processing units 412; 508; 608 may be implemented by a central processing unit (CPU) and / or a graphics processing unit (GPU).
[0124] Through the techniques presented in this paper, synthetic image data can be generated to benchmark and improve tasks such as 3D human pose estimation from a single RGB image. The methods described can be based on deep learning, which typically requires large amounts of annotated data to be effective. However, obtaining (especially real) image data with accurate 3D body pose annotations is a difficult process.
[0125] Using the techniques presented herein, image data of (particularly living) people (and / or animals) can be generated under conditions where the 3D body pose of the people (and / or animals) as well as their appearance and position are controlled. The generated models (particularly including conditional image synthesis models) can be used to create (particularly synthesized image) data for fine-grained evaluation of 3D (e.g., human) pose estimators, in particular to benchmark their robustness and generalization capabilities under different conditions, and to identify their systematic errors and biases. Alternatively or additionally, the generated models (particularly conditional image synthesis models) can also be used to generate training data for 3D human pose estimators.
[0126] The technology presented herein enables text-based image generation of a person (and / or animal) with control of their 3D body pose. In particular, text control and 3D body pose control can be separated.
[0127] Previous text-to-image synthesis methods (such as ControlNet
[11] , which is incorporated herein by reference) are only able to control 2D pose while maintaining fine-grained text control. While it is possible to control 3D (e.g., human) pose using ControlNet generated with depth conditions, its ability to control image content through text is generally severely limited, e.g., Figure 7A 、 7B , 7C and 7D as exemplarily illustrated.
[0128] Figure 7A 、 7B , 7C and 7D exemplarily illustrate the shortcomings of conventional ControlNet
[11] (particularly using at most two visual information, such as 2D human pose conditions and a 2D depth map) compared to techniques for generating synthetic image data using visual information, which includes a triple of a 2D skeleton representation, a 2D projection (particularly dense) semantic encoding and a 2D depth map, as well as a textual prompt (particularly which can be received by both a text-to-image model and an image conditioning model). Figure 7A and 7B The depth condition and the 2D human pose condition are shown respectively, based on which the Figure 7C and 7D Composite image in . Figure 7A The depth image represents the 3D body pose, and Figure 7B The key points in a (eg, including a small number of major joints such as shoulder, elbow, wrist, hip, knee, and ankle) represent their 2D projections. Figure 7C and 7D The text hint for the example is "Man wearing coat in city".
[0129] Figure 7C The image can be generated, for example, by utilizing the same conditional image synthesis model as employed in the technique for generating synthetic image data using triplets of visual information and textual cues, but lacking the 2D (particularly dense) semantic encoding (and / or not using textual cues in the image conditioning model). In particular, a conventional ControlNet-depth Figure 7D Note that the pose of the person does not match the 3D condition. For example, Figure 7C The left arm of the person in the figure is close to the body, but should be pointing forward, and the right hand of the person is holding the railing tightly, and according to Figure 7A The depth image of the person should be open and positioned with the thumb pointing up. In addition, as specified in the text prompt Figure 7DIn the example, deep conditional generation fails when generating a coat. The conditional image synthesis model employed in the technique for generating synthetic image data described in this paper follows the textual cues better than the conventional ControlNet-depth. In particular Figure 7D Neither the coat nor the city is shown.
[0130] An alternative approach is provided by computer graphics-based methods [1], which are incorporated herein by reference. However, pipelines are difficult to set up and assets must be manually designed and created by humans, which limits their overall diversity. In contrast, the techniques presented in this paper are easy to use and are able to generate a diverse dataset of synthetic images that are consistent with their 3D pose ground truth, notably without manual labor.
[0131] Image diffusion models learn to progressively denoise an image to generate samples. Denoising can occur in pixel space or in a “latent” space encoded from training data. Stable Diffusion (SD) uses the latent image as the training domain. In this context, the terms “image”, “pixel”, and “denoising” all refer to the corresponding concepts in the “perceptual latent space” [8], which is incorporated herein by reference. Given an image z0, the diffusion algorithm progressively adds noise to the image and produces a noisy image z t , where t is the number of times noise is added. When t is large enough, the image is close to pure noise. Given a set of conditions, including time step t, text prompt c t (especially related to the appearance of the body and / or environmental information), as well as task-specific conditions c f (especially including visual information, which includes a 2D skeleton representation of the body, 2D projections, especially dense semantic encoding and 2D depth map), the image diffusion algorithm learns a network ∈ θ To predict the noise image z t The noise,
[0132]
[0133] Where L is the overall learning objective of the entire diffusion model. This learning objective can be directly used to fine-tune the diffusion model.
[0134] An embodiment of this technique builds upon ControlNet
[11] and introduces 3D pose control into SD. A 3D mesh of an exemplary human body is encoded using the SMPL model[5]. SMPL is a generative model that decomposes the human body into shape (e.g., how an individual varies in height, weight, and / or body proportions) and pose (e.g., how a 3D surface deforms with joints). Shape β∈R 10 The pose θ∈R is parameterized by the first 10 coefficients of the PCA shape space.3K The relative 3D rotation is modeled by K = 23 joints in axis-angle representation. SMPL is a differentiable function that outputs a triangular mesh M(θ, β)∈R with N = 6980 vertices. 3N , which is obtained by shaping the template body vertices conditioned on β and θ, then connecting the bones via forward kinematics according to the joint rotation θ, and finally deforming the surface using linear blend skinning. The 3D keypoints X(θ, β)∈R are obtained by linear regression from the final mesh vertices 3P , where P (eg, P=24) is the number of 3D joints in the skeleton.
[0135] The techniques presented herein can be used, particularly in downstream AI applications, by operating on digital (and / or analog) image data that can be obtained by receiving sensor signals, such as video, RGB cameras, radar, LiDAR, ultrasound, motion, and / or thermal images for computer control of machines such as robots, (e.g., automated) vehicles, household appliances, power tools, manufacturing machines, personal assistants, access control systems, systems for transmitting information (e.g., surveillance systems), and / or medical (imaging) systems. The downstream AI achieves this by performing tasks related to body detection (e.g., classifying sensor data, detecting the presence of objects in sensor data) and / or performing semantic segmentation on sensor data (e.g., with respect to pedestrians).
[0136] The techniques described herein comprise the upstream portion of a machine learning (ML) and / or AI toolchain. The techniques for generating synthetic image data (and / or the models generated upstream) need not directly but can indirectly improve the ML systems (and / or downstream AI, also referred to as downstream generated models) that can be used for the aforementioned applications. In particular, method 100 generates training data so that method 300 generates test data to check whether the trained ML system can then operate safely.
[0137] The techniques described in this paper can be used for data augmentation as well as domain conversion tasks, such as from synthetic images to realistic images. The generated samples can be used to evaluate and / or train any data-driven method, such as pedestrian detection models (especially as a body detection-related task).
[0138] According to an exemplary embodiment, a conditional image synthesis model based on SD[8] and ControlNet
[11] is trained to generate synthetic image data of a person (and / or animal) under control of the 3D body pose of the person (and / or animal). In addition, control of the image content (e.g., location, appearance of the person, and / or weather) is provided through textual prompts. Subsequently, the conditional image synthesis model is used to generate images for system evaluation of a 3D human pose estimator. This is possible precisely because the techniques described herein have control over the 3D body pose by means of a (e.g., conditional image synthesis) model.
[0139] Figures 8A to 8E 9A to 9E show examples of combining visual information (also called pose conditions) as conditions, i.e., 2D skeleton representations (also represented as 2D keypoints, Figure 8A and 9A ), 2D projection (especially dense) semantic encoding ( Figure 8B and 9B ) and 2D depth map ( Figure 8C and 9C ).
[0140] Figure 8D and 8E The generated synthetic image data in is used as the text prompt "Man in coat on stairs". Figure 9D and 9E The generated synthetic image data in is used as the text prompt "Man wearing coat in the city".
[0141] An exemplary conditional image synthesis model is built upon ControlNet with explicit 3D conditioning, which captures the 3D structure of the human body and its semantics.
[0142] Figure 10 Schematically illustrates the combination of an SD encoder 1004-E, a corresponding SD decoder 1004-D, and a ControlNet 1002 having a trainable copy 1002-E of the SD encoder 1004-E (specifically frozen after initial training). The ControlNet architecture 1002 controls the SD model (which includes the SD encoder 1004-E and the SD decoder 1004-D).
[0143] ControlNet 1002 includes several zero convolution layers 1002-ZC. Figure 10, as an illustrative example, one zero convolutional layer 1002-ZC before the trainable encoder 1002-E and several zero convolutional layers 1002-ZC after the trainable encoder 1002-E are shown. Any other arrangement of zero convolutional layers 1002-ZC (e.g., having more or fewer zero convolutional layers 1002-ZC before and / or after the trainable encoder 1002-E) may be possible.
[0144] The visual information at reference symbol 1010 (also denoted as visual image conditions) is collectively denoted as c f ={c d , c dp , c k} and fed into ControlNet (e.g., concatenated along the channel dimension). The visual information (and / or conditioning) consists of three parts: a depth map of the (e.g., human) body c d , (especially dense) semantic encoding of the (e.g., human) body dp and 2D skeleton representation (abbreviated as unskeleton) c k The depth map provides 3D pose information.
[0145] Before applying the first zero convolution 1002-ZC, the (eg, concatenated) visual information c at reference symbol 1010 f Through multiple layers ( Figure 10 not explicitly shown) for conversion.
[0146] The textual prompt c at reference symbol 1012 is related to the appearance of a (eg, human) body and / or environmental information. t For example, it is fed into the SD encoder 1004-E and the SD decoder 1004-D through a cross-attention mechanism.
[0147] exist Figure 10 In the illustrative example of FIG, the text prompt c at reference symbol 1012 t is additionally fed into ControlNet 1002 .
[0148] Furthermore, the first latent image z at reference symbol 1006 t (also denoted as the corrupted image at time step t, and / or denoted as the noisy latent embedding of the corrupted image at time step t) is used to initialize SD, where the image z at reference symbol 1008 is t-1 (also denoted as the corrupted image at time step t-1) is provided as a (eg, indirect and / or derived) output. The (eg, direct) output of the SD decoder 1004-D may include the addition to the latent image z t The noise prediction of the image z can be determined (e.g., calculated) based on the prediction. t-1.
[0149] In an exemplary embodiment, visual information is generated by rendering a posed SMPL mesh. Given a human pose θ including axis-angle representations of joint rotations, a parameterized body model M is used to infer a 3D mesh of the human body. The depth of the rendered mesh provides a representation c d . A (particularly dense) semantic encoding is created by assigning the 3D position of each vertex of the mesh in T-pose as a color. This provides semantic information about different body parts and allows ControlNet to distinguish, for example, left and right, front and back. Finally, 2D keypoints allow the use of datasets that only provide 2D keypoint annotations. This increases the overall diversity of the generated synthetic image data (particularly as training data for downstream AI).
[0150] To train conditional image synthesis models, a mixture of datasets can be used. Datasets with accurate 3D pose annotations provide the image data and 3D pose pairs required for accurate generation. However, datasets with accurate 3D pose annotations typically lack diversity in environment and appearance (e.g., human or animal bodies). Therefore, the set of training datasets can be supplemented with datasets that only provide 2D keypoint annotations. Using (e.g., human) mesh recovery methods, pseudo 3D labels for 2D datasets can be determined (e.g., calculated), and alignment with 2D keypoints can be used as a quality control metric to prune the data.
[0151] Figure 11A 、 11B 11C and 11D show exemplary downstream applications for an automated (e.g., autonomous) driving vehicle 1102-1, a (particularly autonomous) robot 1102-2, and an access control system 1102-3. Each of the downstream applications includes a controller 1104, on which a corresponding downstream AI is installed and trained, validated, and / or tested according to method 300. Figure 11C Two types of sensors that may be used for access control are further exemplarily shown, namely a camera 1110 and a microphone 1112 .
[0152] To evaluate a 3D pose estimator, the pose can be derived from a public 3D pose estimation benchmark, such as 3DPW
[10] , which is incorporated herein by reference. Performance on generated synthetic image data may be degraded due to distribution shift in the pixel value distribution. Therefore, for system evaluation, a synthetic replica of the 3D pose estimation benchmark is created. The goal is to create a synthetic clone, consisting of an image that is as close to the original image as possible with respect to image content. A visual question answering (VQA) model can be used to extract the content of an image represented as text. The extracted content can be used to generate images and / or synthetic replicas. All experiments on robustness are performed using this replica as a baseline to eliminate distribution shift as a factor.
[0153] An exemplary general setup for experiments on the robustness of a pose estimator to a certain attribute (e.g., information about the appearance of the body and / or the environment) or a set of attributes can take the following form: the textual prompt used to generate the replica can be used (e.g., reused) and the selected attribute or set of attributes can be introduced. Then, starting from the same noise used to generate the base image, new synthetic image data is generated. In this way, the general structure and / or content is preserved and only the expected attributes are introduced to the synthetic image data, which allows for more meaningful comparisons. Furthermore, techniques such as prompt-to-prompt [2] can be used to preserve structure even better.
[0154] For evaluation, clothing, location, lighting, weather, age, race, and / or gender can be targeted as attributes. In particular, (e.g., real) image data with controlled lighting, weather, location, and clothing attributes is often difficult to obtain due to their large variability in the real world. Considering combinations of attributes is also important for evaluation.
[0155] Each attribute may appear a sufficient number of times in the training data. However, some combinations of attributes may never appear in the (especially real) training data and thus lead to a degradation in performance. Exhaustively testing all combinations of attributes is infeasible due to the exponential growth of the result set. For this purpose, combinatorial testing [7] can be applied to select a subset of all attribute configurations. Common 3D human pose metrics such as mean per-joint position error (mpjpe), Procrustes aligned mean per-joint position error (pa-mpjpe) and / or percentage of correct keypoints (pck) [4, 9, 6] can be used at different distance thresholds. Alternatively or additionally, to measure the quality of the synthetic image data, Frechet inception distance (FID) [3] and / or image similarity metrics such as LPIPS
[12] can be used.
[0156] The techniques presented in this paper can be used for conditional image synthesis, for example, using diffusion models. In particular, data synthesis can be used for data augmentation and validation purposes of DL models, such as 2D / 3D pose estimators and / or pedestrian detectors. Using conditional image synthesis is particularly beneficial (and / or very possible) when collecting additional (e.g., real) data is expensive and / or legally impossible for privacy reasons.
[0157] Cited prior art
[0158] [1] Michael J. Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang, BEDLAM: A synthetic dataset of bodies exhibiting detailed lifelike animated motion. In Proceedings IEEE / CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 8726–8737, June 2023.
[0159] [2] Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or, Prompt-to-prompt image editing with cross attention control, 2022.
[0160] [3] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, Günter Klambauer, and Sepp Hochreiter, Gans trained by a two time-scale update rule converge to a nash equilibrium. CoRR, abs / 1706.08500, 2017.
[0161] [4] Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik, End-to-end recovery of human shape and pose. CoRR, abs / 1712.06584, 2017.
[0162] [5] Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black, SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, October 2015.
[0163] [5a]Silvia Zuffi, Angjoo Kanazawa, David W. Jacobs, and Michael J. Black, 3DMenagerie: Modeling the 3D shape and pose of animals. CoRR, abs / 1611.07700, 2016.
[0164] [6]Dushyant Mehta, Helge Rhodin, Dan Casas, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt, Monocular 3d human pose estimation using transfer learning and improved CNN supervision. CoRR, abs / 1611.09813, 2016.
[0165] [7]Changhai Nie and Hareton Leung, A survey of combinatorial testing. ACM Comput. Surv., 43(2), February 2011.
[0166] [8]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Ommer, High-resolution image synthesis with latent diffusion models, 2021.
[0167] [9]Alexander Toshev and Christian Szegedy, DeepPose: Human pose estimation via deep neural networks. CoRR, abs / 1312.4659, 2013.
[0168]
[10] Timo von Marcard, Roberto Henschel, Michael Black, Bodo Rosenhahn, and Gerard Pons-Moll, Recovering accurate 3d human pose in the wild using imus and a moving camera. In European Conference on Computer Vision (ECCV), September 2018.
[0169]
[11] Lvmin Zhang and Maneesh Agrawala, Adding conditional control to text-to-image diffusion models, 2023.
[0170]
[12] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang, The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.
[0171]
[13] Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi, Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. arXiv:2205.11487[cs.CV], 2022.
[0172]
[14] Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen, Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv:2204.06125 [cs.CV], 2022.
Claims
1. A computer-implemented method (100) for generating synthetic image data for use in body detection-related tasks based on sensor data, wherein the synthetic image data can be used for training, validating and / or testing a downstream AI, in particular a downstream neural network (NN), the computer-implemented method (100) comprising the following steps: - receiving (S102) visual information related to the body, wherein the visual information includes a two-dimensional 2D skeleton representation of the body, a 2D projection (especially dense) semantic encoding of the body, and a 2D depth map of the body; - receiving (S104) a text prompt related to at least one of the appearance of the body and / or environmental information related to the body; and - generating (S106) composite image data of the body based on the received (S104) textual prompts conditioned by the received (S102) visual information, wherein the generating (S106) is performed by a conditional image synthesis model.
2. The method (100) according to claim 1, further comprising the steps of: -Determining (S101) visual information related to the body through the generated body model, the visual information including a 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body and a 2D depth map of the body, wherein determining (S101) the visual information includes performing a projection of the generated body model onto an image plane.
3. The method (100) according to the preceding claim, wherein the generated body model generates a model of the body based on a set of shape parameters and pose parameters; optionally The generated body models include skinned multi-person linear SMPL models.
4. The method (100) according to any one of the preceding claims, further comprising the steps of: -Providing (S108) the generated (S106) synthetic image data of the body to a downstream AI, in particular a downstream NN, wherein the downstream AI, in particular the downstream NN is configured for body detection related tasks based on sensor data, and wherein providing (S108) the generated (S106) synthetic image data includes providing visual information, a set of shape parameters and posture parameters of a generated body model, a generated body model based on which the visual information is determined, and / or one or more related quantities as basic facts for the body detection related tasks.
5. The method (100) according to any one of the preceding claims, wherein the conditional image synthesis model comprises: a generated text-to-image model configured to generate ( S106 ) synthetic image data based on the received ( S104 ) text prompt; and an image conditioning model configured to encode the received (S102) visual information to condition and / or control the generated text-to-image model; optionally The image adjustment model is configured to encode the received (S102) visual information in combination with the received (S104) text prompt.
6. The method (100) according to claim 5, wherein the generated text-to-image model comprises a diffusion model, in particular a stabilized diffusion SD network, wherein the SD network comprises a U-Net architecture with an encoder and a decoder, in particular with skip connections.
7. The method (100) according to claim 5 or 6, wherein the image conditioning model comprises a ControlNet, wherein the ControlNet comprises an encoder and convolutional layers, wherein the convolutional layers have a cross-attention mechanism with the generated text-to-image model, in particular with the encoder of the SD network.
8. A computer-implemented method (200) for training a conditional image synthesis model to generate synthetic image data for use in body detection-related tasks based on sensor data, wherein the synthetic image data can be used to train, validate and / or test a downstream AI, in particular a downstream neural network (NN), the computer-implemented method (200) comprising the following steps: - receiving (S202) a synthetic image training dataset, wherein the synthetic image training dataset includes at least one of the following: - an image dataset of a body, in particular a 2D image dataset of a body, annotated by a set of 3D body landmarks applied exogenously to the body when acquiring the image dataset; - a dataset of images of the body, in particular a dataset of 2D images of the body, annotated by human experts, where the annotations include 3D body pose information; and - a dataset of images of the body, in particular a dataset of 2D images of the body, which are annotated with 2D keypoint information and wherein the training of the conditional image synthesis model further includes generating 3D pose information; as well as -Training (S204) a conditional image synthesis model based on the received (S202) synthetic image training dataset, wherein training (S204) the conditional image synthesis model comprises converting annotations of the synthetic image training dataset into at least a portion of visual information related to the body, wherein the visual information comprises a two-dimensional 2D skeletal representation of the body, a 2D projection (particularly dense) semantic encoding of the body and a 2D depth map of the body, and wherein in particular the 2D image dataset is considered as ground truth for the generated synthetic image data.
9. A computer-implemented method (300) for training, validating and / or testing a downstream AI, in particular a downstream neural network (NN), for performing body detection-related tasks based on sensor data, the computer-implemented method (300) comprising the following steps: - receiving (S302) composite image data of a body, wherein the composite image data of the body is generated according to any one of method claims 1 to 7, wherein the composite image data of the body comprises a generated body model, visual information related to the body, and / or related quantities, wherein the visual information comprises a two-dimensional 2D skeletal representation of the body, a 2D projected (in particular dense) semantic encoding of the body, and a 2D depth map of the body, and / or wherein the related quantities comprise information about the body derived from the generated body model and / or from the visual information; and - Training (S304) downstream AI, in particular downstream NN, based on received (S302) synthetic image data of the body, wherein the generated body model, visual information related to the body, and / or related quantities are regarded as ground truth.
10. The method (300) according to the preceding claim, wherein performing body detection related tasks is based on the received sensor data, wherein the tasks comprise at least one of classification, semantic segmentation and detection of objects, in particular detection of bodies.
11. The method (300) of claim 9 or 10, wherein the sensor data is received from at least one of a camera, a radar sensor, a LiDAR sensor, an ultrasonic sensor, a motion sensor, and a thermal image sensor.
12. Use of the method (300) according to any one of claims 9 to 11 for applying downstream artificial intelligence, in particular downstream neural networks, to at least one of the following: -Automated driving; - Plan the robot's movements; -operate household appliances; and -Control access control system.
13. A computing device (400) for generating synthetic image data for use in body detection-related tasks based on sensor data, wherein the synthetic image data can be used to train, validate and / or test a downstream AI, in particular a downstream neural network (NN), the computing device (400) comprising: - a first input interface (402) configured to receive visual information related to a body, wherein the visual information comprises a two-dimensional 2D skeletal representation of the body, a 2D projective (particularly dense) semantic encoding of the body, and a 2D depth map of the body; - a second input interface (404) configured to receive a text prompt related to at least one of the appearance of the body and / or environmental information related to the body; as well as - a generation module (406) comprising a conditional image synthesis model, wherein the conditional image synthesis model is configured to generate composite image data of the body based on received textual cues conditioned by received visual information.
14. A computing device (500) for training a conditional image synthesis model to generate synthetic image data for body detection-related tasks based on sensor data, wherein the synthetic image data can be used to train, validate and / or test downstream AI, in particular downstream neural network (NN), the computing device (500) comprising: - an input interface (502) configured to receive a synthetic image training dataset, wherein the synthetic image training dataset comprises at least one of: - an image dataset of a body, in particular a 2D image dataset of a body, annotated by a set of 3D body landmarks applied exogenously to the body when acquiring the image dataset; - a dataset of images of the body, in particular a dataset of 2D images of the body, annotated by human experts, where the annotations include 3D body pose information; and - a dataset of images of the body, in particular a dataset of 2D images of the body, which are annotated with 2D keypoint information and wherein the training of the conditional image synthesis model further includes generating 3D pose information; as well as - a training module (504) configured for training a conditional image synthesis model based on the received synthetic image training dataset, wherein training the conditional image synthesis model comprises converting annotations of the synthetic image training dataset into at least a portion of visual information related to the body, wherein the visual information comprises a two-dimensional 2D skeletal representation of the body, a 2D projective (in particular dense) semantic encoding of the body and a 2D depth map of the body, and wherein in particular the 2D image dataset is considered as ground truth for the generated synthetic image data.
15. A computing device (600) for training, validating and / or testing a downstream AI, in particular a downstream neural network (NN), for performing body detection related tasks based on sensor data, the computing device (600) comprising: - an input interface (602) configured to receive composite image data of a body, wherein the composite image data of the body is generated according to any one of method claims 1 to 7, wherein the composite image data of the body comprises a generated body model, visual information related to the body, and / or related quantities, wherein the visual information comprises a two-dimensional 2D skeletal representation of the body, a 2D projective (in particular dense) semantic encoding of the body and a 2D depth map of the body, and / or wherein the related quantities comprise information about the body derived from the generated body model and / or from the visual information; as well as - A training module (604) configured for training a downstream AI, in particular a downstream NN, based on received synthetic image data of a body, wherein the generated body model, visual information related to the body, and / or related quantities are regarded as ground truth.