Technique for generating synthetic image data for body detection-related tasks
By generating synthetic image data with 2D skeletal representations and depth maps, the challenge of limited 3D body pose annotations is addressed, improving AI model robustness and performance in diverse real-world scenarios.
Patent Information
- Application Number
- JP2025022550
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-15
- Filing Date
- 2025-02-14
- Publication Date
- 2025-08-27
AI Technical Summary
Acquiring images with accurate 3D body pose annotations is challenging, requiring dedicated capture studios and is time-consuming, limiting the diversity of scene locations, human appearances, and poses in datasets, which affects the performance and robustness of 3D human pose estimators in real-world situations.
Generating synthetic image data using a conditioned image synthesis model that incorporates 2D skeletal representations, dense semantic encoding, and 2D depth maps, along with text prompts, to create diverse and photorealistic images for training and validating downstream AI models for body detection tasks.
Improves the robustness and generalization of AI models by providing diverse training data, enhancing their performance in real-world scenarios and enabling efficient training, validation, and testing for body detection tasks.
Smart Images

Figure 2025125545000002 
Figure 2025125545000003 
Figure 2025125545000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to techniques for generating synthetic image data usable for training, validating, and / or testing downstream artificial intelligence (AI), particularly downstream neural networks (NNs), for body detection-related tasks, techniques for training a conditional image synthesis model to generate synthetic image data usable for training, validating, and / or testing downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data, and techniques for training, validating, and / or testing downstream AI, particularly downstream NNs, for performing body detection-related tasks based on sensor data. In particular, methods, computing devices, systems, computer program products, and computer-readable storage media are provided. [Background technology]
[0002] Background technology In applications such as automated (especially autonomous) driving, movement planning for (especially autonomous) robots, the operation of smart consumer electronics devices and / or the control of surveillance systems, it is of paramount importance to not only recognize the presence of living bodies (e.g., humans or animals), but also to estimate their three-dimensional (3D) body pose and / or their impending motion. In the context of automated (and / or autonomous) driving, this includes extreme situations such as imminent collisions with people or vehicles performing dangerous maneuvers or pedestrians. Summary of the Invention [Problem to be solved by the invention]
[0003] In principle, pre-trained AI models can be used to estimate 3D body pose from images. However, acquiring images with accurate 3D body pose annotations has traditionally been a challenging process, requiring dedicated capture studios and time-consuming data acquisition. This significantly limits the diversity of scene locations, human appearances, and poses represented in such datasets. Therefore, currently, evaluations of 3D human pose estimators are performed on datasets that do not represent the diverse situations occurring in the real world. Therefore, the performance and robustness of current state-of-the-art 3D human pose estimators in actual real-world situations remain questionable. [Means for solving the problem]
[0004] Disclosure of the Invention In the following, the solution according to the invention will be described in terms of the claimed methods and the claimed computing devices. Features, advantages or alternative embodiments herein may be assigned to other claimed subject matter (e.g., systems, computer programs or computer program products), and vice versa. In other words, a claim for a computing device may be modified by features described or claimed in the context of a method. In this case, functional features of the method are embodied by structural units of the respective computing device, and vice versa.
[0005] Regarding a first method aspect, a computer-implemented method is provided for generating synthetic image data usable for training, validation, and / or testing of downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data. The method includes receiving visual information related to a body. The visual information includes a two-dimensional (2D) skeletal representation of the body, a 2D projected (particularly dense) semantic encoding of the body, and a 2D depth map of the body. The method further includes receiving a text prompt related to the appearance of the body and / or environmental information relative to the body. The method further includes generating synthetic image data of the body based on the received text prompt conditioned by the received visual information. The generation is performed by a conditioned image synthesis model.
[0006] Techniques, including computer-implemented methods, provide improved robustness and generality of downstream artificial intelligence (AI, particularly downstream NNs) trained with provided downstream (particularly synthetic) training data. Alternatively or additionally, data augmentation and / or data enrichment (e.g., based on diversity of bodies, 3D body poses, appearances, environments, and / or locations) are provided to efficiently train, validate, and / or test downstream AIs (particularly downstream NNs), particularly when limited (particularly real) data exists. This can improve the generalization ability of the downstream AIs (particularly downstream NNs).
[0007] The techniques may further enable training, validating, and / or testing downstream AI (particularly downstream NNs) for domain transfer tasks (e.g., converting synthetic images to photorealistic images).
[0008] The downstream AI (particularly the downstream NN) may be configured for object (particularly including body) detection, (e.g., 3D) body pose detection, classification and / or semantic segmentation of image data received by sensors (e.g., video, cameras, radar sensors, LiDAR sensors, ultrasonic sensors, motion sensors and / or thermal sensors, among others).
[0009] The downstream AI (particularly the downstream NN) can be trained, validated and / or tested using synthetic image data (simply put, synthetic images), in particular taking into account the received body-related visual information (or any further body-related information determined based on the visual information) as ground truth.
[0010] The body may include a (particularly living) human body and / or a (particularly living) animal body.
[0011] The human and / or animal body may comprise a living being in a traffic situation, for example, for downstream applications of autonomous driving or driver assistance systems (collectively referred to as automated driving, levels L1 to L5, in particular the higher levels L3 to L5). Alternatively or additionally, the human and / or animal body may comprise a living being in a home automation environment.
[0012] The multiple visual information of the body, i.e., the 2D skeletal representation, the 2D (especially dense) semantic coding, and the 2D depth map, may be mutually consistent (and / or include mutual consistency conditions) in that they are associated with the same body and / or similar (e.g., 3D) body poses. Alternatively or additionally, the visual information related to the body may be determined based on input to a generative body model. The input to the generative body model may include a set of shape parameters and pose parameters related to the body. The output of the generative body model may include a representation of the generated body, particularly including a 3D skeleton (e.g., including major joints and / or bones, e.g., including ankle joints, knee joints, and hip joints for determining the position of the lower limbs) and a 3D surface representation (e.g., as a 3D mesh). The shape parameters may take into account body proportions, height, and / or weight. Alternatively or additionally, the pose parameters may take into account the 3D position (and / or articulations) of the body's skeleton (and / or particularly the positions of the major joints and / or bones). The pose and shape parameters are sometimes collectively referred to as the 3D body pose.
[0013] Generative body models can take into account variations in body shape depending on the (e.g., 3D) body pose.
[0014] The generative body model may comprise a Skinned Multi-Person Linear model (SMPL) model, as described in M. Loper et al., SMPL: A skinned multi-person linear model, ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1-248:16, October 2015, which is incorporated herein by reference.
[0015] Alternatively or additionally, the 2D skeletal representation of the body, the 2D projected (particularly dense) semantic coding of the body, and the 2D depth map of the body can be determined based on the generative body model. For example, the 2D skeletal representation of the body may include a projection of the 3D position of the skeleton onto an image plane. Alternatively or additionally, the 2D projected (particularly dense) semantic coding of the body and / or the 2D depth map of the body may include a projection of the 3D surface of the body (e.g., represented by a 3D mesh) onto (particularly the same) image plane.
[0016] The 2D projected encoding of the body can be a semantic encoding that includes semantic information, e.g., for each pixel and / or voxel, in the image plane. For example, a pixel can include semantic information (e.g., represented by one or more colors) that encodes an anatomical structure, such as a limb (e.g., arm or leg), head, or torso, of the body that is to be imaged (e.g., generated) on the synthetic image data.
[0017] Alternatively or additionally, the coding may be dense. Dense coding may indicate that coding is performed independently for each pixel and / or each voxel. Alternatively or additionally, dense (and / or dense semantic) coding may refer to coding (and / or coding vectors) being provided for all points (and / or pixels and / or voxels) on the surface of the body. This is in contrast to semantic coding, which provides coding only for some points (and / or pixels and / or voxels). An example of conventional (particularly non-dense) coding is shown in FIG. 7B. The gray values (and / or gray shades) in FIG. 7B may, for example, provide semantic information only for joints (e.g., left elbow, right elbow). Alternatively, the semantic coding may refer only to limbs (e.g., as a large-scale structure).
[0018] The 2D depth map can encode the (e.g., relative) depth of the imaged information, for example, for each pixel and / or voxel. Lighter colors (e.g., white or light gray) can encode information closer to the viewer, and darker colors (e.g., dark gray or black) can encode information further from the viewer. For example, anatomical structures may have increasingly brighter pixels as they lie in the foreground of the generated composite image data.
[0019] The received textual prompts (also referred to as text prompts) may include data representing information in textual form (e.g., as raw data and / or in language, particularly natural language). Receiving may be performed via a digital interface (e.g., a user interface (UI), particularly a graphic user interface (GUI)) and / or a human-machine interface (HMI). Alternatively or additionally, the received text prompts may be received in acoustic form to be converted into a format processable as input for the conditional image synthesis model. Further alternatively or additionally, the text prompts may be received from a computer-implemented text prompt generator configured to generate a diverse set of text prompts.
[0020] The received text prompts related to body appearance may include information regarding the body's appearance, the body's height, the body's weight, the body's proportions (e.g., including the size relationship between the body's head and torso), age, ethnicity (particularly in the case of a human body), species (particularly in the case of an animal body), e.g., horse, cow, dog and / or cat, biological sex, and / or clothing worn on the (particularly human) body, e.g., outerwear such as a coat.
[0021] Text prompts related to body appearance can supplement the received visual information with further details regarding body appearance for generating a composite image.
[0022] The received text prompts related to environmental information for the body may include information about location (e.g., outdoors, urban and / or suburban surroundings), lighting conditions (e.g., daytime or nighttime), background scene in which the body is located, and / or weather conditions (e.g., sunny, foggy, rainy, and / or snowy).
[0023] Text prompts related to environmental information can be used to generate the body's environment in synthetic image data (e.g., of a photorealistic type). For example, the text prompts can be "woman on a bridge," "man walking in a park in the rain," or "man on stairs wearing a coat."
[0024] In a further development, the received visual information and text prompts may include multiple visual information and associated text prompts, each for one (particularly different) body, so that multiple bodies are included in the generated composite image data.
[0025] The generation of synthetic image data of the body by the conditioned image synthesis model may involve the calculation of a synthetic image by image processing carried out by a digital computer unit. Alternatively or additionally, the generation does not necessarily have to be based on the reception of images by an optical sensor.
[0026] The conditioned image synthesis model can be or comprise a generative AI and / or a generative model. The conditioned image synthesis model can be or comprise a text-to-image model for converting received text prompts into one or more features of synthetic image data. The synthetic image can be generated with control over, for example, 3D body pose (as encoded by visual information), appearance, and location (as an example of an environment).
[0027] The constrained image synthesis model is sometimes referred to as the upstream AI (and / or upstream model).
[0028] The constrained image synthesis model may include a deep learning (DL) model.
[0029] The conditioning may include a combination of a depth map, a (particularly dense) semantic coding of the body (also referred to as semantic information), and a 2D skeleton (and / or 2D projection) of the body. Alternatively or additionally, the content of the generated synthetic image data may be controlled by text, especially in the form of text prompts.
[0030] The downstream AI (also referred to as a downstream model), in particular the downstream NN, may receive as input the generated synthetic image data and, optionally, the received visual information (and / or its root information including 3D body pose and / or pose parameters and shape parameters and / or any further information related to the body determined based on the visual information). The received visual information (and / or its root information including 3D body pose and / or pose parameters and shape parameters and / or any further information related to the body determined based on the visual information) may encode ground truth related to the generated synthetic image data for a downstream task performed by the downstream AI, in particular the downstream NN.
[0031] Downstream tasks may include object detection, (e.g., 3D) body pose detection, classification and / or semantic segmentation of image data. Downstream tasks may enable obstacle detection relevant to safe automated (e.g., autonomous) driving, movement (and / or motion) planning for robots and / or consumer electronic devices in automation systems, and / or detection of human bodies and / or animals for access control.
[0032] The method may further include determining, by the generative body model, visual information related to the body, including a 2D skeletal representation of the body, a 2D projected (especially dense) semantic encoding of the body, and a 2D depth map of the body. Determining the visual information may include performing a projection of the generated body model onto an image plane.
[0033] The 2D skeletal representation (also referred to as information indicating 2D skeletal positions) may include, in particular, 2D keypoint information. The 2D keypoints may refer to the positions of (e.g., major) joints (e.g., a subset of joints representing large joints (and / or motor functions), such as shoulders, hips, elbows, knees, wrists, and / or ankles) of a human or artificial body. The 2d keypoint information may include 2D projections of the (e.g., major) joint positions onto the image plane.
[0034] A 2D depth map can encode 3D body pose in terms of (e.g., simple) 3D surface representations, in particular 3D meshes. The 2D depth map may include an assignment of brightness (and / or color) to each point of the 3D surface representation.
[0035] The 2D keypoints can be located within a volume enclosed by a (for example, simple) 3D surface representation, in particular a 3D mesh.
[0036] The 2D projected (particularly dense) semantic coding (also called semantic information) can assign information about body parts to each point of the 2D depth map. The semantic information can include assigning a color (and / or brightness) to each point of the 2D depth map. The colors can be specific to anatomical structures, such as arms and legs, for example.
[0037] Determining visual information through a generative body model ensures a consistent combination of the 2D skeletal representation, the 2D projected (particularly dense) semantic coding, and the 2D depth map of the body.
[0038] The generative body model may generate a model of the body based on a set of shape parameters and pose parameters. Optionally, the generative body model may include an SMPL model.
[0039] Parameterization of the body in terms of a set of shape and pose parameters allows for the generation of a wide variety of (eg, 3D) body poses, body shapes and / or body-related visual information.
[0040] SMPL comprises a learned model of human body shape and pose-dependent shape variations and / or a skinned vertex-based model of natural human poses, as described in M. Loper et al., SMPL: A Skinned Multi-Person Linear Model, ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1-248:16 (Oct. 2015), which is incorporated herein by reference. The use of SMPL enables accurate representation of human body shape and pose (and / or 3D body pose). SMPL may include consideration of forward kinematics, linear blend skinning, and / or linear regression to obtain 3D keypoints. The 3D keypoints may include the 3D (e.g., coordinates and / or voxels) locations of (e.g., key) joints.
[0041] Alternatively or additionally, for animal bodies, shape and pose models are described in S. Zuffi et al., 3D Menagerie: Modeling the 3D Shape and Pose of Animals, http: / / arxiv.org / abs / 1611.07700, [5a], which is incorporated herein by reference. One model is trained for different species (e.g., including quadrupedal mammals).
[0042] The method may further include providing the generated synthetic image data of the body to a downstream AI, particularly a downstream NN. The downstream AI, particularly a downstream NN, may be configured for a body detection-related task on the sensor data. Providing the generated synthetic image data may include providing the visual information, the set of shape parameters and pose parameters of the generated body model, the generated body model on which the visual information is determined, and / or one or more related quantities as ground truth for the body detection-related task.
[0043] Providing the generated synthetic image data along with ground truth can enable efficient training, validation and / or testing of downstream AI, particularly downstream NNs.
[0044] The conditioned image synthesis model may comprise a generative text-image model configured to generate synthetic image data based on received text prompts, and an image conditioning model configured to encode the received visual information to condition and / or control (and / or modify) the generative text-image model.
[0045] The generative text-image model can enable the generation and / or synthesis of (particularly photorealistic) synthetic images that include appearance and / or environmental information of the text prompt.
[0046] The generative text-image model may comprise an encoder, a (particularly skip-connected) decoder, and optionally one or more intermediate blocks. For example, the generative text-image model may comprise a U-Net architecture.
[0047] The image conditioning model may comprise an encoder, optionally one or more intermediate blocks, followed by a zero convolutional layer, which may comprise a convolutional layer with both weights and biases initialized as zero (e.g., 1×1).
[0048] The encoder of the image conditioning model can be a (particularly trainable) copy of the encoder of the generative text-image model. The output of the optional one or more intermediate blocks and the output of the zero convolutional layer can be fed (e.g., by a cross-attention mechanism) to a decoder of the optional one or more intermediate blocks and the generative text-image model.
[0049] The image conditioning model can encode and / or transform the body-related visual information (i.e., the received 2D skeletal representation, the 2D projected (especially dense) semantic coding, and the 2D depth map) so that it can be fed to the text-image model (e.g., using a cross-attention mechanism for several layers of the text-image model, in particular the decoder layer), thereby enabling the text-image model to generate (especially photorealistic) synthetic image data that is consistent with the generated body model underlying the visual information. Alternatively or additionally, conditioning the generation of synthetic images (especially by the generative text-image model) by visual information encoded (especially by the image conditioning model) advantageously provides (especially photorealistic) synthetic image data including the body together with ground truth data.
[0050] The image conditioning model can be configured to perform encoding of received visual information in combination with (and / or based on) text prompts related to the appearance of the body and / or environmental information relative to the body.
[0051] The generative text-image model can "start from scratch," e.g., start with pure noise (also referred to as a degraded image) as image input. Additionally, the generative text-image model can receive text prompts related to the appearance of the body and / or environmental information relative to the body. The image conditioning model can receive visual information including, for example, a 2D skeletal representation, a 2D projected (particularly dense) semantic coding, and a 2D depth map of the body, concatenated along the channel dimension. Additionally, the image conditioning model can receive image input (e.g., a degraded image and / or pure noise) (particularly identical to that received by the generative text-image model) and / or text prompts related to the appearance of the body and / or environmental information relative to the body.
[0052] The output of the generative text-image model (particularly conditioned by the output of the image conditioning model) may include a prediction of the noise added to the clean image to create the degraded image (and / or pure noise). The (e.g., real and / or clean) image output can be determined (e.g., calculated) from the predicted noise.
[0053] Combined conditioning on 2D skeletal representations, 2D projected (especially dense) semantic coding, and 2D depth maps, all of which are included in the received visual information, can significantly improve the photorealistic accuracy of synthetic image data and, therefore, the quality of downstream training.
[0054] The generative text-image model may comprise a diffusion model, in particular a Stable Diffusion (SD) network, which may comprise a U-Net architecture with an encoder and a decoder (in particular with skip connections).
[0055] The diffusion model (also referred to as a diffusion probability model and / or a score-based generative model) may comprise a machine learning (ML) model and / or a generative model configured to perform image generation, image denoising, inpainting, and / or super-resolution by a diffusion process, in particular by using a forward process (e.g., adding noise, in particular Gaussian noise, to an image), a backward process (e.g., predicting noise, in particular Gaussian noise, in an image and compensating and / or subtracting accordingly), and a sampling procedure.
[0056] A diffusion model can be understood as casting the image generation process as a denoising task. Iteratively transforms pure noise into a real image. A model, specifically a U-Net, can be trained for this denoising task. It is possible to work directly in image space and generate real images directly from pure noise. Alternatively or additionally, a latent diffusion model can perform (and / or perform) the reverse process in the low-dimensional latent space of a variational autoencoder (VAE). Pure noise can be transformed into a latent embedding of an image that must be decoded to generate a real image. The "image in latent space can be transformed back into an image in non-latent (and / or image) space" part can be performed (and / or performed) by the decoder of the VAE. Image generation does not require the use of the VAE's encoder. Alternatively or additionally, the VAE's encoder needs to be used only during training of the diffusion model.
[0057] SD may use or comprise a latent diffusion model. SD may be related to or comprise a neural network architecture (also referred to as a model) for text-to-image (particularly diffusion) generation. SD aims to learn a diffusion process that generates a (e.g., denoised) image dataset from a given probability distribution (and / or the distribution of a given latent embedding dataset). The SD's neural network architecture may comprise a transformer and / or U-Net with an encoder and decoder with skip connections (e.g., between the encoder and decoder layers) and / or cross-attention mechanisms (e.g., within the encoder and / or decoder layers). Alternatively or additionally, the SD's U-Net may comprise an attention layer.
[0058] The SD NN architecture may further include a diffusion process that transforms (e.g., noisy and / or corrupted) input images into images in a latent image space. The transformer and / or U-Net can transform the images in the latent space back into images in the non-latent (and / or image) space (e.g., denoised and / or clean images).
[0059] SD can advantageously generate high-quality (eg, high-resolution and / or photorealistic) synthetic images based on text prompts at low computational cost.
[0060] Alternatively or additionally, other types of diffusion models (e.g., "Imagen," described by the Google Research Brain Team in Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, arXiv:2205.11487,
[13] , incorporated herein by reference; the weights of this model are not published, but open-source copies exist) can operate in image space. In general, the techniques for generating synthetic image data described herein are applicable to other types of diffusion models.
[0061] Alternatively or additionally, the text-image model may comprise DALL·E or any of its successors (e.g., DALL·E2) and / or any text-conditional image generation based on CLIP (e.g., as described in A. Ramesh et al. Hierarchical Text-Conditional Image Generation with CLIP Latents, arXiv:2204.06125 [cs.CV],
[14] , incorporated herein by reference).
[0062] The image conditioning model can comprise a ControlNet, which comprises an encoder and convolutional layers with a cross-attention mechanism for the generative text-to-image model, particularly for the encoder of the SD network.
[0063] Generative models, particularly SD networks, can be conditioned by certain features, particularly through the use of ControlNet. For the purposes of conditioning, see L. Zhang et al. Adding Conditional Control to Text-to-Image Diffusion Models (2023), available at Computer Vision Foundation and IEEE Xplore, which is incorporated herein by reference.
[0064] Connecting the SD network to ControlNet can provide particularly computationally efficient conditioning, saving time and memory (eg, in a graphics processing unit (GPU)).
[0065] The combination of the SD network and the ControlNet can provide particularly varied generated synthetic image data (particularly realistic) that includes bodies consistent with the generated body models.
[0066] The ControlNet can be initialized with weights from the depth ControlNet of Zhang et al.
[11] . The main difference is that, in the technique for generating synthetic image data described herein, the ControlNet is fine-tuned with respect to data (especially for body detection-related tasks) by inputs (especially visual information received related to the body and / or text prompts related to the body's appearance and / or environmental information relative to the body) to allow for better control over appearance and background compared to the original ControlNet.
[0067] Regarding a second method aspect, a computer-implemented method is provided for training a conditioned image synthesis model to generate synthetic image data usable for training, validation, and / or testing of downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data. The method includes receiving a synthetic image training dataset. The synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated with a set of 3D body markers externally applied to the body at the time of image dataset acquisition, a (particularly 2D) image dataset of a body annotated by a human expert, where the annotations include 3D body pose information, and / or a (particularly 2D) image dataset of a body annotated with 2D keypoint information. The method further includes training a conditioned image synthesis model based on the received synthetic image training dataset. Training the conditioned image synthesis model includes converting annotations in the synthetic image training dataset into at least a portion of visual information related to the body. The visual information includes a 2D skeletal representation of the body, a 2D projected (particularly dense) semantic coding of the body, and a 2D depth map of the body. In particular, the 2D image dataset is considered as the ground truth for the generated synthetic image data.
[0068] Training the conditioned image synthesis model may further include generating 3D pose information, in particular from annotated 2D keypoint information of the (especially 2D) image dataset and / or as an intermediate step before converting the annotations into at least a portion of the body-related visual information.
[0069] The (particularly 2D) image dataset may comprise a real image of the body. Alternatively or additionally, the 2D image dataset may comprise a synthetic image of the body.
[0070] The (especially 2D) image dataset can correspond to the ground truth for training a conditioned image synthesis model.
[0071] Training can be applied to a combination of a generative text-image model and an image conditioning model. The image conditioning model (e.g., ControlNet) can be trained using a standard denoising loss function, specifically. ControlNet can also take text prompts as input to predict noise, as illustrated in Equation (1).
[0072] By training a conditional image synthesis model, we can learn the weights of the nodes (e.g., in the hidden layer of a NN, especially a ControlNet).
[0073] The 3D body marker set and / or annotations (also called labels) by human experts may include (particularly accurate) 3D body poses.
[0074] 3D pose information generated from annotated 2D keypoint information is sometimes referred to as pseudo-3D labels.
[0075] The training of the conditioned image synthesis model can be monitored by quality control metrics, which may include 3D (e.g., person) pose metrics, in particular, mean per joint position error (mpjpe), procrustes aligned mean per joint position error (pa-mpjpe), and / or percentage of correct keypoints (pck). A distance threshold can be selected for one of the 3D (e.g., person) pose metrics.
[0076] Alternatively or additionally, the quality control metric may include Frechet inception distance (FID) and / or an image similarity metric, in particular learned perceptual image patch similarity (LPIPS).
[0077] Regarding a third method aspect, a computer-implemented method is provided for training, validating, and / or testing a downstream AI, particularly a downstream neural network (NN), for performing body detection-related tasks based on sensor data. The method includes receiving synthetic image data of a body. The synthetic image data of the body can be generated according to the first method aspect. The synthetic image data of the body includes a generated body model, visual information and / or related quantities related to the body. The visual information may include a 2D skeletal representation of the body, a 2D projected (particularly dense) semantic coding of the body, and a 2D depth map of the body. Alternatively or additionally, the related quantities may include information about the body derived from the generated body model and / or derived from the visual information. The method further includes training the downstream AI, particularly a downstream neural network (NN), based on the received synthetic image data of the body. Any one of the generated body model, the visual information and / or related quantities related to the body can be considered ground truth.
[0078] The synthetic dataset data provided by the technique can be used for training, benchmarking, validating, and / or testing downstream AI (particularly downstream NNs). In particular, the use of synthetic image datasets can make it possible to identify systematic errors and / or biases in downstream AI (particularly downstream NNs).
[0079] By using synthetic image data of the body together with visual information, the underlying generated body model, and / or ground truth based on information derived therefrom to train, validate, and / or test the downstream AI, particularly the downstream NN, the downstream AI, particularly the downstream NN, can be improved in terms of performing body detection-related tasks for a wide variety of bodies and environments. Alternatively or additionally, the speed and / or convergence of training of the downstream AI (particularly the downstream NN) can be improved.
[0080] The performance of body detection related tasks may be performed based on the received sensor data. The tasks may include classification, semantic segmentation and / or object detection, in particular body detection.
[0081] The detection of bodies may in particular include the detection of road users, for example pedestrians and / or (especially wild) animals such as wild boars and deer.
[0082] The sensor data may be received from a (e.g., video) camera, a radar sensor, a LiDAR sensor, an ultrasonic sensor, a motion sensor, and / or a thermal imaging sensor. The sensor data may alternatively or additionally include a sensed belt position of a driver of the vehicle.
[0083] The method according to the third method aspect can be used to apply downstream AI, in particular downstream NN, to automatic (in particular autonomous) driving, robot movement planning, operation of household appliances and / or control of access control systems.
[0084] Applications may include (e.g., receiving sensor data from) video surveillance and / or motion capture. Alternatively or additionally, applications may include assisted driving, autonomous driving, smart home (and / or domestic) appliances, (e.g., robotic) personal assistants, and / or operation of technical devices (e.g., robots, power tools, and / or manufacturing machinery) in manufacturing environments. Further alternatively or additionally, applications may include security systems, for example, for theft prevention, access control, and / or vehicle driver monitoring.
[0085] The downstream AI (particularly the downstream NN) may include, in the inference stage, receiving sensor data as input and providing task-specific output, including, among other things, determined classifications and / or 3D body poses.
[0086] With respect to a first device aspect, there is provided a computing device for generating synthetic image data usable for training, validating, and / or testing downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data. The computing device may be configured to perform any one of the steps and / or comprise any one of the features disclosed in relation to the first method aspect.
[0087] With respect to a second device aspect, there is provided a computing device for training a conditioned image synthesis model for generating synthetic image data usable for training, validating, and / or testing downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data. The computing device may be configured to perform any one of the steps and / or comprise any one of the features disclosed in connection with the second method aspect.
[0088] With respect to a third device aspect, there is provided a computing device for training, validating, and / or testing a downstream AI, in particular a downstream NN, for performing body detection-related tasks based on sensor data. The computing device may be configured to perform any one of the steps and / or comprise any one of the features disclosed in connection with the third method aspect.
[0089] Regarding a system aspect, a system for training, validating, testing, and / or applying a downstream AI (particularly a downstream NN) is provided. The system comprises a computing device according to a first device aspect configured to generate synthetic image data, a computing device according to a second device aspect configured to train a conditioned image synthesis model using the generated synthetic image data, and a computing device according to a third device aspect configured to train, validate, and / or test the downstream AI (particularly the downstream NN). The system further comprises the downstream AI (particularly the downstream NN) configured to perform a body-related detection task, and at least one sensor and / or image capture device configured to provide sensor data on which the body-related detection task is performed.
[0090] According to a further aspect, a computer program product is provided, the computer program product including program elements that, when loaded into a memory of a computing device, direct the computing device to perform method steps according to the first method aspect, the second method aspect and / or the third method aspect.
[0091] In accordance with yet a further aspect, there is provided a computer-readable medium having stored thereon program elements that, when executed by a computing device, can be read and executed by the computing device to perform method steps according to the first method aspect, the second method aspect, and / or the third method aspect. [Brief explanation of the drawings]
[0092] [Figure 1] 1 is a flowchart of a method for generating synthetic image data usable for training, validation, and / or testing of downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data. [Figure 2] 1 is a flowchart of a method for training a conditional image synthesis model to generate synthetic image data usable for training, validation, and / or testing of downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data. [Figure 3] 1 is a flowchart of a method for training, validating, and / or testing a downstream AI, particularly a downstream NN, for performing body detection-related tasks based on sensor data. [Figure 4] 1 is an overview of the structure and architecture of a computing device for generating synthetic image data usable for training, validation and / or testing of downstream AI, particularly downstream NNs, for body detection related tasks based on sensor data. [Figure 5] 1 is an overview of the structure and architecture of a computing device for training a conditional image synthesis model to generate synthetic image data usable for training, validation and / or testing of downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data. [Figure 6] 1 is an overview of the structure and architecture of a computing device for training, validating, and / or testing downstream AI, particularly downstream NNs, for performing body detection related tasks based on sensor data. [Figure 7A]FIG. 1 illustrates conventional conditioning with only a 2D depth map and a 2D skeletal representation, leading to the generation of a synthetic image that is inconsistent with the conditioning and / or the text prompt "a man in a street wearing a coat." [Figure 7B] FIG. 1 illustrates conventional conditioning with only a 2D depth map and a 2D skeletal representation, leading to the generation of a synthetic image that is inconsistent with the conditioning and / or the text prompt "a man in a street wearing a coat." [Figure 7C] FIG. 1 illustrates conventional conditioning with only a 2D depth map and a 2D skeletal representation, leading to the generation of a synthetic image that is inconsistent with the conditioning and / or the text prompt "a man in a street wearing a coat." [Figure 7D] FIG. 1 illustrates conventional conditioning with only a 2D depth map and a 2D skeletal representation, leading to the generation of a synthetic image that is inconsistent with the conditioning and / or the text prompt "a man in a street wearing a coat." [Figure 8A] Figure 1 shows a first example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man on the stairs wearing a coat" to lead to consistent and realistic synthetic image data. [Figure 8B] Figure 1 shows a first example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man on the stairs wearing a coat" to lead to consistent and realistic synthetic image data. [Figure 8C] Figure 1 shows a first example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man on the stairs wearing a coat" to lead to consistent and realistic synthetic image data. [Figure 8D] Figure 1 shows a first example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man on the stairs wearing a coat" to lead to consistent and realistic synthetic image data. [Figure 8E]Figure 1 shows a first example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man on the stairs wearing a coat" to lead to consistent and realistic synthetic image data. [Figure 9A] Figure 2 shows a second example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man in a coat on the street" to result in consistent and realistic synthetic image data. [Figure 9B] Figure 2 shows a second example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man in a coat on the street" to result in consistent and realistic synthetic image data. [Figure 9C] Figure 2 shows a second example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man in a coat on the street" to result in consistent and realistic synthetic image data. [Figure 9D] Figure 2 shows a second example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man in a coat on the street" to result in consistent and realistic synthetic image data. [Figure 9E] Figure 2 shows a second example where a 2D skeletal representation, a 2D (especially dense) semantic coding and a 2D depth map are combined with the text prompt "A man in a coat on the street" to result in consistent and realistic synthetic image data. [Figure 10] FIG. 1 illustrates an exemplary architecture of a constrained image synthesis model utilizing a combination of stable diffusion and ControlNet. [Figure 11A] FIG. 10 shows a schematic diagram of an example of a downstream application. [Figure 11B] FIG. 10 shows a schematic diagram of an example of a downstream application. [Figure 11C] FIG. 10 shows a schematic diagram of an example of a downstream application. DETAILED DESCRIPTION OF THE INVENTION
[0093] Detailed Description FIG. 1 shows a schematic flow chart of a computer-implemented method 100 for generating synthetic image data that can be used for training, validation, and / or testing downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data.
[0094] Method 100 includes step S102 of receiving visual information related to a body. The visual information includes a two-dimensional (2D) skeletal representation of the body, a 2D projected (especially dense) semantic coding of the body, and a 2D depth map of the body. Method 100 further includes step S104 of receiving a text prompt related to the appearance of the body and / or environmental information relative to the body. Method 100 further includes step S106 of generating synthetic image data of the body based on the text prompt received in S104, conditioned by the visual information received in S102. The generating in S106 is performed by a conditioned image synthesis model.
[0095] Optionally, the method 100 comprises a step S101 of determining visual information related to the body, including a 2D skeletal representation of the body, a 2D projected (especially dense) semantic coding of the body, and a 2D depth map of the body, by means of the generative body model. Determining the visual information S101 may comprise performing a projection of the generated body model onto an image plane.
[0096] Further optionally, method 100 includes a step S108 of providing the synthetic image data of the body generated in S106 to a downstream AI, in particular a downstream NN. The downstream AI, in particular a downstream NN, can be configured for a body detection-related task on the sensor data. Providing S108 the synthetic image data generated in S106 may include providing the visual information, a set of shape parameters and pose parameters of the generated body model, the generated body model on which the visual information is determined, and / or one or more related quantities as ground truth for the body detection-related task.
[0097] FIG. 2 shows a schematic flowchart of a computer-implemented method 200 for training a conditional image synthesis model to generate synthetic image data that can be used for training, validating, and / or testing downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data.
[0098] The method 200 includes step S202 of receiving a synthetic image training dataset. The synthetic image training dataset may include a (particularly 2D) image dataset of a body, which is annotated with a set of 3D body markers externally applied to the body at the time of acquiring the image dataset. Alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated by a human expert, the annotation including 3D body pose information. Further, alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated with 2D keypoint information. In the case of a (particularly 2D) image dataset of a body, training the conditioned image synthesis model includes generating 3D pose information.
[0099] Method 200 further includes step S204 of training a conditioned image synthesis model based on the synthetic image training dataset received in S202. Training the conditioned image synthesis model S204 includes converting annotations in the synthetic image training dataset into at least a portion of visual information related to the body. The visual information includes a 2D skeletal representation of the body, a 2D projected (particularly dense) semantic coding of the body, and a 2D depth map of the body. The (particularly 2D) image dataset is considered as ground truth for the generated synthetic image data.
[0100] FIG. 3 shows a schematic flow chart of a computer-implemented method 300 for training, validating, and / or testing downstream AI, particularly downstream NNs, for performing body detection-related tasks based on sensor data.
[0101] Method 300 includes step S302 of receiving synthetic image data of a body. The synthetic image data of the body is generated according to method 100. The synthetic image data of the body includes a generated body model, visual information and / or related quantities related to the body. The visual information includes a 2D skeletal representation of the body, a 2D projected (especially dense) semantic coding of the body, and a 2D depth map of the body. Alternatively or additionally, the related quantities include information about the body derived from the generated body model and / or derived from the visual information. Method 300 further includes step S304 of training a downstream AI, particularly a downstream NN, based on the synthetic image data of the body received in S302. The generated body model, visual information and / or related quantities related to the body are considered ground truth.
[0102] FIG. 4 shows a schematic architecture of a computing device 400 for generating synthetic image data that can be used for training, validation, and / or testing downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data.
[0103] The computing device 400 comprises a first input interface 402 configured to receive visual information related to a body. The visual information includes a 2D skeletal representation of the body, a 2D projected (particularly dense) semantic coding of the body, and a 2D depth map of the body. The computing device 400 further comprises a second input interface 404 configured to receive text prompts related to at least one of the body's appearance and / or environmental information relative to the body. The computing device 400 further comprises a generation module 406 comprising a conditioned image synthesis model. The conditioned image synthesis model is configured to generate synthetic image data of the body based on the received text prompts conditioned by the received visual information.
[0104] Optionally, the computing device 400 comprises a determining module 401 configured to determine visual information related to the body, including a 2D skeletal representation of the body, a 2D projected (especially dense) semantic encoding of the body, and a 2D depth map of the body, by means of the generative body model. Determining the visual information may include performing a projection of the generated body model onto an image plane.
[0105] Further optionally, the computing device 400 comprises an output interface 408 configured to provide the generated synthetic image data of the body to a downstream AI, in particular a downstream NN. The downstream AI, in particular a downstream NN, can be configured for a body detection-related task on the sensor data. Providing the generated synthetic image data may include providing the visual information, a set of shape parameters and pose parameters of the generated body model, the generated body model on which the visual information is determined, and / or one or more related quantities as ground truth for the body detection-related task.
[0106] Any one of the first input interface 402, the second input interface 404, and the optional output interface 408 may be embodied by an input / output interface 410. Alternatively or additionally, the generating module 406 and / or the optional determining module 401 may be embodied by a processing unit. Further alternatively or additionally, the computing device 400 may include at least one memory 414.
[0107] FIG. 5 illustrates a schematic architecture of a computing device 500 for training a conditional image synthesis model to generate synthetic image data that can be used for training, validating, and / or testing downstream AI, particularly downstream NNs, for body detection-related tasks based on sensor data.
[0108] The computing device 500 comprises an input interface 502 configured to receive a synthetic image training dataset. The synthetic image training dataset may include a (particularly 2D) image dataset of a body, which is annotated with a set of 3D body markers externally applied to the body when acquiring the image dataset. Alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated by a human expert, where the annotation includes 3D body pose information. Further alternatively or additionally, the synthetic image training dataset may include a (particularly 2D) image dataset of a body annotated with 2D keypoint information. In particular for a (particularly 2D) image dataset of a body, training the conditioned image synthesis model includes generating 3D pose information, for example, by a 3D pose information generation module (not shown in FIG. 5 ). The computing device 500 further comprises a training module 504 configured to train the conditioned image synthesis model based on the received synthetic image training dataset. Training the conditioned image synthesis model involves converting annotations from a synthetic image training dataset into at least a portion of visual information related to the body, including a 2D skeletal representation of the body, a 2D projected (particularly dense) semantic encoding of the body, and a 2D depth map of the body, where the (particularly 2D) image dataset is considered as ground truth for the generated synthetic image data.
[0109] The input interface 502 may be embodied by an input / output interface 506. Alternatively or additionally, the training module 504 may be embodied by a processing unit 508. Further alternatively or additionally, the computing device 500 may include at least one memory 510.
[0110] FIG. 6 illustrates a schematic architecture of a computing device 600 for training, validating, and / or testing downstream AI, particularly downstream NNs, for performing body detection-related tasks based on sensor data.
[0111] The computing device 600 comprises an input interface 602 configured to receive synthetic image data of a body. The synthetic image data of the body is generated according to the method 100. The synthetic image data of the body includes a generated body model, visual information and / or related quantities related to the body. The visual information includes a 2D skeletal representation of the body, a 2D projected (particularly dense) semantic coding of the body, and a 2D depth map of the body. Alternatively or additionally, the related quantities include information about the body derived from the generated body model and / or derived from the visual information. The computing device 600 further comprises a training module 604 configured to train a downstream AI, particularly a downstream neural network, based on the received synthetic image data of the body. The generated body model, visual information and / or related quantities related to the body are considered ground truth.
[0112] The input interface 602 may be embodied by an input / output interface 606. Alternatively or additionally, the training module 604 may be embodied by a processing unit 608. Further alternatively or additionally, the computing device 600 may include at least one memory 610.
[0113] Any one of the processing units 412, 508, 608 may be embodied by a central processing unit (CPU) and / or a graphics processing unit (GPU).
[0114] The techniques presented herein can generate synthetic image data to benchmark and improve tasks such as 3D human pose estimation from a single RGB image. The methods can be based on deep learning, which traditionally requires large amounts of annotated data to be effective. However, obtaining accurate 3D body pose-annotated (especially real-world) image data is a challenging process.
[0115] The techniques presented herein enable the generation of image data of (especially live) people (and / or animals) with controlled 3D body pose, appearance, and location. Generative models (including, in particular, constrained image synthesis models) can be used to create (especially synthetic image) data for fine-grained evaluation of 3D (e.g., human) pose estimators, particularly to benchmark their robustness and generalization ability to different conditions and identify systematic errors and biases. Alternatively or additionally, generative models (especially constrained image synthesis models) can also be used to generate training data for 3D human pose estimators.
[0116] The techniques presented here enable text-based image generation while controlling the 3D body pose of a person (and / or animal). In particular, text control and 3D body pose control can be separated.
[0117] Previous text-image synthesis methods, such as ControlNet
[11] , incorporated herein by reference, are only able to control 2D pose while maintaining fine-grained text control. While ControlNet, which utilizes depth-conditioned generation, is capable of controlling 3D pose (e.g., of a person), the ability to control image content with text has traditionally been severely limited, as shown illustratively in Figures 7A, 7B, 7C, and 7D.
[0118] Figures 7A, 7B, 7C, and 7D exemplify the shortcomings of the conventional ControlNet
[11] (which uses at most two pieces of visual information, specifically 2D human pose conditions and 2D depth maps), compared to a technique for generating synthetic image data using three pieces of visual information: a 2D skeletal representation, a 2D projected (particularly dense) semantic coding, and a 2D depth map, as well as a text prompt (which may be received by both a text-image model and an image conditioning model). Figures 7A and 7B show the depth conditions and 2D human pose conditions, respectively, based on which the synthetic images shown in Figures 7C and 7D are generated. The depth image in Figure 7A represents a 3D body pose, and the key points in Figure 7B (including a few major joints, such as shoulders, elbows, wrists, hips, knees, and ankles) represent its 2D projection. The text prompt in the example of Figures 7C and 7D is "A man wearing a coat, walking down the street."
[0119] The image in Figure 7C can be generated by utilizing a conditional image synthesis model similar to that employed in techniques for generating synthetic image data using, for example, three visual elements and a text prompt, but lacking 2D (especially dense) semantic coding (and / or not using a text prompt in the image conditioning model). The example in Figure 7D can be obtained using a conventional ControlNet-depth model, among other things. Note that the person's posture does not match the 3D conditions. For example, in Figure 7C, the person's left arm is close to the body but should be facing further forward, and the person's right hand is grasping the banister, but according to the depth image in Figure 7A, it should be positioned open with the thumb facing up. Additionally, generating depth conditions fails to generate the court in the example in Figure 7D, which is specified in the text prompt. The conditional image synthesis model employed in the techniques for generating synthetic image data described herein follows the text prompt better than a conventional ControlNet-depth model. Notably, neither the court nor the street is shown in Figure 7D.
[0120] An alternative is a computer graphics-based approach [1], which is incorporated herein by reference. However, the pipeline setup is difficult and the assets must be manually designed and created, limiting overall versatility. In contrast, the technique presented here is easy to use and has the ability to generate diverse synthetic image datasets aligned with 3D pose ground truth without requiring any manual effort.
[0121] Image diffusion models learn to progressively denoise an image to generate samples. The denoising can be done in pixel space or in a "latent" space encoded from training data. Stable diffusion (SD) uses latent images as the training domain. In this context, the terms "image," "pixel," and "denoising" all refer to the corresponding concepts in "perceptual latent space" [8], which is incorporated herein by reference. Given an image z0, the diffusion algorithm incrementally adds noise to the image to generate samples of the noisy image z0. t where t is the number of times noise is added. When t is large enough, the image approximates pure noise. At time step t, a text prompt c (particularly related to body appearance and / or environmental information) is generated. t , and task-specific conditions c (including, inter alia, 2D skeletal representations, 2D projected particularly dense semantic coding, and visual information including a 2D depth map of the body). f Given a set of conditions, including t Network ε for predicting noise added to θ Learn, i.e.,
number
[0122] An embodiment of the technique builds on ControlNet
[11] and introduces 3D pose control to SD. The 3D mesh of an exemplary human body is encoded using the SMPL model [5]. SMPL is a generative model that factorizes the human body into shape (e.g., how individuals vary in height, weight, and / or body proportions) and pose (e.g., how the 3D surface deforms due to articulation). The shape β∈R 10 is parameterized by the first 10 coefficients of the PCA shape space. 3K is modeled in axis-angle representation by the relative 3D rotations of K = 23 joints. SMPL is a triangular mesh M(θ,β)∈R with N = 6980 vertices. 3N is a differentiable function that outputs a mesh obtained by shaping the vertices of a template body subject to β and θ, then articulating the bones according to the joint rotations θ using forward kinematics, and finally deforming the surface using linear blend skinning. The 3D keypoints X(θ,β)∈R 3P is obtained by linear regression from the vertices of the final mesh, where P (e.g., P=24) is the number of 3D joints in the skeleton.
[0123] The techniques presented herein can be used (particularly in downstream AI applications) for computer control of machines such as robots, (e.g., autonomous) vehicles, household appliances, power tools, manufacturing machines, personal assistants, communication systems such as access control systems, surveillance systems, and / or medical (imaging) systems by operating on digital (and / or analog) image data obtained by receiving sensor signals, e.g., video images, RGB camera images, radar images, LiDAR images, ultrasound images, motion images, and / or thermal images. The downstream AI does this by performing body detection related tasks, e.g., classification of sensor data, detection of the presence of objects in the sensor data, and / or semantic segmentation on the sensor data, e.g., regarding pedestrians.
[0124] The techniques described herein constitute an upstream portion of a machine learning (ML) and / or AI toolchain. Techniques (and / or upstream generative models) for generating synthetic image data can indirectly, though not necessarily directly, improve ML systems (and / or downstream AI, also referred to as downstream generative models) that can be used in the above applications. In particular, method 100 generates training data for method 300, which generates test data for testing whether a trained ML system can be safely operated.
[0125] The techniques described herein can be used not only for data augmentation but also for domain transfer tasks, e.g., from synthetic to photorealistic images. The generated samples can be used for evaluating and / or training any data-driven method, e.g., pedestrian detection models (especially for body detection-related tasks).
[0126] According to an exemplary embodiment, a constrained image synthesis model based on SD [8] and ControlNet
[11] is trained that can control the 3D body pose of a person (and / or animal) to generate synthetic image data. Additionally, control over the image content (e.g., location, person appearance, and / or weather) is provided through text prompts. The constrained image synthesis model is then used to generate images for systematic evaluation of 3D human pose estimators. This is possible because the techniques described herein use a (e.g., constrained image synthesis) model to control the 3D body pose.
[0127] 8A-8E and 9A-9E show examples of combined visual information (also called pose conditions) as conditions: 2D skeletal representation (also denoted as 2D keypoints, FIGS. 8A and 9A), 2D projected (especially dense) semantic coding (FIGS. 8B and 9B), and 2D depth maps (FIGS. 8C and 9C).
[0128] The generated composite image data in Figures 8D and 8E uses "Man on stairs wearing a coat" as the text prompt. The generated composite image data in Figures 9D and 9E uses "Man on street wearing a coat" as the text prompt.
[0129] The exemplary constrained image synthesis model is based on ControlNet with explicit 3D conditioning that captures the 3D structure of the human body and its semantics.
[0130] 10 shows a schematic diagram of the combination of an SD encoder 1004-E, a corresponding SD decoder 1004-D, and a ControlNet 1002 with a trainable copy 1002-E of the SD encoder 1004-E (specifically frozen after initial training). The ControlNet architecture 1002 controls the SD model (comprising the SD encoder 1004-E and the SD decoder 1004-D).
[0131] The ControlNet 1002 includes multiple zero convolutional layers 1002-ZC. In FIG. 10, one zero convolutional layer 1002-ZC before the trainable encoder 1002-E and multiple zero convolutional layers 1002-ZC after the trainable encoder 1002-E are shown by way of example. Any other arrangement of the zero convolutional layers 1002-ZC (e.g., more or fewer zero convolutional layers 1002-ZC before and / or after the trainable encoder 1002-E) may be possible.
[0132] The visual information (also referred to as visual image conditions) in reference numeral 1010 is c f ={c d ,c dp ,c k} and are input to the ControlNet (e.g., concatenated along the channel dimension). Visual information (and / or conditions) are then fed into a depth map of the (e.g., human) body, c d , (e.g., human) body (especially dense) semantic coding c dpand 2D skeletal representation (simply put, skeleton) c k It consists of three parts: The depth map provides 3D pose information.
[0133] (e.g., linked) visual information c in reference numeral 1010 f may be transformed by multiple layers (not explicitly shown in FIG. 10) before using the first zero convolutional layer 1002-ZC.
[0134] Text prompts c related to physical appearance (e.g., human) and / or environmental information at reference numeral 1012 t is supplied to the SD encoder 1004-E and the SD decoder 1004-D by, for example, a cross-attention mechanism.
[0135] Text prompt c in reference number 1012 t is provided in addition to ControlNet 1002 in the example of FIG.
[0136] Additionally, the first latent image z at 1006 t (also denoted as the degraded image at time step t and / or the noisy latent embedding of the degraded image at time step t) is used to initialize the SD, and the image z t-1 The (e.g., direct) output of the SD decoder 1004-D is the latent image z t , based on which the image z t-1 can be determined (e.g., calculated).
[0137] In an exemplary embodiment, visual information is generated by rendering a posed SMPL mesh. Given a human pose θ, including axis-angle representations of joint rotations, a parametric body model M is used to infer a 3D mesh of the human body. Rendering the depth of the mesh yields a representation c d The resulting (particularly dense) semantic coding is created by assigning each vertex of the mesh its 3D position in the T-pose as a color. This provides semantic information about different body parts, allowing ControlNet to distinguish between left and right, front and back, for example. Finally, 2D keypoints make it possible to use datasets that only provide 2D keypoint annotations. This increases the overall diversity of the synthetic image data generated (especially as training data for downstream AI).
[0138] A variety of datasets can be used to train a conditioned image synthesis model. Datasets with accurate 3D pose annotations provide the image data and 3D pose pairs required for accurate generation. However, datasets with accurate 3D pose annotations traditionally lack diversity in environment and appearance (e.g., of human or animal bodies). For this reason, the set of training datasets can be supplemented with datasets that provide only 2D keypoint annotations. Pseudo-3D labels for the 2D dataset can be determined (e.g., calculated) using (e.g., human) mesh recovery methods, and data can be pruned using alignment with 2D keypoints as a quality control metric.
[0139] 11A, 11B, and 11C illustrate exemplary downstream applications for a self-driving (e.g., autonomous) vehicle 1102-1, a (particularly autonomous) robot 1102-2, and an access control system 1102-3. Each downstream application includes a controller 1104 in which a corresponding downstream AI is installed and trained, validated, and / or tested according to method 300. FIG. 11C also illustrates two types of sensors that can be used for access control: a video camera 1110 and a microphone 1112.
[0140] To evaluate the 3D pose estimator, poses can be obtained from a common 3D pose estimation benchmark, such as 3DPW
[10] , which is incorporated herein by reference. Performance on the generated synthetic image data can be degraded by misdistribution of pixel values. Therefore, for systematic evaluation, a synthetic replica of the 3D pose estimation benchmark is created. The goal is to create a synthetic clone containing images that are as close as possible to the original image in terms of image content. A Visual-Question-Answering (VQA) model can be used to extract the image content, represented as text. The extracted content can be used to generate the image and / or synthetic replica. All robustness experiments are performed using this replica as a baseline to eliminate misdistribution as a factor.
[0141] An exemplary general setup for experiments on the robustness of a pose estimator to a particular attribute or set of attributes (e.g., related to body appearance and / or environmental information) can take the following form: The text prompt used to generate the replica can be used (e.g., again) to introduce the selected attribute or set of attributes. New synthetic image data is then generated from the same noise used to generate the base image. In this way, general structure and / or content is preserved, and only the intended attributes are introduced into the synthetic image data, allowing for more meaningful comparisons. Additionally, structure can be further preserved through the use of techniques such as prompt-to-prompt [2].
[0142] For evaluation, attributes such as clothing, location, lighting, weather, age, ethnicity, and / or gender can be targeted. In particular, (e.g., real-world) image data in which attributes such as lighting, weather, location, and clothing are controlled has traditionally been difficult to obtain due to the large variability in the real world. It is also important to consider a combination of attributes for evaluation.
[0143] Each attribute is likely to appear a sufficient number of times in the training data. However, certain combinations of attributes may never appear in the training data (especially in real-world situations), leading to poor performance. Exhaustively testing all combinations of attributes is infeasible because the resulting set grows exponentially. Therefore, combinatorial testing [7] can be applied to select a subset of all attribute configurations. Common 3D human pose metrics, such as mean error per joint position (mpjpe), Procrustes-aligned mean error per joint position (pa-mpjpe), and / or percentage of correct keypoints (pck) [4, 9, 6], can be used at different distance thresholds. Alternatively or additionally, image similarity metrics such as the Fréchet Inception Distance (FID) [3] and / or LPIPS
[12] can be used to measure the quality of the synthetic image data.
[0144] The techniques presented herein can be utilized for constrained image synthesis, such as with diffusion models. Specifically, data synthesis can be used for data augmentation and validation purposes for DL models, such as 2D / 3D pose estimators and / or pedestrian detectors. The use of constrained image synthesis is particularly beneficial when collecting additional (e.g., real) data is expensive and / or legally impossible (and / or likely) for privacy reasons.
[0145] References [1] Michael J. Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang. BEDLAM: A synthetic dataset of bodies exhibiting detailed lifelike animated motion. In Proceedings IEEE / CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pages 8726-8737, June 2023. [2] Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. 2022. [3] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, Guenter Klambauer, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a nash equilibrium. CoRR, abs / 1706.08500, 2017. [4] Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. CoRR, abs / 1712.06584, 2017. [5] Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1-248:16, October 2015. [5a] Silvia Zuffi, Angjoo Kanazawa, David W. Jacobs and Michael J. Black. 3D Menagerie: Modeling the 3D shape and pose of animals. CoRR, abs / 1611.07700, 2016 [6] Dushyant Mehta, Helge Rhodin, Dan Casas, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation using transfer learning and improved CNN supervision. CoRR, abs / 1611.09813, 2016. [7] Changhai Nie and Hareton Leung. A survey of combinatorial testing. ACM Comput. Surv., 43(2), feb 2011. [8] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjoern Ommer. High-resolution image synthesis with latent diffusion models, 2021. [9] Alexander Toshev and Christian Szegedy. Deeppose: Human pose estimation via deep neural networks. CoRR, abs / 1312.4659, 2013.
[10] Timo von Marcard, Roberto Henschel, Michael Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In European Conference on Computer Vision (ECCV), sep 2018.
[11] Lvmin Zhang and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models, 2023.
[12] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.
[13] Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet and Mohammad Norouzi. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. arXiv: 2205.11487 [cs.CV], 2022.
[14] Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu and Mark Chen. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv: 2204.06125 [cs.CV], 2022.
Claims
1. 1. A computer-implemented method (100) for generating synthetic image data usable for training, validation, and / or testing of downstream AI, particularly downstream neural networks (NNs), for body detection-related tasks based on sensor data, comprising: The method comprises: - receiving (S102) visual information related to a body, said visual information comprising a two-dimensional (2D) skeletal representation of said body, a 2D projected particularly dense semantic coding of said body and a 2D depth map of said body; - receiving (S104) a text prompt related to at least one of the appearance of said body and / or environmental information for said body; - generating (S106) synthetic image data of the body based on the received (S104) text prompt conditioned by the received (S102) visual information, said generation (S106) being performed by a conditioned image synthesis model; A method (100) comprising:
2. The method (100) further comprises: - determining (S101) the visual information related to the body, comprising a 2D skeletal representation of the body, a 2D projected particularly dense semantic coding of the body, and a 2D depth map of the body, by means of a generative body model, wherein determining the visual information (S101) comprises performing a projection of the generative body model onto an image plane. Including, The method (100) of claim 1.
3. The generative body model generates a model of the body based on a set of shape parameters and pose parameters, and optionally the generative body model comprises a skinned multi-person linear (SMPL) model; The method (100) of claim 2.
4. The method (100) further comprises: providing (S108) the generated (S106) synthetic image data of the body to the downstream AI, in particular the downstream NN, which is configured for a body detection related task on sensor data, and wherein providing (S108) the generated (S106) synthetic image data comprises providing the visual information, a set of shape and pose parameters of the generated body model, the generated body model on which the visual information is determined, and / or one or more related quantities as ground truth for the body detection related task. Including, The method (100) of any one of claims 1 to 3.
5. the conditioned image synthesis model includes a generative text-image model configured to generate (S106) the synthetic image data based on the received (S104) text prompt, and an image conditioning model configured to encode the received (S102) visual information to condition and / or control the generative text-image model; and optionally: the image conditioning model is configured to encode the received (S102) visual information in combination with the received (S104) text prompt; The method (100) of any one of claims 1 to 4.
6. The generated text-image model comprises a diffusion model, in particular a stable diffusion (SD) network, the SD network comprising a U-Net architecture having an encoder and in particular a skip-connected decoder. The method (100) of claim 5.
7. the image conditioning model includes a ControlNet, the ControlNet comprising an encoder and a convolutional layer having a cross-attention mechanism for the generated text-to-image model, in particular for the encoder of the SD network; 7. The method (100) according to claim 5 or 6.
8. 1. A computer-implemented method (200) for training a conditional image synthesis model to generate synthetic image data usable for training, validation, and / or testing of downstream AI, particularly downstream neural networks (NNs), for body detection-related tasks based on sensor data, comprising: The method comprises: - receiving (S202) a synthetic image training dataset, said synthetic image training dataset comprising: an image dataset, in particular 2D, of a body, annotated by a set of 3D body markers applied externally to said body when acquiring said image dataset, an image dataset, in particular 2D, of the body, annotated by a human expert, said annotations including 3D body pose information, and an image dataset, in particular 2D, of a body annotated by 2D keypoint information, wherein the training of the conditioned image synthesis model includes the generation of 3D pose information; and - training (S204) a conditioned image synthesis model based on the received (S202) synthetic image training dataset, wherein training (S204) of the conditioned image synthesis model comprises transforming annotations of the synthetic image training dataset into at least a part of visual information related to a body, the visual information comprising a two-dimensional (2D) skeletal representation of the body, a 2D projected extra-dense semantic coding of the body and a 2D depth map of the body, the extra-2D image dataset being considered as ground truth for the generated synthetic image data; A method (200) comprising:
9. 1. A computer-implemented method (300) for training, validating, and / or testing downstream AI, in particular a downstream neural network (NN), for performing body detection related tasks based on sensor data, comprising: The method comprises: - receiving (S302) synthetic image data of a body, the synthetic image data of the body being generated according to any one of the methods claims 1 to 7, the synthetic image data of the body comprising a generated body model, visual information and / or related quantities related to the body, the visual information comprising a two-dimensional (2D) skeletal representation of the body, a 2D projected particularly dense semantic coding of the body and a 2D depth map of the body, and / or the related quantities comprising information about the body derived from the generated body model and / or from the visual information; - training (S304) the downstream AI, in particular the downstream NN, based on the received (S302) synthetic image data of the body, wherein the generated body model, the visual information related to the body and / or the related quantities are considered as ground truth; The method (300).
10. the execution of the body detection related tasks is performed based on the received sensor data, the tasks including at least one of classification, semantic segmentation, and object detection, in particular body detection; 10. The method (300) of claim 9.
11. the sensor data is received from at least one of a video camera, a radar sensor, a LiDAR sensor, an ultrasonic sensor, a motion sensor, and a thermal imaging sensor; 11. The method (300) of claim 9 or 10.
12. said downstream artificial intelligence, in particular said downstream neural network, -Autonomous driving, - Robot movement planning, - operation of household appliances, and - Control of access control systems 12. Use of the method (300) according to any one of claims 9 to 11 for applying to at least one of:
13. 1. A computing device (400) for generating synthetic image data usable for training, validation, and / or testing of downstream AI, in particular downstream neural networks (NNs), for body detection related tasks based on sensor data, comprising: The computing device (400) a first input interface (402) configured to receive visual information related to a body, said visual information comprising a two-dimensional (2D) skeletal representation of said body, a 2D projected particularly dense semantic coding of said body, and a 2D depth map of said body; a second input interface (404) configured to receive text prompts related to at least one of the appearance of the body and / or environmental information relative to the body; a generation module (406) comprising a conditioned image synthesis model, the conditioned image synthesis model being configured to generate synthetic image data of the body based on the received text prompt conditioned by the received visual information; A computing device (400) comprising:
14. 1. A computing device (500) for training a conditional image synthesis model for generating synthetic image data usable for training, validation, and / or testing of downstream AI, in particular downstream neural networks (NNs), for body detection related tasks based on sensor data, comprising: The computing device (500) an input interface (502) configured to receive a synthetic image training dataset, said synthetic image training dataset comprising: an image dataset, in particular 2D, of a body, annotated by a set of 3D body markers applied externally to said body when acquiring said image dataset, an image dataset, in particular 2D, of the body, annotated by a human expert, said annotations including 3D body pose information, and - an image dataset, in particular 2D, of a body annotated by 2D keypoint information, wherein the training of the conditioned image synthesis model includes the generation of 3D pose information. an input interface (502) including at least one of: a training module (504) configured to train a conditioned image synthesis model based on the received synthetic image training dataset, wherein the training of the conditioned image synthesis model comprises transforming the annotations of the synthetic image training dataset into at least a portion of visual information related to a body, the visual information comprising a two-dimensional (2D) skeletal representation of the body, a 2D projected extra-dense semantic coding of the body, and a 2D depth map of the body, the extra-2D image dataset being considered as ground truth for the generated synthetic image data; A computing device (500) comprising:
15. 1. A computing device (600) for training, validating and / or testing downstream AI, in particular downstream neural networks (NNs), for performing body detection related tasks based on sensor data, comprising: The computing device (600) an input interface (602) configured to receive synthetic image data of a body, the synthetic image data of the body being generated according to any one of the methods claims 1 to 7, the synthetic image data of the body comprising a generated body model, visual information and / or related quantities related to the body, the visual information comprising a two-dimensional (2D) skeletal representation of the body, a 2D projected particularly dense semantic coding of the body and a 2D depth map of the body, and / or the related quantities comprising information about the body derived from the generated body model and / or from the visual information; a training module (604) configured to train the downstream AI, in particular the downstream NN, based on the received synthetic image data of the body, the generated body model, the visual information related to the body and / or the related quantities being considered as ground truth; A computing device (600) comprising: