Method and apparatus for generating synthetic image data of a person

By generating anonymous metadata through a multi-view camera system and generative artificial intelligence network, the high cost and data protection issues of generating synthetic human image data are solved, achieving anonymous and photorealistic image generation.

CN122199697APending Publication Date: 2026-06-12ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ROBERT BOSCH GMBH
Filing Date
2025-12-12
Publication Date
2026-06-12

Smart Images

  • Figure CN122199697A_ABST
    Figure CN122199697A_ABST
Patent Text Reader

Abstract

Method and device for generating synthetic image data of a person. A method, for example a computer-implemented method, for generating synthetic image data of a person, the method having: providing first information, the first information characterizing anonymous metadata relating to at least one person imaged on at least one input image; providing an artificial neural network, the artificial neural network being designed to generate synthetic image data based on at least the first information; generating, by means of the neural network, synthetic image data based on at least the first information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method for generating synthetic image data of people.

[0002] This disclosure also relates to an apparatus for generating synthetic image data of a person. Summary of the Invention

[0003] Some examples involve a method for generating synthetic image data of people, such as a computer-implemented method, which includes: providing first information, the first information representing anonymized metadata associated with at least one person imaged on at least one input image; providing an artificial neural network designed to generate synthetic image data based at least on the first information; and generating synthetic image data by means of the neural network, at least based on the first information. In this way, in some examples, anonymized and / or photorealistic synthetic image data of people can be generated, for example, in the form of one or more images, which can be used, for example, to train other methods and / or systems, such as machine learning-based methods and / or systems. Thus, in some examples, determining the corresponding real image data with people can be avoided, which can be complex and costly, and may also raise data protection issues.

[0004] For example, the first information has at least one of the following elements: a) depth information associated with at least one person; or b) pose information associated with at least one person, the pose information being, for example, characterizing the pose of the at least one person.

[0005] In some examples, artificial neural networks are generative artificial intelligence systems, such as stable diffusion networks or generative adversarial networks.

[0006] In some examples, the method is specified to include: creating at least one input image, for example, using a multi-view camera system; and determining the first information based on the at least one input image, wherein the at least one person is presented or assumes, for example, a pose, or at least two different poses, on different input images. In some examples, this allows for the relatively flexible provision of a large number of different synthetic image data, each containing at least one person, for example, at least one person in different poses, for example, presented in the form of multiple images, such as digital images.

[0007] A multi-view camera system, for example, consists of multiple cameras that are pointed at the same object or scene from different perspectives. These multiple cameras may acquire images simultaneously, and these images can then be combined to form a more comprehensive representation. In some examples, a multi-view camera system may be able to acquire depth information.

[0008] In other examples, the method includes deleting at least one input image, for example, deleting at least one input image after determining the first information, or deleting at least one input image immediately after determining the first information. In this way, in some examples, a particularly high level of data protection is ensured.

[0009] For example, it may be determined that the first information includes: determining depth information associated with the at least one person in the form of at least one depth image; and / or determining that the first information includes: determining pose information associated with the at least one person in the form of at least one pose image. In other examples, this enables efficient processing of information in the form of image data.

[0010] In some examples, the method may be proposed to: determine depth information associated with the environment of at least one character; and optionally, supplement, for example enhance, the depth information associated with the environment of at least one character based on at least one virtual model of the at least one character. Thus, for example, particularly detailed depth information can be provided.

[0011] For example, the method includes: using a marker-based method to create the at least one input image and / or to determine the first information; and optionally, removing image information associated with the markers of the marker-based method from the first information, for example, replacing the image information associated with the markers of the marker-based method from the first information with other image information, for example, by means of an image inpainting method.

[0012] In some examples, this image inpainting method is a way to reconstruct or fill in missing or damaged image regions in a digital image (e.g., regions from which image information associated with tags has been previously removed). For example, in this image inpainting method, missing or removed regions are supplemented by surrounding image information, so that these regions, for example, blend seamlessly into the entire image. In some examples, this image inpainting method may use algorithmic methods or machine learning, for example, to generate realistic new image information or content, and, for example, to ensure that the newly generated image regions are consistent with the rest of the image.

[0013] For example, the method has: Provide at least one artificial neural control network for a neural network, wherein the at least one artificial neural control network is designed to: receive the first information, for example in the form of at least one depth image and / or in the form of at least one pose image, and control the operation of the neural network based on the received first information; and generate image data, for example in the form of multiple photorealistic synthetic images, by means of the neural network under the control of the at least one control network.

[0014] In some examples, the at least one artificial neural control network may, based on the first information, generate or determine input parameters such as text prompts and / or other input parameters for a stable diffusion network, and feed them to the stable diffusion network to generate image data.

[0015] In other examples, the functionality of the at least one artificial neural control network can also be integrated into the stable diffusion network, for example.

[0016] Some examples relate to an apparatus for generating synthetic image data of people, wherein the apparatus is designed to perform the methods described in accordance with this disclosure.

[0017] Other examples relate to a computer-readable storage medium that includes instructions that, when executed by a computer, cause the computer to perform the methods described in accordance with this disclosure.

[0018] Other examples relate to a computer program that includes instructions that, when executed by a computer, cause the computer to perform the methods described in accordance with this disclosure.

[0019] Other examples involve a data carrier signal that transmits and / or characterizes a computer program as described in this disclosure.

[0020] In some examples, the methods and / or devices and / or computer-readable storage media and / or computer programs and / or data carrier signals described in this disclosure are proposed for use in at least one of the following elements: a) generating anonymous, e.g., photorealistic synthetic images of persons, such as persons with a specifyable pose; or b) providing training data, such as for training other neural networks; or c) providing test data, such as for testing image evaluation methods, such as image-based personal data analysis methods; d) avoiding the generation of real image data of persons; or e) complying with data protection regulations.

[0021] Other features, applications, and advantages derive from the subsequent description of examples presented in the accompanying drawings. All features described or shown herein, either alone or in any combination, form the subject matter of this disclosure, regardless of their generalization in the claims or their references thereto, and regardless of their expression or presentation in the specification or drawings. Attached Figure Description

[0022] In the attached diagram: Figure 1 A simplified flowchart is shown schematically; Figure 2 A simplified flowchart is shown schematically; Figure 3 A simplified flowchart is shown schematically; Figure 4 A simplified flowchart is shown schematically; Figure 5 A simplified flowchart is shown schematically; Figure 6 A simplified flowchart is shown schematically; Figure 7 A simplified block diagram is shown schematically; Figure 8 A simplified flowchart is shown schematically; Figure 9 A simplified flowchart is shown schematically; Figure 10 An example of its use is illustrated. Detailed Implementation

[0023] For some examples, see Figure 1 , Figure 2 The present invention relates to a method for generating synthetic image data BD of persons, such as a computer-implemented method, comprising: providing 100 first information I-1, the first information representing anonymous metadata associated with at least one person P1, P2 imaged on at least one input image EB; and providing 102 artificial neural network NN, the artificial neural network being designed to generate synthetic image data BD based at least on the first information I-1. Using this neural network NN, at least based on the first information I-1, 104 synthetic image data BDs are generated. In this way, in some examples, anonymized and / or photorealistic synthetic image data BDs of people can be generated, for example, in the form of one or more images. These synthetic image data can be used, for example, to train other methods and / or systems (not shown), such as machine learning-based methods and / or systems. Thus, in some examples, it is possible to avoid determining the corresponding real image data with real people, which can be complex and costly, and may also raise data protection issues.

[0024] For example, see Figure 2 The first information I-1 has at least one of the following elements: a) depth information IT associated with at least one character P1, P2; or b) pose information IP associated with at least one character P1, P2, the pose information being, for example, a representation of the pose of the at least one character P1, P2.

[0025] In some examples, the artificial neural network NN is a generative artificial intelligence system, such as a stable diffusing network or a generative adversarial network (“GAN”).

[0026] In some examples, depth information IT may exist or be provided, for example, in the form of at least one depth image I-T'. In some examples, pose information IP may exist or be provided, for example, in the form of at least one pose image I-P'. In other examples, other forms of representation of information IT and / or IP are also possible.

[0027] In some examples, "anonymous metadata" refers to metadata that is anonymized and therefore can no longer be associated with people P1, P2 in the input image EB, for example, beyond the depth information IT associated with at least one person and / or the pose information IP associated with at least one person. This allows, for example, the use of relevant depth and / or pose information without using personally relevant data associated with the at least one person P1, P2 (e.g., in the sense of data protection law).

[0028] In some examples, see Figure 3 The method specifies that it has the following characteristics: creating at least one input image EB, for example, by means of a multi-view camera system MV-KS (…). Figure 2 Based on the at least one input image EB, the first information I-1 is determined 112, wherein, for example, the at least one person P1, P2 is presented or adopts at least two different poses on different input images EB. In some examples, this allows for relatively flexible provision of a large number of different synthetic image data BD, each containing at least one person, for example, at least one person in different poses, for example, presented in the form of multiple images, such as digital images. In some examples, a pose may also be specified for different input images EB, for example, only one pose may be specified.

[0029] MV-KS multi-view camera system, see Figure 2For example, it can consist of multiple cameras that are pointed at the same object or scene from different perspectives (in the current case, for example, two people P1 and P2 inside a vehicle). These multiple cameras can, for example, simultaneously acquire images, which can be combined to form a more comprehensive representation. See some examples. Figure 2 The multi-view camera system MV-KS, for example, is capable of acquiring depth information. Optionally, the multi-view camera system MV-KS may have at least one depth sensor TS ( Figure 2 ).

[0030] In some examples, see Figure 3 The method includes: deleting 114 at least one input image EB, for example, deleting at least one input image after determining 112 the first information I-1, or for example, deleting at least one input image immediately after determining 112 the first information I-1. In this way, in some examples, a particularly high level of data protection is ensured.

[0031] For example, see Figure 3 The determination 112 first information I-1 has: determination 112a depth information IT associated with at least one person P1 in the form of at least one depth image I-T'; and / or determination 112 first information I-1 has: determination 112b pose information IP associated with at least one person P1, P2 in the form of at least one pose image I-P'. In other examples, this enables efficient processing of information in the form of image data.

[0032] exist Figure 2 In this context, block E1 represents, for example, determining depth information IT, for instance, in the form of at least one depth image I-T', based on at least one input image EB, such as using a method based on the principle of Neural Radiation Fields (“NeRF”). Neural Radiation Fields enable, for example, the representation and / or reconstruction of complex 3D scenes based on 2D image data. Here, by using neural networks, the 3D scene is modeled, for example, as a continuous volumetric field in which each point of the scene is assigned specific color and density information. In some examples, the NeRF principle can be used to determine the depth information IT in block E1.

[0033] See other examples. Figure 2 Other methods for determining depth information (IT) are also conceivable, such as using the output data of at least one depth sensor (TS).

[0034] exist Figure 2In this context, block E2 represents, for example, determining pose information IP based on at least one motion-capture method, for example, in the form of at least one pose image I-P'.

[0035] In some examples, see Figure 4 It can be proposed that the method has the following characteristics: determining the environment UM of 120 and at least one person P1, P2. Figure 2 The associated depth information IT-UM; and optionally, based on at least one virtual model of the at least one character P1, P2 (see below, according to Figure 9 The element E39 supplements, for example, enhances, the depth information IT-UM associated with the environment UM of at least one character P1, P2 by 122, or by 122a. Thus, for example, particularly detailed depth information I-T'' can be provided, for example, in the form of a depth map.

[0036] In some examples, see Figure 2 The environment UM of the people P1 and P2 in the input image EB is, for example, the interior space of a vehicle. Therefore, the principles of this disclosure can be used, for example, to generate synthetic, photorealistic, anonymized image data of one or more people in the interior space of a vehicle, such image data, for example, for training or testing driver monitoring and / or occupant monitoring systems. However, the principles of this disclosure are not limited to the interior space of a vehicle, but can also be applied—without limiting generality—to other possible environments UM of the people P1 and P2.

[0037] For example, see Figure 5 The method includes: creating 110 the at least one input image EB using a marker-based method 130 and / or determining 112 the first information I-1; and optionally, removing 132 image information associated with the markers of the marker-based method from the first information I-1, for example, replacing 132a image information associated with the markers of the marker-based method from the first information I-1 with other image information, for example, by means of an image inpainting method 132b. For example, it can be specified that the markers of the marker-based method are removed from depth images, which may constitute at least a portion of the first information I-1.

[0038] In some examples, the image restoration method 132b is a method for reconstructing or filling missing or damaged image regions in a digital image (e.g., regions from which image information associated with markers has been previously removed). For example, in this image restoration method, missing or removed regions are supplemented by surrounding image information, such that these regions are seamlessly integrated into the overall image. In some examples, the image restoration method 132b may use algorithmic methods or machine learning, for example, to generate realistic new image information or content, and to ensure, for example, that the newly generated image regions are consistent with the rest of the image. Thus, in some examples, it can be ensured that the generated image data BD does not contain any information that is itself undesirable and may be related to the markers used to determine the first information I-1.

[0039] For example, see Figure 6 The method comprises: providing at least 140 artificial neural control networks NN-K1 for a neural network NN, wherein the at least one artificial neural control network NN-K1 is designed to: receive the first information I-1, for example in the form of at least one depth image I-T' ( Figure 2 The operation of the neural network NN is controlled based on the received first information I-1 and / or in the form of at least one pose image I-P'. Under the control of the at least one control network NN-K1, 142 image data BD is generated by means of the neural network NN, for example in the form of multiple photorealistic synthetic images B1, B2, ...

[0040] In some examples, see Figure 2 , Figure 6 The at least one artificial neural control network NN-K1 can, based on the first information I-1, generate or determine input parameters such as text prompts and / or other input parameters of the stable diffusion network NN, and feed them to the stable diffusion network NN to generate image data BD.

[0041] In cases where the network NN is designed as a GAN type according to other examples, the at least one artificial neural control network NN-K1 can be omitted or adapted to a GAN network.

[0042] See other examples. Figure 2 The functionality of at least one artificial neural control network NN-K1 can also be integrated into the stable diffusion network NN, for example.

[0043] For some examples, see Figure 7 The present invention relates to an apparatus 200 for generating synthetic image data BD of a person, wherein the apparatus 200 is designed to perform the method described in accordance with the present disclosure.

[0044] In some examples, see Figure 7 The device 200 is defined as having: a computing device (“Computer”) 202 having at least one computing core 202a; and a storage device 204 allocated to the computing device 202 for at least temporarily storing at least one of the following elements: a) data DAT (e.g., data associated with first information I-1 and / or input image EB and / or image data BD); b) a computer program PRG, for example for performing the methods described in accordance with this disclosure.

[0045] See other examples. Figure 7 The storage device 204 has: volatile memory (e.g., working memory (RAM)) 204a; and / or non-volatile (NVM) memory (e.g., flash EEPROM) 204b; or a combination thereof or a combination with other memory types not explicitly mentioned.

[0046] For other examples, see Figure 7 The present invention relates to a computer-readable storage medium SM comprising instructions PRG that, when executed by a computer 202, cause the computer to perform the method described herein.

[0047] For other examples, see Figure 7 The present invention relates to a computer program PRG comprising instructions which, when executed by a computer 202, cause the computer to perform the method described herein.

[0048] For other examples, see Figure 7 This relates to a data carrier signal (DCS) that transmits and / or represents a computer program (PRG) according to the present disclosure. The DCS can be transmitted (e.g., transmitted and / or received) via, for example, an optional data interface 206 of the device 200.

[0049] In some examples, see Figure 7 The device 200 can also be formed purely based on hardware.

[0050] Subsequently, other aspects and examples are described, which—in other examples—can be combined with at least one of the above aspects and / or examples, either individually or in any combination of each other.

[0051] In some examples, the principles of this disclosure enable the generation of photorealistic images BD, B1, B2, ... in a data-driven manner. Figure 2 , Figure 6Furthermore, it also complies with data protection regulations because it synthesizes non-existent virtual characters, for example, only synthesizing non-existent virtual characters. This can be achieved, for example, by using a stable diffusion network (NN). Figure 2 Image synthesis is performed, and anonymous metadata (e.g., motion capture, depth data) of a person in a controlled environment UM (such as inside a vehicle) is automatically generated, for example, in the form of the first information I-1.

[0052] In some examples, the process of generating training data according to this disclosure, such as for a specific scenario like "camera-based person recognition and pose estimation in a vehicle interior space," can be—without limiting its generality—as follows: Multiple cameras, such as a multi-view camera system MV-KS, are installed and calibrated in the vehicle. Figure 2 Multiple cameras. One or more persons P1, P2 are equipped with markers (not shown) on their bodies (e.g., according to known motion capture methods). Persons P1, P2 with these markers are photographed, for example, recorded, in a vehicle, for example, in different poses by cameras of a multi-view camera system MV-KS, which, for example, is obtained according to... Figure 2 The input images are EB. Then, for example, depth information IT is determined from these input images EB, and / or markers in these input images EB are detected and, for example, triangulated, where, for example, 3D point data is obtained, representing the corresponding positions of these markers. In some examples, these input images EB are then no longer needed and are deleted.

[0053] In some examples, for instance, with the aid of 3D labeled data, the corresponding poses of the photographed figures P1 and P2 (e.g., joint rotations and / or positions representing components of a virtual skeleton) can be determined, for instance, calculated, resulting in pose information IP.

[0054] In other examples, depth information IT and / or, for example, 3D pose information IP, is converted to the desired target camera (or to the viewpoint of such a target camera) and then used, for example, as input to at least one control network NN-K1 of the diffusion network NN. In some examples, the diffusion network generates arbitrary new images of people, for example, in the same pose and in an environment semantically similar to the environment UM of the input image EB, but without using the personal data of people P1, P2 of the input image EB, or without the arbitrary new images of these people containing the personal data of people P1, P2 of the input image EB.

[0055] Therefore, in some examples, a wide variety of realistic, photorealistic, and personally relevant training data (i.e., synthetic images containing non-real people) can be generated with relatively little effort (e.g., at most a partial amount of manual work) (e.g., processing of markers) using the principles of this disclosure. Furthermore, the generated image data BDs are not subject to any data protection and / or licensing restrictions because they are synthetically generated. Compared to some conventional, for example, fully synthetic data generation methods using computer graphics, the principles of this disclosure enable, for example, complex interactions between people and scenes and / or other people on the image data BDs, and, for example, enable so-called "corner-case" data generation, where, for example, a person is typing on a mobile phone or putting on a coat. In this context, "corner-case" data generation refers to creating special datasets, in the form of image data BDs depicting rare or unexpected situations. In real-world applications, these scenarios are only occasional, but can be important for the robustness and accuracy of the model.

[0056] In some examples, see Figure 2 The principles of this disclosure are based on a multi-view camera system MV-KS with an optional depth sensor TS, which is mounted, for example, in a closed target environment UM (e.g., the interior space of a vehicle). For example, the multi-view camera system MV-KS can be initially calibrated with extrinsic and intrinsic parameters to determine the position, orientation, and projected imaging of the cameras relative to each other in the 3D environment UM.

[0057] For illustrative purposes, the following statements primarily—but not generally—concern specific use cases for generating image data within the interior space of a vehicle. However, this does not limit the application of the methods described in this disclosure to other application areas, such as in partially enclosed environments.

[0058] In some examples, see Figure 2 The cameras of the multi-view camera system MV-KS are mounted at different locations within the vehicle and, for example, calibrated once. Real-life individuals P1 and P2 wear, for example, tagged suits used in marker-based motion capture methods, where the markings on these suits can be detected by the cameras of the MV-KS system. Alternatively, in other examples, marker-free motion capture methods or systems can be used, which, for example, directly estimate the poses of individuals P1 and P2 based on multiple camera images using artificial intelligence algorithms. However, in some examples, marker-based motion capture systems may be more stable than marker-free methods and, for example, provide more robust results when occluded.

[0059] In some examples, methods or systems based on accelerometers and / or any other form of motion capture can also be applied, for example, as an alternative to or supplement to the methods described above.

[0060] For example, at the beginning, characters P1 and P2 ( Figure 2 They were photographed wearing their tagging suits in an initial calibration pose, for example, to calibrate the motion capture system, such as to determine the size of the pose skeleton and optionally to determine the association between the tag and the skeleton.

[0061] In other examples, such as after capturing a standard pose, images and / or video recordings of individuals P1 and P2 within a vehicle are created using cameras mounted on a multi-view camera system MV-KS, for example, in different poses and / or interactions. The image data acquired in this way, such as the input image EB, is then evaluated using motion capture methods (e.g., based on markers), and, for example, for each image or each video frame, 3D human pose data is output, such as at least similar to pose information IP or I-P' (see below). Figure 3 (Block 112). In some examples, the evaluation of the motion capture method includes, on the one hand, marker detection, and on the other hand, the assignment of markers to the input image EB and the association of markers with the captured persons P1, P2. In other examples using a markerless motion capture system, the above-mentioned marker-related aspects can be omitted.

[0062] In some examples, additional 3D information (e.g., depth information IT, I-T'), such as the 3D information of the scene, can be determined based on existing image data, i.e., input image EB, for example, by means of photogrammetry. This can be achieved using different techniques, such as classical multi-view stereo methods or neural radiation field (NeRF) based methods.

[0063] Then, i.e., after determining the pose information IP, I-P' and depth information IT, I-T', the captured image and / or video data, such as the input image EB, can be deleted. These markers are cropped from the depth or 3D data, for example, and the resulting gaps are refilled, for example, using image inpainting methods. This step is optional and can be omitted in some examples, such as when using a markerless motion capture system that works without markers (e.g., spheres), such as three-dimensional markers, or works using only two-dimensional markers (e.g., coded patterns), such as two-dimensional markers on clothing.

[0064] Alternatively or supplementally, an input image EB designed as a color input image may, in some examples, be transformed, for example, directly into a depth image by means of a so-called “Depth-from-Mono” AI (artificial intelligence) algorithm, without the use of photogrammetry and / or, for example, additional dedicated depth sensors.

[0065] In some examples, such as instead of directly converting the depth information containing characters P1 and P2, for example under the same settings, the environment UM can be determined first. Figure 2 The depth information of the environment UM (e.g., an empty vehicle or the interior space of a vehicle) can then be supplemented, for example, enhanced, by using a virtual character model. See also Figure 4 These virtual character models are positioned, for example, based on determined pose information IP, I-P' relative to the vehicle's interior space. In some examples, depth information determined in this way, such as fused depth information, can cover various body types, for example, in the form of a depth map.

[0066] For example, given depth information (IT) and pose information (IP) relative to a 3D coordinate system, a depth map can be generated, for example, from a new camera perspective (e.g., the target camera perspective). Similarly, the pose information (IP) of characters P1 and P2 can be transformed, for example, "morphed," to the new camera perspective. In some examples, the depth map and pose information thus generated can be, for example, in the form of an image (e.g., see also...). Figure 2 The elements I-T', I-P') are passed as input to at least one, for example, a pre-trained control network NN-K1, which controls, for example, a similarly pre-trained neural diffusion network NN, wherein, for example, the neural diffusion network NN is designed to generate image data based on initial noise and text cues, for example, under the control of the control network NN-K1. Thus, in some examples, based on the input depth map and pose information to the neural diffusion network NN, any number of image data BDs of virtual characters in different environments UM, for example, each in the same pose (e.g., characterized by pose information IP), can be generated, for example, in the form of new synthetic images, for example, by changing the text cues and / or noise, based on which the neural diffusion network NN generates these image data BDs.

[0067] In some examples, the principles of this disclosure may be provided, for example, as an additional feature of a camera or camera system, such as for a security camera, or in camera-based automotive applications, such as anywhere that personal data is analyzed based on images.

[0068] In some examples, the device 200 ( Figure 7 ) or its functions can also be integrated, for example, into the multi-view camera system MV-KS ( Figure 2 )middle.

[0069] Figure 8 The schematic diagram illustrates a simplified flowchart according to certain examples. Element E10 represents, for example, a multi-view camera system (see also...). Figure 2 The elements are: (MV-KS) multiple calibrated cameras, and element E11 represents an optional depth sensor. Element E12 represents one or more input images acquired by means of camera E10. Element E13 represents: determining depth information IT based on input image E12, for example, using the NeRF method. Element E14 represents: determining pose information IP based on input image E12 using a motion capture method. Element E15 represents the depth information determined according to element E13, for example, in the form of a depth image or depth map and / or 3D point cloud. Element E16 represents: optionally, tag repair based on depth information E15 and pose information E17. For example, this optional repair E16 is performed when the capture method according to element E14 is based on the tag, i.e., when information E15 and E17 contain information that may be associated with the tag. Element E18 indicates that, optionally, information from elements E15, E16, and E17 is transformed, for example, "warped," to the target camera, from the viewpoint of that target camera, for example, to generate the desired image data BD; see also element E21. Element E19 indicates a stable diffusing artificial neural network designed to be based on input information from element E18 and, if necessary, on other input information E20 such as text prompts and / or seed values ​​(see also...). Figure 2 Optional control information (I-ST) is used to generate one or more synthetic image data sets, such as image dataset E21. In some examples, as described above and as... Figure 8 The process illustrated in the example can be used to generate any number of synthetic anonymized image datasets E21 based on one or more input images, for example, with real people P1, P2, each of which has synthetically generated people, which have, for example, similar poses to the real people in the input images, and / or are in similar surrounding environments or similar environments UM.

[0070] Figure 9 The schematic diagram illustrates a simplified flowchart according to certain examples. Element E30 represents, for example, a multi-view camera system (see also...). Figure 2 Multiple calibrated cameras (MV-KS elements), for example, at least similar to those according to Figure 8 The element E10, and according to Figure 9Element E31 represents an optional depth sensor. Element E32 represents one or more first input images of a specifyable environment (e.g., the interior space of an empty vehicle), which are acquired by means of camera E30. Element E33 represents determining first depth information based on the first input image E32, for example, using the NeRF method; and element E34 represents the depth information acquired in this way, for example, in the form of a 3D point cloud. Element E35 represents one or more second input images of a specifyable environment, different from those with at least one person P1, P2 (in the interior space of the vehicle) present. Figure 2 The first input image E32 is one or more second input images obtained by means of camera E30. Element E36 indicates that, based on the second input image E35, for example, in the case of using the NeRF method, second depth information is determined; and element E37 indicates that, mainly based on the second input image E35, a marker-based motion capture method is applied to determine 3D pose information.

[0071] according to Figure 9 Element E38 represents aspects of enhancing depth information when using the digital models E39 of at least one character P1, P2 and 3D pose information E41, and optionally transforming, for example, "warping," to a specified target camera viewpoint, see target camera calibration information E40, for example, similar to... Figure 8 Element E18. According to Figure 9 The elements E42, E43, and E44 correspond to the following according to Figure 8 The elements are E19, E20, and E21.

[0072] In some examples, see Figure 10 The present disclosure proposes the use 300 of the method and / or device 200 and / or computer-readable storage medium SM and / or computer program PRG and / or data carrier signal DCS for at least one of the following elements: a) generating 301 anonymized, e.g., photorealistic synthetic images of people, such as people with a specifyable pose; or b) providing 302 training data, such as for training other neural networks; or c) providing 303 test data, such as for testing image evaluation methods, such as image-based personal data analysis methods; d) avoiding 304 generating real image data of people; or e) complying with 305 data protection regulations.

Claims

1. A method for generating synthetic image data (BD) of a person, such as a computer-implemented method, the method comprising: providing (100) first information (I-1), the first information representing anonymous metadata related to at least one person (P1, P2) imaged on at least one input image (EB); providing (102) an artificial neural network (NN), the artificial neural network being designed to generate the synthetic image data (BD) based at least on the first information (I-1); and generating (104) the synthetic image data (BD) by means of the neural network (NN), based at least on the first information (I-1).

2. The method according to claim 1, wherein, The first information (I-1) has at least one of the following elements: a) depth information (IT) associated with the at least one person (P1, P2); or b) pose information (IP) associated with the at least one person (P1, P2), the pose information being, for example, characterizing the pose of the at least one person (P1, P2).

3. The method according to at least one of the preceding claims, wherein, The artificial neural network (NN) is a generative artificial intelligence system, such as a stable diffusion network or a generative adversarial network.

4. The method according to at least one of the preceding claims, the method comprising: creating (110) the at least one input image (EB), for example by means of a multi-view camera system (MV-KS); and determining (112) the first information (I-1) based on the at least one input image (EB).

5. The method according to claim 4, wherein the method comprises: deleting (114) the at least one input image (EB), for example, deleting the at least one input image after determining (112) the first information (I-1).

6. The method according to claim 4 or 5, wherein a) determining (112) the first information (I-1) has: determining (112a) depth information (IT) associated with the at least one person (P1, P2) in the form of at least one depth image (I-T'); and / or b) determining (112) the first information (I-1) has: determining (112b) pose information (IP) associated with the at least one person (P1, P2) in the form of at least one pose image (I-P').

7. The method according to at least one of the preceding claims, the method comprising: determining (120) depth information (IT-UM) associated with the environment (UM) of the at least one character (P1, P2); and optionally, supplementing (122), for example enhancing (122a), the depth information (IT-UM) associated with the environment (UM) of the at least one character (P1, P2) based on at least one virtual model of the at least one character (P1, P2).

8. The method according to any one of claims 4 to 7, the method comprising: using (130) a tag-based method to create (110) the at least one input image (EB) and / or determining (112) the first information (I-1); and optionally, removing (132) image information associated with the tag of the tag-based method from the first information (I-1), for example replacing (132a) the image information associated with the tag of the tag-based method from the first information (I-1) with other image information, for example by means of an image restoration method (132b).

9. The method according to at least one of the preceding claims, the method comprising: providing (140) at least one artificial neural control network (NN-K1) for the neural network (NN), wherein, The at least one artificial neural control network (NN-K1) is designed to: receive the first information (I-1), for example in the form of at least one depth image (I-T') and / or in the form of at least one pose image (I-P'), and control the operation of the neural network (NN) based on the received first information (I-1); and, under the control of the at least one control network (NN-K1), generate (142) the image data (BD) by means of the neural network (NN), for example in the form of multiple photorealistic synthetic images (B1, B2, ...).

10. An apparatus (200) for generating synthetic image data (BD) of a person, wherein, The device (200) is designed to perform the method according to at least one of the preceding claims.

11. A computer-readable storage medium (SM) comprising instructions (PRG) that, when executed by a computer (202), cause the computer to perform the method according to at least one of claims 1 to 9.

12. A computer program (PRG) comprising instructions that, when executed by a computer (202), cause the computer to perform the method according to at least one of claims 1 to 9.

13. A data carrier signal (DCS) that transmits and / or characterizes the computer program (PRG) according to claim 12.

14. The method (300) according to at least one of claims 1 to 9 and / or the device (200) according to claim 10 and / or the computer-readable storage medium (SM) according to claim 11 and / or the computer program (PRG) according to claim 12 and / or the data carrier signal (DCS) according to claim 13 for use (300) of at least one of the following elements: a) generating (301) an anonymous, for example, photorealistic synthetic image of a person, such as a person with a specified pose; or b) providing (302) training data, for example for training other neural networks; or c) providing (303) test data, for example for testing image evaluation methods, such as image-based personal data analysis methods; d) avoiding (304) generating real image data of a person; or e) complying with (305) data protection regulations.