Computer-implemented method and device for determining a spatial arrangement of a three-dimensional object model of an object in a space

The method and device use neural networks and PAFs to determine the spatial arrangement of 3D object models from multiple images, addressing VR's hardware-related issues by providing accurate and immersive interactions without additional hardware.

WO2025223992A1PCT designated stage Publication Date: 2025-10-30UNWARE GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/060656
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2025-04-17
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current virtual reality systems rely on wearable hardware for head and body movement detection, which is uncomfortable, socially isolating, and can cause motion sickness due to latency and inaccuracy, and depth cameras limit frame rates and processing complexity.

Method used

A method and device that determine the spatial arrangement of a three-dimensional object model using a large number of initial images from different positions, employing trained neural networks to assign and triangulate two-dimensional and three-dimensional positions without additional hardware, utilizing Part Affinity Fields (PAFs) for connection and a DBSCAN algorithm for clustering.

Benefits of technology

Enables precise and rapid object capture in VR environments without hardware, allowing for accurate and immersive interactions by determining the spatial arrangement of objects like humans with high frame rates and reduced latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025060656_30102025_PF_FP_ABST
    Figure EP2025060656_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method (100) and a device (200) for determining a spatial arrangement of a three-dimensional object model of an object in a space, the object having a plurality of object portions. The method (100) comprises: providing (110) a plurality of first two-dimensional captured images of the object in the space, which were captured at a first point in time, wherein the first captured images were captured from predetermined different capture positions and the first captured images at least partially comprise the object portions; determining (120) two-dimensional actual positions (201) of object portions of the object in the plurality of first captured images, and associating the 2D actual positions with the object portions using a first trained neural network; determining (130) three-dimensional actual positions of the object portions on the basis of the plurality of first captured images, the 2D actual positions (201) and the capture positions; creating (140) a three-dimensional object model on the basis of the determined 3D actual positions and the associations; and determining (150) a spatial arrangement of the 3D object model in the space on the basis of the determined 3D actual positions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Computer-implemented method and device for determining a spatial arrangement of a three-dimensional object model of an object in a space

[0002] The invention relates to a computer-implemented method and a device for determining a spatial arrangement of a three-dimensional object model of an object in a space, in particular for the application of virtual reality (VR).

[0003] Virtual reality (VR), a computer-generated reality with images and optional sound, can be realized today using various technologies. Currently, compact VR headsets or room-sized CAVEs (Cave Automatic Virtual Environments) are typically used. In both cases, additional hardware is required, which the user wears. This hardware serves to display images on head-mounted displays, track viewing angles and positions using tracking glasses, output sound via headphones, and allow interaction with controllers in virtual worlds. However, the use of additional hardware makes prolonged use of the technology very uncomfortable, tiring, and, in the case of VR headsets, socially isolating. Furthermore, wearing hardware on the body poses hygiene problems, and wearing headsets can cause motion sickness.

[0004] Virtual and augmented reality applications (product configurators, pre-visualization of buildings, gaming, etc.) and other applications require fast and precise real-time detection of three-dimensional head and body movements of one or more users.

[0005] A crucial aspect of immersive presentation and interaction with virtual worlds is user recognition and tracking. This is currently achieved through wearable hardware components (active or passive trackers), as existing systems without such hardware fall far short of meeting the quality requirements of wearable devices. Specifically, head position detection requires low latency, low noise, and high accuracy across a wide range of situations. Failure to meet these criteria can result in dizziness, discomfort, or nausea, as the visual content cannot be quickly or accurately adjusted to the user's perspective in real time.

[0006] Another important aspect of creating compelling virtual realities is interaction within the digital world, both to capture user input and to navigate within it. It is equally important that this input is recognized as quickly and accurately as possible, otherwise a truly immersive experience will not be achieved.

[0007] Well-known systems for detecting head position and body movements include Azure Kinect Body Tracking SDK, Nuitrack SDK, and OpenPose. Apart from OpenPose, these solutions rely on the use of individual depth cameras to enable 3D body movement detection for various applications. Examples of applications for such technologies include human interaction, posture analysis, patient monitoring, and instruction for sports and fitness. Additionally, current technology often limits the frame rate because either the depth camera used does not support higher frame rates (due to technological limitations) or the underlying detection and processing process is too complex.

[0008] It is an object of the present invention to provide a computer-implemented method and a device that improves upon or even eliminates one or more of the aforementioned disadvantages. In particular, it is an object of the present invention to determine the three-dimensional arrangement of an object such as a person in a space quickly, precisely, and resource-efficiently without the need for hardware.

[0009] According to a first aspect, the task is solved by a computer-implemented method for determining the spatial arrangement of a three-dimensional object model of an object in space, where the object has several object sections. The method comprises:

[0010] Providing a large number of initial two-dimensional images of the object in space, taken at a first time point, wherein the initial images were taken from predetermined different recording positions and the initial images at least partially show the object sections;

[0011] Determining the two-dimensional actual positions of object sections in the multitude of initial image captures and assigning the 2D actual positions to the object using a first trained neural network;

[0012] Determining three-dimensional actual positions of the object sections based on the multitude of initial image captures, the 2D actual positions, and the capture positions;

[0013] Creating a three-dimensional object model based on the specified 3D lst positions;

[0014] Determining the spatial arrangement of the 3D object model in space based on the defined 3D LST positions. Using the method described above, the 3D object model and its spatial arrangement in space can be determined quickly and precisely without additional hardware on the object itself. Particularly when used for VR, this method enables precise and rapid object capture, for example, of a person and their head position. The VR environment can be augmented and / or mixed virtual reality.

[0015] Determining the spatial arrangement of the 3D object model in space can be understood as follows: for example, a computer perceives space as digital data and creates a digital copy of the object within this digital space. Thus, the object and, for example, its movement within the digital space can be recreated and / or copied.

[0016] The 3D object model can be a digital representation of the object. The 3D object model can be a digital twin of the object. The object model can be a virtual replica of the object.

[0017] Assigning the 2D LST positions to the object can involve defining Part Affinity Fields (PAFs) to connect the specified 2D LST positions. In particular, when using this method for two or more objects with respective object segments, the object segments can thus be assigned to the correct objects.

[0018] PAFs can be trained to provide a set of flow fields that encode unstructured pairwise relationships between object segments.

[0019] The creation of the three-dimensional object model can be further based on the PAFs.

[0020] The object can be a person, a living object, a non-living object, a moving object, a stationary object, or a vehicle. The object's sections can be at least partially movable and / or rigid relative to one another. The method is described with respect to a single object, but it is not limited to this. The method can be used to determine a spatial arrangement of two or more three-dimensional object models of two or more objects in space. The two or more objects can be captured at least partially simultaneously in one or more of the images.

[0021] Based on the spatial arrangement of the 3D object model, the spatial arrangement of the object itself can be inferred and / or determined.

[0022] The spatial arrangement of the object or the 3D object model can be a three-dimensional position of the object or the 3D object model.

[0023] Providing the multitude of initial images can include storing the multitude of initial images, in particular on a memory of a device described below. Alternatively or additionally, providing the multitude of initial images can include recording the room from the recording positions, in particular by means of a multitude of camera sensors of the device, and outputting the multitude of initial images.

[0024] Features described in relation to the multitude of first images can be applied analogously to further multitudes of images, for example, a subsequently mentioned multitude of second images.

[0025] The procedure may further include deploying the first trained neural network, which is trained to determine the two-dimensional positions of object segments in a two-dimensional image and to map the 2D positions to the object. Mapping the 2D positions to the object may include determining PAFs to connect the determined 2D positions. Deployment may further include storing and / or accessing the first neural network. The procedure may also include training the first neural network to deploy the trained first neural network.

[0026] Training data can be generated and / or provided for training the neural network. This training data can be at least partially or entirely synthetically generated. The training data can be generated using a 3D engine, particularly a 3D game engine. The 3D engine can generate two-dimensional training images that contain the training data. These training images can include one or more objects for which the initial neural network will be used. For example, the object could be a person. The 3D engine can generate various virtual environments or spaces. Furthermore, the 3D engine can adjust and / or vary light intensities, lighting directions, shadows, and other details. Finally, the 3D engine can generate and / or arrange one or more of the virtual objects—in this case, one or more virtual people—within the environment or space.The 3D engine can create multiple training scenarios, capturing each scenario from different positions and providing corresponding two-dimensional training images. Furthermore, the 3D engine can vary the objects by changing object details, such as gender, age, height, clothing, hairstyle, hair color, etc., for people. The environment or space can also be varied. Additionally, the objects, the environment or space, and / or the positions for capturing the training images can be modified. By creating the training images using the 3D engine, the two-dimensional positions of the image information are known. The training data can include the information generated by the 3D engine. Furthermore, the content of each training image is known.Thus, for example, supervised learning can be used to train the first neural network to determine the positions of object segments and their corresponding parts based on the training images. Assigning the 2D positions to the object can involve determining part affinity fields (PAFs) to connect the specific 2D positions. The object segments can be known from the training images, for example, the positions of the hands, feet, head, and torso of a person shown. Furthermore, the first neural network can be trained on the training data to create the 3D object model.

[0027] The training described above can be performed on the object model or object to be determined. If the object is, for example, a vehicle, appropriate training footage of vehicles must be created and / or provided. The type of vehicle can be varied. Varying the scene and / or the object can be based on the performance and / or accuracy of the initial neural network. For example, if cars and trucks are used, the neural network may not be trained accurately enough, and therefore, the training should be limited to cars.

[0028] The method involves determining the two-dimensional actual positions of object segments within a multitude of initial image captures and assigning these 2D actual positions to the object segments. Specifically, it involves determining the PAFs (Periodic Functional Factors) for connecting these determined 2D actual positions using the first trained neural network. This can include determining the actual position of a specific object segment within each image capture. Consequently, for multiple or all object segments included in a given image capture, the corresponding 2D actual positions can be determined, and the assignment process can be carried out, particularly the determination of the associated PAFs.

[0029] Determining the 3D lst positions, the 3D object model and / or the spatial arrangement of the 3D object model can be done using the first neural network.

[0030] The recording positions can be characterized by a two-dimensional, especially three-dimensional, position relative to or in space.

[0031] The first neural network can generate a keypoint heatmap to determine the 2D positions in a captured image. After generation, the keypoint heatmaps can be evaluated and / or processed using an identification algorithm. This identification algorithm can be, in particular, a parallelized identification algorithm. The identification algorithm can be designed to identify one or more properties, especially differences in the properties of keypoint regions and / or individual keypoints in the keypoint heatmap, and specifically to determine one or more peaks of the keypoint heatmap. These properties can be based on the information provided by the keypoint heatmap. The identification algorithm can be a connected component labeling algorithm or a blob detection algorithm.Using the identification algorithm, local high points in the keypoint heatmaps can be converted into pixel coordinates. A high point can be a region or a single pixel in the keypoint heatmap. The high point can be such a region or such a single pixel that exhibits the property Aen above a predetermined threshold. Once one or more regions have been identified, the center of each region can be determined for sub-pixel accurate high point detection and used for further processing. Alternatively, another algorithm can be used to convert the local high points into pixel coordinates.

[0032] A keypoint heatmap can characterize the probability of the presence of an object segment or the probabilities of several object segments in a captured image, where the probabilities are assigned to and / or can be assigned to two-dimensional positions within the captured image.

[0033] Determining the 3D LST positions can include:

[0034] Determining three-dimensional connecting lines from the recording positions to keypoints of the keypoint heatmaps;

[0035] Combining two or more 3D connecting lines to determine 3D lst position candidates;

[0036] Determine a predetermined subset of 3D lst position candidates; cluster the 3D lst position candidates of the subset to determine the 3D lst positions.

[0037] Determining the 3D lst position candidates can involve identifying one or more intersections of the 3D connecting lines of the recording positions with a respective keypoint.

[0038] The connecting lines can extend through each keypoint. Each keypoint initially has a 2D position in the captured image. The connecting line leads from and / or through the captured position and to and / or through the respective keypoint, as the keypoint contains no depth information. Consequently, the 3D connecting line can be infinitely long and / or have a predetermined maximum length.

[0039] The 3D connection lines can be straight. The 3D connection lines can connect a recording position to a keypoint in the keypoint heatmaps.

[0040] By combining two or more 3D connecting lines of a keypoint from different acquisition positions, the 3D LST position candidates can be triangulated. Triangulation can involve at least partial overlap of the 3D connecting lines of a given keypoint. Alternatively or additionally, triangulation can be performed by using those 3D connecting lines that are spaced below a predetermined minimum distance from each other and, in particular, do not overlap. Due to the three-dimensionality, not all 3D connecting lines of a keypoint need to overlap, so non-overlapping 3D connecting lines in the immediate vicinity can also be considered. However, this can also lead to incorrect overlaps or incorrectly spaced 3D connecting lines, especially 3D connecting line segments and / or points.Therefore, a limited number, for example 3, of triangulated 3D actual position candidates can be determined along a 3D connecting line and clustered to final points, in this case the 3D actual positions, using clustering, in particular using a DBSCAN algorithm. The clustering can include or be clustering using the DBSCAN algorithm. The DBSCAN algorithm can be executed once or multiple times.

[0041] The creation of the three-dimensional object model can be further based on a predetermined object skeleton of the object, where the object skeleton describes a relative arrangement of the object's segments to one another. In the case of a human, for example, the object skeleton can characterize an arrangement of the head, shoulders, elbows, hands, hips, legs, and / or feet relative to each other. For instance, predetermined length ratios between the individual segments and / or arrangements of the individual segments to one another can be characterized.

[0042] The procedure may include further:

[0043] Providing a large number of second two-dimensional images of the object in space, taken at a second time, wherein the second images were taken from predetermined different recording positions and the second images at least partially show the object sections;

[0044] Determining the difference between the first and second images taken from the respective shooting positions;

[0045] - Adjusting the spatial arrangement of the 3D object model based on the specified difference.

[0046] By determining the difference and adjusting the spatial arrangement, tracking, particularly cross-frame tracking, can be provided. A frame here can mean that the multitude of first capture images are used at the first time point. The second time point can be after the first, so that the multitude of first capture images can define a first frame and the multitude of second capture images a second frame. The first and second time points can be at a predetermined time interval. This time interval can be 30, 60, 90 FPS or more, especially 91 FPS or more. Consequently, a respective multitude of capture images can be recorded and provided at a predetermined time.At 90 or more FPS, 90 or more multiple images can be captured and provided, since a multiple of images can be captured at 90 or more times.

[0047] Determining the difference and / or adjusting the spatial arrangement can be done using the first neural network.

[0048] Alternatively or additionally, at least some of the steps performed for the large number of initial images can be repeated for the large number of subsequent images. This can include, in particular, re-determining 2D LST positions, assigning the 2D LST positions to the object segments, determining the PAFs, 3D LST positions, the 3D object model, and / or the spatial arrangement of the 3D object model.

[0049] Alternatively or additionally, the procedure can include identifying another object and / or the absence of the object in the multitude of second images. If the additional object is indeed a new object, the aforementioned steps can be performed to determine the 3D object model and / or the spatial arrangement of the 3D object model of the additional object.

[0050] The process can include tracking, in particular cross-frame tracking.

[0051] Determining the difference may further involve applying a filter to clean up noise from the 3D LST positions of the object segments. The filter could be a Kalman filter.

[0052] The procedure may include further:

[0053] Determining 2D LST positions of object detail sections within the multitude of initial capture images and assigning the 2D LST positions to the object detail sections based on the 2D LST positions of the object detail sections, the 2D LST positions of the object sections, the 3D LST positions of the object sections, and / or the 3D object model using a second trained neural network; determining 3D LST positions of the object detail sections based on the 2D LST positions of the object detail sections, the capture positions, the multitude of initial capture images, and the assignments of the object detail sections;

[0054] Adding object detail sections to the 3D object model based on the 3D lst positions.

[0055] Assigning the 2D LST positions to the object detail sections can involve defining PAFs to connect the specified 2D LST positions of the object detail sections. Defining the 3D LST positions of the object detail sections can then be done on the PAFs of the object detail sections.

[0056] Determining the 2D LST positions of the object detail sections can be performed after determining the 2D LST positions of the object sections in the initial set of images and assigning these 2D LST positions to the object. Furthermore, these steps regarding the object detail sections can be performed before and / or concurrently with subsequent steps related to the object sections.

[0057] Based on the enhanced 3D object model, further properties of the object, especially of object detail sections, can be determined. For example, it can be determined whether a person's hand is closed or open, or an emotion of the face can be determined.

[0058] Determining the 3D lst positions of the object detail sections can be further based on the PAFs of the object detail sections.

[0059] The completion of the 3D object model can be further based on assigning the 2D lst positions to the object detail sections, in particular the PAFs of the object detail sections.

[0060] The first and second neural networks can be at least partially different or identical. The first neural network can be the second neural network and / or encompass it. Alternatively, they can be two different neural networks. Determining the 2D LST positions of the object detail segments can be performed in one or more sections of the captured images. If object detail segments of one or more specific object segments are to be determined, the sections can be selected such that they encompass the specific object segments. Consequently, it is not necessary to process the entire captured image with respect to the object detail segments.

[0061] The process can further include the deployment of the second neural network. The second neural network can be configured to:

[0062] - Determining 2D LST positions of object segments in a captured image;

[0063] - Assigning the 2D-lst positions to the object sections, in particular determining PAFs to connect the specified 2D actual positions of the object detail sections based on the specified 2D-lst positions of the object detail sections, the 3D-lst positions of the object sections and / or the 3D object model.

[0064] Provisioning can involve saving to memory and / or accessing the second neural network.

[0065] The second neural network can be trained similarly to the first neural network using appropriate training data. The training data can be generated using the 3D engine, similar to the first neural network, but it will contain more and / or exclusively the object's detail sections.

[0066] The process of adding details to the 3D object model can involve supplementing it with object detail sections. For example, based on the numerous initial images, a person's hands and their positions can be determined, but not the positions of the fingers. The finger positions can be determined using the second neural network and the steps mentioned previously.

[0067] Furthermore, a difference between the 2D LST positions, the 3D LST positions, and / or the 3D object model can be determined based on the multitude of first images compared to the multitude of second images, and the 3D object model and / or its spatial arrangement can be adjusted accordingly. Consequently, tracking of object detail sections can also be provided.

[0068] The object can be a human being. Object segments can be hands, arms, forearms, upper arms, elbows, shoulders, torso, neck, head, pelvis, legs, knees, and / or feet. Object detail segments can be fingers, eyes, toes, mouth, facial muscles, and / or facial features.

[0069] The procedure may include further:

[0070] - Outputting a virtual reality (VR) to the object;

[0071] Controlling the output VR based on the specified 3D lst positions of the object sections and / or object detail sections, the 3D object model, the spatial arrangement of the 3D object model and / or the adapted spatial arrangement of the 3D object model.

[0072] Controlling the output VR can involve receiving one or more object inputs from the object, such as a person waving their hand. Control can be based on this single input or multiple inputs.

[0073] The method can further include outputting the specified 3D object model to an output device. The output device can be a screen. The output device can be configured to output a VR. The method can further include controlling the 3D object model output to the output device based on specific 3D LST positions and / or differences. Consequently, the 3D object model output to the output device can be an avatar of the object. The method can further include controlling the output device.

[0074] Furthermore, the method can include controlling a plurality of camera sensors configured to provide the plurality of captured images. This control can include adjusting the position of at least one of the plurality of camera sensors. The plurality of camera sensors can be height-adjustable.

[0075] According to a second aspect, the problem is solved by a device for determining the spatial arrangement of a three-dimensional object model of an object in a space, wherein the object has several object sections. The device comprises a plurality of camera sensors configured to record the space and output a plurality of two-dimensional images, and arranged at predetermined, different recording positions. The device further comprises a processor configured to:

[0076] - Issuing a recording command to capture a plurality of first two-dimensional images of the space at a first time to the plurality of camera sensors and to output the plurality of first images, wherein the first images at least partially show the object sections;

[0077] Determining the two-dimensional actual positions of object sections in the multitude of initial image captures and assigning the 2D actual positions to the object sections using a first trained neural network;

[0078] Determining three-dimensional actual positions of the object sections based on the multitude of initial image captures, the 2D actual positions, and the capture positions;

[0079] Creating a three-dimensional object model based on the specified 3D lst positions;

[0080] Determining a spatial arrangement of the 3D object model in space based on the determined 3D lst positions.

[0081] Assigning the 2D lst positions to the object sections may involve determining Part Affinity Fields (PAFs) to connect the specified 2D lst positions.

[0082] The creation of the three-dimensional object model can be further based on the assigned 2D lst positions.

[0083] The camera sensors can be cameras and / or include infrared cameras.

[0084] The first neural network and a previously mentioned second neural network can be stored in a memory of the device. Procedure features described in relation to the procedure according to the first aspect can be described as device features of the device according to the second aspect and are therefore not repeated here.

[0085] The device may further include an output means for outputting a virtual reality to the object. Alternatively or additionally, the device may include a memory for storing the multitude of initial capture images, the 2D LST positions, the mappings, the PAFs, the 3D LST positions, the 3D object model, the spatial arrangement of the 3D object model, the capture positions, and / or the initial neural network. Alternatively or additionally, the device may include a multitude of audio units configured for outputting audio to the object. The processor may include a controller for the device. The processor may be configured to control the units of the device. The device may alternatively or additionally include a communication unit for communicating with the units of the device and / or external units and / or external devices.

[0086] The output device can include a curved output surface, at least in sections. The output surface can be a screen. Alternatively or additionally, the output device can include one or more projectors designed to project the VR onto the output surface.

[0087] The numerous audio units can be designed to create a three-dimensional listening experience.

[0088] By determining the 2D LST positions, the 3D LST positions, the 3D object model, and / or the spatial arrangement of the 3D object model, a visual output via the output device and / or an audio output via the multitude of audio units can be controlled in such a way that the object, for example, a person, can perceive the VR. In particular, the output(s) can be adapted to the respective positions.

[0089] The multitude of first images can be stored in memory. Additionally or alternatively, the multitude of second images can also be stored in memory. The 2D LST positions and their assignments, particularly the PAFs, can be determined by accessing the multitude of first and / or second images stored in memory. The processor can be further configured to extract the information characterizing the 2D LST positions and their assignments, especially the PAFs, from memory and store it in the processor's main memory. This division of labor enables a particularly resource-efficient and fast process. Further processing can be performed, at least partially, using the information stored in the main memory.

[0090] To minimize latency, the information mentioned above and below can be stored in memory and / or main memory in such a way that only the information to be provided and / or processed in memory or main memory is stored. Consequently, only the information necessary for each step of the process can be transferred between memory and main memory.

[0091] The task is solved according to a third aspect by a computer program product comprising instructions that, when the program is executed by a processor, cause it to perform the procedure according to the first aspect. Alternatively or additionally, the computer program product may include instructions that, when the program is executed by the device according to the second aspect, cause it to perform the procedure according to the first aspect.

[0092] The computer program product can be stored on a computer-readable memory or storage medium. The computer-readable memory can be the device's memory.

[0093] Preferred embodiments are explained by way of example with reference to the accompanying figures. These show:

[0094] Fig. 1 is a schematic representation of a computer-implemented method for determining the spatial arrangement of a three-dimensional object model of an object in space; Fig. 2 is a schematic representation of determining 2D lst positions of object segments and part affinity fields (PAFs); and

[0095] Fig. 3 shows a schematic representation of a device for determining a spatial arrangement of a three-dimensional object model of an object in a room.

[0096] In the figures, identical or essentially functionally equivalent or similar elements are designated with the same reference symbols.

[0097] Fig. 1 shows a schematic representation of a computer-implemented method 100 for determining the spatial arrangement of a three-dimensional object model of an object in a space, wherein the object has several object sections. For example, a device 300, described below in conjunction with Fig. 3, can be used. A space can be detected by means of a plurality of camera sensors 310 of the device 300. The object, for example a person, can be located within this space and be depicted in images output by means of the camera sensors 310.

[0098] The method 100 can be stored in the form of a computer program product, for example on a memory of the device 300.

[0099] Method 100 can be used particularly for VR applications. Based on the specific spatial arrangement of the object model and the associated position of the object, interaction with VR can be enabled quickly and precisely.

[0100] Method 100 comprises providing 110, using a multitude of camera sensors 310, a multitude of initial two-dimensional images of the object in space, which were captured at a first time point. The initial images were captured from predetermined, different recording positions and show at least partial sections of the object. For a subsequent step of determining 3D first positions, the recording positions must be known. For example, a person can position themselves within the space, and the images are captured from the various recording positions at the first time point. Consequently, the person is captured from different perspectives.

[0101] The procedure 100 further comprises determining 120 two-dimensional actual positions 201 of object segments in the multitude of initial image captures and assigning the 2D actual positions to the object segments, in particular determining part affinity fields, PAFs 202 (see Fig. 2) for connecting the determined 2D actual positions using a first trained neural network. Consequently, the determination of the 2D actual positions 201 and / or the PAFs 202 can be a bottom-up approach.

[0102] Method 100 further comprises determining 130 three-dimensional actual positions of the object sections based on the multitude of initial image captures, the 2D actual positions 201, and the acquisition positions. Method 100 further comprises creating 140 a three-dimensional object model based on the determined 3D actual positions and, in particular, the PAFs 202. Method 100 further comprises determining 150 a spatial arrangement of the 3D object model in space based on the determined 3D actual positions.

[0103] Parts of the process 100 are described below using the example of Fig. 2.

[0104] Fig. 2 shows a single image A from the plurality of first images. Image A was captured from one of the predetermined shooting positions using one of the camera sensors 310.

[0105] The first neural network now determines the 2D LST positions. To do this, the neural network can first create a keypoint heatmap for the captured image A. The keypoint heatmap indicates the probability of an object segment being located at a given position in the captured image A. As shown in Fig. 2, multiple object segments can be present. Circles 201 indicate that an object segment is highly likely to be located in this area. The circles serve only for illustration and explanation. The probability can be specified pixel by pixel, so that a non-circular representation is not used. Such a circle can encompass multiple keypoints. For example, a particularly high probability can be indicated by a different color than a low probability in the captured image A.

[0106] Furthermore, the first neural network determines the assignments, in particular the PAFs 202, each of which connects at least two or more keypoints. A PAF can have two or more vectors; only two vectors are shown as examples.

[0107] This process is performed for a portion, preferably all, of the numerous images captured.

[0108] The three-dimensional known recording positions and the 2D lst positions of the keypoints can be connected in a next step using three-dimensional connecting lines.

[0109] By combining two connecting lines with a single keypoint, a three-dimensional position of this keypoint can be determined. This can be done, for example, using triangulation. However, this can also lead to false intersections of connecting lines. Therefore, a limited number (e.g., 3) of triangulated keypoints are determined along a connecting line and clustered into final keypoints or 3D LST positions using the DBSCAN algorithm.

[0110] Since it's possible that not all object segments are present in every image, missing keypoints from the missing segments can be supplemented with keypoints from the other images during the steps mentioned above. This supplementation can be done, in particular, based on clustering the keypoints. Once the 3D object model has been determined, the object's detail segments can be identified. For example, if the object is a person and a hand position has been determined, the 2D LST positions and the 3D LST positions of the fingers can be determined using a second trained neural network. Alternatively, the object's detail segments can be determined after the 2D LST positions of the object segments have been established.

[0111] Using the second neural network, the 2D LST positions of object detail segments within the numerous initial image captures can be determined, and these 2D LST positions can be assigned to the object detail segments. This assignment can include determining PAFs (Process Action Functions) to connect the determined 2D LST positions of the object detail segments. The determination of the 2D LST positions can be further based on the 2D LST positions of the object detail segments, the 3D LST positions of the object segments, and / or the 3D object model. Furthermore, the second neural network can determine the 3D LST positions of the object detail segments based on the 2D LST positions of the object detail segments, the capture positions, the numerous initial image captures, the assignments, and / or the PAFs of the object detail segments. Alternatively, this step can also be performed using a different algorithm.Finally, using the second neural network or one or the other algorithm, the 3D object model can be supplemented with the object detail sections based on the 3D-I st positions and the assignments, in particular the PAFs of the object detail sections.

[0112] This allows, for example, the hand and its associated fingers of a person to be quickly and precisely identified and used for interaction with VR.

[0113] Fig. 3 shows a schematic representation of a device 300 for determining a spatial arrangement of a three-dimensional object model of an object in a space, wherein the object has several object sections.

[0114] The device 300 comprises a plurality of camera sensors 310, which are designed to record the space and output a plurality of two-dimensional recording images and are arranged at predetermined different recording positions.

[0115] The device 300 further comprises a processor 320, which is configured to issue a recording command for acquiring a plurality of first two-dimensional images of the space at a first time to the plurality of camera sensors and for outputting, by means of the plurality of camera sensors, the plurality of first images. The first images show at least partial object sections. The plurality of camera sensors are advantageously infrared cameras.

[0116] The processor 320 is further designed to determine two-dimensional actual positions of object sections of the object in the multitude of first recording images and to assign the 2D actual positions to the object sections, in particular to determine part affinity fields (PAFs) for connecting the determined 2D actual positions using a first trained neural network.

[0117] Furthermore, the processor 320 is designed to determine three-dimensional actual positions of the object sections based on the multitude of initial recording images, the 2D actual positions and the recording positions.

[0118] The 320 processor is further designed to create a three-dimensional object model based on the specified 3D lst positions and assignments, in particular the PAFs, and to determine a spatial arrangement of the 3D object model in space based on the specified 3D lst positions.

[0119] The device 300 can be further equipped with output means (not shown) for outputting a VR image to the object, for example, a person. The rapid and precise determination of the person's position, especially the head, allows for accurate VR output.

[0120] The processor 320 can be configured to execute the computer program. For this purpose, the computer program can be stored in a memory of the device 100. REFERENCE MARK

[0121] 100 Computer-implemented methods for determining a spatial arrangement of a three-dimensional object model of an object

[0122] 110 Providing a large number of initial two-dimensional

[0123] Images

[0124] 120 Determining two-dimensional actual positions of

[0125] Object sections of the object and assigning the 2D LST positions to the object sections.

[0126] 130 Determining the three-dimensional actual positions of the object sections

[0127] 140 Creating a three-dimensional object model

[0128] 150 Determining a spatial arrangement of the 3D object model

[0129] 201 2D-lst positions

[0130] 202 PAFs

[0131] 300 Device for determining a spatial arrangement of a three-dimensional object model of an object

[0132] 310 Variety of camera sensors

[0133] 320 processor

[0134] A photograph

Claims

REQUIREMENTS 1. Computer-implemented method (100) for determining a spatial arrangement of a three-dimensional object model of an object in a space, wherein the object has multiple object sections, comprising: Providing (110) a plurality of first two-dimensional image images of the object in space, taken at a first time, wherein the first image images were taken from predetermined different recording positions and the first image images show at least part of the object sections; Determining (120) two-dimensional actual positions (201) of object sections of the object in the multitude of first recording images and assigning the 2D actual positions (201) to the object sections using a first trained neural network; Determine (130) three-dimensional actual positions of the object sections based on the multitude of first recording images, the 2D actual positions (201) and the recording positions; Creating (140) a three-dimensional object model based on the specified 3D lst positions; Determine (150) a spatial arrangement of the 3D object model in the space based on the determined 3D lst positions.

2. Method (100) according to claim 1, wherein the first neural network for determining the 2D positions in a captured image generates a keypoint heatmap.

3. Method (100) according to claim 2, wherein determining (130) the 3D lst positions comprises: Determining three-dimensional connecting lines from the recording positions to keypoints of the keypoint heatmaps; Combining two or more 3D connecting lines to determine 3D lst position candidates; Determining a predetermined subset of 3D lst position candidates; Clustering the 3D lst position candidates of the subset to determine the 3D lst positions.

4. Method (100) according to one of the preceding claims, wherein the creation (140) of the three-dimensional object model is further based on a predetermined object skeleton of the object, wherein the object skeleton describes a relative arrangement of the object sections to each other.

5. Method (100) according to any one of the preceding claims, further comprising: providing a plurality of second two-dimensional Images of the object in the room, taken at a second time, wherein the second images were taken from the predetermined different recording positions and the second images show at least some of the object sections; Determining the difference between the first and second images taken from the respective shooting positions; Adjusting the spatial arrangement of the 3D object model based on the specified difference.

6. Method (100) according to claim 5, wherein determining the difference further comprises applying a filter to clean up noise of the 3D lst positions of the object segment.

7. Method (100) according to any one of the preceding claims, further comprising: determining 2D LST positions of object detail sections of the Object sections in the multitude of initial capture images and assignment of the 2D actual positions to the object detail sections based on the 2D actual positions the object detail sections, the 3D lst positions of the object sections and / or the 3D object model using a second trained neural network; Determining 3D lst positions of the object detail sections based on the 2D lst positions of the object detail sections, the capture positions, the multitude of initial capture images, and the assignments; Adding object detail sections to the 3D object model based on the 3D lst positions.

8. Method (100) according to any of the preceding claims, wherein the object is a human being, wherein the object sections are hands, arms, forearms, upper arms, elbows, shoulders, torso, neck, head, pelvis, legs, knees and / or feet, and / or wherein the object detail sections are fingers, eyes, toes, mouth, facial muscles and / or facial sections.

9. Method (100) according to any of the preceding claims, further comprising: outputting a virtual reality, VR to the object; Controlling the output VR based on the specified 3D lst positions of the object sections and / or object detail sections, the 3D object model, the spatial arrangement of the 3D object model and / or the adapted spatial arrangement of the 3D object model.

10. Device (300) for determining a spatial arrangement of a three-dimensional object model of an object in a space, wherein the object has several object sections, comprising: a plurality of camera sensors (310) configured for recording the space and outputting a plurality of two-dimensional recording images and arranged at predetermined different recording positions; a processor (320) configured for: Issuing a recording command to record a plurality of first two-dimensional recording images of the space at a first time to the plurality of camera sensors (210) and to output the plurality of first recording images, wherein the first recording images at least partially show the object sections; Determining two-dimensional actual positions (201) of object sections of the object in the multitude of first recording images and assigning the 2D actual positions to the object sections using a first trained neural network; Determining three-dimensional actual positions of the object sections based on the multitude of initial image captures, the 2D actual positions (201) and the capture positions; Creating a three-dimensional object model based on the specified 3D lst positions; Determining a spatial arrangement of the 3D object model in space based on the determined 3D lst positions.

11. Device (300) according to claim 10, further comprising: an output means for outputting a virtual reality to the object; a memory for storing the plurality of first recording images, the 2D lst positions (201), the assignments, the 3D lst positions, the 3D object model, the spatial arrangement of the 3D object model, the recording positions and / or the first neural network; and / or a plurality of audio units configured for outputting audio to the object.

12. Device (300) according to claim 10 or 11, wherein the device (300) comprises the memory, wherein the plurality of first recording images are stored on the memory, wherein the 2D lst positions (201) and the assignments are determined by accessing the plurality of first recording images stored in the memory, wherein the processor (320) is further configured to extract information characterizing the 2D lst positions (201) and the assignments from the memory and to store it in a working memory of the processor (320).

13. Computer program product comprising instructions which, when the program is executed by a processor, cause the processor to execute the method (100) according to any one of claims 1 to 9.

14. Computer program product according to claim 13, wherein the computer program product is stored on a computer-readable memory.

Citation Information

Patent Citations

  • Mesh topological structure acquisition method and device, electronic equipment and storage medium

    CN114663983A

  • Digital human automatic modeling method and device

    CN117115358A

  • 3D skeletonization using truncated epipolar lines

    US20190139297A1