Neural Expansion of 3D Content in Augmented Reality Environments
A machine learning-based method generates AR content by integrating anchor content with physical space layouts, addressing the inefficiencies of manual AR asset placement, enhancing AR content integration and availability.
Patent Information
- Application Number
- JP2023209931
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-01-17
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2043-12-13
AI Technical Summary
Existing augmented reality (AR) environments that combine conventional media content with physical space layouts are time-consuming and resource-intensive, requiring manual placement of AR assets by multiple teams of artists and developers.
A method using a machine learning model to automatically generate AR content by inputting a physical space layout and anchor content, creating a 3D volume with the placement of 3D representations of anchor content within the space, and outputting views to a computing device.
This approach efficiently and seamlessly integrates anchor content with the physical space layout, reducing time and resource requirements while increasing the diversity and availability of AR content.
Smart Images

Figure 0007698697000001 
Figure 0007698697000002 
Figure 0007698697000003
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to machine learning and augmented reality, and more particularly, to neural augmentation of three-dimensional (3D) content in an augmented reality environment.
Background Art
[0002] Augmented reality (AR) refers to combining the real world with computer-generated content to realize an interactive perceptual experience. For example, an AR system may include a camera, a depth sensor, a microphone, an accelerometer, a gyroscope, and / or another type of sensor that detects events or changes in the environment around the user. The AR system may also include a display, a speaker, and / or another type of output device that combines the data collected by the sensors with additional AR content to create an immersive experience. The AR system also partially modifies the real-world output and / or the AR content in response to changes in the environment, interactions between the user and the AR content, and / or other inputs.
[0003] One use of AR involves combining conventional media content, such as images, audio, and / or video, with the layout of the physical space in the real world. For example, an AR system operating on a wearable device or a portable electronic device can extend an image or video of a scene across a room by overlaying the objects, shapes, colors, and / or textures of the scene on the walls, ceiling, floor, and / or other parts of the room. The AR system can also place various parts of the scene around doors, windows, and / or other types of objects within the room, such that the content appears to flow around rather than block these objects.
[0004] However, an AR environment that combines conventional media content with the layout of a physical space is typically generated by a process that consumes large amounts of time and resources. For example, a team of artists or other content creators can interact with a set of applications to convert objects, shapes, colors, and / or textures from an image or video of a scene into AR assets. Next, a different team of developers can interact with another set of applications to resize the AR assets, place and orient the augmented reality assets within an AR environment that incorporates the layout of the physical space, and / or otherwise customize the placement of the AR assets within the layout of the physical space. This process is repeated for each piece of conventional media content and each physical space in which the conventional media content is to be augmented.
Summary of the Invention
Problems to be Solved by the Invention
[0005] As described above, what is needed in the art is a more effective method for incorporating conventional media content into an AR environment.
Means for Solving the Problems
[0006] One embodiment of the present invention specifies a method for generating augmented reality (AR) content. The method includes inputting a first layout of a physical space and a first set of anchor content into a machine learning model. The method also includes generating, by operation of the machine learning model, a first three-dimensional (3D) volume that includes (1) a first subset of the physical space and (2) placement of one or more 3D representations of the first set of anchor content within a second subset of the physical space. The method further includes outputting, to a computing device, one or more views of the first 3D volume.
[0007] One technical advantage of the disclosed approach over the prior art is the ability to automatically generate seamless AR content that couples anchor content with the layout of the physical space. Thus, the disclosed approach is more time and resource efficient than conventional approaches that use various software components to convert conventional media content into AR assets and manually place the AR assets within the AR environment. Another technical advantage of the disclosed approach is the ability to generate AR content on-the-fly from a particular set of anchor content and a particular physical space. As a result, the disclosed approach increases the diversity and availability of AR content that couples conventional media content with the layout of the physical space. These technical advantages provide one or more technical improvements over prior art approaches.
Brief Description of the Drawings
[0008] In order for the above features of the various embodiments to be understood in detail, a more detailed description of the inventive concept briefly summarized above may be obtained by referring to the various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the accompanying drawings illustrate only typical embodiments of the inventive concept and should in no way be considered as limiting the scope of the present invention, as other equally effective embodiments exist.
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 3C
Figure 4A
Figure 4B
Figure 4C
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0009] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.
[0010] System Overview FIG. 1 shows a computing device 100 configured to execute one or more aspects of various embodiments. In one embodiment, the computing device 100 consists of a desktop computer, a laptop computer, a smartphone, a personal digital assistant (PDA), a tablet computer, or any other type of computing device configured to receive input, process data, and optionally display images, and is suitable for implementing one or more embodiments. The computing device 100 is configured to operate a training engine 122 and an execution engine 124 that reside in a memory 116.
[0011] Note that the computing devices described in this document are examples, and any other technically possible configurations fall within the scope of the present disclosure. For example, multiple instances of the training engine 122 and the execution engine 124 may operate on a set of nodes within a distributed and / or cloud computing system to perform the functions of the computing device 100. In another example, the training engine 122 and / or the execution engine 124 may operate on various sets of hardware, multiple types of devices, or environments, and can be adapted to various use cases or applications. In a third example, the training engine 122 and the execution engine may operate on various computing devices and / or various sets of computing devices.
[0012] In one embodiment, the computing device 100 includes, without limitation, one or more processors 102, an input / output (I / O) device interface 104 coupled to one or more input / output (I / O) devices 108, a memory 116, a storage device 114, and an interconnect (bus) 112 that connects the network interface 106. The processor 102 may be a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), an artificial intelligence (AI) accelerator, any other type of processing device, or a combination of different processing devices, such as a CPU configured to operate with a GPU, and may be any suitable processor implemented as any technically possible hardware device that can generally process data and / or execute software applications. Also, in the context of the present disclosure, the computing elements within the computing device 100 may correspond to a physical computing system (e.g., a system within a data center) or a virtual computing instance operating within a computing cloud.
[0013] The I / O device 108 includes devices that can provide input, such as a keyboard, mouse, touch screen, microphone, etc., and devices that can provide output, such as a display. Also, the I / O device 108 may include devices that can receive input and provide output, such as a touch screen, a Universal Serial Bus (USB) port, etc. The I / O device 108 is configured to receive various types of input from the end user (e.g., designer) of the computing device 100 and to provide various types of output, such as a displayed digital image or digital video or text, to the end user of the computing device 100. In some embodiments, one or more of the I / O devices 108 are configured to couple the computing device 100 to the network 110.
[0014] The network 110 is any technically possible type of communication network that permits the exchange of data between the computing device 100 and an external entity or device such as a web server or another network-connected computing device. For example, the network 110 may include a Wide Area Network (WAN), a Local Area Network (LAN), a wireless (WiFi) network, and / or the Internet, etc.
[0015] The storage device 114 includes non-volatile storage for applications and data and may include a fixed or removable disk drive, a flash memory device, and a CD-ROM, DVD-ROM, Blu-Ray, HD-DVD, or other magnetic, optical, or solid state storage device. The training engine 122 and the execution engine 124 may be stored in the storage device 114 and loaded into the runtime memory 116.
[0016] Memory 116 includes a random access memory (RAM) module, a flash memory device, or any other type of memory device or combinations thereof. Processor 102, I / O device interface 104, and network interface 106 are configured to read data from and write data to memory 116. Memory 116 includes various software programs (including training engine 122 and execution engine 124) executable by processor 102 and application data associated with the software programs.
[0017] In some embodiments, training engine 122 trains one or more machine learning models to generate an augmented reality (AR) environment that incorporates conventional media content into the physical space. For example, training engine 122 trains one or more neural networks to expand a two-dimensional (2D) scene depicted in an image and / or video across the walls, ceiling, floor, and / or other surfaces of a room. Training engine 122 also, or alternatively, trains one or more neural networks to generate a three-dimensional (3D) volume that incorporates objects, colors, shapes, textures, structures, and / or other attributes of the 2D scene into the layout of the room.
[0018] Execution engine 124 uses the trained machine learning model to generate an AR environment that combines conventional media content with the layout of the physical space. For example, execution engine 124 can input the physical layout of a room and anchor content including an image or video depicting a scene into a trained neural network. Execution engine 124 can use the trained neural network to depict chairs, tables, windows, doors, and / or other objects in their normal state within the room and expand the scene across the walls, ceiling, floor, and / or other surfaces of the room. Execution engine 124 also, or alternatively, can use the trained neural network to generate a 3D volume that places 2D and / or 3D representations of objects, colors, shapes, textures, and / or other attributes of the 2D scene within the room.
[0019] Neural Generation of Augmented Reality Environments from Anchor Content FIG. 2A is a more detailed view of the training engine 122 and the execution engine 124 of FIG. 1 according to various embodiments. More specifically, FIG. 2A shows the operation of the training engine 122 and the execution engine 124 when generating an AR environment 290 that uses a machine learning model 200 to extend a set of anchor contents 230 across a layout 232 of a physical space.
[0020] The anchor contents 230 include one or more fragments of any type of media that can be incorporated into the AR environment 290. For example, a particular set of anchor contents 230 can include a single image, multiple images, one or more video frames, one or more visualizations, one or more 3D models, one or more audio files, one or more text strings, and / or another type of digital content that can be overlaid and / or combined with a real-world situation to generate the AR environment 290.
[0021] In one or more embodiments, the anchor contents 230 are output and / or captured within the physical space incorporated into the AR environment 290. For example, the anchor contents 230 can include image, video, sound, and / or another type of media content output by a television, projector, display, speaker, and / or another type of output device in a room corresponding to the physical space. In another example, the anchor contents 230 can include paintings, photographs, murals, sculptures, sound, and / or another type of object or phenomenon present in or detected in that room. The anchor contents 230 can be additionally specified by a user interacting with the AR environment 290 using a bounding box or enclosure, in a calibration process that involves displaying a known image on the output device before displaying the anchor contents 230 and / or in another manner.
[0022] The anchor content 230 may also exist separately from, or instead of, the physical space over which the anchor content 230 extends. For example, the anchor content 230 may be specified as one or more files, which may include one or more images, videos, audio, 3D models, text, and / or other types of content that can be retrieved from data storage and incorporated into the AR environment 290. In another example, a user interacting with the AR environment 290 can generate and / or update the anchor content 230 by drawing the anchor content 230 using a touch screen and / or another type of input device.
[0023] The machine learning model 200 includes a spatial segmentation neural network 202, a content segmentation neural network 204, and an extrapolation neural network 206. In some embodiments, the spatial segmentation neural network 202, the content segmentation neural network 204, and the extrapolation neural network 206 are implemented as neural networks and / or other types of machine learning models. For example, the spatial segmentation neural network 202, the content segmentation neural network 204, and the extrapolation neural network 206 may include, but are not limited to, one or more convolutional neural networks, fully connected neural networks, recurrent neural networks, residual neural networks, transformer neural networks, autoencoders, variational autoencoders, adversarial generative networks, autoregressive models, bidirectional attention models, hybrid models, diffusion models, and / or other types of machine learning models capable of processing and / or generating content.
[0024] More specifically, the machine learning model 200 generates the output 2D content 238 that is incorporated into the AR environment 290 based on the anchor content 230 and the layout 232. The layout 232 includes the positions and / or orientations of the objects 234(1) to 234(X) (each individually referred to as an object 234) within the physical space. For example, the layout 232 may include a 2D or 3D map of a room. The map includes semantic segmentation that divides the room into a plurality of regions corresponding to walls, floors, ceilings, doors, tables, chairs, rugs, windows, and / or other objects within the room.
[0025] In one or more embodiments, the layout 232 is generated by the spatial partitioning network 202 based on sensor data 228 related to a physical space. For example, the sensor data 228 may include an image, depth map, point cloud, and / or another representation of the physical space. The sensor data 228 may be collected by cameras, inertial sensors, depth sensors, and / or other types of sensors on an augmented reality device and / or another type of computing device within and / or near the physical space. The sensor data 228 may also be used to generate a 2D or 3D model corresponding to the virtual twin of the physical space. The sensor data 228 and / or the virtual twin may be input into the spatial partitioning network 202, and predictions of objects and / or object categories for individual components (e.g., pixel positions, points within a point cloud, etc.), positions, or regions within the sensor data 228 and / or the virtual twin may be obtained as the output of the spatial partitioning network 202.
[0026] In some embodiments, the anchor content 230 is similarly processed by the content partitioning network 204 to generate the content partition 294. For example, one or more images within the anchor content 230 may be input into the content partitioning network 204, and the content partition 294 may be obtained as predictions of objects and / or object categories (e.g., foreground, background, cloud, star, object, animal, plant, face, structure, shape, situation, etc.) generated by the content partitioning network 204 for individual pixel positions and / or other subsets of the image.
[0027] The anchor content 230, sensor data 228, layout 232, and / or content partition 294 are provided as input to the extrapolation network 206. In response to the input, the extrapolation network 206 generates latent representations 236(1) - 236(Y) (each individually referred to as a latent representation 236) of various portions of the input data. The extrapolation network 206 also converts the latent representations 236 into an output 2D content 238 that includes a plurality of images 240(1) - 240(Z) (each individually referred to as an image 240).
[0028] In some embodiments, each image 240 within the output 2D content 238 represents one or more portions of physical space and depicts a semantically meaningful extension of the anchor content 230 within the physical space. For example, the output 2D content 238 may include six images 240 corresponding to the six faces of a cube representing a standard box-shaped room. In another example, the output 2D content 238 may include one or more images 240 depicting a 360-degree, spherical, and / or another type of panoramic view of physical space not limited to a box-shaped room. In both examples, each image 240 may include real-world objects within that room, such as (but not limited to) doors, windows, furniture, and / or decorations. Each image 240 may also include various subsets of the anchor content 230 (identified in the content segment 294) overlaid on the walls, floors, ceilings, and / or other surfaces of the room. These components of the anchor content 230 may also be arranged or distributed within the corresponding image 240, avoiding obscuring and / or partially overlapping doors, windows, furniture, decorations, and / or other types of objects within the room.
[0029] The training engine 122 trains the machine learning model 200 using training data 214 that includes a set of ground truth segmentations 208, a set of training sensor data 210, and a set of training anchor content 212. The training sensor data 210 includes images, point clouds, and / or other digital representations of visual and / or spatial attributes of various types of physical space. For example, the training sensor data 210 may include 2D and / or 3D representations of rooms or buildings of various architectural styles and layouts, outdoor urban spaces, underground spaces, natural environments, and / or other physical environments.
[0030] The training anchor content 212 includes images, videos, audio, and / or other content that can be combined with the training sensor data 210 to generate an AR environment (e.g., the AR environment 290). Similar to the anchor content 230, the training anchor content 212 can be depicted (e.g., as part of the corresponding physical space) and / or incorporated within the training sensor data 210, and / or can be separately extracted from the training sensor data 210 (e.g., like a digital file from data storage).
[0031] The ground truth data segment 208 includes labels related to the training sensor data 210 and / or the training anchor content 212. For example, the ground truth data segment 208 can include labels representing floors, walls, ceilings, lighting fixtures, furniture, decorations, doors, windows, and / or other objects that can be found within the physical space represented by the training sensor data 210. These labels can be assigned to pixel regions, 3D points, grids, sub-grids, and / or other data elements within the training sensor data 210. In another example, the ground truth data segment 208 can include labels representing foregrounds, backgrounds, textures, objects, shapes, structures, people, faces, animals, plants, situations, and / or other entities found or represented within the training anchor content 212. These labels can be assigned to pixel regions, audio samples, 3D models, and / or other elements or parts of the training anchor content 212. The ground truth data segment 208 can be available for all sets of the training sensor data 210 and / or the training anchor content 212, enabling supervised training of one or more components of the machine learning model 200, or the ground truth data segment 208 can be available for a subset of the training sensor data 210 and / or the training anchor content 212, enabling semi-supervised and / or weakly-supervised training of the components.
[0032] As shown in FIG. 2A, the training engine 122 executes a forward pass that inputs training sensor data 210 into the spatial segmentation network 202 to obtain a set of training spatial segmentations 222 as the corresponding output of the spatial segmentation network 202. During the forward pass, the training engine 122 also inputs training anchor content 212 into the content segmentation network 204 to obtain a set of training content segmentations 224 as the corresponding output of the content segmentation network 204. The training spatial segmentations 222 include predictions of classes related to data elements within the training sensor data 210, and the training content segmentations 224 include predictions of classes related to data elements within the training anchor content 212. For example, the training spatial segmentations 222 can include predicted probabilities of classes representing floors, walls, ceilings, lighting fixtures, furniture, decorations, doors, windows, and / or other objects that may be found within the physical space for pixel regions, 3D points, and / or other portions of the training sensor data 210 representing the physical space. The training content segmentations 224 can include predicted probabilities of classes representing foregrounds, backgrounds, textures, objects, shapes, structures, people, persons, faces, body parts, animals, plants, and / or other entities found or represented within various regions or portions of the training anchor content 212.
[0033] During the forward pass, the training engine 122 also uses the extrapolation network 206 to convert a plurality of pairs of the training space segment 222 and the training content segment 224 into a 2D training output 226 representing an augmented reality view. The augmented reality view combines the attributes of the physical space represented by the training sensor data 210 and the corresponding training space segment 222 with the attributes of the training anchor content 212 and the corresponding training content segment 224. For example, the training engine 122 inputs into the extrapolation network 206 a set of training sensor data 210 for a room, the corresponding training space segment generated from that set of training sensor data 210 by the space segment network 202, a piece of the training anchor content 212, and / or the training content segment generated from that piece of the training anchor content 212 by the content segment network 204. In response to the input, the extrapolation network 206 can generate a plurality of images (e.g., image 240) that combine the visual and semantic attributes associated with the input training sensor data 210 with the visual and semantic attributes associated with the input training anchor content 212. Each image can represent one or more portions of a panoramic view associated with different surfaces of a room with flat walls and / or a physical space of any shape (e.g., a non-box-shaped room, a undulating space, an outdoor space, etc.).
[0034] Once the forward pass is complete, the training engine 122 calculates a plurality of losses based on the training data 214 and the output generated by the spatial segmentation network 202, the content segmentation network 204, and / or the extrapolation network 206. More specifically, the training engine 122 calculates one or more segmentation losses 218 between the training spatial segmentation 222 generated from the training sensor data 210 by the spatial segmentation network 202 and the corresponding ground truth data segmentation 208. The training engine 122 also calculates, or alternatively, one or more segmentation losses 218 between the training content segmentation 224 generated from the training anchor content 212 by the content segmentation network 204 and the corresponding ground truth data segmentation 208. These segmentation losses 218 can include, but are not limited to, cross-entropy loss, dice loss, boundary loss, Tversky loss, and / or another measure of the error between a particular segmentation generated by the spatial segmentation network 202 and / or the content segmentation network 204 and the corresponding ground truth data segmentation.
[0035] The training engine 122 also calculates, or alternatively, one or more similarity losses 216 between a fragment of the training anchor content 212 and the 2D training output 226 generated by the extrapolation network 206 from that fragment of the training anchor content 212 and a set of training sensor data 210. In some embodiments, the similarity loss 216 measures the visual similarity between the training anchor content 212 and the portion of the 2D training output 226 corresponding to the extension of the training anchor content 212 in the physical space.
[0036] For example, the training engine 122 can use the ground truth data segment 208 to identify an area of physical space (e.g., one or more walls, ceilings, floors, etc.) over which the training anchor content 212 is overlaid or extended. These areas can be specified and / or selected to control the manner in which the resulting 2D training output 226 is generated. The training engine 122 can apply a mask associated with these areas to the 2D training output 226 to remove portions of the 2D training output 226 that are outside of these areas. The training engine 122 can calculate the L1 loss, L2 loss, mean squared error, Huber loss, and / or other similarity losses 216 as a measure of the similarity or difference between the visual attributes (e.g., shape, color, pattern, pixel values, line thickness, contour, etc.) of various components of the training anchor content 212 and the visual attributes of the remaining 2D training output 226. The training engine 122 can also, or instead, convert the remaining 2D training output 226 into a first set of latent representations (e.g., by applying one or more components of the extrapolation network 206 and / or a pre-trained feature extractor to the remaining 2D training output 226), and calculate one or more similarity losses 216 as a measure of the similarity or difference between the first set of latent representations and a second set of latent representations associated with the training anchor content 212 (e.g., latent representations 236 generated from the training anchor content 212 by the extrapolation network 206 and / or a pre-trained feature extractor). As a result, the similarity loss 216 can be used to ensure that the machine learning model 200 learns to generate an extension of the training anchor content 212 to a portion or area of physical space.
[0037] The training engine 122 can also, or instead, calculate one or more layout losses 220 between the 2D training output 226 and the corresponding training sensor data 210 and / or the ground truth data segment 208 associated with the training anchor content 212. In one or more embodiments, the layout loss 220 measures the extent to which the 2D training output 226 depicts a semantically meaningful extension of the training anchor content 212 across the physical space represented by the training sensor data 210.
[0038] For example, the training engine 122 can identify various objects (e.g., walls, ceilings, floors, furniture, decorations, windows, doors, etc.) within a room represented by a set of training sensor data 210 using the ground truth data segment 208. The training engine 122 can also use the ground truth data segment 208 to generate a mask that identifies regions of objects (e.g., doors, windows, furniture, and / or other objects that the training anchor content 212 should not occlude or replace) to be depicted in the 2D training output 226 for that room. These regions can be specified and / or selected to control how the resulting 2D training output 226 is generated. The training engine 122 can apply the mask to the 2D training output 226 to remove portions of the 2D training output 226 that are outside of these regions. The training engine 122 can then calculate one or more layout losses 220, such as L1 loss, L2 loss, mean squared error, Huber loss, and / or other layout losses, as a measure of similarity or difference between the visual attributes (e.g., shape, color, pattern, pixel values, etc.) of the remaining 2D training output 226 and the visual attributes of the corresponding objects within the training sensor data 210. The training engine 122 can also, or instead, convert the remaining 2D training output 226 into a first set of latent representations (e.g., by applying one or more components of the extrapolation network 206 and / or a pre-trained feature extractor to the remaining 2D training output 226), and calculate one or more layout losses 220 as a measure of similarity or difference between the first set of latent representations and a second set of latent representations associated with the corresponding objects within the training sensor data 210 (e.g., latent representations 236 generated from portions of the training sensor data 210 that depict or represent these objects by the extrapolation network 206 and / or the pre-trained feature extractor). In other words, the layout losses 220 can be used to ensure that the 2D training output 226 includes accurate and / or complete depictions of these objects at the corresponding locations, and that any overlap or extension of the training anchor content 212 within the AR view of the room does not occlude these objects.
[0039] In one or more embodiments, the layout loss 220 is used to ensure semantic and / or spatial consistency across frames of the 2D training output 226. For example, the 2D training output 226 may include a plurality of frames depicting an extension of corresponding video frames within the training anchor content 212 across the physical space represented by the training sensor data 210. The training engine 122 may provide one or more previous frames within the 2D training output 226, the corresponding training space segment 222, the corresponding training content segment 224, and / or the corresponding frame of the training anchor content 212 as additional inputs to the space segment network 202, the content segment network 204, and / or the extrapolation network 206 to inform the generation of the current frame of the 2D training output 226 from the current frame of the training anchor content 212. The training engine 122 may also calculate one or more layout losses 220 between or across consecutive frames of the 2D training output 226 to ensure that objects or videos within the video frames of the training anchor content 212 are depicted or extended at generally the same positions within the physical space without jumping around or varying erratically.
[0040] Next, the training engine 122 executes a backward pass to update the parameters of the spatial segmentation network 202, the content segmentation network 204, and / or the extrapolation network 206 using various permutations or combinations of the similarity loss 216, the segmentation loss 218, and / or the layout loss 220. For example, the training engine 122 can use a training technique (e.g., gradient descent and backpropagation) to update the parameters of the spatial segmentation network 202 based on the segmentation loss 218 calculated between the training spatial segmentation 222 and the corresponding ground truth data segmentation 208. The training engine 122 can also update the parameters of the content segmentation network 204 based on the segmentation loss 218 calculated between the training content segmentation 224 and the corresponding ground truth data segmentation 208. When the training of the spatial segmentation network 202 and the content segmentation network 204 is complete, the training engine 122 freezes the parameters of the spatial segmentation network 202 and the content segmentation network 204 and can train the extrapolation network 206 based on the layout loss 220 and / or the similarity loss 216. The training engine 122 can also alternate the training of the spatial segmentation network 202 and the content segmentation network 204 based on the segmentation loss 218 and continue to train the extrapolation network 206 based on the layout loss 220 and / or the similarity loss 216 until the segmentation loss 218, the layout loss 220, and / or the similarity loss 216 are below the corresponding thresholds.
[0041] In another example, the training engine 122 can perform end-to-end training of the spatial segmentation network 202, the content segmentation network 204, and the extrapolation network 206 using the layout loss 220, the similarity loss 216, and / or other losses calculated based on the 2D training output 226. This end-to-end training can be performed alternately with and / or in another way than the training based on the segmentation loss 218 of the spatial segmentation network 202 and the content segmentation network 204 after the spatial segmentation network 202 and the content segmentation network 204 are trained based on the corresponding segmentation loss 218.
[0042] The operation of the training engine 122 was described above with respect to a certain type of loss. However, it will be understood that the spatial segmentation network 202, the content segmentation network 204, and / or the extrapolation network 206 can be trained using other types of techniques, losses, and / or machine learning components. For example, the training engine 122 can train the spatial segmentation network 202, the content segmentation network 204, and / or the extrapolation network 206 in an adversarial manner using one or more discriminator neural networks (not shown) and one or more discriminator losses calculated from the ground truth data segmentation 208, the training spatial segmentation 222, the training content segmentation 224, and / or the 2D training output 226 based on the predictions generated by the discriminator neural networks. In another example, the training engine 122 can train the spatial segmentation network 202, the content segmentation network 204, and / or the extrapolation network 206 using one or more losses that reflect artist guidelines or parameters in the 2D training output 226. These artist guidelines or parameters relate to symmetry, balance, color, orientation, scale, style, the movement of the training anchor content 212 across the physical space represented by the training sensor data 210, the size of the objects within the training anchor content 212 drawn or placed within the physical space represented by the training sensor data 210, and / or other attributes that affect the way the training anchor content 212 is drawn or extended across the physical space represented by the training sensor data 210. In a third example, the training engine 122 can train the spatial segmentation network 202, the content segmentation network 204, and / or the extrapolation network 206 to learn to extend specific fragments of the training anchor content 212 across various physical spaces.
[0043] After the training of the machine learning model 200 is completed, the execution engine 124 generates output 2D content 238 using one or more components of the trained machine learning model 200. In particular, the execution engine 124 uses the spatial segmentation circuit network 202 to convert sensor data 228 (e.g., images, point clouds, depth maps, etc.) collected from the physical space into a semantic layout 232 that identifies the positions of objects 234 within that physical space. The execution engine 124 also receives anchor content 230 (e.g., images, paintings, videos, etc.) as a region of the sensor data 228 selected by the user and / or as one or more image, video, or audio files that exist separately from the sensor data 228. The execution engine 124 uses the content segmentation circuit network 204 to convert the anchor content 230 into a corresponding content segmentation 294. The execution engine 124 then uses the extrapolation circuit network 206 to generate a potential representation 236 of the object 234 within the portion of the anchor content 230 identified using the layout 232 and / or the content segmentation 294. The execution engine 124 also uses the extrapolation circuit network 206 to generate the output 2D content 238 based on the potential representation 236 and / or additional inputs, such as the position and orientation of the computing device providing the AR environment 290 (e.g., by sampling, transforming, correlating, or otherwise processing the potential representation 236). As described above, the output 2D content 238 includes an image 240 depicting a certain type of object 234 (e.g., doors, windows, furniture, decorations, etc.) at their respective positions within the physical space and an extension and / or overlay of the anchor content 230 across other types of objects 234 (e.g., walls, floors, ceilings, etc.) within the physical space. The image 240 can also depict the object 234 from a certain viewpoint, e.g., the viewpoint from which the sensor data 228 was collected.
[0044] The execution engine 124 also incorporates the output 2D content 238 into the AR environment 290. For example, the execution engine 124 can incorporate the image 240 into the visual, audio, and / or other outputs generated by the AR system that provides the AR environment 290, such that the AR environment 290 appears to depict a semantically meaningful extension of the anchor content 230 across the physical space from the perspective of the user interacting with the AR system. When the user changes the perspective of the AR system and / or updates the anchor content 230 (e.g., by depicting additional portions of the anchor content 230, trimming, shrinking / enlarging, and / or deforming the anchor content 230, changing the color balance, saturation, color temperature, exposure, luminance, and / or other color-related attributes of the anchor content 230, changing the view of the anchor content 230, loading the anchor content 230 from a different file, playing a video or movie that includes multiple frames of the anchor content 230, etc.), the execution engine 124 receives the updated anchor content 230 and / or additional sensor data 228 that reflects the latest view or representation of the physical space, generates a new layout 232 from the additional sensor data 228 and / or a new content segment 294 from the new anchor content 230, and can generate new output 2D content 238 that depicts the objects 234 associated with the new perspective and the extension of the anchor content 230. As a result, the execution engine 124 generates an immersive AR environment 290 that continuously responds to changes in the sensor data 228 and / or the anchor content 230.
[0045] In some embodiments, the output 2D content 238 is used in other types of immersive environments, such as virtual reality (VR) and / or mixed reality (MR) environments. This content can depict a virtual world that can be continuously experienced by any number of users synchronously, while providing continuity of data such as an individual's identity, user history, credentials, assets, and / or rewards. Note that this content can include conventional audiovisual content and hybrids of fully immersive VR, AR, and / or MR experiences, such as two-way video.
[0046] FIG. 2B is a more detailed view of the training engine 122 and the execution engine 124 of FIG. 1 according to various embodiments. More specifically, FIG. 2B shows the operation of the training engine 122 and the execution engine 124 when generating an AR environment 292 that includes an output 3D volume 284 using a machine learning model 280. Within the output 3D volume 284, a set of 3D objects 286(1)-286(C) (each individually referred to as a 3D object 286) obtained from a set of anchor contents 296 are arranged in a semantically meaningful way within the layout 276 of the physical space.
[0047] Similar to the anchor content 230 of FIG. 2A, a particular set of anchor contents 296 includes one or more fragments of any type of media that can be incorporated into the AR environment 292. For example, the anchor content 296 can include images, image sequences, videos, audio, text, and / or other types of digital content that can be overlaid and / or combined with the real-world situation to form the AR environment 292. In another example, the anchor content 296 can include one or more 3D videos, 3D models, and / or other types of 3D content that can be used with the AR environment 292.
[0048] In one or more embodiments, the anchor content 296 is output and / or captured within the physical space incorporated into the AR environment 292. For example, the anchor content 296 can include photos, paintings, and / or 2D or 3D videos output by a television, projector, display, and / or another visual output device within a room. In another example, the anchor content 296 can include paintings, photos, murals, sculptures, sounds, and / or other physical entities present or detected within the room. The anchor content 296 can be specified using a bounding box or enclosure by a user who interacts with the AR environment 292 in a calibration process that involves displaying a known image on the output device before displaying the anchor content 296 and / or in another way.
[0049] The anchor content 296 may also exist separately from, or instead of, the physical space in which the anchor content 296 is disposed. For example, the anchor content 296 may be designated as a file containing another type of content that can be retrieved from an image, video, audio, text, 3D model, and / or data storage and incorporated into the AR environment 292. In another example, a user interacting with the AR environment 292 can generate and / or update the anchor content 296 by one or more trimming, zooming, rotating, translating, color adjustment, and / or rendering operations.
[0050] The machine learning model 280 includes a spatial segmentation network 242, a content segmentation network 244, and a 3D synthesis network 246. In some embodiments, the spatial segmentation network 242, the content segmentation network 244, and the 3D synthesis network 246 are implemented as neural networks and / or other types of machine learning models. For example, the spatial segmentation network 242, the content segmentation network 244, and the 3D synthesis network 246 can include, but are not limited to, one or more convolutional neural networks, fully connected neural networks, recurrent neural networks, residual neural networks, transformer neural networks, autoencoders, variational autoencoders, adversarial generative networks, autoregressive models, bidirectional attention models, hybrid models, diffusion models, neural radiance field models, and / or other types of machine learning models that can process and / or generate content.
[0051] In some embodiments, the machine learning model 280 generates an output 3D volume 284 based on the anchor content 296 and the layout 276. In some embodiments, the layout 276 includes the 3D positions, orientations, and / or representations of the objects 278(1) to 278(A) (each individually referred to as an object 278) within the physical space. For example, the layout 276 can include one or more grids, point clouds, textures, and / or other 3D representations of a room captured by one or more cameras, depth sensors, and / or other types of sensors. Various parts or subsets of the 3D representation can be further segmented or labeled to represent objects, such as walls, floors, ceilings, doors, tables, chairs, carpets, windows, and / or other objects within the room.
[0052] In one or more embodiments, the layout 276 is generated based on sensor data 274 related to physical space by the spatial partitioning network 242. For example, the sensor data 274 may include an image of the physical space, a depth map, a point cloud, a grid, a texture, and / or other 3D representations. The sensor data 274 may be collected by sensors of an augmented reality device and / or another type of computing device within or near the physical space. The sensor data 274 may also be used to generate a 3D model corresponding to the virtual twin of the physical space. The sensor data 274 and / or the virtual twin may be input into the spatial partitioning network 242, and predictions of objects and / or object categories for individual elements (e.g., pixel positions, points in a point cloud, etc.), positions, or regions within the data may be obtained as corresponding outputs of the spatial partitioning network 242.
[0053] In some embodiments, the anchor content 296 is similarly processed by the content partitioning network 244 to generate the content partition 282. For example, one or more images within the anchor content 296 may be input into the content partitioning network 244, and the content partition 282 may be obtained as predictions of objects and / or object categories (e.g., foreground, background, cloud, star, object, animal, face, structure, shape, etc.) generated by the content partitioning network 244 for individual pixel positions and / or other data elements within the image.
[0054] As described above, the anchor content 296 may also include individual 3D objects 286 and / or various other components that can be individually placed or incorporated within the output 3D volume 284. For example, the anchor content 296 may include separate 3D models for various virtual people, faces, trees, buildings, furniture, clouds, stars, and / or other objects. In this example, the content partition 282 may be omitted or performed for each 3D model to further identify corresponding sub-components of the object or entity (e.g., the handle, window, chassis, door, and / or other parts of a 3D model of a car).
[0055] Anchor content 296, sensor data 274, layout 276, and / or content segments 282 are provided as input to a 3D synthesis network 246. In response to the input, the 3D synthesis network 246 generates potential representations 288(1)-288(B) (each individually referred to as a potential representation 288) of various portions of the input data. The 3D synthesis network 246 also converts the potential representations 288 into an output 3D volume 284.
[0056] In some embodiments, the output 3D volume 284 includes a 3D representation of the physical space associated with the sensor data 274. Within the 3D representation, 3D objects 286 obtained from the anchor content 296 are arranged in a semantically meaningful way within the physical space. For example, the output 3D volume 284 can be represented by a neural radiance field generated by the 3D synthesis network 246. The neural radiance field can include 3D representations of real-world objects 278 within the physical space, such as doors, windows, furniture, and / or decorations. The neural radiance field can also include 3D objects 286 corresponding to various components of the anchor content 296 (identified by the content segments 282). These components of the anchor content 296 can also be placed or distributed within the output 3D volume 284, avoiding obscuring and / or partially overlapping doors, windows, furniture, decorations, and / or other real-world objects within the room. These components of the anchor content 296 can also be arranged and / or animated to interact with real-world objects within the room.
[0057] A training engine 122 trains a machine learning model 280 using training data 298 that includes a set of ground truth data segments 248, a set of training sensor data 250, a set of training anchor content 252, and / or a set of training 3D objects 254. The training sensor data 250 includes images, point clouds, and / or other digital representations of the visual and / or spatial attributes of rooms, buildings, urban environments, natural environments, underground environments, and / or other types of physical spaces. The training sensor data 250 can also include, or instead of, a digital twin of a physical space constructed using images, point clouds, grids, textures, and / or other visual or spatial attributes of the physical space.
[0058] The training anchor content 252 includes images, videos, audio, and / or other content that can be combined with the training sensor data 250 to generate an AR environment (e.g., the AR environment 292). Similar to the anchor content 296, the training anchor content 252 can be depicted in and / or captured by (e.g., as part of the corresponding physical space) the training sensor data 250 and / or retrieved separately from the training sensor data 250 (e.g., as a digital file from a data storage).
[0059] The ground truth data segment 248 includes labels related to the training sensor data 250 and / or the training anchor content 252. For example, the ground truth data segment 248 can include labels representing floors, walls, ceilings, lighting fixtures, furniture, decorations, doors, windows, and / or other objects that can be found within the physical space represented by the training sensor data 250. These labels can be assigned to pixel regions, 3D points, grids, sub-grids, and / or other data elements within the training sensor data 250. In another example, the ground truth data segment 248 can include labels representing foregrounds, backgrounds, textures, objects, shapes, structures, people, persons, faces, body parts, animals, plants, and / or other entities found in or represented by the training anchor content 252. These labels can be assigned to pixel regions, point clouds, grids, sub-grids, audio tracks or channels, and / or other elements or parts of the training anchor content 252. The ground truth data segment 248 is available for all sets of the training sensor data 250 and / or the training anchor content 252 and can enable supervised training of one or more components of the machine learning model 280, or the ground truth data segment 248 is available for some sets of the training sensor data 250 and / or the training anchor content 252 and can enable semi-supervised and / or weakly-supervised training of those components.
[0060] The training 3D object 254 includes a 3D representation of the training anchor content 252. For example, the training anchor content 252 may include an image or video depicting a 2D representation of a 3D model or scene, and the training 3D object 254 may include a 3D model or scene. In other words, the training 3D object 254 can be used as a 3D representation of the ground truth data of the objects within the training anchor content 252.
[0061] In some embodiments, some or all of the training 3D objects 254 are included in the training anchor content 252. For example, the training 3D object 254 can be input into one or more components of the machine learning model 280, allowing the machine learning model 280 to learn to place the training 3D object 254 in a semantically meaningful way within the 3D representation of the physical space.
[0062] As shown in FIG. 2B, the training engine 122 executes a forward pass that inputs the training sensor data 250 into the spatial segmentation network 242 and obtains a set of training spatial segmentations 266 as the corresponding output of the spatial segmentation network 242. During the forward pass, the training engine 122 also inputs the training anchor content 252 into the content segmentation network 244 and obtains a set of training content segmentations 268 as the corresponding output of the content segmentation network 244. The training spatial segmentations 266 include predictions of classes related to the data elements within the training sensor data 250, and the training content segmentations 268 include predictions of classes related to various subsets or parts of the training anchor content 252. For example, the training spatial segmentations 266 can include the predicted probabilities of classes representing pixel regions of the training sensor data 250 representing the physical space, 3D points, and / or other types of floors, walls, ceilings, lighting fixtures, furniture, decorations, doors, windows, and / or other objects that can be found within the physical space. The training content segmentations 268 can include the predicted probabilities of classes representing foreground, background, texture, objects, shapes, structures, people, figures, faces, animals, plants, and / or other entities found or represented within the training anchor content 252 for various regions or parts of the training anchor content 252.
[0063] During the forward pass, the training engine 122 also uses the 3D synthesis network 246 to convert a pair of the training space segment 266 and the training content segment 268 into a 3D training output 264 representing a 3D scene, where the 3D scene combines the attributes of the physical space represented by the training sensor data 250 and the corresponding training space segment 266 with the attributes of the training anchor content 252 and the corresponding training content segment 268. For example, the training engine 122 can input a set of training sensor data 250 about the physical space, the corresponding training space segment generated from that set of training sensor data 250 by the space segment network 242, a set of training anchor content 252, and / or the training content segment generated from that set of training anchor content 252 by the content segment network 244 into the 3D synthesis network 246. In response to the input, the 3D synthesis network 246 can generate a 3D training output 264 including a neural radiance field and / or another representation of the 3D scene volume. The 3D scene volume can combine the visual and semantic attributes related to the input training sensor data 250 with the visual and semantic attributes related to the input training anchor content 252.
[0064] In one or more embodiments, the 3D synthesis network 246 generates the 3D training output 264 in multiple stages. For example, the 3D synthesis network 246 can include a first set of neural network layers that are given the 2D training anchor content 252 and the corresponding training content segment 268 and convert the 2D training anchor content 252 into a 3D representation of the same object. The 3D synthesis network 246 can also include a second set of neural network layers that are given an input including the 3D representation, a set of training sensor data 250, and / or the training space segment 266 generated from the set of training sensor data 250 and place the 3D representation within a 3D model of the physical space represented by the set of training sensor data 250. If the training anchor content 252 includes 3D representations of objects, these 3D representations can be input directly into the second set of neural network layers without the need for additional processing by the first set of neural network layers.
[0065] Once the forward pass is complete, the training engine 122 calculates a plurality of losses based on the training data 298 and the output generated by the spatial segmentation circuit network 242, the content segmentation circuit network 244, and / or the 3D synthesis circuit network 246. More specifically, the training engine 122 calculates one or more segmentation losses 262 between the training spatial segmentation 266 generated from the training sensor data 250 by the spatial segmentation circuit network 242 and the corresponding ground truth data segmentation 248. The training engine 122 also calculates, or alternatively, one or more segmentation losses 262 between the training content segmentation 268 generated from the training anchor content 252 by the content segmentation circuit network 244 and the corresponding ground truth data segmentation 248. These segmentation losses 262 may include, but are not limited to, cross-entropy loss, dice loss, boundary loss, Tversky loss, and / or another measure of the error between a particular segmentation generated by the spatial segmentation circuit network 242 and / or the content segmentation circuit network 244 and the corresponding ground truth data segmentation.
[0066] The training engine 122 also, or instead, calculates one or more similarity losses 256 between a fragment of the training anchor content 252 and a corresponding view associated with the 3D training output 264 generated from that fragment of the training anchor content 252 and a set of training sensor data 250 by the extrapolation network 206. In some embodiments, the similarity loss 256 measures the visual similarity between the training anchor content 252 and the depicted views of the 3D training output 264 generated based on the training anchor content 252. For example, the training engine 122 can use the ground truth data segment 248 to identify an area of the physical space (e.g., one or more walls, ceilings, floors, etc.) over which the training anchor content 252 is placed or shown. These areas can be specified and / or selected to control the manner in which the resulting 3D training output 264 is generated. The training engine 122 can apply a mask associated with these areas to the 3D training output 264 to remove portions of the 3D training output 264 that are outside of these areas. The training engine 122 can then calculate the L1 loss, L2 loss, mean squared error, Huber loss, and / or other similarity losses 256 as a measure of the similarity or difference between the visual attributes (e.g., shape, color, pattern, pixel values, etc.) of various components of the training anchor content 252 and the visual attributes of the depicted views of the 3D training output 264. The training engine 122 can also, or instead, convert the remaining 3D training output 264 and / or the depicted views to a first set of latent representations (e.g., by applying one or more components of the 3D synthesis network 246 and / or a pre-trained feature extractor to the remaining 3D training output 264), and calculate one or more similarity losses 256 as a measure of the similarity or difference between the first set of latent representations and a second set of latent representations associated with the training anchor content 252 (e.g., latent representations 288 generated from the training anchor content 252 by the 3D synthesis network 246 and / or a pre-trained feature extractor).
[0067] The training engine 122 also or alternatively calculates one or more layout losses 260 based on the training sensor data 250 corresponding to the 3D training output 264 and / or the ground truth data segment 248 associated with the training anchor content 252. In one or more embodiments, the layout loss 260 measures the extent to which the 3D training output 264 depicts a semantically meaningful placement of the training anchor content 252 within the physical space represented by the training sensor data 250.
[0068] For example, the training engine 122 can identify various objects in a room (e.g., walls, ceilings, floors, furniture, decorations, windows, doors, etc.) represented by a set of training sensor data 250 using the ground truth data segment 248. The training engine 122 can generate a mask that identifies areas of the objects in that room (e.g., doors, door frames, windows, window frames, furniture, decorations, object boundaries, and / or other areas that should not be obscured by the placement of training anchor content 252) to be depicted in the 3D training output 264. These areas can control the manner in which the 3D training output 264 that is specified and / or selected is generated. The training engine 122 can apply the mask to the 3D training output 264 and remove portions of the 3D training output 264 that are outside of these areas. The training engine 122 can then calculate the L1 loss, L2 loss, mean squared error, Huber loss, and / or one or more other layout losses as a measure of the similarity or difference between the visual attributes (e.g., shape, color, pattern, pixel values, etc.) of the remaining 3D training output 264 (or the depicted views of the remaining 3D training output 264) and the visual attributes of the corresponding objects in the training sensor data 250. The training engine 122 can also, or instead, convert the remaining 3D training output 264 (or the depicted views of the remaining 3D training output 264) into a first set of latent representations (e.g., by applying one or more components of the 3D synthesis network 246 and / or a pre-trained feature extractor to the remaining 3D training output 264), and calculate the one or more layout losses 260 as a measure of the similarity or difference between the first set of latent representations and a second set of latent representations associated with the corresponding objects in the training sensor data 210 (e.g., latent representations 288 generated from portions of the training sensor data 250 that depict or represent these objects by the 3D synthesis network 246 and / or the pre-trained feature extractor). As a result, the layout loss 260 can be used to ensure that the 3D training output 264 includes accurate and / or complete depictions of these objects at the corresponding locations and that the overlay or placement of the training anchor content 252 within the 3D volume representing the room does not obscure or distort these objects.
[0069] In some embodiments, one or more layout losses 260 can be used to ensure that the training anchor content 252 is placed at a location within the physical space represented by a particular set of training sensor data 250. These locations can include, but are not limited to, plain surfaces such as walls or floors, empty volumes that exceed a threshold size (e.g., the empty portion of a room that exceeds a certain volume or set of dimensions), surfaces on which objects can be placed (e.g., tables, desks, fireplaces, etc.), and / or certain types of objects that can be decorated with the training anchor content 252 (e.g., lighting fixtures, windows, or doorways through which the training anchor content 252 can flow, textures or surfaces that can be combined with corresponding elements of the training anchor content 252 to form a pattern or video, etc.). To train the machine learning model 280 to place various types of training anchor content 252 at these locations, the training engine 122 can generate a mask that identifies these regions or locations and apply that mask to the 3D training output 264 to remove portions of the 3D training output 264 that are outside of these regions or locations. The training engine 122 can also calculate the L1 loss, L2 loss, mean squared error, Huber loss, and / or one or more other layout losses 260 as a measure of the similarity or difference between the visual attributes of the remaining 3D training output 264 (or the depicted view of the remaining 3D output) and the visual attributes of the training anchor content 252.
[0070] In other words, the layout loss 260 can be defined to cause the machine learning model 280 to generate a 3D training output 264 that includes a certain type of training anchor content 252 arranged in a certain portion of the physical space represented by a set of training sensor data 250 and / or a virtual twin generated from the set of training sensor data 250. For example, the layout loss 260 can be defined to cause the machine learning model 280 to generate a 3D training output 264 that includes appliances, dishes, glassware, clocks, people, animals, and / or other individual objects (identified by the training content section 268 and / or the ground truth data section 248) within the training anchor content 252 placed on an object having a table, coffee table, desk, fireplace, rug, floor, and / or other horizontal surface (identified by the training space section 266 and / or the ground truth data section 248). In another example, the layout loss 260 can be defined to cause the machine learning model 280 to generate a 3D training output 264 that includes clouds, rain, stars, light rays, vines, or other celestial or hanging objects (identified by the training content section 268 and / or the ground truth data section 248) within the training anchor content 252 along or near the ceiling or upper part of the room (identified by the training space section 266 and / or the ground truth data section 248). In a third example, the layout loss 260 can be defined to cause the machine learning model 280 to generate a training output 264 that includes a river or waterfall (identified by the training content section 268 and / or the ground truth data section 248) within the training anchor content 252 flowing through an empty or undecorated space (identified by the training space section 266 and / or the ground truth data section 248).
[0071] In one or more embodiments, the layout loss 260 is used to ensure semantic and / or spatial consistency over time steps associated with the 3D training output 264. For example, the 3D training output 264 may include a series of 3D volumes depicting the placement of objects from corresponding video frames or time steps within the training anchor content 252 in the physical space represented by the training sensor data 210. The training engine 122 can provide one or more previous frames within the 3D training output 264, the corresponding training space segment 266, the corresponding training content segment 268, and / or the corresponding frame or time step within the training anchor content 252 as additional inputs to the space segment network 242, the content segment network 244, and / or the 3D synthesis network 246, notifying the generation of the current frame of the 3D training output 264 from the current frame of the training anchor content 252. The training engine 122 can also calculate one or more layout losses 260 between or over consecutive 3D volumes within the 3D training output 264, ensuring that objects or videos within consecutive frames or time steps within the training anchor content 252 are placed in generally the same position within the physical space (without bouncing or fluctuating irregularly).
[0072] The training engine 122 can also, or alternatively, calculate one or more reconstruction losses 258 between the 3D representation of the training anchor content 252 output by the 3D synthesis network 246 and the corresponding training 3D object 254. For example, the training engine 122 can calculate one or more reconstruction losses 258 between the 3D model within the training 3D object 254 and the 3D representation generated from the corresponding portion of the 2D training anchor content 252 by the 3D synthesis network 246. Thus, the reconstruction loss 258 enables the machine learning model 280 to learn to transform the depiction of 2D objects within the training anchor content 252 into the corresponding training 3D objects 254.
[0073] After calculating the similarity loss 256, the reconstruction loss 258, the layout loss 260, and / or the segmentation loss 262 for a particular forward pass, the training engine 122 performs a backward pass that updates the parameters of the spatial segmentation network 242, the content segmentation network 244, and / or the 3D synthesis network 246 using various permutations or combinations of the similarity loss 256, the reconstruction loss 258, the layout loss 260, and / or the segmentation loss 262. For example, the training engine 122 can use a training technique (e.g., gradient descent and backpropagation) to update the parameters of the spatial segmentation network 242 based on the segmentation loss 262 calculated between the training spatial segmentation 266 and the corresponding ground truth data segmentation 248. The training engine 122 can also update the parameters of the content segmentation network 244 based on the segmentation loss 262 calculated between the training content segmentation 268 and the corresponding ground truth data segmentation 248. After the training of the spatial segmentation network 242 and the content segmentation network 244 is complete, the training engine 122 can freeze the parameters of the spatial segmentation network 242, the content segmentation network 244, and the 3D synthesis network 246 based on the layout loss 260, the similarity loss 256, and / or the reconstruction loss 258. The training engine 122 can also continue to alternately train the spatial segmentation network 242 and the content segmentation network 244 based on the segmentation loss 262 and the 3D synthesis network 246 based on the layout loss 260, the similarity loss 256, and / or the reconstruction loss 258 until the segmentation loss 262, the layout loss 260, the similarity loss 256, and / or the reconstruction loss 258 are below their corresponding thresholds.
[0074] In another example, the training engine 122 can perform end-to-end training of the spatial segmentation network 242, the content segmentation network 244, and the 3D synthesis network 246 using the layout loss 260, the similarity loss 256, the reconstruction loss 258, and / or other losses calculated based on the 3D training output 264. This end-to-end training can be performed alternately with the training of the spatial segmentation network 242 and the content segmentation network 244 based on the corresponding segmentation loss 262 and / or in another manner after the spatial segmentation network 242 and the content segmentation network 244 have been trained based on the segmentation loss 262.
[0075] The operation of the training engine 122 has been described above with respect to a certain type of loss. However, it will be understood that the spatial segmentation network 242, the content segmentation network 244, and / or the 3D synthesis network 246 can be trained using other types of techniques, losses, and / or machine learning components. For example, the training engine 122 can train the spatial segmentation network 242, the content segmentation network 244, and / or the 3D synthesis network 246 in an adversarial manner using one or more discriminator neural networks (not shown) and using one or more discriminator losses calculated from the ground truth data segmentation 248, the training spatial segmentation 266, the training content segmentation 268, and / or the 3D training output 264 based on the predictions generated by the discriminator neural network. In another example, the training engine 122 can train the spatial segmentation network 242, the content segmentation network 244, and / or the 3D synthesis network 246 using one or more losses that reflect artist guidance or parameters in the 3D training output 264. This artist guidance or parameter is related to symmetry, balance, color, orientation, scale, style, the movement of the training anchor content 252 across the physical space represented by the training sensor data 250, the size of the objects within the training anchor content 252 drawn or placed within the physical space represented by the training sensor data 250, and / or other attributes that affect the way the training anchor content 252 is drawn or placed in the physical space represented by the training sensor data 250. In a third example, the training engine 122 can train the spatial segmentation network 242, the content segmentation network 244, and / or the 3D synthesis network 246 to learn to expand specific fragments of the 2D or 3D training anchor content 252 across various physical spaces.
[0076] When the training of the machine learning model 280 is completed, the execution engine 124 generates an output 3D volume 284 using one or more components of the trained machine learning model 280. In particular, the execution engine 124 uses the spatial segmentation network 242 to convert sensor data 274 (e.g., images, point clouds, grids, depth maps, etc.) collected from the physical space into a semantic layout 276 that identifies the position, shape, and orientation of an object 278 within that physical space. The execution engine 124 also uses the content segmentation network 244 to convert a particular piece of anchor content 296 (e.g., an image, painting, video, etc.) captured in the sensor data 274 or provided separately from the sensor data 274 into a corresponding content segmentation 282. The execution engine 124 then uses the 3D synthesis network 246 to generate a latent representation 288 of the object 278 within the layout 276 and / or a portion of the anchor content 296 identified using the content segmentation 282. The execution engine 124 also uses the 3D synthesis network 246 to generate the output 3D volume 284 based on the latent representation 288 (e.g., by sampling, transforming, correlating, or otherwise processing the latent representation 288). As described above, the output 3D volume 284 includes a set of 3D objects 286 (e.g., doors, windows, furniture, decorations, etc.) from the physical space in their respective positions and includes a second set of 3D objects 286 corresponding to 3D representations of various portions of the anchor content 296 arranged within the physical space.
[0077] The execution engine 124 also incorporates various views of the output 3D volume 284 into the AR environment 292. For example, the execution engine 124 can estimate and / or determine the perspective related to the AR system that provides the AR environment 292 based on an image of the physical space captured by the AR system, inertial sensor data from the AR system, the position and orientation of the computing device corresponding to the AR system, and / or other sensor data 228 related to or captured by the AR system. The execution engine 124 can use the machine learning model 280 to depict views of the output 3D volume 284 from the perspective of the AR system that provides the AR environment 292. As a result, the AR environment 290 can depict a semantically meaningful placement of the 3D representation of the anchor content 296 across the physical space from the perspective of the user interacting with the AR system. When the user changes the perspective of the AR system and / or updates the anchor content 296 (e.g., by depicting additional portions of the anchor content 296, trimming, shrinking / enlarging, and / or deforming the anchor content 296, changing the color balance, saturation, color temperature, exposure, luminance, and / or other color-related attributes of the anchor content 296, changing the view of the anchor content 296, loading the anchor content 296 from a different file, playing a video or animation containing multiple frames of the anchor content 296, etc.), the execution engine 124 receives the updated anchor content 296 and / or updated sensor data 274 reflecting the latest view or representation of the physical space, generates a new layout 276 from the additional sensor data 274 and / or new content segments 282 from the new anchor content 296, generates a new output 3D volume 284 depicting 3D objects 286 related to the new anchor content 296, and can depict that new output 3D volume from the latest view. As a result, the execution engine 124 generates an immersive AR environment 292 that continuously responds to changes in the sensor data 274 and / or the anchor content 296.
[0078] In some embodiments, the output 3D volume 284 is used in other types of immersive environments, such as virtual reality (VR) and / or mixed reality (MR) environments. This content can depict a virtual world that can be continuously experienced synchronously by any number of users while providing continuity of data such as an individual's identity, user history, credentials, property, and / or rewards. Note that this content can include conventional audiovisual content and a hybrid of fully immersive VR, AR, and / or MR experiences, such as two-way video.
[0079] FIG. 3A shows an example layout of a physical space according to various embodiments. As shown in FIG. 3A, the layout example includes an object 302 depicting a piece of anchor content. For example, the object 302 can include a painting hung on a wall and / or an image or video displayed on a monitor or television screen.
[0080] The layout of FIG. 3A also includes a plurality of additional objects 304, 306, 308, 310, 312, and 314 arranged within that physical space. For example, the object 304 can consist of a set of picture frames, the object 306 can consist of a bookshelf, the object 308 can consist of a speaker system or soundbar, the object 310 can consist of a door, the object 312 can consist of a fireplace, and the object 314 can consist of a wall.
[0081] FIG. 3B shows an example of an AR environment that extends 2D anchor content within the physical space of FIG. 3A according to various embodiments. More specifically, FIG. 3B shows an AR environment that extends the anchor content depicted on the object 302 across the wall corresponding to the object 314. As shown in FIG. 3B, the extended anchor content does not obscure the real-world objects 302, 304, 306, 308, 310, 312 within the physical space. Instead, the anchor content is shown on the unadorned portion of the wall, and the objects 302, 304, 306, 308, 310, 312, and 314 are shown in their respective positions within the physical space. As a result, the anchor content is mixed with the layout of the physical space in a semantically meaningful way.
[0082] FIG. 3C shows an example of an AR environment that includes a view of a 3D volume related to the physical space of FIG. 3A according to various embodiments. In particular, FIG. 3C shows an AR environment that includes a view of a 3D volume. The 3D volume includes 3D representations of anchor content depicted on object 302 at various positions within the physical space.
[0083] More specifically, the 3D volume includes a plurality of 3D objects 322, 324, and 326 generated from various portions of the anchor content. Each of the 3D objects 322, 324, and 326 is disposed in an empty portion of the physical space within the 3D volume. The 3D volume also includes objects 302, 304, 306, 308, 310, 312, and 314 from the layout of FIG. 3A.
[0084] The 3D volume can also be used to render the 3D objects 322, 324, and 326 generated from the anchor content and the objects 302, 304, 306, 308, 310, 312, and 314 from various views. For example, a user can walk around the physical space with a computing device that provides the AR environment and view the 3D objects 322, 324, and 326 and / or the objects 302, 304, 306, 308, 310, 312, and 314 from various angles, view the 3D objects 322, 324, and 326 and / or the objects 302, 304, 306, 308, 310, 312, and 314 at various levels of detail, change the position, size, color, and / or other attributes of a particular 3D object 322, 324, or 326 and / or a particular real-world object 302, 304, 306, 308, 310, 312, or 314, and / or interact with the AR environment in other ways.
[0085] FIG. 4A shows an example layout of a physical space according to various embodiments. Similar to the layout of FIG. 3A, the example layout of FIG. 4A includes an object 402 that depicts a piece of anchor content. For example, object 402 can include a painting hung on a wall and / or an image displayed on a monitor or television screen.
[0086] The layout of FIG. 4A also includes a plurality of additional objects 404, 406, 408, 410, 412, and 414 arranged within its physical space. For example, object 404 may consist of a set of picture frames, object 406 may consist of a bookshelf, object 408 may consist of a speaker system or a soundbar, object 410 may consist of a door, object 412 may consist of a fireplace, and object 414 may consist of a wall.
[0087] FIG. 4B shows an example of an AR environment that extends 2D anchor content within the physical space of FIG. 4A according to various embodiments. More specifically, FIG. 4B shows an AR environment that extends the anchor content depicted on object 402 across the wall corresponding to object 414. As shown in FIG. 4B, the extended anchor content does not obscure the real-world objects 402, 404, 406, 408, 410, 412, and 414 within the physical space. Instead, the anchor content is shown on the undecorated portion of the wall, and the objects 402, 404, 406, 408, 410, 412, and 414 are shown at their respective positions within the physical space. As a result, the anchor content is mixed with the layout of the physical space in a semantically meaningful way.
[0088] FIG. 4C shows an example of an AR environment that includes a view of a 3D volume related to the physical space of FIG. 4A according to various embodiments. In particular, FIG. 4C shows an AR environment that includes a view of a 3D volume. The 3D volume includes 3D representations of the anchor content depicted on object 402 at various positions within the physical space.
[0089] More specifically, the 3D volume includes a plurality of 3D objects 422, 424, 426, and 428 generated from various portions of the anchor content. Each 3D object 422, 424, 426, and 428 is arranged in the empty portion of the physical space within the 3D volume. The 3D volume also includes a 2D extension of the anchor content across the wall corresponding to object 414.
[0090] The 3D volume can also be used to render 3D objects 422, 424, 426, and 428 generated from anchor content and objects 402, 404, 406, 408, 410, 412, and 414 from various views. For example, a user can walk around a physical space with an AR device that provides an AR environment, view 3D objects 422, 424, 426, and 428 and / or objects 402, 404, 406, 408, 410, 412, and 414 from various angles, view 3D objects 422, 424, 426, and 428 and / or objects 402, 404, 406, 408, 410, 412, and 414 at various levels of detail, change the position, size, color, and / or other attributes of a particular 3D object 422, 424, 426, or 428 and / or a particular real-world object 402, 404, 406, 408, 410, 412, or 414, and / or interact with the AR environment in other ways.
[0091] The AR environment examples of FIGS. 3B, 3C, 4B, and 4C depict an extension or placement of anchor content within the corresponding physical space, but it will be understood that the anchor content can be extended in other ways across the physical space. For example, a painting, image, or video depicted on object 302 or 402 can be extended across an entire wall corresponding to object 314 or 414 and overlaid on one or more of objects 302, 304, 306, 308, 310, 312, 402, 404, 406, 408, 410, and / or 412 visible from a viewpoint associated with the AR system providing the AR environment and / or used to obscure one or more of objects 302, 304, 306, 308, 310, 312, 402, 404, 406, 408, 410, and / or 412 visible from a viewpoint associated with the AR system providing the AR environment.
[0092] Generally, the way in which anchor content is extended, projected, or disposed within a particular space can be controlled based on the loss used to train the corresponding machine learning model (e.g., machine learning models 200 and / or 280) to generate AR content. For example, the machine learning model can be trained to extend, project, or dispose various fragments of the anchor content within the physical space such that as a result, one object within the physical space is not affected by the anchor content, one object is depicted as interacting with (e.g., supporting, containing, merging with, etc.) the anchor content, one object has the anchor content overlaid thereon, and / or one object is obscured by the anchor content.
[0093] FIG. 5 is a flow diagram of method steps for training a machine learning model according to various embodiments and generating AR content incorporating anchor content within the physical space. Although the method steps are described with the system of FIGS. 1 - 2B, those skilled in the art will understand that any system configured to execute the method steps in any order falls within the scope of the present disclosure.
[0094] As shown, in step 502, the training engine 122 operates a first set of neural networks that generate semantic partitions related to one or more physical spaces and / or one or more sets of anchor content. For example, the training engine 122 can input images, point clouds, grids, depth maps, and / or other sensor data or representations of the physical space into the first partition neural network. The training engine 122 operates the first partition neural network to obtain predictions of objects (e.g., walls, ceilings, floors, doorways, doors, windows, heaters, lighting fixtures, various types of furniture, various types of decorations, etc.) for various regions of the sensor data as the output of the first partition neural network. The training engine 122 can also, or instead, input one or more images, video frames, 3D objects, and / or other representations of a specific set of anchor content into the second partition neural network. The training engine 122 operates the second partition neural network to obtain predictions of objects (e.g., people, animals, plants, faces, figures, backgrounds, foregrounds, structures, etc.) for various regions or subsets of the anchor content as the output of the second partition neural network.
[0095] In step 504, the training engine 122 updates the parameters of the first set of neural networks based on one or more partition losses related to the semantic partition. Continuing with the above example, the training engine 122 can obtain a first set of ground truth data partitions related to the sensor data and / or a second set of ground truth data partitions related to the anchor content. Each ground truth data partition can include labels that identify objects corresponding to various regions, parts, or subsets of the corresponding sensor data and / or anchor content. The training engine 122 can calculate cross-entropy loss, dice loss, boundary loss, Tversky loss, and / or another measure of the error between a specific partition output by the partition neural network and the corresponding ground truth data partition. The training engine 122 can then use gradient descent and backpropagation to update the neural network weights within its partition neural network to reduce the measure of the error.
[0096] In step 506, the training engine 122 determines whether to continue training the first set of neural networks. For example, the training engine 122 may determine that each partition neural network should continue to be trained using the corresponding partition loss until one or more conditions are met. These conditions include, but are not limited to, convergence of the parameters of the partition neural network, reduction of the partition loss below a threshold, and / or a certain number of training steps, iterations, batches, and / or events. While training the first set of neural networks continues, the training engine 122 repeats steps 502 and 504.
[0097] When the training engine 122 determines that the training of the first set of neural networks is complete (step 506), the training engine 122 begins to train the second set of neural networks to generate AR content. More specifically, in step 508, the training engine 122 operates a second set of neural networks that generate 2D outputs and / or 3D outputs based on additional data related to semantic partitions and / or physical spaces and / or anchor content. For example, the training engine 122 may use one or more trained partition neural networks to generate semantic partitions of sensor data for a physical space and / or one or more sets of anchor content. The training engine 122 can input the semantic partitions along with the corresponding sensor data and anchor content into an extrapolation network and operate the extrapolation network to generate a 2D output including one or more images. Each image can combine the attributes of the physical space represented by the sensor data with the attributes of a set of anchor content. The training engine 122 can also, or alternatively, input the semantic partitions along with the corresponding sensor data and anchor content into a 3D synthesis network and operate the 3D synthesis network to generate a 3D volume. The 3D volume can include another representation of a 3D scene that combines the visual and semantic attributes related to the neural radiance field and / or the input sensor data with the visual and semantic attributes related to the input anchor content.
[0098] In step 510, the training engine 122 updates the parameters of the second set of neural networks based on one or more losses associated with the 2D and / or 3D output. Continuing with the above example, the training engine 122 can train the extrapolation network by calculating a similarity loss as a measure of the difference between the visual attributes of the expansion of the anchor content across a portion of the physical space depicted in the 2D output generated by the extrapolation network and the corresponding portion of the anchor content. The training engine 122 can also, or instead, calculate a layout loss as a measure of the difference between a depiction of a portion of the physical space within the 2D output and the corresponding portion of the sensor data of the physical space. The training engine 122 can then use gradient descent and backpropagation to update the neural network weights within its extrapolation network to reduce the similarity loss and / or the layout loss.
[0099] The training engine 122 can train the 3D synthesis network by calculating a reconstruction loss based on one or more 3D representations of the anchor content generated by the 3D synthesis network and one or more ground truth data 3D objects associated with the anchor content. The training engine 122 can also, or instead, calculate a layout loss as a measure of the difference between a representation of a subset of the physical space within the first 3D volume and the corresponding subset of the sensor data of the physical space. The training engine 122 can also, or instead, calculate another layout loss based on the placement of one or more 3D representations of the anchor content within the first 3D volume. The training engine 122 can also, or instead, calculate a similarity loss as a measure of the difference between the visual attributes of one or more rendered views of the anchor content within the 3D volume and the corresponding portion of the anchor content. The training engine 122 can then use gradient descent and backpropagation to update the neural network weights within its 3D synthesis network to reduce the reconstruction loss, the similarity loss, and / or the layout loss.
[0100] In step 512, the training engine 122 determines whether to continue training the second set of neural networks. For example, the training engine 122 may determine that the extrapolation network and / or the 3D synthesis network should continue to be trained using the corresponding loss until one or more conditions are met. These conditions include, but are not limited to, the convergence of the neural network parameters, the decrease of the loss below a threshold, and / or a certain number of training steps, iterations, batches, and / or events. While the training of the second set of neural networks continues, the training engine 122 repeats steps 508 and 510. Next, when the conditions are met, the training engine 122 ends the process of training the second set of neural networks.
[0101] The training engine 122 can also repeat steps 502, 504, 506, 508, 510, and 512 one or more times to continue training the first and / or second sets of neural networks. For example, the training engine 122 can alternate training the first and second sets of neural networks until all losses meet their respective thresholds, a certain number of executions of each type of training are performed, and / or another condition is met. The training engine 122 can also, or instead, perform one or more rounds of end-to-end training of the first and second sets of neural networks to optimize the operation of all neural networks for the task of generating AR content from the description of the anchor content and the physical space.
[0102] FIG. 6 is a flowchart of method steps for generating an AR environment that incorporates a set of anchor content according to various embodiments into the layout of a physical space. Although the method steps are described in conjunction with the systems of FIGS. 1 - 2B, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.
[0103] As shown in the figure, at step 602, the execution engine 124 determines the layout of the physical space based on sensor data related to the physical space. For example, the execution engine 124 may receive sensor data as a set of images, point clouds, grids, depth maps, and / or other representations of visual or spatial attributes of the physical space. The execution engine 124 may also use the first segmentation neural network to generate the layout as a semantic segmentation of the sensor data. The semantic segmentation may include predictions of real-world objects (such as floors, ceilings, walls, decorations, desks, chairs, sofas, rugs, paintings, etc.) for various regions or subsets of the sensor data.
[0104] At step 604, the execution engine 124 determines a semantic segmentation related to a set of anchor content. The set of anchor content may include one or more fragments of the anchor content, and a particular fragment of the anchor content may include, but is not limited to, images, video frames, 2D or 3D shapes, textures, audio clips, and / or other types of content incorporated into the AR environment. The execution engine 124 can input one or more images, video frames, 3D models, and / or other types of anchor content into the second segmentation neural network. The anchor content may be selected by the user as a subset of the sensor data of the physical space and / or provided separately from the sensor data. The execution engine 124 can also obtain the semantic segmentation as the output of the second segmentation neural network. The semantic segmentation includes predictions of objects (such as faces, people, scenes, situations, animals, plants, buildings, vehicles, structures, shapes, etc.) based on the content for various regions or subsets of the anchor content.
[0105] At step 606, the execution engine 124 inputs the sensor data, anchor content, layout, and / or semantic segmentation into a machine learning model. For example, the execution engine 124 can input the sensor data, anchor content, layout, and / or semantic segmentation into an extrapolation neural network that generates a 2D output. The execution engine 124 can also, or instead, input the sensor data, anchor content, layout, and / or semantic segmentation into a 3D synthesis neural network that generates a 3D output.
[0106] In step 608, the execution engine 124 generates one or more images and / or 3D volumes that include a depiction of a first portion of the physical space and a representation of anchor content within a second subset of the physical space by operation of the machine learning model. For example, the execution engine 124 may use an extrapolation network to generate six images corresponding to the six faces of a cube representing a standard box-shaped room. The execution engine 124 may also, or instead, use an extrapolation network to generate one or more images depicting a 360-degree, spherical, and / or another type of panoramic view of the physical space not limited to a box-shaped room. Each image may depict real-world objects within the room, such as doors, windows, furniture, and / or decorations. Each image may also depict various components of the anchor content superimposed on the walls, floor, ceiling, and / or other surfaces within the room. These components of the anchor content may also be arranged or distributed within the corresponding images so as to avoid obscuring and / or partially overlapping doors, windows, furniture, decorations, and / or other objects within the room.
[0107] In another example, the execution engine 124 may use a 3D synthesis network to generate a 3D output that includes a neural radiance field and / or another representation of a 3D scene. The 3D output may include 3D representations of real-world objects from the physical space, such as doors, windows, furniture, and / or decorations. The 3D output may also include 2D or 3D objects corresponding to various components of the anchor content and located within the empty portions of the room and / or arranged on a certain type of surface. These components of the anchor content may also be located within the 3D volume and be able to avoid obscuring and / or partially overlapping doors, windows, furniture, decorations, and / or other objects within the room. These components of the anchor content may also, or instead, be arranged to allow the components to interact or blend with the objects within the room.
[0108] In step 610, the execution engine 124 causes one or more views of the image and / or 3D volume to be output within the AR environment provided by the computing device. For example, the execution engine 124 incorporates the image into the visual output generated by the display of the computing device, such that the AR environment may appear to the user interacting with the AR system to depict a semantically meaningful extension of the anchor content across the physical space from the user's perspective. In another example, the execution engine 124 may use a 3D compositing circuitry to render a view of the 3D volume from the perspective of the computing device and cause the computing device to output that view to the user.
[0109] In step 612, the execution engine 124 determines whether to continue providing the AR environment. For example, the execution engine 124 may determine that the AR environment should be provided while an application that realizes the AR environment is operating on the computing device and / or while the user is interacting with the AR environment. If the AR environment should be provided, the execution engine 124 repeats steps 602, 604, 606, 608, and 610 to update the AR environment in response to changes in the perspective of the computing device, the physical space, and / or the anchor content. For example, the execution engine 124 can receive updates to the anchor content as a "draw" input from the user; trimming, shrinking / enlarging, and / or other deformations of the anchor content; changing color balance, saturation, color temperature, exposure, luminance, and / or other color-related attributes of the anchor content; applying sharpening, blurring, noise removal, distortion, or other changes to the anchor content; changing the view of the anchor content; selecting the anchor content from one or more files; and / or playing a video including multiple frames of the anchor image. The execution engine 124 can also, or instead, receive additional sensor data that reflects the movement of the computing device and / or changes in the physical space. The execution engine 124 can execute step 602 to generate a new layout from the additional sensor data and execute step 604 to generate a semantic segmentation from the anchor content. The execution engine 124 can also execute steps 606 and 608 to generate a new 2D or 3D output that combines the physical space from the current perspective with the latest anchor content. As a result, the execution engine 124 generates an immersive AR environment that enables the user to explore a semantically meaningful combination with the anchor content in the physical space.
[0110] In summary, the disclosed approach generates AR content that extends or places 2D or 3D representations of photos, paintings, video frames, drawn scenes, video scenes, or other types of anchor content within a physical space. The physical space is represented by one or more images, point clouds, grids, depth maps, and / or other types of sensor data. A machine learning model is used to generate a first semantic segmentation of the sensor data and a second semantic segmentation of the anchor content. The first semantic segmentation includes predictions of objects typically found in a room (or other type of physical space) for various regions or subsets of the sensor data. The second semantic segmentation includes predictions of objects typically associated with the anchor content for various regions or subsets of the anchor content.
[0111] Another part of the machine learning model is used to convert the semantic segmentations, sensor data, and anchor content into AR content. The AR content can include one or more images depicting the physical space from a certain viewpoint. Those images include representations of certain types of real-world objects (e.g., doors, windows, furniture, etc.) within the physical space and an extension of the AR content across other types of real-world objects (e.g., walls, ceilings, floors, etc.) within the physical space. The AR content can also, or alternatively, include one or more views of a 3D volume generated by the machine learning model. Those views can include a depiction of the physical space and the placement of 3D representations of the AR content at various positions within the physical space.
[0112] The AR content is output to an AR, VR, and / or mixed reality environment provided by a mobile electronic device, a wearable device, and / or another type of computing device. When the position and orientation of the computing device, the physical space, and / or the anchor content change, the machine learning model is used to generate updated AR content that reflects the changes in the sensor data and / or the anchor content. The updated AR content is also output to the AR environment and allows a user of the computing device to explore, modify, or interact with the AR environment.
[0113] One technical advantage of the disclosed approach over the prior art is the ability to automatically and seamlessly generate AR content that couples anchor content with the layout of the physical space. Thus, the disclosed approach is more time and resource efficient than conventional approaches that use various software components to convert conventional media content into AR assets and manually place the AR assets within the AR environment. Another technical advantage of the disclosed approach is the ability to generate AR content on-the-fly from a particular set of anchor content and a particular physical space. As a result, the disclosed approach increases the diversity and availability of AR content that couples conventional media content with the layout of the physical space. These technical advantages provide one or more technical improvements over prior art techniques.
[0114] 1. In some embodiments, a computer-executable method for generating augmented reality content includes inputting a first layout of a physical space and a first set of anchor content into a machine learning model, and generating, by operation of the machine learning model, a first three-dimensional (3D) volume including (1) a first subset of the physical space and (2) placement of one or more 3D representations of the first set of anchor content within a second subset of the physical space, and outputting one or more views of the first 3D volume to a computing device.
[0115] 2. The step of generating the first 3D volume includes applying a first set of neural network layers included in the machine learning model to the first set of anchor content to generate a semantic segmentation of the first set of anchor content, and applying a second set of neural network layers included in the machine learning model to the first layout, the first set of anchor content, and the semantic segmentation to generate the first 3D volume, the computer-executable method of claim 1.
[0116] 3. The step of generating the first 3D volume includes applying a first set of neural network layers included in the machine learning model to the first set of anchor contents to generate the one or more 3D representations, and applying a second set of neural network layers included in the machine learning model to the one or more 3D representations and the first layout to determine the placement of the one or more 3D representations within the second subset of the physical space. The computer-executed method according to claim 1 or 2.
[0117] 4. The computer-executed method according to any one of claims 1 to 3, further including the step of generating the first layout as a semantic segmentation of sensor data related to the physical space by the operation of the machine learning model.
[0118] 5. The computer-executed method according to any one of claims 1 to 4, wherein the sensor data includes at least one of an image, a point cloud, a grid, or a depth map of the physical space.
[0119] 6. The computer-executed method according to any one of claims 1 to 5, further including the step of training the machine learning model based on a set of training layouts, a set of training anchor images, and one or more losses related to the first 3D volume.
[0120] 7. The computer-executed method according to any one of claims 1 to 6, wherein the one or more losses include a layout loss calculated based on the representation of the first subset of the physical space within the first 3D volume and the corresponding subset of the physical space.
[0121] 8. The computer-executed method according to any one of claims 1 to 7, wherein the one or more losses include a layout loss calculated based on the first layout and the placement of the one or more 3D representations of the first set of anchor contents within the first 3D volume.
[0122] 9. The computer-executed method according to any one of claims 1 to 8, wherein the first 3D volume includes a neural radiance field.
[0123] 10. The computer-executed method according to any one of items 1 to 9, wherein the first set of anchor content includes at least one of an image, a video, or a 3D object.
[0124] 11. In some embodiments, one or more persistent computer-readable media store a set of instructions that, when executed by one or more processors, cause the one or more processors to input a first layout of a physical space and a first set of anchor content into a machine learning model, and generate, by operation of the machine learning model, a first three-dimensional (3D) volume including (1) a first subset of the physical space and (2) the placement of one or more 3D representations of the first set of anchor content within a second subset of the physical space, and output one or more views of the first 3D volume to a computing device.
[0125] 12. The one or more persistent computer-readable media according to item 11, wherein the set of instructions further cause the one or more processors to apply a set of neural circuit layers included in the machine learning model to sensor data related to the physical space to generate the first layout, and the first layout includes predictions of a plurality of objects for a plurality of regions of the sensor data.
[0126] 13. The one or more persistent computer-readable media according to item 11 or 12, wherein the set of instructions further cause the one or more processors to generate, by operation of the machine learning model, a second 3D volume including (1) a third subset of the physical space and (2) the placement of one or more 3D representations of a second set of anchor content within a fourth subset of the physical space, and output one or more views of the second 3D volume to the computing device.
[0127] 14. The one or more persistent computer-readable media according to any one of items 11 to 13, wherein the first set of anchor content and the second set of anchor content include at least one of a plurality of different video frames included in a video, a depiction of two different scenes, or a plurality of different sets of 3D objects.
[0128] 15. The one or more persistent computer-readable media according to any one of items 11 to 14, wherein the set of instructions further causes the one or more processors to perform a step of training the machine learning model based on one or more losses associated with the first 3D volume.
[0129] 16. The one or more persistent computer-readable media according to any one of items 11 to 15, wherein the one or more losses include a similarity loss calculated based on the first set of anchor content and a depiction of the second subset of the physical space within the 3D volume.
[0130] 17. The one or more persistent computer-readable media according to any one of items 11 to 16, wherein the one or more losses include a reconstruction loss calculated based on the one or more 3D representations of the first set of anchor content generated by the machine learning model and one or more 3D objects.
[0131] 18. The one or more persistent computer-readable media according to any one of items 11 to 17, wherein the one or more losses include a segmentation loss calculated based on a semantic segmentation of the first set of anchor content and a ground truth data segmentation associated with the first set of anchor content.
[0132] 19. The step of outputting one or more views of the first 3D volume to the computing device includes rendering the first 3D volume from the one or more views and outputting the one or more views within an augmented reality environment provided by the computing device. The one or more persistent computer-readable media according to any one of items 11 to 18.
[0133] 20. In some embodiments, the system comprises one or more memories storing a set of instructions and one or more processors coupled to the one or more memories, and when executing the set of instructions, the one or more processors are configured to: input a first representation of a physical space and a first set of anchor content into a machine learning model; generate, by operation of the machine learning model, a first three-dimensional (3D) volume including (1) a first subset of the physical space and (2) the placement of one or more 3D representations of the first set of anchor content within a second subset of the physical space; and cause one or more views of the first 3D volume to be output to a computing device.
[0134] Any and / or all combinations of any of the claimed elements and / or any of the elements described herein fall within the scope of the invention and the scope of protection, in any manner.
[0135] The descriptions of the various embodiments are presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0136] Aspects of the present embodiments may be embodied in a system, a method, or a computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software aspects and hardware aspects that may generally be referred to herein as a "module," "system," or "computer." Any of the hardware and / or software techniques, processes, functions, components, engines, modules, or systems described in this disclosure may also be implemented as a circuit or a set of circuits. Additionally, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code incorporated therein.
[0137] Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this specification, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0138] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / acts specified in the flowchart and / or implement the blocks in the block diagram. Such a processor may be, by way of non-limiting example, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.
[0139] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. Note that in other embodiments, the functions recited in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the functions may sometimes be executed in the reverse order, depending upon the functionality involved. Also, each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by special purpose hardware systems or combinations of special purpose hardware and computer instructions that perform the specified functions or operations.
[0140] Although the foregoing is directed to embodiments of the present disclosure, other and additional embodiments of the present disclosure may occur to those skilled in the art without departing from the basic scope of the present disclosure. The scope of the present disclosure is determined by the appended claims.
Explanation of Reference Numerals
[0141] 100 Computing device 102 Processor 104 I / O device interface 106 Network interface 108 I / O device 110 Network 112 Interconnection (bus) 114 Storage device 116 Memory 122 Training engine 124 Execution engine 302 Object (painting hung on a wall or image displayed on a monitor) 304 Object (a set of picture frames) 306 Object (bookshelf) 308 Object (Speaker system or sound bar) 310 Object (Door) 312 Object (Furnace) 314 Object (Wall)
Claims
1. A computer-executed method for generating augmented reality content, comprising: inputting a first layout of a physical space and a first set of anchor content represented within the physical space into a machine learning model; generating, by operation of the machine learning model, a first three-dimensional (3D) volume including (1) a first subset of the physical space that includes the first set of anchor content and (2) an arrangement of one or more 3D representations of the first set of anchor content within a second subset of the physical space, the arrangement of the one or more 3D representations being placed at different positions within the physical space relative to the positions of the first set of anchor content within the first subset of the physical space; causing a computing device to output one or more views of the first 3D volume within an augmented reality environment provided by the computing device; and a computer-executed method including the above.
2. The step of generating the first 3D volume includes: applying a first set of neural network layers included in the machine learning model to the first set of anchor content to generate a semantic classification of the first set of anchor content; applying a second set of neural network layers included in the machine learning model to the first layout, the first set of anchor content, and the semantic classification to generate the first 3D volume. The computer-executed method according to claim 1, including the above.
3. The step of generating the first 3D volume includes: applying a first set of neural network layers included in the machine learning model to the first set of anchor content to generate the one or more 3D representations; applying a second set of neural network layers included in the machine learning model to the one or more 3D representations and the first layout to determine the arrangement of the one or more 3D representations within the second subset of the physical space. The computer-executed method according to claim 1, including the above.
4. The computer-executed method according to claim 1, further including generating the first layout as a semantic classification of sensor data related to the physical space by operation of the machine learning model.
5. The computer-executed method according to claim 4, wherein the sensor data includes at least one of an image, a point cloud, a grid, or a depth map of the physical space.
6. The computer-executed method according to claim 1, further comprising training the machine learning model based on a set of training layouts, a set of training anchor images, and one or more losses associated with the first 3D volume.
7. The computer-executed method according to claim 1, wherein the one or more losses include a layout loss calculated based on a representation of the first subset of the physical space within the first 3D volume and a corresponding subset of the physical space.
8. The computer-executed method according to claim 1, wherein the one or more losses include a layout loss calculated based on the first layout and the arrangement of the one or more 3D representations of the first set of anchor contents within the first 3D volume.
9. The computer-executed method according to claim 1, wherein the first 3D volume includes a neural radiance field.
10. The computer-executed method according to claim 1, wherein the first set of anchor contents includes at least one of an image, a video, or a 3D object.
11. One or more persistent computer-readable media storing a set of instructions, which, when executed by one or more processors, cause the one or more processors to input a first layout of a physical space and a first set of anchor contents represented within the physical space into a machine learning model; generate, by operation of the machine learning model, a first three-dimensional (3D) volume including (1) a first subset of the physical space including the first set of anchor contents and (2) an arrangement of one or more 3D representations of the first set of anchor contents within a second subset of the physical space, the arrangement of the one or more 3D representations being placed at different positions within the physical space relative to the positions of the first set of anchor contents within the first subset of the physical space; output one or more views of the first 3D volume within an augmented reality environment provided by a computing device; and perform the steps. One or more persistent computer-readable media.
12. The command group further causes the one or more processors to apply a set of neural circuit layers included in the machine learning model to sensor data related to the physical space to generate the first layout, the first layout including predictions of a plurality of objects for a plurality of regions of the sensor data, the one or more persistent computer-readable media of claim 11.
13. The command group causes the one or more processors to generate a second 3D volume including (1) a third subset of the physical space and (2) the placement of one or more 3D representations of a second set of anchor content within a fourth subset of the physical space by operation of the machine learning model; output one or more views of the second 3D volume to the computing device; and further cause the one or more persistent computer-readable media of claim 11.
14. The first set of anchor content and the second set of anchor content include at least one of a plurality of different video frames included in a video, a depiction of two different scenes, or a plurality of different sets of 3D objects, the one or more persistent computer-readable media of claim 13.
15. The command group further causes the one or more processors to train the machine learning model based on one or more losses related to the first 3D volume, the one or more persistent computer-readable media of claim 11.
16. The one or more losses include a similarity loss calculated based on the first set of anchor content and the depiction of the second subset of the physical space within the 3D volume, the one or more persistent computer-readable media of claim 15.
17. The one or more losses include a reconstruction loss calculated based on the one or more 3D representations of the first set of anchor content generated by the machine learning model and one or more 3D objects, the one or more persistent computer-readable media of claim 15.
18. The one or more losses include a segmentation loss calculated based on a semantic segmentation of the first set of anchor content and a ground truth data segmentation related to the first set of anchor content, the one or more persistent computer-readable media of claim 15.
19. The step of causing the computing device to output one or more views of the first 3D volume comprises: drawing the first 3D volume from the one or more views; and outputting the one or more views within an augmented reality environment provided by the computing device. One or more persistent computer-readable media according to claim 11. **Claim 20** A system comprising: one or more memories storing a set of instructions; and one or more processors coupled to the one or more memories, wherein the one or more processors, when executing the set of instructions, perform steps of: inputting a first representation of a physical space and a first set of anchor content represented within the physical space into a machine learning model; generating, by operation of the machine learning model, a first three-dimensional (3D) volume including (1) a first subset of the physical space that includes the first set of anchor content and (2) an arrangement of one or more 3D representations of the first set of anchor content within a second subset of the physical space, the arrangement of the one or more 3D representations being placed at different positions within the physical space relative to positions of the first set of anchor content within the first subset of the physical space; and causing the computing device to output one or more views of the first 3D volume within an augmented reality environment provided by the computing device. A system configured to perform the steps.
Citation Information
Patent Citations
Theme-based extension of realistically represented views
JP2014515130A
Location-based virtual element modality in three-dimensional content
JP2020042802A
Matching content to spatial 3D environments
JP2021527247A
Image classification model training method, image processing method and device, and computer program
JP2022505775A
Rendering new images of scenes using geometry-aware neural networks conditioned on latent variables
WO2022167602A2