A platform for enabling multiple users to generate and use a neural luminance field model
The computing system addresses the limitations of existing technologies by training neural luminance field models to generate view synthesis images of user objects, allowing for virtual representation, cataloging, and comparison of objects with enhanced visualization capabilities.
Patent Information
- Application Number
- JP2023212189
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-02-15
- Filing Date
- 2023-12-15
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2043-12-15
AI Technical Summary
Existing technologies limit users' ability to access 3D modeling, object segmentation, and novel view rendering, making it difficult for users to visualize and compare objects without physically rearranging them.
A computing system that trains neural luminance field models using user image data to generate view synthesis images of user objects, enabling users to create virtual representations and catalogs of objects, and compare them with uniform lighting and pose.
Enables users to learn 3D representations of objects and generate photorealistic view renderings from multiple viewpoints, facilitating virtual placement, comparison, and visualization of objects without physical rearrangement.
Smart Images

Figure 0007682985000001 
Figure 0007682985000002 
Figure 0007682985000003
Abstract
Description
Technical Field
[0001] Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 433,111, filed on Dec. 16, 2022, and U.S. Provisional Patent Application No. 63 / 433,559, filed on Dec. 19, 2022. U.S. Provisional Patent Application No. 63 / 433,111 and U.S. Provisional Patent Application No. 63 / 433,559 are hereby incorporated by reference in their entireties.
[0002] The present disclosure generally relates to a platform that enables multiple users to generate and use neural luminance field models to generate virtual representations of user objects. More particularly, the present disclosure relates to training one or more neural luminance field models to acquire user images and generate one or more novel view synthesis images of one or more objects depicted in the user images.
Background Art
[0003] 3D modeling, object segmentation, and novel view rendering may not be accessible to users. Such features can be useful for searching, visualizing rearranged environments, understanding objects, and comparing objects without the need to physically arrange them. Previous techniques for virtually viewing objects relied heavily on large amounts of data, including photographs and / or videos. Photographs involve viewing two-dimensionally from a single and / or limited number of views. Videos are similarly limited to explicitly captured data. User access to 3D modeling techniques may be inaccessible to users due to time costs and / or lack of knowledge of modeling programs.
[0004] Furthermore, photos may only provide the user with limited amounts of information. Size and compatibility with the new environment can be difficult to understand from an image. For example, a user may want to rearrange their room, but physically rearranging the room can be cumbersome just to see the possibilities. Users who use images may rely heavily on their imagination to understand size, lighting, and orientation. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0005] Aspects and advantages of embodiments of the present disclosure may be described in part in the following description, or may be apparent from the description, or may be learned through the practice of the embodiments.
[0006] One exemplary aspect of the present disclosure is directed to a computing system. The system may include one or more processors and one or more non-transitory computer-readable media that jointly store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include obtaining user image data and request data. The user image data may depict one or more images that include one or more user objects. The one or more images may have been generated by a user computing device. The operations may include training one or more neural luminance field models based on the user image data. The one or more neural luminance field models may be trained to generate view syntheses of one or more objects. The operations may include generating one or more view synthesis images with the one or more neural luminance field models based on the request data. In some implementations, the one or more view synthesis images may include one or more renderings of one or more objects.
[0007] Another exemplary aspect of the present disclosure is directed to a method implemented by a computer for generating a virtual closet. The method may include obtaining, by a computing system including one or more processors, a plurality of user images. Each of the plurality of user images may include one or more items of clothing. In some implementations, the plurality of user images may be associated with a plurality of different items of clothing. The method may include training, by the computing system, a respective neural luminance field model for each of the plurality of different items of clothing. Each respective neural luminance field model may be trained to generate one or more view synthesis renderings of a particular respective item of clothing. The method may include storing, by the computing system, each respective neural luminance field model in a collection database. The method may include providing, by the computing system, a virtual closet interface. The virtual closet interface may be capable of providing a plurality of clothing view synthesis renderings for display based on the plurality of respective neural luminance field models. The plurality of clothing view synthesis renderings may be associated with at least a subset of the plurality of different items of clothing.
[0008] Another exemplary aspect of the present disclosure is directed to one or more non-transitory computer-readable media that jointly store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations may include obtaining a plurality of user image data sets. Each user image data set of the plurality of user image data sets may depict one or more images that include one or more objects. In some implementations, the one or more images may have been generated on a user computing device. The operations may include processing the plurality of user image data sets with one or more classification models to determine a subset of the plurality of user image data sets that includes features describing one or more specific object types. The operations may include training a plurality of neural luminance field models based on the subset of the plurality of user image data sets. In some implementations, each neural luminance field model may be trained to generate a view synthesis of one or more specific objects of each user image data set of the subset of the plurality of user image data sets. The operations may include generating a plurality of view synthesis renderings with the plurality of neural luminance field models. The plurality of view synthesis renderings may depict a plurality of different objects of a specific object type. The operations may include providing a user interface for viewing the plurality of view synthesis renderings.
[0009] Systems and methods can be utilized to learn 3D representations of user objects, and then those 3D representations can be utilized to generate a virtual catalog of user objects. Additionally and / or alternatively, the systems and methods can be utilized to compare user objects and / or other objects. The comparison may be assisted by rendering view synthesis renderings of different objects with uniform illumination and / or a uniform pose. For example, images of different objects may depict objects with different illumination, positions, and / or distances. The systems and methods disclosed herein can be utilized to learn 3D representations of objects and generate view synthesis renderings of different objects with uniform illumination, a uniform pose, and / or a uniform scaling.
[0010] Other aspects of the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0011] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.
[0012] A detailed examination of the embodiments directed to those skilled in the art referring to the accompanying drawings is described herein.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 9C
Figure 10
Figure 11
Figure 12
Figure 13
DETAILED DESCRIPTION OF THE INVENTION
[0014] Reference numbers repeated across multiple figures are intended to identify the same features in various implementations.
[0015] Generally, the present disclosure is directed to systems and methods for providing a platform for a user to train and / or utilize a neural radiance field model for rendering user objects. In particular, the systems and methods disclosed herein can leverage one or more neural radiance field models and user images to learn a three-dimensional representation of a user object. The trained neural radiance field model can enable a user to reposition their environment through augmented reality, which can avoid the physically burdensome process of manually repositioning a room only to revert the object if its appearance is not what the user desires. Additionally and / or alternatively, the trained neural radiance field model can be utilized to generate a virtual catalog (e.g., a virtual closet) of user objects. The virtual catalog can include renderings of the user objects in a uniform pose and / or uniform lighting that can enable a uniform depiction of the user objects (e.g., for comparison of objects). Additionally and / or alternatively, generation of novel view synthesis images can be utilized to view various objects from various positions and orientations without physically traversing the environment.
[0016] A trained neural radiance field model can enable a user to create a unique live try-on experience with awareness of geometry. Utilization at an individual-based level can enable each individual user and / or a collection of users to train a neural radiance field that may be personalized to their own objects and / or objects within their own environment. Personalization can enable virtual placement changes, virtual comparisons, and / or geometry-aware and position-aware visualizations anywhere. Additionally and / or alternatively, the systems and methods can include view synthesis rendering of different objects with uniform lighting, uniform pose, uniform positioning, and / or uniform scaling to provide a platform for comparing objects with context-aware rendering.
[0017] 3D modeling, object segmentation, and novel view rendering may not be normally accessible to users based on previous techniques. Such features can be useful for searching, visualizing an environment with placement changes, understanding objects, and comparing objects without the need to physically arrange them.
[0018] The systems and methods disclosed herein can utilize a platform for providing a user with neural radiance field (NERF) technology to enable the user to create, store, share, and view high-quality 3D content at a broad level. The systems and methods can be useful for reformation, clothing design, object comparison, and / or catalog generation (e.g., a merchant can build high-quality 3D content of a product and add it to their website).
[0019] In some implementations, the systems and methods disclosed herein can be utilized to generate a three-dimensional model of a user object and render a synthetic image of the user object. The systems and methods may utilize objects from a collection of user photos to learn the three-dimensional model. The learned three-dimensional model can be utilized to render a particular combination of an object and an environment that can be controlled by the user and / or based on a context such as the user's search history. Additionally and / or alternatively, the user can manipulate the rendering (e.g., "scroll"). For example, novel view synthesis using a trained neural radiance field model can be utilized to enable the user to view the object from different views without physically traversing the environment.
[0020] A platform that enables the user to generate and utilize a neural radiance field model can enable the user to visualize "their" object along with other objects and / or other environments. Additionally and / or alternatively, a particular combination of an object and / or an environment can be a function of unique inputs such as object characteristics like search history, availability, price, etc., which can provide a context-aware user-specific experience.
[0021] Additionally and / or alternatively, the systems and methods disclosed herein may include a platform that provides an interface for a user to train a neural radiance field model with user-generated content (e.g., images) provided by the user. The trained neural radiance field model may be utilized to generate a virtual representation of user objects that can be added to a collection of users that can be utilized for things like composition, comparison, sharing, etc. The collection of users may include a virtual closet, a virtual furniture catalog, a virtual collectibles catalog (e.g., a user may generate a virtual representation of their physical collection (e.g., a collection of bobbleheads)), and / or a virtual trophy collection. The platform can enable the user to generate photorealistic view renderings from multiple different viewpoints that can be accessed and displayed even when the user is not in proximity to the physical object. The platform may include sharing among users that can be utilized for social media, marketplace, and / or messaging.
[0022] The systems and methods of the present disclosure provide several technical effects and advantages. As an example, the systems and methods can learn a three-dimensional representation of user objects based on user images to provide view synthesis rendering of user objects. In particular, images captured on a user computing device can be processed to classify and / or segment the object. Then, by training a neural radiance field model, a three-dimensional modeling representation can be learned for one or more objects within the user image. The trained neural radiance field model can then be utilized for augmented reality rendering, novel view synthesis, and / or instance interpolation.
[0023] Another technical advantage of the systems and methods of the present disclosure is the ability to utilize one or more view synthesis images to provide a virtual catalog of user objects. For example, multiple neural radiance field models can be utilized to generate multiple view synthesis renderings of multiple user objects. The view synthesis renderings may be generated and / or extended based on uniform lighting, uniform scaling, and / or uniform pose. The multiple view renderings may be provided by a user interface for a user to easily view their objects from their phone or other computing device.
[0024] Another example of technical effects and advantages relates to improved computational efficiency and improved functionality of computing systems. For example, the systems and methods disclosed herein can utilize user images to reduce the computational cost of searching for images of objects online and can ensure that the correct objects are modeled.
[0025] Referring now to the figures, exemplary embodiments of the present disclosure will be considered in more detail.
[0026] FIG. 1 depicts a block diagram of an exemplary view synthesis image generation system 10 according to an exemplary embodiment of the present disclosure. In particular, the view synthesis image generation system 10 can obtain user image data 14 and / or request data 18 from a user 12 (e.g., from a user computing system). The user image data 14 and / or request data 18 can be obtained in response to a time event, one or more user inputs, application downloads and profile settings, and / or determination of a trigger event. The user image data 14 and / or request data 18 can be obtained via one or more interactions with a platform (e.g., a web platform). In some implementations, an application programming interface associated with the platform can obtain and / or generate the user image data 14 and / or request data 18 in response to one or more inputs. The user 12 can be an individual, a retailer, a manufacturer, a service provider, and / or another entity.
[0027] The user image data 14 can be utilized to generate a three-dimensional model of the user object depicted in the user image data (16). Generating the three-dimensional model 16 can include learning a three-dimensional representation of each object by training one or more neural luminance field models with the user image data 14.
[0028] The rendering block 20 can process the request data 18 and, using the generated 3D model, can render one or more view composite images 22 of the object. The request data 18 can describe an explicit user request to generate view composite rendering (e.g., augmented reality rendering) in the user's environment and / or a user request to render one or more objects in combination with one or more additional objects or features. The request data 18 can describe a context and / or parameters (e.g., lighting, size of environmental objects, time, position and orientation of other objects in the environment, and / or other context related to generation) that may affect how the object is rendered. The request data 18 may be generated and / or acquired according to the user's context.
[0029] The view composite image 22 of the object can be provided via a viewfinder, a still image, a catalog user interface, and / or a virtual reality experience. The generated view composite image 22 may be stored locally and / or on the server in relation to the user profile. In some implementations, the view composite image 22 of the object may be stored by the platform via one or more server computing systems related to the platform. Additionally and / or alternatively, the view composite image 22 of the object may be provided and / or interacted with for display via a user interface related to the platform. The user may add the view composite image 22 of the object to one or more collections related to the user, and the one or more collections may be viewed as an aggregate via a collection user interface.
[0030] Figure 2 depicts a block diagram of exemplary virtual object collection generation 200 according to an exemplary embodiment of the present disclosure. In particular, virtual object collection generation 200 can include obtaining user image data 214 (e.g., images from a user-specific image gallery that may be stored in a local and / or server computing system) and / or request data 218 (e.g., manual requests, context-based requests, and / or requests initiated by an application) from user 212 (e.g., from a user computing system that may include a user computing device (e.g., a mobile computing device)). User image data 214 and / or request data 218 can be obtained in response to a time event (e.g., a given interval for updating a virtual object catalog by thoroughly searching within a user-specific image database), one or more user inputs (e.g., one or more inputs to a user interface for learning a three-dimensional representation of a particular object in the environment and / or one or more user inputs for importing or exporting images from user-specific storage), download and profile setup of an application (e.g., a virtual closet application and / or an object modeling application), and / or determination of a trigger event (e.g., the user's location, obtaining a search query, and / or a knowledge trigger event). User 212 can be an individual, a retailer, a manufacturer, a service provider, and / or another entity.
[0031] User image data 214 may include images generated and / or acquired by user 212 (e.g., via an image capture device (e.g., the camera of a mobile computing device)). Alternatively and / or additionally, user image data 214 may include data selected by the user. For example, the data selected by the user may include one or more images and / or image data sets selected by the user via one or more user inputs. The data selected by the user may include images from a web page, image data posted to a social media platform, images within the user's "camera roll", image data stored locally in an image folder, and / or data stored in one or more other databases. The data selected by the user may be selected via one or more user inputs that may include gesture inputs, tap inputs, cursor inputs, text inputs, and / or any other form of input. User image data 214 may be stored locally and / or stored in one or more server computing systems. User image data 214 may be specifically associated with a particular user and / or may be shared data selected (e.g., set, shared among a set of configured groups, and / or shared via a network and / or web page) by the user for generating virtual objects, which can then be stored in a collection and / or provided for display. User image data 214 may include automatically selected image data. The automatic selection may be based on one or more object detections, one or more object classifications, and / or one or more image classifications. For example, multiple image data sets may be processed to determine a subset of image data sets that include image data depicting one or more objects of a particular object type. The subset may be selected for processing.
[0032] The user image data 214 can be used (216) to generate a three-dimensional model of a user object depicted in the user image data 214 (for example, by training the parameters of a neural luminance field model to learn a three-dimensional representation of the color and density values of the user object). Generating the three-dimensional model 216 can include learning a three-dimensional representation of each object by training one or more neural luminance field models with the user image data 214.
[0033] The rendering block 220 (for example, one or more layers for stimulating one or more neural luminance field models and / or one or more application programming interfaces for obtaining and / or utilizing the neural luminance field models) can process the request data 218 and render one or more view synthesis images 222 of the object using the generated three-dimensional model. The request data 218 can describe an explicit user request to generate view synthesis rendering (for example, augmented reality rendering) in the user's environment and / or a user request to render one or more objects in combination with one or more additional objects or features. The request data 218 can describe a context and / or parameters (for example, lighting, the size of environmental objects, time, the position and orientation of other objects in the environment, and / or other context related to generation) that may affect how the object is rendered. The request data 218 may be generated and / or obtained according to the user's context.
[0034] The view synthesis image 222 of the object can be provided via a viewfinder, a still image, a catalog user interface, and / or a virtual reality experience. The generated view synthesis image 222 may be stored locally and / or on the server in relation to the user profile.
[0035] In some implementations, the view composite image 222 may be rendered by the rendering block 220 based on one or more uniform parameters 224. The uniform parameters 224 may include a uniform pose (facing a particular direction (e.g., facing forward)), a uniform position (e.g., an object positioned at the center of the image), a uniform lighting (e.g., no shadows, front lit, natural lit, etc.), and / or a uniform scale (e.g., the object being rendered may be scaled based on a uniform scaling such that the rendering has a uniform one inch to two pixel scale). The uniform parameters 224 may be utilized to provide a coherent rendering of the object that can provide a comparison based on more complete information of the object and / or a comparison of groupings based on more complete information.
[0036] Additionally and / or alternatively, one or more view synthesis images 222 may be added to the catalog 226. For example, one or more view synthesis images 222 can depict one or more clothing objects and may be added to a virtual closet catalog. The virtual closet catalog can include renderings of clothing for multiple users that can be utilized for clothing planning, shopping, and / or comparing clothing. Catalog 226 may be a user-specific catalog, a product database for retailers and / or manufacturers, and / or a group-specific catalog for group sharing. In some implementations, the generated catalog may be processed to determine object proposals for a user and / or a group of users. For example, a user's preferences, style, and / or deficiencies may be determined based on the depictions in the user-specific catalog. Damage to clothing, color palettes, styles, amounts of specific object types, and / or collections of objects may be determined and utilized to determine proposals to offer to the user. The system and method may provide a selection of specific objects based on existing objects with heavy damage. Additionally and / or alternatively, a user's style may be determined and proposals for other objects of that style may be proposed.
[0037] FIG. 3 depicts a flowchart of an exemplary method operating in accordance with an exemplary embodiment of the present disclosure. FIG. 3 depicts steps that are performed in a specific order for purposes of illustration and discussion, but the methods of the present disclosure are not limited to the particularly shown order or arrangement. The various steps of method 300 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0038] In 302, a computing system can obtain user image data and request data. The user image data can depict one or more images including one or more user objects. The one or more images may have been generated on a user computing device. Alternatively and / or additionally, the user image data can include data selected by the user (e.g., one or more images obtained from a web page and / or web platform for use in generating virtual objects). In some implementations, the request data can describe a request to generate a collection specific to an object type. The request data can be associated with a context. In some implementations, the context can describe at least one of an object context or an environment context.
[0039] In 304, a computing system can train one or more neural luminance field models based on the user image data. The one or more neural luminance field models can be trained to generate view synthesis of one or more objects. The one or more neural luminance field models can be configured to predict color values and / or density values associated with an object to generate a view synthesis rendering of the image, and the view synthesis rendering can include a view rendering of the object from a new view not depicted in the user image data.
[0040] In some implementations, the computing system can process the user image data to determine that one or more objects are of a specific object type and store the one or more neural luminance field models in a collection database. The collection database can be associated with a collection specific to an object type. The specific object type can be associated with one or more pieces of clothing.
[0041] At 306, a computing system can generate one or more view synthesis images with one or more neural luminance field models based on request data. The one or more view synthesis images can include one or more renderings of one or more objects.
[0042] In some implementations, the one or more view synthesis images can be generated by processing position and viewing direction with one or more neural luminance field models to generate one or more predicted density values and one or more color values, and generating one or more view synthesis images based on the one or more predicted density values and the one or more color values.
[0043] In some implementations, the request data can describe one or more adjustment settings. Generating one or more view synthesis images with one or more neural luminance field models based on the request data can include adjusting one or more color values of a set of predicted values generated by the one or more neural luminance field models.
[0044] Additionally and / or alternatively, the request data can describe a specific position and a specific viewing direction. Generating one or more view synthesis images with one or more neural luminance field models based on the request data can include processing the specific position and the specific viewing direction with one or more neural luminance field models to generate a view rendering of one or more objects depicting a view associated with the specific position and the specific viewing direction.
[0045] In some implementations, a computing system can provide one or more view synthesis images to a user computing system for display. For example, view synthesis rendering can be provided for display via one or more user interfaces, and may include a grid view, a carousel view, a thumbnail view, and / or a magnified view.
[0046] Additionally and / or alternatively, the computing system can provide a virtual object user interface to the user computing system. The virtual object user interface can provide one or more view synthesis images for display. One or more objects can be separated from the original environment depicted in the user image data.
[0047] In some implementations, a computing system can obtain a plurality of additional user image data sets. Each of the plurality of additional user image data sets may have been generated on a user computing device. The computing system can process each of the plurality of additional user image data sets with one or more object determination models to determine that a subset of the plurality of additional user image data sets includes respective objects of a particular object type. The computing system can train respective additional neural luminance field models for each of the respective additional user image data sets of the subset of the plurality of additional user image data sets and store the respective additional neural luminance field models in a collection database.
[0048] Figure 4 depicts a block diagram of exemplary view synthesis image generation 400 according to an exemplary embodiment of the present disclosure. View synthesis image generation 400 may include obtaining a user-specific image 402 of a user object (e.g., one or more images of an object captured by a user computing device). In some implementations, a particular user interface may be utilized to instruct the user on how to capture the images to be used to train the neural luminance field model. Additionally and / or alternatively, the user interface and / or one or more applications may be utilized to capture a number of images of the object to be used to train one or more neural luminance field models to generate a view rendering of one or more objects. Alternatively and / or additionally, the user-specific image 402 of the user object may be obtained from a user-specific image database (e.g., an image gallery associated with the user).
[0049] The user-specific image 402 of the user object may be utilized to generate a 3D model of the user object (404) (e.g., the user-specific image 402 of the user object may be utilized to train one or more neural luminance field models to learn one or more 3D representations of the object). The generated 3D model may be utilized to generate rendering data 406 of the object (e.g., a trained neural luminance field model and / or parameter data). The rendering data may be stored in relation to object-specific data that may include classification data (e.g., a label of the object), source image data, metadata, and / or user annotations.
[0050] Subsequently, the stored rendering data may be selected based on the context information 410 (408). For example, the rendering data may be selected based on context information 410 that may include the user's search history, the user's search query, budget, other selected objects, the user's location, time, and / or aesthetic values related to the user's current environment (408).
[0051] Then, the selected rendering data can be processed by the rendering block 412 to render one or more view composite images 414 of the selected object. In some implementations, multiple rendering data sets related to multiple different user objects may be obtained to render one or more images of the multiple user objects. Additionally and / or alternatively, the user object may be rendered in a user environment, a template environment, and / or an environment selected by the user. One or more user objects may be rendered adjacent to a proposed object (e.g., an item proposed for purchase).
[0052] FIG. 5 depicts a block diagram of the training and utilization 500 of an exemplary neural luminance field model according to an exemplary embodiment of the present disclosure. In particular, a plurality of images 502 may be obtained. The plurality of images 502 may include a first image, a second image, a third image, a fourth image, a fifth image, and / or an nth image. The plurality of images 502 may be obtained from a user-specific database (e.g., local storage and / or cloud storage associated with the user). In some implementations, the plurality of images 502 may be obtained via a capture device associated with the user computing system.
[0053] A plurality of images 502 can be processed by one or more classification models 504 (as well as / or one or more detection models and / or one or more segmentation models) to determine a subset 506 of images that includes one or more objects related to one or more object types. For example, the classification model 504 may determine that the subset 506 of images includes objects of a particular object type (e.g., a clothing object type, a furniture object type, and / or a particular product type). The subset 506 of images can include the first image, the third image, and the nth image. Different images in the subset 506 of images may depict different objects. In some implementations, images that depict the same object may be determined and used to generate an object-specific dataset for improved training.
[0054] And the subset 506 of images can be used to train a plurality of neural radiance field models 508 (e.g., a first NeRF model related to an object (e.g., the first object) in the first image, a third NeRF model related to an object (e.g., the third object) in the third image, and an nth NeRF model related to an object (e.g., the nth object) in the nth image). Each neural radiance field model can be trained to generate view synthesis renderings of different objects. Different neural radiance field datasets (e.g., neural radiance field models 508 and / or learned parameters) may be stored.
[0055] Then, the user may interact with the user interface 510. Based on the interaction with the user interface, one or more of the neural luminance field datasets may be acquired. One or more selected neural luminance field datasets may be utilized by the rendering block 512 to generate one or more view synthesis images 514 depicting one or more user objects.
[0056] One or more additional user interface interactions that may prompt the rendering of additional view synthesis rendering based on one or more adjustments related to one or more inputs may be received.
[0057] FIG. 6 depicts a diagram of an exemplary virtual closet interface 600 according to an exemplary embodiment of the present disclosure. In particular, a plurality of images 602 may be acquired and / or processed to generate a plurality of rendering datasets 604 associated with the clothing depicted in the plurality of images 602. The plurality of images 602 may be acquired from a user-specific database (e.g., local storage and / or cloud storage associated with the user). In some implementations, one or more of the plurality of images and / or one or more of the plurality of rendering datasets 604 may be acquired from a database associated with one or more other users (e.g., a rendering dataset associated with a retailer and / or manufacturer selling a product (e.g., a dress or a shirt)).
[0058] A plurality of rendering data sets 604 may be selected and / or accessed in response to one or more interactions with the user interface 606. One or more of the rendering data sets may be selected and processed by the rendering block 608 to generate one or more view composite renderings. For example, the system and method may be utilized to assemble clothing to be worn. The user may be presented with and / or select clothing that is rendered in a coherent pose and lighting for review. The generated renderings may include a rendering 610 of a virtual closet segmented from the user, and / or a rendering 612 of a virtual closet rendered on the user, or a virtual closet rendering rendered on a template person. In some implementations, the user may scroll through different view composite renderings of clothing, determine the clothing to visualize on a particular user (e.g., augmented reality try-on and / or a template image of the user or another individual), and render the selected clothing on the user.
[0059] In some implementations, each clothing subtype may include a plurality of view composite renderings associated with different objects of that clothing subtype. A carousel interface may be provided for each clothing subtype, and multiple carousel interfaces may be provided simultaneously to scroll through each subtype individually and / or together, which may enable a coherent set of clothing. The user may then select a try-on rendering user interface element and render the selected clothing on the user and / or a template person.
[0060] A similar user interface may be implemented for interior design, landscaping, and / or game design. The virtual closet interface and / or other similar user interfaces may include one or more proposals determined based on user objects, the user's search history, the user's browsing history, and / or the user's preferences. The rendering dataset for the proposals may be obtained from a server database. The server database may include rendering datasets generated by other users (e.g., retailers, manufacturers, and / or peer-to-peer sellers). The proposals may be based on availability, size, and / or price range for a particular user.
[0061] Figure 7 depicts a flowchart of an exemplary method operating in accordance with an exemplary embodiment of the present disclosure. Figure 7 depicts steps that are performed in a particular order for purposes of explanation and discussion, but the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0062] At 702, a computing system can obtain a plurality of user images. Each of the plurality of user images can include one or more items of clothing. The plurality of user images can be associated with a plurality of different items of clothing. In some implementations, the plurality of user images may be automatically obtained from a storage database associated with a particular user based on the acquired request data. Additionally and / or alternatively, the plurality of user images may be selected from a collection of user images based on metadata, one or more user inputs, and / or one or more classifications.
[0063] In some implementations, a computing system can access a storage database associated with a user and process an aggregation of user images with one or more classification models to determine a plurality of user images that include one or more objects classified as clothing.
[0064] At 704, the computing system can train a respective neural luminance field model for each of a plurality of different clothing items. Each neural luminance field model can be trained to generate one or more view synthesis renderings of a particular respective clothing item.
[0065] At 706, the computing system can store each neural luminance field model in a collection database. The collection database may be associated with an object type (e.g., clothing, furniture, plants, etc.) and / or an object subtype (e.g., pants, shirt, shoes, table, lamp, chair, lily, orchid, bush, etc.). The collection database may be associated with a particular user and / or a particular marketplace. The collection database may be supplemented with products discovered by the user via one or more online marketplaces, social media posts from social media platforms, and / or a proposed rendering dataset associated with a proposed object (or product). The proposal may be based on determined user needs, determined user style, determined user aesthetic values, and / or user context. The proposed products may be associated with known sizes, known availability, and / or known price suitability.
[0066] In 708, a computing system can provide a virtual closet interface. The virtual closet interface can provide multiple clothing view syntheses for display based on each of a plurality of neural luminance field models. The multiple clothing view syntheses can be associated with at least a plurality of different subsets of clothing. In some implementations, the virtual closet interface can include one or more interface features for viewing an array of clothing including two or more pieces of clothing displayed simultaneously. The multiple clothing view syntheses can be generated based on one or more uniform pose parameters and one or more uniform lighting parameters.
[0067] FIG. 8 depicts a flowchart of an exemplary method operating in accordance with an exemplary embodiment of the present disclosure. FIG. 8 depicts steps performed in a particular order for purposes of explanation and discussion, but the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0068] In 802, a computing system can obtain a plurality of user image datasets. Each user image dataset of the plurality of user image datasets can depict one or more images including one or more objects. The one or more images may have been generated by a user computing device (e.g., a mobile computing device having an image capture component). The capture and / or generation of the user image datasets can be facilitated by one or more user interface elements for capturing a number of images of each particular object.
[0069] At 804, a computing system can process a plurality of user image datasets with one or more classification models to determine a subset of the plurality of user image datasets that includes features describing one or more specific object types. In some implementations, the determination may include one or more additional machine learning models (e.g., one or more detection models, one or more segmentation models, and / or one or more feature extractors). The one or more classification models may be trained to classify one or more specific object types (e.g., clothing types, furniture types, etc.). The one or more classified objects may be segmented to generate a plurality of segmented images.
[0070] At 806, a computing system can train a plurality of neural radiance field models based on a subset of the plurality of user image datasets. Each neural radiance field model can be trained to generate a view synthesis of one or more specific objects of each user image dataset in the subset of the plurality of user image datasets. In some implementations, the subset may be processed to generate a plurality of training patches associated with each specific user image dataset in the subset. The patches may be utilized to train the neural radiance field models.
[0071] In some implementations, a computing system can determine a first set of user image datasets that includes features describing a first object subtype. The computing system can associate each first set of neural radiance models with a first object subtype label, can determine a second set of user image datasets that includes features describing a second object subtype, and can associate each second set of neural radiance models with a second object subtype label.
[0072] At 808, a computing system can generate multiple view synthesis renderings with multiple neural luminance field models. The multiple view synthesis renderings can depict multiple different objects (e.g., different clothing) of a particular object type (e.g., clothing object type and / or furniture object type).
[0073] At 810, a computing system can provide a user interface for viewing multiple view synthesis renderings. The user interface can include a rendering pane for viewing the multiple view synthesis renderings. The multiple view synthesis renderings may be provided via a carousel interface, multiple thumbnails, and / or a compiled rendering in which multiple objects are displayed within a single environment.
[0074] In some implementations, a computing system can receive an ensemble rendering request. The ensemble rendering request can describe a request to generate view renderings of a first object of a first object subtype and a second object of a second object subtype. The computing system can generate an ensemble view rendering with a respective first set of first neural luminance field models of the neural luminance field models and a respective second set of second neural luminance field models of the neural luminance field models. The ensemble view rendering can include image data depicting the first object and the second object within a shared environment.
[0075] In some implementations, one or more neural radiance field models can be utilized to generate augmented reality assets and / or virtual reality experiences. For example, one or more neural radiance field models can be utilized to generate multiple view synthesis renderings of one or more objects and / or environments for providing augmented reality experiences and / or virtual reality experiences to a user. The augmented reality experience can be utilized for a user to view an object (e.g., a user object) at different locations and / or positions within the environment where the user is currently located, which can be an environment different from the environment where the current physical objects exist. The virtual reality experience can be utilized to provide a virtual walkthrough experience, which can be utilized for renovation, virtual visits (e.g., virtually visiting a haunted house or an escape room), viewing the interior of an apartment, and / or for social media sharing of an environment that a user can share for friends and / or family to view the environment. Additionally and / or alternatively, the view synthesis rendering can be utilized for video game development and / or other content generation.
[0076] FIG. 9A depicts a block diagram of an exemplary computing system 100 that performs view synthesis image generation, according to an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled via a network 180.
[0077] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0078] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0079] In some implementations, the user computing device 102 can store or include one or more neural luminance field models 120. For example, the neural luminance field model 120 can be various machine learning models such as a neural network (e.g., a deep neural network), or other types of machine learning models including non-linear models and / or linear models. The neural network can include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. An exemplary neural luminance field model 120 is discussed with reference to FIGS. 1-2 and FIGS. 4-6.
[0080] In some implementations, one or more neural luminance field models 120 are received from the server computing system 130 via the network 180, stored in the memory 114 of the user computing device, and then can be used or otherwise implemented by one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single neural luminance field model 120 (e.g., to perform 3D modeling of user objects in parallel across multiple instances of objects in the user image).
[0081] More specifically, the neural luminance field model 120 can be configured to process a three-dimensional position and a two-dimensional line-of-sight direction to determine one or more predicted color values and one or more predicted density values, and generate a view synthesis of one or more objects from the position and the line-of-sight direction. A particular neural luminance field model may be associated with one or more labels. A particular neural luminance field model can be obtained based on an association with a given label and / or a given object. The neural luminance field model 120 can be utilized for synthesizing an image with a plurality of objects, virtually viewing an object, and / or for augmented reality rendering.
[0082] Additionally or alternatively, one or more neural luminance field models 140 can be included in or otherwise stored and implemented in a server computing system 130 that communicates with the user computing device 102 by a client-server relationship. For example, the neural luminance field model 140 can be implemented by the server computing system 130 as part of a web service (e.g., a view synthesis image generation service). Thus, one or more models 120 can be stored and implemented in the user computing device 102, and / or one or more models 140 can be stored and implemented in the server computing system 130.
[0083] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 can be a touch-sensitive component (such as a touch display screen or a touch pad) that can sense the touch of a user input object (such as a finger or a stylus). The touch-sensitive component can act to implement a virtual keyboard. Other exemplary user input components include a microphone, a regular keyboard, or other means by which a user can provide user input.
[0084] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (such as a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 that causes the server computing system 130 to perform operations.
[0085] In some implementations, the server computing system 130 includes one or more server computing devices or is implemented by one or more server computing devices. When the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0086] As described above, the server computing system 130 can store or otherwise include one or more machine learning neural luminance field models 140. For example, the model 140 can be various machine learning models or can otherwise include various machine learning models. Exemplary machine learning models include neural networks or other multi-layer non-linear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Exemplary model 140 is considered with reference to FIGS. 1-2 and 4-6.
[0087] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 through interaction with a training computing system 150 coupled to be communicable via the network 180. The training computing system 150 can be separate from the server computing system 130 or can be part of the server computing system 130.
[0088] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is implemented by one or more server computing devices.
[0089] The training computing system 150 can include a model trainer 160 that trains the machine learning models 120 and / or 140 stored in the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, backpropagation. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions can be used. Gradient descent can be used to iteratively update the parameters over a number of training iterations.
[0090] In some implementations, performing backpropagation may include performing truncated backpropagation through time. The model trainer 160 can perform several generalization techniques (e.g., weight decay, dropout, etc.) to enhance the generalization ability of the model being trained.
[0091] In particular, the model trainer 160 can train the neural luminance field models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, user images, metadata, additional training images, ground truth labels, exemplary training renderings, exemplary feature annotations, exemplary anchors, and / or training video data.
[0092] In some implementations, when the user gives consent, training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 with user-specific data received from the user computing device 102. In some cases, this process can be referred to as personalizing the model.
[0093] The model trainer 160 includes computer logic utilized to provide desired functionality. The model trainer 160 may be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.
[0094] The network 180 can be any type of communication network such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication via the network 180 can be carried over any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0095] The machine learning models described herein may be used in a variety of tasks, applications, and / or use cases.
[0096] In some implementations, the input to the machine learning model of the present disclosure can be image data. The machine learning model can process the image data to generate an output. By way of example, the machine learning model can process the image data to generate an image recognition output (e.g., recognition of image data, latent embedding of image data, encoded representation of image data, hash of image data, etc.). As another example, the machine learning model can process the image data to generate an image segmentation output. As another example, the machine learning model can process the image data to generate an image classification output. As another example, the machine learning model can process the image data to generate an image data modification output (e.g., modification of image data, etc.). As another example, the machine learning model can process the image data to generate an encoded image data output (e.g., encoded and / or compressed representation of image data, etc.). As another example, the machine learning model can process the image data to generate an upscaled image data output. As another example, the machine learning model can process the image data to generate a prediction output.
[0097] In some implementations, the input to the machine learning model of the present disclosure can be text or natural language data. The machine learning model can process the text or natural language data to generate an output. As an example, the machine learning model can process natural language data to generate a language encoding output. As another example, the machine learning model can process text or natural language data to generate a latent text embedding output. As another example, the machine learning model can process text or natural language data to generate a translation output. As another example, the machine learning model can process text or natural language data to generate a classification output. As another example, the machine learning model can process text or natural language data to generate a text segmentation output. As another example, the machine learning model can process text or natural language data to generate a semantic intent output. As another example, the machine learning model can process text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine learning model can process text or natural language data to generate a prediction output.
[0098] In some implementations, the input to the machine learning model of the present disclosure can be latent encoded data (e.g., a latent space representation of the input). The machine learning model can process the latent encoded data to generate an output. As an example, the machine learning model can process the latent encoded data to generate a recognition output. As another example, the machine learning model can process the latent encoded data to generate a reconstruction output. As another example, the machine learning model can process the latent encoded data to generate a search output. As another example, the machine learning model can process the latent encoded data to generate a reclustering output. As another example, the machine learning model can process the latent encoded data to generate a prediction output.
[0099] In some implementations, the input to the machine learning model of the present disclosure can be statistical data. The machine learning model can process the statistical data to generate an output. As an example, the machine learning model can process the statistical data to generate a recognition output. As another example, the machine learning model can process the statistical data to generate a prediction output. As another example, the machine learning model can process the statistical data to generate a classification output. As another example, the machine learning model can process the statistical data to generate a segmentation output. As another example, the machine learning model can process the statistical data to generate a visualization output. As another example, the machine learning model can process the statistical data to generate a diagnostic output.
[0100] In some implementations, the input to the machine learning model of the present disclosure can be sensor data. The machine learning model can process the sensor data to generate an output. As an example, the machine learning model can process the sensor data to generate a recognition output. As another example, the machine learning model can process the sensor data to generate a prediction output. As another example, the machine learning model can process the sensor data to generate a classification output. As another example, the machine learning model can process the sensor data to generate a segmentation output. As another example, the machine learning model can process the sensor data to generate a visualization output. As another example, the machine learning model can process the sensor data to generate a diagnostic output. As another example, the machine learning model can process the sensor data to generate a detection output.
[0101] In some cases, the machine learning model may be configured to perform tasks that include encoding (and / or corresponding decoding) of input data for reliable and / or efficient transmission or storage. For example, the task may be an audio compression task. The input may include audio data, and the output may include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a compression task of visual data. In another example, the task may include generating an embedding for input data (e.g., input audio or visual data).
[0102] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, and the output is a set of scores, where each score corresponds to a different object class and represents the likelihood that one or more images depict an object belonging to the object class. The image processing task can be object detection, and the image processing output identifies one or more regions in one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, and the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, and the image processing output defines, for each pixel in one or more images, a respective depth value. As another example, the image processing task can be motion estimation, the network input includes a plurality of images, and the image processing output defines, for each pixel in one of the input images, the motion of the scene depicted in the pixels between the input images of the network input.
[0103] FIG. 9A shows one exemplary computing system that can be used to implement the present disclosure. Other computing systems can also be used. For example, in some implementations, user computing device 102 can include model trainer 160 and training dataset 162. In such an implementation, model 120 can be trained and used locally on user computing device 102. In some of such implementations, user computing device 102 can implement model trainer 160 to personalize model 120 based on user-specific data.
[0104] FIG. 9B depicts a block diagram of an exemplary computing device 40 operating in accordance with an exemplary embodiment of the present disclosure. Computing device 40 can be a user computing device or a server computing device.
[0105] Computing device 40 includes several applications (e.g., applications 1 through N). Each application includes its own machine learning library and machine learning model. For example, each application can include a machine learning model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.
[0106] As shown in FIG. 9B, each application can communicate with some other components of a computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with the components of its respective device using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0107] FIG. 9C depicts a block diagram of an exemplary computing device 50 operating in accordance with an exemplary embodiment of the present disclosure. The computing device 50 can be a user computing device or a server computing device.
[0108] The computing device 50 includes several applications (e.g., applications 1 through N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API spanning all applications).
[0109] The central intelligence layer includes several machine learning models. For example, as shown in FIG. 9C, each machine learning model (e.g., a model) can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model (e.g., a single model) for all of the applications. In some implementations, the central intelligence layer is included in or otherwise implemented by the operating system of the computing device 50.
[0110] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As shown in FIG. 9C, the central device data layer can communicate with some other components of the computing device, such as, for example, one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with the components of each device using an API (e.g., a private API).
[0111] FIG. 10 depicts a block diagram of a training 1000 of an exemplary neural luminance field model according to an exemplary embodiment of the present disclosure. Training the neural luminance field model 1006 may include processing one or more training data sets. The one or more training data sets may be specific to one or more objects and / or one or more environments. For example, the neural luminance field model 1006 can process a training position 1002 (e.g., a three-dimensional position) and a training viewing direction 1004 (e.g., a two-dimensional viewing direction and / or a vector) to generate one or more predicted color values 1008 and / or one or more predicted density values 1010. The one or more predicted color values 1008 and the one or more predicted density values 1010 can be used to generate a view rendering 1012.
[0112] A training image 1014 related to the training position 1002 and the training viewing direction 1004 can be obtained. The training image 1014 and the view rendering 1012 can be used to evaluate a loss function 1016. Then, the evaluation can be used to adjust one or more parameters of the neural luminance field model 1006. For example, the training image 1014 and the view rendering 1012 can be used to evaluate the loss function 1016 to generate a gradient descent, and the gradient descent can be backpropagated to adjust one or more parameters. The loss function 1016 can include an L2 loss function, a perceptual loss function, a mean squared loss function, a cross-entropy loss function, and / or a hinge loss function.
[0113] In some implementations, the systems and methods disclosed herein can be used to generate and / or render an augmented environment based on user image data, one or more neural luminance field models, a mesh, and / or a proposed data set.
[0114] FIG. 11 depicts a block diagram of an exemplary augmented environment generation system 1100 according to an exemplary embodiment of the present disclosure. The augmented environment generation system 1100. In particular, FIG. 11 depicts an augmented environment generation system 1100 that includes obtaining user data 1102 (e.g., a search query, search parameters, preference data, historical user data, data selected by the user, and / or image data) and outputting to the user an interactive user interface 1110 that depicts a three-dimensional representation of an augmented environment 1108 that includes a plurality of objects 1104 rendered in an environment 1106.
[0115] For example, user data 1102 related to the user may be obtained. The user data 102 may include a search query (e.g., one or more keywords and / or one or more query images), historical data (e.g., the user's search history, the user's browser history, and / or the user's purchase history), preference data (e.g., explicitly entered preferences, learned preferences, and / or weighted adjustments of preferences), data selected by the user, filtering parameters (e.g., price range, location, brand, rating, and / or size), user image data, and / or a generated collection (e.g., a collection generated by the user that may include a shopping cart, a virtual object catalog, and / or a virtual interest board).
[0116] The user data 1102 may be utilized to determine one or more objects 1104. The one or more objects 1104 may be responsive to the user data 1102. For example, the one or more objects 1104 may be associated with search results responsive to a search query and / or one or more filtering parameters. In some implementations, the one or more objects 1104 may be determined by processing the user data 1102 with one or more machine learning models trained to propose objects.
[0117] One or more rendering data sets associated with one or more objects 1104 may be obtained to extend environment 1106 to generate an extended environment 1108 that may be provided in the interactive user interface 1110. The one or more rendering data sets may include one or more meshes for each particular object and one or more neural luminance field data sets (e.g., one or more neural luminance field models having one or more learned parameters associated with the object).
[0118] The extended environment 1108 may be provided as a mesh that is rendered in the environment 1106 during an instance of environment navigation and may be provided by neural luminance field rendering in the environment 1106 during an instance of a threshold time that is obtained while looking at the extended environment 1108 from a particular position and viewing direction.
[0119] Navigation and stalling may occur in response to an interaction with the interactive user interface 1110. The interactive user interface 1110 may include a pop-up element for providing additional information about one or more objects 1104 and / or may be utilized to replace / add / delete an object 1104.
[0120] The environment 1106 may be a template environment and / or may be a user environment generated based on one or more user inputs (e.g., virtual model generation and / or one or more input images).
[0121] FIG. 12 depicts a block diagram of an exemplary virtual environment generation 1200 according to an exemplary embodiment of the present disclosure. In particular, FIG. 12 depicts where user data 1202 is processed to generate a virtual environment 1216 that may be provided for display via an interactive user interface 1218.
[0122] User data 1202 may be obtained from a user computing system. User data 1202 may include a search query, historical data (e.g., search history, browsing history, purchase history, and / or interaction history), preference data, data selected by the user, and / or user profile data. User data 1202 may be processed by a proposal block 1204 to determine one or more objects 1206 associated with the user data 1202. The one or more objects 1206 may be associated with one or more products for purchase. Then, one or more rendering data sets 1210 may be obtained from a rendering asset database 1208 based on the one or more objects 1206. The one or more rendering data sets 1210 may be obtained by querying the rendering asset database 208 with data associated with the one or more objects 1206. In some implementations, the one or more rendering data sets 1210 may be pre-associated with the one or more objects 1206 (e.g., by one or more labels).
[0123] Then, one or more templates 1212 may be obtained. The one or more templates 1212 may be associated with one or more exemplary environments (e.g., an exemplary room, an exemplary lawn, and / or an exemplary vehicle). The one or more templates 1212 may be determined based on user data 1202 and / or based on the one or more objects 1206. The template 1212 may include image data, mesh data, a trained neural radiance field model, a three-dimensional representation, and / or a virtual reality experience.
[0124] One or more templates 1212 and one or more rendering data sets 1210 can be processed by a rendering model 1214 to generate a virtual environment 1216. The rendering model 1214 can include one or more neural radiance field models (e.g., one or more neural radiance field models trained with other user data sets and / or one or more neural radiance field models trained with the user's image data set), one or more augmentation models, and / or one or more mesh models.
[0125] The virtual environment 1216 can depict one or more objects 1206 rendered in a template environment. The virtual environment 1216 can be generated based on one or more templates 1212 and one or more rendering data sets 1210. The virtual environment 1216 can be provided for display on an interactive user interface 1218. In some implementations, a user may be able to interact with the interactive user interface 1218 to view the virtual environment 1216 from different angles and / or at different scalings.
[0126] FIG. 13 depicts a block diagram of exemplary augmented image data generation 1300 according to an exemplary embodiment of the present disclosure. In particular, FIG. 13 depicts where user data 1302 and image data 1312 are processed to generate augmented image data 1316 that can be provided for display via an interactive user interface 1318.
[0127] User data 1302 can be obtained from a user computing system. User data 1302 can include a search query, historical data (e.g., search history, browsing history, purchase history, and / or interaction history), preference data, and / or user profile data. User data 1302 can be processed by a proposal block 1304 to determine one or more objects 1306 related to the user data 1302. The one or more objects 1306 can be associated with one or more products for purchase. Then, one or more rendering data sets 1310 can be obtained from a rendering asset database 1308 based on the one or more objects 1306. The one or more rendering data sets 1310 can be obtained by querying the rendering asset database 1308 using data related to the one or more objects 1306. In some implementations, the one or more rendering data sets 1310 can be pre-associated with the one or more objects 1306 (e.g., by one or more labels).
[0128] Then, image data 1312 can be obtained. The image data 1312 can be associated with one or more user environments (e.g., the user's living room, the user's bedroom, the current environment where the user is located, the user's lawn, and / or a specific vehicle related to the user). The image data 1312 may be obtained in response to one or more selections by the user. The image data 1312 can include one or more images of the environment. In some implementations, the image data 1312 can be utilized to train one or more machine learning models (e.g., one or more neural radiance field models).
[0129] Image data 1312 and one or more rendering data sets 1310 can be processed by a rendering model 1314 to generate extended image data 1316. The rendering model 1314 can include one or more neural luminance field models, one or more extension models, and / or one or more mesh models.
[0130] The extended image data 1316 can depict one or more objects 1306 rendered in a user environment. The extended image data 1316 can be generated based on the image data 1312 and one or more rendering data sets 1310. The extended image data 1316 may be provided for display on an interactive user interface 1318. In some implementations, a user may interact with the interactive user interface 1318 to view one or more various renderings of the extended image data 1316 that depict different angles and / or different scalings for the extended user environment.
[0131] The technologies discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a very wide variety of possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components working in combination. The database and application can be implemented on a single system or distributed across multiple systems. The distributed components can operate sequentially or in parallel.
[0132] Although the subject matter has been described in detail in connection with various specific exemplary embodiments, each example is provided for illustrative purposes rather than limitation of the disclosure. Those skilled in the art will readily appreciate that modifications to such embodiments, changes to such embodiments, and equivalents of such embodiments can be readily made. Accordingly, the disclosure of the subject matter is not intended to exclude such modifications, changes, and / or additions to the subject matter as would be readily apparent to those skilled in the art. For example, features shown or described as part of one embodiment can be used by another embodiment to generate further embodiments. Accordingly, the disclosure is intended to embrace such modifications, changes, and equivalents.
Description of Reference Numerals
[0133] 10 View synthesis image generation system 12 User 14 User image data 16 Generating a 3D model 18 Request data 20 Rendering block 22 View synthesis image 40 Computing device 50 Computing device 100 Computing system 102 User computing device 112 Processor 114 Memory 116 Data 102 User computing device 120 Neural luminance field model 122 User input component 130 Server computing system 132 Processor 134 Memory 136 Data 138 Instruction 140 Neural luminance field model 150 Training computing system 152 Processor 154 Memory 156 Data 158 Instruction 160 Model Trainer 162 Training Data 180 Network 200 Virtual Object Collection Generation 212 User 214 User Image Data 216 Generating a 3D Model 218 Request Data 220 Rendering Block 222 View Composite Image 224 Uniform Parameter 226 Catalog 300 Method 400 View Composite Image Generation 402 User-Specific Image of User Object 406 Rendering Data of Object 410 Context Information 412 Rendering Block 414 One or More View Composite Images of Selected Object 500 Training and Utilization of Neural Radiance Field Model 502 Image 504 Classification Model 506 Subset of Images 508 Neural Radiance Field Model 510 User Interface 512 Rendering Block 514 View Composite Image 600 Virtual Closet Interface 602 Image 604 Rendering Dataset 606 User Interface 608 Rendering Block 610 Rendering of User-Segmented Virtual Closet from User Rendering of a virtual closet rendered on a user 700 Method 800 Method 1000 Training of a neural luminance field model 1002 Training position 1004 Training line-of-sight direction 1006 Neural luminance field model 1008 Predicted color value 1010 Predicted density value 1012 View rendering 1014 Training image 1016 Loss function 1100 Extended environment generation system 1102 User data 1104 Object 1106 Environment 1108 Extended environment 1110 Interactive user interface 1200 Virtual environment generation 1202 User data 1204 Proposal block 1206 Object 1208 Rendering asset database 1210 Rendering dataset 1212 Template 1214 Rendering model 1216 Virtual environment 1218 Interactive user interface 1300 Extended image data generation 1302 User data 1304 Proposal block 1306 Object 1308 Rendering asset database 1310 Rendering dataset 1312 Image data 1314 Rendering model 1316 Extended image data 1318 Interactive User Interface
Claims
1. A computing system comprising: one or more processors; and one or more computer-readable storage media that, when executed by the one or more processors, cause the computing system to: obtain a plurality of user image datasets and request data, wherein each user image dataset of the plurality of user image datasets depicts one or more images including one or more user objects, and the one or more images are generated by a user computing device; process the plurality of user image datasets with one or more classification models to determine a subset of the plurality of user image datasets that includes features describing one or more specific object types; train a plurality of neural radiance field models based on the subset of the plurality of user image datasets, wherein each of the plurality of neural radiance field models is trained to generate a view synthesis of a specific object of each user image dataset of the subset of the plurality of user image datasets; generate a plurality of view synthesis images with the plurality of neural radiance field models based on the request data, wherein the plurality of view synthesis images include one or more renderings of a plurality of different objects of the one or more specific object types; and jointly store instructions for causing the computing system to perform the operations in one or more computer-readable storage media. A computing system comprising the above.
2. The request data describes a request to generate a collection specific to an object type, and the operations further comprise: processing the user image datasets to determine that the one or more objects are of a specific object type; and storing the plurality of neural radiance field models in a collection database, wherein the collection database is associated with the collection specific to the object type. The system according to claim 1.
3. The operations further comprise: obtaining a plurality of additional user image datasets, wherein each of the plurality of additional user image datasets is generated by the user computing device. Processing each of the plurality of additional user image datasets with one or more object determination models to determine that a subset of the plurality of additional user image datasets includes respective objects of the specific object type; Training respective additional neural luminance field models for respective additional user image datasets of the subset of the plurality of additional user image datasets; Storing each of the respective additional neural luminance field models in the collection database; The system according to claim 2, further comprising.
4. The system according to claim 2, wherein the specific object type is associated with one or more items of clothing.
5. The operation is The system according to claim 1, further comprising providing the one or more view synthesis images to a user computing system for display.
6. The system according to claim 1, wherein the request data is associated with a context, and the context describes at least one of an object context or an environmental context.
7. The operation is Providing a virtual object user interface to a user computing system, the virtual object user interface providing the one or more view synthesis images for display, and the one or more objects being separated from an original environment depicted in the one or more images of the user image dataset. The system according to claim 1, further comprising.
8. The one or more view synthesis images are Processing a position and a viewing direction with the plurality of neural luminance field models to generate one or more predicted density values and one or more color values; and Generating the one or more view synthesis images based on the one or more predicted density values and the one or more color values. The system according to claim 1, generated by.
9. The request data describes one or more adjustment settings, Generating the one or more view synthesis images with the plurality of neural luminance field models based on the request data includes adjusting one or more color values of a set of predicted values generated by the plurality of neural luminance field models. The system according to claim 1.
10. wherein the requested data describes a specific position and a specific line-of-sight direction, and generating the one or more view synthesis images with the plurality of neural luminance field models based on the requested data includes processing the specific position and the specific line-of-sight direction with the plurality of neural luminance field models to generate a view rendering of the one or more objects depicting a view related to the specific position and the specific line-of-sight direction. The system according to claim 1.
11. A method implemented by a computer for generating a virtual closet, comprising: acquiring, by a computing system including one or more processors, a plurality of user images, each of the plurality of user images including one or more items of clothing, the plurality of user images being associated with a plurality of different items of clothing; training, by the computing system, a plurality of neural luminance field models for the plurality of different items of clothing based on the plurality of user images, each of the plurality of neural luminance field models being trained to generate one or more view synthesis renderings of a respective specific item of clothing; storing, by the computing system, each neural luminance field model in a collection database; providing, by the computing system, a virtual closet interface, the virtual closet interface providing a plurality of clothing view synthesis renderings for display based on the plurality of neural luminance field models, the plurality of clothing view synthesis renderings being associated with at least a subset of the plurality of different items of clothing; and a method.
12. The method according to claim 11, wherein the plurality of user images are automatically acquired from a storage database associated with a specific user based on acquired requested data.
13. The method according to claim 11, wherein the plurality of user images are selected from an aggregation of user images based on at least one of metadata, one or more user inputs, or one or more classifications.
14. Accessing, by the computing system, a storage database associated with a user; Processing, by the computing system, an aggregation of user images with one or more classification models to determine the plurality of user images including one or more objects classified as clothing; The method according to claim 11, further comprising.
15. The method according to claim 11, wherein the virtual closet interface includes features of one or more interfaces for viewing a set of clothing including two or more pieces of clothing displayed simultaneously.
16. The method according to claim 11, wherein the plurality of clothing view synthesis renderings are generated based on one or more uniform pose parameters and one or more uniform lighting parameters.
17. One or more computer-readable storage media that, when executed by one or more computing devices, cause the one or more computing devices to Obtain a plurality of user image datasets, wherein each user image dataset of the plurality of user image datasets depicts one or more images including one or more objects, and the one or more images are generated by a user computing device; Process the plurality of user image datasets with one or more classification models to determine a subset of the plurality of user image datasets including features describing one or more specific object types; Train a plurality of neural luminance field models based on the subset of the plurality of user image datasets, wherein each neural luminance field model is trained to generate a view synthesis of a specific object of each user image dataset of the subset of the plurality of user image datasets; Generate a plurality of view synthesis renderings with the plurality of neural luminance field models, wherein the plurality of view synthesis renderings depict a plurality of different objects of the specific object type; Provide a user interface for viewing the plurality of view synthesis renderings One or more computer-readable storage media jointly storing instructions for performing operations including.
18. One or more computer-readable storage media according to claim 17, wherein the user interface includes a rendering pane for viewing the plurality of view composite renderings.
19. The operation includes determining a first set of user image data sets including features describing a first object subtype; associating each first set of neural luminance field models with a first object subtype label; determining a second set of user image data sets including features describing a second object subtype; associating each second set of neural luminance field models with a second object subtype label One or more computer-readable storage media according to claim 17, further comprising.
20. The operation includes receiving an ensemble rendering request, the ensemble rendering request describing a request to generate view renderings of a first object of the first object subtype and a second object of the second object subtype; generating an ensemble view rendering with a first neural luminance field model of each first set of neural luminance field models and a second neural luminance field model of each second set of neural luminance field models, the ensemble view rendering including image data depicting the first object and the second object in a shared environment; One or more computer-readable storage media according to claim 19, further comprising.
Citation Information
Patent Citations
Virtual clothing output system
JP2018041459A
SYSTEM AND METHOD FOR SYNTHESIS OF APPAREL ENSEMBLES FOR MODELS - Patent application
JP2022534082A
Rendering new images of scenes using geometry-aware neural networks conditioned on latent variables
WO2022167602A2