Generative scene modeling
By using machine learning models to generate custom background images in real-time based on 2D content analysis, the method addresses the impracticality of static libraries, achieving immersive and scene-specific enhancements for 2D media viewing.
Patent Information
- Application Number
- PCT/US2024/016078
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-16
- Publication Date
- 2025-08-21
AI Technical Summary
Existing 2D media viewing experiences lack immersion and customization, as traditional libraries of background images are impractical and inefficient for dynamic, scene-by-scene adaptation to diverse content.
Employing a combination of machine learning models to generate custom background images in real-time, analyzing 2D content frames for semantic and aesthetic characteristics, and generating immersive backgrounds that match the content's themes and moods.
Provides highly immersive and customized background imagery that enhances the 2D content viewing experience, dynamically adapting to each scene, overcoming the limitations of static image libraries.
Smart Images

Figure US2024016078_21082025_PF_FP_ABST
Abstract
Description
GENERATIVE SCENE MODELINGBACKGROUND
[0001] Mixed reality is an umbrella term referring to various technologies that serve to augment, virtualize, or otherwise extend a user’s experience of reality in a variety of ways. For example, virtual reality, augmented reality, and other similar technologies refer to different types of mixed reality that have been developed and deployed for use with entertainment, educational, vocational, and other types of applications. In certain cases, mixed reality experiences may be presented on handheld devices such as smartphones or laptop computers viewed from a few feet away. In other cases, however, mixed reality experiences may be presented by way of head -mounted display devices that free up the users’ hands and more fully immerse the users in the mixed reality worlds.SUMMARY
[0002] Methods and systems for generative scene modeling are described herein. In particular, as will be described in detail below, generative content tools including various machine learning models may be configured to work in tandem with one another to generate and continuously update a custom background image that may be presented together with content (e.g., including 2D content such as a conventional movie, television show, video game, etc.) in a highly immersive mixed reality experience. For example, generative scene modeling could produce an outer space environment (e.g., where the user is surrounded by stars, planets, etc.) during one scene of a science fiction movie, an environment of an alien world during another scene of the movie, and other immersive environments during other parts of the movie. Each of these virtual backgrounds may be dynamically customized to match various semantic and aesthetic aspects of the 2D content, including themes, moods, color palettes, lighting, and so forth.
[0003] In one implementation, a method comprises steps including: 1) generating, using a first machine learning model (e.g., a visual -language model), a textual description of an image (e.g., a frame of a 2D content instance such as a movie, television show, video game, etc.); 2) converting, using a second machine learning model (e.g., a large-language model), the textual description of the image into a text prompt for a custom background image; 3) generating, using a third machine learning model (e.g., an image-generation model), the custom background image based on the text prompt; and 4) presenting, on adisplay of a display device, the image and the custom background image.
[0004] In another implementation, a non-transitory computer-readable medium stores instructions that, when executed, cause a processor to perform a process comprising: 1) accessing a video content instance (e.g., a 2D movie, television show, video game, etc.) that includes a sequence of 2D frames; 2) generating, using a first machine learning model (e.g., a visual -language model), a textual description of a designated frame included within the sequence of 2D frames; 3) converting, using a second machine learning model (e.g., a large- language model), the textual description of the designated frame into a text prompt for a custom background image; 4) generating, using a third machine learning model (e.g., an image-generation model), the custom background image based on the text prompt; and 5) presenting, on a display of a display device, the designated frame and the custom background image.
[0005] In yet another implementation, a generative scene modeling system comprises: 1) a visual -language model configured to generate a textual description of an image (e.g., a designated frame of a 2D content instance such as a 2D movie, television show, video game, etc.); 2) a large-language model configured to convert the textual description of the image into a text prompt for a custom background image; 3) an image-generation model configured to generate the custom background image based on the text prompt; and 4) a display device including a display configured to present the image and the custom background image.
[0006] The details of these and other implementations are set forth in the accompanying drawings and the description below. Other features will also be made apparent from the following description, drawings, and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. 1 shows an illustrative generative scene modeling system configured to facilitate immersive viewing of content in accordance with principles described herein.
[0008] FIG. 2 shows an illustrative method for generative scene modeling to facilitate immersive viewing of content in accordance with principles described herein.
[0009] FIG. 3 shows an illustrative video content instance and a sequence of 2D frames included in the video content instance in accordance with principles described herein.
[0010] FIG. 4 shows an illustrative image alongside a custom background mask configured for use in generating a custom background image to be presented with the image in accordance with principles described herein.
[0011] FIG. 5 shows illustrative aspects of how a text prompt for a custombackground image may be generated based on an image and using a variety of machine learning models in accordance with principles described herein.
[0012] FIG. 6 shows illustrative aspects of how an outpainting technique may be used in the generation of a custom background image in accordance with principles described herein.
[0013] FIG. 7 shows illustrative aspects of how an inpainting technique may be used in the generation of a custom background image in accordance with principles described herein.
[0014] FIGS. 8 A and 8B show illustrative aspects of how elements of a custom background image may be modified using example optical transport interpolations in accordance with principles described herein.
[0015] FIG. 9 shows an illustrative custom background image that may be generated based on an illustrative image for presentation with the image on a display device in accordance with principles described herein.
[0016] FIG. 10 shows an illustrative computing system that may be used to implement various devices and / or systems described herein.DETAILED DESCRIPTION
[0017] Methods and systems for generative scene modeling are described herein. In particular, methods and systems described herein may generatively model scenes in which content can be presented, such as by producing a custom background over which 2D content may be viewed.
[0018] Despite technological advances that have led to new forms of media such as 3D media, 360° virtual reality media, and so forth, traditional 2D media formats (e.g., including conventional movies, television shows, video games, etc.) remain highly popular. For one thing, such 2D media is significantly easier and cheaper to produce than newer forms of media, leading to an enormous amount of 2D content being available for a variety of interests and at a variety of price points (including user-generated content created using ubiquitous devices such as smartphones). Additionally, 2D media tends to be significantly more convenient to watch (or otherwise experience), as it may easily be enjoyed on a variety of screen types and sizes from a variety of locations and in a variety of circumstances.
[0019] The fact that traditional 2D content is still commonly consumed using legacy viewing technologies does not, however, mean that the viewing experience of such content cannot be improved and even transformed by newer, more sophisticated technologies. Indeed,while a highly immersive 2D content viewing experience might have once required entry into a dark theater with state-of-the-art visual and audio systems, mixed reality devices employing head -mounted displays now allow for a personal viewing theater (e.g., complete with a dark, private viewing environment, surround sound, etc.) to be simulated right on the head of a user, wherever that user may be. Moreover, because the entire 360° visual field of the user is under the control of a head-mounted display device, 2D content may be presented in ways that are even more immersive and advanced than would be practical (or even possible) in traditional theaters. In other examples, display devices may be used that are not headmounted or do not encompass an entire 360° visual field, but that nonetheless involve presenting content (e.g., 2D content) in connection with a background that may be customized by way of principles of generative scene modeling described herein.
[0020] As used herein, generative scene modeling refers to mixed reality scenes that may be generatively modeled based on certain aspects (e.g., aesthetic characteristics, semantic characteristics, etc.) of 2D images that are presented within the scenes being created. 2D images such as frames of a video content instance (e.g., an instance of conventional 2D media such as a movie, a television show, a video game, etc.) may be analyzed by a variety of machine learning models (e.g., visual -language models, visualquestion-answering models, large-language models, image-generation models, other artificial intelligence networks, etc.). Based on this analysis, systems described herein may then generate highly immersive and customized background imagery to augment the viewing experience of the 2D content. This background imagery may be generated in real time and kept updated as the 2D content changes (e.g., from frame to frame, from shot to shot, from scene to scene, etc.). Referring to an example mentioned above, for instance, a user watching a science fiction movie could be virtually placed in outer space (e.g., surrounded by stars, planets, etc.) for one scene of the movie, virtually placed on an alien world for another scene of the movie, and so forth. As will be described in detail below, each generatively modeled background scene produced in this way may be customized to match various semantic and aesthetic aspects of the 2D content (e.g., themes, moods, color schemes, etc.). In this way, a fully customized, highly immersive, and enjoyable aspect may be added to the conventional 2D content viewing experience.
[0021] Certain benefits that have been described could be achieved, to some extent, from selecting a background image from a predefined library of such images. For instance, a generic background image could be completely black or could replicate a movie theater, a high-end home theater, or another such location from which 2D content has been traditionallyenjoyed. The library could even include predesignated background images for certain genres (e.g., an outer space background for the science fiction genre, etc.) or even for specific video content instances (e.g., an ocean background for the movie Titanic, etc.). It would be extremely impractical and inefficient, if not impossible, however, to use this library -based approach to present truly customized background images on a show-by-show and even scene- by-scene basis. Each content instance, and indeed each scene of a given content instance, may not only have its own unique characters, sets, props, and themes, but may further have its own unique aesthetic (e.g., color palette, lighting style, etc.), its own unique mood, and so forth. To create and organize a library of customized backgrounds for each scene of even a few popular content instances (e.g., classic movies, popular television shows, etc.) would present a significant technical challenge. To attempt to do this for any 2D content that any user might be interested in would be virtually impossible.
[0022] Methods and systems for generative scene modeling described herein, however, present a technical solution to this technical problem. Specifically, rather than trying to achieve the immersive benefits described above by a priori construction of an impractically immense library of background artwork, methods and systems described herein employ a combination of emerging machine learning models to generate custom background images dynamically and in real time for any 2D content. For example, as will be described and illustrated in more detail below, certain machine learning models (also referred to as artificial intelligence networks, tools, etc.) may be used to analyze images such as frames of a video, other machine learning models may be used to generate a text prompt for a desirable background based on that analysis, and still other machine learning models may generate the background imagery so that it can be presented by the display device. In this way, the technical challenge of having an unwieldy library of background artwork for only certain supported content is mitigated, while customized, bespoke artwork for any 2D content is generated and updated dynamically as the user experiences the content.
[0023] Various implementations will now be described in more detail with reference to the figures. It will be understood that the particular implementations described below are provided as non-limiting examples and may be applied in various situations. Additionally, it will be understood that other implementations not explicitly described herein may also fall within the scope of the claims set forth below. Systems and methods described herein for generative scene modeling for immersive viewing of 2D content may result in any or all of the technical and non-technical benefits mentioned above, as well as various additional benefits that will be described and / or made apparent below.
[0024] FIG. 1 shows an illustrative generative scene modeling system 100 configured to facilitate immersive viewing of 2D content in accordance with principles described herein. As shown, the generative scene modeling system implementation represented by system 100 includes a first machine learning model implemented as a visual -language model 102 that may be configured to generate a textual description 104 of an image 106 that is received as input by system 100. System 100 further includes a second machine learning model implemented as a large-language model 108 that may be configured to convert textual description 104 of image 106 into a text prompt 110 for a custom background image that is to be customized to the image. To this end, system 100 is shown to further include a third machine learning model implemented as an image-generation model 112 that may be configured to generate this custom background image 114 based on text prompt 110. System 100 also includes or has control of a display device 116 that includes a display 118 configured to present image 106 and custom background image 114.
[0025] Each of the machine learning models included in system 100 (i.e., visuallanguage model 102, large-language model 108, and image-generation model 112, as well as other models that may be included in other implementations) will be understood to represent models or networks that have been trained to perform particular tasks described herein. While each machine learning model may be described or referred to herein using terminology indicative of one type of machine learning model that may be used for the stated functions, it will be understood that the machine learning models may be implemented in a variety of ways to perform the intended functionality described herein. In some examples, multiple machine learning models described for a particular implementation may be combined. For instance, rather than a separate visual -language model and large-language model as shown in FIG. 1, a hybrid machine learning model could implement the functionality described for both visual-language model 102 and large-language model 108.
[0026] These models may be constructed using any suitable machine learning or artificial intelligence technologies and may be implemented using any hardware and / or software as may serve a particular implementation. For example, in some implementations, one or more of these models may be implemented using processing units integrated within a user equipment device (e.g., display device 116 or an associated portable device connected to display device 116). In these cases, the software associated with the models may be stored and executed directly on the user equipment device. In the same or other implementations, one or more of the models may be implemented using hardware integrated with servers or other computing systems (e.g., including non-portable computing systems) that arecommunicatively coupled to the display device 116 (e.g., by way of one or more communication networks) so as to provide processing support for the display device 116.
[0027] The framework provided by models 102, 108, and 112 (as well as other models that may be included, as will be described in more detail below) may enable system 100 to perform generative scene modeling by extending and / or supplementing image 106 in a customized, semantically-relevant way to make the mixed reality experience of engaging the 2D content highly immersive. As will be described in more detail below, this framework, in which textual description 104 may be generated and used to form text prompt 110 for creation of custom background image 114, is one way that system 100 may perform generative scene modeling. Additionally, as will be described in more detail below, image-to- photosphere outpainting techniques (e.g., using diffusion masks), inpainting techniques, and other such processes may be used to produce generative scenes (e.g., custom background image 114). In some implementations, these techniques may be used in connection with the types of models and data shown in the implementation of FIG. 1, while in other implementations these techniques could be used to create custom background images without further reliance on the types of models and / or data shown in this particular implementation.
[0028] The presentation by display device 116 on display 118 may be performed in any suitable way. For example, in some implementations, display device 116 may be implemented as a head-mounted display device, such as a mixed reality headset that will be understood to integrate one or more display screens positioned in front of the user’s eyes (implementing display 118), a strap to mount the unit on the user’s head, speakers for audio reproduction, processing units, pose detection sensors, and so forth. As shown, display device 116 may provide a presentation 120 (e.g., on display 118) in which image 106 is positioned at a central location of the display and surrounded by custom background image 114. As will be described in more detail below, image 106 and custom background image 114 may collectively form a 360° generative scene that fully immerses the user as the user experiences the 2D content of image 106. For example, as mentioned above, image 106 could be a frame of a science fiction movie while custom background image 114 could depict settings, characters, or other aesthetically or semantically relevant content associated with the movie (e.g., an outer space setting, etc.) to help immerse the user in the universe of the movie as the user watches it. To further illustrate this concept, FIG. 1 shows a field of view 122 of the user at a particular moment in time. As shown, view 122 may be sufficiently large to incorporate an entirety of image 106 while not being wide enough to include all of custom background image 114. Some of custom background image 114 may be visible (e.g., in the user’speripheral vision, etc.) while watching the 2D content of image 106, but the user would have to look up or turn around to see other parts of custom background image 114 in examples where it provides a full 360° (e.g., spherical) generative scene around the user.
[0029] While FIG. 1 illustrates a general example of a head-mounted display device that may be used to perform operations described herein (i.e., display device 116), it will be understood that a variety of user equipment devices could serve this role in different implementations. For example, display device 116 may be implemented using any type or form factor of extended reality or mixed reality headset (e.g., virtual reality goggles, augmented reality glasses, etc.) or other device (e.g., a smartphone, a projection device, etc.) that is configured to provide an immersive (e.g., 360°) experience for a user engaging with 2D content. In some examples, image 106 may be received as image data to be rendered on display 118 within display device 116, such that no real -world version of the 2D content is present in the environment (e.g., no real-world display screen is presenting the 2D content). In other examples, however, a real-world screen such as a television or movie screen could present image 106 such that the image is received by a front facing camera on display device 116 and is presented to the user by being passed through a transparent display (e.g., augmented reality glasses) or a non-transparent display (e.g., mixed reality goggles with image passthrough). In these latter examples, a user wearing a head-mounted display device could enjoy the added immersion of the generative scene modeling described herein while watching 2D content with other people who may not be wearing headsets or experiencing the added layer of immersion provided by the headset.
[0030] FIG. 2 shows an illustrative method 200 for generative scene modeling to facilitate immersive viewing of 2D content in accordance with principles described herein. While FIG. 2 shows illustrative operations 202-208 according to one implementation, other implementations of method 200 may omit, add to, reorder, and / or modify any of the operations 202-208 shown in FIG. 2. In some examples, multiple operations shown in FIG. 2 or described in relation to FIG. 2 may be performed concurrently (e.g., in parallel) with one another, rather than being performed sequentially as illustrated and / or described. Each of operations 202-208 of method 200 will now be described in more detail as the operations may be performed by an implementation of generative scene modeling system 100.
[0031] At operation 202, system 100 may generate a textual description of an image using a first machine learning model (e.g., a visual -language model or another suitable model configured to generate the textual description of the image in ways described herein). For example, the textual description may indicate certain semantic characteristics of the image,such as a 2D content instance (e.g., movie, television show, video game, etc.) that the image is part of, what characters or objects are shown in the image, themes or settings associated with what is depicted in the image, or the like. As another example, the textual description may indicate certain aesthetic characteristics of the image, such as a lighting characteristic of a scene depicted in the image (e.g., indoor versus outdoor lighting, brighter versus darker lighting, etc.), a color palette used by the image (e.g., bright and loud shades versus more subdued shades and grays, etc.), or the like. The textual description generated at operation 202 may not conform to any particular format or standard of grammar or language usage. Rather, for example, the textual description generated at this stage may be a list of relatively disparate observations about the image in a relatively raw form that has not been synthesized into a particular coherent description of the image. As will be described in more detail below, other models such as a visual-question-answering model may assist the first machine learning model (e.g., the visual -language model) at this stage to generate a textual description having the characteristics that have been described. For instance, the first machine learning model may be configured to caption the image based only on what is visible in the image itself, while the visual-question-answering model may be used to gather more context about relative semantic and aesthetic information. The image can be an image captured by a camera associated with the system, or can be an image provided to the system by another entity. The image can be frame of a 2D content instance such as a movie, television show, video game, or the like.
[0032] At operation 204, system 100 may use a second machine learning model (e.g., a large-language model or another suitable model) to convert the textual description of the image generated at operation 202 into a text prompt for a custom background image that is to be customized to the image. In contrast to the relatively disjointed textual description that may be generated at operation 202, the text prompt produced at this stage may be synthesized to relay the various characteristics that have been observed and represented in the textual description in a syntactically correct, natural prompt with high degree of coherency for the custom background image that is desired. Examples of textual description and text prompts are described in more detail below.
[0033] At operation 206, system 100 may use a third machine learning model (e.g., an image-generation model such as an image-diffusion model or another suitable model) to generate the custom background image based on the text prompt produced at operation 204. For example, the text prompt may indicate certain semantic and aesthetic properties that are desirable for the custom background image to have and may indicate these properties in amanner that the third machine learning model is configured to effectively process. In some examples, other information may also be provided to the third machine learning model such as a mask that is to be filled in to create or complete the custom background image (e.g., using outpainting and / or inpainting techniques described herein), information about features of the image that might be made to extend from the sides of the image into the background, and so forth. As a result, the custom background image generated at this stage may be highly customized and adapted specifically for presentation with the particular image (e.g., the particular frame currently being presented), in contrast to conventional backgrounds that could be selected (e.g., from a library) to roughly correspond to a certain type of content, but that would not be generatively customized and adapted to the content to the same extent.
[0034] At operation 208, system 100 may present the image and the custom background image on a display of a display device (e.g., a head-mounted display device, etc.). For example, as described and illustrated in relation to FIG. 1, the image may be presented at a central location where the user is likely to direct most of their attention, while the custom background image may surround the image to provide ambiance, mood, character, atmosphere, and / or immersiveness to the overall experience of engaging with the 2D content of which the image is a part.
[0035] As will be described in more detail below, some implementations may include a non-transitory computer-readable medium storing instructions that, when executed, cause a processor to perform a process comprising method 200. Additionally, it will be understood that other operations not explicitly shown or described in relation to method 200 may be included in other methods performed by systems described herein. For instance, prior to the generating of operation 202, system 100 may access a video content instance including a sequence of 2D frames and may select, identify, or otherwise designate a frame from that sequence such that the image being used in the later operations is this designated frame included within the sequence of 2D frames. These and other features of potential implementations will now be described and illustrated in more detail.
[0036] FIG. 3 shows an illustrative video content instance 302 and a sequence 304 of 2D frames included in the video content instance in accordance with principles described herein. More particularly, sequence 304 of 2D frames is shown along a timeline 306 that is broken into several sections that wrap around in accordance with the dimensions of the figure (as illustrated by the dashed line connecting the end of one section of the timeline to the beginning of the next section below it). All of the 2D frames in sequence 304 will be referred to as frames 308 and most of them are labeled in this way. However, as shown, certain of the2D frames 308 (arbitrarily drawn at the beginning of each section of timeline 306 in this example) are drawn to be larger than the others, as well as to have a different style of border. These 2D frames 308 are labeled as designated frames 310 and will be understood to represent frames that have been identified, selected, or otherwise set apart from the other frames 308 for various uses including in the generation of custom background images in accordance with principles described herein. In some examples, designated frames may be key frames that are pre-marked in the video content or that may be identified using a key frame identification analysis that analyzes intrinsic properties of the frames in the sequence (e.g., to identify which frames meet certain criteria of key frames). In other examples, frames may be designated more arbitrarily, such as with minimal or no frame analysis. For instance, every 100th frame (or some other arbitrary number) could be designated as a designated frame in a particular implementation. Details about certain ways that designations may be made are described in more detail below. However frame designation may be performed in a given implementation, a designated frame may be treated differently than a non-designated frame.
[0037] Video content instance 302 may represent any suitable type of video content. For instance, video content instance 302 may be implemented by a movie, an episode of a television show, a video clip (e.g., a video captured by a user, a video provided by a social networking or video application such as YouTube, etc.), an instance of interactive video content such as a video game, or any other suitable instance of video content (e.g., 2D video content) that may be presented to a user by way of a mixed reality presentation device. Whatever its type, video content instance 302 may include a sequence of frames (e.g., sequence 304) in which most frames are closely related to neighboring frames in the sequence. In certain situations, however, such as when a new shot or new scene begins in the video content, that general rule does not hold, and one frame may be significantly different from the frame that immediately preceded it in time. Frames that are representative of a group of other frames near them (e.g., for a particular portion of the video such as one shot or one scene) are referred to herein as key frames and may serve as one example of designated frames 310 that may be used in a particular implementation.
[0038] Since it may not be practical (or even desirable) for a unique custom background image to be generated from scratch for every single frame 308 of a sequence 304, designated frames 310 may serve as aesthetic and semantic signposts for custom background images to be generated so that the background presented with the video content remains relevant as a video progresses. To this end, as mentioned above, images 106 being processed,manipulated, presented, and otherwise used by system 100 and method 200 in the ways described above may be designated frames from a 2D frame sequence such as designated frames 310 of sequence 304. As such, system 100 may be configured to access a video content instance 302 that includes a sequence 304 of 2D frames and to designate (e.g., based on a rule or heuristic, based on a key frame identification analysis, based on metadata or inherent characteristics of the frames, etc.) one or more designated frames 310 within the sequence.
[0039] The designation of frames 308 as designated frames 310 may be performed using any suitable rule, heuristic, analysis, metadata attached to the frames, inherent characteristics of the frame, or other basis that may serve to help select, identify, or otherwise distinguish designated frames 310 from other frames 308. For example, one analysis technique may involve generating respective color histograms for each frame 308 and setting one or more thresholds to determine if the overall color palette from a current frame 308 is sufficiently different from another frame (e.g., from a frame previous to the current frame 308, from the most recent designated frames 310, etc.) to merit designation as a designated frame 310. Other analysis techniques may involve using optical flow to analyze movement from frame to frame to help identify when a new designated frame is to be identified, using dead time analysis to determine that a new shot or scene is likely to have begun with a new set of frames for which a designated frame 310 is to be selected, or other such techniques. In some implementations, a fusion of different frame designation analyses and / or techniques may be employed in a combination frame designation analysis. Additionally, confidence ranking may be utilized with any of these individual or combination analysis techniques to help ensure the quality of the designated frame identification.
[0040] As the identification (i.e., selection, designation, etc.) of designated frames 310 from the other 2D frames 308 is performed, it may be undesirable to select two frames 308 as designated frames 310 if they are close to one another in sequence 304 (i.e., if there is not a certain number of other, non-designated frames 308 between them). As such, the designation of frames performed in certain implementations may use hysteresis (or other similar rules or heuristics) to prevent a designated frame 310 from being identified within a threshold number of frames 308 after a previous designated frame within the sequence 304 of 2D frames. In other words, once a designated frame 310 is identified, the hysteresis may disallow another designated frame 310 from being identified within a certain number of the next number of frames 308 in the sequence, even if there is a frame that would otherwise meet a criteria of a designated frame (e.g., having a color histogram that crosses thethresholds that would merit its designation as a designated frame, in an example using that type of analysis). In this way, system 100 may ensure that the custom background image does not change too rapidly (e.g., so as to become more distracting than immersive) and is able to be updated and revised in natural, indiscreet ways to facilitate the immersiveness of the viewing experience. Additionally, using hysteresis, confidence ranking, and / or other such mechanisms to limit the frequency of new designated frames that are to be fully processed as new images 106 may help manage the amount of compute power that is used by system 100 so that it can keep up with the processing load being requested while providing a high-quality experience for the user in real-time as they experience video content instance 302.
[0041] Once an image 106 is identified (e.g., selected, designated, received, etc.) for generative scene modeling (e.g., whether it is identified as a designated frame 310 of a frame sequence or in another way), system 100 may perform operations described herein to generate a custom background image 114 and to present the image 106 and the custom background image 114 together in an immersive presentation such as the presentation 120 described above. Additional aspects involved in performing such operations will now be described with respect to an illustrative image 106 that depicts certain example 2D content.
[0042] FIG. 4 shows this illustrative image 106 alongside a custom background mask 402 that is configured for use in generating an implementation of custom background image 114 to be presented with the image 106 in accordance with principles described herein. As shown, this example image 106 shows a road cutting through a wooded area in an outdoor scene. For example, the image 106 shown in FIG. 4 may be a frame (e.g., a designated frame) from a video content instance for a shot in which car is driving on this road. To create a custom background image 114 for this particular image 106, custom background mask 402 shows an example canvas onto which custom background content is to be generatively produced by system 100 in any of the ways described herein.
[0043] In the example of FIG. 4, the custom background image 114 produced using custom background mask 402 is configured to combine with the image 106 (e.g., the image being presented in the blank area shown within custom background mask 402) to form a 360° image for presentation on the display of the display device (e.g., display 118 of display device 116). This 360° image is illustrated in FIG. 4 by a line under mask that positions the actual image 106 (the blank area) at 0° and has the custom background mask 402 extending 180° in each direction from there (+180° to the right, -180° to the left) to form a full 360° image. While this example is shown to be rectangular in shape for illustrative purposes, it will be understood that custom background mask 402 and custom background image 114may be spherical to be presented by display device 116 as an immersive scene that appears to envelop the user in all directions.
[0044] Given a mask to be filled in, such as custom background mask 402, there are various techniques and operations that may be used to generate the custom background image 114. As one example, generative imagery may be created using the various models of system 100 to fill in some or all of custom background mask 402 with aesthetically and / or semantically relevant content. As another example (which may be part of this process or used as an alternative), a generative outpainting technique that extends from image 106 may be performed until nearing portions of the mask that have already been filled in (e.g., edges where the mask wraps around to meet other parts of the mask, portions that have already been filled in with generative images using the models of system 100, etc.), whereupon an inpainting technique may be used to minimize visible seams and create a smooth and coherent image. Various details for these processes of generating custom background image 114 will now be described in more detail in relation to FIGS. 5-7.
[0045] FIG. 5 shows illustrative aspects of how a text prompt for a custom background image may be generated based on an image and using a variety of machine learning models in accordance with principles described herein. FIG. 5 further shows how such a text prompt, once generated, may be used in the creation of a custom background image 114. More particularly, FIG. 5 depicts the example image 106 with the road and forested landscape (which was described in relation to FIG. 4) first being received and processed by an implementation of visual -language model 102 and a visual-questionanswering model 502. Visual-question-answering model 502 is shown to employ an example series of questions 504 to help generate, in connection with visual -language model 102, the example textual description 104 of the image. An implementation of large-language model 108 is then shown to be used to convert the textual description 104 into an example text prompt 110, which is used by an implementation of image-generation model 112 as the basis for a custom background image 114. Each of these elements will now be described in more detail.
[0046] As mentioned above, visual -language model 102 may be a machine learning model that is trained to caption images such as image 106. As such, visual -language model 102 may not necessarily have access to contextual information about image 106 (e.g., what video content instance the image comes from, etc.) but may be configured to generate language that describes what is shown in the image when taken alone. Information provided by visual -language model 102 may reflect some degree of understanding of what is depictedin image 106, but this may not reflect certain useful context that may come from a broader understanding of the scene and the video content instance from which the frame is taken. For example, if two characters were shown crossing the street in image 106, visual -language model 102 might identify them as “an adult woman” and a “male child,” and may recognize that they are “holding hands.” Depending on what is happening in the video content instance at this point, however, these observations alone may be insufficient to create a background image with an appropriate mood. For instance, if image 106 comes from a holiday comedy in which a mother is holding her son’s hand as they go to search for a Christmas tree, a very different mood may be extrapolated than if the image 106 comes from a thriller in which a kidnapper is pulling a child along while attempting to flee the police.
[0047] Accordingly, it may be desirable in certain implementations for textual description 104 to capture a more holistic accounting of what is being depicted in the image. For example, it may be desirable for textual description 104 to include information beyond a mere captioning of what image 106 depicts in and of itself. It may be desirable for the description to provide enough thematic, semantic, and aesthetic context that the ultimate ambiance set by the resulting custom background image 114 will match the feeling that the viewer is likely to be experiencing as they are immersed in this portion of the video content instance.
[0048] To this end, FIG. 5 shows that, along with visual -language model 102, the generating of textual description 104 of the image may be performed in this example further using visual-question-answering model 502 to determine the visual characteristics of image 106 for the textual description. For example, certain visual characteristics determined using visual-question-answering model 502 may be semantic characteristics of the image. Such semantic characteristics could include, for example, the video content instance that image 106 comes from (e.g., what movie or show the image is from), the genre of the video content instance (e.g., holiday comedy versus thriller in the example above), the mood of the particular scene (e.g., happy and light versus tense and scary, etc.), the identify of characters or settings shown in the image, context about what is transpiring in the image, and so forth. Other visual characteristics determined using visual-question-answering model 502 may be aesthetic characteristics of the image. Such aesthetic characteristics could include, for example, the color palette exhibited by the image (e.g., a bright, upbeat color scheme; a subdued, gray color scheme; etc.), the lighting style (e.g., bright, outdoor lighting; more intimate indoor lighting; dark, tense lighting; etc.), and other such characteristics.
[0049] The set of questions 504 shown in FIG. 5 to be connected to visual -question-answering model 502 illustrate a few examples of the types of questions that visual-questionanswering model 502 may use to discover relevant information that can be used in textual description 104 (along with more generic captioning of the image that visual -language model 102 may perform). For example, as shown, questions 504 may be configured to direct visualquestion-answering model 502 to determine semantic characteristics such as “What show does the image belong to?” and “What is pictured in the scene?”, as well as aesthetic characteristics such as “Does the pictured scene use natural light?” and “What color is most prominent in the image?” In other examples, the visual-question-answering model 502 could use additional questions about characters depicted in the image (or the actors playing them), events transpiring in the image, settings depicted in the image, and so forth as may serve a particular implementation.
[0050] Together, visual-language model 102 and visual-question-answering model 502 (using questions such as questions 504) may generate a thorough textual description 104 of the image, but one that may not necessarily be fully synthesized, coherent, and / or streamlined for use by downstream processes. For instance, as shown, textual description 104 may include a relatively incoherent list of observations or characteristics that have been observed without yet synthesizing them into a description of a custom background image that is desired. The textual description 104 may, however, include enough information to provide a holistic description of both the aesthetic and semantic meaning of the image. For example, semantic parsers may be used to combine information received from both models 102 and 502 to compile the description into a form that large-language model 108 can work with to synthesize a more coherent text prompt 110.
[0051] Once the textual description 104 of image 106 has been generated, large- language model 108 may be used to convert the textual description 104 into a text prompt 110 for the custom background image by, for example, 1) identifying a plurality of disparate characteristics of the image from the textual description 104 of the image; and 2) synthesizing the text prompt 110 by incorporating the plurality of disparate characteristics into a unified prompt featuring natural and syntactically-correct language. In other words, as shown, large- language model 108 may convert the relatively incoherent description of the various aspects of image 106 that may have been observed using models 102 and / or 502 into a more coherent text prompt that is configured to produce a suitable custom background image when used to prompt an image-generation model such as image-generation model 112. Large-language model 108 may be trained to understand and synthesize a variety of disparate and / or incoherent information (e.g., a list of unrelated characteristics, etc.) so as to produce a well-conceived prompt likely to produce a desirable outcome from image-generation model 112, which itself is trained to intake such prompts in syntactically-correct (i.e., applying applicable rules of grammatical syntax, etc.) and natural (i.e., human-like) language. In this way, the text prompt 110 into which textual description 104 is converted may be configured to cause a custom background image 114 to be customized to image 106 with various customizations as reflected in the characteristics and observations that have been identified in the construction of the prompt. For example, the custom background image 114 may be customized with one or more semantic customizations (e.g., so as to include thematic imagery related to the media content image that the image is a frame of, etc.), one or more aesthetic customizations (e.g., so as to match the mood, lighting, color palette, and other aesthetic aspects of the image), and so forth.
[0052] In some examples, large-language model 108 (or an additional model or layer, not shown in FIG. 5, that is subsequent to and associated with large-language model 108) may be configured to generalize text prompt 110 so as not to over-describe or over-specify what is wanted for the custom background image 114. For instance, in the example mentioned above in which the image frame included characters such as an adult woman and a boy holding hands to cross the street, it may not be desirable for text prompt 110 to specify so much of that detail that the resultant custom background image 114 includes an array of mothers and sons and people holding hands across the background. Rather, since the objective may be to capture the general mood of the frame and produce a background that immerses the viewer more fully into the emotion of the content, a more generalized text prompt may be generated that dispenses with detail about these characters (or at least the detail about what they are currently doing in this particular frame) to request that imagegeneration model 112 produce a more immersive background that is on theme but not overly customized to where it could be more distracting than immersive. In some examples, this generalization could be performed by preconfigured, static text configured to verbally instruct the image-generation model 112 to create a generalized image to capture the vibe of a more specific text prompt that would include the particular details of the image.
[0053] Image-generation model 112 may be implemented using any suitable machine learning model that is configured to receive a natural language prompt (such as has been generated for text prompt 110) and generatively construct the custom background image 114 based on that prompt, as shown in FIG. 5. As one example, image-generation model 112 may be implemented by an image-diffusion model that is configured to convert natural language into an image in a way that accounts for various conditions and constraints that may bedesirable to impose on the image generation. For example, an image diffusion model may be configured to accept constraints and boundary conditions that facilitate the creation of a panorama such as is desired for custom background image 114. An image diffusion model may be well-adapted to create an image, for instance, that exhibits certain characteristics on one edge and then transitions to other characteristics on an opposite edge so as to help make a plurality of images that can be seamlessly stitched into a larger panorama of generative content for a particular custom background image. Other image-generation models that could be used to implement image-generation model 112 may include, for instance, a generative adversarial network, a variational autoencoder model, a pixel-based convolutional neural network (PixelCNN) model, an autoregressive model, a text-to-image model, a hybrid of two or more of these, or the like.
[0054] In some cases (e.g., for certain implementations, for certain images being processed by a given implementation, etc.), it may be desirable to extend 2D content depicted in the image into the background. For example, for a scenic shot such as depicted in the image 106 illustrated in FIG. 5, it may be desirable for certain aspects of the image to extend past the boundaries of image 106 into the custom background image where content may not be carefully examined by the viewer, but where it may be observed in the viewer’s peripheral vision. In this way, the custom background image could further immerse the viewer in the scenic landscape as the image appears to open up and fill the viewer’s entire field of view.
[0055] To achieve this type of image extension, FIG. 6 shows illustrative aspects of how an outpainting technique may be used in the generation of a custom background image in accordance with principles described herein. Specifically, a view 600-A shows the same image 106 shown above (with the road through the forested mountain landscape) with arrows representing a generative outpainting technique 602 that may be performed to extend the edges of the image into the background. In this example, the generating of custom background image 114 includes performing this outpainting technique 602 to produce the custom background image, in accordance with the text prompt, as an outward extension of the image 106. As a result, a view 600-B shows an example of the generative content that image-generation model 112 may produce using the outpainting technique 602. As shown, certain objects such as trees are made to extend beyond the borders of image 106, additional mountainous scenery is included, a sky above the landscape is added, and so forth.
[0056] In some examples, outpainted extensions such as shown in view 600-B may comprise the entirety of a custom background image. For instance, outpainting technique 602 may extend from each edge of image 106 as shown and may ultimately connect back to oneanother to form the full sphere of custom background image 114. It will be understood in this type of example that the further from image 106 the background extends, the more speculative the outpainted content may become as compared to the actual setting being depicted in the video content instance. In other examples, outpainted extensions such as produced by outpainting technique 602 may be used to open up the image to fill more of the user’s field of vision and create a larger, more expansive space, but then may be configured to merge with one or more patches of other imagery (e.g., generative imagery) that are placed in the custom background mask 402 as the custom background image is filled out. Such patches may include imagery semantically and / or aesthetically related to the current image and / or the content instance such as, for example, a depiction of a character, logo, object, setting, symbol, or the like. As one example, the sky generated in the outpainted extensions of view 600-B could merge eventually with an image of the name and logo of the media content instance that is placed on the custom background photosphere to be directly above the user.
[0057] Whether relying exclusively on extending image 106 or connecting generative imagery on image patches placed on the mask, it may be desirable for background content to be smoothly stitched together (e.g., fused, blended, merged, etc.) for the final custom background image 114 that is produced. As such, specific constraints may be placed on image generation to create a smooth and seamless background sphere. One way such constraints may be implemented is by performing, as part of the generating of the custom background image, an inpainting technique. Such an inpainting technique may be used to merge a first portion of the custom background image (e.g., a patch that has been laid down on the mask, an outpainted extension from the image, etc.) with either a second portion of the custom background image (e.g., another patch, another outpainted extension that wraps around to the first one, etc.) or with a portion of the image itself (e.g., an edge that the custom background image is connecting to).
[0058] To illustrate, FIG. 7 shows example aspects of how an inpainting technique may be used in the generation of a custom background image in accordance with principles described herein. Specifically, a view 700-A shows arrows representing a generative inpainting technique 702 that may be performed to smoothly fuse or merge various content as described in text near each edge 704 of the portion of the custom background image that the inpainting technique 702 is being used to fill in. For instance, this example portion of a custom background image may implement a portion of the sky that, as indicated, includes no clouds along a top edge 704-1 and thick clouds along a bottom edge 704-2. Additionally, asfurther indicated for this example portion of the sky, a left edge 704-3 may depict blue sky while a right edge 704-4 depicts the pink sky of a sunset.
[0059] Accounting for each of these aspects of the custom background image that have already been generated (along each edge 704), a view 700-B shows how inpainting technique 702 may be used to blend the existing portions to form inpainted content 706. As shown, for instance, thin clouds may grow thicker as the content moves from top edge 704-1 down to edge 704-2, while blue sky fades into pink sky moving from left edge 704-3 to right edge 704-4. In some examples, inpainting technique 702 may involve additional steps such as blurring and aligning borders, adding padding content (e.g., solid colors such as black or a prominent or average color from image 106) between different portions or patches of a custom background image, and so forth. Ultimately, inpainting technique 702 may facilitate a custom background image that lacks noticeable seams or discontinuities between the image and the custom background image, as well as between different portions of the custom background image itself (e.g., where different patches connect, where the custom background image wraps around, etc.).
[0060] As mentioned above, an input image for which a custom background image is being generated (i.e., an image 106 for which a custom background image 114 is being created) may be a frame (e.g., a designated frame) included within a sequence of 2D frames included in a video content instance such as a movie. In this type of example, system 100 may generate anchor custom background images associated with each designated frame and generate intermediate custom background images for the frames between the designated frames so as to smoothly transition from one anchor custom background image to another. More particularly, for example, a first designated frame may precede a second designated frame in a set of designated frames (e.g., designated frames 310 included within sequence 304 of 2D frames 308), and system 100 may be configured to modify, between the first designated frame and the second designated frame, the custom background image to morph the custom background image from having at least one of a semantic or visual customization with respect to the first designated frame to having at least one of the semantic or visual customization with respect to the second designated frame.
[0061] To illustrate, FIGS. 8A and 8B show example aspects of how certain elements of a custom background image may be dynamically modified. Specifically, example transitions 802 and 804 (i.e., transitions 802-A and 804-A in FIG. 8A and transitions 802-B and 804-B in FIG. 8B) show different ways that a circular shape 806 (e.g., associated with a first anchor custom background image) could be transformed, over the course of severalintermediate frames (e.g., intermediate custom background images), to produce other shapes 808 and / or 810 (e.g., for a second anchor custom background image). In the example transitions 802-A and 804-A of FIG. 8 A, a Euclidean barycenter transport interpolation is shown as one way that circular shape 806 can be smoothly transformed into either of shapes 808 (for transition 802-A) or 810 (for transition 804-A). In contrast, corresponding transitions 802-B and 804-B of FIG. 8B show a Wasserstein barycenter transport interpolation being employed as an alternative way that the same transformations could be performed.
[0062] As shown, the Euclidean barycenter transport interpolation of FIG. 8 A may fade gradually from one shape to another (it will be understood that the different fill patterns shown in FIG. 8A represent different shades of color as circular shape 806 fades out and a target shape 808 or 810 fades in). Conversely, the Wasserstein barycenter transport interpolation of FIG. 8B is shown to not rely on this type of fading, but, rather, gradually morphs circular shape 806 until it becomes the desired target shape 808 or 810. Using these types of transport interpolations (including other suitable transport interpolations not explicitly shown), system 100 may not only reduce or eliminate spatial discontinuities in generative custom background images being produced (e.g., noticeable seams between image 106 and / or different parts of the custom background image 114), but also may reduce or eliminate temporal discontinuities from frame to frame so that the custom background image does not change so abruptly as to distract or take away from the immersiveness of the user experience.
[0063] Combining the various principles that have been described, FIG. 9 shows an illustrative custom background image 114 that may be generated based on the illustrative image 106 that has been used as a primary example throughout this description. This custom background image 114 may be configured for presentation with image 106 on a display device (e.g., a head-mounted display device, etc.) in accordance with principles described herein. In this example, the road is shown to not only be included in image 106 but also to continue beneath the image 106 (e.g., if the user looks down) and to pick up again if the user turns 180° to the right or left to look at the scene behind them. Because the custom background image 114 is understood to be a 360° image that wraps around, the road behind them is split half on the left side and half on the right side of the illustration in FIG. 9. While no patches are shown in this particular example, it will be understood that various objects (e.g., symbols, etc.), text (e.g., title, logos, etc.), likenesses (e.g., characters, etc.), and / or other content relevant to image 106 and / or the video content instance from which image 106 comes could also be integrated into this custom background image 114.
[0064] As has been mentioned, various methods and processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer- readable medium and executable by one or more computing devices. In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium (e.g., a memory, etc.), and executes those instructions, thereby performing one or more operations such as the operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.
[0065] A computer-readable medium (also referred to as a processor-readable medium) includes any non-transitory medium that participates in providing data (e.g., instructions) that may be read by a computer (e.g., by a processor of a computer). Such a medium may take many forms, including, but not limited to, non-volatile media, and / or volatile media. Non-volatile media may include, for example, optical or magnetic disks and other persistent memory. Volatile media may include, for example, dynamic random-access memory (DRAM), which typically constitutes a main memory. Common forms of computer- readable media include, for example, a disk, hard disk, magnetic tape, any other magnetic medium, a compact disc read-only memory (CD-ROM), a digital video disc (DVD), any other optical medium, random access memory (RAM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EPROM), FLASH- EEPROM, any other memory chip or cartridge, or any other tangible medium from which a computer can read.
[0066] FIG. 10 shows an illustrative computing system 1000 that may be used to implement various devices and / or systems described herein. For example, computing system 1000 may include or implement (or partially implement) generative scene modeling systems such as generative scene modeling system 100 and / or any components thereof, such as the display device 116, the computing resources used to implement the machine learning models, and so forth.
[0067] As shown in FIG. 10, computing system 1000 may include a communication interface 1002, a processor 1004, a storage device 1006, and an input / output (I / O) module 1008 communicatively connected via a communication infrastructure 1010. While an illustrative computing system 1000 is shown in FIG. 10, the components illustrated in FIG. 10 are not intended to be limiting. Additional or alternative components may be used in other embodiments. Components of computing system 1000 shown in FIG. 10 will now be described in additional detail.
[0068] Communication interface 1002 may be configured to communicate with oneor more computing devices. Examples of communication interface 1002 include, without limitation, a wired network interface (such as a network interface card), a wireless network interface (such as a wireless network interface card), a modem, an audio / video connection, and any other suitable interface.
[0069] Processor 1004 generally represents any type or form of processing unit capable of processing data or interpreting, executing, and / or directing execution of one or more of the instructions, processes, and / or operations described herein. Processor 1004 may direct execution of operations in accordance with one or more applications 1012 or other computer-executable instructions such as may be stored in storage device 1006 or another computer-readable medium.
[0070] Storage device 1006 may include one or more data storage media, devices, or configurations and may employ any type, form, and combination of data storage media and / or device. For example, storage device 1006 may include, but is not limited to, a hard drive, network drive, flash drive, magnetic disc, optical disc, RAM, dynamic RAM, other non-volatile and / or volatile data storage units, or a combination or sub-combination thereof. Electronic data, including data described herein, may be temporarily and / or permanently stored in storage device 1006. For example, data representative of one or more executable applications 1012 configured to direct processor 1004 to perform any of the operations described herein may be stored within storage device 1006. In some examples, data may be arranged in one or more databases residing within storage device 1006.
[0071] I / O module 1008 may include one or more I / O modules configured to receive user input and provide user output. One or more VO modules may be used to receive input for a single virtual experience. I / O module 1008 may include any hardware, firmware, software, or combination thereof supportive of input and output capabilities. For example, I / O module 1008 may include hardware and / or software for capturing user input, including, but not limited to, a keyboard or keypad, a touchscreen component (e.g., touchscreen display), a receiver (e.g., an RF or infrared receiver), motion sensors, and / or one or more input buttons.
[0072] I / O module 1008 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, I / O module 1008 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.
[0073] The following examples describe systems and methods for generative scene modeling for immersive viewing of 2D content in accordance with principles described herein:
[0074] 1. A method comprising: generating, using a first machine learning model, a textual description of an image; converting, using a second machine learning model, the textual description of the image into a text prompt for a custom background image; generating, using a third machine learning model, the custom background image based on the text prompt; and presenting, on a display of a display device, the image and the custom background image.
[0075] 2. The method of any of the preceding examples, wherein: the first machine learning model is a visual-language model; the second machine learning model is a large- language model; and the third machine learning model is an image-generation model.
[0076] 3. The method of any of the preceding examples, wherein the image is a frame included within a sequence of 2D frames.
[0077] 4. The method of any of the preceding examples, further comprising: accessing a video content instance that includes the sequence of 2D frames; and designating the frame to be a designated frame of the sequence of 2D frames.
[0078] 5. The method of any of the preceding examples, wherein the designating is performed using hysteresis to prevent the designated frame from being identified within a threshold number of frames after a previous designated frame within the sequence of 2D frames.
[0079] 6. The method of any of the preceding examples, wherein: the designated frame is a first designated frame that precedes a second designated frame in a set of designated frames included within the sequence of 2D frames; and the method further comprises modifying, between the first designated frame and the second designated frame, the custom background image to morph the custom background image from having at least one of a semantic or visual customization with respect to the first designated frame to having at least one of the semantic or visual customization with respect to the second designated frame.
[0080] 7. The method of any of the preceding examples, wherein the modifying the custom background image includes applying an optical transport interpolation to morph the custom background image smoothly from frame to frame between the first designated frame and the second designated frame within the sequence of 2D frames.
[0081] 8. The method of any of the preceding examples, wherein the generating the custom background image includes performing an outpainting technique to produce the custom background image, in accordance with the text prompt, as an outward extension of the image.
[0082] 9. The method of any of the preceding examples, wherein the generating the custom background image includes performing an inpainting technique to merge a first portion of the custom background image with at least one of a second portion of the custom background image or a portion of the image.
[0083] 10. The method of any of the preceding examples, wherein the generating the textual description of the image is performed further using a visual-question-answering model to determine a visual characteristic of the image for the textual description.
[0084] 11. The method of any of the preceding examples, wherein: the visual characteristic determined using the visual-question-answering model is a semantic characteristic of the image; and the text prompt into which the textual description is converted is configured to cause the custom background image to be customized to the image with a semantic customization.
[0085] 12. The method of any of the preceding examples, wherein: the visual characteristic determined using the visual-question-answering model is an aesthetic characteristic of the image; and the text prompt into which the textual description is converted is configured to cause the custom background image to be customized to the image with an aesthetic customization.
[0086] 13. The method of any of the preceding examples, wherein the converting the textual description of the image into the text prompt for the custom background image includes: identifying a plurality of disparate characteristics of the image from the textual description of the image; and synthesizing the text prompt by incorporating the plurality of disparate characteristics into a unified prompt featuring natural and syntactically correct language.
[0087] 14. The method of any of the preceding examples, wherein: the display device is a head-mounted display device; and the custom background image is configured to combine with the image to form a 360-degree image when presented on the display of the head-mounted display device.
[0088] 15. A non-transitory computer-readable medium storing instructions that, when executed, cause a processor to perform a process comprising: accessing a video content instance including a sequence of 2D frames; generating, using a first machine learning model,a textual description of a designated frame included within the sequence of 2D frames; converting, using a second machine learning model, the textual description of the designated frame into a text prompt for a custom background image; generating, using a third machine learning model, the custom background image based on the text prompt; and presenting, on a display of a display device, the designated frame and the custom background image.
[0089] 16. The non-transitory computer-readable medium of any of the preceding examples, wherein: the first machine learning model is a visual -language model; the second machine learning model is a large-language model; and the third machine learning model is an image-generation model.
[0090] 17. The non-transitory computer-readable medium of any of the preceding examples, wherein the process further comprises designating, within the sequence of 2D frames, the designated frame using hysteresis to prevent the designated frame from being identified within a threshold number of frames after a previous designated frame within the sequence of 2D frames.
[0091] 18. The non-transitory computer-readable medium of any of the preceding examples, wherein: the designated frame is a first designated frame that precedes a second designated frame in a set of designated frames included within the sequence of 2D frames; the process further comprises modifying, between the first designated frame and the second designated frame, the custom background image to morph the custom background image from having at least one of a semantic or visual customization with respect to the first designated frame to having at least one of the semantic or visual customization with respect to the second designated frame; and the modifying the custom background image includes applying an optical transport interpolation to morph the custom background image smoothly from frame to frame between the first designated frame and the second designated frame within the sequence of 2D frames.
[0092] 19. The non-transitory computer-readable medium of any of the preceding examples, wherein the generating the custom background image includes: performing an outpainting technique to produce the custom background image, in accordance with the text prompt, as an outward extension of the designated frame; and performing an inpainting technique to merge a first portion of the custom background image with at least one of a second portion of the custom background image or a portion of the designated frame.
[0093] 20. A generative scene modeling system comprising: a visual -language model configured to generate a textual description of an image; a large-language model configured to convert the textual description of the image into a text prompt for a custom backgroundimage; an image-generation model configured to generate the custom background image based on the text prompt; and a display device including a display configured to present the image and the custom background image.
[0094] 21. The generative scene modeling system of any of the preceding examples, wherein the image is a frame included within a sequence of 2D frames of a video content instance presented on the display within the head-mounted display device.
[0095] 22. The generative scene modeling system of any of the preceding examples, wherein: the frame is a first designated frame that precedes a second designated frame in a set of designated frames included within the sequence of 2D frames; and the image-generation model is further configured to modify, between the first designated frame and the second designated frame, the custom background image to morph the custom background image from having at least one of a semantic or visual customization with respect to the first designated frame to having at least one of the semantic or visual customization with respect to the second designated frame.
[0096] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0097] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the description and claims. In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
[0098] Specific structural and functional details disclosed herein are merely representative for purposes of describing example implementations. Example implementations, however, may be embodied in many alternate forms and should not be construed as limited to only the implementations set forth herein.
[0099] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. A first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of the implementations of the disclosure. As used herein, the term and / or includes any and all combinations of one or more of the associated listed items.
[0100] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the implementations. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used in this specification, specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0101] It will be understood that when an element is referred to as being “coupled,” “connected,” or “responsive” to, or “on,” another element, it can be directly coupled, connected, or responsive to, or on, the other element, or intervening elements may also be present. In contrast, when an element is referred to as being “directly coupled,” “directly connected,” or “directly responsive” to, or “directly on,” another element, there are no intervening elements present. As used herein the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0102] Spatially relative terms, such as “beneath,” “below,” “lower,” “above,” “upper,” and the like, may be used herein for ease of description to describe one element or feature in relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as “below” or “beneath” other elements or features would then be oriented “above” the other elements or features. Thus, the term “below” can encompass both an orientation of above and below. The device may be otherwise oriented (rotated 130 degrees or at other orientations) and the spatially relative descriptors used herein may be interpreted accordingly.
[0103] Unless otherwise defined, the terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in theart to which these concepts belong. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or the present specification and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0104] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover such modifications and changes as fall within the scope of the implementations. It will be understood that they have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described. As such, the scope of the present disclosure is not limited to the particular combinations hereafter claimed, but instead extends to encompass any combination of features or example implementations described herein irrespective of whether or not that particular combination has been specifically enumerated in the accompanying claims at this time.
Claims
WHAT IS CLAIMED IS:
1. A method compri sing : generating, using a first machine learning model, a textual description of an image; converting, using a second machine learning model, the textual description of the image into a text prompt for a custom background image; generating, using a third machine learning model, the custom background image based on the text prompt; and presenting, on a display of a display device, the image and the custom background image.
2. The method of claim 1, wherein: the first machine learning model is a visual -language model; the second machine learning model is a large-language model; and the third machine learning model is an image-generation model.
3. The method of any one of claims 1-2, wherein the image is a frame included within a sequence of 2D frames.
4. The method of claim 3, further comprising: accessing a video content instance that includes the sequence of 2D frames; and designating the frame to be a designated frame of the sequence of 2D frames.
5. The method of claim 4, wherein the designating is performed using hysteresis to prevent the designated frame from being identified within a threshold number of frames after a previous designated frame within the sequence of 2D frames.
6. The method of claim 4 or 5, wherein: the designated frame is a first designated frame that precedes a second designated frame in a set of designated frames included within the sequence of 2D frames; and the method further comprises modifying, between the first designated frame and the second designated frame, the custom background image to morph the custom background image from having at least one of a semantic or visual customization with respect to the firstdesignated frame to having at least one of the semantic or visual customization with respect to the second designated frame.
7. The method of claim 6, wherein the modifying the custom background image includes applying an optical transport interpolation to morph the custom background image smoothly from frame to frame between the first designated frame and the second designated frame within the sequence of 2D frames.
8. The method of any one of claims 1-7, wherein the generating the custom background image includes performing an outpainting technique to produce the custom background image, in accordance with the text prompt, as an outward extension of the image.
9. The method of any one of claims 1-8, wherein the generating the custom background image includes performing an inpainting technique to merge a first portion of the custom background image with at least one of a second portion of the custom background image or a portion of the image.
10. The method of any one of claims 1-9, wherein the generating the textual description of the image is performed further using a visual -question-answering model to determine a visual characteristic of the image for the textual description.
11. The method of claim 10, wherein: the visual characteristic determined using the visual-question-answering model is a semantic characteristic of the image; and the text prompt into which the textual description is converted is configured to cause the custom background image to be customized to the image with a semantic customization.
12. The method of claim 10, wherein: the visual characteristic determined using the visual-question-answering model is an aesthetic characteristic of the image; and the text prompt into which the textual description is converted is configured to cause the custom background image to be customized to the image with an aesthetic customization.
13. The method of any one of claims 1-12, wherein the converting the textual description of the image into the text prompt for the custom background image includes: identifying a plurality of disparate characteristics of the image from the textual description of the image; and synthesizing the text prompt by incorporating the plurality of disparate characteristics into a unified prompt featuring natural and syntactically correct language.
14. The method of any one of claims 1-13, wherein: the display device is a head-mounted display device; and the custom background image is configured to combine with the image to form a 360- degree image when presented on the display of the head-mounted display device.
15. A non-transitory computer-readable medium storing instructions that, when executed, cause a processor to perform a process comprising: accessing a video content instance including a sequence of 2D frames; generating, using a first machine learning model, a textual description of a designated frame included within the sequence of 2D frames; converting, using a second machine learning model, the textual description of the designated frame into a text prompt for a custom background image; generating, using a third machine learning model, the custom background image based on the text prompt; and presenting, on a display of a display device, the designated frame and the custom background image.
16. The non-transitory computer-readable medium of claim 15, wherein: the first machine learning model is a visual -language model; the second machine learning model is a large-language model; and the third machine learning model is an image-generation model.
17. The non-transitory computer-readable medium of any one of claims 15-16, wherein the process further comprises designating, within the sequence of 2D frames, the designated frame using hysteresis to prevent the designated frame from being identified within a threshold number of frames after a previous designated frame within the sequence of 2D frames.
18. The non-transitory computer-readable medium of any one of claims 15-17, wherein: the designated frame is a first designated frame that precedes a second designated frame in a set of designated frames included within the sequence of 2D frames; the process further comprises modifying, between the first designated frame and the second designated frame, the custom background image to morph the custom background image from having at least one of a semantic or visual customization with respect to the first designated frame to having at least one of the semantic or visual customization with respect to the second designated frame; and the modifying the custom background image includes applying an optical transport interpolation to morph the custom background image smoothly from frame to frame between the first designated frame and the second designated frame within the sequence of 2D frames.
19. The non-transitory computer-readable medium of any one of claims 15-18, wherein the generating the custom background image includes: performing an outpainting technique to produce the custom background image, in accordance with the text prompt, as an outward extension of the designated frame; and performing an inpainting technique to merge a first portion of the custom background image with at least one of a second portion of the custom background image or a portion of the designated frame.
20. A generative scene modeling system comprising: a visual -language model configured to generate a textual description of an image; a large-language model configured to convert the textual description of the image into a text prompt for a custom background image; an image-generation model configured to generate the custom background image based on the text prompt; and a display device including a display configured to present the image and the custom background image.
21. The generative scene modeling system of claim 20, wherein the image is a frame included within a sequence of 2D frames of a video content instance presented on the display within the display device.
22. The generative scene modeling system of claim 21, wherein: the frame is a first designated frame that precedes a second designated frame in a set of designated frames included within the sequence of 2D frames; and the image-generation model is further configured to modify, between the first designated frame and the second designated frame, the custom background image to morph the custom background image from having at least one of a semantic or visual customization with respect to the first designated frame to having at least one of the semantic or visual customization with respect to the second designated frame.
Citation Information
Patent Citations
Video processing method and related equipment
CN116193275A
Image generation method, electronic equipment and storage medium
CN116485943A