Generating immersive augmented reality experiences from existing images and videos
By combining a volumetric content presentation system with IoT devices, the system tracks user location and environmental changes in real time and dynamically renders volumetric content, solving the problem of insufficient immersive experience in existing augmented reality technologies and achieving a more immersive and interactive AR experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SNAP INC
- Filing Date
- 2023-08-23
- Publication Date
- 2026-07-24
Smart Images

Figure CN119816868B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This patent application claims the benefit of U.S. Application Serial No. 18 / 170,271, filed February 16, 2023, which claims the benefit of U.S. Provisional Patent Application No. 63 / 402,914, filed August 31, 2022, entitled “Generating Immersive Augmented Reality Experiences Based on Existing Images and Videos,” the entire contents of each of which are incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to mobile and wearable computing technologies. In particular, exemplary embodiments of this disclosure relate to systems, methods, and user interfaces for creating immersive augmented reality experiences based on existing images and videos. Background Technology
[0004] Augmented reality (AR) experiences include applying virtual content to a real-world environment, whether by presenting virtual content through a transparent display through which the real-world environment can be seen, or by augmenting image data to include virtual content overlaid on the real-world environment depicted therein. Virtual content can include one or more AR content items. AR content items can include audio content, visual content, or visual effects. Examples of audio and visual content include photographs, text, logos, animations, and sound effects. Devices that enable AR experiences in any of these methods are referred to herein as "AR devices."
[0005] For some example AR devices, audio and visual content or visual effects are applied to media data, such as live image streams. Other example AR devices include head-mounted displays that can be implemented with transparent or semi-transparent displays, through which users can view their surroundings. Such devices allow users to view their surroundings through the transparent or semi-transparent display and also to see objects generated for display that appear as part of and / or superimposed on the surroundings (e.g., virtual objects such as 3D renderings, images, videos, text, etc.). Users of head-mounted devices can access and use computer software applications to perform various tasks or engage in entertainment activities. To use the computer software applications, users interact with a 3D user interface provided by the head-mounted device.
[0006] The so-called "Internet of Things" or "IoT" is a network of physical objects (called "smart devices" or "IoT devices") embedded with sensors, software, and other technologies, used to connect and exchange data with other devices via the Internet. For example, IoT devices are used in home automation to control lighting, heating and air conditioning, media and security systems, and camera systems. Several IoT-enabled devices are already available, serving as smart home hubs to connect different smart home products. IoT devices are also used in several other applications. Application layer protocols and supporting frameworks for implementing such IoT applications have been provided. Artificial intelligence has also been integrated with IoT infrastructure to enable more efficient IoT operations, improve human-computer interaction, and enhance data management and analysis. Attached Figure Description
[0007] To facilitate identification of any discussion of a particular element or action, one or more of the highest significant digits in the figure references refer to the figure number in which the element is first introduced.
[0008] Figure 1 It is a graphical representation of a network environment based on some example implementations, in which a volumetric content rendering system can be deployed.
[0009] Figure 2A This is a perspective view of a head-mounted device according to some example implementations.
[0010] Figure 2B The following are illustrated according to some example implementations. Figure 2A Other views of the head-mounted device.
[0011] Figure 3 This illustrates, according to some example implementations, including Figure 1 A block diagram showing the details of the networked system for the head-mounted device.
[0012] Figure 4 This is a graphical representation of a volumetric content presentation system with both client-side and server-side functionality, based on some example implementations.
[0013] Figure 5 It is a graphical representation of a data structure maintained in a database according to some example implementations.
[0014] Figure 6 This is a conceptual diagram illustrating an implementation of a volumetric content rendering system that identifies one or more two-dimensional elements from image data for rendering in augmented reality, according to some example embodiments.
[0015] Figure 7A This is a conceptual diagram illustrating the user's view before one or more volumetric content items are overlaid on a real-world environment, according to some example implementations.
[0016] Figure 7B This is a conceptual diagram illustrating the user's field of vision after one or more volumetric content items are overlaid on a real-world environment, according to some example implementations.
[0017] Figures 8A to 8D This is a flowchart illustrating the operation of a volumetric content rendering system when performing a method for creating an immersive AR experience based on existing image data, according to some example implementations.
[0018] Figure 9 This is a block diagram illustrating representative software architectures that can be used in conjunction with various hardware architectures described herein, based on some example implementations.
[0019] Figure 10 This is a block diagram illustrating components of a machine, according to some example embodiments, capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any or more of the methods discussed herein. Detailed Implementation
[0020] The following description includes systems, methods, techniques, instruction sequences, and computer program products for implementing illustrative embodiments of the present disclosure. In the following description, numerous specific details are set forth for illustrative purposes in order to provide an understanding of various embodiments of the subject matter of the invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the invention can be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques need not be shown in detail.
[0021] Volumetric content is an example of an augmented reality (AR) experience. Volumetric content can include volumetric video and three-dimensional spatial images captured in 3D (along with audio signals recorded along with the volumetric video and images). Recording volumetric content involves capturing elements of three-dimensional space, such as objects and people, volumetrically using a combination of camera devices and sensors. Volumetric content includes volumetric representations of one or more three-dimensional elements (e.g., objects or people) in three-dimensional space. A volumetric representation of an element (e.g., an AR content item) refers to a visual representation of a three-dimensional element in three dimensions. Rendering volumetric content can include displaying one or more AR content items superimposed on a real-world space, which may be the same space as the three-dimensional space from which the volumetric video was captured or a different space. Rendering volumetric content can include: displaying one or more content items in motion; displaying one or more content items performing movement or other actions; displaying one or more content items in a static position; or a combination thereof. Content items may be displayed during the rendering of volumetric content or portions thereof.
[0022] The presentation of volumetric content can include tracking a user's position and movement within their physical real-world environment, and using that tracked position and movement to allow the user to move around and interact with the volumetric content. Therefore, the presentation of volumetric content can include displaying content items from multiple perspectives based on the user's movement and changes in position. In this way, the presentation of volumetric content provides users with an immersive AR experience.
[0023] This disclosure encompasses various aspects including systems, methods, techniques, instruction sequences, and computer program products for creating immersive augmented reality experiences from existing images and videos. Whether it's a single video from a single scene or a series of images and videos about travel, the volumetric content rendering systems described herein allow users to create immersive experiences using metadata and visual details.
[0024] In the example, the existing image depicts a dog playing with a ball in a park. The volumetric content rendering system 100 segments the different components of the image to render them in AR. The grass from the park becomes green, the background of the image turns the wall into a similar shadow, and the dog and ball are segmented and placed in the physical world as "clones" in AR. In some implementations, the system performs the segmentation automatically. In some implementations, if the image contains too many key elements, the segmentation can rely on user input. Furthermore, using, for example, weather data, the system can render AR clouds for a cloudy day or AR snow if it's snowing that day. If it's windy, using AI, the system can create additional effects to make the dog's fur blow as if the dog is experiencing wind. Although the current example deals with a location-independent scene, the creation of immersive AR experiences can be location-dependent and may also be more contextual. In some implementations, IoT / smart devices can be used to make the experience more immersive, for example, by triggering a smart fan to add a wind effect.
[0025] Figure 1 This is a block diagram illustrating an example volumetric content rendering system 100 for rendering volumetric content. The volumetric content rendering system 100 includes a client device 102. The client device 102 hosts several applications including rendering clients 104. Each rendering client 104 is communicatively coupled to a rendering server system 108 via a network 106 (e.g., the Internet). In this example, the client device 102 is a wearable device (e.g., smart glasses) worn by a user 124, which includes a camera device and optics including a transparent display through which the user 124 can see the real-world environment.
[0026] The rendering client 104 can communicate and exchange data with another rendering client 104 and the rendering server system 108 via the network 106. The data exchanged between rendering clients 104 and between another rendering client 104 and the rendering server system 108 includes features (e.g., commands to enable features) and payload data (e.g., text, audio, video, or other multimedia data).
[0027] The rendering server system 108 provides server-side functionality to a specific rendering client 104 via network 106. While some functions of the volumetric content rendering system 100 are described herein as being performed by either the rendering client 104 or the rendering server system 108, the location of some functions within the rendering client 104 or the rendering server system 108 is a design choice. For example, it may be technically preferred that certain technologies and functions are initially deployed within the rendering server system 108, but that technology and functions are later migrated to the rendering client 104, which has sufficient processing power on the client device 102.
[0028] The rendering server system 108 supports various services and operations provided to the rendering client 104. Such operations include sending data to and receiving data from the rendering client 104, and processing data generated by the rendering client 104. Specifically, turning now to the communication server system 108, the application programming interface (API) server 110 is coupled to the application server 112 and provides a programming interface to the application server 112. As an example, this data may include volumetric content (e.g., volumetric video), message content, device information, geolocation information, media annotations and overlays, message content persistence conditions, social network information, and real-time event information. Data exchange within the volumetric content rendering system 100 is invoked and controlled via the user interface (UI) and the available functions of the rendering client 104.
[0029] Specifically, now turning to presentation server system 108, application programming interface (API) server 110 is coupled to application server 112 and provides a programming interface to application server 112. Application server 112 is communicatively coupled to database server 114, which facilitates access to database 116, in which data associated with messages processed by application server 112 is stored.
[0030] Application Programming Interface (API) server 110 receives and sends message data (e.g., commands and message payloads) between client device 102 and application server 112. Specifically, API server 110 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by presentation client 104 to invoke functions of application server 112. API server 110 exposes various functions supported by application server 112, including: account registration, login functionality, sending messages from one presentation client 104 to another presentation client 104 via application server 112, sending media files (e.g., volumetric video) to presentation client 104, setting a collection of media data (e.g., a story), retrieving a user's friend list for client device 102, retrieving such a collection, retrieving messages and content, adding and deleting friends in a social graph, the location of friends within the social graph, and opening application events (e.g., related to presentation client 104).
[0031] Application server 112 hosts several server applications and subsystems, including rendering server 118, image processing server 120, and social networking server 122. Rendering server 118 is typically responsible for managing volumetric content and facilitating its rendering by client device 102. Image processing server 120 is dedicated to performing various image processing operations, typically related to images or videos generated and displayed by client device 102. Rendering server 118 and image processing server 120 may work together to provide user 124 with one or more AR experiences. For example, rendering server 118 and image processing server 120 may work together to support the rendering of volumetric content by client device 102. Further details regarding the rendering of volumetric content will be discussed below.
[0032] Social networking server 122 supports various social networking features and services and makes these features and services available to presentation server 118. To this end, social networking server 122 maintains and accesses an entity graph within database 116. Examples of features and services supported by social networking server 122 include identifiers of other users associated with or “following” a particular user within volumetric content presentation system 100, as well as identifiers of the particular user’s interests and other entities.
[0033] Application server 112 is communicatively coupled to database server 114, which facilitates access to database 116, which stores data associated with content rendered by rendering server 118 and image processing server 120.
[0034] The presentation server system 108 can also communicate and exchange data with one or more network-connected devices 126. For some devices, the presentation server system 108 can communicate and exchange data directly with the network-connected devices 126(126), while in other instances, the presentation server system 108 can communicate and exchange data with the network-connected devices 126(126) via one or more web servers 128 (e.g., third-party applications). For example, the web servers 128(126) may expose one or more APIs for communicating with the network-connected devices 126(126). Examples of data communicated between the presentation server system 108 and one or more network-connected devices include device status data and sensor data, as well as various requests and commands, or as part of various requests and commands. Figure 1 As shown, in some embodiments, the client device 102 (e.g., Figure 2A The display device (such as glasses 200) is different from (one or more) network connection devices 126.
[0035] As used herein, the term "network-connected device" includes devices known to those skilled in the art as "IoT devices." Therefore, network-connected device 126 may include common household and other devices that a standard end user might encounter, such as smart lights and bulbs, thermostats, smart TVs, smart speakers, smart switches, smart appliances (e.g., washing machines, dryers, stoves, and microwave ovens), navigation systems, etc.
[0036] Figure 2A This is a perspective view based on some example head-mounted display devices (e.g., glasses 200). Glasses 200 is... Figure 1 An example of client device 102. Glasses 200 are capable of displaying content and are therefore an example of a display device, as will be mentioned below. Furthermore, the display capabilities of glasses 200 support AR experiences, and therefore glasses 200 are an example of an AR device. As described above, AR experiences include: applying virtual content to a real-world environment, whether by presenting virtual content through a transparent display through which a real-world environment can be seen, or by enhancing image data to include virtual content overlaid on the real-world environment depicted therein.
[0037] The eyeglasses 200 may include a frame 202 made of any suitable material, such as plastic or metal, including any suitable shape memory alloy. In one or more examples, the frame 202 includes a first optical element holder or left optical element holder 204 (e.g., a display or lens holder) and a second optical element holder or right optical element holder 206 connected by a bridging portion 212. The first optical element or left optical element 208 and the second optical element or right optical element 210 may be disposed within the left optical element holder 204 and the right optical element holder 206, respectively. The right optical element 210 and the left optical element 208 may be a lens, a display, a display assembly, or a combination thereof. Any suitable display assembly may be disposed in the eyeglasses 200.
[0038] Frame 202 further includes a left arm or left temple 222 and a right arm or right temple 224. In some examples, frame 202 may be formed from a single piece of material to have a monolithic or integrated construction.
[0039] The eyeglasses 200 may include a computing device such as computer 220, which may be of any suitable type for carrying by frame 202, and in one or more examples, the computing device may have a suitable size and shape to be partially housed in one of temple pieces 222 or 224. Computer 220 may include one or more processors, as well as memory, wireless communication circuitry, and a power supply. As discussed below, computer 220 includes low-power circuitry, high-speed circuitry, and a display processor. Various other examples may include these elements in different configurations or integrated in different ways. Additional details of various aspects of computer 220 may be implemented as shown in data processor 302, which is discussed below.
[0040] The computer 220 additionally includes a battery 218 or other suitable portable power source. In some examples, the battery 218 is disposed in the left temple 222 and electrically coupled to the computer 220 disposed in the right temple 224. The glasses 200 may include a connector or port (not shown) for charging the battery 218, a wireless receiver, a transmitter or transceiver (not shown), or a combination of such devices.
[0041] The glasses 200 include a first or left camera device 214 and a second or right camera device 216. Although two camera devices are depicted, other examples contemplate the use of a single or additional (i.e., more than two) camera devices. In one or more examples, in addition to the left camera device 214 and the right camera device 216, the glasses 200 also includes any number of input sensors or other input / output devices. Such sensors or input / output devices may additionally include biometric sensors, position sensors, motion sensors, etc.
[0042] In some examples, the left camera device 214 and the right camera device 216 provide video frame data for the glasses 200 to use to extract 3D information from the real-world scene.
[0043] The glasses 200 may also include a touchpad 226, which is mounted to or integrated with one or both of the left temple 222 and the right temple 224. The touchpad 226 is generally arranged vertically, and in some examples is approximately parallel to the user's temple. As used herein, generally vertical alignment means that the touchpad is more vertical than horizontal, but may be more vertical than described above. Additional user input can be provided via one or more buttons 228, which in the illustrated example are located on the outer upper edges of the left optics retainer 204 and the right optics retainer 206. The one or more touchpads 226 and buttons 228 provide a means by which the glasses 200 can receive input from the user of the glasses 200.
[0044] Figure 2B The glasses 200 are shown from the user's perspective. For clarity, Figure 2A Several components shown have been omitted. For example... Figure 2A As described in Figure 2B The glasses 200 shown include a left optical element 208 and a right optical element 210, which are respectively fixed in a left optical element holder 204 and a right optical element holder 206.
[0045] The glasses 200 include: a forward optical assembly 230, which includes a right projector 232 and a right near-eye display 234; and a forward optical assembly 238, which includes a left projector 240 and a left near-eye display 244.
[0046] In some examples, the near-eye display is a waveguide. The waveguide includes reflective or diffractive structures (e.g., gratings and / or optical elements such as mirrors, lenses, or prisms). Light 236 emitted by projector 232 encounters the diffractive structure of the waveguide of near-eye display 234, which directs the light toward the user's right eye to provide an image superimposed on or within the right optical element 210, representing a view of the real world seen by the user. Similarly, light 242 emitted by projector 240 encounters the diffractive structure of the waveguide of near-eye display 244, which directs the light toward the user's left eye to provide an image superimposed on or within the left optical element 208, representing a view of the real world seen by the user. A combination of the GPU, forward optics 230, left optical element 208, and right optical element 210 provides the optical engine for glasses 200. Glasses 200 uses the optical engine to generate an overlay of the user's real-world view, including displaying a 3D user interface to the user of glasses 200.
[0047] However, it should be understood that other display technologies or configurations can be used within the optical engine to display images to the user in their field of view. For example, instead of the projector 232 and the waveguide, an LCD, LED, or other display panel or surface can be provided.
[0048] In use, the user of glasses 200 will be presented with information, content, and various 3D user interfaces on a near-eye display. As described in more detail herein, the user can then use touchpad 226 and / or buttons 228, associated devices (e.g., Figure 3 The client device 328 shown interacts with the glasses 200 via voice input or touch input and / or hand movements, positions and orientations detected by the glasses 200.
[0049] Figure 3 This is a block diagram illustrating details of a networking system 300, including glasses 200, according to some examples. The networking system 300 includes glasses 200, client device 328, and server system 332. Client device 328 may be a smartphone, tablet computer, tablet phone, laptop computer, access point, or any other such device capable of connecting to glasses 200 using both low-power wireless connection 336 and / or high-speed wireless connection 334. Client device 328 is connected to server system 332 via network 330. Network 330 may include any combination of wired and wireless connections. Server system 332 may be one or more computing devices as part of a service or network computing system. (The last sentence appears to be incomplete and possibly refers to separate components.) Figure 9 and Figure 10 The details of the software architecture 920 or machine 1000 described herein are used to implement any elements of the client device 328, server system 332, and network 330.
[0050] The glasses 200 include a data processor 302, a display 310, one or more camera devices 308, and additional input / output elements 316. The input / output elements 316 may include a microphone, an audio speaker, a biometric sensor, additional sensors, or an additional display element integrated with the data processor 302. (About...) Figure 9 and Figure 10 Examples of input / output element 316 are further discussed. For example, input / output element 316 may include any I / O component 1018, which may include output component 1026, moving component 1034, etc. Figure 2B An example of display 310 is discussed herein. In the specific example described herein, display 310 includes displays for the user's left and right eyes.
[0051] The data processor 302 includes an image processor 306 (e.g., a video processor), a GPU and a display driver 338, a tracking module 340, an interface 312, a low-power circuit system 304, and a high-speed circuit system 320. The components of the data processor 302 are interconnected via a bus 342.
[0052] Interface 312 refers to any source of user commands provided to data processor 302. In one or more examples, interface 312 is a physical button that, when pressed, sends a user input signal from interface 312 to low-power processor 314. Low-power processor 314 can process pressing such a button and then immediately releasing it as a request to capture a single image, and vice versa. Low-power processor 314 can process pressing such a button for a first time period as a request to capture video data while the button is pressed and to stop video capture when the button is released, wherein the video captured while the button is pressed is stored as a single video file. Alternatively, pressing the button for a longer time period can capture a still image. In some examples, interface 312 can be any mechanical switch or physical interface capable of accepting user input associated with requesting data from camera device 308. In other examples, interface 312 can have software components or can be associated with commands received wirelessly from another source, such as client device 328.
[0053] Image processor 306 includes circuitry for receiving signals from camera device 308 and processing those signals into a format suitable for storage in memory 324 or for transmission to client device 328. In one or more examples, image processor 306 (e.g., video processor) includes a microprocessor integrated circuit (IC) customized for processing sensor data from camera device 308, and volatile memory used by the microprocessor in operation.
[0054] The low-power circuit system 304 includes a low-power processor 314 and a low-power wireless circuit system 318. These elements of the low-power circuit system 304 can be implemented as separate components or as part of a single-chip system on a single IC. The low-power processor 314 includes logic for managing other elements of the glasses 200. As described above, for example, the low-power processor 314 can accept user input signals from interface 312. The low-power processor 314 can also be configured to receive input signals or command communications from client device 328 via low-power wireless connection 336. The low-power wireless circuit system 318 includes circuit elements for implementing a low-power wireless communication system, Bluetooth. TM Smart, also known as Bluetooth TM Low power consumption is a standard implementation method for low-power wireless communication systems that can be used to implement the low-power wireless circuit system 318. In other examples, other low-power communication systems can be used.
[0055] The high-speed circuit system 320 includes a high-speed processor 322, a memory 324, and a high-speed wireless circuit system 326. The high-speed processor 322 can be any processor capable of managing high-speed communication and operation of any general-purpose computing system used by the data processor 302. The high-speed processor 322 includes processing resources used by the high-speed wireless circuit system 326 to manage high-speed data transmission over the high-speed wireless connection 334. In some examples, the high-speed processor 322 executes an operating system such as LINUX or another such operating system. Among other responsibilities, the high-speed processor 322, which executes the software architecture of the data processor 302, manages data transmission with the high-speed wireless circuit system 326. In some examples, the high-speed wireless circuit system 326 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as Wi-Fi. In other examples, the high-speed wireless circuit system 326 may implement other high-speed communication standards.
[0056] Memory 324 includes any storage device capable of storing camera data generated by camera device 308 and image processor 306. While memory 324 is shown as integrated with high-speed circuitry 320, in other examples, memory 324 may be a separate, independent element of data processor 302. In some such examples, electrical wiring may provide a connection from image processor 306 or low-power processor 314 to memory 324 via a chip including high-speed processor 322. In other examples, high-speed processor 322 may manage addressing of memory 324 such that low-power processor 314 will bootstrap high-speed processor 322 whenever a read or write operation involving memory 324 is desired.
[0057] The tracking module 340 estimates the pose of the glasses 200. For example, the tracking module 340 uses image data and corresponding inertial data from the camera device 308 and the positioning component, as well as GPS data, to track the position and determine the pose of the glasses 200 relative to a reference frame (e.g., the real-world environment). The tracking module 340 continuously collects and uses updated sensor data describing the motion of the glasses 200 to determine the updated three-dimensional pose of the glasses 200, which indicates changes in the relative position and orientation with respect to physical objects in the real-world environment. The tracking module 340 allows the glasses 200 to be visually positioned relative to physical objects within the user's field of vision via the display 310.
[0058] The GPU and display driver 338 can use the pose of the glasses 200 to generate frames of virtual content or other content to be displayed on the display 310 when the glasses 200 is running in a conventional augmented reality mode. In this mode, the GPU and display driver 338 generate updated frames of virtual content based on the updated 3D pose of the glasses 200, which reflects changes in the user's position and orientation relative to physical objects in the user's real-world environment.
[0059] One or more functions or operations described herein can also be performed in an application residing on glasses 200, client device 328, or a remote server. Glasses 200 can be a standalone client device capable of independent operation, or it can be a companion device that works with a host device to transfer intensive processing and / or exchange data with presentation server system 108 via network 106. Glasses 200 can also be communicatively coupled to companion devices such as smartwatches and can be configured to exchange data with companion devices. Glasses 200 may also include various components common to mobile electronic devices such as smart glasses or smartphones (e.g., a display controller for controlling the display of visual media on a display mechanism incorporated in the device).
[0060] Figure 4This is a block diagram illustrating further details of a volumetric content rendering system 100, based on some examples. Specifically, the volumetric content rendering system 100 is shown as including a rendering client 104 and an application server 112. The volumetric content rendering system 100 includes several subsystems supported on the client side by the rendering client 104 and on the server side by the application server 112. For example, these subsystems include: a collection management system 402, a rendering control system 404, an enhancement system 406, and an immersive experience creation system 408.
[0061] The collection management system 402 is responsible for managing sets or collections of content (e.g., sets of text, images, video, and audio data). These sets of content can be organized into "event libraries" or "event stories." Such sets can be made available for a specified time period (e.g., the duration of an event related to the content). For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 402 can also be responsible for displaying icons that notify the user interface of the client 104 of the existence of a specific set.
[0062] Furthermore, the collection management system 402 includes a curation interface 412, which allows collection managers to manage and curate specific collections of content. For example, the curation interface 412 enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 402 employs machine vision (or image recognition technology) and content rules to automatically curate content collections.
[0063] The presentation control system 404 is responsible for facilitating and controlling the presentation of volumetric content. Therefore, the presentation control system 404 provides a mechanism that allows users to specify control operations for controlling the presentation of volumetric content. For example, control operations may include: a stop operation to halt presentation; a pause operation to pause presentation; a fast-forward operation to advance presentation at a higher speed; a rewind operation to rewind presentation; a zoom-in operation to increase the zoom level of presentation; a zoom-out operation to decrease the zoom level of presentation; and a playback speed modification operation to change the speed of presentation (e.g., generating slow-motion presentation of volumetric video).
[0064] In some implementations, the user can achieve this via one or more I / O components (examples of I / O components are referenced below). Figure 10(Further detailed description) One or more inputs are provided to specify inputs that indicate control actions for controlling the presentation of volumetric content. In some embodiments, the presentation control system 404 may provide an interactive control interface including one or more interactive elements (e.g., virtual buttons) to trigger control actions, and the presentation control system 404 monitors interactions with the interactive interface to detect inputs indicating control actions. In some embodiments, a user may trigger a control action using a gesture (e.g., a hand gesture or head gesture) that can be associated with a specific control action.
[0065] Enhancement system 406 provides various functions that enable users to enhance (e.g., annotate or otherwise modify or edit) media content. For example, enhancement system 406 provides functions related to the generation, publication, and application of enhancement data, such as media overlays on volumetric content (e.g., image filters). Media overlays may include audio and visual content as well as visual effects. Examples of audio and visual content include photographs, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photographs) at client device 102. For example, media overlays may include text or images that can be overlaid on top of a photograph taken by client device 102. Enhancement system 406 provides one or more media overlays to presentation client 104 during operation based on the geographic location of client device 102 or based on other information such as the social network information of the user of client device 102. Media overlays may be stored in database 116 and accessed through database server 114.
[0066] A filter is an example of a media overlay displayed on an image or video during presentation to a user. Filters can be of various types, including user-selected filters from a set of filters presented to the user by presentation client 104. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the user based on geographic location. For example, presentation client 104 can present geolocation filters specific to nearby or particular locations within the user interface based on geographic location information determined by the Global Positioning System (GPS) unit of client device 102.
[0067] Another type of filter is a data filter, which can be selectively presented to the user by the presentation client 104 based on other inputs or information collected by the client device 102. Examples of data filters include the current temperature at a specific location, the current speed of the user's movement, the battery life of the client device 102, or the current time.
[0068] AR content items are another example of media overlay. AR content items can be real-time special effects and / or sounds that can be added to images or videos, including volumetric images and volumetric videos.
[0069] Generally, AR content items, overlays, image transformations, images, and similar terms refer to modifications that can be applied to image data (e.g., videos or images) that include volumetric content. This includes real-time modifications, which modify images as they are captured using the device sensors (e.g., one or more cameras) of client device 102, and then displayed and modified by the display device of client device 102 (e.g., an embedded display of the client device). This also includes modifications to stored content, such as volumetric videos in a gallery or collection that can be modified. For example, in client device 102, which has access to multiple AR content items, a user can use a single video with multiple AR content items to see how different AR content items will modify the stored video. For example, by selecting different AR content items for the content, multiple augmented reality content items with different pseudo-random motion models can be applied to the same content. Similarly, real-time video capture can be used with the illustrated modifications to show how the video image currently captured by the sensors of client device 102 will modify the captured data. Such data can simply be displayed on a screen without being stored in memory, or the content captured by the device sensors can be recorded and stored in memory with or without modification (or both). In some systems, the preview feature can show how different augmented reality content items will be displayed simultaneously in different windows on the monitor. For example, this can make it possible to view multiple windows with different pseudo-random animations on the monitor at the same time.
[0070] Therefore, using augmented reality content items and various systems, or other augmented systems that use augmented data to modify content, can involve detecting objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in video frames, tracking such objects as they leave, enter, and move within the field of view, and modifying or transforming them while tracking them. In various examples, different methods can be used to achieve such transformations. Some examples may involve generating 3D mesh models of one or more objects; and using transformations of the models and animated textures within the video to achieve transformations. In other examples, tracking points on the object can be used to place an image or texture (which can be two-dimensional or three-dimensional) at the tracking location. In yet another example, neural network analysis of video frames can be used to place images, models, or textures within content (e.g., frames of images or videos). Therefore, AR content items refer both to the images, models, and textures used to create transformations within content and to the additional modeling and analysis information required to achieve such transformations using object detection, tracking, and placement.
[0071] Real-time video processing can be performed using any type of video data (e.g., video streams, video files, volumetric video, etc.) stored in the memory of any type of computerized system. For example, a user can load video files and store them in the device's memory, or the device's sensors can be used to generate video streams. Furthermore, computer-animated models can be used to process any object, such as a human face and parts of the human body, animals, or inanimate objects (e.g., chairs, cars, or other objects).
[0072] In some examples, when a specific modification is selected along with content to be enhanced (e.g., edited), the element to be transformed is identified by a computing device, and then, if the element to be transformed exists in a frame of the video, it is detected and tracked. The elements of the object are modified according to the request for modification, thus transforming the frames of the video stream. Different methods can be used to perform frame transformations of the video stream for different types of transformations. For example, for frame transformations that primarily involve changing the form of elements of an object, feature points for each element of the object are calculated (e.g., using an Active Shape Model (ASM) or other known methods). Then, for each element in at least one element of the object, a feature-point-based mesh is generated. This mesh can be used in subsequent stages of tracking the elements of the object in the video stream. In the tracking process, the aforementioned mesh for each element is aligned with the position of each element. Additional points are then generated on the mesh. A first set of points is generated for each element based on the modification request, and a second set of points is generated for each element based on this first set of points and the modification request. The frames of the video stream can then be transformed by modifying the elements of the object based on this first set of points and this second set of points, as well as the mesh. In this method, the background of the object being modified can also be changed or distorted by tracking and modifying the background.
[0073] In some examples, transformations that alter certain regions of an object using its elements can be performed by calculating characteristic points for each element of the object and generating a mesh based on those calculated characteristic points. Points are generated on the mesh, and then various regions are generated based on those points. The elements of the object are then tracked by aligning the regions of each element with the positions of at least one of the elements, and the properties of the regions can be modified based on modification requests, thereby transforming frames of the video stream. Depending on the specific modification request, the properties of the mentioned regions can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least some portions of the region from the frames of the video stream; including one or more new objects into the region based on the modification request; and modifying or distorting the elements of the region or object. In various examples, any combination of such modifications or other similar modifications can be used. For certain models to be animated, some characteristic points can be selected as control points to determine the entire state space of the model's animation options.
[0074] In some examples of computer animation models that use face detection to transform image data, a specific face detection algorithm (such as Viola-Jones) is used to detect faces in the image. Then, the Active Shape Model (ASM) algorithm is applied to the facial regions of the image to detect facial feature reference points.
[0075] Other methods and algorithms suitable for face detection can be used. For example, in some examples, landmarks are used to locate features; landmarks represent distinguishable points present in most of the images considered. For example, for facial markers, the location of the left pupil could be used. If the initial landmarks are not recognizable (e.g., if the person is wearing an eye patch), secondary landmarks can be used. Such a boundary recognition process can be used for any such object. In some examples, a set of landmarks forms a shape. The shape can be represented as a vector using the coordinates of the points in the shape. One shape is aligned with another shape using a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the points of the shape. The meanshape is the average of the training shapes used for alignment.
[0076] In some examples, a search for landmarks begins with an average shape aligned with the position and size of a face determined by a global face detector. This search then repeats the following steps: proposing provisional shapes by adjusting the positions of shape points using template matching of the image texture around each point, and then conforming the provisional shapes to a global shape model until convergence occurs. In some systems, individual template matching is unreliable, and the shape model pools the results of weak template matching to form a stronger overall classifier. The entire search is repeated at each level of the image pyramid, from coarse to fine resolution.
[0077] The enhancement system 406 can capture image or video streams on a client device (e.g., client device 102) and perform complex image manipulations locally on client device 102 while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, mood shifts (e.g., changing a face from frowning to smiling), state shifts (e.g., aging an object, reducing its apparent age, or changing its gender), style shifts, application of graphical elements, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on client device 102.
[0078] In some examples, enhancement system 406 may use a computer animation model for transforming video and image content, wherein the neural network operates as part of a presentation client 104 operating on client device 102. Enhancement system 406 determines the presence of faces within an image or video stream and provides interactive modification elements (e.g., icons) associated with the computer animation model used to transform the image data, or the computer animation model may exist as an interface described herein. Interactive modification elements include changes to the underlying user's face within the image or video content that can be modified as part of a modification operation. Once an interactive modification element is selected, the transformation system initiates a process to transform the user's image to reflect the selected interactive modification element (e.g., generating a smiley face on the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video content can be presented in a graphical user interface displayed on client device 102. Enhancement system 406 may implement complex convolutional neural networks on a portion of the image or video content to generate and apply the selected modification. That is, the modified content can be presented to the user in real-time or near real-time. Furthermore, the modification can persist while the content is being presented. Machine learning neural networks can be used to implement such modifications.
[0079] The immersive experience creation system 408 is responsible for creating immersive AR experiences based on existing images and videos. In doing so, the immersive experience creation system 408 can utilize one or more known machine learning or artificial intelligence image processing techniques to segment images and videos to identify key elements. The immersive experience creation system 408 generates volumetric content items for each key element that can be utilized in creating the immersive AR experience. Further details regarding the creation and presentation of immersive AR experiences are discussed below.
[0080] Figure 5 This is a diagrammatic representation of the data structures maintained in database 116, based on some examples. Although the contents of database 116 are shown as including several tables, it should be understood that the data can be stored in other types of data structures (e.g., as an object-oriented database).
[0081] Entity table 506 stores entity data and (e.g., by reference) links to entity diagram 508 and profile data 502. Entities whose records are maintained within entity table 506 can include individuals, company entities, organizations, objects, locations, events, etc. Regardless of entity type, any entity whose data is stored in presentation server system 108 can be an identifiable entity. Each entity is provided with a unique identifier and an entity type identifier (not shown). Entity table 506 allows various enhancements from enhancement table 510 to be associated with various images and videos stored in image table 514 and video table 512.
[0082] Entity graph 508 stores the relationships and associated information between entities. Such relationships can be, for example, social, professional (e.g., working in the same company or organization), based on interests, or based on activities.
[0083] Profile data 502 stores various types of profile data about a specific entity. Based on the privacy settings specified by the specific entity, profile data 502 can be selectively used and presented to other users of the volume content presentation system 100. In the case that the entity is an individual, profile data 502 includes, for example, a username, phone number, address, settings (e.g., notification and privacy settings), and an avatar (or a set of such avatars) selected by the user.
[0084] Database 116 also stores augmented data (such as overlays of AR content items and filters) in augmentation table 510. The augmented data is associated with and applied to videos (whose data is stored in video table 512) and images (whose data is stored in image table 514), including volumetric videos and volumetric images.
[0085] Story table 516 stores data about collections of content, including associated image, video, or audio data, compiled into collections (e.g., stories or libraries). The creation of a specific collection can be initiated by a specific user (e.g., each user whose records are maintained in entity table 506). A user can create a "personal story" in the form of a collection of content that has already been created and sent / broadcast by that user. For this purpose, the user interface of client 104 may include an icon that is selectable by the user, allowing the sending user to add specific content to his or her personal story.
[0086] As mentioned above, video table 512 stores video data including volumetric videos. Similarly, image table 514 stores image data including volumetric images.
[0087] Figure 6 This is a conceptual diagram illustrating an example in which a volumetric content rendering system 100 identifies one or more two-dimensional elements from a two-dimensional image 610 to be presented as part of an augmented reality experience. As shown, the two-dimensional image 610 includes multiple two-dimensional elements: a sun 602, two birds 604, a cloud 606, a dog 608, grass, and trees. To individually identify elements in the two-dimensional image 610, the volumetric content rendering system 100 can perform image segmentation on the two-dimensional image 610 using various known segmentation algorithms. Furthermore, in this example, the cloud 606 is identified based on user input.
[0088] The volumetric content rendering system 100 generates volumetric content items based on two-dimensional elements identified in the two-dimensional image 610. For example, the volumetric content rendering system 100 generates a volumetric content item 612 based on a sun 602 identified in the two-dimensional image 610. When generating the volumetric content item, the volumetric content rendering system 100 can adjust the size of the volumetric content item based on image analysis, user input, or automatically. That is, the volumetric content rendering system 100 can utilize one or more scaling techniques when generating the volumetric content item 612.
[0089] Figure 7A and Figure 7B This is a concept diagram illustrating an example augmented reality experience provided to user 702 on user device 704. User device 704 is... Figure 1 Example of client device 102. Figure 7A This illustrates an example viewpoint of a user standing in a real-world environment, which is the room in this example. As shown, several real-world elements in the user's viewpoint include the floor, walls, windows, and bookshelves.
[0090] Figure 7B An example view of user 702 is shown after the volumetric content rendering system 100 renders volumetric content items as overlays onto the real-world environment on user device 704. Volumetric content items representing grass are rendered as overlays on the ground surface, volumetric content items representing the sun, trees, and clouds are rendered as overlays on the wall, and a volumetric content item representing a dog 608 is placed as a "clone" in the AR on the floor of the real-world environment (i.e., the room), thus providing the user with an AR experience 702. The size of the volumetric content items 612 has also been scaled (e.g., adjusted) to fit the real-world environment.
[0091] Figure 8A This is a flowchart illustrating the operation of a volumetric content rendering system when performing a method for providing an immersive AR experience based on existing image data, according to some examples. Method 800 can be implemented with computer-readable instructions for execution by one or more processors, such that the operation of method 800 can be performed partially or wholly by functional components of volumetric content rendering system 100; therefore, method 800 is described below by way of example with reference to it. However, it should be understood that at least some of the operation of method 800 can be deployed on various other hardware configurations besides volumetric content rendering system 100.
[0092] At operation 802, the volumetric content rendering system 100 accesses and includes image data comprising one or more two-dimensional images. The image data may include any one or more physical copies of digital images, video frames, or photographs. The image data may be stored in and accessed from a database (e.g., database 116) or the memory of a wearable device (e.g., client device 102) or mobile device.
[0093] In the first example, the image data includes a two-dimensional image depicting a puppy running in a grassy field. The image data is stored in the memory component of the client device 102. The volumetric content rendering system 100 accesses the image data from the memory component based on user input.
[0094] In the second example, the image data includes a two-dimensional image based on a physical copy of a photograph depicting a running puppy. The two-dimensional image can be generated by capturing a photograph that is a physical copy of the running puppy, in preparation for further operations.
[0095] At operation 804, the volumetric content rendering system 100 recognizes two-dimensional elements depicted in one or more two-dimensional images. These two-dimensional image elements can correspond to objects in the virtual or real world. For example, two-dimensional image elements can correspond to animals, people, plants, landscapes, buildings, the sun, the moon, or the sky.
[0096] In some implementations, to identify two-dimensional image elements, the volumetric content rendering system 100 performs image segmentation on the image data to identify and extract the two-dimensional elements. When performing image segmentation, the volumetric content rendering system 100 may utilize any one or more known digital image processing and computer vision techniques, including known machine learning and artificial intelligence image segmentation and classification techniques. In some implementations, to segment or extract two-dimensional image elements from one or more two-dimensional images, the volumetric content rendering system 100 relies on user input to delineate at least some of the two-dimensional elements in the image.
[0097] At operation 806, the volumetric content rendering system 100 generates volumetric content items based on two-dimensional elements identified from the image data. The volumetric content item includes a volumetric representation of the two-dimensional elements identified from the two-dimensional image. The format of the volumetric content item is compatible with the format used for rendering on a client device, such that the volumetric content item appears to the user using the client device as part of and / or superimposed on their surrounding environment (e.g., a real-world environment). Returning to the example of a puppy running in grass, the puppy, grass, and sky are identified two-dimensional image elements. Volumetric content items are created for each or more of the puppy, grass, and sky.
[0098] Therefore, in some instances, a set of two-dimensional elements can be identified from image data, and volumetric content items can be generated only for a subset of these elements. In some implementations, the volumetric content rendering system 100 can automatically perform the selection of one or more specific two-dimensional elements from the two-dimensional image elements identified through image segmentation. Returning to the example of the image of a puppy running in the grass, although all of the puppy, grass, and blue sky can be identified based on image segmentation, the volumetric content rendering system 100 can select only the puppy for the creation of the volumetric content item.
[0099] In some implementations, to select a specific two-dimensional element from a plurality of identified two-dimensional elements, the volumetric content rendering system 100 provides an interactive interface operable to receive a selection of one of a plurality of two-dimensional image elements for generating a volumetric content item. For example, the volumetric content rendering system 100 may enable a display device to present the interactive interface. The volumetric content rendering system 100 may receive input indicating the selection of two-dimensional image elements for generating a volumetric content item based on user interaction with the interactive interface. The interactive interface presents a plurality of two-dimensional image elements segmented from image data and allows the user to select one or more two-dimensional elements for generating a volumetric content item. In an example, the image data includes a two-dimensional image depicting a road full of cars. The volumetric content rendering system 100 may identify each car as a segmented element of the image. In this example, the interactive interface presents the image in such a way that the user can select a subset of cars on the road as the basis for generating one or more volumetric content items.
[0100] At operation 808, the volumetric content rendering system 100 causes the display device to render volumetric content items as overlaid on the real-world environment visible to the user of the display device. In this way, the rendering of volumetric content items makes them appear as if they are in a real-world environment.
[0101] In some implementations, volumetric content items are rendered as overlays on the real-world environment based on the association between the volumetric content item and an element (or surface) in the real-world environment. Consistent with these implementations, the reasons for rendering the volumetric content item include: identifying or detecting real-world elements in the real-world environment, and generating or identifying the association between the volumetric content item and the identified real-world element. Based on the association between the volumetric content item and the identified real-world element, the volumetric content rendering system 100 renders the volumetric content item as overlays on the real-world environment. In some implementations, the volumetric content rendering system 100 causes a display device to render the volumetric content item as overlays on the real-world environment based on the location associated with a real-world element (e.g., the location of the real-world element in the real-world environment). Real-world elements include real-world objects or people. Some example real-world elements are floors, walls, and tables. By rendering the volumetric content item as overlays on associated real-world elements, the realism and immersion of the AR experience are enhanced. For example, a volumetric content item representing grass can be associated with a ground surface and rendered as an overlay on the ground to provide an AR experience in which a user stands on the grass and can walk on the grass.
[0102] To generate associations between volumetric content items and identified real-world elements, the volumetric content rendering system 100 can provide an interactive interface operable to receive these associations. For example, the volumetric content rendering system 100 can cause a display device to present the interactive interface. The volumetric content rendering system 100 receives input from the interactive interface indicating a selection of the association between the volumetric content item and the real-world element.
[0103] In some implementations, the rendering of a volumetric content item can be automatically triggered in response to detecting that the location of the wearable device is within a threshold distance of a location depicted by one or more images. For example, the image data may include one or more images depicting a first location. For example, the volumetric content rendering system 100 may determine the first location based on metadata associated with the image data. The volumetric content rendering system 100 may determine a second location associated with a display device (e.g., based on location data from the display device, a complementary device of its primary device, or a complementary device of the display device). In response to determining that the second location is within a threshold distance of the first location, the volumetric content rendering system 100 may render the volumetric content item.
[0104] In some implementations, the rendering of volumetric content items may include presenting visual effects along with the volumetric content items. For example, a volumetric content rendering system 100 may acquire metadata included in or associated with image data and determine the visual effects to be presented along with the volumetric content items based on the acquired metadata. The metadata provides information about the image data, including information about the real-world environment associated with the image data. The real-world environment associated with the image data may be the same real-world environment in which the volumetric content items are overlaid, or a different real-world environment. In some examples, the information may include weather conditions, date, time of day, people, and locations. The information corresponds to the creation time of the image data. In a specific example, a two-dimensional image depicts a dog running in the rain. The metadata associated with the image includes weather data indicating the rain conditions at the time the image was created. The metadata may include machine-generated data and / or user input. The metadata may be acquired from multiple sources, including image data, one or more camera devices that generated the image data, and third-party servers.
[0105] Visual effects are determined based on the acquired metadata and rendered along with volumetric content items. Visual effects include one or more enhancements to be applied to the real-world environment. Visual effects may include animations. In the example, a rain animation is rendered along with volumetric content items, based on weather data from metadata indicating rain conditions.
[0106] The volumetric content rendering system 100 applies visual effects generated by a display device along with volumetric content items to a real-world environment. The rendering of these visual effects makes the real-world environment appear to the user as if it were in a situation similar to a two-dimensional image being captured. In the example, the real-world environment is an indoor office, and metadata indicates rainy weather conditions. The visual effects, including a rain animation, are rendered as an overlay on the view of a user in the indoor office, along with the volumetric content items.
[0107] In some implementations, the rendering of volumetric content items includes the rendering of animation effects associated with the volumetric content items based on defined visual effects. The animation effects associated with volumetric content items include visual enhancements that make the volumetric content items appear more consistent with the metadata-generated visual effects. In an example, the rendering of volumetric representations of a dog, grass, and sky is overlaid on a real-world environment along with a rain animation visual effect. In this example, the volumetric content rendering system 100 can render three animation effects combined with the volumetric representations of the dog, grass, and sky. The animation effect associated with the dog's volumetric representation makes the dog appear wet; the animation effect associated with the grass's volumetric representation makes the grass appear wet and start to form small ponds; and the animation effect associated with the sky's volumetric representation makes the sky appear darker. As another example, when a windy visual effect is rendered based on metadata, the animation effect associated with the dog's volumetric representation makes an animated wind blow through the dog's fur.
[0108] In some implementations, the rendering of volumetric content items includes determining real-world effects and enabling network-connected devices (e.g., IoT devices or other smart devices) to provide real-world effects. Real-world effects can be associated with functionality provided by the network-connected devices. For example, example network-connected devices include smart fans, smart thermostats, and smart light bulbs. Realistic effects are created by enabling network-connected devices in a real-world environment to perform certain functions to create conditions that can be associated with visual effects in a real-world environment. Real-world effects can be determined based on metadata associated with image data. For example, windy conditions can be a real-world effect determined based on metadata associated with image data (e.g., weather data included in the metadata). Based on the determined windy conditions associated with the image data, the volumetric content rendering system 100 can activate a smart fan in the real-world environment to achieve a wind-like condition in the real-world environment, thereby adding a realistic effect to the rendering of the volumetric content item. In another example, a sunny condition associated with the image data is determined based on metadata. In response, the volumetric content rendering system 100 can activate a smart thermostat to operate at a high temperature and cause a smart light in the real-world environment to change its color to a bright yellow.
[0109] In some implementations, the rendering of volumetric content items includes the rendering of one or more virtual objects combined with volumetric content items overlaid on a real-world environment. Virtual objects are imaginary, fictional, or computer-generated objects that no longer exist in the real world. For example, a virtual object could be a dinosaur, a fire-breathing dragon, a zombie, or a wizard. Volumetric virtual items generated based on virtual objects can be stored in a library of volumetric virtual items (e.g., in database 116).
[0110] At operation 810, the volumetric content rendering system 100 generates volumetric content based on the rendering of volumetric content items overlaid on the real-world environment. The volumetric content may also include audio data containing one or more audio signals, and therefore the rendering of the volumetric content may include the rendering of one or more audio signals. The new volumetric content includes: volumetric content items, one or more audio signals, and one or more elements of the real-world environment. In some instances, the volumetric content includes volumetric representations of the volumetric content items and one or more elements of the real-world environment. At operation 812, the volumetric content rendering system 100 stores the volumetric content for subsequent rendering.
[0111] like Figure 8B As shown, in some embodiments, method 800 may further include operation 814. Consistent with these embodiments, operation 814 may be performed as part of the volumetric content rendering system 100 in operation 804 recognizing two-dimensional elements depicted in one or more two-dimensional images. As described above, the two-dimensional image elements may correspond to virtual or real-world objects, such as animals, people, plants, landscapes, buildings, the sun, the moon, or the sky. At operation 814, the volumetric content rendering system 100 performs image segmentation on the image data to recognize the two-dimensional elements. When performing image segmentation, the volumetric content rendering system 100 may utilize any one or more known digital image processing and computer vision techniques, including known machine learning and artificial intelligence image segmentation and classification techniques.
[0112] like Figure 8C As shown, in some embodiments, method 800 may further include operations 816 and 818. Consistent with these embodiments, operations 816 and 818 may be performed as part of the volumetric content rendering system 100 in operation 804 recognizing two-dimensional elements depicted in one or more two-dimensional images. At operation 816, the volumetric content rendering system 100 provides an interactive interface that enables a user to recognize two-dimensional elements from image data. Thus, the interactive interface may include the presentation of image data with one or more interactive elements, allowing the user to select and / or specify two-dimensional elements described in the image data. When providing the interactive interface, the volumetric content rendering system 100 may use a user device (e.g., client device 102) to present the interactive interface.
[0113] At operation 818, the volumetric content rendering system 100 receives user-recognized input indicating two-dimensional elements. In some examples, the input is generated by the user 124 using touch, gesture, or one or more interactive elements to outline two-dimensional elements in an image.
[0114] like Figure 8DAs shown, in some embodiments, method 800 may further include operations 820, 822, and 824. Consistent with these embodiments, operations 820, 822, and 824 may be performed as part of operations 806. At operation 820, the volumetric content rendering system 100 identifies real-world elements in the real-world environment. At operation 822, the volumetric content rendering system 100 associates a volumetric content item with a real-world element. At operation 824, the volumetric content rendering system 100 renders the volumetric content item based on the association with the real-world element. In one example, the volumetric content item is rendered at a predefined distance from the real-world element based on the association. In another example, the volumetric content item is rendered as an overlay on the real-world element based on the association. When rendering the volumetric content item, the volumetric content rendering system 100 may perform one or more scaling operations to adjust the size of the volumetric content item. The adjustment may be determined automatically based on the characteristics of the real-world environment and the volumetric content item, or it may be determined based on user input. As an example, the size of volumetric content items can be adjusted (scaled) to match the size of real-world elements.
[0115] By returning to the reference Figure 6 and Figure 7A as well as Figure 7B In the example shown, the volume content rendering system 100 recognizes the wall 708 as a real-world element (at operation 820), associates the volume content item 612 with the wall 708 (at operation 822), and renders the volume content item 612 as overlaid on the wall 708 (at operation 824).
[0116] Figure 9 This is a block diagram illustrating an example software architecture 920 that can be used in conjunction with various hardware architectures described herein. Figure 9 This is a non-limiting example of a software architecture, and it will be recognized that many other architectures can be implemented to facilitate the functionality described herein. Software architecture 920 can be implemented in, for example... Figure 10 The execution is performed on the hardware of machine 1000, which includes processor 1004, memory / storage device 1006, and I / O components 1018, etc. A representative hardware layer 928 is shown and can represent, for example... Figure 10 The machine 1000. A representative hardware layer 928 includes a processing unit 922 having associated executable instructions 930. The executable instructions 930 represent executable instructions of the software architecture 920, including implementations of the methods, components, etc., described herein. Hardware layer 928 also includes a memory and / or storage module 924, which also has executable instructions 930. Hardware layer 928 may also include other hardware 926.
[0117] exist Figure 9 In the exemplary architecture shown, software architecture 920 can be conceptualized as a layer stack, where each layer provides specific functionality. For example, software architecture 920 may include layers such as operating system 940, libraries 918, framework / middleware 910, applications 904, and a presentation layer 902. Operationally, applications 904 and / or other components within a layer can invoke API calls 942 via the software stack and receive responses to API calls 942 as messages 944. The layers shown are representational in nature, and not all software architectures have all layers. For example, some mobile or dedicated operating systems may not provide a framework / middleware 910, while others may provide such a layer. Other software architectures may include additional or different layers.
[0118] Operating system 940 can manage hardware resources and provide public services. For example, operating system 940 may include core 934, server 936, and driver 938. Core 934 can act as an abstraction layer between hardware and other software layers. For example, core 934 can be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. Server 936 can provide other public services to other software layers. Driver 938 is responsible for controlling or connecting to the underlying hardware. For example, depending on the hardware configuration, driver 938 may include display driver, camera driver, etc. Drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Drivers, audio drivers, power management drivers, etc.
[0119] Library 918 provides common infrastructure used by application 904 and / or other components and / or layers. Library 918 provides functionality that allows other software components to perform tasks more easily than by directly interfacing with the underlying operating system 940 functions (e.g., kernel 934, server 936, and / or driver 938). Library 918 may include system libraries 912 (e.g., the C standard library), which provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, library 918 may include API libraries 914, such as media libraries (e.g., libraries supporting the rendering and manipulation of various media formats such as MPEG4, H.2104, MP3, AAC, AMR, JPG, and PNG), graphics libraries (e.g., OpenGL frameworks that can be used to render 2D and 3D graphics content on a display), database libraries (e.g., SQLite that provides various relational database functions), web libraries (e.g., WebKit that provides web browsing functionality), etc. Library 918 may also include various other libraries 916 to provide many other APIs to application 904 and other software components / modules.
[0120] The framework / middleware 910 provides higher-level common infrastructure that can be used by application 904 and / or other software components / modules. For example, the framework / middleware 910 can provide various GUI functions, advanced resource management, advanced location services, etc. The framework / middleware 910 can provide a wide range of other APIs that can be utilized by application 904 and / or other software components / modules, some of which are specific to a particular operating system 940 or platform.
[0121] Application 904 includes built-in applications 906 and / or third-party applications 908. Examples of representative built-in applications 906 may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, and / or game applications. Third-party applications 908 may include applications using Android by entities other than the platform-specific vendor. TM or iOS TM Applications developed using a Software Development Kit (SDK) can be used on platforms such as iOS. TM ANDROID TM , Mobile software running on the phone's mobile operating system or other mobile operating systems. Third-party applications 908 can call API calls 942 provided by the mobile operating system (such as operating system 940) to facilitate the functions described herein.
[0122] Application 904 can use built-in operating system functions (e.g., core 934, server 936, and / or driver 938), libraries 918, and frameworks / middleware 910 to create a user interface for interacting with the system's user. Alternatively or additionally, in some systems, user interaction may occur through a presentation layer such as presentation layer 902. In these systems, the application / component "logic" can be separated from the application / component's user-interacting aspects.
[0123] Figure 10 This is a block diagram illustrating components of a machine 1000, according to some example embodiments, capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and executing any or more of the methods discussed herein. Specifically, Figure 10 A graphical representation of a machine 1000 in the form of an example computer system is shown, within which instructions 1010 (e.g., software, programs, applications, applets, or other executable code) can be executed to cause the machine 1000 to perform any or more of the methods discussed herein. Therefore, instructions 1010 can be used to implement the modules or components described herein. Instructions 1010 transform a general, unprogrammed machine 1000 into a specific machine 1000 programmed to perform the described and illustrated functions in the described manner. In alternative embodiments, machine 1000 operates as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1000 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1000 may include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 1010 specifying actions to be taken by machine 1000. Furthermore, although only a single machine 1000 is shown, the term "machine" should also be considered to include a collection of machines that individually or jointly perform method 800 to perform any one or more of the methods discussed herein.
[0124] Machine 1000 may include processor 1004, memory / storage device 1006, and I / O unit 1018, which may be configured to communicate with each other, for example, via bus 1002. In an example embodiment, processor 1004 (e.g., CPU, Reduced Instruction Set Computing (RISC) processor, Complex Instruction Set Computing (CISC) processor, GPU, Digital Signal Processor (DSP), ASIC, Radio Frequency Integrated Circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 1008 and processor 1016 capable of executing instructions 1010. Although Figure 10 Multiple processors 1004 are shown, but machine 1000 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0125] Memory / storage device 1006 may include memory 1012, such as main memory or other storage devices, and storage cells 1014, which processor 1004 can access, for example, via bus 1002. Storage cells 1014 and memory 1012 store instructions 1010 embodying any one or more of the methods or functions described herein. During execution of instructions 1010 by machine 1000, instructions 1010 may also reside wholly or partially within memory 1012, within storage cells 1014, within at least one processor of processor 1004 (e.g., within the processor's cache memory), or any suitable combination thereof. Therefore, the memories of memory 1012, storage cells 1014, and processor 1004 are examples of machine-readable media.
[0126] I / O component 1018 may include a wide variety of components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, etc. The specific I / O component 1018 included in a particular machine 1000 will depend on the type of machine. For example, a portable machine such as a mobile phone is likely to include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It will be appreciated that I / O component 1018 may include... Figure 10Many other components are not shown. The I / O components 1018 are grouped by function only for the sake of simplicity in the following discussion, and this grouping is by no means limiting. In various example embodiments, the I / O components 1018 may include output components 1026 and input components 1028. Output components 1026 may include visual components (e.g., displays, such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), auditory components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. Input components 1028 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreen displays that provide position and / or force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.
[0127] In other example implementations, I / O component 1018 may include biometric component 1030, motion component 1034, environmental component 1036 or positioning component 1038, and various other components. For example, biometric component 1030 may include components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 1034 may include: accelerometer component (e.g., accelerometer), gravity sensor component, rotation sensor component (e.g., gyroscope), etc. Environmental component 836 may include, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas sensor that detects the concentration of hazardous gases for safety or measures pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. Positioning component 1038 may include a position sensor component (e.g., a Global Positioning System (GPS) receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure to obtain altitude), an orientation sensor component (e.g., a magnetometer), etc.
[0128] A wide variety of technologies can be used to implement communication. I / O component 1018 may include communication component 1040, which is operable to couple machine 1000 to network 1032 or device 1020 via coupling 1024 and coupling 1022, respectively. For example, communication component 1040 may include network interface component or other suitable device to interface with network 1032. In further examples, communication component 1040 may include wired communication component, wireless communication component, cellular communication component, near field communication (NFC) component, Bluetooth component (e.g., Bluetooth Low Energy), Wi-Fi component, and other communication components that provide communication via other modalities. Device 1020 may be another machine or any peripheral device from a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0129] Furthermore, the communication component 1040 can detect identifiers or may include components operable to detect identifiers. For example, the communication component 1040 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes, such as Universal Product Code (UPC) barcodes; multi-dimensional barcodes, such as Quick Response (QR) codes, Aztec codes, data matrices, dataglyphs, MaxiCodes, PDF4114, supercodes, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying audio signals of the tag). Additionally, various information can be obtained via the communication component 1040, such as location obtained via Internet Protocol (IP) geolocation, location obtained via Wi-Fi signal triangulation, location obtained by detecting NFC beacon signals that can indicate a specific location, etc.
[0130] Glossary
[0131] In this context, "carrier signal" refers to any intangible medium capable of storing, encoding, or carrying instructions to be executed by a machine, and includes digital or analog communication signals or other intangible media to facilitate the communication of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device and employing any of a number of well-known transmission protocols.
[0132] In this context, "client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, PDAs, smartphones, tablet computers, ultrabooks, netbooks, laptop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics systems, game consoles, set-top boxes, or any other communication device that a user can use to access the network.
[0133] In this context, "communication network" refers to one or more parts of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a POTS (Plain Old-Style Telephone Service) network, a cellular telephone network, a wireless network, etc. A network, another type of network, or a combination of two or more such networks. For example, a network or part of a network may include a wireless or cellular network, and the coupling to the network may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, the coupling can implement any data transmission technology of various types, such as Single Carrier Radio Transmission (1xRTT), Evolved Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data Rate Evolution of GSM (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other standards defined by various standards-setting organizations, other telematics protocols, or other data transmission technologies.
[0134] In this context, "machine-readable medium" means a component, device, or other tangible medium capable of temporarily or permanently storing instructions and data, and may include, but is not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage devices (e.g., erasable programmable read-only memory (EPROM)) and / or any suitable combination thereof. The term "machine-readable medium" should be considered to include a single medium or multiple media capable of storing instructions (e.g., a centralized or distributed database or associated cache and server). The term "machine-readable medium" should also be considered to include any medium or combination of media capable of storing instructions (e.g., code) executable by a machine, such that when executed by one or more processors of the machine, the instructions cause the machine to perform any or more of the methods described herein. Therefore, "machine-readable medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network comprising multiple storage devices or apparatuses. The term "machine-readable medium" does not include signals themselves.
[0135] In this context, a "component" refers to a logical or physical entity that has boundaries defined by functional or subroutine calls, branch points, APIs, or other technologies that provide partitioning or modularity for specific processing or control functions. Components can interface with other components via their interfaces to perform machine processes. Components can be encapsulated functional hardware units designed for use with other components, as well as part of a program that typically performs a specific function related to that function. Components can constitute software components (e.g., code embodied on a machine-readable medium) or hardware components.
[0136] A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in a physical manner. In various example implementations, one or more hardware components (e.g., standalone computer systems, client computer systems, or server computer systems) of a computer system, or one or more hardware components (e.g., processors or processor groups) of a computer system, can be configured by software (e.g., an application or application portion) to operate to perform certain operations as described herein. Hardware components can also be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as a field-programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry systems temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor.
[0137] Once configured via such software, the hardware component becomes a specific machine (or a specific part of a machine) uniquely tailored to perform the configured function, rather than a general-purpose processor. It will be understood that the hardware component may be implemented mechanically in a dedicated and permanently configured circuit system or in a temporarily configured (e.g., software-configured) circuit system, for cost and time considerations. Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include tangible entities, i.e., entities physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate or perform certain operations described herein.
[0138] Consider implementations where hardware components are temporarily configured (e.g., programmed), eliminating the need to configure or instantiate each hardware component at any given time. For example, where the hardware components include a general-purpose processor that is configured as a dedicated processor via software, the general-purpose processor can be configured as different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures one or more specific processors to constitute a specific hardware component at one time and different hardware components at different times.
[0139] Hardware components can provide information to and receive information from other hardware components. Accordingly, the described hardware components can be considered communicatively coupled. In the presence of multiple hardware components, communication can be achieved through signal transmission between two or more hardware components (e.g., via appropriate circuitry and buses). In embodiments where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, by storing information in a memory structure accessed by the multiple hardware components and retrieving information from said memory structure. For example, one hardware component can perform an operation and store the output of that operation in a communicatively coupled memory device. Another hardware component can then access the memory device at a subsequent time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., information collection).
[0140] The various operations of the example methods described herein can be performed, at least in part, by one or more processors configured, either temporarily (e.g., by software) or permanently, to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute a component of a processor implementation that performs operations to execute one or more of the operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented, at least in part, by processors, wherein one or more specific processors are examples of hardware. For example, at least some of the operations of the methods can be performed by one or more processors or processor-implemented components.
[0141] Furthermore, one or more processors may operate to support the execution of related operations in a “cloud computing” environment or as “Software as a Service” (SaaS). For example, at least some operations may be executed by a group of computers (as an example of a machine including processors), wherein these operations are accessible via a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application programming interfaces (APIs)). The execution of some operations may be distributed among processors, rather than residing within a single machine, but deployed across multiple machines. In some example implementations, the processor or processor-implemented components may be located in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other example implementations, the processor or processor-implemented components may be distributed across multiple geographic locations.
[0142] In this context, "processor" refers to any circuit or virtual circuit (physical circuitry simulated by logic executed on an actual processor) that manipulates data values according to control signals (e.g., "commands," "opcodes," "machine codes," etc.) and generates corresponding output signals applied to operate a machine. For example, a processor can be a CPU, RISC processor, CISC processor, GPU, DSP, ASIC, RFIC, or any combination thereof. A processor can also be a multi-core processor with two or more independent processors (sometimes called "cores") that can execute instructions simultaneously.
[0143] In this context, a "timestamp" refers to a sequence of characters or encoded information that identifies when an event occurred (e.g., giving a date and time of day), sometimes accurate to a fraction of a second.
Claims
1. A method comprising: Access image data that includes one or more two-dimensional images; By performing image segmentation, multiple two-dimensional elements are identified from one or more two-dimensional images; Select a two-dimensional element from the plurality of two-dimensional elements; A volume content item is generated based on the two-dimensional elements identified from the one or more two-dimensional images, the volume content item representing the two-dimensional element; Identify real-world elements in a real-world environment; Associate the volume contents with the real-world elements in the real-world environment; An interactive interface is presented, which includes the volume content item and one or more real-world elements detected in the real-world environment. The interactive interface is operable to receive the association between the volume content item and the real-world elements. Receive input indicating the association between the volume content item and the real-world element; as well as This enables the display device to present the volumetric content item as an overlay on the real-world environment at the location of the identified real-world element, based on the association, and the real-world environment is within the user's field of view of the display device.
2. The method according to claim 1, wherein, The interactive interface is a second interactive interface, and the method further includes: Provides a first interactive interface operable to receive user recognition of the two-dimensional elements from the one or more two-dimensional images; and Receive input indicating the user identification of the two-dimensional element, wherein the identification of the two-dimensional element is based on the input.
3. The method according to claim 1, wherein, Identifying the two-dimensional elements from the one or more two-dimensional images includes performing image segmentation on the one or more two-dimensional images.
4. The method according to claim 1, wherein, Presenting the volumetric content item as an overlay on the real-world environment includes: This causes one or more virtual objects to be presented together with the volume contents.
5. The method according to claim 4, wherein, The one or more virtual objects include one or more imagined computer-generated objects.
6. The method according to claim 1, further comprising: Obtain metadata associated with the one or more two-dimensional images; Visual effects are generated based on the metadata, and the visual effects include one or more enhancements to be applied to the real-world environment; as well as This enables the visual effects applied to the real-world environment to be displayed via a display device.
7. The method according to claim 6, wherein, The presentation of the visual effects applied to the real-world environment includes rendering animation effects together with the volumetric content items.
8. The method according to claim 1, further comprising: Obtain metadata associated with the one or more two-dimensional images; The real-world effects are determined based on the metadata, and these real-world effects are associated with the functionality provided by the network-connected devices. as well as This enables the network-connected device in the real-world environment to provide the real-world effect.
9. The method according to claim 1, further comprising: Volumetric content is generated by overlaying the volumetric content items onto the real-world environment. as well as The volume content is stored for subsequent presentation.
10. The method according to claim 1, wherein: The one or more images depict the first location; as well as The method further includes: Determine a second location associated with the display device; Determine that the second position is within a threshold distance of the first position; and In response to determining that the second position is within the threshold distance of the first position, the volumetric content item is rendered.
11. A system comprising: One or more hardware processors; as well as At least one memory storing instructions that cause the one or more hardware processors to perform operations, the operations including: Access image data that includes one or more two-dimensional images; By performing image segmentation, multiple two-dimensional elements are identified from one or more two-dimensional images; Select a two-dimensional element from the plurality of two-dimensional elements; A volume content item is generated based on the two-dimensional elements identified from the one or more two-dimensional images; Identify real-world elements in a real-world environment; Associate the volume contents with the real-world elements in the real-world environment; An interactive interface is presented, which includes the volume content item and one or more real-world elements detected in the real-world environment. The interactive interface is operable to receive the association between the volume content item and the real-world elements. Receive input indicating the association between the volume content item and the real-world element; and This enables the display device to present the volumetric content item as an overlay on the real-world environment at the location of the identified real-world element, based on the association, and the real-world environment is within the user's field of view of the display device.
12. The system according to claim 11, wherein, The interactive interface is a second interactive interface, and the operation further includes: Provide a first interactive interface, operable to receive user recognition of the two-dimensional elements from the one or more two-dimensional images; and Receive input indicating the user identification of the two-dimensional element, wherein the identification of the two-dimensional element is based on the input.
13. The system according to claim 11, wherein, Identifying the two-dimensional elements from the one or more two-dimensional images includes performing image segmentation on the one or more two-dimensional images.
14. The system according to claim 11, wherein, Presenting the volumetric content item as an overlay on the real-world environment includes: This causes one or more virtual objects to be presented together with the volume contents.
15. The system according to claim 14, wherein, The one or more virtual objects include one or more imagined computer-generated objects.
16. The system according to claim 11, wherein, The operation also includes: Obtain metadata associated with the one or more two-dimensional images; Visual effects are generated based on the metadata, including one or more enhancements to be applied to the real-world environment; and This enables the display device to present the visual effects applied to the real-world environment.
17. The system according to claim 16, wherein, The presentation of the visual effects applied to the real-world environment includes rendering animation effects together with the volumetric content items.
18. The system according to claim 11, wherein, The operation also includes: Obtain metadata associated with the one or more two-dimensional images; Based on the metadata, a real-world effect is determined, which is associated with functionality provided by the network-connected device; and This enables the network-connected device in the real-world environment to provide the real-world effect.
19. The system according to claim 11, wherein, The operation also includes: Volumetric content is generated by overlaying the volumetric content items onto the real-world environment; and The volume content is stored for subsequent presentation.
20. A machine-readable medium storing instructions that, when executed by a computer system, cause the computer system to perform operations, the operations including: Access image data that includes one or more two-dimensional images; By performing image segmentation, multiple two-dimensional elements are identified from one or more two-dimensional images; Select a two-dimensional element from the plurality of two-dimensional elements; A volume content item is generated based on the two-dimensional elements identified from the one or more two-dimensional images; Identify real-world elements in a real-world environment; Associate the volume contents with the real-world elements in the real-world environment; An interactive interface is presented, which includes the volume content item and one or more real-world elements detected in the real-world environment. The interactive interface is operable to receive the association between the volume content item and the real-world elements. Receive input indicating the association between the volume content item and the real-world element; as well as This enables the display device to present the volumetric content item as an overlay on the real-world environment at the location of the identified real-world element, based on the association, and the real-world environment is within the user's field of view of the display device.