Independent scene movement based on mask layer
By using the mask layer to adjust the image in an extended reality system to consider the movement of physical objects, the problem of high power and large bandwidth requirements is solved, and independent scene movement with low latency and low power is achieved, improving the portability and image quality of the device.
Patent Information
- Application Number
- CN202380081071.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-11-15
- Publication Date
- 2025-07-04
AI Technical Summary
The existing extended reality systems have short battery life and heavy equipment, which are inconvenient to carry due to high power requirements and large bandwidth requirements.
By using mask layers to determine independent scene movement, adjusting images to take into account movement of physical objects, reducing the burden on the processor, using low latency and low power image processing techniques.
Effectively reduces the power and bandwidth requirements of extended reality systems, while maintaining image quality, improving device portability and user experience.
Smart Images

Figure CN120266159A_ABST
Abstract
Description
Technical Field
[0001] This application relates to processing one or more images for an extended reality system. For example, aspects of this application relate to systems and techniques for determining independent scene movement based on one or more mask layers. Background Art
[0002] Degree of freedom (DoF) refers to the number of fundamental ways a rigid object can move in three-dimensional (3D) space. In some examples, six different DoFs can be tracked. The six DoFs include three translational DoFs corresponding to translational movement along three perpendicular axes, which can be referred to as the x-axis, y-axis, and z-axis. The six DoFs include three rotational DoFs corresponding to rotational movement about three axes, which can be referred to as pitch, yaw, and roll. Some extended reality (XR) devices, such as virtual reality (VR) or augmented reality (AR) headsets, can track some or all of these degrees of freedom. For example, 3DoF XR headsets typically track three rotational DoFs and can thus track whether a user turns and / or tilts their head. 6DoF XR headsets track all six DoFs and can thus also track a user's translational movement.
[0003] XR systems typically use powerful processors to perform feature analysis (e.g., extraction, tracking, etc.) and other complex functions fast enough to display an output based on those functions to their users. Powerful processors typically draw power at a high rate. Similarly, transmitting large amounts of data to a powerful processor typically draws power at a high rate. Headsets and other portable devices typically have small batteries so as not to be uncomfortably heavy for the user. Thus, some XR systems must be plugged into an external power source and are thus not portable. Portable XR systems typically have a short battery life and / or are uncomfortably heavy due to including large batteries. Summary of the Invention
[0004] Systems and techniques for displaying augmented reality enhanced media content are described herein. For example, aspects of the present disclosure relate to systems and techniques for reducing the power and bandwidth of an XR system and simultaneously maintaining image quality using low-latency and lower-power independent scene movement. A simplified summary of the invention related to one or more aspects disclosed herein is presented below. Accordingly, the following summary of the invention should neither be considered an exhaustive overview of all contemplated aspects nor be considered to identify key or critical elements related to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary of the invention presents in simplified form certain concepts related to one or more aspects regarding the mechanisms disclosed herein prior to the detailed description presented below.
[0005] Systems and techniques for an apparatus for image generation are provided. The apparatus includes: a memory that includes instructions; and a processor coupled to the memory. The processor is configured to: render an image of a scene that includes a first virtual object associated with a first physical object that moves independently of the scene; identify a first set of pixels associated with the first virtual object in the image of the scene; generate a mask layer based on the first set of pixels, the mask layer indicating the location of the first set of pixels in the image; and send the mask layer and the image to a display.
[0006] As another example, an apparatus for image generation is provided. The apparatus includes: a memory that includes instructions; and a processor coupled to the memory. The at least one processor is configured to: obtain a mask layer and an image associated with the mask layer, where the mask layer is associated with a first virtual object; generate updated pose information for the apparatus, where the updated pose information indicates movement between the time the image was generated and the current time; obtain updated pose information for a first physical object that moves independently of the apparatus and that is associated with the first virtual object; determine a first amount to distort the image based on the updated pose information for the apparatus; determine a second amount to distort a first portion of the image based on the mask layer and the updated pose information for the first physical object; distort the image based on the first amount to distort the image based on the updated pose information and distort the first portion of the image based on the second amount to distort the first virtual object by the first physical object; and display a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0007] In another example, a method for image generation is provided. The method includes: rendering an image of a scene that includes a first virtual object associated with a first physical object that moves independently of the scene; identifying a first set of pixels associated with the first virtual object in the image of the scene; generating a mask layer based on the first set of pixels, the mask layer indicating the location of the first set of pixels in the image; and sending the mask layer and the image to a display.
[0008] For another example, a method for image generation is provided. The method includes: obtaining a mask layer and an image associated with the mask layer, where the mask layer is associated with a first virtual object; generating updated pose information of a device, where the updated pose information indicates a movement between a time when the image was generated and the current time; obtaining updated pose information of a first physical object, where the first physical object moves independently of the device and where the first physical object is associated with the first virtual object; determining a first amount of distorting the image based on the updated pose information of the device; determining a second amount of distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; distorting the image based on the first amount of distorting the image based on the updated pose information, and distorting the first portion of the image based on the second amount of distorting the first virtual object by the first physical object; and displaying a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0009] As another example, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: render an image of a scene, the scene including a first virtual object associated with a first physical object that moves independently of the scene; identify a first set of pixels associated with the first virtual object in the image of the scene; generate a mask layer based on the first set of pixels, the mask layer indicating a position of the first set of pixels in the image; and send the mask layer and the image to a display.
[0010] As another example, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: obtain a mask layer and an image associated with the mask layer, where the mask layer is associated with a first virtual object; generate updated pose information of a device, where the updated pose information indicates a movement between a time when the image was generated and the current time; obtain updated pose information of a first physical object, where the first physical object moves independently of the device and where the first physical object is associated with the first virtual object; determine a first amount of distorting the image based on the updated pose information of the device; determine a second amount of distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; distorting the image based on the first amount of distorting the image based on the updated pose information, and distorting the first portion of the image based on the second amount of distorting the first virtual object by the first physical object; and display a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0011] In another example, an apparatus for image generation is provided. The apparatus includes: components for rendering an image of a scene, the scene including a first virtual object associated with a first physical object that moves independently of the scene; components for identifying a first set of pixels associated with the first virtual object in the image of the scene; components for generating a mask layer based on the first set of pixels, the mask layer indicating the positions of the first set of pixels in the image; and components for sending the mask layer and the image to a display.
[0012] In another example, an apparatus for image generation is provided. The apparatus includes: components for obtaining a mask layer and an image associated with the mask layer, wherein the mask layer is associated with a first virtual object; components for generating updated pose information of the apparatus, wherein the updated pose information indicates a movement between the time when the image was generated and the current time; components for obtaining updated pose information of a first physical object, wherein the first physical object moves independently of the apparatus and wherein the first physical object is associated with the first virtual object; components for determining a first amount for distorting the image based on the updated pose information of the apparatus; components for determining a second amount for distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; components for distorting the image based on the first amount for distorting the image based on the updated pose information and distorting the first portion of the image based on the second amount for distorting the first virtual object by the first physical object; and components for displaying a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0013] In some aspects, the apparatus may include an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device (e.g., a mobile phone or other mobile device), a wearable device (e.g., a network-connected watch or other wearable device), a personal computer, a laptop computer, a server computer, a television, a video game console, or other device or as part thereof. In some aspects, the apparatus further includes at least one camera for capturing one or more images or video frames. For example, the apparatus may include one camera (e.g., an RGB camera) or multiple cameras for capturing one or more images and / or one or more videos including video frames. In some aspects, the apparatus includes a display for displaying one or more images, videos, notifications, or other displayable data. In some aspects, the apparatus includes a transmitter configured to send data or information to at least one device over a transmission medium. In some aspects, the processor includes a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), or other processing device or component.
[0014] The disclosure is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood in reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0015] The foregoing, as well as other features and examples, will become more apparent after reference to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Exemplary examples of the present application are described in detail below with reference to the following drawings:
[0017] Figure 1 is a block diagram illustrating the architecture of an image capture and processing system in accordance with aspects of the present disclosure.
[0018] Figure 2 is a diagram illustrating the architecture of an example extended reality (XR) system in accordance with some aspects of the present disclosure.
[0019] Figure 3 is a block diagram illustrating the architecture of a simultaneous localization and mapping (SLAM) system in accordance with aspects of the present disclosure.
[0020] Figure 4 illustrates an example of an augmented reality enhancement application engine in accordance with aspects of the present disclosure.
[0021] Figure 5 illustrates an example of a main application engine and a secondary application engine that can provide augmented reality enhancement to the main application engine in accordance with aspects of the present disclosure.
[0022] Figure 6A and Figure 6B are images illustrating example scenarios for an XR system in accordance with aspects of the present disclosure.
[0023] Figure 7 illustrates an example mask layer in accordance with aspects of the present disclosure.
[0024] Figure 8 is a block diagram illustrating an example of a process for low-latency and low-power independent scene movement in an XR system in accordance with aspects of the present disclosure.
[0025] Figure 9 is a flowchart illustrating an example of a process for performing image generation in accordance with some aspects.
[0026] Figure 10 is a flowchart illustrating another example of a process for performing image generation in accordance with some aspects.
[0027] Figure 11AIs a perspective view of a head-mounted display (HMD) 1110 that exemplifies performing feature tracking and / or visual simultaneous localization and mapping (VSLAM) according to some examples.
[0028] Figure 11B Is an illustration according to some examples Figure 11A of a head-mounted display (HMD) being worn by a user.
[0029] Figure 12A Is a perspective view of the front surface of a mobile device that exemplifies performing feature tracking and / or visual simultaneous localization and mapping (VSLAM) using one or more front cameras according to some examples.
[0030] Figure 12B Is a perspective view of the back surface of a mobile device 1250 according to aspects of the present disclosure.
[0031] Figure 13 Is a diagram that exemplifies an example of a system for implementing certain aspects of the present technology. Detailed Description
[0032] Certain aspects and examples of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and examples may be applied independently, and some of them may be applied in combination. In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of the subject matter of the present application. However, it will be apparent that the various examples may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0033] The following description merely provides exemplary examples and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description will provide a viable description for those skilled in the art to implement the exemplary examples. It should be understood that various changes may be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0034] A camera (e.g., an image capture device) is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). The terms “image,” “image frame,” and “frame” may be used interchangeably herein. A camera may be configured with various image capture and image processing settings. Different settings produce images with different appearances. Some camera settings, such as ISO, exposure time, aperture size, aperture value, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters may be applied to the image sensor used to capture one or more image frames. Other camera settings may configure post-processing of one or more image frames, such as changes in contrast, brightness, saturation, sharpness, levels, curves, or color. For example, settings or parameters may be applied to a processor (e.g., an image signal processor or ISP) used to process one or more image frames captured by the image sensor.
[0035] Degree of freedom (DoF) refers to the number of basic ways in which a rigid object can move in three-dimensional (3D) space. In some cases, six different DoFs may be tracked. The six degrees of freedom include three translational degrees of freedom corresponding to translational movement along three perpendicular axes. These three axes may be referred to as the x-axis, y-axis, and z-axis. The six degrees of freedom include three rotational degrees of freedom corresponding to rotational movement about three axes, which may be referred to as pitch, yaw, and roll.
[0036] An extended reality (XR) system or device may provide virtual content to a user and / or may combine the real world and the physical environment and a virtual environment (composed of virtual content) to provide an XR experience to the user. The real world environment may include real world objects (also referred to as physical objects), such as people, vehicles, buildings, tables, chairs, and / or other real world or physical objects. An XR system or device may facilitate interaction with different types of XR environments (e.g., a user may use the XR system or device to interact with an XR environment). An XR system may include a VR system that facilitates interaction with a virtual reality (VR) environment, an AR system that facilitates interaction with an augmented reality (AR) environment, an MR system that facilitates interaction with a mixed reality (MR) environment, and / or other XR systems. Examples of XR systems or devices include head-mounted displays (HMDs), smart glasses, etc. In some cases, an XR system may track parts of a user (e.g., the user's hands and / or fingertips) to allow the user to interact with items of virtual content.
[0037] AR is a technology that provides virtual or computer-generated content (referred to as AR content) on a user's view of a physical, real-world scene or environment. AR content can include virtual content such as videos, images, graphic content, location data (e.g., Global Positioning System (GPS) data or other location data), sounds, any combination thereof, and / or other augmented content. An AR system or device is designed to enhance (or augment) rather than replace a person's current perception of reality. For example, a user can see a physical object that is real and stationary or moving through an AR device display, but the user's visual perception of the physical object can be enhanced or augmented by a virtual image of the object (e.g., a real-world car replaced by a virtual image of a DeLorean), by AR content added to the physical object (e.g., virtual wings added to a living animal), by AR content displayed relative to the physical object (e.g., information virtual content displayed near a sign on a building, a virtual coffee cup virtually anchored to a real-world table in one or more images (e.g., placed on top of the real-world table)), and / or by displaying other types of AR content. Various types of AR systems can be used for gaming, entertainment, and / or other applications.
[0038] In some cases, an XR system can include an optical "see-through" display or "passthrough" display (e.g., a see-through or passthrough AR HMD or AR glasses), allowing the XR system to directly display XR content (e.g., AR content) onto a real-world view without displaying video content. For example, a user can view a physical object through a display (e.g., glasses or lenses), and the AR system can display AR content onto the display to provide the user with an enhanced visual perception of one or more real-world objects. In one example, the display of an optical see-through AR system can include lenses or glasses in front of each eye (or a single lens or glasses above both eyes). The see-through display can allow the user to directly see the real-world object or physical object and can display (e.g., project or otherwise display) an enhanced image of the object or additional AR content to enhance the user's visual perception of the real world.
[0039] Visual Simultaneous Localization and Mapping (VSLAM) is a computational geometry technique for use in devices having a camera, such as robots, head-mounted displays (HMDs), mobile phones, and autonomous vehicles. In VSLAM, a device can build and update a map of an unknown environment based on images captured by the device's camera. As the device updates the map, the device can maintain track of its pose (e.g., position and / or orientation) within the environment. For example, a device can be initiated in a particular room of a building and can move throughout the interior of the building, capturing images. The device can build a map of the environment and track its position within the environment based on tracking the positions at which different objects in the environment appear in different images.
[0040] In the context of a system that tracks movement through an environment (such as an XR system and / or a VSLAM system), degrees of freedom can refer to which of the six degrees of freedom the system is able to track. A 3DoF system typically tracks three rotational DoFs - pitch, yaw, and roll. For example, a 3DoF head-mounted headset can track a user of the headset turning their head left or right, tilting their head up or down, and / or tilting their head left or right. A 6DoF system can track three translational DoFs as well as three rotational DoFs. Thus, for example, a 6DoF head-mounted headset can track a user moving forward, backward, laterally, and / or vertically in addition to tracking the three rotational DoFs.
[0041] In some cases, an XR system can include an HMD display wearable by a user of the XR system, such as an AR HMD or AR glasses. Generally, it is desirable to keep the HMD display as light and as small as possible. To help reduce the weight and size of the HMD display, the HMD display can be a relatively low-power system (e.g., in terms of battery and computational power) compared to a device (e.g., a companion device, such as a mobile phone, server device, or other device) to which the HMD display is connected (e.g., wired or wirelessly).
[0042] In some cases, split rendering can be implemented. In split rendering, a companion device can perform certain tasks with respect to one or more images to be displayed by an HMD display and send the task results to the HMD display. The HMD display can then perform additional image processing tasks and display the one or more images. In some cases, there may be an inherent latency when certain tasks are remotely executed at the companion device and the results are sent to the HMD display for additional processing. In an XR system, a user of the XR system may move the HMD display during such latency. This movement can cause an overall shift in the field of view of the HMD display. Additionally, one or more objects in the field of view of the HMD display can move independently of the movement of the HMD display, such as the user's hand, arm, or other objects. Using the same transformation for different movements (e.g., hand movement, arm movement, etc.) can result in objects being displayed at incorrect positions on the display, which can lead to an inconsistent immersive XR experience for the user of the HMD display. Furthermore,
[0043] This disclosure describes systems, devices, methods (also referred to as processes), and computer-readable media (collectively referred to herein as "systems and techniques") for determining independent scene movement based on one or more mask layers. For example, the systems and techniques can adjust (e.g., warp) portions of a rendered image to account for the movement of an XR device (e.g., an HMD display, AR glasses, etc.) and objects that can move independently of the XR device. In one illustrative example, a companion device can determine that a virtual object is associated with a physical object (e.g., the user's hand) in the field of view of the XR device that can move independently of the XR device. The companion device can generate a mask layer indicating the pixels associated with the object. The companion device can send the mask layer to the XR device along with the image to be displayed by the XR device. The XR device can warp (e.g., adjust) the image based on the mask layer to account for the movement of the XR device and the physical object. The XR device can then display the warped image. In addition to split rendering scenarios, the techniques discussed herein can be applicable to other situations where there may be a latency between the time of rendering and the display of the rendered image, such as server-based rendering.
[0044] Various aspects of the present application will be described with reference to the drawings. Figure 1is a block diagram illustrating the architecture of an exemplary image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing images of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture individual images (or photos), and / or can capture video including multiple images (or video frames) in a particular sequence. In some cases, the lens 115 and the image sensor 130 can be associated with an optical axis. In one exemplary example, both the photosensitive area of the image sensor 130 (e.g., a photodiode) and the lens 115 can be centered on the optical axis. The lens 115 of the image capture and processing system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the incoming light from the scene towards the image sensor 130. The light received by the lens 115 passes through an aperture. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 120 and is received by the image sensor 130. In some cases, the aperture can have a fixed size.
[0045] One or more control mechanisms 120 can control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. One or more control mechanisms 120 can include multiple mechanisms and components; for example, the control mechanism 120 can include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 can also include additional control mechanisms other than those illustrated, such as control mechanisms for controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.
[0046] The focusing control mechanism 125B of the control mechanism 120 can obtain a focusing setting. In some examples, the focusing control mechanism 125B stores the focusing setting in a memory register. Based on the focusing setting, the focusing control mechanism 125B can adjust the positioning of the lens 115 relative to the positioning of the image sensor 130. For example, based on the focusing setting, the focusing control mechanism 125B can move the lens 115 closer to or farther away from the image sensor 130 by actuating a motor or a servo system (or other lens mechanism) to adjust the focus. In some cases, additional lenses may be included in the image capture and processing system 100, such as one or more microlenses located above each photodiode of the image sensor 130, each of the one or more microlenses bending the light received from the lens 115 towards the corresponding photodiode before the light reaches the photodiode. The focusing setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The control mechanism 120, the image sensor 130, and / or the image processor 150 can be used to determine the focusing setting. The focusing setting can be referred to as an image capture setting and / or an image processing setting. In some cases, the lens 115 can be fixed relative to the image sensor, and the focusing control mechanism 125B can be omitted without departing from the scope of the present disclosure.
[0047] The exposure control mechanism 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control mechanism 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 125A can control the size of the aperture (e.g., aperture size or f / stop), the duration for which the aperture is open (e.g., exposure time or shutter speed), the duration for which the sensor collects light (e.g., exposure time or electronic shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting can be referred to as an image capture setting and / or an image processing setting.
[0048] The zoom control mechanism 125C of the control mechanism 120 can obtain a zoom setting. In some examples, the zoom control mechanism 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control mechanism 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly can include a focusing lens (in some cases, the focusing lens can be the lens 115), which first receives light from the scene 110, and the light then passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system can include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference from each other), with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control mechanism 125C moves one or more of the lenses in the afocal zoom system, such as the negative lens and one or both of the positive lenses. In some cases, the zoom control mechanism 125C can control the zoom by capturing an image from an image sensor (e.g., including the image sensor 130) among a plurality of image sensors at a zoom corresponding to the zoom setting. For example, the image processing system 100 can include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a greater zoom. In some cases, based on the selected zoom setting, the zoom control mechanism 125C can capture an image from the corresponding sensor.
[0049] The image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image produced by the image sensor 130. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters and may thus measure light that matches the color of the filter covering the photodiode. Various color filter arrays can be used, including a Bayer color filter array, a four-color filter array (also known as a four-color Bayer filter array or QCFA), and / or any other color filter array. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, where each pixel of the image is generated based on red light data from at least one photodiode covered in the red color filter, blue light data from at least one photodiode covered in the blue color filter, and green light data from at least one photodiode covered in the green color filter.
[0050] Return to Figure 1 , yellow, magenta, and / or cyan (also known as "emerald") filters can be used to replace or supplement the red, blue, and / or green filters. In some cases, some photodiodes may be configured to measure infrared (IR) light. In some embodiments, the photodiodes that measure IR light may not be covered by any filter, thus allowing the IR photodiodes to measure both visible light (e.g., color) and IR light. In some examples, the IR photodiodes may be covered by an IR filter, thus allowing IR light to pass through and blocking light from other parts of the spectrum (e.g., visible light, color). Some image sensors (e.g., the image sensor 130) may lack filters altogether (e.g., color, IR, or any other part of the spectrum) and may instead use different photodiodes (vertically stacked in some cases) throughout the pixel array. The different photodiodes throughout the pixel array may have different spectral sensitivity curves, thereby responding to light of different wavelengths. Monochrome image sensors may also lack filters and thus lack color depth.
[0051] In some cases, image sensor 130 may alternatively or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, the opaque and / or reflective masks may be used for phase detection autofocus (PDAF). In some cases, the opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cut-off filter, UV cut-off filter, band-pass filter, low-pass filter, high-pass filter, etc.). Image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output by the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output by the photodiodes (and / or the analog signal amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed with respect to one or more of the control mechanisms 120 may alternatively or additionally be included in image sensor 130. Image sensor 130 may be a charge-coupled device (CCD) sensor, an electron multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), complementary metal-oxide semiconductor (CMOS), N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0052] Image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or with respect to Figure 10One or more processors among any other type of processor 1010 discussed in the computing system 1000. The host processor 152 can be a digital signal processor (DSP) and / or other types of processors. In some specific implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), a memory, connectivity components (e.g., BluetoothTM, Global Positioning System (GPS), etc.), any combination thereof and / or other components. The I / O port 156 can include any suitable input / output port or interface according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 Physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof and / or other input / output ports. In an illustrative example, the host processor 152 can communicate with the image sensor 130 using an I2C port, and the ISP 154 can communicate with the image sensor 130 using a MIPI port.
[0053] The image processor 150 can perform multiple tasks, such as demosaicking, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving inputs, managing outputs, managing memory, or some combination thereof. The image processor 150 can store the image frames and / or the processed images in a random access memory (RAM) 140 / 1025, a read-only memory (ROM) 145 / 1020, a cache, a memory cell, another storage device, or some combination thereof.
[0054] A variety of input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device 1035, any other input device 1045, or some combination thereof. In some cases, captions may be input into the image processing device 105B via the physical keyboard or keypad of the I / O device 160, or via the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that implement a wired connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The I / O 160 may include one or more wireless transceivers that implement a wireless connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The peripheral devices may include any type of I / O device 160 discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they may themselves be considered I / O devices 160.
[0055] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some specific embodiments, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly via one or more wireless transceivers. In some specific embodiments, the image capture device 105A and the image processing device 105B may be disconnected from each other.
[0056] As Figure 1 shown, the vertical dashed line divides Figure 1 the image capture and processing system 100 into two parts, representing the image capture device 105A and the image processing device 105B, respectively. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, certain components illustrated in the image capture device 105A (such as the ISP 154 and / or the host processor 152) may be included in the image capture device 105A.
[0057] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some embodiments, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.
[0058] Although the image capture and processing system 100 is shown as including certain components, those of ordinary skill in the art should understand that the image capture and processing system 100 may include more components than Figure 1 those shown therein. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, the components of the image capture and processing system 100 may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture and processing system 100.
[0059] In some examples, Figure 2 the extended reality (XR) system 200 may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof. In some examples, Figure 3 the simultaneous localization and mapping (SLAM) system 300 may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.
[0060] Figure 2FIG. 0 is a schematic diagram illustrating the architecture of an example extended reality (XR) system 200 in accordance with some aspects of the present disclosure. The XR system 200 may run (or execute) XR applications and implement XR operations. In some examples, as part of an XR experience, the XR system 200 may perform tracking and positioning, mapping of an environment in the physical world (e.g., a scene), and / or positioning and rendering of virtual content on a display 209 (e.g., a screen, a visible plane / region, and / or other display). For example, the XR system 200 may generate a map of an environment in the physical world (e.g., a three-dimensional (3D) map), track the pose (e.g., position and orientation) of the XR system 200 relative to the environment (e.g., relative to the 3D map of the environment), position and / or anchor virtual content at a specific location on the map of the environment, and render the virtual content on the display 209 such that the virtual content appears to be at a location in the environment corresponding to the specific location on the map of the scene, where the virtual content is positioned and / or anchored at the map of the scene. The display 209 may include glass, a screen, a lens, a projector, and / or other display mechanisms that allow a user to see the real-world environment and also allow XR content to be overlaid, overlapped, blended, or otherwise displayed thereon.
[0061] In this illustrative example, the XR system 200 includes one or more image sensors 202, an accelerometer 204, a gyroscope 206, a storage device 207, a computing component 210, an XR engine 220, an image processing engine 224, a rendering engine 226, and a communication engine 228. It should be noted that Figure 2 the components 202-228 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples may include more, fewer, or different components compared to Figure 2 the components shown. For example, in some cases, the XR system 200 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 2 one or more other software and / or hardware components not shown. Although the various components of the XR system 200 (such as the image sensor 202) may be referred to herein in the singular form, it should be understood that the XR system 200 may include multiple of any of the components discussed herein (e.g., multiple image sensors 202).
[0062] The XR system 200 includes an input device 208 or communicates (wired or wirelessly) with the input device. The input device 208 can include any suitable input device, such as a touch screen, a pen or other pointer device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, a video game controller, a steering wheel, a joystick, a set of buttons, a trackball, a remote control, any other input device 1045 discussed herein, or any combination thereof. In some cases, the image sensor 202 can capture images that can be processed for interpreting gesture commands.
[0063] The XR system 200 can also communicate (wired or wirelessly) with one or more other electronic devices. For example, the communication engine 228 can be configured to manage connections and communicate with one or more electronic devices. In some cases, the communication engine 228 can correspond to Figure 10 the communication interface 1040.
[0064] In some specific implementations, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, XR engine 220, image processing engine 224, and rendering engine 226 can be part of the same computing device. For example, in some cases, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, XR engine 220, image processing engine 224, and rendering engine 226 can be integrated into a head-mounted display (HMD), extended reality glasses, a smart phone, a laptop computer, a tablet computer, a gaming system, and / or any other computing device. However, in some specific implementations, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, XR engine 220, image processing engine 224, and rendering engine 226 can be part of two or more separate computing devices. For example, in some cases, some of the components 202 - 226 can be part of or implemented by one computing device, and the remaining components can be part of or implemented by one or more other computing devices.
[0065] The storage device 207 can be any storage device for storing data. Additionally, the storage device 207 can store data from any of the components in the XR system 200. For example, the storage device 207 can store data from the image sensor 202 (e.g., images or video data), data from the accelerometer 204 (e.g., measurements), data from the gyroscope 206 (e.g., measurements), data from the computing component 210 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, face recognition data, occlusion data, etc.), data from the XR engine 220, data from the image processing engine 224, and / or data from the rendering engine 226 (e.g., output frames). In some examples, the storage device 207 can include a buffer for storing frames to be processed by the computing component 210.
[0066] One or more computing components 210 can include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, an image signal processor (ISP) 218, and / or other processors (e.g., a neural processing unit (NPU) that implements one or more trained neural networks). The computing component 210 can perform various operations, such as image enhancement, computer vision, graphics rendering, extended reality (e.g., tracking, positioning, pose estimation, mapping, content anchoring, content rendering, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, the computing component 210 can implement (e.g., control, operate, etc.) the XR engine 220, the image processing engine 224, and the rendering engine 226. In other examples, the computing component 210 can also implement one or more other processing engines.
[0067] The image sensor 202 can include any image and / or video sensor or capture device. In some examples, the image sensor 202 can be part of a multi-camera assembly (such as a dual-camera assembly). The image sensor 202 can capture image and / or video content (e.g., raw image and / or video data), which can then be processed by the computing component 210, the XR engine 220, the image processing engine 224, and / or the rendering engine 226, as described herein. In some examples, the image sensor 202 can include an image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination thereof.
[0068] In some examples, the image sensor 202 may capture image data and may generate an image (also referred to as a frame) based on the image data, and / or may provide the image data or the frame to the XR engine 220, the image processing engine 224, and / or the rendering engine 226 for processing. The image or frame may include a video frame or a static image in a video sequence. The image or frame may include an array of pixels representing a scene. For example, the image may be a Red, Green, Blue (RGB) image having red, green, and blue color components per pixel; a Luminance, Chroma Red, Chroma Blue (YCbCr) image having a luminance component and two chroma (color) components (chroma red and chroma blue) per pixel; or any other suitable type of color or monochrome image.
[0069] In some cases, the image sensor 202 (and / or other cameras of the XR system 200) may also be configured to capture depth information. For example, in some embodiments, the image sensor 202 (and / or other cameras) may include a Red Green Blue-Depth (RGB-D) camera. In some cases, the XR system 200 may include one or more depth sensors (not shown) that are separate from the image sensor 202 (and / or other cameras) and may capture depth information. For example, such depth sensors may obtain depth information independently of the image sensor 202. In some examples, the depth sensor may be physically mounted in the same general location as the image sensor 202, but may operate at a different frequency or frame rate than the image sensor 202. In some examples, the depth sensor may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in a scene. Depth information may then be obtained by exploiting the geometric deformation of the projected pattern caused by the surface shape of the object. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).
[0070] The XR system 200 may also include other sensors in one or more of its sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to the computing component 210. For example, the accelerometer 204 may detect the acceleration of the XR system 200 and generate an acceleration measurement based on the detected acceleration. In some cases, the accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward), which may be used to determine the position or pose of the XR system 200. The gyroscope 206 may detect and measure the orientation and angular velocity of the XR system 200. For example, the gyroscope 206 may be used to measure the pitch, roll, and yaw of the XR system 200. In some cases, the gyroscope 206 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 202 and / or the XR engine 220 may use the measurements obtained by the accelerometer 204 (e.g., one or more translation vectors) and / or the measurements obtained by the gyroscope 206 (e.g., one or more rotation vectors) to calculate the pose of the XR system 200. As previously mentioned, in other examples, the XR system 200 may also include other sensors, such as an inertial measurement unit (IMU), a magnetometer, a gaze and / or eye tracking sensor, a machine vision sensor, a smart scene sensor, a speech recognition sensor, an impact sensor, a vibration sensor, a position sensor, a tilt sensor, etc.
[0071] As described above, in some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure the specific force, angular velocity, and / or orientation of the extended reality system 200. In some examples, the one or more sensors may output information measured in association with the capture of an image captured by the image sensor 202 (and / or other cameras of the XR system 200) and / or depth information obtained using one or more depth sensors of the XR system 200.
[0072] The XR engine 220 can use the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) to determine the pose (also referred to as the head pose) of the XR system 200 and / or the pose of the image sensor 202 (or other cameras of the XR system 200). In some cases, the pose of the XR system 200 and the pose of the image sensor 202 (or other cameras) can be the same. The pose of the image sensor 202 refers to the position and orientation of the image sensor 202 relative to a reference frame (e.g., with respect to the scene 110). In some specific implementations, the camera pose can be determined for six degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by the X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some specific implementations, the camera pose can be determined for three degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).
[0073] In some cases, a device tracker (not shown) can use the measurements from one or more sensors and the image data from the image sensor 202 to track the pose of the XR system 200 (e.g., 6DoF pose). For example, the device tracker can fuse the visual data from the image data (e.g., using a visual tracking solution) with the inertial data from the measurements to determine the position and movement of the XR system 200 relative to the physical world (e.g., the scene) and the map of the physical world. As described below, in some examples, when tracking the pose of the XR system 200, the device tracker can generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate an update to the 3D map for the scene. The 3D map update can include, for example but not limited to, new or updated features and / or feature or fiducial points associated with the scene and / or the 3D map of the scene, and / or a localization update that identifies or updates the position of the XR system 200 within the scene and the 3D map of the scene. The 3D map can provide a digital representation of the scene in the real / physical world. In some examples, the 3D map can anchor location-based objects and / or content to real-world coordinates and / or objects. The XR system 200 can use the mapped scene (e.g., the scene in the physical world represented and / or associated with the 3D map) to merge the physical and virtual worlds and / or to merge virtual content or objects with the physical environment.
[0074] In some aspects, the computing component 210 may use a visual tracking solution to determine and / or track the pose of the image sensor 202 and / or the XR system 200 as a whole based on images captured by the image sensor 202 (and / or other cameras of the XR system 200). For example, in some examples, the computing component 210 may use computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques to perform the tracking. For example, the computing component 210 may perform SLAM or may communicate (wired or wirelessly) with a SLAM system (not shown) such as Figure 3 the SLAM system 300. SLAM refers to a class of techniques that create a map of an environment (e.g., a map of the environment modeled by the XR system 200) while simultaneously tracking the camera (e.g., the image sensor 202) and / or the pose of the XR system 200 relative to that map. This map may be referred to as a SLAM map and may be three-dimensional (3D). SLAM techniques may use color or grayscale image data captured by the image sensor 202 (and / or other cameras of the XR system 200) and may be used to generate an estimate of the 6DoF pose measurements of the image sensor 202 and / or the XR system 200. Such SLAM techniques configured to perform 6DoF tracking may be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and / or other sensors) may be used to estimate, correct, and / or otherwise adjust the estimated pose.
[0075] In some cases, 6DoF SLAM (e.g., 6DoF tracking) may associate features observed from certain input images from the image sensor 202 (and / or other cameras) to the SLAM map. For example, 6DoF SLAM may use feature point associations from the input images to determine the pose (position and orientation) of the image sensor 202 and / or the XR system 200 of the input image. 6DoF mapping may also be performed to update the SLAM map. In some cases, the SLAM map maintained using 6DoF SLAM may contain 3D feature points triangulated from two or more images. For example, key frames may be selected from the input images or video stream to represent the observed scene. For each key frame, the corresponding 6DoF camera pose associated with the image may be determined. The pose of the image sensor 202 and / or the XR system 200 may be determined by projecting features from the 3D SLAM map into the image or video frame and updating the camera pose based on the verified 2D-3D correspondences.
[0076] In an illustrative example, computing component 210 may extract feature points from certain input images (e.g., each input image, a subset of input images, etc.) or from each key frame. Feature points (also referred to as registration points) as used herein are unique or identifiable portions of an image, such as a part of a hand, an edge of a table, and other examples. Features extracted from the captured images may represent different feature points along a three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and each feature point may have an associated feature location. Feature points in a key frame match (are the same as or correspond to) or fail to match feature points of a previously captured input image or key frame. Feature detection may be used to detect feature points. Feature detection may include image processing operations for examining one or more pixels of an image to determine whether a feature exists at a particular pixel. Feature detection may be used to process an entire captured image or certain portions of the image. For each image or key frame, once features have been detected, local image patches around the feature may be extracted. Any suitable technique may be used to extract features, such as Scale-Invariant Feature Transform (SIFT) (which localizes features and generates their descriptions), Learned Invariant Feature Transform (LIFT), Speeded-Up Robust Features (SURF), Gradient Location-Orientation Histogram (GLOH), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoint (FREAK), KAZE, Accelerated KAZE (AKAZE), Normalized Cross-Correlation (NCC), descriptor matching, another suitable technique, or a combination thereof.
[0077] As an illustrative example, computing component 210 may extract feature points corresponding to a mobile device (e.g., Figure 4 mobile device 440 of Figure 5 mobile device 540 of) etc. In some cases, feature points corresponding to the mobile device may be tracked to determine the pose of the mobile device. As described in more detail below, the pose of the mobile device may be used to determine the location of a projection of AR media content that may enhance media content displayed on a display of the mobile device.
[0078] In some cases, XR system 200 may also track a user's hand and / or fingers to allow the user to interact with and / or control virtual content in a virtual environment. For example, XR system 200 may track the pose and / or movement of a user's hand and / or fingertips to identify or translate the user's interaction with the virtual environment. User interaction may include, for example but not limited to, moving a virtual content item, resizing a virtual content item, selecting an input interface element in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), providing input through a virtual user interface, etc.
[0079] Figure 3is a block diagram illustrating the architecture of a Simultaneous Localization and Mapping (SLAM) system 300. In some examples, the SLAM system 300 can be or can include an Extended Reality (XR) system (such as Figure 2 the XR system 200). In some examples, the SLAM system 300 can be a wireless communication device, a mobile device or a cell phone (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device, a personal computer, a laptop computer, a server computer, a portable video game console, a portable media player, a camera device, a manned or unmanned ground vehicle, a manned or unmanned aerial vehicle, a manned or unmanned water vehicle, a manned or unmanned underwater vehicle, a manned or unmanned vehicle, an autonomous vehicle, a vehicle, a computing system of a vehicle, a robot, another device, or any combination thereof.
[0080] Figure 3 Each sensor of the one or more sensors 305 of the SLAM system 300 is included in or coupled to each sensor. The one or more sensors 305 can include one or more cameras 310. Each camera of the one or more cameras 310 can include an image capture device 105A, an image processing device 105B, an image capture and processing system 100, another type of camera, or a combination thereof. Each camera of the one or more cameras 310 can respond to light from a specific spectrum. The spectrum can be a subset of the electromagnetic (EM) spectrum. For example, each camera of the one or more cameras 310 can be a visible light (VL) camera that responds to the VL spectrum, an infrared (IR) camera that responds to the IR spectrum, an ultraviolet (UV) camera that responds to the UV spectrum, a camera that responds to light of another spectrum from another part of the electromagnetic spectrum, or some combination thereof.
[0081] One or more sensors 305 may include one or more other types of sensors in addition to camera 310, such as one or more of the following: accelerometers, gyroscopes, magnetometers, inertial measurement units (IMUs), altimeters, barometers, thermometers, radio detection and ranging (RADAR) sensors, light detection and ranging (LIDAR) sensors, sound navigation and ranging (SONAR) sensors, sound detection and ranging (SODAR) sensors, global navigation satellite system (GNSS) receivers, global positioning system (GPS) receivers, Beidou navigation satellite system (BDS) receivers, Galileo receivers, Globalnaya Navigazionnaya Sputnikovaya Sistema (GLONASS) receivers, Navigation Indian Constellation (NavIC) receivers, quasi-zenith satellite system (QZSS) receivers, Wi-Fi positioning system (WPS) receivers, cellular network positioning system receivers, beacon positioning receivers, short-range wireless beacon positioning receivers, personal area network (PAN) positioning receivers, wide area network (WAN) positioning receivers, wireless local area network (WLAN) positioning receivers, other types of positioning receivers, other types of sensors discussed herein, or combinations thereof. In some examples, one or more sensors 305 may include Figure 2 any combination of sensors of XR system 200.
[0082] Figure 3 The SLAM system 300 of includes a visual-inertial odometry (VIO) tracker 315. The term visual-inertial odometry may also be referred to as visual odometry herein. The VIO tracker 315 receives sensor data 365 from one or more sensors 305. For example, the sensor data 365 may include one or more images captured by one or more cameras 310. The sensor data 365 may include other types of sensor data from one or more sensors 305, such as data from any type of sensor 305 listed herein. For example, the sensor data 365 may include IMU data from one or more inertial measurement units (IMUs) of one or more sensors 305.
[0083] After receiving sensor data 365 from one or more sensors 305, the VIO tracker 315 performs feature detection, extraction, and / or tracking using the feature tracking engine 320 of the VIO tracker 315. For example, in a case where the sensor data 365 includes one or more images captured by one or more cameras 310 of the SLAM system 300, the VIO tracker 315 can identify, detect, and / or extract features in each image. Features can include visually distinct points in an image, such as portions of the image depicting edges and / or corners. The VIO tracker 315 can periodically and / or continuously receive sensor data 365 from one or more sensors 305, such as by continuously receiving more images from one or more cameras 310 when the one or more cameras 310 capture video, where the images are video frames of the video. The VIO tracker 315 can generate descriptors of the features. The feature descriptors can be generated at least in part by generating a description of the features, as depicted in local image patches extracted around the features. In some examples, the feature descriptors can describe the features as a collection of one or more feature vectors. The VIO tracker 315 (optionally having a map building engine 330 and / or a relocalization engine 355 in some cases) can associate multiple features with a map of the environment based on such feature descriptors. The feature tracking engine 320 of the VIO tracker 315 can perform feature tracking by identifying features in each image that the VIO tracker 315 has previously identified in one or more previous images (in some cases, based on identifying features with matching feature descriptors in different images). The feature tracking engine 320 can track changes in the one or more positions depicting the features in each different image. For example, the feature extraction engine can detect a specific corner of a room depicted in the left side of a first image captured by a first camera among the cameras 310. The feature extraction engine can detect the same feature (e.g., the same specific corner of the same room) depicted in the right side of a second image captured by the first camera. The feature tracking engine 320 can identify that the features detected in the first image and the second image are two depictions of the same feature (e.g., the same specific corner of the same room), and that the feature appears at two different positions in the two images. The VIO tracker 315 can determine that the first camera has moved based on the same feature appearing on the left side of the first image and the right side of the second image, e.g., if the feature (e.g., the specific corner of the room) depicts a static portion of the environment.
[0084] The VIO tracker 315 may include a sensor integration engine 325. The sensor integration engine 325 may use sensor data from other types of sensors 305 (in addition to the camera 310) to determine information that the feature tracking engine 320 may use when performing feature tracking. For example, the sensor integration engine 325 may receive IMU data from the IMU of one or more sensors 305 (e.g., which may be included as part of the sensor data 365). Based on the IMU data in the sensor data 365, the sensor integration engine 325 may determine that from the acquisition or capture of the first image by the first camera in the camera 310 to the acquisition or capture of the second image, the SLAM system 300 has rotated 15 degrees in the clockwise direction. Based on this determination, the sensor integration engine 325 may identify that a feature depicted at a first position in the first image is expected to appear at a second position in the second image, and the second position is expected to be a predetermined distance to the left of the first position (e.g., a predetermined number of pixels, inches, centimeters, millimeters, or another distance metric). The feature tracking engine 320 may consider this expectation when tracking features between the first image and the second image.
[0085] Based on feature tracking performed by the feature tracking engine 320 and / or sensor integration performed by the sensor integration engine 325, the VIO tracker 315 can determine the 3D feature position 373 of a specific feature. The 3D feature position 373 can include one or more 3D feature positions and can also be referred to as 3D feature points. The 3D feature position 373 can be a set of coordinates along three different axes perpendicular to each other, such as an X coordinate along the X axis (e.g., in the horizontal direction), a Y coordinate along the Y axis perpendicular to the X axis (e.g., in the vertical direction), and a Z coordinate along the Z axis perpendicular to both the X axis and the Y axis (e.g., in the depth direction). The VIO tracker 315 can also determine one or more key frames 370 (hereinafter referred to as key frames 370) corresponding to a specific feature. The key frame (from one or more key frames 370) corresponding to a specific feature can be an image that clearly depicts the specific feature. In some examples, the key frame (from one or more key frames 370) corresponding to a specific feature can be an image that clearly depicts the specific feature. In some examples, the key frame corresponding to a specific feature can be an image that reduces the uncertainty of the 3D feature position 373 of the specific feature when considered by the feature tracking engine 320 and / or the sensor integration engine 325 for determining the 3D feature position 373. In some examples, the key frame corresponding to a specific feature also includes data associated with the pose 385 of the SLAM system 300 and / or the camera 310 during the capture of the key frame. In some examples, the VIO tracker 315 can transmit the 3D feature position 373 and / or the key frame 370 corresponding to one or more features to the map building engine 330. In some examples, the VIO tracker 315 can receive a map slice 375 from the map building engine 330. The VIO tracker 315 can use the information within the map slice 375 as features for feature tracking using the feature tracking engine 320.
[0086] Based on feature tracking performed by the feature tracking engine 320 and / or sensor integration performed by the sensor integration engine 325, the VIO tracker 315 can determine the pose 385 of the SLAM system 300 and / or the camera 310 during the capture of each image in the sensor data 365. The pose 385 can include the position of the SLAM system 300 and / or the camera 310 in 3D space, such as a set of coordinates along three different axes perpendicular to each other (e.g., X coordinate, Y coordinate, and Z coordinate). The pose 385 can include the orientation of the SLAM system 300 and / or the camera 310 in 3D space, such as pitch, roll, yaw, or some combination thereof. In some examples, the VIO tracker 315 can transmit the pose 385 to the relocalization engine 355. In some examples, the VIO tracker 315 can receive the pose 385 from the relocalization engine 355.
[0087] The SLAM system 300 further includes a map construction engine 330. The map construction engine 330 generates a 3D map of the environment based on the 3D feature positions 373 and / or key frames 370 received from the VIO tracker 315. The map construction engine 330 may include a map densification engine 335, a key frame remover 340, a beam adjuster 345, and / or a loop closure detector 350. The map densification engine 335 may perform map densification, which in some examples increases the number and / or density of 3D coordinates describing the map geometry. The key frame remover 340 may remove key frames and / or, in some cases, add key frames. In some examples, the key frame remover 340 may remove key frames 370 corresponding to regions in the map to be updated and / or having a low corresponding confidence value. In some examples, the beam adjuster 345 may refine the 3D coordinates describing the scene geometry, the parameters of the relative motion, and / or the optical characteristics of the image sensor used to generate the frames, according to an optimality criterion involving the corresponding image projections of all points. The loop closure detector 350 may identify when the SLAM system 300 has returned to a previously mapped area and may use such information to update map slices and / or reduce the uncertainty of certain 3D feature points or other points in the map geometry. The map construction engine 330 may output map slices 375 to the VIO tracker 315. The map slices 375 may represent a 3D portion or subset of the map. The map slices 375 may include map slices 375 representing new, previously unmapped areas of the map. The map slices 375 may include map slices 375 representing updates (or modifications or revisions) to previously mapped areas of the map. The map construction engine 330 may output map information 380 to the relocalization engine 355. The map information 380 may include at least a portion of the map generated by the map construction engine 330. The map information 380 may include one or more 3D points constituting the geometry of the map, such as one or more 3D feature positions 373. The map information 380 may include one or more key frames 370 corresponding to certain features and certain 3D feature positions 373.
[0088] The SLAM system 300 also includes a relocalization engine 355. The relocalization engine 355 can perform relocalization, for example, when the VIO tracker 315 fails to identify a threshold number of features in an image, and / or when the VIO tracker 315 loses track of the pose 385 of the SLAM system 300 within a map generated by the map construction engine 330. The relocalization engine 355 can perform relocalization by performing extraction and matching using the extraction and matching engine 360. For example, the extraction and matching engine 360 can extract features from an image captured by the camera 310 of the SLAM system 300 when the SLAM system 300 is in the current pose 385, and can match the extracted features with features depicted in different key frames 370, identified by 3D feature positions 373, and / or identified in the map information 380. By matching these extracted features with previously identified features, the relocalization engine 355 can identify that the pose 385 of the SLAM system 300 is the pose 385 at which the previously identified features are visible to the camera 310 of the SLAM system 300, and thus is similar to one or more previous poses 385 at which the previously identified features were visible to the camera 310. In some cases, the relocalization engine 355 can perform relocalization based on wide baseline map construction or the distance between the current camera position and the camera position at which the features were initially captured. The relocalization engine 355 can receive information about the pose 385 (e.g., information about one or more recent poses of the SLAM system 300 and / or the camera 310) from the VIO tracker 315, and the relocalization engine 355 can base its relocalization determination on this information. Once the relocalization engine 355 relocalizes the SLAM system 300 and / or the camera 310 and thus determines the pose 385, the relocalization engine 355 can output the pose 385 to the VIO tracker 315.
[0089] In some examples, the VIO tracker 315 may modify the image in the sensor data 365 before performing feature detection, extraction, and / or tracking on the modified image. For example, the VIO tracker 315 may rescale and / or resample the image. In some examples, rescaling and / or resampling the image may include shrinking, downsampling, sub-sampling, and / or quadratic sampling the image one or more times. In some examples, the VIO tracker 315 that modifies the image may include converting the image from color to grayscale, or from color to black and white, for example, by desaturating the colors in the image, removing certain color channels, reducing the color depth in the image, replacing the colors in the image, or a combination thereof. In some examples, the VIO tracker 315 modifying the image may include the VIO tracker 315 masking certain regions of the image. A dynamic object may include an object that has a changing appearance between one image and another. For example, a dynamic object may be an object that moves within the environment, such as a person, a vehicle, or an animal. A dynamic object may be an object that has a changing appearance at different times, such as a display screen that may show different things at different times. A dynamic object may be an object that has a changing appearance based on the pose of the camera 310, such as a reflective surface, a prism, or a mirror surface that reflects, refracts, and / or scatters light in different ways depending on the position of the camera 310 relative to the dynamic object. The VIO tracker 315 may detect dynamic objects using face detection, face recognition, face tracking, object detection, object recognition, object tracking, or a combination thereof. The VIO tracker 315 may detect dynamic objects using one or more artificial intelligence algorithms, one or more trained machine learning models, one or more trained neural networks, or a combination thereof. The VIO tracker 315 may mask one or more dynamic objects in the image by overlaying a mask over a region of the image that includes a depiction of one or more dynamic objects. The mask may be an opaque color (such as black). The region may be a bounding box having a rectangular or other polygonal shape. The region may be determined on a pixel-by-pixel basis.
[0090] Figure 4 Illustrates an example of the augmented reality enhancement application engine 400. In the illustrative example, the augmented reality enhancement application engine 400 includes a simulation engine 405, a rendering engine 410, a main rendering module 415, and an AR rendering module 460. As illustrated, the main rendering module 415 may include an effect rendering engine 420, a post-processing engine 425, and a user interface (UI) rendering engine 430. The AR rendering module 460 may include an AR effect rendering engine 465 and an AR UI rendering engine 470. It should be noted that Figure 4 the components 405-470 shown are provided as non-limiting examples for illustrative and explanatory purposes, and other examples may include more, fewer, or different components compared to Figure 4 the components shown.
[0091] In some cases, the augmented reality enhancement application engine 400 is included in the electronic device 440 and / or communicates (wired or wirelessly) with the electronic device. In some examples, the augmented reality enhancement application engine 400 is included in the XR system 450 and / or communicates (wired or wirelessly) with the XR system.
[0092] In Figure 4 In the illustrated example, the simulation engine 405 may generate a simulation for the augmented reality enhancement application engine 400. In some cases, the simulation may include, for example, one or more images, one or more videos, one or more strings (e.g., alphanumeric characters, numbers, text, Unicode characters, symbols, and / or icons), one or more two-dimensional (2D) shapes (e.g., circles, ellipses, squares, rectangles, triangles, other polygons, circular polygons with one or more rounded corners, portions thereof, or combinations thereof), one or more three-dimensional (3D) shapes (e.g., spheres, cylinders, cubes, pyramids, triangular prisms, rectangular prisms, tetrahedrons, other polyhedra, circular polyhedra with one or more rounded edges and / or corners, portions thereof, or combinations thereof), shape textures, bump mapping of shapes, lighting effects, or combinations thereof. In some examples, the simulation may include at least a portion of an environment. The environment may be a real-world environment, a virtual environment, and / or a mixed environment including elements of a real-world environment and a virtual environment.
[0093] In some cases, the simulation generated by the simulation engine 405 may be dynamic. For example, the simulation engine 405 may update the simulation based on different triggers, which include but are not limited to physical contact, sound, gestures, input signals, passage of time, and / or any combination thereof. As used herein, at a particular moment, the application state of the augmented reality enhancement application engine 400 may include any information associated with the following: the simulation engine 405, the rendering engine 410, the main rendering module 415, the effect rendering engine 420, the post-processing engine 425, the UI rendering engine 430, the AR rendering module 460, the AR effect rendering engine 465, the AR UI rendering engine 470, inputs to the augmented reality enhancement application engine 400, outputs from the augmented reality enhancement application engine 400, and / or any combination thereof.
[0094] As illustrated, the simulation engine 405 may obtain a mobile device input 441 from the mobile device 440. In some cases, the simulation engine 405 may obtain an XR system input 451 from the XR system 450. The mobile device input 441 and / or the XR system input 451 may include, for example, user input through a user interface of an application displayed on the display of the mobile device 440, from an input device (e.g., Figure 2input device 208), one or more sensors (e.g., Figure 2 image sensor 202, accelerometer 204, gyroscope 206) of the user input. In some cases, the simulation engine 405 may update the application state of the augmented reality enhancement application engine 400 based on the mobile device input 441, the XR system input 451, and / or any combination thereof.
[0095] In Figure 4 an illustrative example, the rendering engine 410 may obtain application state information from the simulation engine 405. In some cases, the rendering engine 410 may determine the portion of the application state information that will be rendered by the display that can be used by the augmented reality enhancement application engine 400. For example, the rendering engine 410 may determine whether a connection (wired or wireless) has been established between the XR system 450 and the mobile device 440. In some cases, the rendering engine 410 may determine the application state information that will be rendered by the main rendering module 415 and the AR rendering module 460. In some cases, the rendering engine 410 may determine that the XR system 450 is not connected (wired or wirelessly) to the mobile device 440. In some cases, the rendering engine 410 may determine the application state information of the main rendering module 415 and discard the determination of the application state information that will be rendered by the AR rendering module 460 and will not be displayed. Thus, the rendering engine 410 may facilitate an adaptive rendering configuration for the augmented reality enhancement application engine 400 based on the availability and / or type of the available display. In some specific implementations, a separate rendering engine 410 as shown in Figure 4 may be excluded. In an illustrative example, the main rendering module 415 and / or the AR rendering module 460 may include at least a portion of the functionality of the above-described rendering engine 410.
[0096] The main rendering module 415 may include an effects rendering engine 420, a post-processing engine 425, and a UI rendering engine 430. In some cases, the main rendering module 415 may render image frames configured for display on the display of the mobile device 440. As illustrated, the main rendering module 415 may output the generated image frames (e.g., media content) for display on the display of the mobile device 440. In some cases, the effects rendering information may render the application state information generated by the simulation engine 405. For example, the effects rendering engine may generate a 2D projection of a portion of the 3D environment included in the application state information. For example, the rendering engine 420 may generate a perspective projection of the 3D environment through a virtual camera. In some cases, the application state information may include the pose of the virtual camera within the environment. In some cases, the effects rendering engine 420 may generate additional visual effects not included in the 3D environment. For example, the rendering engine 420 may apply texture mapping to enhance the visual appearance of the effects generated by 420. In some cases, the rendering engine 420 may exclude portions of the application state information specified by the rendering engine 410 for the AR rendering module 460. For example, the main rendering module 415 may exclude the effects present in the simulation environment.
[0097] In some cases, the post-processing engine 425 may provide additional processing to the rendered effects generated by the effects rendering engine 420. For example, the post-processing engine 425 may perform scaling, image smoothing, z-buffering, contrast enhancement, gamma, color mapping, any other image processing, and / or any combination thereof.
[0098] In some embodiments, the UI rendering engine 430 may render the UI. In some cases, in addition to the effects rendered based on the application environment (e.g., 3D environment), the user interface may also provide application state information. In some cases, the UI may be generated as an overlay on a portion of the image frames output by the post-processing engine 425.
[0099] The AR rendering module 460 may include an AR effects rendering engine 465 and an AR UI rendering engine 470. In some cases, the AR effects rendering engine 465 may render the application state information generated by the simulation engine 405. For example, the AR effects rendering engine 465 may generate a 2D projection of the 3D environment included in the application state information. In some cases, the AR effects rendering engine 465 may generate effects that appear to protrude from the display surface of the display of the mobile device 440.
[0100] In some cases, the display of the XR system 450 may have different display parameters (e.g., different resolution, frame rate, aspect ratio, and / or any other display parameter) from the display of the mobile device 440. In some cases, the display parameters may also vary between different types of output devices (e.g., different HMD models, other XR systems, etc.). Thus, rendering the display data for 450 with the AR rendering module 460 may affect the performance of the main rendering module 415 (e.g., by consuming computing resources of the GPU, CPU, memory, etc.). In some cases, including the AR rendering module 460 within the augmented reality enhancement application engine 400 may require periodic updates to provide compatibility with different devices.
[0101] Figure 5 An example of the main application engine 500 and a secondary application engine 560 that can provide augmented reality enhancement to the main application engine 500 is illustrated. In Figure 5 the illustrative example, the main application engine 500 includes a simulation engine 505, a rendering engine 515, an encoding engine 520, and a communication engine 525. In the illustrated example, the secondary application engine 560 includes a tracking engine 565 (e.g., Figure 2 the XR engine 220 of Figure 3 , the VIO tracker 315 of Figure 5 ), an AR rendering engine 575, a decoding engine 580, and a communication engine 585. As illustrated, the main application engine 500 and the secondary application engine 560 can communicate via a (wired or wireless) communication link 530. It should be noted that Figure 5 the components 505 - 525 shown for the main application engine 500 of Figure 5 are non - restrictive examples provided for illustrative and explanatory purposes, and other examples may include more, fewer, or different components compared to the components shown in Figure 5 . Similarly, it should be noted that
[0102] In Figure 5In the illustrated example, the simulation engine 505 of the main application engine 500 can generate simulations for applications on the mobile device 540. In some cases, the simulation can include, for example, one or more images, one or more videos, one or more strings (e.g., alphanumeric characters, numbers, text, unified character encoding characters, symbols, and / or icons), one or more two-dimensional (2D) shapes (e.g., circles, ellipses, squares, rectangles, triangles, other polygons, circular polygons with one or more rounded corners, parts thereof, or combinations thereof), one or more three-dimensional (3D) shapes (e.g., spheres, cylinders, cubes, pyramids, triangular prisms, rectangular prisms, tetrahedrons, other polyhedrons, circular polyhedrons with one or more rounded edges and / or corners, parts thereof, or combinations thereof), shape textures, bump mapping of shapes, lighting effects, or combinations thereof. In some examples, the simulation can include at least a portion of an environment. The environment can be a real-world environment, a virtual environment, and / or a hybrid environment that includes elements of the real-world environment and virtual environment elements.
[0103] In some cases, the simulations generated by the simulation engine 505 can be dynamic. For example, the simulation engine 505 can update the simulation based on different triggers, which include but are not limited to physical contact, sound, gestures, input signals, passage of time, and / or any combination thereof. As used herein, at a particular moment, the application state of the main application engine 500 can include any information associated with the simulation engine 505, the effect rendering engine 515, the communication engine 525, and / or any combination thereof.
[0104] The rendering engine 515 can correspond to Figure 5 the main rendering module 415 and perform similar functions. For example, the rendering engine 515 can include a module for effect rendering (e.g., Figure 4 the rendering engine 420), a module for post-processing (e.g., Figure 4 the post-processing engine 425), and / or a module for UI rendering (e.g., Figure 4 the UI rendering engine 430).
[0105] The communication engine 525 of the main application engine 500 and the communication engine 585 of the auxiliary application engine 560 can communicate via a communication link 530. In some cases, the communication link 530 can be bidirectional. In some examples, the communication engine 525 can send application state information (e.g., from the simulation engine 505) to the communication engine 585 of 560. In some cases, the application state information can include information that can be used to generate AR effects. In some examples, the application state information can include data that can be used by the auxiliary application engine 560 to generate an AR UI. In some cases, the communication engine 525 can also send inputs obtained from the mobile device 540 to the communication engine 585 via the communication link 530. In some cases, the communication engine 585 of the auxiliary application engine 560 can send pose information, connectivity status, user input, etc. to the communication engine 525 of the main application engine 500. The communication engine 525 and the communication engine 585 can also send and / or receive synchronization signals for synchronizing the displays between the display of the mobile device 540 and the display of the HMD 550. The examples of the communication between the communication engine 525 and the communication engine 585 provided herein are non-limiting and are provided as examples. In some cases, more, less, and / or different information can be transmitted over the communication link 530 without departing from the scope of the present disclosure. Although the HMD 550 is used herein as an illustrative example of an XR device, the systems and techniques can be used for any type of XR device, such as AR, VR, or MR glasses.
[0106] Referring to the auxiliary application engine 560, the tracking engine 565 can use information captured by sensors (e.g., Figure 2 the image sensor 202, accelerometer 204, gyroscope 206, Figure 3 one or more sensors 305, cameras 310, etc.) of
[0107] The AR rendering module 460 can be similar to Figure 4 the AR rendering module 460 of Figure 4The AR effect rendering engine 465) and / or the AR UI rendering engine (e.g., Figure 4 the AR UI rendering engine 470). In some cases, the AR rendering engine 575 may output AR media content having display parameters (e.g., different resolutions, frame rates, aspect ratios, and / or any other display parameters) different from the media content output from the rendering engine 515 to the mobile device 540 to the HMD 550. In some cases, by dividing the rendering functionality between the main application engine 500 and the auxiliary application engine 560, the computing resources for providing the AR enhanced application experience can be shared among the computing resources of multiple devices such as the mobile device 540 and the HMD 550. Additionally, providing the independent AR rendering engine 575 in the auxiliary application engine 560 can simplify the development of the main application engine 500. For example, the rendering engine 515 of the main application engine 500 may not need to maintain compatibility with various different mobile devices having different display configurations.
[0108] In some cases, to allow the HMD 550 to be wearable, the HMD 550 may be relatively constrained in terms of battery and processing power compared to the mobile device 540. To reduce the processing requirements on the HMD 550, the images to be displayed by the HMD 550 may be rendered by the mobile device 540 and sent to the HMD 550 via the communication link 530. In some cases, the HMD may receive multiple images simultaneously for display to the user. For example, the rendering engine 515 of the mobile device 540 may render the left-eye image, the right-eye image, and in some cases provide depth information. For example, the depth information may include information indicating the distance of a point in the scene (e.g., a point corresponding to the surface of an object) from the viewpoint (such as the camera viewpoint). In some cases, the depth information may be inferred based on, for example, the difference between the left-eye image and the right-eye image received by the HMD 550. In some cases, the depth information may be used to distort (e.g., apply a displacement vector to) portions of the image to help adjust for the movement of an object that can move independently of the camera (such as the camera on the HMD 550) between the time the image is rendered by the rendering engine 515 and the time the image is received by the HMD 550. The rendered images may be in any known image or video format. In some cases, the images may include only the objects that will be overlaid on the environment visible through the HMD 550. The encoding engine 520 may encode the rendered images to reduce the size of the images for transmission. The encoded images may be sent to the HMD 550 via the communication engine 525 and the communication link 530.
[0109] The HMD may receive the encoded images via the communication engine 585. For example, these received images may then be decoded by the decoding engine 580. In some cases, there may be latency (e.g., display latency) introduced by the rendering, encoding, sending, receiving, and decoding processes, and during this display latency, the user may, for example, move the HMD 550. This movement may not be taken into account by the images rendered by the rendering engine 515, and any objects in the rendered images may be displayed in a different position than expected due to this movement. To account for the potential movement of the HMD 550, the AR rendering engine 575 may distort the received images based on the pose and / or tracking information from the tracking engine 565 that describes the movement of the HMD 550. In some cases, the techniques discussed herein may be applicable to other scenarios where there is a sufficient lag during which noticeable movement may have occurred between when the images are rendered and when those images are displayed.
[0110] Figure 6A and Figure 6B are images illustrating example scenarios for an XR system in accordance with aspects of the present disclosure. Figure 6A Illustrates an image 600 that may be partially rendered by a mobile device. For example, the mobile device may render a first virtual object (such as a tree 602) and overlay the tree 602 in a first position in the environment. The mobile device may also render a second virtual object (such as a globe 604) that is overlaid on the user's hand 606 in the environment. In some cases, the hand 606 may physically exist within the environment and be within the field of view of the image 600. In some cases, the hand may be rendered or overlaid in the image. Once rendered, the image 600 may be encoded and sent to the HMD 550 for display.
[0111] While the image 600 is being rendered, encoded, sent, received, and decoded, the HMD may move in a rightward direction (e.g., the user may move). Since this movement did not occur when the image 600 was rendered, the image 600 does not account for the movement. In some cases, the HMD may distort the image 600 and display Figure 6B the distorted image 650. In some cases, the image 600 may be distorted into the image 650 based at least in part on the pose and / or tracking information indicating how the HMD has moved. In the distorted image 650, the tree 652 corresponding to the tree 602 has been distorted such that the tree 602 is displayed in the first position regardless of the movement of the HMD.
[0112] In some cases, while movement of the HMD may be considered, information about how the HMD has moved may not be considered for all possible movements in the environment. In some cases, the user's limbs may be visible in the environment, and the movement of these limbs may not align with how the HMD has moved. As an example, the user's hand 656 may be visible in the environment, and while the HMD may move to the right during the latency period, the user's hand 656 may have moved to the left during the latency period. If an object (such as the earth 654) is distorted based on information indicating how the HMD has moved, the earth 654 will no longer appropriately overlay the user's hand 656. That is, the user may move their hand in a different direction than their head (and thus the HMD), and the information describing how the HMD has moved may not apply to their hand.
[0113] In some cases, a mask layer may be used to allow adjustment of the rendered image for the independent movement of physical objects in the scene (e.g., objects visible in the environment and the field of view) and the movement of the HMD. Figure 7 An example mask layer 700 in accordance with aspects of the present disclosure is illustrated. In some cases, the mask layer may include bits corresponding to each pixel of the image. The value of the mask layer may indicate the portion of the image that includes virtual objects associated with physical objects in the field of view of the image that may move independently of the HMD. For example, if a virtual object in the image (such as Figure 6A the earth 604) should be projected on / into / around a physical object (e.g., a hand) in the field of view of the image to be displayed on the HMD (e.g., in the image), the portion of the mask layer corresponding to the pixels of the virtual object may be indicated in the mask layer. Since this physical object (e.g., a hand) may move independently of the HMD, the virtual object (e.g., the earth 604) may also move independently of the HMD.
[0114] In the mask layer 700, a value of 1 may correspond to pixels of the image that form a virtual object (e.g., the earth 604) that may move independently of the HMD, and a value of 0 may correspond to pixels that do not form a virtual object that may move independently of the HMD. In some cases, multiple mask layers may be generated. For example, a mask layer may be generated for each object associated with a physical object in the field of view of the image to be displayed by the HMD that may move independently of the HMD. Thus, if the user's two hands are visible in the image and both hands are associated with virtual objects, two mask layers may be generated, one for each hand.
[0115] In some cases, the mask layer may include an object identifier that indicates a virtual object for which the mask layer is referenced. For example, the object identifier may be multiple bits, where different bits among the multiple bits correspond to different virtual objects. For example, the mask layer may include a 2-bit value, where the first bit corresponds to a first virtual object and the second bit corresponds to a second virtual object. In such an example, the bit value 01 may correspond to pixels forming a first virtual object (e.g., the earth 604) that can move independently of the HMD as the left hand moves, and the bit value 10 may correspond to pixels forming a second virtual object that can move independently of the HMD as the right hand moves. Similarly, the value 00 may correspond to pixels forming a virtual object that can move with the HMD. As another example, the object identifier may be a unique code (e.g., a number) for each object, and the mask layer may include this number corresponding to different virtual objects. The bit values included in the mask layer may be based on the number of different virtual objects that can move independently of the HMD. For example, two bits can be used to represent four virtual objects, three bits can be used to represent eight objects, etc. In this way, the mask layer can account for multiple independent movements in a scene (e.g., in a single mask layer).
[0116] Figure 8 is a block diagram 800 illustrating an example process for determining low-latency and low-power independent scene movement in an XR system based on a mask layer. In some cases, the XR system may include an accessory device 802 and an XR display device 804. In some cases, the accessory device 802 can be any device capable of executing an application engine (such as Figure 5 the main application engine 500). Examples of the accessory device 802 may include a mobile device, a personal computer, a tablet computer, etc. The accessory device 802 is coupled to the XR display device 804. The XR display device 804 can be any XR device (such as the HMD 550) capable of displaying images and executing an auxiliary application engine (such as the auxiliary application engine 560). The accessory device 802 may be coupled to the XR display device 804 via a wired or wireless network (such as Ethernet, Wi-Fi, etc.).
[0117] In some cases, compared to the XR display device 804, the companion device 802 can be a relatively higher-power device, and the companion device 802 can generate images for the XR display device 804 to display. To generate images for the XR device to display, the companion device can render 806 the set of images to be displayed. The rendered images can be based on the positions of the tracked physical objects. For example, the XR display device 804 can include one or more sensors or SLAM systems (such as the SLAM system 300), which are capable of tracking and determining the pose and / or tracking information of the XR display device 804, as well as the pose and / or tracking information of the user's limbs (such as hands, fingers, legs, feet, etc.). In some cases, it can be determined that the user's limbs can move independently of the movement of the HMD. In some cases, it can also be determined that other physical objects separate from the user tracked by the XR system can move independently of the movement of the HMD. In some cases, the determination of whether a physical object can move independently of the movement of the HMD can be based on whether pose and / or tracking information is available for the physical object.
[0118] If a virtual object is to be rendered and the virtual object is associated with a physical object that can move independently, a mask layer associated with the physical object can be generated, for example, as part of the rendering 806. For example, the virtual object can be represented as a mesh of vertices in a virtual 3D space, and these vertices can be rasterized into pixels of a 2D image for display. If a part of the virtual object that can move independently is rasterized into pixels, the corresponding positions in the mask layer can be updated to indicate that the corresponding pixels can move independently. In some cases, the XR display device can be capable of generating a 3D view by displaying separate left and right images. In such cases, the companion device can render the left image 808A and the right image 808B (collectively referred to as the images 808). In addition, when a virtual object that can move independently is to be displayed, the companion device can also render a set of left mask layers 810A and a set of left mask layers 810B (collectively referred to as the mask layers 810). In some cases, the left depth information 812A and the right depth information 812B (collectively referred to as the depth information 812) can also be generated as part of the rendering 806. After the rendering 806, the generated image information (such as images, mask layers, and depth information) can be encoded 814 and packetized 816 (such as split into data packets) for transmission to the XR display device 804.
[0119] The XR display device 804 may receive the generated image information, demultiplex 818 it, and decode 820 it. In some cases, after decoding 820, the generated image information may be passed to the AR rendering engine 822 for processing. In some cases, the AR rendering engine 822 may include a warping engine 824 and a motion compensation engine 826. In some cases, the AR rendering engine 822 may receive updated HMD pose information 828 and updated object pose information 830. The updated HMD pose information 828 and the updated object pose information 830 may include pose information relative to the HMD pose and object pose at the time of rendering the image 808.
[0120] In some cases, the motion compensation engine 826 may determine the amount of warping of the warped image 808 based on the updated HMD pose information 828, the updated object pose information 830, and the depth information 812, and the warping engine 824 may perform the warping of the image. For example, the motion compensation engine 826 may determine where an object should move in the image 808 given a change in the position of the HMD determined based on the HMD pose information 828, a change in the position of the object based on the object pose information 830, and how far the object is from the HMD based on the depth information.
[0121] In some cases, the XR device or its components (such as the XR display device 804) may use temporal warping (e.g., asynchronous reprojection) to correct rotational motion between frames and use optical flow to correct translational motion between frames. For example, warping can generally correct rotational motion efficiently and accurately. Thus, the XR device may use warping to correct rotational motion in an image. The XR device may reproject an image that has translational motion but no rotational motion. For example, the XR device may subtract the rotational motion between images from the estimated motion and use the resulting translational motion to estimate the warped image. The XR device may use warping to correct rotational motion in the warped image. In some examples, the XR device may determine the pose of the XR device when the rendered image is rendered and use the pose to infer the warped image. Before determining and / or applying optical flow for the warped image, the XR device may use the pose information to reproject the rendered image with no rotational motion.
[0122] In some cases, an XR device may correct and / or project different types of motion. For example, the XR device may correct and / or project rotational head movement (e.g., three degrees of freedom or "3DOF"). Rotational movement may include rotation about each axis. The XR device may use warping to correct rotational movement by applying a rotation transformation to the rendered image. Additionally, the XR device may correct and / or project translational head movement (e.g., six degrees of freedom or "6DOF"). Translational movement may include translation along each axis. The XR device may also correct and / or project object movement in a scene. Object movement may correspond to an object that is in motion in the scene and has changed its position between frames. In some examples, the motion estimation described herein may cover rotational movement (e.g., 3DOF), translational movement (e.g., 6DOF), and object movement.
[0123] In some cases, for image extrapolation, such as to warp certain objects, the XR device may texture map the rendered image to a mesh of geometry. The XR device may move the positions of vertices in the mesh based on, for example, updated object pose information 830 and depth information 812 of the object. This may shift and manipulate them according to the motion of the region of the camera frame corresponding to the object to move the object into its position indicated by the pose information.
[0124] In some cases, the XR device may perform image extrapolation by moving the geometry or by moving the UV texture coordinates. For example, the XR device may perform an extrapolation rendering process by moving the geometry of the mesh or manipulating the UV texture coordinates to change how the source rendered image is mapped to the geometry. In some cases, the XR device may distort the mesh by moving the geometry of the mesh. In some examples, when moving the geometry, the XR device may move each vertex during the vertex shader process based on the magnitude and direction of the motion vector at the corresponding point. In some examples, when moving the texture coordinates, the XR device may determine the inverse movement and use the inverse to adjust the texture coordinates.
[0125] In some cases, the AR rendering engine 822 may include one or more GPUs. In some cases, the AR rendering engine 822 may construct a pseudo 3D representation of the scene shown in the image 808 based on the image 808 and depth information 812 via one or more GPUs. In some examples, the vertices of the objects in the image 808 may be determined based on the detected edges in the image 808, and those vertices may be associated with the objects based on the mask layer 810. In some cases, one or more GPUs may include vertex and fragment shaders. In some cases, the fragment shader may texture sample a portion of the image 808 based on the mask layer 810. These sampled textures may act as filters to select the relevant fragment coordinates corresponding to the virtual object or HMD movement. These fragment coordinates may be used to calculate the motion vectors for warping the image. The vertex shader may apply a transformation to the vertices associated with the object based on the amount of motion compensation determined for the object (e.g., from the motion compensation engine 826). The vertex shader may also apply a transformation to the vertices associated with those objects not associated with independent movement based on the amount of motion compensation determined for the HMD. In some cases, the vertex shader may texture sample the mask layer 810. These sampled textures may act as filters to select the relevant model transformations corresponding to the virtual object or HMD movement. These transformed vertex coordinates may be used to calculate the motion vectors for warping the image. In some cases, one or more GPUs may use the mask layer to merge multiple motion maps generated by other computational engines in the AR rendering engine 822 into a single motion map. After warping the image 808 based on the mask layer 810, the image may be sent to the display 832 for display.
[0126] Figure 9 FIG. 900 is a flow chart illustrating a process 900 for image generation in accordance with aspects of the present disclosure. The process 900 may be executed by a computing device (or apparatus) or a component of a computing device (e.g., a chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, or other types of computing devices. The operations of the process 900 may be implemented as software components executed and run on one or more processors.
[0127] At block 902, a computing device (or its component) may render an image of a scene that includes a first virtual object associated with a first physical object that moves independently of the scene. In some cases, a display is coupled to the computing device but separate from the computing device. The computing device (or its component) may determine that the first physical object can move independently of the scene. To determine that the first physical object can move independently of the scene, the computing device (or its component) may receive an indication that the display device has stored pose information associated with the first physical object. The computing device (or its component) may determine that the first physical object can move independently of the scene based on the received indication.
[0128] At block 904, the computing device (or its component) may identify a first set of pixels associated with the first virtual object in the image of the scene.
[0129] At block 906, the computing device (or its component) may generate a mask layer based on the first set of pixels that indicates the location of the first set of pixels in the image. In some cases, the mask layer includes object identifiers corresponding to the pixels of the image. To generate the mask layer, the computing device (or its component) may set the locations in the mask layer corresponding to the identified first set of pixels with the object identifier associated with the first virtual object. The computing device (or its component) may generate depth information for the first virtual object. The computing device (or its component) may send the depth information, along with the mask layer and the image, to the display. In some cases, the scene includes a second virtual object associated with a second physical object that moves independently of the scene. The computing device (or its component) may render an image of the scene that includes the second virtual object. The computing device (or its component) may identify a second set of pixels associated with the second virtual object in the image of the scene. The computing device (or its component) may generate a mask layer based on the first set of pixels and the second set of pixels.
[0130] At block 908, the computing device (or its component) may send the mask layer and the image to the display. In some cases, the computing device may be a mobile phone.
[0131] Figure 10 FIG. 13 is a flowchart illustrating a process 1000 for image generation in accordance with aspects of the present disclosure. Process 1000 may be performed by a computing device (or apparatus) or a component of a computing device (such as a chipset, codec, etc.). The computing device may be a mobile device (such as a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, or other types of computing devices. The operations of process 1000 may be implemented as software components executed and run on one or more processors.
[0132] At block 1002, a computing device (or its components) may obtain a mask layer and an image associated with the mask layer, where the mask layer is associated with a first virtual object. In some cases, the mask layer includes object identifiers corresponding to pixels of the image.
[0133] At block 1004, a computing device (or its components) may generate updated pose information for the device, where the updated pose information indicates a movement between the time the image was generated and the current time.
[0134] At block 1006, a computing device (or its components) may obtain updated pose information for a first physical object, where the first physical object moves independently of the device and where the first physical object is associated with the first virtual object. The computing device (or its components) may obtain first pose information for the first physical object. The computing device (or its components) may generate second pose information for the device. The computing device (or its components) may send the first pose information and the second pose information to a companion device for rendering an image for the device.
[0135] At block 1008, a computing device (or its components) may determine a first amount of distortion for the image based on the updated pose information for the device.
[0136] At block 1010, a computing device (or its components) may determine a second amount of distortion for a first portion of the image based on the mask layer and the updated pose information for the first physical object.
[0137] At block 1012, a computing device (or its components) may distort the image based on the first amount of distortion for the image and distort a first portion of the image based on the second amount of distortion of the first virtual object by the first physical object. The computing device (or its components) may receive depth information associated with the first virtual object along with the mask layer and the image. The computing device (or its components) may further determine the second amount of distortion for the first portion of the image based on the depth information. In some cases, the mask layer is associated with a second virtual object. The computing device (or its components) may obtain updated pose information for a second physical object. The computing device (or its components) may determine a third amount of distortion for a second portion of the image based on the mask layer and the updated pose information for the second physical object. The computing device (or its components) may distort the second portion of the image based on the third amount.
[0138] At block 1014, a computing device (or its components) may display a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0139] Figure 11AFIG. 1100 is a perspective view illustrating a head-mounted display (HMD) 1110 that performs feature tracking and / or visual simultaneous localization and mapping (VSLAM) according to some examples. The HMD 1110 can be, for example, an augmented reality (AR) headset, a virtual reality (VR) headset, a mixed reality (MR) headset, an extended reality (XR) headset, or some combination thereof. The HMD 1110 can be an example of an XR system 200, a SLAM system 300, or a combination thereof. The HMD 1110 includes a first camera 1130A and a second camera 1130B along a front portion of the HMD 1110. The first camera 1130A and the second camera 1130B can be both of one or more cameras 310. In some examples, the HMD 1110 can have only a single camera. In some examples, in addition to the first camera 1130A and the second camera 1130B, the HMD 1110 can include one or more additional cameras. In some examples, in addition to the first camera 1130A and the second camera 1130B, the HMD 1110 can include one or more additional sensors.
[0140] Figure 11B illustrating a head-mounted display (HMD) 1110 that performs feature tracking and / or visual simultaneous localization and mapping (VSLAM) according to some examples Figure 11A FIG. 1130 is a perspective view of the HMD 1110 being worn by a user 1120. The user 1120 wears the HMD 1110 on the user 1120's head above the user 1120's eyes. The HMD 1110 can capture images via the first camera 1130A and the second camera 1130B. In some examples, the HMD 1110 displays one or more display images based on the images captured by the first camera 1130A and the second camera 1130B toward the user 1120's eyes. The display images can provide a stereoscopic view of the environment, in some cases with overlaid information and / or with other modifications. For example, the HMD 1110 can display a first display image to the user 1120's right eye, the first display image being based on the image captured by the first camera 1130A. The HMD 1110 can display a second display image to the user 1120's left eye, the second display image being based on the image captured by the second camera 1130B. For example, the HMD 1110 can provide overlay information in the display images that is overlaid on the images captured by the first camera 1130A and the second camera 1130B.
[0141] The HMD 1110 may not include wheels, propellers, or other conveyance means of its own. Instead, the HMD 1110 relies on the movement of the user 1120 to move the HMD 1110 back and forth in the environment. In some cases, such as when the HMD 1110 is a VR headset, the environment can be fully or partially virtual. If the environment is at least partially virtual, the movement through the virtual environment can also be virtual. For example, the movement through the virtual environment can be controlled by the input device 208. The movement actuator can include any such input device 208. The movement through the virtual environment may not require wheels, propellers, legs, or any other form of conveyance means. Even if the environment is virtual, SLAM technology may still be valuable because the virtual environment can be deconstructed and / or generated by devices other than the HMD 1110 (such as a remote server or console associated with a video game or video game platform). In some cases, feature tracking and / or SLAM can even be performed by a vehicle or other device in the virtual environment, where the vehicle or other device has its own physical conveyance system that allows it to move physically back and forth in the physical environment. For example, SLAM can be performed in the virtual environment to test whether the SLAM system 300 is working properly without wasting time or energy during movement and without wearing out the physical conveyance system.
[0142] Figure 12AFIG. 1200 is a perspective view of a front surface 1255 of a mobile device 1250 that illustrates using one or more front cameras 1230A-B to perform features described herein, including, for example, feature tracking and / or visual simultaneous localization and mapping (VSLAM). The mobile device 1250 can be, for example, a cellular phone, a satellite phone, a portable game console, a music player, a health tracking device, a wearable device, a wireless communication device, a laptop device, a mobile device, any other type of computing device or computing system 1300 discussed herein, or a combination thereof. The front surface 1255 of the mobile device 1250 includes a display screen 1245. The front surface 1255 of the mobile device 1250 includes a first camera 1230A and a second camera 1230B. The first camera 1230A and the second camera 1230B are illustrated in a bezel around the display screen 1245 on the front surface 1255 of the mobile device 1250. In some examples, the first camera 1230A and the second camera 1230B can be positioned in a notch or cutout cut out from the display screen 1245 on the front surface 1255 of the mobile device 1250. In some examples, the first camera 1230A and the second camera 1230B can be under-display cameras located between the display screen 1245 and the remainder of the mobile device 1250 such that light passes through a portion of the display screen 1245 before reaching the first camera 1230A and the second camera 1230B. The first camera 1230A and the second camera 1230B of the perspective view 1200 are front cameras. The first camera 1230A and the second camera 1230B face a direction that is perpendicular to a planar surface of the front surface 1255 of the mobile device 1250. The first camera 1230A and the second camera 1230B can be both of one or more cameras 310. In some examples, the front surface 1255 of the mobile device 1250 can have only a single camera. In some examples, in addition to the first camera 1230A and the second camera 1230B, the mobile device 1250 can include one or more additional cameras. In some examples, in addition to the first camera 1230A and the second camera 1230B, the mobile device 1250 can include one or more additional sensors.
[0143] Figure 12BFIG. 12120 is a perspective view of a rear surface 1265 of an exemplary mobile device 1250. The mobile device 1250 includes a third camera 1230C and a fourth camera 1230D on the rear surface 1265 of the mobile device 1250. The third camera 1230C and the fourth camera 1230D in FIG. 1290 are rear-facing. The third camera 1230C and the fourth camera 1230D are oriented in a direction perpendicular to a planar surface of the rear surface 1265 of the mobile device 1250. Although the rear surface 1265 of the mobile device 1250 does not have a display screen 1245 as illustrated in FIG. 1290, in some examples, the rear surface 1265 of the mobile device 1250 may have a second display screen. If the rear surface 1265 of the mobile device 1250 has a display screen 1245, any positioning of the third camera 1230C and the fourth camera 1230D relative to the display screen 1245 may be used, as discussed with respect to the first camera 1230A and the second camera 1230B at the front surface 1255 of the mobile device 1250. The third camera 1230C and the fourth camera 1230D may be both of one or more cameras 310. In some examples, the rear surface 1265 of the mobile device 1250 may have only a single camera. In some examples, in addition to the first camera 1230A, the second camera 1230B, the third camera 1230C, and the fourth camera 1230D, the mobile device 1250 may further include one or more additional cameras. In some examples, in addition to the first camera 1230A, the second camera 1230B, the third camera 1230C, and the fourth camera 1230D, the mobile device 1250 may further include one or more additional sensors.
[0144] Like the HMD 1110, the mobile device 1250 does not include wheels, propellers, or other conveyance means of its own. Instead, the mobile device 1250 relies on the movement of a user who holds or wears the mobile device 1250 to move the mobile device 1250 back and forth in the environment. In some cases, such as when the mobile device 1250 is used for AR, VR, MR, or XR, the environment may be fully or partially virtual. In some cases, the mobile device 1250 may be inserted into a head-mounted device (HMD) (e.g., inserted into a cradle of the HMD) such that the mobile device 1250 serves as a display of the HMD, where the display screen 1245 of the mobile device 1250 serves as a display of the HMD. If the environment is at least partially virtual, the movement through the virtual environment may also be virtual. For example, the movement through the virtual environment may be controlled by one or more joysticks, buttons, video game controllers, mice, keyboards, touchpads, and / or other input devices coupled to the mobile device 1250 in a wired or wireless manner.
[0145] Figure 13FIG. is an example diagram illustrating a system for implementing certain aspects of the present technology. Specifically, Figure 13 An example of a computing system 1300 is illustrated, which can be any computing device that constitutes an internal computing system, a remote computing system, a camera, or any component thereof, where components of the system communicate with each other using connection 1305. Connection 1305 can be a physical connection using a bus or a direct connection into a processor 1310, such as in a chipset architecture. Connection 1305 can also be a virtual connection, a networking connection, or a logical connection.
[0146] In some examples, computing system 1300 is a distributed system, where the functions described in this disclosure can be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some examples, one or more of the described system components represent many such components, each component performing some or all of the functions that the component is described for. In some cases, these components can be physical or virtual devices.
[0147] Example system 1300 includes at least one processing unit (CPU or processor) 1310 and connection 1305, which couples various system components including system memory 1315 (such as read-only memory (ROM) 1320 and random access memory (RAM) 1325) to processor 1310. Computing system 1300 can include a cache 1312 that is directly connected to, in close proximity to, or integrated as part of processor 1310 of high-speed memory.
[0148] Processor 1310 can include any general-purpose processor and hardware services or software services, such as services 1332, 1334, and 1336 stored in storage device 1330, which are configured to control processor 1310 and a dedicated processor in which software instructions are incorporated into the actual processor design. Processor 1310 can be a fully self-contained computing system that includes multiple cores or processors, buses, memory controllers, caches, etc. A multi-core processor can be symmetric or asymmetric.
[0149] To enable user interaction, computing system 1300 includes an input device 1345 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, a camera, an accelerometer, a gyroscope, etc. Computing system 1300 may also include an output device 1335 that can be one or more of a plurality of output mechanisms. In some cases, a multimodal system may enable a user to provide multiple types of input / output to communicate with computing system 1300. Computing system 1300 may include a communication interface 1340 that generally may govern and manage user input and system output. The communication interface may perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including using audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, port / plug, Ethernet port / plug, fiber optic port / plug, dedicated wired port / plug, wireless signaling, low energy (BLE) wireless signaling, wireless signaling, radio frequency identification (RFID) wireless signaling, near field communication (NFC) wireless signaling, dedicated short range communication (DSRC) wireless signaling, 802.11 Wi-Fi wireless signaling, wireless local area network (WLAN) signaling, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), infrared (IR) communication wireless signaling, public switched telephone network (PSTN) signaling, integrated services digital network (ISDN) signaling, 3G / 4G / 5G / LTE cellular data network wireless signaling, ad hoc network signaling, radio wave signaling, microwave signaling, infrared signaling, visible light signaling, ultraviolet light signaling, wireless signaling along the electromagnetic spectrum, or some combination thereof. Communication interface 1340 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of computing system 1300 based on receiving one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0150] The storage device 1330 can be a non-volatile and / or non-transitory and / or computer-readable memory device, and can be a hard disk or other types of computer-readable media that can store data accessible by a computer, such as a tape cassette, a flash memory card, a solid-state memory device, a digital versatile disc, a cassette tape, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage media, flash memory, memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, a digital video disc (DVD) optical disc, a Blu-ray disc (BDD) optical disc, a holographic optical disc, another optical media, a secure digital (SD) card, a micro secure digital (microSD) card, a card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or cartridge and / or a combination thereof.
[0151] The storage device 1330 can include software services, servers, services, etc., and when the code defining such software is executed by the processor 1310, the code causes the system to perform functions. In some examples, the hardware services that perform specific functions can include software components stored in a computer-readable medium connected to the necessary hardware components (such as the processor 1310, the connection 1305, the output device 1335, etc.) to perform functions.
[0152] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium can include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media can include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. The computer-readable medium can have code and / or machine-executable instructions stored thereon, which can represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. The information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.
[0153] In some examples, computer-readable storage devices, media, and memories can include cables or wireless signals that contain bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0154] Specific details are provided in the above description to provide a thorough understanding of the examples provided herein. However, those of ordinary skill in the art will understand that the examples can be implemented without these specific details. For clarity, in some instances, the present technology may be presented as including separate functional blocks, including functional blocks that contain devices, device components, steps or routines in a method embodied in software or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein can be used. For example, circuits, systems, networks, processes, and other components can be shown as components in block diagram form so as not to obscure the examples with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details so as to avoid obscuring the examples.
[0155] The above may describe various examples as processes or methods, which are depicted as flowcharts, process flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many of the operations in the operations can be performed in parallel or concurrently. In addition, the order of the operations can be rearranged. When the operations of a process are completed, the process is terminated, but the process may have additional steps not included in the figures. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.
[0156] The processes and methods according to the above examples can be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. Such instructions can include, for example, instructions and data that cause a general-purpose computer, special-purpose computer, or processing device to configure a general-purpose computer, special-purpose computer, or processing device to perform a certain function or group of functions in other ways. Parts of the computer resources used can be accessed through a network. The computer-executable instructions can be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic disks or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0157] Devices implementing the processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and can take any form factor among a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program products) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. The processor can execute the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein can also be embodied in peripheral devices or plug-in cards. By additional examples, such functionality can also be implemented on a circuit board among different chips or different processes executed on a single device.
[0158] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functions described in this disclosure.
[0159] In the foregoing description, aspects of the present application have been described with reference to specific examples of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while illustrative examples of the present application have been described in detail herein, it should be understood that the inventive concept may be implemented and adopted in various other ways, and the appended claims are intended to be construed to include such variations, unless limited by the prior art. The various features and aspects of the foregoing application may be used singly or in combination. Further, the examples may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be understood that in alternative examples, the methods may be performed in an order different from that described.
[0160] One of ordinary skill in the art will understand that, without departing from the scope of the specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced, respectively, with the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols.
[0161] Where a component is described as “configured to” perform certain operations, such a configuration may be implemented, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0162] The phrase “coupled to” means that any component is directly or indirectly physically connected to another component, and / or any component directly or indirectly communicates with another component (e.g., is connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0163] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0164] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0165] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in a wireless communication device handset and other devices. Any feature described as a module or component may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially implemented by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0166] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor; however, in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" can refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within a dedicated software module or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0167] Exemplary aspects of the present disclosure include:
[0168] Aspect 1. An apparatus for image generation, the apparatus comprising: a memory including instructions; and a processor coupled to the memory and configured to: render an image of a scene, the scene including a first virtual object associated with a first physical object that moves independently of the scene; identify a first set of pixels associated with the first virtual object in the image of the scene; generate a mask layer based on the first set of pixels, the mask layer indicating the positions of the first set of pixels in the image; and send the mask layer and the image to a display.
[0169] Aspect 2. The apparatus according to aspect 1, wherein the display is coupled to the apparatus but separate from the apparatus.
[0170] Aspect 3. The apparatus according to any one of aspects 1 to 2, wherein the mask layer includes object identifiers corresponding to the pixels of the image.
[0171] Aspect 4. The apparatus according to aspect 3, wherein, in order to generate the mask layer, the processor is configured to set the positions in the mask layer corresponding to the identified first set of pixels by means of the object identifiers associated with the first virtual object.
[0172] Aspect 5. The apparatus according to any one of aspects 1 to 4, wherein the processor is configured to: generate depth information of the first virtual object; and send the depth information, together with the mask layer and the image, to the display.
[0173] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein the processor is configured to determine that the first physical object is movable independently of the scene.
[0174] Aspect 7. The apparatus according to Aspect 6, wherein, in order to determine that the first physical object is movable independently of the scene, the processor is configured to: receive an indication that the display has stored pose information associated with the first physical object; and determine that the first physical object is movable independently of the scene based on the received indication.
[0175] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein the scene includes a second virtual object associated with a second physical object that is movable independently of the scene, and wherein the processor is configured to: render the image of the scene including the second virtual object; identify a second set of pixels associated with the second virtual object in the image of the scene; and generate the mask layer based on the first set of pixels and the second set of pixels.
[0176] Aspect 9. The apparatus according to any one of Aspects 1 to 8, wherein the apparatus includes a mobile phone.
[0177] Aspect 10. An apparatus for image generation, the apparatus including: a memory that includes instructions; and a processor coupled to the memory and configured to: obtain a mask layer and an image associated with the mask layer, wherein the mask layer is associated with a first virtual object; generate updated pose information of the apparatus, wherein the updated pose information indicates a movement between a time when the image was generated and the current time; obtain updated pose information of a first physical object, wherein the first physical object moves independently of the apparatus, and wherein the first physical object is associated with the first virtual object; determine a first amount for distorting the image based on the updated pose information of the apparatus; determine a second amount for distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; distort the image based on the first amount for distorting the image based on the updated pose information, and distort the first portion of the image based on the second amount for distorting the first virtual object by the first physical object; and display a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0178] Aspect 11. The apparatus according to Aspect 10, wherein the mask layer includes object identifiers corresponding to pixels of the image.
[0179] Aspect 12. The apparatus according to any one of aspects 10 to 11, wherein the processor is configured to: receive depth information associated with the first virtual object together with the mask layer and the image; and further determine the second amount for distorting the first portion of the image based on the depth information.
[0180] Aspect 13. The apparatus according to any one of aspects 10 to 12, wherein the processor is configured to: obtain first pose information of the first physical object; generate second pose information of the apparatus; and send the first pose information and the second pose information to a companion device for rendering an image for the apparatus.
[0181] Aspect 14. The apparatus according to any one of aspects 10 to 13, wherein the mask layer is associated with a second virtual object, and wherein the processor is configured to: obtain updated pose information of the second physical object; determine a third amount for distorting a second portion of the image based on the mask layer and the updated pose information of the second physical object; and distort the second portion of the image based on the third amount.
[0182] Aspect 15. A method for image generation, the method comprising: rendering an image of a scene, the scene including a first virtual object associated with a first physical object that moves independently of the scene; identifying a first set of pixels associated with the first virtual object in the image of the scene; generating a mask layer based on the first set of pixels, the mask layer indicating the positions of the first set of pixels in the image; and sending the mask layer and the image to a display.
[0183] Aspect 16. The method according to aspect 15, wherein the display is coupled to a companion device but separate from the companion device.
[0184] Aspect 17. The method according to any one of aspects 15 to 16, wherein the mask layer includes object identifiers corresponding to pixels of the image.
[0185] Aspect 18. The method according to aspect 17, wherein generating the mask layer includes setting positions in the mask layer corresponding to the identified first set of pixels with the object identifiers associated with the first virtual object.
[0186] Aspect 19. The method according to any one of aspects 15 to 18, the method further comprising: generating depth information of the first virtual object; and sending the depth information together with the mask layer and the image to the display.
[0187] Aspect 20. The method according to any one of aspects 15 to 19, the method further comprising determining that the first physical object is movable independently of the scene.
[0188] Aspect 21. The method according to aspect 20, wherein determining that the first physical object is movable independently of the scene comprises: receiving an indication that the display has stored pose information associated with the first physical object; and determining, based on the received indication, that the first physical object is movable independently of the scene.
[0189] Aspect 22. The method according to any one of aspects 15 to 21, wherein the scene includes a second virtual object associated with a second physical object that is movable independently of the scene, and the method further comprises: rendering the image of the scene including the second virtual object; identifying a second set of pixels associated with the second virtual object in the image of the scene; and generating the mask layer based on the first set of pixels and the second set of pixels.
[0190] Aspect 23. A method for image generation, the method comprising: obtaining a mask layer and an image associated with the mask layer, wherein the mask layer is associated with a first virtual object; generating updated pose information of a device, wherein the updated pose information indicates a movement between the time when the image is generated and the current time; obtaining updated pose information of a first physical object, wherein the first physical object moves independently of the device, and wherein the first physical object is associated with the first virtual object; determining a first amount of distorting the image based on the updated pose information of the device; determining a second amount of distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; distorting the image based on the first amount of distorting the image based on the updated pose information, and distorting the first portion of the image based on the second amount of distorting the first virtual object by the first physical object; and displaying a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0191] Aspect 24. The method according to aspect 23, wherein the mask layer includes object identifiers corresponding to pixels of the image.
[0192] Aspect 25. The method according to any one of aspects 23 to 24, the method further comprising: receiving depth information associated with the first virtual object together with the mask layer and the image; and further determining the second amount of distorting the first portion of the image based on the depth information.
[0193] Aspect 26. The method according to any one of aspects 23 to 25, the method further comprising: obtaining first pose information of the first physical object; generating second pose information of the device; and sending the first pose information and the second pose information to a companion device for rendering an image.
[0194] Aspect 27. The method according to any one of aspects 23 to 26, wherein the mask layer is associated with a second virtual object, and the method further comprising: obtaining updated pose information of the second physical object; determining a third amount for distorting a second portion of the image based on the mask layer and the updated pose information of the second physical object; and distorting the second portion of the image based on the third amount.
[0195] Aspect 28. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: render an image of a scene, the scene including a first virtual object associated with a first physical object that moves independently of the scene; identify a first set of pixels associated with the first virtual object in the image of the scene; generate a mask layer based on the first set of pixels, the mask layer indicating the positions of the first set of pixels in the image; and send the mask layer and the image to a display.
[0196] Aspect 29. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: obtain a mask layer and an image associated with the mask layer, wherein the mask layer is associated with a first virtual object; generate updated pose information of a device, wherein the updated pose information indicates a movement between a time when the image was generated and a current time; obtain updated pose information of a first physical object, wherein the first physical object moves independently of the device and wherein the first physical object is associated with the first virtual object; determine a first amount for distorting the image based on the updated pose information of the device; determine a second amount for distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; distort the image based on the first amount that distorts the image based on the updated pose information and distort the first portion of the image based on the second amount that distorts the first virtual object by the first physical object; and display a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
[0197] Aspect 32: The device according to any one of aspects 1 to 15, wherein the device is a mobile device.
[0198] Aspect 33. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to perform one or more of the operations according to any one of aspects 16 to 30.
[0199] Aspect 34: An apparatus for image generation, the apparatus including components for performing one or more of the operations according to any one of aspects 16 to 30.
Claims
1. An apparatus for image generation, the apparatus comprising: a memory including instructions; and a processor coupled to the memory and configured to: render an image of a scene, the scene including a first virtual object associated with a first physical object that moves independently of the scene; identify a first set of pixels in the image of the scene that are associated with the first virtual object; generate a mask layer based on the first set of pixels, the mask layer indicating the positions of the first set of pixels in the image; and send the mask layer and the image to a display.
2. The apparatus according to claim 1, wherein the display is coupled to the apparatus but separate from the apparatus.
3. The apparatus according to claim 1, wherein the mask layer includes object identifiers corresponding to the pixels of the image.
4. The apparatus according to claim 3, wherein, to generate the mask layer, the processor is configured to set, via the object identifier associated with the first virtual object, the positions in the mask layer corresponding to the identified first set of pixels.
5. The apparatus according to claim 1, wherein the processor is configured to: generate depth information of the first virtual object; and send the depth information, together with the mask layer and the image, to the display.
6. The apparatus according to claim 1, wherein the processor is configured to determine that the first physical object can move independently of the scene.
7. The apparatus according to claim 6, wherein, to determine that the first physical object can move independently of the scene, the processor is configured to: receive an indication that the display has stored pose information associated with the first physical object; and determine, based on the received indication, that the first physical object can move independently of the scene.
8. The apparatus according to claim 1, wherein the scene includes a second virtual object associated with a second physical object that moves independently of the scene, and wherein the processor is configured to: render the image of the scene including the second virtual object; identify a second set of pixels in the image of the scene that are associated with the second virtual object; and generate the mask layer based on the first set of pixels and the second set of pixels.
9. The apparatus according to claim 1, wherein the apparatus includes a mobile phone.
10. An apparatus for image generation, the apparatus comprising: a memory including instructions; and a processor coupled to the memory and configured to: obtain a mask layer and an image associated with the mask layer, wherein the mask layer is associated with a first virtual object; generate updated pose information of the apparatus, wherein the updated pose information indicates the movement between the time when the image was generated and the current time; obtain updated pose information of a first physical object, wherein the first physical object moves independently of the apparatus and is associated with the first virtual object; Determine a first amount for distorting the image based on the updated pose information of the device; Determine a second amount for distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; Distort the image based on the first amount for distorting the image, and distort the first portion of the image based on the second amount for distorting a first virtual object by the first physical object; And Display a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
11. The apparatus according to claim 10, wherein the mask layer includes object identifiers corresponding to pixels of the image.
12. The apparatus according to claim 10, wherein the processor is configured to: Receive depth information associated with the first virtual object together with the mask layer and the image; and Further determine the second amount for distorting the first portion of the image based on the depth information.
13. The apparatus according to claim 10, wherein the processor is configured to: Obtain first pose information of the first physical object; Generate second pose information of the apparatus; and Send the first pose information and the second pose information to a companion device for rendering an image for the apparatus.
14. The apparatus according to claim 10, wherein the mask layer is associated with a second virtual object associated with a second physical object, and wherein the processor is configured to: Obtain updated pose information of the second physical object; Determine a third amount for distorting a second portion of the image based on the mask layer and the updated pose information of the second physical object; and Distort the second portion of the image based on the third amount.
15. A method for image generation, the method comprising: Render an image of a scene, the scene including a first virtual object associated with a first physical object that moves independently of the scene; Identify a first set of pixels associated with the first virtual object in the image of the scene; Generate a mask layer based on the first set of pixels, the mask layer indicating the positions of the first set of pixels in the image; And Send the mask layer and the image to a display.
16. The method according to claim 15, wherein the display is coupled to a companion device but separate from the companion device.
17. The method according to claim 15, wherein the mask layer includes object identifiers corresponding to pixels of the image.
18. The method according to claim 17, wherein generating the mask layer includes setting positions in the mask layer corresponding to the identified first set of pixels by the object identifiers associated with the first virtual object.
19. The method according to claim 15, the method further comprising: Generate depth information of the first virtual object; And Send the depth information together with the mask layer and the image to the display.
20. The method according to claim 15, wherein the method further comprises determining that the first physical object is movable independently of the scene.
21. The method according to claim 20, wherein determining that the first physical object is movable independently of the scene comprises: receiving an indication that the display has stored pose information associated with the first physical object; and determining, based on the received indication, that the first physical object is movable independently of the scene.
22. The method according to claim 15, wherein the scene includes a second virtual object associated with a second physical object movable independently of the scene, and the method further comprises: rendering the image of the scene including the second virtual object; identifying a second set of pixels associated with the second virtual object in the image of the scene; and generating the mask layer based on the first set of pixels and the second set of pixels.
23. A method for image generation, the method comprising: obtaining a mask layer and an image associated with the mask layer, wherein the mask layer is associated with a first virtual object; generating updated pose information of the device, wherein the updated pose information indicates a movement between the time when the image is generated and the current time; obtaining updated pose information of a first physical object, wherein the first physical object moves independently of the device, and wherein the first physical object is associated with the first virtual object; determining a first amount for distorting the image based on the updated pose information of the device; determining a second amount for distorting a first portion of the image based on the mask layer and the updated pose information of the first physical object; distorting the image based on the first amount for distorting the image, and distorting the first portion of the image based on the second amount for distorting the first virtual object by the first physical object; and displaying a final distorted image based on the distortion of the image and the distortion of the first portion of the image.
24. The method according to claim 23, wherein the mask layer includes object identifiers corresponding to the pixels of the image.
25. The method according to claim 23, wherein the method further comprises: receiving depth information associated with the first virtual object together with the mask layer and the image; and further determining the second amount for distorting the first portion of the image based on the depth information.
26. The method according to claim 23, wherein the method further comprises: obtaining first pose information of the first physical object; generating second pose information of the device; and sending the first pose information and the second pose information to a companion device for rendering an image.
27. The method according to claim 23, wherein the mask layer is associated with a second virtual object, and the method further comprises: obtaining updated pose information of the second physical object; determining a third amount for distorting a second portion of the image based on the mask layer and the updated pose information of the second physical object; and Distort the second portion of the image based on the third quantity.