Augmented reality enhanced media
By generating media content elements associated with the application engine in the XR system and displaying corresponding media content on multiple displays, the problems of short battery life and poor portability of the existing XR system are solved, and more efficient battery usage and portability are achieved.
Patent Information
- Application Number
- CN202380076879.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-02
- Filing Date
- 2023-11-03
- Publication Date
- 2025-06-13
AI Technical Summary
Existing Extended Reality (XR) systems require a powerful processor to perform complex functions, resulting in short battery life and inconvenient device portability.
By generating media content elements associated with the application engine and displaying on a first display of the first device, the media content elements displayed on the second display of the second device relative to the first display posture of the first device are output.
It achieves improving the battery life and portability of the XR system without increasing the weight and size of the device, while enhancing the display effect of media content.
Smart Images

Figure CN120153339A_ABST
Abstract
Description
[0001] Field
[0002] This application relates to augmented media utilizing augmented reality. For example, aspects of the application relate to systems and techniques for augmenting media content displayed on a first display of a first device with augmented reality media content displayed on a second display of a second device.
[0003] Background
[0004] Degree of freedom (DoF) refers to the number of fundamental ways in which a rigid body can move in three-dimensional (3D) space. In some examples, six different DoFs can be tracked. The six DoFs include three translational DoFs corresponding to translational movement along three perpendicular axes (which may be referred to as the x, y, and z axes). The six DoFs include three rotational DoFs corresponding to rotational movement about these three axes, which may be referred to as pitch, yaw, and roll. Some extended reality (XR) devices, such as virtual reality (VR) or augmented reality (AR) head-mounted devices, can track some or all of these degrees of freedom. For example, 3DoF XR head-mounted devices typically track three rotational DoFs and can thus track whether a user turns and / or tilts their head. 6DoF XR head-mounted devices can track all six degrees of freedom and can thus also track a user's translational movement.
[0005] XR systems typically use powerful processors to perform feature analysis (e.g., extraction, tracking, etc.) and other complex functions quickly enough to display an output to their users based on these functions. Powerful processors typically draw power at a high rate. Similarly, sending a large amount of data to a powerful processor typically draws power at a high rate. Head-mounted devices and other portable devices typically have small batteries so as not to be uncomfortably heavy for the user. Thus, some XR systems must be plugged into an external power source and are thus not portable. Portable XR systems typically have a short battery life and / or are uncomfortably heavy due to containing large batteries.
[0006] Overview
[0007] Systems and techniques for displaying augmented reality augmented media content are described herein. According to at least one example, a method for displaying media content on one or more displays is provided. The method includes: generating a first media content element for an application engine associated with an application state of the application engine; generating a second media content element for the application engine associated with the application state of the application engine; displaying the first media content element on a first display of a first device; and outputting the second media content element for display on a second display of a second device relative to a pose of the first display of the first device.
[0008] In another example, a device for displaying media content on one or more displays is provided. The device includes at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: generate a first media content element for an application engine associated with an application state of the application engine; generate a second media content element for the application engine associated with the application state of the application engine; display the first media content element on a first display of a first device; and output the second media content element for display on a second display of a second device relative to a pose of the first display of the first device.
[0009] In another example, a non-transitory computer-readable medium storing instructions is provided. When executed by one or more processors, the instructions cause the one or more processors to: generate a first media content element for an application engine associated with an application state of the application engine; generate a second media content element for the application engine associated with the application state of the application engine; display the first media content element on a first display of a first device; and output the second media content element for display on a second display of a second device relative to a pose of the first display of the first device.
[0010] In another example, a device for displaying media content on one or more displays is provided. The device includes: means for generating a first media content element for an application engine associated with an application state of the application engine; means for generating a second media content element for the application engine associated with the application state of the application engine; means for displaying the first media content element on a first display of a first device; and means for outputting the second media content element for display on a second display of a second device relative to a pose of the first display of the first device.
[0011] According to at least one example, a method for displaying media content on one or more displays is provided. The method includes: obtaining an application state of an application engine from a first device including a first display; generating a media content element associated with the application state of the application engine for a second device including a second display; and displaying the media content element on the second display of the second device relative to a pose of the first display of the first device.
[0012] In another example, a device for displaying media content on one or more displays is provided. The device includes at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: obtain an application state of an application engine from a first device including a first display; generate media content elements associated with the application state of the application engine for a second device including a second display; and display the media content elements on the second display of the second device relative to a pose of the first display of the first device.
[0013] In another example, a non-transitory computer-readable medium storing instructions is provided. When executed by one or more processors, the instructions cause the one or more processors to: obtain an application state of an application engine from a first device including a first display; generate media content elements associated with the application state of the application engine for a second device including a second display; and display the media content elements on the second display of the second device relative to a pose of the first display of the first device.
[0014] In another example, a device for displaying media content on one or more displays is provided. The device includes: means for obtaining an application state of an application engine from a first device including a first display; means for generating media content elements associated with the application state of the application engine for a second device including a second display; and means for displaying the media content elements on the second display of the second device relative to a pose of the first display of the first device.
[0015] In some aspects, the device includes: a camera, a mobile device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wireless communication device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other devices. In some aspects, one or more processors include an image signal processor (ISP). In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device includes an image sensor for capturing image data. In some aspects, the device further includes a display for displaying images, one or more notifications (e.g., associated with the processing of the images), and / or other displayable data.
[0016] This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0017] The foregoing, as well as other features and embodiments, will become more apparent when referring to the following specification, claims, and accompanying drawings. Brief Description of the Drawings
[0019] Exemplary embodiments of the present application are described in detail below with reference to the following drawings:
[0020] Figure 1 is a block diagram illustrating the architecture of an image capture and processing device according to some examples of the present disclosure;
[0021] Figure 2 is a block diagram illustrating the architecture of an exemplary extended reality (XR) system according to some examples of the present disclosure;
[0022] Figure 3 is a block diagram illustrating the architecture of a simultaneous localization and mapping (SLAM) device according to some examples of the present disclosure;
[0023] Figure 4 is a block diagram illustrating the architecture of an extended reality enhanced application engine according to some examples of the present disclosure;
[0024] Figure 5 is a block diagram illustrating the architecture of an application engine and an augmented reality companion engine according to some examples of the present disclosure;
[0025] Figure 6 is a view illustrating an example of a video game with augmented reality enhancement according to some examples of the present disclosure;
[0026] Figure 7 is a flowchart illustrating an example of a technique for displaying media content according to some examples;
[0027] Figure 8 is a flowchart illustrating an example of a technique for displaying media content according to some examples;
[0028] Figure 9A is a perspective view of a head-mounted display (HMD) performing feature tracking and / or visual simultaneous localization and mapping (VSLAM) according to some examples;
[0029] Figure 9B is an illustration according to some examples of Figure 9A a perspective view of a head-mounted display (HMD) worn by a user;
[0030] Figure 10A is a perspective view of the front surface of a mobile handheld device according to some examples;
[0031] Figure 10B is a perspective view of the back surface of a mobile handheld device according to some examples;
[0032] Figure 11 is a diagram illustrating an example of a system for implementing certain aspects of the techniques herein.
[0033] Detailed Description
[0034] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments may be applied independently and some of them may be applied in combination, which will be obvious to those skilled in the art. In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it is obvious that the embodiments may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0035] The following description only provides exemplary embodiments and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application set forth in the appended claims.
[0036] An image capture device (e.g., a camera) is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). The terms “image,” “image frame,” “video frame,” and “frame” may be used interchangeably herein. An image capture device typically includes at least one lens that receives light from a scene and bends the light toward the image sensor of the image capture device. The light received by the lens passes through an aperture controlled by one or more control mechanisms and is received by the image sensor. The one or more control mechanisms may control exposure, focus, and / or zoom based on information from the image sensor and / or based on information from an image processor (e.g., a host or application processing and / or image signal processor). In some examples, the one or more control mechanisms include a motor or other control mechanism that moves the lens of the image capture device to a target lens position.
[0037] Degree of freedom (DoF) refers to the number of basic ways in which a rigid body can move in three-dimensional (3D) space. In some cases, six different DoFs can be tracked. The six degrees of freedom include three translational degrees of freedom corresponding to translational movement along three perpendicular axes. The three axes may be referred to as the x-axis, y-axis, and z-axis. The six degrees of freedom include three rotational degrees of freedom corresponding to rotational movement about these three axes, which may be referred to as pitch, yaw, and roll.
[0038] An extended reality (XR) system or device can provide virtual content to a user and / or can combine the real world or physical environment with a virtual environment (including virtual content) to provide an XR experience to the user. The real world environment can include real world objects (also referred to as physical objects), such as people, vehicles, buildings, tables, chairs, and / or other real world or physical objects. The XR system or device can facilitate interaction with different types of XR environments (e.g., a user can use the XR system or device to interact with an XR environment). The XR system can include a virtual reality (VR) system that facilitates interaction with a VR environment, an augmented reality (AR) system that facilitates interaction with an AR environment, a mixed reality (MR) system that facilitates interaction with an MR environment, and / or other XR systems. As used herein, the terms XR system and XR device can be used interchangeably. Examples of XR systems or devices include head-mounted displays (HMDs), smart glasses, etc. In some cases, the XR system can track parts of the user (e.g., the user's hands and / or fingertips) to allow the user to interact with virtual content items.
[0039] AR is a technology that provides virtual or computer-generated content (referred to as AR content) on a user's view of a physical, real world scene or environment. The AR content can include virtual content such as video, images, graphic content, location data (e.g., global positioning system (GPS) data or other location data), sound, any combination thereof, and / or other augmented content. AR systems or devices are designed to enhance (or augment) rather than replace a person's current perception of reality. For example, a user can see a stationary or moving real physical object through an AR device display, but the user's visual perception of the physical object can be augmented or enhanced by: a virtual image of the object (e.g., a real world car replaced by a virtual image of a DeLorean), AR content added to the physical object (e.g., virtual wings added to a live animal), AR content displayed relative to the physical object (e.g., information virtual content displayed near a sign on a building, a virtual coffee cup virtually anchored to (e.g., placed on) a real world table in one or more images, etc.), and / or by displaying other types of AR content. Various types of AR systems can be used for gaming, entertainment, and / or other applications.
[0040] Visual Simultaneous Localization and Mapping (VSLAM) is a computational geometry technique used in devices with cameras, such as robots, head-mounted displays (HMDs), mobile handheld devices, and autonomous vehicles. In VSLAM, the device can build and update a map of an unknown environment based on images captured by the device's camera. As the device updates the map, the device can track its pose (e.g., position and / or orientation) within the environment. For example, the device can be activated in a specific room of a building and can move throughout the building, capturing images. The device can map the environment and track its position within the environment based on tracking the positions where different objects appear in different images.
[0041] In the context of systems that track movement in an environment, such as XR systems and / or VSLAM systems, degrees of freedom can refer to which of the six degrees of freedom the system is capable of tracking. 3DoF systems typically track three rotational DoFs - pitch, yaw, and roll. For example, a 3DoF head-mounted device can track a user of the head-mounted device turning their head left or right, tilting their head up or down, and / or tilting their head left or right. 6DoF systems can track three translational DoFs as well as three rotational DoFs. Thus, for example, a 6DoF head-mounted device can track a user moving forward, backward, laterally, and / or vertically in addition to tracking the three rotational DoFs.
[0042] Systems that track movement in an environment, such as XR systems and / or VSLAM systems, typically include powerful processors. These powerful processors can be used to perform complex operations quickly enough to display up-to-date outputs based on these operations to users of these systems. Such complex operations can involve feature tracking, 6DoF tracking, VSLAM, rendering virtual objects to overlay a user's environment in XR, animating virtual objects, and / or other operations discussed herein. Powerful processors typically draw power at a high rate. Sending large amounts of data to powerful processors typically draws power at a high rate, and such systems tend to capture large amounts of sensor data (e.g., images, position data, and / or other sensor data) per second. Head-mounted devices and other portable devices typically have small batteries so as not to be uncomfortably heavy for the user. Thus, a typical XR head-mounted device must be plugged into an external power source, be uncomfortably heavy due to including a large battery, or have a very short battery life.
[0043] Currently, the visual elements of games on mobile devices are restricted to what can be displayed on the physical display of the device itself. In most cases, the mobile device display is relatively small and only occupies a small portion of the player's field of view. Modern games need to present both the game user interface (UI) and the visual representation of the game itself on the same display surface, which often results in a cluttered, cramped, and difficult-to-interact-with user interface, obscuring the game and frustrating the player. Additionally, the small display of mobile devices limits the player's immersion in the game experience. Since there is too little of the user's field of view related to the game being played, it is easier for game players to disengage from the game experience and become distracted or bored.
[0044] As described in more detail herein, systems, devices, methods (also referred to as processes and computer-readable media, collectively referred to herein as "systems and technologies") for enhancing the display of media content using augmented reality media content are described.
[0045] A user of an XR system (e.g., an AR device) may sometimes wear the XR system while the user is interacting with another electronic device (e.g., a mobile device, smartphone, tablet device, etc.). In some cases, the XR system and the electronic device may be configured to display media content (e.g., display still images, video frames, etc.) simultaneously on both the XR system and the electronic device. In an illustrative example, a user may wear an XR system while playing a video game on a mobile device (e.g., a smartphone). In some cases, in addition to the game content displayed on the display of the smartphone, the XR system may also display game content. In some cases, the XR system may display the game content oriented relative to the display of the smartphone to create the appearance that the display of the smartphone extends beyond the physical display area. In some cases, the XR system may display effects (e.g., fog, explosions, particles, etc.). For example, an explosion displayed by the XR system may have an appearance that originates from the display of the mobile device and extends beyond the display area of the mobile device to create a three-dimensional effect.
[0046] In some examples, the XR system may be coupled to the mobile device (e.g., via wired, wireless, or any combination thereof). In some implementations, the media content (e.g., rendered frames) of the game on both the mobile device and the XR system may be locally generated by a game engine. For example, the game engine may include parallel pipelines for rendering elements of the game simulation for each of the XR system and the mobile device.
[0047] In some implementations, the gaming media content of the mobile device and the XR system can be generated by rendering different applications. For example, a gaming application can generate media content for the mobile device (e.g., via a game engine renderer). In parallel, an accompanying application can generate media content for the XR system (e.g., via XR rendering). In some cases, the gaming application and the accompanying application can include communication modules for communication between these applications. For example, game data and / or game assets can be transferred from the gaming application to the accompanying application. In some cases, input from the XR system can be transferred from the accompanying application to the gaming application.
[0048] Aspects of the present application will be described with reference to the accompanying drawings. Figure 1 FIG. 5 is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing an image of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture a stand-alone image (or photo) and / or can capture a video including a plurality of images (or video frames) in a particular sequence. In some cases, the lens 115 and the image sensor 130 can be associated with an optical axis. In an illustrative example, both the photosensitive area of the image sensor 130 (e.g., a photodiode) and the lens 115 can be centered on the optical axis. The lens 115 of the image capture and processing system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the incident light from the scene towards the image sensor 130. The light received by the lens 115 passes through an aperture. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 120 and is received by the image sensor 130. In some cases, the aperture can have a fixed size.
[0049] One or more control mechanisms 120 can control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. One or more control mechanisms 120 can include multiple mechanisms and components; for example, the control mechanism 120 can include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 can also include additional control mechanisms other than the illustrated control mechanisms, such as control mechanisms for controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.
[0050] The focus control mechanism 125B of the control mechanism 120 can obtain a focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or a servo system (or other lens mechanism) to adjust the focus. In some cases, additional lenses may be included in the image capture and processing system 100, such as one or more microlenses above each photodiode of the image sensor 130, each microlens individually bending the light received from the lens 115 towards the corresponding photodiode before the light reaches the photodiode. The focus setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The control mechanism 120, the image sensor 130, and / or the image processor 150 can be used to determine the focus setting. The focus setting can be referred to as an image capture setting and / or an image processing setting. In some cases, the lens 115 can be fixed relative to the image sensor and the focus control mechanism 125B can be omitted without departing from the scope of the present disclosure.
[0051] The exposure control mechanism 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control mechanism 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 125A can control the size of the aperture (e.g., aperture size or f / stop), the time duration for which the aperture is open (e.g., exposure time or shutter speed), the time duration for which the sensor collects light (e.g., exposure time or electronic shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting can be referred to as an image capture setting and / or an image processing setting.
[0052] The zoom control mechanism 125C of the control mechanism 120 can obtain a zoom setting. In some examples, the zoom control mechanism 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control mechanism 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a varifocal zoom lens. In some examples, the lens assembly can include a focusing lens (which can be the lens 115 in some cases), which first receives light from the scene 110, and then the light passes through an afocal zoom system between the focusing lens (e.g., lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system can include: two positive (e.g., converging, convex) lenses having equal or similar (e.g., within a threshold difference from each other) focal lengths, and a negative (e.g., diverging, concave) lens therebetween. In some cases, the zoom control mechanism 125C moves one or more lenses in the afocal zoom system, such as the negative lens, and one or both of the positive lenses. In some cases, the zoom control mechanism 125C can control the zoom by capturing an image of an image sensor (e.g., including the image sensor 130) with a zoom corresponding to the zoom setting. For example, the image processing system 100 can include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a greater zoom. In some cases, based on the selected zoom setting, the zoom control mechanism 125C can capture an image from the corresponding sensor.
[0053] The image sensor 130 includes one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a particular pixel in the image generated by the image sensor 130. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered by color filters and can thus measure light that matches the color of the filter covering the photodiode. A variety of color filter arrays can be used, including a Bayer color filter array, a quad color filter array (also known as a quad Bayer color filter array or QCFA), and / or any other color filter array. For example, a Bayer color filter includes a red filter, a blue filter, and a green filter, where each pixel of the image is generated based on red light data from at least one photodiode covered by the red filter, blue light data from at least one photodiode covered by the blue filter, and green light data from at least one photodiode covered by the green filter.
[0054] Return to Figure 1 , as an alternative or supplement to red, blue, and / or green filters, other types of color filters can use yellow, magenta, and / or cyan (also known as "emerald") color filters. In some cases, some photodiodes may be configured to measure infrared (IR) light. In some implementations, the photodiodes that measure IR light may not be covered by any filter, allowing the IR photodiodes to measure both visible light (e.g., color) and IR light. In some examples, the IR photodiodes may be covered by an IR filter, allowing IR light to pass through and blocking light from other parts of the spectrum (e.g., visible light, color). Some image sensors (e.g., image sensor 130) may lack filters altogether (e.g., color, IR, or any other part of the spectrum) and can instead use different photodiodes throughout the (in some cases vertically stacked) pixel array. Different photodiodes throughout the pixel array can have different spectral sensitivity curves and thus respond to light of different wavelengths. Monochrome image sensors may also lack filters and thus lack color depth.
[0055] In some cases, the image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching a particular photodiode or portions of a particular photodiode at a particular time and / or from a particular angle. In some cases, the opaque and / or reflective mask may be used for phase detection autofocus (PDAF). In some cases, the opaque and / or reflective mask may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cut-off filter, UV cut-off filter, band-pass filter, low-pass filter, high-pass filter, etc.). The image sensor 130 may also include: an analog gain amplifier for amplifying the analog signal output by the photodiode and / or an analog-to-digital converter (ADC) for converting the analog signal output by the photodiode (and / or amplified by the analog gain amplifier) into a digital signal. In some cases, specific components or functions discussed with respect to one or more of the control mechanisms 120 may alternatively or additionally be included in the image sensor 130. The image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), complementary metal-oxide semiconductor (CMOS), N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0056] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more processors 1110 of any other type discussed with respect to Figure 11 the computing system 1100. The host processor 152 may be a digital signal processor (DSP) and / or some other type of processor. In some implementations, the image processor 150 is a single integrated circuit or chip that includes the host processor 152 and the ISP 154 (e.g., referred to as a system-on-chip or SoC). In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth TM, such as the Global Positioning System (GPS), any combination thereof, and / or other components. The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an Inter-Integrated Circuit (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General-Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as an MIPI CSI-2 Physical (PHY) layer port or interface), an Advanced High-Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port, and the ISP 154 may communicate with the image sensor 130 using an MIPI port.
[0057] The image processor 150 may perform several tasks, such as demosaicking, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving inputs, managing outputs, managing memory, or some combination thereof. The image processor 150 may store the image frames and / or the processed images in a Random Access Memory (RAM) 140 / 1125, a Read-Only Memory (ROM) 145 / 1120, a cache, a memory cell, another storage device, or some combination thereof.
[0058] A variety of input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device 1135, any other input device 1145, or some combination thereof. In some cases, captions may be input into the image processing device 105B via the physical keyboard or keypad of the I / O device 160, or via the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that implement a wired connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or transmit data to one or more peripheral devices. The I / O 160 may include one or more wireless transceivers that implement a wireless connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or transmit data to one or more peripheral devices. The peripheral devices may include any type of I / O device 160 discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they themselves may be considered I / O devices 160.
[0059] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including: an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, the image capture device 105A and the image processing device 105B may be wirelessly coupled together, for example, via one or more wires, cables, or other electrical connectors and / or via one or more wireless transceivers. In some implementations, the image capture device 105A and the image processing device 105B may be disconnected from each other.
[0060] As Figure 1 shown, the vertical dashed line separates Figure 1The image capture and processing system 100 is divided into two parts representing an image capture device 105A and an image processing device 105B, respectively. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, certain components illustrated in the image capture device 105A, such as the ISP 154 and / or the host processor 152, may be included in the image capture device 105A.
[0061] The image capture and processing system 100 may include an electronic device, such as a mobile or stationary telephone handset (e.g., a smartphone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device and the image processing device 105B may include a computing device, such as a mobile handset, a desktop computer, or other computing device.
[0062] Although the image capture and processing system 100 is shown as including certain components, those of ordinary skill in the art will appreciate that the image capture and processing system 100 may include more components than Figure 1 those shown therein. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of the image capture and processing system 100 may include and / or may use electronic circuitry or other electronic hardware (which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuitry)) to implement, and / or may include and / or may use computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture and processing system 100.
[0063] In some examples, Figure 2The extended reality (XR) system 200 can include an image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination thereof. In some examples, Figure 3 The simultaneous localization and mapping (SLAM) system 300 can include an image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination thereof.
[0064] Figure 2 FIG. is a diagram illustrating the architecture of an example XR system 200 in accordance with some aspects of the present disclosure. The XR system 200 can run (or execute) an XR application and implement XR operations. In some examples, the XR system 200 can perform tracking and localization, mapping, and / or positioning and rendering of virtual content on a display 209 (e.g., a screen, a visible plane / region, and / or other display) of an environment in the physical world as part of an XR experience. For example, the XR system 200 can generate a map of the environment in the physical world (e.g., a three-dimensional (3D) map), track the pose (e.g., localization and position) of the XR system 200 relative to the environment (e.g., relative to the 3D environment map), position and / or anchor virtual content at (a) specific location(s) on the environment map, and render the virtual content on the display 209 such that the virtual content appears to be at a location in the environment corresponding to the specific location(s) on the scene map where the virtual content is positioned and / or anchored. The display 209 can include glass, a screen, lenses, a projector, and / or other display mechanisms that allow a user to see the real-world environment and also allow XR content to be superimposed, overlapped, blended, or otherwise displayed thereon.
[0065] In this illustrative example, the XR system 200 includes one or more image sensors 202, an accelerometer 204, a gyroscope 206, a storage 207, a computing component 210, an XR engine 220, an image processing engine 224, a rendering engine 226, and a communication engine 228. It should be noted that Figure 2 the components 202-228 shown in are provided as non-limiting examples for illustrative and explanatory purposes, and other examples can include more, fewer, or different components than Figure 2 the components shown in. For example, in some cases, the XR system 200 can include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 2One or more other software and / or hardware components not shown. Although various components of XR system 200, such as image sensor 202, may be referred to herein in the singular, it should be understood that XR system 200 may include multiple components of any of the components discussed herein (e.g., multiple image sensors 202).
[0066] System 200 includes input device 208 or is in (wired or wireless) communication with input device 208. Input device 208 may include any suitable input device, such as a touch screen, a pen or other pointer device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, a video game controller, a steering wheel, a joystick, a set of buttons, a trackball, a remote control, any other input device 1145 discussed herein, or any combination thereof. In some cases, image sensor 202 may capture images that can be processed to interpret pose commands.
[0067] XR system 200 may also communicate (wired or wirelessly) with one or more other electronic devices. For example, communication engine 228 may be configured to manage connections and communicate with one or more electronic devices. In some cases, communication engine 228 may correspond to Figure 11 communication interface 1140.
[0068] In some implementations, one or more of image sensors 202, accelerometer 204, gyroscope 206, storage 207, computing components 210, XR engine 220, image processing engine 224, and rendering engine 226 may be part of the same computing device. For example, in some cases, one or more of image sensors 202, accelerometer 204, gyroscope 206, storage 207, computing components 210, XR engine 220, image processing engine 224, and rendering engine 226 may be integrated into a head-mounted display (HMD), extended reality glasses, a smartphone, a laptop device, a tablet device, a gaming system, and / or any other computing device. However, in some implementations, one or more of image sensors 202, accelerometer 204, gyroscope 206, storage 207, computing components 210, XR engine 220, image processing engine 224, and rendering engine 226 may be part of two or more separate computing devices. For example, in some cases, some of components 202 - 226 may be part of or implemented by one computing device, and the remaining components may be part of or implemented by one or more other computing devices.
[0069] The storage 207 can be any (one or more) storage devices for storing data. Additionally, the storage 207 can store data from any component of the XR system 200. For example, the storage 207 can store data from the image sensor 202 (e.g., image or video data), data from the accelerometer 204 (e.g., measurements), data from the gyroscope 206 (e.g., measurements), data from the computing component 210 (e.g., processing parameters, preferences, virtual content, rendered content, scene mapping, tracking and positioning data, object detection data, privacy data, XR application data, face recognition data, occlusion data, etc.), data from the XR engine 220, data from the image processing engine 224, and / or data from the rendering engine 226 (e.g., output frames). In some examples, the storage 207 can include a buffer for storing frames to be processed by the computing component 210.
[0070] One or more computing components 210 can include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, an image signal processor (ISP) 218, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). The computing component 210 can perform various operations such as image enhancement, computer vision, graphics rendering, extended reality operations (e.g., tracking, positioning, pose estimation, mapping, content anchoring, content rendering, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, the computing component 210 can implement (e.g., control, operate, etc.) the XR engine 220, the image processing engine 224, and the rendering engine 226. In other examples, the computing component 210 can also implement one or more other processing engines.
[0071] The image sensor 202 can include any image and / or video sensor or capture device. In some examples, the image sensor 202 can be part of a multi-camera assembly (such as a dual-camera assembly). The image sensor 202 can capture image and / or video content (e.g., raw image and / or video data), which can then be processed by the computing component 210, the XR engine 220, the image processing engine 224, and / or the rendering engine 226 as described herein. In some examples, the image sensor 202 can include an image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination thereof.
[0072] In some examples, the image sensor 202 can capture image data and can generate an image (also referred to as a frame) based on the image data and / or can provide the image data or the frame to the XR engine 220, the image processing engine 224, and / or the rendering engine 226 for processing. The image or frame can include a video frame having a video sequence, or a still image. The image or frame can include an array of pixels representing a scene. For example, the image can be a Red, Green, Blue (RGB) image having red, green, and blue components per pixel; a Luminance, Chroma Red, Chroma Blue (YCbCr) image having a luminance component and two chrominance (color) components (chroma red and chroma blue) per pixel; or any other suitable type of color or monochrome image.
[0073] In some cases, the image sensor 202 (and / or other cameras of the XR system 200) can also be configured to capture depth information. For example, in some implementations, the image sensor 202 (and / or other cameras) can include a Red Green Blue-Depth (RGB-D) camera. In some cases, the XR system 200 can include one or more depth sensors (not shown) that are separate from the image sensor 202 (and / or other cameras) and can capture depth information. For example, such depth sensors can obtain depth information independently of the image sensor 202. In some examples, the depth sensor can be physically mounted at the same general location as the image sensor 202, but can operate at a different frequency or frame rate than the image sensor 202. In some examples, the depth sensor can take the form of a light source that can project a structured light pattern or a textured light pattern (the pattern can include one or more narrow light bands) onto one or more objects in the scene. Subsequently, the depth information can be obtained by exploiting the geometric distortion of the projected pattern caused by the shape of the object surface. In one example, the depth information can be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).
[0074] The XR system 200 may also include other sensors in one or more of its sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to the computing component 210. For example, the accelerometer 204 may detect the acceleration of the XR system 200 and may generate an acceleration measurement based on the detected acceleration. In some cases, the accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, front / back) that can be used to determine the position or pose of the XR system 200. The gyroscope 206 may detect and measure the orientation and angular velocity of the XR system 200. For example, the gyroscope 206 may be used to measure the pitch, roll, and yaw of the XR system 200. In some cases, the gyroscope 206 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 202 and / or the XR engine 220 may use the measurements obtained by the accelerometer 204 (e.g., one or more translation vectors) and / or the gyroscope 206 (e.g., one or more rotation vectors) to calculate the pose of the XR system 200. As previously mentioned, in other examples, the XR system 200 may also include other sensors, such as an inertial measurement unit (IMU), a magnetometer, a gaze and / or eye tracking sensor, a machine vision sensor, a smart scene sensor, a speech recognition sensor, a collision sensor, a vibration sensor, a position sensor, a tilt sensor, etc.
[0075] As mentioned above, in some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that measures the specific force, angular rate, and / or orientation of the XR system 200 using a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers. In some examples, the one or more sensors may output measured information associated with the capture of an image captured by the image sensor 202 (and / or other cameras of the XR system 200) and / or depth information obtained using one or more depth sensors of the XR system 200.
[0076] The XR engine 220 can use the outputs of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) to determine the pose of the XR system 200 (also referred to as the head pose) and / or the pose of the image sensor 202 (or other cameras of the XR system 200). In some cases, the pose of the XR system 200 and the pose of the image sensor 202 (or other cameras) can be the same. The pose of the image sensor 202 refers to the position and orientation of the image sensor 202 relative to a reference frame (e.g., relative to the scene 110). In some implementations, the camera pose can be determined for six degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some implementations, the camera pose can be determined for three degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).
[0077] In some cases, a device tracker (not shown) can use measurements from one or more sensors and image data from the image sensor 202 to track the pose of the XR system 200 (e.g., 6DoF pose). For example, the device tracker can fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and movement of the XR system 200 relative to the physical world (e.g., the scene) and the mapping of the physical world. As described below, in some examples, when tracking the pose of the XR system 200, the device tracker can generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate an update to the 3D map of the scene. The 3D map update can include, for example but not limited to, new or updated features and / or features or landmark points associated with the scene and / or the 3D map of the scene, a localization update that identifies or updates the position of the XR system 200 within the scene and the 3D map of the scene, and so on. The 3D map can provide a digital representation of the scene in the real / physical world. In some examples, the 3D map can anchor location-based objects and / or content to real-world coordinates and / or objects. The XR system 200 can use the mapped scene (e.g., the scene represented and / or associated with the 3D map in the physical world) to merge the physical and virtual worlds and / or to merge virtual content or objects with the physical environment.
[0078] In some aspects, the pose of the image sensor 202 and / or the XR system 200 as a whole can be determined and / or tracked by the computing component 210 using a visual tracking solution based on images captured by the image sensor 202 (and / or other cameras of the XR system 200). For example, in some examples, the computing component 210 can perform tracking using computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques. For example, the computing component 210 can perform SLAM or can be in (wired or wireless) communication with a SLAM system (not shown) such as Figure 3 the SLAM system 300). SLAM refers to a class of techniques in which a camera (e.g., the image sensor 202) and / or the XR system 200 are simultaneously tracked relative to a map while creating a map of the environment (e.g., an environmental map modeled by the XR system 200). This map can be referred to as a SLAM map and can be three-dimensional (3D). SLAM techniques can be performed using color or grayscale image data captured by the image sensor 202 (and / or other cameras of the XR system 200) and can be used to generate an estimate of the 6DoF pose measurement of the image sensor 202 and / or the XR system 200. Such SLAM techniques configured to perform 6DoF tracking can be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and / or other sensors) can be used to estimate, correct, and / or otherwise adjust the estimated pose.
[0079] In some cases, 6DoF SLAM (e.g., 6DoF tracking) can associate features observed from a particular input image from the image sensor 202 (and / or other cameras) with the SLAM map. For example, 6DoF SLAM can use feature point associations from the input image to determine the pose (position and orientation) of the image sensor 202 and / or the XR system 200 for the input image. 6DoF mapping can also be performed to update the SLAM map. In some cases, the SLAM map maintained using 6DoF SLAM can contain 3D feature points triangulated from two or more images. For example, key frames can be selected from the input image or video stream to represent the observed scene. For each key frame, the corresponding 6DoF camera pose associated with the image can be determined. The pose of the image sensor 202 and / or the XR system 200 can be determined by projecting features from the 3D SLAM map into the image or video frame and updating the camera pose based on the verified 2D-3D correspondences.
[0080] In an illustrative example, the computing component 210 may extract feature points from a particular input image (e.g., each input image, a subset of the input images, etc.) or from each key frame. Feature points (also referred to as registration points) as used herein are distinct or identifiable portions of an image, such as a part of a hand, an edge of a table, etc. Features extracted from the captured images may represent different feature points along a three-dimensional space (e.g., coordinates on the X-axis, Y-axis, and Z-axis), and each feature point may have an associated feature location. Feature points in a key frame may match (be the same or corresponding) or not match feature points of a previously captured input image or key frame. Feature detection may be used to detect the respective feature points. Feature detection may include image processing operations for examining one or more pixels of an image to determine whether a feature exists at a particular pixel. Feature detection may be used to process the entire captured image or certain portions of the image. For each image or key frame, once a feature is detected, a local image patch surrounding the feature may be extracted. Any suitable technique may be used to extract features, such techniques as Scale-Invariant Feature Transform (SIFT) (which locates features and generates their descriptions), Learned Invariant Feature Transform (LIFT), Speeded-Up Robust Features (SURF), Gradient Location-Orientation Histogram (GLOH), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoint (FREAK), KAZE, Accelerated-KAZE (AKAZE), Normalized Cross-Correlation (NCC), descriptor matching, another suitable technique, or a combination thereof.
[0081] As an illustrative example, the computing component 210 may extract feature points corresponding to a mobile device (e.g., Figure 4 mobile device 440 of, Figure 5 mobile device 540 of) etc. In some cases, feature points corresponding to the mobile device may be tracked to determine the pose of the mobile device. As described in more detail below, the pose of the mobile device may be used to determine the projection location of AR media content, which may enhance the media content displayed on the display of the mobile device.
[0082] In some cases, the XR system 200 may also track a user's hand and / or fingers to allow the user to interact with and / or control virtual content in the virtual environment. For example, the XR system 200 may track the pose and / or movement of a user's hand and / or fingertips to identify or translate a user interaction with the virtual environment. User interactions may include, for example but not limited to: moving a virtual content item, resizing a virtual content item, selecting an input interface element in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), providing input via the virtual user interface, etc.
[0083] Figure 3 is a block diagram illustrating the architecture of a Simultaneous Localization and Mapping (SLAM) system 300. In some examples, the SLAM system 300 can be or can include an Extended Reality (XR) system, such as Figure 2 the XR system 200. In some examples, the SLAM system 300 can be a wireless communication device, a mobile device or a handheld device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device, a personal computer, a laptop computer, a server computer, a portable video game console, a portable media player, a camera device, a manned or unmanned ground vehicle, a manned or unmanned aerial vehicle, a manned or unmanned water vehicle, a manned or unmanned underwater vehicle, a manned or unmanned vehicle, an autonomous vehicle, a vehicle, a computing system of a vehicle, a robot, another device, or any combination thereof.
[0084] Figure 3 Each of the one or more sensors 305 of the SLAM system 300 includes or is coupled to. Each of the one or more sensors 305 can include one or more cameras 310. Each of the one or more cameras 310 can include an image capture device 105A, an image processing device 105B, an image capture and processing system 100, another type of camera, or a combination thereof. Each of the one or more cameras 310 can respond to light from a specific spectrum. The spectrum can be a subset of the electromagnetic (EM) spectrum. For example, each of the one or more cameras 310 can be a VL camera that responds to the visible light (VL) spectrum, an IR camera that responds to the infrared (IR) spectrum, a UV camera that responds to the ultraviolet (UV) spectrum, a camera that responds to light of another spectrum from another part of the electromagnetic spectrum, or a certain combination thereof.
[0085] The one or more sensors 305 can include one or more other types of sensors in addition to the cameras 310, such as one or more of each of the following: an accelerometer, a gyroscope, a magnetometer, an inertial measurement unit (IMU), an altimeter, a barometer, a thermometer, a radio detection and ranging (RADAR) sensor, a light detection and ranging (LIDAR) sensor, a sound navigation and ranging (SONAR) sensor, a sound detection and ranging (SODAR) sensor, a global navigation satellite system (GNSS) receiver, a global positioning system (GPS) receiver, a Beidou satellite navigation system (BDS) receiver, a Galileo receiver, a Global Navigation Satellite System (GLONASS) receiver of Russia, an Indian Constellation Navigation (NavIC) receiver, a Quasi-Zenith Satellite System (QZSS) receiver, a Wi-Fi positioning system (WPS) receiver, a cellular network positioning system receiver, A beacon positioning receiver, a short-range wireless beacon positioning receiver, a personal area network (PAN) positioning receiver, a wide area network (WAN) positioning receiver, a wireless local area network (WLAN) positioning receiver, other types of positioning receivers, other types of sensors discussed herein, or combinations thereof. In some examples, one or more sensors 305 may include Figure 2 any combination of sensors of the XR system 200.
[0086] Figure 3 The SLAM system 300 of Figure 3 includes a visual inertial odometry (VIO) tracker 315. The term visual inertial odometry may also be referred to as visual odometry herein. The VIO tracker 315 receives sensor data 365 from one or more sensors 305. For example, the sensor data 365 may include one or more images captured by one or more cameras 310. The sensor data 365 may include other types of sensor data from one or more sensors 305, such as data from any of the types of sensors 305 listed herein. For example, the sensor data 365 may include IMU data from one or more inertial measurement units (IMUs) in one or more sensors 305.
[0087] When sensor data 365 is received from one or more sensors 305, the VIO tracker 315 uses the feature tracking engine 320 of the VIO tracker 315 to perform feature detection, extraction, and / or tracking. For example, in a case where the sensor data 365 includes one or more images captured by one or more cameras 310 of the SLAM system 300, the VIO tracker 315 can identify, detect, and / or extract features in each image. Features can include visually distinct points in an image, such as portions depicting the edges and / or corners of the image. The VIO tracker 315 can receive sensor data 365 from one or more sensors 305 periodically and / or continuously, for example, by continuing to receive more images from one or more cameras 310 when the one or more cameras 310 capture video, where the images are the respective video frames of the video. The VIO tracker 315 can generate descriptors of the features. The feature descriptors can be generated at least in part by generating a description of the feature as depicted in a local image patch extracted around the feature. In some examples, the feature descriptors can describe the features as a set of one or more feature vectors. In some cases, the VIO tracker 315, together with the mapping engine 330 and / or the relocalization engine 355, can associate multiple features with a map of the environment based on such feature descriptors. The feature tracking engine 320 of the VIO tracker 315 can perform feature tracking by identifying features in each image that the VIO tracker 315 has previously identified in one or more previous images based on identifying features with matching feature descriptors in different images. The feature tracking engine 320 can track changes in the positions of one or more of the features depicted in each of the different images. For example, the feature extraction engine can detect a specific corner of a room depicted in the left side of a first image captured by a first camera among the cameras 310. The feature extraction engine can detect the same feature (e.g., the same specific corner of the same room) depicted in the right side of a second image captured by the first camera. The feature tracking engine 320 can identify that the features detected in the first image and the second image are two depictions of the same feature (e.g., the same specific corner of the same room) and that the feature appears in two different positions in the two images. The VIO tracker 315 can determine that the first camera has moved based on the same feature appearing on the left side of the first image and the right side of the second image, e.g., whether the feature (e.g., the specific corner of the room) depicts a static part of the environment.
[0088] The VIO tracker 315 may include a sensor integration engine 325. The sensor integration engine 325 may use sensor data from other types of sensors 305 (in addition to the camera 310) to determine information that the feature tracking engine 320 may use when performing feature tracking. For example, the sensor integration engine 325 may receive IMU data from the IMU of one or more sensors 305 (e.g., it may be included as part of the sensor data 365). The sensor integration engine 325 may determine, based on the IMU data in the sensor data 365, that the SLAM system 300 has rotated 15 degrees in the clockwise direction from obtaining or capturing a first image by the first camera in the camera 310 to obtaining or capturing a second image. Based on this determination, the sensor integration engine 325 may identify that a feature depicted at a first location in the first image is expected to appear at a second location in the second image, and the second location is expected to be to the left of the first location by a predetermined distance (e.g., a predetermined number of pixels, inches, centimeters, millimeters, or another distance metric). The feature tracking engine 320 may consider this expectation when tracking features between the first image and the second image.
[0089] Based on the feature tracking by the feature tracking engine 320 and / or the sensor integration by the sensor integration engine 325, the VIO tracker 315 can determine the 3D feature position 373 of a particular feature. The 3D feature position 373 can include one or more 3D feature positions and can also be referred to as 3D feature points. The 3D feature position 373 can be a set of coordinates along three different axes perpendicular to each other, such as an X coordinate along the X-axis (e.g., in the horizontal direction), a Y coordinate along the Y-axis perpendicular to the X-axis (e.g., in the vertical direction), and a Z coordinate along the Z-axis perpendicular to both the X-axis and the Y-axis (e.g., in the depth direction). The VIO tracker 315 can also determine one or more key frames 370 (hereinafter referred to as key frames 370) corresponding to a particular feature. The key frame (from one or more key frames 370) corresponding to a particular feature can be an image in which the particular feature is clearly depicted. In some examples, the key frame (from one or more key frames 370) corresponding to a particular feature can be an image in which the particular feature is clearly depicted. In some examples, the key frame corresponding to a particular feature can be an image that reduces the uncertainty of the 3D feature position 373 of the particular feature when the feature tracking engine 320 and / or the sensor integration engine 325 consider it for determining the 3D feature position 373. In some examples, the key frame corresponding to a particular feature also includes data about the pose 385 of the SLAM system 300 and / or the (a) camera 310 during the capture of the key frame. In some examples, the VIO tracker 315 can send the 3D feature position 373 and / or the key frame 370 corresponding to one or more features to the mapping engine 330. In some examples, the VIO tracker 315 can receive a mapping slice 375 from the mapping engine 330. The VIO tracker 315 can characterize the information features within the mapping slice 375 for feature tracking using the feature tracking engine 320.
[0090] Based on the feature tracking by the feature tracking engine 320 and / or the sensor integration by the sensor integration engine 325, the VIO tracker 315 can determine the pose 385 of the SLAM system 300 and / or the camera 310 during the capture of each image in the sensor data 365. The pose 385 can include the position of the SLAM system 300 and / or the camera 310 in 3D space, such as a set of coordinates along three different axes perpendicular to each other (e.g., an X coordinate, a Y coordinate, and a Z coordinate). The pose 385 can include the orientation of the SLAM system 300 and / or the camera 310 in 3D space, such as pitch, roll, yaw, or some combination thereof. In some examples, the VIO tracker 315 can send the pose 385 to the relocalization engine 355. In some examples, the VIO tracker 315 can receive the pose 385 from the relocalization engine 355.
[0091] The SLAM system 300 further includes a mapping engine 330. The mapping engine 330 generates a 3D map of the environment based on the 3D feature positions 373 and / or key frames 370 received from the VIO tracker 315. The mapping engine 330 may include a mapping densification engine 335, a key frame remover 340, a bundle adjuster 345, and / or a loop detector 350. The mapping densification engine 335 may perform mapping densification, in some examples, increasing the number and / or density of 3D coordinates describing the mapping geometry. The key frame remover 340 may remove key frames, and / or in some cases add key frames. In some examples, the key frame remover 340 may remove key frames 370 corresponding to regions of the mapping to be updated and / or regions with relatively low corresponding confidence values. In some examples, the bundle adjuster 345 may refine the 3D coordinates describing the scene geometry, the parameters of relative motion, and / or the optical characteristics of the image sensor used to generate the frames according to an optimal criterion involving the corresponding image projections of all points. The loop detector 350 may identify when the SLAM system 300 has returned to a previously mapped region and may use such information to update the map slices and / or reduce the uncertainty of certain 3D feature points or other points in the mapping geometry. The mapping engine 330 may output a map slice 375 to the VIO tracker 315. The map slice 375 may represent a 3D portion or subset of the map. The map slice 375 may include a map slice 375 representing a new, previously unmapped region of the map. The map slice 375 may include a map slice 375 representing an update (or modification or revision) of a previously mapped region of the map. The mapping engine 330 may output mapping information 380 to the relocalization engine 355. The mapping information 380 may include at least a portion of the map generated by the mapping engine 330. The mapping information 380 may include one or more 3D points constituting the geometry of the map, such as one or more 3D feature positions 373. The mapping information 380 may include one or more key frames 370 corresponding to certain features and certain 3D feature positions 373.
[0092] The SLAM system 300 further includes a relocalization engine 355. The relocalization engine 355 can perform relocalization, for example, when the VIO tracker 315 fails to identify more than a threshold number of features in an image and / or the VIO tracker 315 loses track of the pose 385 of the SLAM system 300 within the map generated by the mapping engine 330. The relocalization engine 355 can perform relocalization by performing extraction and matching using the extraction and matching engine 360. For example, the extraction and matching engine 360 can extract features from an image captured by the camera 310 of the SLAM system 300 when the SLAM system 300 is in the current pose 385, and can match the extracted features with features depicted in different keyframes 370 identified by the 3D feature locations 373 and / or identified in the mapping information 380. By matching these extracted features with previously identified features, the relocalization engine 355 can identify that the pose 385 of the SLAM system 300 is the pose 385 at which the previously identified features are visible to the camera 310 of the SLAM system 300, and thus is similar to one or more previous poses 385 at which the previously identified features were visible to the camera 310. In some cases, the relocalization engine 355 can perform relocalization based on wide baseline mapping or the distance between the current camera position and the camera position at which the features were originally captured. The relocalization engine 355 can receive information about the pose 385 (e.g., information about one or more recent poses of the SLAM system 300 and / or the camera 310) from the VIO tracker 315, and the relocalization engine 355 can base its relocalization determination on this information. Once the relocalization engine 355 relocalizes the SLAM system 300 and / or the camera 310 and thus determines the pose 385, the relocalization engine 355 can output the pose 385 to the VIO tracker 315.
[0093] In some examples, the VIO tracker 315 may modify the images in the sensor data 365 and then perform feature detection, extraction, and / or tracking on the modified images. For example, the VIO tracker 315 may rescale and / or resample the image. In some examples, rescaling and / or resampling the image may include downscaling, subsampling, subscaling, and / or subsampling the image one or more times. In some examples, the VIO tracker 315 modifying the image may include converting the image from color to grayscale, or from color to black and white, such as by desaturating the colors in the image, removing certain color channels, reducing the color depth in the image, replacing the colors in the image, or a combination thereof. In some examples, the VIO tracker 315 modifying the image may include the VIO tracker 315 masking certain regions of the image. Dynamic objects may include objects that may have a changing appearance between one image and another. For example, a dynamic object may be an object that moves within the environment, such as a person, a vehicle, or an animal. A dynamic object may be an object that has a changing appearance at different times, such as a display screen that may display different content at different times. A dynamic object may be an object that has a changing appearance based on the pose of the camera(s) 310, such as a reflective surface, a prism, or a mirror surface that reflects, refracts, and / or scatters light in different ways depending on the position of the camera 310 relative to the dynamic object. The VIO tracker 315 may detect dynamic objects using face detection, face recognition, face tracking, object detection, object recognition, object tracking, or a combination thereof. The VIO tracker 315 may detect dynamic objects using one or more artificial intelligence algorithms, one or more trained machine learning models, one or more trained neural networks, or a combination thereof. The VIO tracker 315 may mask the one or more dynamic objects by covering an image region that includes the depiction(s) of the one or more dynamic objects in the image with a mask. The mask may be an opaque color, such as black. The region may be a bounding box having a rectangular or other polygonal shape. The region may be determined on a pixel-by-pixel basis.
[0094] Figure 4 An example of an augmented reality enhanced application engine 400 is illustrated. In the illustrative example, the augmented reality enhanced application engine 400 includes a simulation engine 405, a rendering engine 415, a main rendering module 420, and an AR rendering module 460. As illustrated, the main rendering module 420 may include an effects rendering engine 425, the rendering engine 415, a post-processing engine 430, and a UI rendering engine 435. The AR rendering module 460 may include an AR effects rendering engine 465 and an AR UI rendering engine 470. It should be noted that Figure 4 the components 405 - 470 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples may include more than Figure 4more, fewer, or different components than those shown.
[0095] In some cases, the augmented reality enhanced application engine 400 is included in the electronic device 440 and / or in (wired or wireless) communication with the electronic device 440. In some examples, the augmented reality enhanced application engine 400 is included in the XR system 450 and / or in (wired or wireless) communication with the XR system 450. In an illustrative example, the XR system can be an XR headset as Figure 4 illustrated therein.
[0096] In Figure 4 the example illustrated therein, the simulation engine 405 can generate a simulation for the augmented reality enhanced application engine 400. In some implementations, the simulation engine 405 can optionally include at least one or more of a physical properties module 406, an animation module 407, an audio module 408, an artificial intelligence (AI) module 409, or a gameplay module 411. In some cases, the simulation can include, for example, one or more images, one or more videos, one or more strings of characters (e.g., alphanumeric characters, numbers, text, Unicode characters, code points, and / or icons), one or more two-dimensional (2D) shapes (e.g., circles, ellipses, squares, rectangles, triangles, other polygons, rounded polygons with one or more rounded corners, parts thereof, or combinations thereof), one or more three-dimensional (3D) shapes (e.g., spheres, cylinders, cubes, pyramids, triangular prisms, rectangular prisms, tetrahedrons, other polyhedrons, rounded polyhedrons with one or more rounded edges and / or corners, parts thereof, or combinations thereof), textures of shapes, bump mapping of shapes, lighting effects, or combinations thereof. In some examples, the simulation can include at least a portion of an environment. The environment can be a real-world environment, a virtual environment, and / or a mixed environment that includes elements of a real-world environment and elements of a virtual environment.
[0097] In some cases, the physical properties module 406 can simulate the physical properties of objects within a virtual environment. For example, the physical properties module 406 can simulate the effects of gravity, collisions, fluid dynamics, any other physical properties, and / or any combination thereof for the simulation. In some cases, the animation module 407 can generate animation effects for the simulation. In some aspects, the animation module 407 can be used to provide the appearance of movement of an object by changing the texture. For example, the animation module 407 can apply different textures that create the appearance of movement to a character model in a 2D environment (e.g., by applying sprites in which the limbs of the character model are in different positions). In some cases, the animation module 407 can change the geometry of an object to provide the appearance of movement of the object.
[0098] In some examples, the audio module 408 of the simulation engine 405 can generate audio for the simulation. For example, the audio module 408 can play music, sound effects, voices, any other sounds, and / or any combination thereof.
[0099] In some aspects, the AI module 409 can generate the behaviors of objects (such as people, animals, vehicles, electronic devices, etc.) for the simulation. In an illustrative example, in the context of a video game simulation, the module 409 can generate the behaviors of non-player characters. Some example behaviors can include movement patterns, dialogue content, etc.
[0100] In some cases, the gameplay module 411 can generate, retrieve, store, and / or update the gameplay state for the simulation. In some cases, the gameplay module 411 can track gameplay progress, player character position, player inventory, any other gameplay content, and / or any combination thereof.
[0101] In some cases, the simulation generated by the simulation engine 405 can be dynamic. For example, the simulation engine 405 can update the simulation based on different triggers, including but not limited to physical contact, sound, posture, input signals, the passage of time, and / or any combination thereof. As used herein, the application state of the augmented reality enhanced application engine 400 can include any information associated with the simulation engine 405, the rendering engine 415, the main rendering module 420, the effect rendering engine 425, the post-processing engine 430, the UI rendering engine 435, the AR rendering module 460, the AR effect rendering engine 465, the AR UI rendering engine 470, the input to the augmented reality enhanced application engine 400, the output from the augmented reality enhanced application engine 400, and / or any combination thereof at a specific moment.
[0102] As illustrated, the simulation engine 405 can obtain the mobile device input 441 from the mobile device 440. In some cases, the simulation engine 405 can obtain the XR system input 451 from the XR system 450. The mobile device input 441 and / or the XR system input 451 can include, for example, user input through the user interface of an application displayed on the display of the mobile device 440, user input from an input device (such as Figure 2 the input device 208), user input from one or more sensors (such as Figure 2 the image sensor 202, the accelerometer 204, the gyroscope 206). In some cases, the simulation engine 405 can update the application state of the augmented reality enhanced application engine 400 based on the mobile device input 441, the XR system input 451, and / or any combination thereof.
[0103] In Figure 4In the illustrative example, the rendering engine 415 can obtain application state information from the simulation engine 405. In some implementations, the rendering engine 415 can optionally include at least one or more of an environment module 416, a character module 417, a decal module 418, or a prop module 419. In some implementations, the environment module can render the environment (e.g., buildings, landscapes, trees, etc.) being simulated by the application (e.g., game, virtual world). In some aspects, the character module 417 can render playable and / or non-player characters for simulation. In some examples, the decal module 418 can render details (e.g., footprints, damage to objects, etc.) superimposed on portions of the environment. In some cases, the prop module 419 can render dynamic elements related to the simulation. For example, in the case of a game, the prop module 419 can render inventory items (e.g., health packs, weapons, tools, etc.).
[0104] In some cases, the rendering engine 415 can determine portions of the application state information to be rendered by the displays available to the augmented reality enhanced application engine 400. In some implementations, the modules 416 - 419 of the rendering engine 415 can be configured to determine different types of content to be rendered from the simulation for the displays available to the augmented reality enhanced application engine 400. For example, the rendering engine 415 can determine whether a connection (wired or wireless) has been established between the XR system 450 and the mobile device 440. In some cases, the rendering engine 415 can determine the application state information to be rendered by the main rendering module 420 and the AR rendering module 460. In some cases, the rendering engine 415 can determine that the XR system 450 is not (wired or wirelessly) connected to the mobile device 440. In some cases, the rendering engine 415 can determine the application state information of the main rendering module 420 and forego determining the application state information to be rendered by the AR rendering module 460 that will not be displayed. Accordingly, the rendering engine 415 can facilitate an adaptive rendering configuration for the augmented reality enhanced application engine 400 based on the availability and / or type of the available displays. In some implementations, as Figure 4 shown, the separate rendering engine 415 can be excluded. In an illustrative example, the main rendering module 420 and / or the AR rendering module 460 can include at least a portion of the functionality of the rendering engine 415 described above.
[0105] The main rendering module 420 may include an effects rendering engine 425, a post-processing engine 430, and a UI rendering engine 435. In some cases, the main rendering module 420 may render image frames configured for display on the display of the mobile device 440. As illustrated, the main rendering module 420 may output the generated image frames (e.g., media content) for display on the display of the mobile device 440. In some cases, the effects rendering information may render the application state information generated by the simulation engine 405. For example, the effects rendering engine 425 may generate a 2D projection of a portion of the 3D environment included in the application state information. For example, the main rendering module 420 may generate a perspective projection of the 3D environment via a virtual camera. In some cases, the application state information may include the pose of the virtual camera within the environment. In some cases, the effects rendering engine 425 may generate additional visual effects not included within the 3D environment. For example, the effects rendering engine 425 may apply texture mapping to enhance the visual appearance of the effects generated by the rendering engine 415. In some cases, the main rendering module 420 may exclude portions of the application state information specified by the rendering engine 415 for the AR rendering module 460. For example, the main rendering module 420 may exclude effects present in the simulated environment.
[0106] In some cases, the post-processing engine 430 may provide additional processing to the rendering effects generated by the effects rendering engine 425. For example, the post-processing engine 430 may perform scaling, image smoothing, z-buffering, contrast enhancement, gamma correction, color mapping, any other image processing, and / or any combination thereof.
[0107] In some implementations, the UI rendering engine 435 may render the UI (e.g., Figure 6 the UI 615). In some cases, in addition to the effects rendered based on the application environment (e.g., the 3D environment), the user interface may provide application state information. In some cases, the UI may be generated to overlay a portion of the image frames output by the post-processing engine 430.
[0108] The AR rendering module 460 may include an AR effects rendering engine 465 and an AR UI rendering engine 470. In some cases, the AR effects rendering engine 465 may render the application state information generated by the simulation engine 405. For example, the AR effects rendering engine 465 may generate a 2D projection of the 3D environment included in the application state information. In some cases, the AR effects rendering engine 465 may generate effects that appear to protrude from the display surface of the display of the mobile device 440.
[0109] In some cases, the display of the XR system 450 may have different display parameters (e.g., different resolution, frame rate, aspect ratio, and / or any other display parameter) from the display of the mobile device 440. In some cases, the display parameters may also vary between different types of output devices (e.g., different HMD models, other XR systems, etc.). In some cases, including the AR rendering module 460 within the augmented reality enhanced application engine 400 may require periodic updates to provide compatibility with different devices.
[0110] Figure 5 An example of a main application 500 and an accompanying application 560 that can provide augmented reality enhancement to the main application 500 is illustrated. In Figure 5 the illustrative example, the main application 500 includes a simulation engine 505, a rendering engine 515, and a communication engine 525. In the illustrated example, the accompanying application 560 includes a tracking engine 565 (e.g., Figure 2 the XR engine 220 of Figure 3 , the VIO tracker 315 of Figure 5 ), an AR rendering engine 575, and a communication engine 585. As illustrated, the main application 500 and the accompanying application 560 can communicate over a (wired or wireless) communication link 530. It should be noted that Figure 5 the components 505 - 525 shown in the main application 500 of Figure 5 are non - restrictive examples provided for illustrative and explanatory purposes, and other examples may include more, fewer, or different components than those shown in Figure 5 . Similarly, it should be noted that
[0111] In Figure 5In the illustrated example, the simulation engine 505 of the main application 500 may generate a simulation for an application on the mobile device 540. In some cases, the simulation may include, for example, one or more images, one or more videos, one or more strings of characters (e.g., alphanumeric characters, numbers, text, Unicode characters, code points, and / or icons), one or more two-dimensional (2D) shapes (e.g., circles, ellipses, squares, rectangles, triangles, other polygons, rounded polygons with one or more rounded corners, portions thereof, or combinations thereof), one or more three-dimensional (3D) shapes (e.g., spheres, cylinders, cubes, pyramids, triangular prisms, rectangular prisms, tetrahedrons, other polyhedrons, rounded polyhedrons with one or more rounded edges and / or corners, portions thereof, or combinations thereof), textures of the shapes, bump maps of the shapes, lighting effects, or combinations thereof. In some examples, the simulation may include at least a portion of an environment. The environment may be a real-world environment, a virtual environment, and / or a mixed environment including elements of the real-world environment and the virtual environment.
[0112] In some cases, the simulation generated by the simulation engine 505 may be dynamic. For example, the simulation engine 505 may update the simulation based on different triggers, including but not limited to physical contact, sound, gesture, input signals, the passage of time, and / or any combination thereof. As used herein, the application state of the main application 500 may include any information associated with the simulation engine 505, the effect rendering engine 515, the communication engine 525, and / or any combination thereof at a particular moment.
[0113] In some cases, the rendering engine 515 may correspond to Figure 4 the rendering engine 415, the main rendering module 420, the AR rendering module 460, and / or any combination thereof and perform similar functions. For example, the rendering engine 515 may include a module for effect rendering (e.g., Figure 4 the effect rendering engine 425), a module for post-processing (e.g., Figure 4 the post-processing engine 430), and / or a module for UI rendering (e.g., Figure 4 the UI rendering engine 435).
[0114] The communication engine 525 of the main application 500 and the communication engine 585 of the companion application 560 can communicate over the communication link 530. In some cases, the communication link 530 can be bidirectional. In some examples, the communication engine 525 can transmit application state information (e.g., from the analog engine 505) to the communication engine 585 of 560. In some cases, the application state information can include information that can be used to generate AR effects. In some examples, the application state information can include data that can be used by the companion application 560 to generate an AR UI. In some cases, the communication engine 525 can also transmit inputs obtained from 540 over the communication link 530 to the communication engine 585. In some cases, the communication engine 585 of the companion application 560 can transmit pose information, connectivity status, user input, etc. to the communication engine 525 of the main application 500. The communication engine 525 and the communication engine 585 can also transmit and / or receive synchronization signals for synchronized display between the display of the mobile device 540 and the display of the HMD 550. The examples of communication between the communication engine 525 and the communication engine 585 provided herein are non-limiting and are provided as examples. In some cases, more, fewer, and / or different information can be communicated over the communication link 530 without departing from the scope of the present disclosure.
[0115] Referring to the companion application 560, the tracking engine 565 can use information captured by sensors (e.g., Figure 2 the image sensor 202, accelerometer 204, gyroscope 206, one or more sensors 305, Figure 3 the camera 310, etc.) of to perform tracking (e.g., SLAM, VIO). In some cases, the tracking engine 565 can determine the pose of the mobile device 540, the pose of the HMD 550, environmental mapping, etc. In some aspects, the tracking engine 565 can determine the outline of the display of the mobile device 540. In some cases, the outline of the display of the mobile device 540 can include boundaries. In some cases, the pose of the mobile device 540 and / or the outline and / or boundaries of the display of the mobile device 540 can be output to the AR rendering engine 575 to provide a target for displaying AR information (e.g., AR effects, AR UI) on the display screen of the HMD 550.
[0116] The AR rendering module 460 can be similar to Figure 4 the AR rendering module 460 of and perform similar functions. For example, in some implementations, the HMD 550 can include an AR effect rendering engine (e.g., Figure 4 the AR effect rendering engine 465 of ) and / or an AR UI rendering engine (e.g., Figure 4The AR UI rendering engine 470). In some cases, the AR rendering engine 575 may output AR media content to the HMD 550 using display parameters (e.g., different resolutions, frame rates, aspect ratios, and / or any other display parameters) different from the media content output from the rendering engine 515 to the mobile device 540. In some cases, by dividing the rendering functionality between the main application 500 and the companion application 560, the computing resources for providing an AR-enhanced application experience can be shared among the computing resources of multiple devices (such as the mobile device 540 and the HMD 550). Additionally, providing a separate AR rendering engine 575 in the companion application 560 can simplify the development of the main application 500. For example, the rendering engine 515 of the main application 500 may not need to maintain compatibility with various different mobile devices having different display configurations.
[0117] Figure 6 Illustrates an example configuration 600 of a video game application 610 with augmented reality enhancements. In Figure 6 the illustrative example, the video game application 610 is displayed on the display 605 of a mobile device 640 (e.g., Figure 4 the mobile device 440 of Figure 5 the mobile device 540). The user is depicted as wearing an HMD 650 (e.g., Figure 4 the XR system 450 of Figure 5 the HMD 550) while playing the game application 610 on the mobile device 640. The video game application may include a simulation of a game environment (e.g., via a simulation engine 405). The game environment may include, for example, a 3D simulation environment. The 3D simulation environment may include, for example, 3D models corresponding to people (e.g., players), structures, objects, terrain, etc. The game simulation may also include user 2D elements, such as a UI 615. In some cases, the UI 615 may include user input controls (e.g., movement controls, inventory item selection, etc.), status information (e.g., health status, mapped location, inventory item content and / or quantity). However, as shown, the user interface information displayed on the display of the mobile device 640 may obscure the visibility of the game simulation of the game application 610. As Figure 6 illustrated, the video game application 610 may be enhanced by augmented reality media elements projected by the HMD 650. For example, an extended user interface 625 may be projected relative to the display of the mobile device 640. As illustrated, the extended user interface 625 may be projected to appear parallel to the surface of the display 605 and extend beyond the game area of the game application 610 (e.g., directly below). In some cases, a tracking engine (e.g., Figure 2 the XR engine 220 of Figure 3 the VIO tracker 315 of Figure 5The tracking engine 565) can determine the pose of the mobile device 640 and maintain the projection of the extended user interface 625 relative to the display 605 of the mobile device 640 as the HMD and / or the mobile device 640 moves.
[0118] As Figure 6 illustrated, the game application 610 and / or the extended UI 625 can enhance the user experience of the video game application 610. For example, the extended UI 625 can provide additional information to the user without obscuring the display 605. Additionally, the special effects 620 can increase the realism of the video game application, provide a more immersive experience, enhance the interaction between the user and the game environment, etc.
[0119] Figure 7 is a flowchart illustrating an example of a process 700 for displaying media on one or more displays. At block 702, the process 700 includes generating, by (e.g., via Figure 4 the rendering engine 415, the main rendering module 420, the AR rendering module 460) of, a first media content element associated with an application state of an application engine for the application engine. In some examples, the application state includes a 3D representation of an environment. In some aspects, generating the first media content element includes rendering the 3D representation of the environment in a first display configuration. In some examples, the first display configuration includes at least one or more of resolution, refresh rate, or color depth.
[0120] At block 704, the process 700 includes generating, by (e.g., via Figure 4 the rendering engine 415, the main rendering module 420, the AR rendering module 460) of, a second media content element associated with an application state of an application engine for the application engine. In some cases, generating the second media content element includes rendering the 3D representation of the environment in a second display configuration. In some examples, the second display configuration is different from the first display configuration.
[0121] At block 706, the process 700 includes displaying the first media content element on a first display of a first device (e.g., Figure 4 the mobile device 440) of.
[0122] At block 708, the process 700 includes outputting the second media content element for display on a second display of a second device relative to a pose of the first display of the first device (e.g., on a display of the Figure 4 XR system 450). In some cases, the second media content element includes a stereoscopic projection for the second display of the second device. In some cases, outputting the second media content element to the second device includes transmitting the second media content element from the first device to the second device via a communication link.
[0123] In some examples, process 700 includes obtaining pose information including at least one or more of first pose information associated with a first device or second pose information associated with a second device. In some cases, process 700 includes determining, based on the pose information, a display area for a second media content element on a second display outside a display area of a first display. In some cases, process 700 includes outputting the display area for the second media content element. In some aspects, process 700 includes one or more of the following: determining, by the first device, a pose of the first device based on data from one or more sensors of the first device; obtaining data from one or more sensors of the second device and determining, by the first device, a pose of the first device based on the data from one or more sensors of the second device; obtaining a pose of the first device from the second device; or obtaining a pose of the second device from the second device. In some cases, process 700 includes determining a profile associated with a first display of the first device. In some examples, the profile associated with the first display includes a boundary of the profile. In some aspects, the display area for the second media content element extends from the first display of the first display. In some examples, the display area for the second media content element extends from a boundary of the profile associated with the first display of the first device. In some cases, the display area for the second media content element is configured to at least partially occlude the first display of the first device when being displayed on a second display of a second device (e.g., Figure 4 the XR system 450).
[0124] In some cases, process 700 includes obtaining a first input at the application engine; generating an additional application state different from the application state based on the first input; generating a third media content element for the application engine associated with the additional application state of the application engine; generating a fourth media content element for the application engine associated with the additional application state of the application engine; displaying the third media content element on a first display of the first device; and outputting the fourth media content element for display on a second display of the second device relative to an additional pose of the first device.
[0125] In some examples, process 700 includes generating a first projection of a first media content element including a first portion of the application state; and generating a second projection of a second media content element including a second portion of the application state, the second portion of the application state being at least partially different from the first portion of the application state. In some aspects, process 700 includes the second media content element including user interface elements.
[0126] In some cases, process 700 includes generating the first media content element and the second media content element by a renderer of the application engine.
[0127] In some implementations, the application engine includes a video game application. In some aspects, the first media content element includes a first part of the gameplay simulation of the video game application. In some cases, the first device includes a handheld electronic device. In some examples, the second media content element includes a second part of the gameplay simulation of the video game application. In some cases, the second device includes an HMD. In some cases, the second display of the HMD is configured to project the second media content element relative to the first display.
[0128] Figure 8 is a flowchart illustrating an example of process 800 for displaying media on one or more displays. At block 802, process 800 includes obtaining the application state of an application engine (e.g., Figure 5 the main application 500) from a first device that includes a first display.
[0129] At block 804, process 800 includes generating, for a second device that includes a second display, media content elements associated with the application state of the application engine (e.g., Figure 5 by the companion application 560, the AR rendering engine 575 of
[0130] At block 806, process 800 includes displaying the media content elements on the second display of the second device (e.g., Figure 5 the HMD 550 of
[0131] relative to the pose of the first display of the first device. In some cases, determining the pose of the first display of the first device includes at least one or more of the following: determining, by one or more sensors of the second device, one or more of the pose of the first device or the pose of the second device; or obtaining pose information associated with the first device from the first device.
[0132] Figure 7 The process 700 illustrated in Figure 11 may also include any operations illustrated in or discussed with respect to the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the XR system, the SLAM system 300, or combinations thereof.
[0133] In some cases, at least a subset of the techniques illustrated in process 700 may be performed remotely by one or more web servers of a cloud service. In some examples, the processes described herein (e.g., process 70 and / or other processes described herein) may be performed by a computing device or apparatus. In some examples, process 700 may be performed by Figure 1 image capture device 105A of Figure 1 image processing device 105B of Figure 1 image capture and processing system 100 of Figure 2 XR system of Figure 3 SLAM system 300 of Figures 9A to 9B head-mounted display (HMD) 810 of Figures 10A to 10B mobile device 1050 of Figure 11 its variants, or a combination thereof. Process 700 may also be performed by a computing device having the architecture of computing system 1100 shown in
[0134] The components of a computing device may be implemented in circuitry. For example, the components may include and / or may use electronic circuits or other electronic hardware (which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits)), and / or may include and / or may use computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0135] by Figure 1(Image capture and processing system 100), Figure 2 (XR system 200), Figure 3 (SLAM system 300) and Figure 11 (Computing system 1100) are illustrated or organized as a logical flow chart, and the operations of the logical flow chart represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents computer-executable instructions stored on one or more computer-readable storage media, which, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order of the described operations is not intended to be construed as limiting, and any number of the described operations can be combined and / or performed in parallel in any order to implement the processes.
[0136] Additionally, the processes illustrated by the block diagrams 100, 200, 300 and the flow chart of the illustrated process 700 and / or other processes described herein can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed jointly on one or more processors, implemented by hardware, or a combination thereof. As mentioned above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program including multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0137] Figure 9A is a perspective view 900 of a head-mounted display (HMD) 910 that illustrates execution feature tracking and / or visual simultaneous localization and mapping (VSLAM) according to some examples. The HMD 910 can be, for example, an augmented reality (AR) head-mounted device, a virtual reality (VR) head-mounted device, a mixed reality (MR) head-mounted device, an extended reality (XR) head-mounted device, or some combination thereof. The HMD 910 can be an example of the XR system 200, the SLAM system 300, or a combination thereof. The HMD 910 includes a first camera 930A and a second camera 930B along the front of the HMD 910. The first camera 930A and the second camera 930B can be two cameras among one or more cameras 310. In some examples, the HMD 910 can have only a single camera. In some examples, the HMD 910 can include one or more additional cameras in addition to the first camera 930A and the second camera 930B. In some examples, the HMD 910 can include one or more additional sensors in addition to the first camera 930A and the second camera 930B.
[0138] Figure 9B The explanation is based on some examples Figure 9A 930 of a head mounted display (HMD) 910 worn by a user 920. The user 920 wears the HMD 910 on the head of the user 920, above the eyes of the user 920. The HMD 910 can capture images using a first camera 930A and a second camera 930B. In some examples, the HMD 910 displays one or more display images based on the images captured by the first camera 930A and the second camera 930B toward the eyes of the user 920. The display images can provide a stereoscopic view of the environment, with superimposed information and / or with other modifications in some cases. For example, the HMD 910 can display a first display image to the right eye of the user 920, the first display image being based on the image captured by the first camera 930A. The HMD 910 can display a second display image to the left eye of the user 920, the second display image being based on the image captured by the second camera 930B. For example, the HMD 910 can provide superimposed information in a display image superimposed on the image captured by the first camera 930A and the second camera 930B.
[0139] The HMD 910 itself does not include wheels, propellers, or other means of transportation. Instead, the HMD 910 relies on the movement of the user 920 to move the HMD 910 around the environment. Thus, in some cases, the HMD 910 may skip path planning using a path planning engine and / or movement actuation using movement actuators when performing SLAM technology. In some cases, the HMD 910 may still use the path planning engine to perform path planning and may indicate to the user 920 the direction to follow the recommended path to guide the user along the recommended path planned using the path planning engine. In some cases, such as when the HMD 910 is a VR headset, the environment may be completely or partially virtual. If the environment is at least partially virtual, movement through the virtual environment may also be virtual. For example, movement through the virtual environment may be controlled by the input device 208. The movement actuator may include any such input device 208. Movement through the virtual environment may not require wheels, propellers, legs, or any other form of transportation. If the environment is a virtual environment, the HMD 910 may still use the path planning engine and / or movement actuation to perform path planning. If the environment is a virtual environment, the HMD 910 may perform movement actuation by performing virtual movement within the virtual environment and using movement actuators. Even if the environment is virtual, SLAM technology may still be valuable because the virtual environment may be unmapped and / or may have been generated by a device other than the HMD 910, such as a remote server or console associated with a video game or video game platform. In some cases, feature tracking and / or SLAM may even be performed in a virtual environment by a vehicle or other device having its own physical conveyance system that allows it to physically move around the physical environment. For example, SLAM may be performed in a virtual environment to test whether the SLAM system 300 is working properly without wasting time or energy on movement and without wearing out the physical conveyance system.
[0140] Figure 10AFIG. 1000 is a perspective view of a front surface 1055 of a mobile device 1050 according to some examples, the mobile device 1050 using one or more front cameras 1030A to 1030B to perform feature tracking and / or visual simultaneous localization and mapping (VSLAM). The mobile device 1050 can be, for example, a cellular phone, a satellite phone, a portable game console, a music player, a health tracking device, a wearable device, a wireless communication device, a laptop device, a mobile device, any other type of computing device or computing system 1100 discussed herein, or a combination thereof. The front surface 1055 of the mobile device 1050 includes a display screen 1045. The front surface 1055 of the mobile device 1050 includes a first camera 1030A and a second camera 1030B. The first camera 1030A and the second camera 1030B are illustrated in a bezel around the display screen 1045 on the front surface 1055 of the mobile device 1050. In some examples, the first camera 1030A and the second camera 1030B can be located in a notch or cutout cut out from the display screen 1045 on the front surface 1055 of the mobile device 1050. In some examples, the first camera 1030A and the second camera 1030B can be under-display cameras located between the display screen 1045 and the remainder of the mobile device 1050 such that light passes through a portion of the display screen 1045 before reaching the first camera 1030A and the second camera 1030B. The first camera 1030A and the second camera 1030B in FIG. 1000 are front cameras. The first camera 1030A and the second camera 1030B face a direction perpendicular to the planar surface of the front surface 1055 of the mobile device 1050. The first camera 1030A and the second camera 1030B can be two cameras among one or more cameras 310. In some examples, the front surface 1055 of the mobile device 1050 can have only a single camera. In some examples, the mobile device 1050 can include one or more additional cameras in addition to the first camera 1030A and the second camera 1030B. In some examples, the mobile device 1050 can include one or more additional sensors in addition to the first camera 1030A and the second camera 1030B.
[0141] Figure 10Bis a perspective view 1090 showing the rear surface 1065 of the mobile device 1050. The mobile device 1050 includes a third camera 1030C and a fourth camera 1030D located on the rear surface 1065 of the mobile device 1050. The third camera 1030C and the fourth camera 1030D in the perspective view 1090 are rear-facing. The third camera 1030C and the fourth camera 1030D face a direction perpendicular to the plane of the rear surface 1065 of the mobile device 1050. Although the rear surface 1065 of the mobile device 1050 does not have a display screen 1045 (as illustrated in the perspective view 1090), in some examples, the rear surface 1065 of the mobile device 1050 may have a second display screen. If the rear surface 1065 of the mobile device 1050 has a display screen 1045, any positioning of the third camera 1030C and the fourth camera 1030D relative to the display screen 1045 may be used, as discussed with respect to the first camera 1030A and the second camera 1030B located at the front surface 1055 of the mobile device 1050. The third camera 1030C and the fourth camera 1030D may be two cameras among one or more cameras 310. In some examples, the rear surface 1065 of the mobile device 1050 may have only a single camera. In some examples, the mobile device 1050 may include one or more additional cameras in addition to the first camera 1030A, the second camera 1030B, the third camera 1030C, and the fourth camera 1030D. In some examples, the mobile device 1050 may include one or more additional sensors in addition to the first camera 1030A, the second camera 1030B, the third camera 1030C, and the fourth camera 1030D.
[0142] Similar to the HMD 910, the mobile device 1050 itself does not include wheels, propellers, or other means of transportation. Instead, the mobile device 1050 relies on the movement of the user who holds or wears the mobile device 1050 to move the mobile device 1050 around the environment. Thus, in some cases, the mobile device 1050 may skip path planning using a path planning engine and / or movement actuation using movement actuators when performing SLAM techniques. In some cases, the mobile device 1050 may still use the path planning engine to perform path planning and may indicate to the user the direction to follow the recommended path to guide the user along the recommended path planned using the path planning engine. In some cases, such as when the mobile device 1050 is used for AR, VR, MR, or XR, the environment may be completely or partially virtual. In some cases, the mobile device 1050 may be inserted into a head-mounted device (HMD) (e.g., inserted into a bracket of the HMD) such that the mobile device 1050 serves as a display of the HMD, where the display screen 1045 of the mobile device 1050 serves as the display of the HMD. If the environment is at least partially virtual, movement through the virtual environment may also be virtual. For example, movement through the virtual environment may be controlled by one or more joysticks, buttons, video game controllers, mice, keyboards, touchpads, and / or other input devices coupled to the mobile device 1050 in a wired or wireless manner. Movement actuators may include any such input devices. Movement through the virtual environment may not require wheels, propellers, legs, or any other form of transportation. If the environment is a virtual environment, the mobile device 1050 may still use the path planning engine and / or movement actuation to perform path planning. If the environment is a virtual environment, the mobile device 1050 may perform virtual movement within the virtual environment and use movement actuators to perform movement actuation.
[0143] Figure 11 is a diagram illustrating an example of a system for implementing certain aspects of the techniques herein. Specifically, Figure 11 illustrates an example of a computing system 1100, which may be any computing device that makes up, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the XR system, the SLAM system 300, the remote computing system, or any component thereof, where the components of the system are in communication with each other using the connection 1105. The connection 1105 may be a physical connection using a bus or a direct connection to the processor 1110 (such as in a chipset architecture). The connection 1105 may also be a virtual connection, a network connection, or a logical connection.
[0144] In some embodiments, computing system 1100 is a distributed system in which the functions described in this disclosure may be distributed within a data center, multiple data centers, a peer-to-peer network, and so on. In some embodiments, one or more of the described system components represent many such components, each performing some or all of the functions described for that component. In some embodiments, the components may be physical or virtual devices.
[0145] Example computing system 1100 includes at least one processing unit (CPU or processor) 1110 and connection 1105 that couples various system components, including system memory 1115 (such as read-only memory (ROM) 1120 and random access memory (RAM) 1125), to processor 1110. Computing system 1100 may include cache 1112 of high-speed memory that is directly connected to, adjacent to, or integrated as part of processor 1110.
[0146] Processor 1110 may include any general-purpose processor and hardware services or software services, such as services 1132, 1134, and 1136 stored in storage device 1130 and configured to control processor 1110, as well as specialized processors in which software instructions are incorporated into the actual processor design. Processor 1110 may be substantially a fully self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multicore processors may be symmetric or asymmetric.
[0147] To enable user interaction, computing system 1100 includes input device 1145, which may represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, and so on. Computing system 1100 may also include output device 1135, which may be one or more of several output mechanisms. In some instances, a multimodal system may enable a user to provide multiple types of input / output to communicate with computing system 1100. Computing system 1100 may include communication interface 1140, which generally may manage and control user input and system output. The communication interface may perform or facilitate receiving and / or transmitting wired or wireless communications using a wired and / or wireless transceiver, including using an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, port / plug, an Ethernet port / plug, a fiber optic port / plug, a dedicated wired port / plug, wireless signal transmission, low energy (BLE) wireless signal transmission, Those communications that are wireless signal transmissions such as radio frequency identification (RFID) wireless signal transmissions, near field communication (NFC) wireless signal transmissions, dedicated short range communication (DSRC) wireless signal transmissions, 802.11 Wi-Fi wireless signal transmissions, wireless local area network (WLAN) signal transmissions, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), infrared (IR) communication wireless signal transmissions, public switched telephone network (PSTN) signal transmissions, integrated services digital network (ISDN) signal transmissions, 3G / 4G / 5G / LTE cellular data network wireless signal transmissions, ad hoc network signal transmissions, radio wave signal transmissions, microwave signal transmissions, infrared signal transmissions, visible light signal transmissions, ultraviolet light signal transmissions, wireless signal transmissions along the electromagnetic spectrum, or some combination thereof. The communication interface 1140 may also include one or more global navigation satellite system (GNSS) receivers or transceivers that are used to determine the location of the computing system 1100 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0148] The storage device 1130 can be a non-volatile and / or non-transitory and / or computer-readable memory device and can be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as cassette tapes, flash memory cards, solid state memory devices, digital versatile discs, cartridges, floppy disks, hard disks, magnetic tapes, magnetic strips / magnetic stripes, any other magnetic storage medium, flash memory, memristor memory, any other solid state memory, compact disc read only memory (CD-ROM) optical discs, rewritable compact discs (CD) optical discs, digital video disc (DVD) optical discs, Blu-ray disc (BDD) optical discs, holographic optical discs, another optical medium, secure digital (SD) cards, micro secure digital (microSD) cards, Memory Stick Smart card chips, EMV chips, subscriber identity module (SIM) cards, mini / micro / nano / pico SIM cards, other integrated circuit (IC) chips / cards, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), other memory chips or cartridges, and / or combinations thereof.
[0149] The storage device 1130 may include software services, servers, services, etc., which when the code defining such software is executed by the processor 1110 cause the system to perform functions. In some embodiments, the hardware services that perform a particular function may include software components stored in a computer-readable medium connected to the necessary hardware components (such as the processor 1110, connection 1105, output device 1135, etc.) to perform the function.
[0150] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a procedure, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. Code segments may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted using any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0151] In some embodiments, computer-readable storage devices, media, and memories may include wires or wireless signals that contain bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0152] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, one of ordinary skill in the art will understand that these embodiments may be practiced without these specific details. For clarity of illustration, in some instances, the techniques of the present invention may be presented as including various functional blocks, which include functional blocks containing devices, device components, steps or routines in a method implemented in software or a combination of hardware and software. Additional components other than those shown and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components to avoid obscuring these embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.
[0153] The various embodiments may be described above as a process or method, which is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process terminates when its operations are completed, but the process may have additional steps not included in the figures. The process may correspond to a method, function, procedure, subroutine, subprogram, etc. When the process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0154] The processes and methods according to the above examples may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. These instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or a group of functions. Portions of the computer resources used may be accessed via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used during the methods according to the described examples, and / or information created include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, etc.
[0155] Devices implementing the various processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may execute the necessary tasks. Typical examples of form factors include: laptop devices, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be implemented with peripheral devices or plug-in cards. As a further example, such functionality may also be implemented on a circuit board among different chips or different processes executing on a single device.
[0156] Instructions, a medium for conveying these instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0157] In the foregoing description, aspects of the present application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the illustrative embodiments of the present application have been described in detail herein, it is to be understood that the various inventive concepts may be implemented and employed in other various ways, and the appended claims are not to be construed as including such variations, unless limited by the prior art. The various features and aspects of the foregoing application may be used singly or in combination. In addition, the embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, this specification and the drawings are to be regarded as illustrative rather than limiting. For purposes of illustration, the methods are described in a particular order. It should be appreciated that in alternative embodiments, the methods may be performed in a different order than that described.
[0158] Those of ordinary skill in the art will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this specification.
[0159] In cases where components are described as “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0160] The phrase "coupled to" means that any component is physically connected to another component directly or indirectly, and / or any component is in communication with another component directly or indirectly (e.g., connected to the other component through a wired or wireless connection and / or other suitable communication interface).
[0161] The phrase "at least one of" in a claim language or set of recitations and / or other language indicating "one or more of" in the set means that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting "at least one of A and B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" in a set and / or "one or more of" in a set does not limit the set to the items enumerated in the set. For example, claim language reciting "at least one of A and B" can mean A, B, or A and B, and can additionally include items not enumerated in the set of A and B.
[0162] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of this application.
[0163] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. These techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device with multiple uses, including applications in a wireless communication device handset and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially realized by a computer-readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. These techniques may additionally or alternatively be at least partially realized by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that may be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0164] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such processors may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within a special purpose software module or hardware module configured for encoding and decoding, or incorporated into a combined video codec (CODEC).
[0165] Exemplary aspects of this disclosure include:
[0166] Aspect 1. A method for displaying media on one or more displays, comprising: generating a first media content element for an application engine associated with an application state of the application engine; generating a second media content element for the application engine associated with the application state of the application engine; displaying the first media content element on a first display of a first device; and outputting the second media content element for display on a second display of a second device relative to a pose of the first display of the first device.
[0167] Aspect 2. The method of Aspect 1, further comprising: obtaining pose information including at least one or more of first pose information associated with the first device or second pose information associated with the second device; determining, based on the pose information, a display area for the second media content element on the second display outside a display area of the first display; and outputting the display area for the second media content element.
[0168] Aspect 3. The method of any one of Aspects 1-2, wherein obtaining the first pose information includes at least one or more of the following: determining, by the first device, a pose of the first device based on data from one or more sensors of the first device; determining, by the first device, a pose of the first device based on data from one or more sensors of the second device; obtaining, from the second device, a pose of the first device; or obtaining, from the second device, a pose of the second device.
[0169] Aspect 4. The method of any one of Aspects 1 to 3, further comprising: determining a contour associated with the first display of the first device, wherein the contour associated with the first display includes a boundary of the contour, and wherein the display area for the second media content element extends from the first display of the first display.
[0170] Aspect 5. The method of any one of Aspects 1 to 4, wherein the display area for the second media content element extends from a boundary of the contour associated with the first display of the first device.
[0171] Aspect 6. The method of any one of Aspects 1 to 5, wherein the display area for the second media content element is configured to at least partially occlude the first display of the first device when displayed on the second display of the second device.
[0172] Aspect 7. The method of any one of Aspects 1 to 6, wherein the second media content element includes: a stereoscopic projection for the second display of the second device.
[0173] Aspect 8. The method of any one of Aspects 1 to 7, wherein the application state includes a three-dimensional (3D) representation of the environment.
[0174] Aspect 9. The method according to any one of aspects 1 to 8, wherein generating the first media content element includes: rendering a 3D representation of the environment in a first display configuration, wherein the first display configuration includes at least one or more of resolution, refresh rate, or color depth.
[0175] Aspect 10. The method according to any one of aspects 1 to 9, wherein generating the second media content element includes: rendering a 3D representation of the environment in a second display configuration, wherein the second display configuration is different from the first display configuration.
[0176] Aspect 11. The method according to any one of aspects 1 to 10, further comprising: obtaining a first input at the application engine; generating an additional application state different from the application state based on the first input; generating a third media content element for the application engine associated with the additional application state of the application engine; generating a fourth media content element for the application engine associated with the additional application state of the application engine; displaying the third media content element on a first display of the first device; and outputting the fourth media content element for display on a second display of a second device relative to an additional pose of the first device.
[0177] Aspect 12. The method according to any one of aspects 1 to 11, wherein generating the first media content element includes a first projection of a first part of the application state; and generating the second media content element includes a second projection of a second part of the application state, the second part of the application state being at least partially different from the first part of the application state.
[0178] Aspect 13. The method according to any one of aspects 1 to 12, wherein the second media content element includes user interface elements.
[0179] Aspect 14. The method according to any one of aspects 1 to 13, further comprising: generating the first media content element and the second media content element by a renderer of the application engine.
[0180] Aspect 15. The method according to any one of aspects 1 to 14, wherein outputting the second media content element to the second device includes: transmitting the second media content element from the first device to the second device via a communication link.
[0181] Aspect 16. The method as in any one of Aspects 1 to 15, wherein: the application engine includes a video game application; the first media content element includes a first part of the game simulation of the video game application; the first device includes a handheld electronic device; the second media content element includes a second part of the game simulation of the video game application; and the second device includes a head-mounted display (HMD), wherein the second display of the HMD is configured to project the second media content element relative to the first display. Aspect 17. A method for displaying media on one or more displays, including: obtaining the application state of an application engine from a first device including a first display; generating a media content element associated with the application state of the application engine for a second device including a second display; and displaying the media content element on the second display of the second device relative to the pose of the first display of the first device.
[0182] Aspect 18. The method as in Aspect 17, wherein determining the pose of the first display of the first device includes at least one or more of the following: determining, by one or more sensors of the second device, one or more of the pose of the first device or the pose of the second device; or obtaining pose information associated with the first device from the first device.
[0183] Aspect 19. A device for displaying media on one or more displays, including: a memory; and one or more processors coupled to the memory and configured to: generate a first media content element associated with the application state of an application engine for the application engine; generate a second media content element associated with the application state of the application engine for the application engine; display the first media content element on the first display of the first device; and output the second media content element for display on the second display of the second device relative to the pose of the first display of the first device.
[0184] Aspect 20. The device as in Aspect 19, wherein in order to obtain pose information including at least one or more of first pose information associated with the first device or second pose information associated with the second device, the one or more processors are configured to; determine, based on the pose information, a display area for the second media content element on the second display outside the display area of the first display; and output the display area for the second media content element.
[0185] Aspect 21. The apparatus of any one of aspects 19 to 20, wherein, to obtain the first pose information, the one or more processors are configured to determine the pose of the first device by the first device based on data from one or more sensors of the first device; determine the pose of the first device by the first device based on data from one or more sensors of the second device; obtain the pose of the first device from the second device; or obtain the pose of the second device from the second device.
[0186] Aspect 22. The apparatus of any one of aspects 19 to 21, wherein the one or more processors are configured to determine a profile associated with a first display of the first device, wherein the profile associated with the first display includes a boundary of the profile, and wherein a display area for a second media content element extends from the first display of the first display.
[0187] Aspect 23. The apparatus of any one of aspects 19 to 22, wherein the display area for the second media content element extends from a boundary of a profile associated with a first display of the first device.
[0188] Aspect 24. The apparatus of any one of aspects 19 to 23, wherein the display area for the second media content element is configured to at least partially occlude the first display of the first device when being displayed on a second display of the second device.
[0189] Aspect 25. The apparatus of any one of aspects 19 to 24, wherein the second media content element includes: a stereoscopic projection for a second display of the second device.
[0190] Aspect 26. The apparatus of any one of aspects 19 to 25, wherein the application state includes a 3D representation of the environment.
[0191] Aspect 27. The apparatus of any one of aspects 19 to 26, wherein, to generate a first media content element, the one or more processors are configured to render a 3D representation of the environment in a first display configuration, wherein the first display configuration includes at least one or more of resolution, refresh rate, or color depth.
[0192] Aspect 28. The apparatus of any one of aspects 19 to 27, wherein, to generate a second media content element, the one or more processors are configured to render a 3D representation of the environment in a second display configuration, wherein the second display configuration is different from the first display configuration.
[0193] Aspect 29. The apparatus of any one of aspects 19 to 28, wherein the one or more processors are configured to: obtain a first input at the application engine; generate an additional application state different from the application state based on the first input; generate a third media content element for the application engine associated with the additional application state of the application engine; generate a fourth media content element for the application engine associated with the additional application state of the application engine; display the third media content element on a first display of the first device; and output the fourth media content element for display on a second display of a second device relative to an additional pose of the first device.
[0194] Aspect 30. The apparatus of any one of aspects 19 to 29, wherein the one or more processors are configured to: generate a first media content element, wherein the first media content element includes a first projection of a first part of the application state; and generate a second media content element, wherein the second media content element includes a second projection of a second part of the application state, the second part of the application state being at least partially different from the first part of the application state.
[0195] Aspect 31. The apparatus of any one of aspects 19 to 30, wherein the second media content element includes user interface elements.
[0196] Aspect 32. The apparatus of any one of aspects 19 to 31, wherein the one or more processors are configured to generate the first media content element and the second media content element by a renderer of the application engine.
[0197] Aspect 33. The apparatus of any one of aspects 19 to 32, wherein, in order to output the second media content element to the second device, the one or more processors are configured to: transmit the second media content element from the first device to the second device via a communication link.
[0198] Aspect 34. The apparatus of any one of aspects 19 to 33, wherein the application engine includes a video game application; the first media content element includes a first part of a game simulation of the video game application; the first device includes a handheld electronic device; the second media content element includes a second part of the game simulation of the video game application; and the second device includes a head-mounted display (HMD), wherein the second display of the HMD is configured to project the second media content element relative to the first display.
[0199] Aspect 35. An apparatus for displaying media on one or more displays, comprising: a memory; and one or more processors coupled to the memory and configured to: obtain an application state of an application engine from a first device including a first display; generate media content elements associated with the application state of the application engine for a second device including a second display; and display the media content elements on the second display of the second device relative to a pose of the first display of the first device.
[0200] Aspect 36. The apparatus of aspect 35, wherein, to determine a pose of the first display of the first device, the one or more processors are configured to: determine, by one or more sensors of the second device, one or more of a pose of the first device or a pose of the second device; or obtain pose information associated with the first device from the first device.
[0201] Aspect 37: A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform any operation of any one of aspects 1 to 36.
[0202] Aspect 38: A device comprising means for performing any operation of aspects 1 to 36.
[0203] Aspect 39: A method comprising operations according to any one of aspects 1 to 16 and any one of aspects 17 to 18.
[0204] Aspect 40: An apparatus for displaying media on one or more displays. The apparatus includes a configured memory (e.g., implemented in circuitry) and one or more processors (e.g., one processor or multiple processors) coupled to the memory. The one or more processors are configured to perform operations according to any one of aspects 1 to 16 and any one of aspects 17 to 18.
[0205] Aspect 41: A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations according to any one of aspects 1 to 16 and any one of aspects 17 to 18.
[0206] Aspect 42: A device comprising means for performing operations according to any one of aspects 1 to 16 and any one of aspects 17 to 18.
Claims
1. A method for displaying media on one or more displays, comprising: generating a first media content element for an application engine associated with an application state of the application engine; generating a second media content element for the application engine associated with the application state of the application engine; displaying the first media content element on a first display of a first device; and outputting the second media content element for display on a second display of a second device in a pose relative to the first display of the first device.
2. The method according to claim 1, further comprising: obtaining pose information including at least one or more of first pose information associated with the first device or second pose information associated with the second device; determining, based on the pose information, a display area for the second media content element on the second display outside a display area of the first display; and outputting the display area for the second media content element.
3. The method according to claim 2, wherein obtaining the first pose information includes at least one or more of the following: determining, by the first device, a pose of the first device based on data from one or more sensors of the first device; obtaining data from one or more sensors of the second device and determining, by the first device, a pose of the first device based on the data from the one or more sensors of the second device; obtaining the pose of the first device from the second device; or obtaining the pose of the second device from the second device.
4. The method according to claim 2, further comprising: determining a contour associated with the first display of the first device, wherein the contour associated with the first display includes a boundary of the contour, and wherein the display area for the second media content element extends from the first display of the first display.
5. The method according to claim 4, wherein the display area for the second media content element extends from the boundary of the contour associated with the first display of the first device.
6. The method according to claim 2, wherein the display area for the second media content element is configured to at least partially occlude the first display of the first device when displayed on the second display of the second device.
7. The method according to claim 1, wherein the second media content element comprises: a stereoscopic projection for the second display of the second device.
8. The method according to claim 1, wherein the application state includes a three-dimensional (3D) representation of an environment.
9. The method according to claim 8, wherein generating the first media content element comprises: rendering the 3D representation of the environment in a first display configuration, wherein the first display configuration includes at least one or more of resolution, refresh rate, or color depth.
10. The method according to claim 9, wherein generating the second media content element comprises: Render the 3D representation of the environment in a second display configuration, where the second display configuration is different from the first display configuration.
11. The method according to claim 1, further comprising Obtaining a first input at the application engine; Generating an additional application state different from the application state based on the first input; Generating a third media content element for the application engine associated with the additional application state of the application engine; Generating a fourth media content element for the application engine associated with the additional application state of the application engine; Displaying the third media content element on the first display of the first device; And Outputting the fourth media content element for display on the second display of the second device relative to an additional pose of the first device.
12. The method according to claim 1, Wherein: Generating the first media content element includes a first projection of a first part of the application state; and Generating the second media content element includes a second projection of a second part of the application state, the second part of the application state being at least partially different from the first part of the application state.
13. The method according to claim 1, wherein the second media content element includes user interface elements.
14. The method according to claim 1, further comprising generating the first media content element and the second media content element by a renderer of the application engine.
15. The method according to claim 14, wherein outputting the second media content element to the second device Includes: Transmitting the second media content element from the first device to the second device via a communication link.
16. A method for displaying media on one or more displays, Comprising: Obtaining an application state of an application engine from a first device including a first display; Generating a media content element associated with the application state of the application engine for a second device including a second display; And Displaying the media content element on the second display of the second device relative to a pose of the first display of the first device.
17. The method according to claim 16, wherein determining the pose of the first display of the first device includes at least one or more of the following: Determining at least one or more of the pose of the first device or the pose of the second device by one or more sensors of the second device; or Obtaining pose information associated with the first device from the first device.
18. An apparatus for displaying media on one or more displays, Comprising: A memory; And One or more processors coupled to the memory and configured to: Generate a first media content element for an application engine associated with an application state of the application engine; Generate a second media content element for the application engine associated with the application state of the application engine; Display the first media content element on a first display of a first device; And Output the second media content element for display on the second display of the second device relative to the pose of the first display of the first device.
19. The apparatus of claim 18, wherein, to obtain pose information including at least one or more of first pose information associated with the first device or second pose information associated with the second device, the one or more processors are configured to; Determine a display area for the second media content element on the second display outside the display area of the first display based on the pose information; And Output the display area for the second media content element.
20. The apparatus of claim 19, wherein, to obtain the first pose information, the one or more processors are configured to: Determine the pose of the first device by the first device based on data from one or more sensors of the first device; Determine the pose of the first device by the first device based on data from one or more sensors of the second device; Obtain the pose of the first device from the second device; Or Obtain the pose of the second device from the second device.
21. The apparatus of claim 19, wherein the one or more processors are configured to determine a contour associated with the first display of the first device, wherein the contour associated with the first display includes a boundary of the contour, and wherein the display area for the second media content element extends from the first display of the first display.
22. The apparatus of claim 21, wherein the display area for the second media content element extends from the boundary of the contour associated with the first display of the first device.
23. The apparatus of claim 19, wherein the display area for the second media content element is configured to at least partially occlude the first display of the first device when displayed on the second display of the second device.
24. The apparatus of claim 18, wherein the application state includes a 3D representation of the environment.
25. The apparatus of claim 24, wherein, to generate the first media content element, the one or more processors are configured to: render the 3D representation of the environment in a first display configuration, wherein the first display configuration includes at least one or more of resolution, refresh rate, or color depth.
26. The apparatus of claim 25, wherein, to generate the second media content element, the one or more processors are configured to: render the 3D representation of the environment in a second display configuration, wherein the second display configuration is different from the first display configuration.
27. The apparatus of claim 18, wherein the one or more processors are configured to: Obtain a first input at the application engine; Generate an additional application state different from the application state based on the first input; Generate a third media content element for the application engine associated with the additional application state of the application engine; Generate a fourth media content element for the application engine associated with the additional application state of the application engine; Display the third media content element on the first display of the first device; And Output the fourth media content element for display on the second display of the second device relative to an additional pose of the first device.
28. The apparatus of claim 18, wherein the one or more processors are configured to: Generate the first media content element, wherein generating the first media content element, wherein the first media content element includes a first projection of a first portion of the application state; and Generate the second media content element, wherein generating the second media content element, wherein the second media content element includes a second projection of a second portion of the application state, the second portion of the application state being at least partially different from the first portion of the application state.
29. An apparatus for displaying media on one or more displays, Comprising: A memory; And One or more processors, the one or more processors being coupled to the memory and configured to: Obtain an application state of an application engine from a first device including a first display; Generate a media content element associated with the application state of the application engine for a second device including a second display; And Display the media content element on the second display of the second device relative to a pose of the first display of the first device.
30. The apparatus of claim 29, wherein to determine the pose of the first display of the first device, the one or more processors are configured to: Determine at least one or more of the pose of the first device or the pose of the second device by one or more sensors of the second device; or Obtain pose information associated with the first device from the first device.