Local environmental input sensing for electronic devices
By using a light estimation process that combines local ambient light sensors and saliency information in electronic devices, the problem of inaccurate local ambient light estimation in extended reality environments is solved, improving the accuracy of environmental condition estimation and user experience.
Patent Information
- Application Number
- CN202480030691.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2024-05-03
- Publication Date
- 2025-12-05
AI Technical Summary
In the prior art, when electronic devices generate extended reality environments, the estimation of local ambient lighting conditions is not accurate enough, resulting in visual and audio output artifacts and a degraded user experience.
A method combining local ambient light sensors and saliency information is adopted to estimate local lighting conditions through a light estimation process, generate a saliency map, and guide subsequent processing chains to improve the accuracy of environmental condition estimation.
It improves the estimation accuracy of local ambient lighting in extended reality environments, reduces visual and audio output artifacts, and enhances the user experience.
Smart Images

Figure CN121079656A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 466,224, filed May 12, 2023, entitled “Localized Environmental Input Sensing for Electronic Devices,” the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This specification relates in its entirety to electronic devices, including, for example, to local environmental input sensing for electronic devices. Background Technology
[0004] Electronic devices are typically equipped with cameras for capturing images. Images captured by the camera can be displayed, stored in the electronic device's memory, transmitted to other electronic devices, and / or used to detect objects in those images. Some electronic devices include ambient light sensors that sense the total amount of ambient light in the physical environment of the electronic device. Attached Figure Description
[0005] Specific features of this subject matter are set forth in the appended claims. However, for purposes of explanation, several embodiments of this subject matter are illustrated in the following figures.
[0006] Figure 1 Example system architectures are illustrated based on one or more specific implementations of various electronic devices that can implement the system of this subject.
[0007] Figure 2 A block diagram illustrating example features of an electronic device according to one or more specific implementations is shown.
[0008] Figure 3 A schematic diagram illustrating a processing chain that can be estimated using environmental conditions according to one or more specific implementations is shown.
[0009] Figure 4 Examples of physical environments, including various environmental conditions, are illustrated according to one or more specific implementations.
[0010] Figure 5 An example of saliency-based environmental condition estimation based on one or more specific implementations is illustrated.
[0011] Figure 6 Example saliency diagrams based on one or more specific implementations are shown.
[0012] Figure 7Another example saliency map in accordance with one or more implementations is illustrated.
[0013] Figure 8 A flowchart of example operations for performing local environmental input sensing for an electronic device in accordance with one or more implementations is illustrated.
[0014] Figure 9 A flowchart of example operations for performing local environmental input sensing in a salient region of a physical environment in accordance with one or more implementations is illustrated.
[0015] Figure 10 An electronic system that can be used for implementing one or more implementations of the subject technology is illustrated. DETAILED DESCRIPTION
[0016] The detailed description set forth below is intended as a description of various configurations of the subject technology and is not intended to represent the only configurations in which the subject technology can be practiced. The appended drawings are incorporated herein and constitute a part of the detailed description. The detailed description includes specific details for the purpose of providing a thorough understanding of the subject technology. However, the subject technology is not limited to the specific details set forth herein and can be practiced with one or more other implementations. In one or more implementations, structures and components are shown in block diagram form in order to avoid obscuring the concepts of the subject technology.
[0017] A physical environment refers to the physical world that people are able to sense and / or interact with without the aid of electronic devices. A physical environment can include physical features such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People are able to directly sense and / or interact with a physical environment such as through sight, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, and / or virtual reality (VR) content, among others. In the case of an XR system, a subset of a person’s physical motions, or representations thereof, are tracked, and, in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that complies with at least one physical law. As one example, an XR system can detect head movements and, in response, adjust graphical content and a sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. As another example, an XR system can detect movements of an electronic device (e.g., a mobile phone, a tablet computer, a laptop computer, etc.) presenting an XR environment and, in response, adjust graphical content and a sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), an XR system can adjust characteristics of graphical content in an XR environment in response to representations of physical motions (e.g., voice commands).
[0018] There are many different types of electronic systems that enable a person to sense various XR environments and / or to interact with various XR environments. Examples include head-mountable systems, projection-based systems, heads-up displays (HUDs), vehicle windshields having integrated display capability, windows having integrated display capability, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mountable system can have one or more speakers and an integrated opaque display. Alternatively, a head-mountable system can be configured to accept an external opaque display (e.g., a smartphone). A head-mountable system can incorporate one or more imaging sensors to capture images or video of a physical environment and / or one or more microphones to capture audio of the physical environment. A head-mountable system can have a transparent or translucent display instead of an opaque display. A transparent or translucent display can have a medium through which light representative of images is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, hologram medium, optical combiner, optical reflector, or any combination thereof. In some implementations, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology that projects graphical images onto a person's retinas. Projection systems can also be configured to project virtual objects into a physical environment, for example, as a hologram or on a physical surface.
[0019] Aspects of the subject technology can provide local environmental input information that can be used by various processing stages of various processing pipelines of an electronic device, such as an electronic device that generates a three-dimensional scene (e.g., an XR environment) and / or performs spatial computing operations. For example, the local environmental input information can include local environmental conditions, such as local lighting conditions in a region of a physical environment that is within a field of view of a three-dimensional scene generated by the electronic device and is smaller than the field of view. The local lighting conditions can include local ambient light levels and / or lighting directions in respective local portions of the physical environment. These local lighting conditions can be provided to various processing stages of various processing pipelines for computer vision, hand tracking, three-dimensional immersion effect generation, spatial computing, image pre-processing, image and / or video processing (e.g., including spatial and / or immersive video capture, generation, and / or processing), spatial mapping, surface and / or texture estimation, shape and / or geometry estimation, material estimation, object tracking, scene reconstruction, and / or spatial audio processes, for example, that can otherwise be negatively impacted by global lighting condition estimates that are insufficiently detailed to represent the lighting environment. For example, these local lighting conditions can allow various processing pipelines to properly account for brightness discontinuities, directional lighting features, and / or other discontinuities in the physical environment that would not be reflected in global lighting condition estimates.
[0020] Figure 1 An example system architecture 100 including various electronic devices that can implement various electronic devices of the subject systems is illustrated in accordance with one or more implementations. However, not all of the depicted components can be used in all implementations, and one or more implementations can include additional or different components than those shown in the figure. Variations in the arrangement and type of the components can be made without departing from the spirit or scope of the claims as set forth herein. Additional components, different components, or fewer components can be provided.
[0021] The system architecture 100 includes an electronic device 105, a handheld electronic device 104, an electronic device 110, an electronic device 115, and a server 120. For explanatory purposes, the system architecture 100 is illustrated in Figure 1 as including the electronic device 105, the handheld electronic device 104, the electronic device 110, the electronic device 115, and the server 120; however, the system architecture 100 can include any number of electronic devices and any number of servers or data centers including multiple servers.
[0022] The electronic device 105 may, for example, be implemented as a tablet device, a smartphone, or a portable system that is wearable (e.g., by the user 101). The electronic device 105 includes a display system capable of presenting a visualization of an extended reality environment to a user. The electronic device 105 can be powered with a battery and / or another power source. In one example, the display system of the electronic device 105 provides a stereoscopic presentation of an extended reality environment to a user, enabling a three-dimensional visual display of a particular scene rendering. In one or more implementations, instead of or in addition to utilizing the electronic device 105 to access an extended reality environment, a user can use a handheld electronic device 104, such as a tablet, a watch, a mobile device, etc.
[0023] The electronic device 105 can include one or more cameras, such as a camera 150 (e.g., a visible light camera, an infrared camera, etc.). For example, the electronic device 105 can include multiple cameras 150. For example, the multiple cameras 150 can include a left-facing camera, a front-facing camera, a right-facing camera, a downward-facing camera, a left-downward-facing camera, a right-downward-facing camera, an upward-facing camera, one or more eye-to-eye cameras, and / or other cameras. Each of the cameras 150 can include one or more image sensors (e.g., a charge-coupled device (CCD) image sensor, a complementary metal-oxide-semiconductor (CMOS) image sensor, etc.).
[0024] In addition, the electronic device 105 can include various sensors 152, including but not limited to other cameras, other image sensors, touch sensors, ambient light sensors, microphones (e.g., sound level microphones), inertial measurement units (IMUs), heart rate sensors, temperature sensors, depth sensors (e.g., lidar sensors, radar sensors, sonar sensors, time-of-flight sensors, etc.), GPS sensors, Wi-Fi sensors, near-field communication sensors, radio frequency sensors, etc. In addition, the electronic device 105 can include hardware elements that can receive user input, such as hardware buttons or switches. User input detected by such cameras, sensors, and / or hardware elements can correspond to, for example, various input modalities. For example, such input modalities can include, but are not limited to, facial tracking, eye tracking (e.g., gaze direction), hand tracking, gesture tracking, biometric readings (e.g., heart rate, pulse, pupil dilation, respiration, temperature, electroencephalogram, olfaction), recognizing speech or audio (e.g., particular hot words), and activating buttons or switches, etc. In one or more implementations, facial tracking, gaze tracking, hand tracking, gesture tracking, object tracking, and / or physical environment mapping processes (e.g., system processes and / or application processes) can utilize images (e.g., image frames) captured by one or more image sensors of the cameras 150 and / or the sensors 152.
[0025] In one or more implementations, the electronic device 105 can be communicatively coupled to a base device, such as the electronic device 110 and / or the electronic device 115. Generally, such a base device can include more computational resources and / or available power than the electronic device 105. In one example, the electronic device 105 can operate in various modes. For example, the electronic device 105 can operate in an independent mode independent of any base device. When the electronic device 105 operates in the independent mode, the number of input modalities can be constrained by power and / or processing limitations of the electronic device 105, such as available battery power of the device. In response to the power limitations, the electronic device 105 can deactivate certain sensors within the device itself to preserve battery power and / or free up processing resources.
[0026] The electronic device 105 can also operate in a tethered mode (e.g., connected to a base device via a wireless connection) to work in conjunction with a given base device. The electronic device 105 can also operate in a connected mode in which the electronic device 105 is physically connected to a base device (e.g., via a cable or some other physical connector) and can utilize power resources provided by the base device (e.g., in cases in which the base device charges and / or powers the electronic device 105 while physically connected).
[0027] When the electronic device 105 operates in the tethered or connected mode, processing user inputs and / or rendering at least a portion of an extended reality environment can be offloaded to the base device, thereby reducing the processing burden on the electronic device 105. For example, in one implementation, the electronic device 105 works in conjunction with the electronic device 110 or the electronic device 115 to generate an extended reality environment that includes physical and / or virtual objects that enable different forms of interaction (e.g., visual, auditory, and / or physical or haptic interaction) between a user and the generated extended reality environment in a real-time manner. In one example, the electronic device 105 provides a rendering of a scene corresponding to the extended reality environment that can be perceived by a user and interacted with in a real-time manner, such as a host environment for a group session with another user. Additionally, as part of presenting the rendered scene, the electronic device 105 can provide sound and / or haptic or tactile feedback to the user. The content of the given rendered scene can depend on available processing power, network availability and capacity, available battery power, and current system workloads. The electronic device 105 can be and / or can include all or part of the electronic system discussed below with respect to Figure 10 .
[0028] Network 106 may communicatively (directly or indirectly) couple electronic devices 105, 110, and / or 115 to each other device and / or server 120. In one or more embodiments, network 106 may be an interconnection network of devices that may include the Internet or be communicatively coupled to the Internet.
[0029] Electronic device 110 may include one or more cameras 150 (e.g., multiple cameras 150) and may be, for example, a smartphone, a portable computing device (such as a laptop computer), an accessory (e.g., a digital camera, headphones), a tablet device, a wearable device (such as a watch, wristband, etc.), or any other suitable device including, for example, one or more speakers 211, a touchscreen, and / or a touchpad. In one or more embodiments, electronic device 110 may not include a touchscreen but may support touchscreen-like gestures, such as in extended reality environments. In one or more embodiments, electronic device 110 may include a touchpad. Figure 1 In this embodiment, by way of example, electronic device 110 is described as a mobile smartphone device. In one or more specific embodiments, electronic device 110, handheld electronic device 104, and / or electronic device 105 may be and / or may include, as described below with respect to... Figure 10 The electronic system under discussion may be all or part of it. In one or more embodiments, electronic device 110 may be another device, such as an Internet Protocol (IP) camera, a tablet computer, or an accessory such as an electronic stylus.
[0030] Electronic device 115 may be, for example, a desktop computer, portable computing devices such as laptops, smartphones, accessories (e.g., digital cameras, headphones), tablet devices, wearable devices such as watches, wristbands, etc. Figure 1 For example, electronic device 115 is described as a desktop computer having one or more cameras 150 (e.g., multiple cameras 150). Electronic device 115 may be and / or may include, as described below, relative to... Figure 10 The electronic system under discussion, in whole or in part.
[0031] Server 120 may form all or part of a computer network or server group 130, such as in a cloud computing or data center implementation. For example, server 120 stores data and software and includes specific hardware (e.g., processors, graphics processors, and other dedicated or custom processors) for rendering and generating extended reality content such as graphics, images, videos, audio, and multimedia files. In one implementation, server 120 may function as a cloud storage server that stores any of the aforementioned extended reality content generated by the aforementioned devices and / or server 120.
[0032] Figure 2 Block diagrams illustrating various components that may be included in electronic device 105 according to various aspects of this disclosure are shown. Figure 2 As shown, electronic device 105 may include one or more cameras, such as camera 150 (e.g., multiple cameras 150, each including one or more image sensors 215), for capturing images and / or video of the physical environment surrounding the electronic device 105, and one or more sensors 152 for obtaining environmental information (e.g., depth information) associated with the physical environment surrounding the electronic device 105. Sensor 152 may include depth sensors (e.g., time-of-flight sensors, infrared sensors, radar, sonar, lidar, etc.), one or more microphones (e.g., including sound level microphones), light sensors (such as ambient light sensors), and / or other types of sensors for sensing various properties of the physical environment (such as one or more colors, one or more brightness levels, one or more depths, one or more sound levels, etc.). For example, one or more microphones included in sensor 152 may be operable to capture audio input from a user of electronic device 105 (such as voice input corresponding to the user speaking into the microphone) and / or determine the ambient sound level in one or more parts of the physical environment. Figure 2 In some implementations, electronic device 105 further includes communication circuitry 208 for communicating with electronic device 110, electronic device 115, server 120, and / or other devices and / or systems. Communication circuitry 208 may include radio frequency (RF) communication circuitry for detecting radio frequency identification (RFID) tags, Bluetooth Low Energy (BLE) communication circuitry, other near field communication (NFC) circuitry, WiFi communication circuitry, cellular communication circuitry, and / or other wired and / or wireless communication circuitry.
[0033] As shown, the electronic device 105 includes processing circuitry 204 (e.g., one or more processors and / or integrated circuits) and memory 206. The memory 206 can store (e.g., temporarily or permanently) content generated and / or otherwise obtained by the electronic device 105. In some operational scenarios, the memory 206 can temporarily store images of a physical environment captured by the camera 150, depth information corresponding to images generated using, for example, a depth sensor in the sensors 152, a mesh and / or texture corresponding to the physical environment, virtual objects such as virtual objects generated by the processing circuitry 204 to include virtual content, and / or virtual depth information for the virtual objects. The memory 206 can store (e.g., temporarily or permanently) intermediate images and / or information generated by the processing circuitry 204 for combining images of the physical environment and virtual objects and / or virtual images to form (e.g., composite) images for display by the display 200 such as by compositing one or more virtual objects onto a pass-through video stream obtained from one or more of the cameras 150.
[0034] As shown, the electronic device 105 can include one or more speakers 211. The speakers can be operable to output audio content including audio content stored and / or generated at the electronic device 105, and / or audio content received from a remote device or server via the communication circuitry 208.
[0035] The memory 206 can store instructions or code for execution by the processing circuitry 204 such as, for example, operating system code corresponding to an operating system installed on the electronic device 105 and application code corresponding to one or more application programs installed on the electronic device 105. The operating system code and / or application code, when executed, can correspond to one or more operating system-level processes and / or application-level processes such as processes that support capturing images, obtaining and / or processing environmental condition information, and / or determining inputs to and / or outputs from the electronic device 105 (e.g., display content on the display 200).
[0036] Figure 3 Aspects of various processing chains that can utilize environmental condition estimates such as lightning condition estimates are illustrated. As Figure 3 As illustrated, multiple cameras and / or sensors can generate images and / or sensor data. In Figure 3In the example, a first camera / sensor 301 (e.g., camera 150 or sensor 152), a second camera / sensor 303 (e.g., camera 150 or sensor 152), and a third camera / sensor 305 (e.g., camera 150 or sensor 152) can generate sensor data (e.g., visible spectrum images, infrared (IR) images, depth information, and / or other sensor data). The first camera / sensor 301, second camera / sensor 303, third camera / sensor 305, and / or other cameras / sensors of the electronic device 105 can be configured differently from each other (e.g., with different color spaces, quantization, exposure times, sampling frequencies) and / or configured for different XR and / or spatial calculation outputs generated by the electronic device 105. Furthermore, as... Figure 3 As shown, the outputs of the first camera / sensor 301, the second camera / sensor 303, and the third camera / sensor 305 can be fed to any software and / or hardware process in one or more processing chains.
[0037] exist Figure 3 In the example, the output of the first camera / sensor 301 is provided to the light estimation process 302. The light estimation process 302 can determine one or more local lighting conditions (e.g., ambient light level, light direction, etc.) and / or one or more other local environmental conditions (e.g., color, texture, etc.) of one or more corresponding portions of the physical environment of the electronic device 105, as discussed in further detail below. Local lighting conditions and / or other local environmental conditions may be local with respect to a spatial portion of the physical environment and / or local with respect to a specific observation point or nearby time. In one or more embodiments, local lighting conditions and / or other local environmental conditions may be smoothed, filtered, or averaged over time. In one or more embodiments, local lighting conditions and / or other local environmental conditions may be stored for various different observation times (e.g., with various different lighting conditions and / or various different viewing angles).
[0038] In one or more embodiments, the light estimation process 302 may perform local light estimation in part based on saliency information indicating that one or more parts of the physical environment are salient to the user of the electronic device 105. This saliency information may include the position of the user's hand, the position of the user's gaze, and / or objects detected and / or identified in the physical environment, or may be derived from the foregoing (e.g., through the light estimation process or through another process at the electronic device), as discussed in further detail below. As described in further detail below, this saliency information may be provided as one or more saliency maps. These saliency maps may be provided as input to the light estimation process 302, or may be generated by the light estimation process 302 based on other saliency information (e.g., user actions and / or gaze information). In one or more embodiments, local lighting conditions and / or other local environmental conditions may be recursively determined at multiple resolutions and saliency levels. In one or more specific implementations, multiple local lighting conditions (e.g., ambient light level, light color, and / or light direction) and / or multiple other local environmental conditions (e.g., color, texture, and / or other characteristics of a physical object) may be obtained for each of one or more local portions of the physical environment.
[0039] like Figure 3 As shown, the output of the light estimation process 302 (e.g., one or more local illumination conditions) may be provided to the image preprocessing operation 304. For example, the image preprocessing operation 304 may include noise reduction, signal-to-noise ratio improvement, image enhancement, color transformation, compression, and / or refinement of image features for subsequent sensor-based processes. As shown, the output of the image preprocessing operation 304 may be provided to the sensor-based process 306. For example, the sensor-based process 306 may include a computer vision process that performs object detection, object recognition, and / or object tracking of one or more physical objects in the physical environment of the electronic device 105. In one or more embodiments, the sensor-based process 306 may also or alternatively include material estimation (e.g., surface texture estimation).
[0040] As shown, the output of the sensor-based process 306 can be provided to one or more other sensor-based processes (e.g., sensor-based process 308). For example, the sensor-based process 308 can include a gesture detection process that utilizes the hand tracking output of the sensor-based process 306. As another example, the sensor-based process 308 can be a scene reconstruction process that utilizes one or more meshes, textures, and / or other three-dimensional environmental information generated by the sensor-based process 306 to generate a three-dimensional reconstruction of the physical environment. As another example, the sensor-based process 308 can include a three-dimensional immersion effect generation process that operates based on sensor data and / or lighting condition estimates to generate immersive effects for a three-dimensional scene. As another example, the sensor-based process 308 and / or another sensor-based process subsequent to the sensor-based process 308 can use the local lighting conditions estimated by the light estimation process 302, the light estimation process 310, and / or the light estimation process 314, as well as sensor data from the camera sensor 301, the camera / sensor 303, and / or the camera / sensor 305 (e.g., one or more images from one or more cameras 150) to generate a three-dimensional scene (e.g., an XR scene) for display by the electronic device that includes a view of a region of the physical environment based on the one or more images from the one or more cameras, as well as virtual content overlaid on the view of the region of the physical environment.
[0041] As another example, the sensor-based process 308 can include a spatial audio generation process that utilizes the three-dimensional geometry and / or surface textures derived by the sensor-based process 306. For example, the output of the sensor-based process 306 (e.g., in examples where the sensor-based process 306 is implemented as a vision algorithm, such as a scene reconstruction algorithm or computer vision algorithm that can identify walls, tables, chairs, fabric, wood, glass, or other materials) can be used by the sensor-based process 308 to generate spatial audio that is provided directly from speakers to the ears of a user, but apparently in a manner that sounds that are propagated through the physical environment itself would be reflected and / or absorbed by the physical features of the physical environment. In one or more implementations, the sensor-based process 306 can convert identified environmental aspects, such as detected objects, textures, and / or materials, into a mesh that can be fed into the sensor-based process 308 to perform spatial audio rendering. In one or more implementations, the spatial audio rendering can include determining reverberation and / or other spatial characteristics of audio output based on calculations of reflections or absorption of environmental conditions identified in the mesh.
[0042] As another example, sensor-based process 306 and / or sensor-based process 308 can be or include a spatial computing process. For example, a spatial computing process can be performed for detecting and / or facilitating user interaction with virtual objects and / or real objects (e.g., in an XR environment such as an AR or MR environment) via an electronic device. Spatial computing can be performed, for example, for creating memories such as spatial images, spatial video, and / or spatial audio recordings, playing spatial games, etc. Spatial computing can include computing processes associated with XR, VR, AR, and / or MR experiences, machine learning, neural networks, and / or artificial intelligence, and / or any computing process involving user input to and / or user output from an electronic device that obtains, stores, and / or processes spatial information associated with physical objects and / or environments. As an example, a spatial computing process can include a user interaction process and / or a memory creation process (e.g., spatial imaging, spatial video recording, and / or immersive video and / or other immersive experiences).
[0043] As shown, image pre-processing operation 304, sensor-based process 306, and / or sensor-based process 308 can each receive output of an earlier (e.g., previous) processing stage in a processing chain that includes light estimation process 302, image pre-processing operation 304, sensor-based process 306, and sensor-based process 308. As shown, image pre-processing operation 304, sensor-based process 306, and / or sensor-based process 308 can also each optionally receive output of camera / sensor 301 and / or light estimation process 302.
[0044] Figure 3 It is also shown how the processing chain including light estimation process 302, image pre-processing operation 304, sensor-based process 306, and sensor-based process 308 can be integrated with one or more other processing chains performed by electronic device 105. For example, sensor-based process 306 can operate on output of image pre-processing operation 304 as well as on output of image pre-processing operation 312, which receives output of light estimation process 310 in a parallel processing chain that uses output of camera / sensor 303. As shown, sensor-based process 308 can operate on output of sensor-based process 306 as well as on output of image pre-processing operation 316, which receives output of light estimation process 314 in yet another parallel processing chain that uses output of camera / sensor 305.
[0045] Figure 3The example inter-connected processing chain illustrates how environmental condition estimation (such as illumination condition estimation (e.g., ambient light level, directional lighting estimation, color estimation, texture estimation, and / or material estimation, etc.)) in the first stage or early stages of a processing chain can influence that processing chain and / or ultimately result in multiple subsequent processing stages in other processing chains that determine user input and / or generate output from the electronic device. For example, estimation of directional light sources and / or ambient light levels can play a major role in inferring other properties of the physical environment in XR and / or spatial computing experiences. Thus, inaccurate environmental condition estimation (such as inaccurate illumination condition estimation) can result in propagated and / or accumulated errors through the processing chain, which can result in visual and / or audio output artifacts and degraded user experience.
[0046] As described in further detail below, aspects of the subject technology provide a robust environmental condition estimator (e.g., light estimation processes 302, 310, and / or 314) that can handle spatially and temporally variable lighting environments and / or other spatially and temporally variable environmental conditions (e.g., color, texture, sound, etc.). In one or more implementations, the light estimation processes described herein take into account saliency. For example, light estimation processes 302, 310, and / or 314 can use saliency information to estimate local lighting conditions and / or other environmental conditions in particular portions of the physical environment. For example, light estimation processes 302, 310, and / or 314 can construct environmental condition estimates based on a saliency map (e.g., derived from user gaze information, user history information, user gestures, environmental sensing algorithms, detected objects, identified objects, and / or user annotations) that guides the sampling strategy used by these light estimation processes. In various implementations, light estimation processes 302, 310, and / or 314 can generate local environmental condition estimates, such as local illumination condition estimates, that take into account: (a) local discontinuities in lighting and geometry, (b) temporal variations in the physical environment (e.g., due to device and / or user motion), and / or (c) prior information about the environment. In one or more implementations, light estimation processes 302, 310, and / or 314 can use a combination of local probability maps (e.g., saliency probability maps), geometric information, and / or temporal variance.
[0047] In one or more implementations, sensor-based process 306, sensor-based process 308, and / or other processes and / or processing chains that can be performed based on sensor information (e.g., including images from one or more cameras 150) and based on local environmental conditions (e.g., based on local lighting conditions) can be implemented as a neural network. In one or more implementations, the light estimation processes disclosed herein can be implemented as a pre-processing to a neural network (e.g., as one or more initial layers) that includes sensor-based process 306, sensor-based process 308, and / or other processes and / or processing chains that can be performed based on sensor information and based on local environmental conditions. In one or more implementations, sensor-based process 306, sensor-based process 308, and / or other processes and / or processing chains that can be performed based on sensor information (e.g., including images from one or more cameras 150) and based on local environmental conditions (e.g., based on local lighting conditions) can be encapsulated in one or more neural networks that can be trained and / or learned in context (e.g., during user operation of electronic device 105). For example, ongoing training of a neural network in context can allow various sensor-based processes and / or other processes and / or processing chains to be adaptable in different environments (e.g., through user feedback or action guidance to adjust saliency).
[0048] Figure 3 The example includes three cameras / sensors, three light estimation processes, three image pre-processors, and two sensor-based processes. However, this is merely illustrative, and in one or more other implementations, electronic device 105 can implement more or fewer processing chains than depicted, which can obtain sensor information from more or fewer than three cameras / sensors, and which can incorporate more or fewer than three light estimators, more or fewer than three image pre-processors, and / or more or fewer than two sensor-based processes. It should also be understood that Figure 3 Figure 4 Sensor-based processes 306 and sensor-based processes 308 can operate based on sensor data received directly from one or more cameras and / or other sensors, and / or can operate on data that has been derived from sensor data from one or more cameras and / or other sensors. It will also be appreciated that processing chains used by an electronic device, such as electronic device 105, can be modified and / or changed to generate varying XR experiences. For example, various different processing stages of one or more processing chains can be linked together across modalities and / or application-specific operations, such as by “wiring together” any of various sensors (e.g., image sensors, video sensors, audio sensors, and / or motion sensors such as IMU sensors), hardware pre-processors (e.g., for image enhancement and / or signal processing), computer vision algorithms (e.g., for power capture and / or playback), 3D immersive processes, and / or spatial audio processes in various arrangements.
[0049] Figure 4 An example physical environment 400 in which an electronic device, such as electronic device 105, can operate is illustrated. In the example of FIG. 4, the physical environment 400 includes a physical wall 402, a physical floor 404, a physical table 406, a physical lamp 408, and a physical mirror 410. In other examples, the physical environment 400 can include any other physical objects, which can include more or fewer physical objects than those depicted. Figure 4 Figure 4 The physical environment 400 is depicted as including a physical table 406 having a surface 412. In the example of FIG. 4, the surface 412 is depicted as including a physical cup 420. In other examples, the surface 412 can include any other physical objects, which can include more or fewer physical objects than those depicted. Figure 4 An example field of view 401 corresponding to one or more of the cameras 150 of the electronic device 105 is also illustrated. As an example, the field of view 401 can be a field of view of a single camera 150 of the electronic device 105, or a combined field of view of multiple cameras 150 of the electronic device 105. In one or more implementations, the field of view 401 can represent a field of view of a pass-through video stream displayed by a display (e.g., display 200) of the electronic device 105 to represent a corresponding portion of the physical environment.
[0050] As discussed herein, in various use cases (e.g., XR, spatial computing), it can be helpful to have an estimate of one or more environmental conditions (e.g., lighting conditions, colors, textures, etc.) of the physical environment 400. For example, in one or more use cases, the electronic device 105 can render virtual content (e.g., using a display 200 of the device) that is overlaid on a view of the physical environment 400 such that the virtual content appears to be located at a location within the physical environment 400 that is remote from the electronic device 105. In such examples, it can be helpful to have an estimate of one or more environmental conditions of the physical environment 400 to accurately render the virtual content. Figure 4 In the example of FIG. 4, a virtual cup 420 has been rendered by the electronic device 105 to appear to a viewer of a display (e.g., display 200) of the electronic device 105 to be on the surface 412 of the physical table 406.
[0051] For example, to render the virtual cup 420 on the surface 412 of the physical table 406, it can be helpful to have an estimate of the lighting conditions in the physical environment 400. For example, the electronic device 105 can render the virtual cup 420 with a brightness that depends on the brightness of the ambient light in the physical environment 400. For example, if the ambient light in the physical environment is relatively low, the electronic device 105 can render the virtual cup 420 with a relatively low brightness, or if the ambient light in the physical environment 400 is relatively high, the electronic device can render the virtual cup with a relatively high brightness. As another example, the electronic device 105 can modify the color of the virtual cup 420 based on the color of the ambient light in the physical environment 400 and / or based on the color of the surface 412 and / or the physical wall 402.
[0052] In some systems, a global lighting condition, such as a global ambient light level, can be obtained by the electronic device by averaging the brightness across the entire field of view 401 or by obtaining an ambient light level with a single-pixel ambient light sensor. However, as discussed herein, this type of global ambient light level can be misleading in representing the brightness of some or all of the physical environment 400, especially in use cases in which, for example, the physical environment 400 includes a very bright region or a very dark region. For example, the physical environment 400 can include one or more brightness discontinuities between bright regions and dark regions of the physical environment.
[0053] In one or more use cases, the physical environment 400 can include a directional light source, such as the physical lamp 408 (e.g., or another directional light source, such as a wall or ceiling mounted light source, or a window through which sunlight or streetlight light enters the physical environment 400). As Figure 4 As illustrated, the physical lamp 408 can generate directional light 415. The directional light 415 can interact with the surface 412 of the physical table 406 to generate a shadow 414 on the physical floor 404. As an example, this can generate a bright region corresponding to the surface 412 of the physical table, a bright region corresponding to the physical lamp 408 itself, and a dark region corresponding to the shadow 414, all of which are within the field of view 401. In Figure 4 In the example physical environment 400, the physical mirror 410 can also reflect light from the physical lamp 408, generating another bright region in the field of view 401.
[0054] In the example in which the virtual cup 420 is rendered on the surface 412 (within the bright region of the physical environment 400 generated by the directional light 415), a global estimate of the light level in the physical environment 400 can result in difficulty viewing the rendering of the virtual cup 420 in the bright region corresponding to the surface 412. In the example in which the virtual cup 420 is rendered on the physical floor 404 (outside of the bright region of the physical environment 400 generated by the directional light 415), a global estimate of the light level in the physical environment 400 can result in difficulty viewing the rendering of the virtual cup 420 in the dark region corresponding to the shadow 414. Figure 4 In the example in which the virtual cup 420 is rendered on the surface 412 (within the bright region of the physical environment 400 generated by the directional light 415), a global estimate of the light level in the physical environment 400 can result in difficulty viewing the rendering of the virtual cup 420 in the bright region corresponding to the surface 412. In the example in which the virtual cup 420 is rendered on the physical floor 404 (outside of the bright region of the physical environment 400 generated by the directional light 415), a global estimate of the light level in the physical environment 400 can result in difficulty viewing the rendering of the virtual cup 420 in the dark region corresponding to the shadow 414.Figure 4 In the example of FIG. 4A, if the color of surface 412 is different from the global color estimate (e.g., if the color of virtual cup 420 is determined to be a color that contrasts with the global color estimate), the global estimate of the color of physical environment 400 can also result in a rendering of virtual cup 420 that is difficult to view.
[0055] Physical environment 400 can include light discontinuities, such as brightness discontinuities, color discontinuities, texture discontinuities, or other visual discontinuities. For example, as shown in field of view 401, physical environment 400 can include a light discontinuity, such as brightness discontinuity 418 between shadow 414 and illuminated portion 416 of physical floor 404. As another example, as shown in field of view 401, physical environment 400 can include a light discontinuity, such as brightness discontinuity 421 between brightly illuminated surface 412 of physical table 406 and a side surface 419 of physical table 406 that is in the shadow of directional light 415. As another example, as shown in field of view 401, physical environment 400 can include a light discontinuity, such as brightness discontinuity 423 between a surface of physical wall 402 and a surface of physical mirror 410. Figure 5
[0056] In one or more implementations, electronic device 105 can perform operations (e.g., sensor-based process 306 or sensor-based process 308) to obtain information (e.g., user interactions) from or render virtual content into portions of physical environment 400 on one side or the other of brightness discontinuity 418 and / or on or above brightness discontinuity 418. For example, a user of electronic device 105 can perform a hand gesture while the user’s hand is partially within shadow 414 and partially within a portion of physical environment 400 that is directly illuminated by physical light 408, resulting in a brightness discontinuity across the user’s hand while the hand gesture is performed. In one or more implementations, electronic device 105 can perform operations that depend on the direction of light in physical environment 400. For example, rendering virtual cup 420 can include rendering virtual cup 420 and a virtual shadow of virtual cup 420, where the positioning, size, and / or shape of the virtual shadow depends on the direction of light at the apparent location of virtual cup 420 in physical environment 400.
[0057] For example, to provide the electronic device 105 with the ability to obtain information (e.g., object tracking information, gesture information, acoustic environment information, etc.) from portions of the physical environment 400 on one side or the other of a luminance discontinuity and / or above or on a luminance discontinuity 418, the ability to render virtual content over portions of the physical environment 400 on one side or the other of a luminance discontinuity and / or above or on a luminance discontinuity 418, the ability to perform operations dependent on the direction of light in the physical environment, and / or the ability to perform operations dependent on the color, texture, or material in the physical environment 400, the electronic device 105 determines a plurality of local lighting conditions and / or other local environmental conditions in a plurality of respective local portions 403 of the physical environment 400.
[0058] For example, the electronic device 105 can determine the ambient light level, light direction, color, texture, material, and / or other environmental aspects of each of the portions 403 of the physical environment 400. The environmental conditions obtained in each of the portions 403 of the physical environment 400 can then be provided to any processing stage of any processing pipeline (e.g., image pre-processing operations 304, image pre-processing operations 312, image pre-processing operations 316, sensor-based processes 306, and / or sensor-based processes 308) that utilize the environmental condition information and / or obtain information from and / or render virtual content into those respective areas. In this way, in one or more implementations, inaccurate, biased, and / or otherwise erroneous environmental condition estimates can be prevented from being provided to and / or propagated by the processing pipelines and / or sub-stages thereof at the electronic device 105.
[0059] In the example of FIG. 4A, the local portions 403 of the physical environment 400 can be obtained by the electronic device 105 using a plurality of cameras 402. In various implementations, the cameras 402 can be positioned in a variety of ways. For example, the cameras 402 can be positioned in a manner that is similar to the cameras 102 of FIG. 1. In various implementations, the cameras 402 can be positioned in a manner that is similar to the cameras 202 of FIG. 2. In various implementations, the cameras 402 can be positioned in a manner that is similar to the cameras 302 of FIG. 3. In various implementations, the cameras 402 can be positioned in a manner that is similar to the cameras 402 of FIG. 4A. Figure 5 In the example of FIG. 4A, the local portions 403 of the physical environment 400 can be obtained by the electronic device 105 using a plurality of cameras 402. In various implementations, the cameras 402 can be positioned in a variety of ways. For example, the cameras 402 can be positioned in a manner that is similar to the cameras 102 of FIG. 1. In various implementations, the cameras 402 can be positioned in a manner that is similar to the cameras 202 of FIG. 2. In various implementations, the cameras 402 can be positioned in a manner that is similar to the cameras 302 of FIG. 3. In various implementations, the cameras 402 can be positioned in a manner that is similar to the cameras 402 of FIG. 4A.
[0060] In some use cases, obtaining local environmental condition estimates for portions 403 of the entire field of view 401 of the physical environment 400 can be an inefficient use of device resources (e.g., processing resources, memory resources, and / or power resources). For example, in a use case in which a user of the electronic device 105 performs a hand gesture to interact with the virtual cup 420 on the surface 412 of the physical table 406, the processing pipeline and / or its sub-stages active at the electronic device can not use local environmental condition estimates in the illuminated portion 416 of the physical floor 404 when the user is interacting with the virtual cup 420.
[0061] As Figure 3 In one or more implementations, the electronic device 105 can obtain local environmental condition estimates (e.g., local lighting condition estimates) only in a salient region 500 of the physical environment 400, as illustrated. For example, the salient region 500 can be identified by the electronic device 105 as being salient to a user of the electronic device 105 based on user actions within the salient region 500. For example, the salient region 500 can be identified by the electronic device 105 based on a user gesture (e.g., a hand gesture) within the salient region 500, a user gaze location within the salient region 500, and / or one or more objects detected and / or identified in the scene. For example, as Figure 6 As illustrated, the user of the electronic device can be gazing at a gaze location 502 corresponding to the location of the virtual cup 420. In response to determining the gaze location 502, the electronic device 105 can identify the salient region 500 as a region around the gaze location 502 (e.g., by identifying a region of a particular size or radius around the gaze location 502). In other examples, the salient region 500 can be determined based on a user gesture corresponding to the virtual cup 420 and / or based on historical user interactions within the salient region 500 (e.g., history of the user placing virtual content on or interacting with the surface 412 of the physical table 406, as learned by the electronic device 105 during use of the electronic device 105).
[0062] In one or more implementations, the electronic device 105 can determine environmental condition estimates for multiple portions 403 of the physical environment 400 within the salient region 500. In one or more other implementations, the electronic device 105 can determine a single environmental condition estimate for the entire salient region 500. The environmental condition estimates within the salient region 500 can be provided to one or more processing pipelines and / or its sub-stages (e.g., as described in connection with Figure 7The environment condition estimates can be used to generate a three-dimensional scene for display by the electronic device 105 using one or more images from one or more cameras 150, including a view of a region (e.g., a field of view 401) of the physical environment 400 based on the one or more images from the one or more cameras 150, and virtual content (e.g., the virtual cup 420) overlaid on the view of the region of the physical environment 400. In one or more implementations, the electronic device 105 can determine a portion of a physical environment (e.g., the physical environment 400) that is salient to a user of the electronic device (e.g., a salient region 500); obtain an estimate of an environment condition (e.g., a local lighting condition or other local environment condition) of the physical environment locally made in the portion of the physical environment; and determine an input to the electronic device (e.g., a hand gesture) or an output of the electronic device (e.g., virtual content such as the virtual cup 420) based at least in part on the estimate of the environment condition of the physical environment locally made in the portion of the physical environment.
[0063] In one or more implementations, the salient region 500 of the physical environment 400 can be identified using a saliency map (e.g., identified to or by an environment condition estimator such as the light estimation process 302, the light estimation process 310, and / or the light estimation process 314). Figure 6 and Figure 4 An illustrative example of a saliency map is shown. For example, Figure 5 The saliency map 600 is illustrated as a binary saliency map having a salient portion 602 and a non-salient portion 604. For example, the saliency map 600 can have a region corresponding to Figure 5 and Figure 6 the field of view 401, and the salient portion 602 can have a region corresponding to (e.g., identifying) Figure 7 the salient region 500.
[0064] As Figure 7As illustrated, when saliency map 600 is provided to an environmental condition estimator (e.g., light estimation process 302, light estimation process 310, and / or light estimation process 314) (or generated therefrom), the environmental condition estimator can obtain environmental condition estimates for one or more portions 403 of the physical environment 400, which are located within salient regions of the physical environment as identified by the salient portion 602 of saliency map 600. In one or more embodiments, the salient portion 602 of saliency map 600 may be centered on or otherwise based on the location of the user's gaze. In one or more embodiments, the salient portion 602 of saliency map 600 may be centered on or otherwise based on the location of the user's gesture. In one or more embodiments, the salient portion 602 of saliency map 600 may be centered on or otherwise based on an object in the physical environment that has been determined to be salient to the user. In some examples, the environmental condition estimator can obtain a single environmental condition estimate for the entire salient region identified by the salient portion 602.
[0065] Figure 8 Another example of a saliency plot that can be provided to (or generated by) an environmental condition estimator is shown. Figure 1 In the example, saliency map 700 includes a high saliency portion 702, a medium saliency portion 703, and a low saliency portion 704. In one or more embodiments, an environmental condition estimator (e.g., light estimation process 302, light estimation process 310, and / or light estimation process 314) that generates or receives saliency map 700 can obtain environmental condition estimates in one or more portions 403 of physical environment 400, which are located within the salient region of physical environment identified by the high saliency portion 702 of saliency map 700.
[0066] The environmental condition estimator may also obtain, for example, a single environmental condition estimate for the region corresponding to the entire moderately salient portion 703 of the physical environment 400, and may not obtain an environmental condition estimate for the region corresponding to the lowly salient portion 704 of the physical environment 400. In one or more embodiments, saliency map 600 may be referred to as a first-order saliency map, and saliency map 700 may be referred to as a second-order saliency map. In one or more embodiments, saliency maps at other saliency levels may be generated (e.g., a zero-order saliency map covering the entire field of view 401, a third-order saliency map with multiple moderately salient regions and / or more fine-grained regions of interest, etc.).
[0067] In one or more embodiments, electronic device 105 may recursively apply a local saliency mask to a region of interest in a physical environment. For example, in one or more embodiments, electronic device 105 may obtain an estimate of local lighting conditions and / or another estimate of local environmental conditions by computing estimates at multiple saliency levels (e.g., using multiple corresponding saliency maps) and determining the optimal estimate based on the estimates at the multiple saliency levels. For example, recursively determining the estimate may include determining the estimate for that specific saliency level based on the estimate in the highly saliency portion of the saliency map for that particular saliency level and based on the estimate for previous saliency levels.
[0068] In one or more embodiments, the estimation of local lighting conditions or other local environmental conditions may include estimations with a time dimension. For example, in one or more embodiments, the estimation may be weighted based on adjacent time estimates (e.g., using temporal locality). For example, in some embodiments, a time vector estimate SBE(L, T) may be created, where L indicates the spatial dimension (e.g., a spatial region in field of view 401) and T indicates the time index of the estimate (e.g., a timestamp).
[0069] Figure 2 An example process 800 for local environment input sensing for an electronic device, according to one or more specific implementations, is illustrated. For illustrative purposes, this document primarily refers to... Figure 1 and Figure 2 The process 800 is described using electronic device 105. However, the process 800 is not limited to... Figure 8 and Figure 4 The electronic device 105, and one or more blocks (or operations) of process 800 may be performed by one or more other components of other suitable devices (including electronic device 110, electronic device 115, and / or server 120). Further for illustrative purposes, some blocks of process 800 are described herein as occurring sequentially or linearly. However, multiple blocks of process 800 may occur in parallel. Furthermore, the blocks of process 800 need not be performed in the order shown, and / or one or more blocks of process 800 need not be performed and / or may be replaced by other operations.
[0070] exist Figure 4In the example of FIG. 8, at block 802, an electronic device (e.g., electronic device 105) can determine one or more (e.g., multiple) local lighting conditions for one or more (e.g., multiple) respective local portions (e.g., local portion 403) of a physical environment (e.g., physical environment 400). At least a portion (e.g., field of view 401) of the physical environment can be visible to one or more cameras of the electronic device. For example, the one or more respective local portions of the physical environment can be within a field of view (e.g., field of view 401) of the one or more cameras of the electronic device and smaller than the field of view. In one or more implementations, each local lighting condition of the one or more local lighting conditions can include an ambient light level in the respective local portion of the physical environment.
[0071] In one or more implementations, at least one local lighting condition of the one or more local lighting conditions can include a direction corresponding to a light source (e.g., direction of directional light 415 of physical light 408). Figure 3 For example, the direction of the light source can be inferred by the electronic device based on a detected position and / or direction of a shadow of a physical object in the physical environment. In one or more implementations, the one or more local lighting conditions can include a first local lighting condition on a first side of a light discontinuity in the physical environment (e.g., in shadow 414 on a first side of luminance discontinuity 418) and a second local lighting condition on a second, opposite side of the light discontinuity (e.g., in illuminated portion 416 of physical floor 404 on another side of luminance discontinuity 418). Figure 7
[0072] At block 804, the electronic device can determine an input (e.g., a hand gesture or a hand position) to the electronic device or an output (e.g., a rendering of a scene such as a three-dimensional scene or an XR scene that includes virtual content such as virtual cup 420) of the electronic device using the one or more local lighting conditions. For example, the electronic device can identify a position of a hand or a hand gesture based on the one or more local lighting conditions and at least one image from the one or more cameras. As another example, the electronic device can generate a three-dimensional scene for display by the electronic device using the one or more local lighting conditions and at least one image from the one or more cameras, the three-dimensional scene including a view of a region of the physical environment based on the at least one image from the one or more cameras and virtual content overlaid on the view of the region of the physical environment. For example, the virtual content can be generated based on one or more of the one or more local lighting conditions. For example, a color, a brightness, a shading, a shadow, and / or other attributes of the virtual content can be generated based on the local lighting conditions (e.g., ambient light levels, light directions, colors, textures, materials, etc.) in the physical environment at a location where the virtual content is displayed to appear in the physical environment.
[0073] In one or more implementations, generating the three-dimensional scene can include providing the one or more local lighting conditions to a processing stage (e.g., image pre-processing operation 304, image pre-processing operation 312, image pre-processing operation 316, or sensor-based process 306) in a processing chain (e.g., one or more of the depicted processing chains) that includes at least one subsequent processing stage (e.g., sensor-based process 306 or sensor-based process 308) after the processing stage. For example, the processing chain can be configured to perform at least one of: image pre-processing of the at least one image from the one or more cameras, computer vision operations using the at least one image from the one or more cameras, three-dimensional immersion effect generation for the three-dimensional scene, gesture detection, surface texture estimation, scene reconstruction, object tracking, six degrees of freedom (6DOF) viewing of the virtual content, or spatial audio processing. Figure 9
[0074] In one or more implementations, the process 800 can further include determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to a user of the electronic device prior to obtaining the one or more local lighting conditions. For example, determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to the user can include identifying a salient region of the physical environment and selecting the one or more respective local portions from within the salient region. In one or more other implementations, a single local lighting condition can be determined for an entire salient region of a physical environment. In one or more implementations, obtaining the one or more local lighting conditions can include recursively obtaining the one or more local lighting conditions using saliency maps at multiple levels of saliency (e.g., saliency map 600 and / or saliency map 700) (e.g., as described herein in connection with FIGS. 6 and 7). Figure 1
[0075] In one or more implementations, determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to the user of the electronic device can include detecting a body part of the user in at least one of the one or more respective local portions of the physical environment. In one or more implementations, determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to the user of the electronic device can include determining that at least one of the one or more respective local portions of the physical environment is associated with a gaze of the user. In one or more implementations, determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to the user of the electronic device can include determining that at least one of the one or more respective local portions of the physical environment is associated with a gesture of the user. In one or more implementations, determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to the user of the electronic device can include determining that at least one of the one or more respective local portions of the physical environment is associated with an object in the physical environment (e.g., an object that is frequently interacted with by the user, an object that is frequently used as an anchor point for virtual content, or an object that is otherwise determined to be salient to the user).
[0076] In one or more embodiments, process 800 may further include: determining, before obtaining the one or more local lighting conditions, that an area of the physical environment including the one or more corresponding local portions of the physical environment is salient to the user of the electronic device. For example, the electronic device may determine that the area is salient to the user of the electronic device based on user actions corresponding to the area. For example, the electronic device may determine that the area is salient to the user of the electronic device based on detecting a body part (e.g., a hand) of the user in the area of the physical environment. Alternatively, the electronic device may determine that the area is salient to the user of the electronic device based on detecting the user's gaze position in the area of the physical environment. Or, the electronic device may determine that the area is salient to the user of the electronic device based on detecting a specific object in the area of the physical environment.
[0077] Figure 2 An example process 900 for sensing local environmental inputs in a salient region of a physical environment, based on one or more specific implementations, is illustrated. For illustrative purposes, this document primarily refers to... Figure 1 and Figure 2 The process 900 is described using electronic device 105. However, process 900 is not limited to... Figure 9 and Figure 5 The electronic device 105, and one or more blocks (or operations) of process 900 may be performed by one or more other components of other suitable devices (including electronic device 110, electronic device 115, and / or server 120). Further for illustrative purposes, some blocks of process 900 are described herein as occurring sequentially or linearly. However, multiple blocks of process 900 may occur in parallel. Furthermore, the blocks of process 900 need not be performed in the order shown, and / or one or more blocks of process 900 need not be performed and / or may be replaced by other operations.
[0078] exist Figure 7 In the example, at box 902, an electronic device (e.g., electronic device 105) can determine a significant portion (e.g., significant area 500) of the physical environment (e.g., physical environment 400) for the user of the electronic device (e.g., significant area 500) (e.g., as combined herein). Figure 3 (As described).
[0079] At block 904, the electronic device can obtain an estimate of an environmental condition (e.g., a local lighting condition or other local environmental condition) of the physical environment locally in the portion of the physical environment. In one or more implementations, the electronic device can obtain multiple estimates of an environmental condition in multiple local regions (e.g., portions 403) within the portion of the physical environment that is significant to the user. In one or more implementations, the electronic device can recursively obtain estimates of the environmental condition using multiple resolutions and levels of significance (e.g., as described herein in connection with Figure 8
[0080] At block 906, the electronic device can determine an input (e.g., a hand gesture) to the electronic device or an output (e.g., a rendering of a scene such as a three-dimensional scene or an XR scene that includes virtual content such as virtual cup 420) of the electronic device based at least in part on the estimate of the environmental condition of the physical environment locally in the portion of the physical environment. For example, the electronic device can determine the input or the output by providing the estimate of the environmental condition to an image pre-processing operation and / or one or more sensor-based processes (e.g., as described herein in connection with Figure 10 In one or more implementations, determining the input or the output can include generating a three-dimensional scene for display by the electronic device based at least in part on the estimate of the environmental condition of the physical environment locally in the portion of the physical environment, the three-dimensional scene including: a view of a region of the physical environment based on at least one image from one or more cameras, and virtual content overlaid on the view of the region of the physical environment (e.g., as described herein in connection with block 804 of FIG. 8). Figure 1
[0081] In one or more implementations, a method can be provided that includes: determining, by an electronic device, one or more local lighting conditions of one or more respective local portions of a physical environment, the one or more respective local portions of the physical environment being within a field of view of one or more cameras of the electronic device and smaller than the field of view; and using the one or more local lighting conditions (e.g., and at least one image from the one or more cameras) to perform a spatial computing operation (e.g., detecting user interactions and / or generating spatial images, spatial audio, and / or spatial video).
[0082] In one or more implementations, a method can be provided that includes: determining, by an electronic device, one or more local lighting conditions of one or more respective local portions of a physical environment, the one or more respective local portions of the physical environment being within a field of view of one or more cameras of the electronic device and smaller than the field of view; detecting a user input to the electronic device based on the one or more local lighting conditions.
[0083] As described above, one aspect of the present technology is the collection and use of data available from specific and legitimate sources for providing local environmental input sensing for electronic devices. The present disclosure contemplates that, in some instances, this collected data can include personal information data that uniquely identifies or can be used to identify a specific person. Such personal information data can include audio data, voice recordings, demographic data, location-based data, online identifiers, telephone numbers, email addresses, home addresses, cryptographic information, data or records relating to a user’s health or exercise performance (e.g., vital signs measurements, medication information, exercise information), birth date, or any other personal information.
[0084] The present disclosure recognizes that the use of personal information data in the present technology can be used to the benefit of users. For example, the personal information data can be used to provide local environmental input sensing for electronic devices.
[0085] The present disclosure contemplates that the entities responsible for the collection, analysis, disclosure, transfer, storage, or other use of such personal information data will comply with well-established privacy policies and / or privacy practices. In particular, such entities should implement and consistently apply privacy policies and practices that are generally recognized as meeting or exceeding industry or governmental requirements for maintaining privacy and security, including, for example, principles published by GAPP, the Network Advertising Initiative, and the like. Such policies should be applied to all data held by the entity in order to maintain individual privacy and security. In addition to being consistent with industry standards that are repeatedly modified in order to hermetic standards, such policies should be specifically adapted to the activities of the entity collecting and / or using the personal information data. Such entities should also cover themselves from third-party evaluations on compliance with these established privacy policies and / or privacy practices. Furthermore, the policies should be adapted with the specific type of personal information data collected and / or accessed as well as the specific manner that such personal information data is collected and / or accessed. Additionally, the policies should be compliant with applicable laws and standards including specific considerations of jurisdictions that can impose more stringent standards than industry norms. For example, in the United States, the collection or obtaining of certain health data can be subject to federal- and / or state-specific laws, such as the Health Insurance Portability and Accountability Act (HIPAA) and the rules and / or regulations issued thereunder, and privacy of health information can be more specifically protected. Similarly, in areas outside the United States, the collection and use of personal information data about an individual can be subject to more stringent processing requirements, such as the European Union’s General Data Protection Regulation.
[0086] Notwithstanding the foregoing, the present disclosure also contemplates embodiments in which users selectively block the use or access of personal information data. That is, the present disclosure contemplates that hardware and / or software elements can be provided to prevent or block access to such personal information data. For example, in the case of local environment input sensing for electronic devices, the present technology can be configured to allow users to "opt in" or "opt out" of participation in the collection and / or sharing of personal information data during registration for services or anytime thereafter. In addition to providing the "opt in" and "opt out" options, the present disclosure contemplates providing notifications relating to the access or use of personal information. For instance, a user can be notified of the access of their personal information data when the application is downloaded, and then again just prior to the application accessing the personal information data.
[0087] Moreover, it is the intent of the present disclosure that personal information data should be managed and processed at a level of granularity that minimizes the risk of unintended or unauthorized access or use. Data can be minimized by limiting collection and deletion of data, once it is no longer needed. In addition, and when applicable, data de-identification can be used to protect the privacy of the user. De-identification can be facilitated, when appropriate, through removal of identifiers, control of the amount or
[0088] Thus, although the present disclosure broadly covers technologies using personal information data to implement one or more various disclosed embodiments, the present disclosure also contemplates embodiments that can be implemented without the need for accessing such personal information data. That is, the various embodiments of the present technology can not operate in this capacity without the need of accessing the above- described personal information data.
[0089] Figure 10 An electronic system 1000 is illustrated that can be used to implement one or more specific implementations of the subject technology. The electronic system 1000 can be, and / or can be a part of, the electronic device 105, the handheld electronic device 104, the electronic device 110, the electronic device 115, and / or the server 120 as shown in Figure 1 The electronic system 1000 can include various types of computer readable media and interfaces for various other types of computer readable media. The electronic system 1000 includes a bus 1008, one or more processing unit(s) 1012, a system memory 1004 (and / or buffer(s)), a ROM 1010, the permanent storage device 1002, an input device interface 1014, an output device interface 1006, and one or more network interfaces 1016, or subsets and variations thereof.
[0090] Bus 1008 generally represents any type or form of interconnection mechanism for transmitting information between various internal devices, including control signals, data signals, or other suitable information between multiple internal devices, such as processing unit(s) 1012, ROM 1010, system memory 1004, and permanent storage device 1002. In one or more specific embodiments, bus 1008 can be a street address bus over which address addresses can be transmitted, a data bus over which data can be transmitted, a control bus over which control signals can be transmitted, or any combination thereof. Via bus 1008, processing unit(s) 1012 can be communicatively coupled to ROM 1010, system memory 1004, and permanent storage device 1002. Processing unit(s) 1012 can retrieve instructions from various memory units to execute and process data stored in various memory units, in order to execute processes of the subject disclosure. In different embodiments, processing unit(s) 1012 can be a single processor, or a multi-core processor.
[0091] ROM 1010 stores static data and instructions that are needed by processing unit(s) 1012 and other modules of electronic system 1000. Permanent storage device 1002, on the other hand, can be a read-and-write memory device. Permanent storage device 1002 can be a non-volatile memory unit that stores instructions and data even when electronic system 1000 is off. In one or more specific embodiments, a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) can be used as permanent storage device 1002.
[0092] In one or more specific embodiments, a removable storage device (such as a floppy disk, flash drive, and its corresponding disk drive) can be used as permanent storage device 1002. Like permanent storage device 1002, system memory 1004 can be a read-and-write memory device. However, unlike permanent storage device 1002, system memory 1004 can be a volatile read-and-write memory, such as a random access memory. System memory 1004 can store any of the instructions and data that processing unit(s) 1012 can need at runtime. In one or more specific embodiments, the processes of the subject disclosure are stored in system memory 1004, permanent storage device 1002, and / or ROM 1010 (each of which is implemented as a non-transitory computer-readable medium), which are each accessed by processing unit(s) 1012 to retrieve instructions and data to execute and process. In different embodiments, system memory 1004 can be described as a computer-readable medium.
[0093] Bus 1008 also connects to input and output devices 1012 and 1010. Input device(s) 1012 enable the user to communicate information and select commands to electronic system 1000. Input devices 1012 can include, for example, alphanumeric input device(s) and cursor control device(s), such as a keyboard and a pointing device (also called a "cursor control device"). Output device(s) 1010 can include, for example, a printer and display device(s), such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a flexible display, a flat-panel display, a solid-state display, a projector, or any other device for outputting information. One or more specific implementations can include a device that functions as both an input and an output device, such as a touch screen. In these implementations, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0094] Finally, as shown, bus 1008 also couples electronic system 1000 to one or more networks and / or to one or more network nodes, such as electronic device 110, through one or more network interfaces 1016. In this manner, electronic system 1000 can be a part of a network of computers (such as a LAN, a wide area network ("WAN"), or an Intranet, for example), or a part of the Internet, for example. Any or all components of electronic system 1000 can be used in conjunction with the subject disclosure.
[0095] These functions described above can be implemented in computer software, firmware, or hardware. The techniques can be implemented using one or more computer program products. Programmable processors and computers can be included in or packaged as mobile devices. The processes and logic flows can be performed by one or more programmable processors and by one or more programmable logic circuitry. General and special purpose computing devices and storage devices can be interconnected through communication networks.
[0096] Some implementations include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine -readable or computer-readable medium (also referred to as computer-readable storage media, machine-readable media, or machine-readable storage media). Some examples of such computer- readable media include RAM, ROM, read-only compact discs (CD-ROM), recordable compact discs (CD-R), rewritable compact discs (CD-RW), read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM), a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory devices (e.g., SD cards, mini-SD cards, micro-SD cards, etc.), magnetic and / or solid-state hard drives, read-only and recordable Blu-ray® discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media can store the computer program instructions which can be executed by at least one processing unit and which include a set of instructions that can be executed by the at least one processing unit for performing various operations. ® Examples of computer program instructions include both machine code, such as produced by a compiler, and files containing a high-level code that can be executed by the computer, the electronic component, or the microprocessor using an interpreter.
[0097] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, some implementations are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some implementations, such integrated circuits execute instructions stored on the circuit itself.
[0098] As used in this description and claims, the terms “computer”, “server”, “processor”, and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. For purposes of this description, the term display or displaying means display on an electronic device. As used in this description and claims, the terms “computer readable medium” and “computer readable media” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.
[0099] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0100] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0101] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). The data generated at the client device can be received from the client device at the server (e.g., as a result of the user interacting with the client device).
[0102] Implementations within the scope of the disclosure can be partially or fully implemented with tangible computer-readable storage media (or multiple tangible computer-readable storage media of one or more types) encoding one or more instructions. Tangible computer-readable storage media as used in this context can also be non-transitory.
[0103] A computer-readable storage medium can be any storage medium that can be read, written, or otherwise accessed by a general or special purpose computing device, including any processing electronics and / or processing circuitry capable of executing instructions. For example, without limitation, a computer-readable medium can include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. A computer-readable medium can also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, flash, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and Millipede memory.
[0104] Furthermore, a computer-readable storage medium can include any non- semiconductor memory, such as optical disk storage, magnetic disk storage, magnetic tape, other magnetic storage devices, or any other storage medium that can store one or more instructions. In one or more implementations, a tangible computer-readable storage medium can be directly coupled to a computing device, while in other implementations, a tangible computer-readable storage medium can be indirectly coupled to a computing device, e.g., via one or more wired connections, one or more wireless connections, or any combination thereof.
[0105] Instructions can be directly executable, or can be used to develop executable instructions. For example, instructions can be implemented as executable or non-executable machine code, or as high-level language instructions that can be compiled to produce executable or non-executable machine code. Furthermore, instructions can also be implemented as data, or can include data. Computer-executable instructions can also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As recognized by those of skill in the art, details including, but not limited to, the number, structure, sequence, and organization of instructions can vary significantly without altering the underlying logic, function, processing, and output.
[0106] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, one or more implementations are performed by one or more integrated circuits such as ASICs or FPGAs. In one or more implementations, such integrated circuits execute instructions stored on the integrated circuits themselves.
[0107] Those skilled in the art will appreciate that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein can be implemented as electronic hardware, computer software, or combinations of both. To illustrate the interchangeability of hardware and software, various illustrative blocks, modules, elements, components, methods, and algorithms have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application. Various components and blocks can be arranged differently or eliminated from the configuration described above, and the disclosed aspects of the subject technology can be implemented in a variety of other ways.
[0108] It should be understood that any particular order or sequence of blocks in the processes disclosed herein are illustrative. Based upon design preferences, it is understood that the particular order or sequence of blocks in the processes can be rearranged, or that all illustrated blocks can be performed. Any of the blocks can be performed simultaneously. In one or more particular implementations, multitasking and parallel processing can be advantageous. Additionally, the division of various system components described in the implementations above should not be understood as requiring such division in all implementations and it should be understood that program components and systems can generally be integrated together in a single software product or encapsulated into multiple software products.
[0109] As used in this specification and any claims of this application, the terms "base station", "receiver", "computer", "server", "processor", and "memory" all refer to electronic or other technological devices. These terms exclude human beings or groups of human beings. For purposes of this specification, the term "display" or "displaying" means displaying on an electronic device.
[0110] As used herein, the phrase "at least one of" followed by a listing of two or more items is used to denote that at least one of the listed items is required, but that more than one of the listed items can be included. For example, if a composition is described as containing "at least one of X, Y, and Z," the composition can contain X alone; Y alone; Z alone; two of X, Y, and Z; all three of X, Y, and Z; or any combination thereof. In other words, the phrase "at least one of" is used to indicate that the listed items are alternatives for each other, but is not limited to requiring at least one of each item listed.
[0111] The predicate words “configured to,” “operable to,” and “programmed to” do not imply any specific tangible or intangible modification of a subject, but, rather, are intended to be synonymous with the phrase “caused to,” with respect to an operation, or with respect to an element. In one or more specific embodiments, a processor configured to monitor and control an operation or element can also mean the processor is caused to monitor and control the operation or element. Likewise, a processor configured to execute code can be interpreted as a processor programmed to execute code or a processor operable to execute code.
[0112] The phrases “one or more of the following aspects,” “some of the aspects,” “one or more aspects,” “one or more of the aspects,” “some aspects,” “one or more of the embodiments,” “some embodiments,” “one or more embodiments,” “aspects,” “embodiments,” “an aspect,” “an implementation,” “some aspects,” “one or more aspects,” “the aspect,” “the implementation,” “another aspect,” “some implementations,” “one or more implementations,” “the embodiment,” “the implementation,” “another embodiment,” “some embodiments,” “one or more embodiments,” “configuration,” “the configuration,” “other configurations,” “some configurations,” “one or more configurations,” “subject technology,” “disclosure,” “the disclosure,” “other variations thereof,” and the like are merely used herein to facilitate discussion of various example aspects, one or more of which can he useful in one or more implementations, and do not necessarily all refer to the same one or more aspects, or the same one or more implementations. An aspect, or some aspects, of the disclosure can be implemented alone, or in combination with one or more other aspects, or some other aspects, of the disclosure. An aspect, or some aspects, of the disclosure can be implemented with or without reference to one or more other aspects, or some other aspects, of the disclosure. The disclosure is not limited to the aspects described herein, but can include any appropriate aspect, or combination thereof, that is or can be developed.
[0113] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any implementation described herein as “exemplary” or as an “example” is not necessarily to be construed as preferred or advantageous over other implementations. Furthermore, to the extent that the term “includes” is used in either the detailed description or the claims, such term is intended to be inclusive in a manner similar to the term “comprising” as “comprising” is interpreted when used as a transitional term in a claim.
[0114] All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether these disclosure is explicitly recited in the claims. No claim element is to be construed as a means plus function unless the element is expressly recited using the phrase “means for” or “step for” and only to the extent that it is expressed as a means or step plus function. The following examples are provided to further illustrate various aspects of the disclosure and are not intended to limit the scope of the disclosure.
[0115] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean "one and only one" unless specifically so stated, and wherein advantages offered are not necessarily equally applicable to all aspects. Unless otherwise specifically indicated, the term "some" refers to one or more. Pronouns in the masculine (e.g., his) include the feminine and neuter gender (e.g., her and its) and vice versa, so that the articles "a," "an," and "the" are not limited to a singular referring object but include the plural as well, unless the context clearly indicates otherwise. Headings and subheadings (if any) are provided for convenience and do not interpret the subject disclosure.
Claims
1. A method comprising: determining, by an electronic device, one or more local lighting conditions for one or more respective local portions of a physical environment, wherein the one or more respective local portions of the physical environment are within a field of view of one or more cameras of the electronic device and are smaller than the field of view; and generating, for display by the electronic device, a three-dimensional scene using the one or more local lighting conditions and at least one image from the one or more cameras, the three-dimensional scene comprising: a view of an area of the physical environment based on the at least one image from the one or more cameras, and virtual content overlaid on the view of the area of the physical environment.
2. The method of claim 1, wherein each of the one or more local lighting conditions comprises an ambient light level in the respective local portion of the physical environment.
3. The method of claim 2, wherein at least one of the one or more local lighting conditions comprises a direction corresponding to a light source.
4. The method of claim 1, wherein generating the three-dimensional scene comprises providing the one or more local lighting conditions to a processing stage in a processing chain, the processing chain comprising at least one subsequent processing stage after the processing stage.
5. The method of claim 4, wherein the processing chain is configured to perform at least one of: image pre-processing on the at least one image from the one or more cameras, a computer vision operation using the at least one image from the one or more cameras, three-dimensional immersion effect generation for the three-dimensional scene, gesture detection, surface texture estimation, scene reconstruction, object tracking, spatial computation, six degrees of freedom display of the virtual content, or spatial audio processing.
6. The method of claim 4, wherein the one or more local lighting conditions comprise a first local lighting condition on a first side of a light discontinuity in the physical environment and a second local lighting condition on a second, opposite side of the light discontinuity.
7. The method of claim 1, further comprising: determining, prior to obtaining the one or more local lighting conditions, that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to a user of the electronic device.
8. The method of claim 7, wherein determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to the user of the electronic device comprises: detecting, in at least one of the one or more respective local portions of the physical environment, a body part of the user.
9. The method of claim 7, wherein determining that the one or more respective local portions of the physical environment are one or more portions of the physical environment that are salient to the user of the electronic device comprises: determining that at least one of the one or more respective local portions of the physical environment is associated with a gaze of the user.
10. A non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: determining, by an electronic device, a plurality of local lighting conditions for a plurality of respective local portions of a physical environment, at least a portion of the physical environment being visible to one or more cameras of the electronic device; and generating, for display by the electronic device, a three-dimensional scene using the plurality of local lighting conditions and one or more images from the one or more cameras, the three-dimensional scene including: a view of a region of the physical environment based on the one or more images from the one or more cameras, and virtual content overlaid on the view of the region of the physical environment.
11. The non-transitory machine-readable medium of claim 10, wherein each of the local lighting conditions comprises an ambient light level in the respective local portion of the physical environment.
12. The non-transitory machine-readable medium of claim 11, wherein at least one of the local lighting conditions comprises a direction corresponding to a light source in the respective local portion of the physical environment.
13. The non-transitory machine readable medium of claim 10, wherein the operations further comprise: determining one of a local color or a local texture using at least one of the plurality of local lighting conditions.
14. The non-transitory machine-readable of claim 10, wherein generating the three- dimensional scene comprises providing the plurality of local lighting conditions to a processing stage in a processing chain, the processing chain comprising at least one subsequent processing stage after the processing stage.
15. The non-transitory machine-readable of claim 10, wherein the plurality of local lighting conditions comprises a first local lighting condition on a first side of a light discontinuity in the physical environment and a second local lighting condition on a second, opposite side of the light discontinuity.
16. The non-transitory machine readable of claim 10, the operations further comprising: determining, prior to obtaining the plurality of local lighting conditions, that the plurality of respective local portions of the physical environment are portions of the physical environment that are salient to a user of the electronic device.
17. An electronic device, the electronic device comprising: a memory; and at least one processor configured to: determine a plurality of local lighting conditions for a plurality of respective local portions of a physical environment, at least a portion of the physical environment being visible to one or more cameras of the electronic device; and generate, for display by the electronic device, a three-dimensional scene using the plurality of local lighting conditions and one or more images from the one or more cameras, the three-dimensional scene comprising: a view of a region of the physical environment based on the one or more images from the one or more cameras, and virtual content overlaid on the view of the region of the physical environment.
18. The electronic device of claim 17, wherein the at least one processor is further configured to determine, prior to obtaining the plurality of local lighting conditions, that a region of the physical environment comprising the plurality of respective local portions of the physical environment is salient to a user of the electronic device.
19. The electronic device of claim 18, wherein the at least one processor is further configured to determine that the region is salient to the user of the electronic device based on a user action corresponding to the region.
20. The electronic device of claim 17, wherein the plurality of local lighting conditions comprises a first local lighting condition on a first side of a light discontinuity in the physical environment and a second local lighting condition on a second, opposite side of the light discontinuity.