Environmental modeling based camera adaptation for

By using environmental modeling and sensor data processing, the camera parameters were adjusted to address the issue of image sensors failing to adapt to the physical environment in pass-through video, resulting in more stable video quality.

CN121844571APending Publication Date: 2026-04-10APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
APPLE INC
Filing Date
2024-09-06
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing pass-through video technologies, image sensors fail to properly or optimally adapt to the surrounding physical environment and other conditions, resulting in poor video quality.

Method used

By using environmental modeling and sensor data processing, camera parameters such as exposure, gain, tone mapping, color balance, and noise reduction are adjusted to adapt to the optical characteristics of the physical environment and user information, providing stable pass-through video.

Benefits of technology

It improves the appearance of video images, reduces flickering and blurring, and provides a more stable user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844571A_ABST
    Figure CN121844571A_ABST
Patent Text Reader

Abstract

Various implementations provide unvarnished video based on adjusting camera parameters according to environmental modeling. Environmental characteristics may be determined based on modeling a physical environment from sensor data captured via one or more sensors. For example, this may involve determining ambient light source optical characteristics, ambient surface optical characteristics, 3D mapping of the environment, user behavior, prediction of optical characteristics of light entering the camera, and the like. The method may involve determining camera parameters of an image captured via an image sensor based on environmental characteristics. For example, the method may determine exposure, gain, tone mapping, color balance, noise reduction, sharpness enhancement. The method may determine camera parameters based on user information (e.g., user preferences, user activity, etc.). The method may involve providing a transparent video of the physical environment based on the determined camera parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 541,104, filed September 28, 2023, which is incorporated herein in its entirety. TECHNICAL FIELD

[0002] The present disclosure relates generally to improving the appearance of video images captured and displayed by electronic devices, and more particularly to systems, methods, and devices that adaptively adjust pass-through video characteristics to account for the environment and other circumstances depicted in the video. BACKGROUND

[0003] Extended reality (XR) devices can provide a depiction or other view of the physical environment surrounding the device. In some cases, such a depiction or other view is provided by providing pass-through video (e.g., video images captured by outward-facing image sensors and used to provide a live view of the surrounding physical environment). Image sensors used in such pass-through video techniques can not be properly or optimally adapted to account for the surrounding physical environment and other circumstances. SUMMARY

[0004] Various implementations disclosed herein are based on providing pass-through video by adapting camera parameters (e.g., exposure, gain, tone mapping, color balance, noise reduction, sharpness enhancement) according to environment modeling and / or other circumstances. In some example implementations, a processor executes instructions stored in a computer-readable medium to perform a method. The processor can be included in an electronic device having one or more sensors including an image sensor. The method can involve capturing sensor data corresponding to a physical environment via the one or more sensors. The method can involve, for example, capturing images, depth data, motion data, temperature data, humidity data, audio data, etc.

[0005] The method can involve determining environment characteristics based on modeling the physical environment according to the sensor data captured via the one or more sensors. For example, this can involve determining ambient light source optical characteristics, ambient surface optical characteristics, a 3D mapping of the environment, user behavior in the environment, a prediction of optical characteristics of light entering the camera, etc.

[0006] The method can involve determining camera parameters for images captured via the image sensor based on the environment characteristics. For example, the method can determine exposure, gain, tone mapping, color balance, noise reduction, sharpness enhancement. The method can additionally or alternatively determine camera parameters based on user information (e.g., user preferences, user activity, etc.). The method can involve providing (e.g., capturing, modifying, presenting, etc.) pass-through video of the physical environment including one or more captured images based on the determined camera parameters.

[0007] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and components for performing or causing to perform any of the methods described herein. Attached Figure Description

[0008] To enable those skilled in the art to understand this disclosure, more detailed descriptions can be made with reference to aspects of some exemplary embodiments, some of which are shown in the accompanying drawings.

[0009] Figure 1 Examples of electronic devices used in a physical environment according to some specific implementations are shown.

[0010] Figure 2 It shows some specific implementations Figure 1 An example of environmental modeling of the physical environment.

[0011] Figure 3 The parameter adjustment process is shown according to some specific implementations.

[0012] Figure 4 This is a flowchart illustrating an exemplary method for providing pass-through video based on environmental modeling and other factors, according to some specific implementations.

[0013] Figure 5 An exemplary device configured according to some specific implementations is shown.

[0014] By convention, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Furthermore, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation

[0015] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.

[0016] Figure 1 An example of an electronic device 120 used by a user within a physical environment 100 is shown. The physical environment refers to the physical world that people can interact with and / or sense without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell. Figure 1 In the context, physical environment 100 is a room that includes sofa 130, table 125, light bulb 140, and ceiling light 135.

[0017] exist Figure 1 In the example, electronic device 120 is shown as a single device. In some specific implementations, electronic device 120 is worn by a user. For example, electronic device 120 may be as follows: Figure 1 The illustrated head-mounted device (HMD) is shown. Some embodiments of electronic device 120 are handheld. For example, electronic device 120 may be a mobile phone, tablet computer, laptop computer, etc. In some embodiments, the functionality of electronic device 120 is implemented via two or more devices (e.g., optionally including a base station). Other examples include laptop computers, desktop computers, servers, or other such devices that include additional capabilities in terms of power, CPU capability, GPU capability, storage capability, memory capability, etc. Multiple devices that can be used to implement the functionality of electronic device 120 may communicate with each other via wired or wireless communication.

[0018] Electronic device 120 captures and displays pass-through video of physical environment 100. In this example, exemplary frame 145 of the video is captured and displayed at electronic device 120. Frame 145 (and additional frames) may be captured and displayed sequentially, for example, as part of a sequence of frames captured in the same order as the frames were captured. In some implementations, frame 145 is displayed approximately simultaneously with its capture, for example, during a live video feed. In some implementations, frame 145 is displayed at a time after a delay period or after the video recording. Frame 145 includes a depiction 160 of sofa 130, a depiction 165 of table 125, a depiction 170 of lamp 140, and a depiction 180 of ceiling light 135. One or both light sources (e.g., lamp 140 and ceiling light 135) can contribute to the amount of light, flicker, and other characteristics affecting the appearance of the captured pass-through video. Furthermore, movement of device 120 (e.g., when a user rotates their head) may cause motion blur. Cameras that capture pass-through video can be adapted, for example, to take into account the physical environment based on modeling of the 3D environment, and / or to be adapted based on other factors such as user information.

[0019] In some implementations, the HMD is configured for video pass-through and is enabled to adaptively change camera parameters to reduce undesirable appearance properties of the video pass-through. Some implementations model and / or determine environmental characteristics (e.g., 3D mappings of light sources / surfaces in the environment) and use these environmental characteristics to adjust exposure to reduce flicker, blur, and / or otherwise improve video appearance. In some implementations, a 3D map, such as a SLAM map, is generated based on sensor data on the HMD and used to adjust camera parameters. The 3D map can be updated continuously (e.g., at a specific flexible update rate). Additional data, such as time of day, ambient brightness, semantic segmentation, planar maps, and / or occlusion understanding, may be used additionally or alternatively. Environmental characteristics, time of day, ambient brightness, semantic segmentation, planar maps, occlusion understanding, motion data, and / or any other relevant factors can be used to calculate optimal exposure parameters during video capture (e.g., during pass-through video delivery on the HMD).

[0020] In some implementations, environmental characteristics are persistently stored. 3D information about the environment can be updated and become more accurate over time as users interact with devices (and / or other devices) within the environment. The data persists over time and is used in different user sessions occurring at different times and dates.

[0021] Persistent mappings can be adjusted over time (e.g., based on sensor data acquired over time). Persistent maps can effectively store various types of information about the physical environment for later use, such as adjusting exposure at later points in time based on an understanding of the environment gained from previous user experiences. For example, at a later point in time, the HMD can determine its 3D position within and relative to the environment in a 3D mapping of light sources, and determine whether and how to adjust camera parameters to provide the desired user experience based on environmental characteristics from the 3D mapping.

[0022] Various techniques can be used to incorporate light source information into environmental characteristic data (e.g., 3D mapping). In one example, an image of the physical environment is acquired and evaluated to identify bright areas, and those bright areas are processed to identify and evaluate the characteristics of light sources within the 3D environment.

[0023] In some implementations, the device operates in one or more camera capture modes. For example, the device may operate according to a first flicker compensation mode (e.g., based solely on real-time flicker sensor data) until environmental characteristic data (e.g., 3D mapping) is developed, and then operate in a second flicker compensation mode (e.g., based on 3D mapping and / or real-time flicker sensor data).

[0024] In some implementations, a 3D mapping of the physical environment identifies the 3D position, shape, flicker rate, brightness, and / or other attributes of one or more light sources in the physical environment. Flicker attributes may be based on previously acquired sensor data. In some implementations, the 3D mapping is based on a SLAM process (e.g., a VIO SLAM process) that is initially performed periodically and does not necessarily provide flicker compensation during the experience. The 3D mapping may be updated over time based on flicker sensor data.

[0025] 3D mapping provides a persistent digital map of the physical environment. It can include information about the location and properties of light sources, the location and properties of surfaces, semantic data (e.g., identifying object types, material types, etc.), and / or information about the environment from which flicker insensitivity and other lighting characteristics can be estimated.

[0026] In some implementations, based on current conditions, a score is determined for each of one or more light sources in the physical environment, and this score is used to determine which (zero or more) light sources are providing unpleasant flicker that should be adjusted, for example, by adjusting exposure. A given light source may be associated with an exposure time / threshold at which flicker is expected to become unpleasant to the average observer or to a particular user.

[0027] In some implementations, the location the user is looking at in a perspective view of the physical environment is used to determine whether and how the exposure should be adjusted to account for flicker. Flicker adjustment can therefore occur in a given environment when the user is looking at a first area, but not in the same environment when the user is looking at a second, different area. Similarly, flicker adjustment (e.g., for dim lighting) may occur when overhead lights are off and the overall brightness of the environment is low, but may not occur when overhead lights are on and the overall brightness of the environment is high. Overall brightness reduces the visibility / aversion of flicker from the lighting.

[0028] In some implementations, user movement, such as head rotation, is evaluated when determining how to adjust camera parameters. For example, it can balance the appearance of flicker and blur in pass-through video. For instance, in cases where the head (and device) is stationary, flicker compensation can be performed to account for flickering light sources. However, in cases of rapid head (and device) rotation, such flicker compensation may be unnecessary (considering the difficulty of perceiving flicker when rotating a person's head) or outweighed by the requirement to reduce blur, for example, by adjusting exposure according to blur reduction parameters. Other factors can also be considered when balancing flicker and blur. For example, in an environment with striped walls, head movement may result in a lot of unpleasant blur (and therefore blur reduction may take precedence over flicker reduction), while in an environment with solid-color walls and surfaces with limited spatial contrast, head movement may not result in much more unpleasant blur (and therefore flicker reduction may take precedence over blur reduction).

[0029] Other factors include whether virtual content is added to the passthrough content, such as adding virtual user interface menus or partially opaque virtual content, how the content of the physical environment is occluded in the passthrough video, display brightness (e.g., in areas where flickering occurs), weighted measures that quantify the relative importance of blur and flicker, the size of the physical environment and / or the distance of areas in the view, the user's current task, the user's preference for flicker, blur, or other passthrough attributes, and the device's exposure capabilities (e.g., whether it is even possible to compensate for specific flicker or motion), etc.

[0030] Figure 2 Showing the target Figure 1 An example of 3D data generated for a physical environment 100. In this example, a 3D map 200 of the physical environment 100 is generated based on image, depth, or other sensor data from device 120 (or another device). The 3D map 200 includes 3D representations 205, 210, 220a-d of the ceiling, floor, and walls of the physical environment 100. The 3D map 200 also includes a 3D representation 225 of a table 125, a 3D representation 230 of a sofa 130, a 3D representation 240 of a lamp 140, and a 3D representation 235 of a ceiling light 135.

[0031] 3D mapping 200 stores information about the 3D location of light sources (e.g., 3D representation 240 of lamp 140 and 3D representation 235 of ceiling light 135) and information about those light sources, such as flicker rate, size, shape, bounding box area, light color, spectral range, brightness, current state (e.g., on, off, dimmed, etc.), type (e.g., window light, lamp, ceiling light, LED, LCD, OLED, incandescent, shadow, no shadow, diffuse, directional, illumination angle, illumination cone, etc.) and / or other attributes. Similarly, 3D mapping stores information about non-light surfaces (e.g., 3D representations 205, 210, 220a-d of ceiling, floor, and walls, 3D representation 225 of table 125, 3D representation 230 of sofa 130, etc.) and information about those surfaces, such as size, shape, reflectivity, texture, spatial contrast, etc.

[0032] The 3D mapping 200 information is used to determine whether and / or how to adjust the parameters (e.g., exposure) of one or more cameras providing pass-through video to provide the desired or optimal pass-through view. Observation points relative to the light sources and / or surfaces represented in the 3D mapping 200 can be determined and used in such determinations. For example, an observation point can be determined (based on the HMD's location) to be closer to the 3D representation 240 of the lamp 140 than to the 3D representation 235 of the overhead lamp 135, and the flickering of the lamp 240 is prioritized over the flickering of the overhead lamp 135. In some embodiments, the score for each light source is determined based on the observation point and the 3D mapping 200, for example, based on how unpleasant the flickering from each light source is expected to be to the observer at the observation point based on the environmental characteristics represented in the 3D mapping 200. In some embodiments, device movement relative to the light source location is also considered; for example, a light source that the user is moving closer to is prioritized relative to a light source that the user is moving away from.

[0033] The various embodiments disclosed herein improve the appearance of video images captured and displayed by electronic devices. This may involve adaptively adjusting the exposure or other parameters of the video capture to account for various factors related to the video appearance, including but not limited to light-based flicker, motion-induced blur, and noise. Some embodiments create a more perceptually stable pass-through video experience, thereby providing image brightness stability, flicker reduction, and / or color stability. In some embodiments, 3D VIO / SLAM or other environmental mappings are extended, for example, by adding a fourth dimension (time). Multimodal information can be used to create or update 3D maps, which can be used to inform camera / ISP parameters and / or decisions.

[0034] Figure 3The parameter adjustment is illustrated. In this example, environmental modeling and characteristic determination are performed (box 310) and used in conjunction with user information (box 350) to provide parameter adjustment (box 360), which is used to provide pass-through video based on the adjusted parameters (box 370). Figure 3 The parameter adjustment process can be configured to improve camera parameter adjustment by replacing (or supplementing) image statistics with environmental modeling information and / or information about other aspects of the experience.

[0035] The parameter tuning process may involve adjusting camera and image processing parameters, including but not limited to exposure, gain, tone mapping, and color balance. In one example, tone mapping is adjusted based on modeling the environment and understanding (in part based on that modeling) the state of brightness adaptation for the user's vision (e.g., how wide the user's eyes might be, how much light is currently entering the user's eyes, etc.). Such information can be used, for example, to crop and shift camera settings as the user adapts to a bright environment. The parameter tuning process may also involve adjusting auxiliary image processing parameters such as noise reduction and sharpness enhancement parameters.

[0036] In some specific implementations, parameter tuning is based on camera characteristics (e.g., camera calibration), environmental characteristics (e.g., image statistics, ambient light sensor (ALS) information, environmental modeling, flicker detection, etc.), and user information (e.g., human preferences, manual tuning, user research, and modeling).

[0037] In some implementations, parameters are adjusted based on the criterion of attempting to provide a relatively stable color experience over time (e.g., a consistent color temperature). In other implementations, parameters are adjusted based on the principle of attempting to balance dynamic range, image brightness stability, motion blur visibility, flicker visibility, stripe visibility, and / or noise visibility.

[0038] Various techniques can be used to adjust parameters based on such criteria. Some exemplary techniques detect light source characteristics (box 330). This may involve detecting the optical characteristics of one or more ambient light sources. This may involve detecting brightness, color / spectrum, temporal brightness / flicker distribution, physical techniques (e.g., sequential RGB or non-sequential RGB), spatial emission distribution (e.g., spotlight versus global), light source classification (e.g., identifying whether a light source is a window light, table lamp, etc.), etc. Detection may involve detection methods that include detecting ALS over time, using a flicker sensor to detect flicker at any time, using a spatial resolution flicker sensor, using a spatial resolution ALS, using ray tracing, performing a database lookup of known light sources (e.g., based on light source brand, model, etc.), using information available through smart home integration (e.g., smart home mapping information, smart device information, etc.), display / monitor detection (e.g., detecting displays based on computer vision or device-to-device communication), window detection, weather / forecast detection, use of date / time, time of year, year or other timing information, detection of glare and / or flare, etc.

[0039] Some exemplary techniques for detecting surface properties (box 340). This may involve detecting the optical properties of one or more environmental surfaces (e.g., wall surfaces, ceiling surfaces, floor surfaces, table surfaces, sofa surfaces, etc.). This may involve detecting surface color, whether the surface's optical properties are specular or diffuse, the surface's bidirectional reflectance distribution function (BRDF), transparency, etc. Detection may involve using image / camera data, ALS data, ray tracing data, machine learning, etc.

[0040] Some exemplary techniques are used to model the environment (box 320). This may involve performing 3D mapping of the environment, detecting 3D light source / surface poses (i.e., 3D position and orientation), and classification (e.g., identifying the type of object or room and / or characteristics such as object separation). Modeling (box 320) can be used to determine light source characteristics (box 330) and / or surface characteristics (box 340). 3D mapping of the environment can be based on a variety of information sources. For example, 3D mapping can be based on IMU data and camera data, VIO SLAM data, alternative environment mapping techniques, scene understanding techniques, semantic tagging / mapping techniques, etc.

[0041] For example, 3D mapping of light source and surface attitude can be integrated with the device's tracking system. The process can utilize device attitude information, ALS data, time-varying scintillation sensor data, spatial resolution scintillation sensor data, spatial resolution ALS data, ray tracing data, and multiple ALS / scintillation / cameras pointing in different directions.

[0042] In some implementations, camera parameters are adjusted based on image statistics. For example, historical data of image statistics when the device is pointed in a specific direction from a specific location can be used to adjust current camera parameters (e.g., in similar situations). Some implementations scale historical data based on image statistics in certain situations (e.g., when natural light is greater than a threshold proportion of light entering the camera).

[0043] User information (box 350) can be based on user behavior models. Such models can detect or classify user movement, for example, detecting whether a user is static or nearly static (e.g., when the user sits down). Models can detect whether a user is in a mobile experience, such as walking, running, playing a game, etc. Models can detect whether a user is moving to a different location or performing a specific type of movement (e.g., opening a door, walking along a corridor, etc.). Models can detect whether a user is on a mobile platform (e.g., a car, bus, train, airplane, elevator, etc.). Various methods can be used to determine user information. Some implementations use machine learning models that determine the type of use (e.g., use case detection) based on inputs such as, but not limited to, historical posture data, live posture data, eye-tracking data, determining which applications are being executed, where / what the user is moving towards, what the user is holding, where the user is looking, etc.

[0044] Some implementations predict the optical properties of light entering the camera in a reasonable pose within the environment. A reasonable pose can be the pose a user is likely to move into within a threshold time period (e.g., within the next 10 seconds, 30 seconds, 1 minute, 2 minutes, etc.), as determined by the device. Such a reasonable pose can be based on a 3D mapping of the environment, user behavior models, time ranges, etc. Some implementations perform ray tracing / rendering based on the camera's physical design (e.g., focus characteristics, f#, field of view (FOV), vignetting, responsivity, spectral QE of different color channels, transmittance, sensor timing / readout time, etc.). Some implementations predict light entering the camera based on far-field light mapping, occlusion mitigation / illusion, flicker distribution blending, camera shutter simulation (e.g., with a global flicker distribution), and / or using chromatographic blending of the camera's spectral responsivity from multiple sources.

[0045] Some implementations use high-level metrics (e.g., across multiple time scales) to determine brightness. In the short term, the relationship between motion blur and flicker on noise visibility can be evaluated. In the medium to long term, the relationship between noise visibility and dynamic range on brightness stability can be evaluated. In one example, there might be a bright light to the user's left. When the light is not in direct view, the camera's exposure may have to be adjusted each time the user looks to the left to avoid saturating the area of ​​the image around the light. This would cause the overall image brightness to spike and sag, which could be unpleasant for the user. Depending on how often the user looks in that direction, ambient-aware exposure can do one of three things. Some implementations reduce the scene's dynamic range when the user looks forward to leave clearance (e.g., a reserved portion of the range) for the area around the light when the user looks to the left. This might be a good option if the user frequently looks to the left and the device considers the content on the left important. Some implementations maintain full dynamic range when the user looks forward and do not adjust the exposure when they look to the left. This would be a good decision if the user frequently looks to the left and the content around the light is considered unimportant. Some implementations maintain full dynamic range and adjust exposure differently when the user looks forward and to the left. This is a good option if the user doesn't often look to the left and this is a new situation.

[0046] Some implementations utilize cost function optimization to determine optimal camera parameters. As an example, this might involve increasing exposure time to reduce noise, increasing motion blur, and / or increasing saturation areas. The cost function can be configured to consider the aversion to motion blur (B), noise (N), and / or highlight clipping (C). The cost function can be generated based on various parameters and information, such as user learning, environment, and scene semantics. An exemplary cost function is: Cost = b B + n N + c C, where b, n, and c are tuning weights and B, N, and C are insensitivity functions of exposure (and other parameters). The optimal exposure can be determined, for example, by argmin_on_exposure(cost).

[0047] Some implementations perform mitigation for unmapped environments, such as ignoring certain areas by using virtual windows to block pass-through or cropping light sources. Other implementations ignore certain areas within the environment, such as areas with significant (e.g., exceeding a threshold amount) variation corresponding to a TV, screen, etc. Still other implementations utilize camera-defined stable areas, for example, based on 3D position, device orientation, and / or image statistics from sensors and the detection process.

[0048] Some specific implementations adjust camera color parameters to control room color temperature, support room transitions, mitigate unmapped environments, and consider skin and / or skin tone stability, etc.

[0049] Some specific implementations utilize environmental modeling (e.g., box 320). Such modeling can utilize one or more mapping processes. Exemplary mapping processes can use various contextual and sensor sources to obtain information. Such data can be provided to the front-end encoder. Time sources can provide time / date information. Algorithm sources can provide scene semantics, 2D / 3D object pose, and / or lighting estimates. ALS can provide brightness / color information. One or more IMUs can provide acceleration / gyroscope information. A flicker sensor (or other light sensor) can provide flicker or other light attribute information. An RGB sensor can provide pass-through image data. A grayscale sensor can provide grayscale image data of the environment.

[0050] Context- and sensor-based information can be used, for example, by a front-end encoder to generate a spatiotemporal map of the environment. Such a spatiotemporal map can be localized. If it corresponds to a known scene, the information is used for updates, such as updating an existing map. If it corresponds to an unknown scene, new embeddings, sensor information (e.g., grayscale sensor data, etc.) can be used for mapping, such as VIO / SLAM mapping, and 3D device pose and / or feature maps (e.g., sparse feature maps) can be provided. The spatiotemporal map and related information can be provided to a back-end encoder, which can use this map for display control, camera control, system control, algorithm control, etc., to address image quality issues such as brightness, color, exposure, flicker, tone mapping, sharpening, noise reduction, motion blur, etc.

[0051] In one example, a user enters a new room and the user's device cannot locate itself, so the device operates in a continuous control loop. As the user looks around the room, new lighting observations are added to the VIO map (e.g., the addition of 3D pose markers). Once enough samples have been added to the lighting map, the device changes its update control loop (e.g., to use less frequent updates).

[0052] In one example, a user enters the room a second time or at a subsequent time, where the room has an existing light / VIO map. The device uses sensor data to locate itself in the 3D map (e.g., VIO) and obtain lighting, surface, or other environmental information. This may involve finding the nearest-neighbor stored key poses and locating lighting or surface observations from that viewpoint. The camera control loop may run with additional context to optimize the system, for example, based on determining the presence of an existing flicker source from that angle.

[0053] Figure 4 This is a flowchart illustrating an exemplary method 400 for providing pass-through video based on adjusting camera parameters. In some specific implementations, method 400 is provided by a device (e.g.,Figure 1 The method 400 may be executed by an electronic device 120. The method 400 may be executed using an electronic device or multiple devices communicating with each other. In some embodiments, the method 400 is executed by a processing logic component, including hardware, firmware, software, or a combination thereof. In some embodiments, the method 400 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory). The method 400 may be executed at a head-mounted device (HMD) having a processor and one or more outward-facing cameras (e.g., one or more left-eye outward-facing cameras associated with a left-eye viewpoint and / or one or more right-eye outward-facing cameras associated with a right-eye viewpoint).

[0054] At box 410, method 400 involves capturing sensor data corresponding to the physical environment via one or more sensors. This may involve capturing image data, capturing depth data, capturing motion data, etc.

[0055] At box 420, method 400 relates to determining environmental characteristics based on modeling the physical environment using sensor data captured via one or more sensors. This may involve determining ambient light source optical characteristics, ambient surface optical characteristics, a 3D mapping of the environment, user behavior within the environment, predictions of the optical characteristics of light entering a camera, etc. Determining environmental characteristics may include detecting ambient light source optical characteristics, including brightness, color, temporal brightness, flicker distribution, physical techniques, spatial emission distribution, or light source classification.

[0056] The optical characteristics of ambient light sources are determined by: analyzing ALS data received over time; analyzing flicker sensor data received over time; analyzing spatial resolution data; analyzing spatial ALS data; analyzing ray tracing data; accessing a database of known light sources; accessing smart home light data; detecting displays or monitors; detecting windows; accessing weather data; and / or evaluating glare or flare data.

[0057] Determining environmental characteristics may include detecting the optical properties of environmental surfaces, such as color, specular properties, diffuse properties, reflective properties, or transparency properties. Environmental surface optical properties can be determined by analyzing image data of the physical environment; analyzing ALS data; and / or analyzing ray tracing data.

[0058] Modeling a physical environment may include generating a 3D mapping of the physical environment based on: inertial measurement unit (IMU) data; camera image data; visual inertial odometry (VIO) data; and / or scene understanding processes. Modeling a physical environment may include: using motion tracking data to generate 3D poses of light sources in the physical environment; and / or generating 3D poses of surfaces in the physical environment. Modeling may involve determining far-field light background, for example, for more uncertain cases. Modeling may be based on: analyzing ALS data received over time; analyzing scintillation sensor data received over time; analyzing spatial resolution data; analyzing spatial ALS data; and / or analyzing ray tracing data. Modeling a physical environment may include room-based classification or spatially separated classification of the physical environment. Modeling a physical environment may be based on: previously obtained image statistics corresponding to one or more viewpoints within the physical environment; and / or scaling historical data based on the proportion of light in the physical environment corresponding to natural light. Modeling a physical environment may include analyzing one or more images to determine image statistics or semantics corresponding to the types of objects depicted in the images. Modeling the physical environment can include identifying the location of one or more light sources based on data from one or more ambient light sensors or flicker sensors.

[0059] Determining environmental characteristics by modeling the physical environment can include predicting the optical properties of light entering the image sensor in one or more plausible poses. One or more plausible poses are determined based on a 3D mapping of the physical environment, a user behavior model, or a time frame. Predicting the optical properties of light entering the image sensor can include ray tracing based on the image sensor's physical design, focal length, aperture size, field of view, vignetting, responsivity, spectral quantum efficiency (QE) of different color channels, transmittance, or sensor timing or readout time. Predicting the optical properties of light entering the image sensor can also include image sensor calibration; far-field light mapping; occlusion mitigation or illusion; far-field light mapping; flicker distribution blending; image sensor shutter simulation; and / or chromatographic blending using spectral responsivity.

[0060] Camera parameters can be determined based on cost function-based optimization.

[0061] Determining camera parameters can be based on mitigation of unmapped environments. Determining camera parameters may include determining one or more areas where camera parameters will be stabilized. Camera parameters can be determined based on room color temperature, room transition treatment, or skin tone stabilization.

[0062] At box 430, based on environmental characteristics, method 400 involves determining camera parameters for an image captured via an image sensor. Camera parameters include, but are not limited to, exposure, gain, tone mapping, color balance, noise reduction, sharpness enhancement, etc. Camera characteristics and / or human preferences may also be used when determining / adjusting camera parameters.

[0063] At box 440, method 400 relates to providing (e.g., capturing, modifying, supplementing, displaying, etc.) pass-through video, including images, of the physical environment based on determined camera parameters. The pass-through may be provided based on a user behavior model. Such a user behavior model may be based on detecting whether the user is approximately stationary, moving, transitioning to a different location, or on a mobile platform. The user behavior model may be based on: historical user pose data; current user pose data; eye tracking; and / or identifying one or more currently executing applications.

[0064] Providing pass-through video may take into account motion blur, flicker visibility, noise visibility, dynamic range, and brightness stability.

[0065] The provision of pass-through video can be further based on camera characteristics and / or user preferences.

[0066] Pass-through video can be provided based on color temperature stabilization over time. Pass-through video can also be provided based on dynamic range, image brightness stability, motion blur visibility, flicker visibility, stripe visibility, and noise visibility.

[0067] Figure 5 This is a block diagram illustrating exemplary components of an electronic device 120 configured according to some specific embodiments. Although certain specific features are illustrated, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some specific implementations, electronic device 120 includes one or more processing units 802 (e.g., DSP, microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 806, one or more communication interfaces 808 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 810, one or more displays 812, one or more internal and / or external image sensor systems 814, memory 820, and one or more communication buses 804 for interconnecting these components and various other components.

[0068] In some embodiments, one or more communication buses 804 include circuitry for interconnecting system components and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 806 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.), etc.

[0069] In some embodiments, one or more displays 812 are configured to present a view of a physical or graphical environment (e.g., a 3D environment) to a user. In some embodiments, one or more displays 812 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 812 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. In one example, electronic device 120 includes a single display. In another example, electronic device 120 includes displays for each of the user's eyes.

[0070] In some embodiments, one or more image sensor systems 814 are configured to acquire image data corresponding to at least a portion of the physical environment 100. For example, one or more image sensor systems 814 include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, etc. In various embodiments, one or more image sensor systems 814 also include an illumination source emitting light, such as a flash. In various embodiments, one or more image sensor systems 814 also include an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data. In various embodiments, the one or more image sensor systems include an optical image stabilization (OIS) system configured to facilitate optical image stabilization according to one or more techniques disclosed herein.

[0071] Memory 820 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 820 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 820 optionally includes one or more storage devices remotely located to one or more processing units 802. Memory 820 includes a non-transitory computer-readable storage medium.

[0072] In some embodiments, memory 820 or a non-transitory computer-readable storage medium of memory 820 stores an optional operating system 830 and one or more instruction sets 840. Operating system 830 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 840 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 840 is software executable by one or more processing units 802 to perform one or more of the techniques described herein.

[0073] Instruction set 840 includes registration instruction set 842, adaptation instruction set 844, and rendering instruction set 846. Instruction set 840 may be embodied as a single software executable or multiple software executables. In alternative embodiments, the software is replaced by dedicated hardware (e.g., silicon). In some embodiments, environment instruction set 842 may be executed by processing unit 802 (e.g., CPU) to create, update, or use 3D mappings or other environmental features described herein. In some embodiments, adaptation instruction set 844 may be executed by processing unit 802 (e.g., CPU) to determine parameters of one or more cameras of electronic device 120 to improve image capture as described herein. In some embodiments, rendering instruction set 846 may be executed by processing unit 802 (e.g., CPU) to render captured video content (e.g., as one or more live video feeds or other pass-through video) as described herein. For these purposes, in various embodiments, these units include instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0074] Although instruction set 840 is shown as residing on a single device, it should be understood that in other specific implementations, any combination of elements may reside in separate computing devices. Furthermore, Figure 4This is intended more as a functional description of various features present in a particular specific implementation than as a structural diagram of the specific implementation described herein. As will be appreciated by those skilled in the art, the items shown individually can be combined, and some items can be separated. The actual number of instruction sets and how features are allocated therein will vary depending on the specific implementation and may depend in part on the specific combination of hardware, software, and / or firmware chosen for that particular implementation.

[0075] This document sets forth numerous specific details to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to a person of ordinary skill have not been described in detail so as not to obscure the claimed subject matter.

[0076] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “calculating,” “determining,” and “identifying” refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, which manipulate or convert data representing physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.

[0077] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein may be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.

[0078] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the above examples can be changed; for example, the boxes can be reordered, grouped, or divided into sub-boxes. Some boxes or procedures can be executed in parallel.

[0079] The use of "applies to" or "configured to" in this document implies open and inclusive language, which does not exclude applicability to or configuration for performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, as processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to be restrictive.

[0080] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.

[0081] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and in the appended claims, the singular forms “a,” “an,” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “or,” as used herein, refers to and covers any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprising” or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, objects, or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, objects, components, or groups thereof.

[0082] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" can be interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when it is detected that the prerequisite is true" or "in response to detection" that the prerequisite is true, depending on the context.

[0083] The foregoing specific embodiments and summary of the invention should be understood as illustrative and exemplary in every respect, and not restrictive. Furthermore, the scope of the invention disclosed herein is determined not only by the detailed description of the illustrative embodiments, but also by the full extent permitted by patent law. It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention, and various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention.

Claims

1. A method, the method comprising: In an electronic device having a processor and one or more sensors, the one or more sensors include an image sensor: Sensor data corresponding to the physical environment is captured via the one or more sensors; Environmental characteristics are determined by modeling the physical environment based on sensor data captured by the one or more sensors. Based on the environmental characteristics, determine the camera parameters of the image captured by the image sensor; as well as Based on the determined camera parameters, a pass-through video including the image of the physical environment is provided.

2. The method according to claim 1, wherein determining the environmental characteristics includes detecting the optical characteristics of an ambient light source, the optical characteristics of the ambient light source including brightness, color, temporal brightness, flicker distribution, physical technique, spatial emission distribution, or light source classification.

3. The method according to claim 2, wherein the optical characteristics of the ambient light source are determined by the following means: Analyze ambient light sensor (ALS) data received over time; Analyze the scintillation sensor data received over time; Analyze spatial resolution data; Analyze spatial ALS data; Analyze ray tracing data; Access a database of known light sources; Access to smart home optical data; Monitor or display detection; Window detection; Access weather data; or Assess glare or flare data.

4. The method of claim 1, wherein determining the environmental characteristics includes detecting the optical characteristics of an environmental surface, the environmental surface optical characteristics including color, specular characteristics, diffuse characteristics, reflective characteristics, or transparency characteristics.

5. The method of claim 4, wherein the optical properties of the ambient surface are determined by: Analyze the image data of the physical environment; Analyze ambient light sensor (ALS) data; or Analyze ray tracing data.

6. The method according to any one of claims 1 to 5, wherein modeling the physical environment comprises generating a 3D mapping of the physical environment based on: Inertial Measurement Unit (IMU) data; Camera image data; Visual inertial odometry (VIO) data; Scene understanding process; or Simultaneous Localization and Mapping (SLAM).

7. The method according to any one of claims 1 to 6, wherein modeling the physical environment comprises using motion tracking data to: Generate the 3D pose of the light source in the physical environment; and Generate the 3D pose of the surface in the physical environment.

8. The method of claim 7, wherein the modeling is based on: Analyze ambient light sensor (ALS) data received over time; Analyze the scintillation sensor data received over time; Analyze spatial resolution data; Analyze spatial ALS data; or Analyze ray tracing data.

9. The method according to any one of claims 1 to 8, wherein modeling the physical environment comprises: Room-based classification of the physical environment or spatial separation classification of the physical environment, or Far-field light map modeling.

10. The method according to any one of claims 1 to 9, wherein modeling the physical environment is based on: Previously obtained image statistics corresponding to one or more viewpoints within the physical environment; and Historical data is scaled based on the proportion of light corresponding to natural light in the physical environment.

11. The method according to any one of claims 1 to 10, wherein the pass-through is provided based on a user behavior model, the user behavior model being based on: Detect whether the user is approximately stationary, moving, changing to a different location, or on a mobile platform.

12. The method of claim 11, wherein the user behavior model is based on: Historical user gesture data; Real-time user posture data; Eye tracking; or Identify one or more applications currently running.

13. The method of any one of claims 1 to 12, wherein determining the environmental characteristics based on modeling the physical environment includes predicting the optical characteristics of light entering the image sensor in one or more reasonable orientations within the physical environment.

14. The method of claim 13, wherein the one or more reasonable poses are determined based on a 3D mapping of the physical environment, a user behavior model, or a time range.

15. The method of claim 13, wherein predicting the optical properties of the light entering the image sensor comprises: Ray tracing is performed based on the physical design of the image sensor, focal length, aperture size, field of view, vignetting, responsivity, spectral quantum efficiency of different color channels, transmittance, sensor timing, sensor readout time, or far-field light map modeling. Image sensor calibration; Obscuring may reduce or alleviate hallucinations; Far-field light pattern; Scintillation distribution mixing; Image sensor shutter simulation; or Chromatographic mixing using spectral response.

16. The method according to any one of claims 1 to 15, wherein providing the transparent video includes taking into account motion blur, flicker visibility, noise visibility, dynamic range, and brightness stability.

17. The method of any one of claims 1 to 16, wherein the camera parameters are determined based on cost function-based optimization.

18. The method according to any one of claims 1 to 17, wherein the camera parameters are determined based on mitigation taking into account the unmapped environment.

19. The method according to any one of claims 1 to 18, wherein determining the camera parameters includes determining one or more regions where the camera parameters will be stabilized.

20. The method according to any one of claims 1 to 19, wherein the camera parameters are determined based on room color temperature, room transition treatment, or skin tone stabilization.

21. The method according to any one of claims 1 to 20, wherein the camera parameters are exposure, gain, tone mapping, or color balance.

22. The method according to any one of claims 1 to 20, wherein the camera parameters are noise reduction parameters or sharpness enhancement parameters.

23. The method according to any one of claims 1 to 22, wherein the transmission video is provided further based on camera characteristics.

24. The method according to any one of claims 1 to 23, wherein the transmission video is provided further based on user preferences.

25. The method according to any one of claims 1 to 24, wherein modeling the physical environment comprises analyzing one or more images to determine image statistics or semantics corresponding to the types of objects depicted in the images.

26. The method according to any one of claims 1 to 25, wherein modeling the physical environment includes identifying the location of one or more light sources based on data from one or more ambient light sensors or flicker sensors.

27. The method according to any one of claims 1 to 26, wherein the transparent video is provided based on stabilizing the color temperature over time.

28. The method according to any one of claims 1 to 27, wherein the transparent video is provided based on dynamic range, image brightness stability, motion blur visibility, flicker visibility, stripe visibility, and noise visibility.

29. A head-mounted device (HMD), the head-mounted device (HMD) comprising: Non-transitory computer-readable storage medium; as well as One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the system to perform an operation corresponding to any one of the methods according to claims 1 to 28.

30. A non-transitory computer-readable storage medium storing program instructions that are executable by a computer to perform operations including the method according to any one of claims 1 to 28.