Machine Learning Inference for Gravity-Aligned Images

By detecting device orientation and movement changes on mobile devices, rotating images to achieve gravity alignment, and inputting them into machine learning models, the problems of image alignment and lighting estimation in AR devices are solved, improving the authenticity and accuracy of the AR experience.

CN113016008BActive Publication Date: 2025-06-10GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980004400.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-21
Publication Date
2025-06-10
Estimated Expiration
2039-10-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of image alignment and lighting estimation in augmented reality (AR) devices, especially when capturing images on mobile devices, where rotation angles and lighting conditions of images are difficult to accurately capture.

Method used

By detecting changes in orientation and movement of the device at the processor, the rotation angle of the image is determined and the image is rotated to an appropriate angle to generate a gravity-aligned image. The image is provided as input to the machine learning model to trigger the relevant AR features and lighting estimation.

Benefits of technology

Accurate alignment and lighting estimation of captured images on mobile devices is achieved, improving the authenticity and accuracy of the AR experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113016008B_ABST
    Figure CN113016008B_ABST
Patent Text Reader

Abstract

Describes a system, method, and computer program product that includes: obtaining a first image at a processor from an image capture device on a computing device; using the processor and at least one sensor to detect a device orientation of the computing device associated with the capture of the first image; determining a rotation angle for rotating the first image based on the device orientation and a tracking stack associated with the computing device; rotating the first image to the rotation angle to generate a second image; and generating a neural network-based estimate associated with the first image and the second image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to virtual reality (VR) and / or augmented reality (AR) experiences and aspects of determining the alignment of images captured by a mobile device. Background Art

[0002] An augmented reality (AR) device is configured to display one or more images and / or objects over a physical space to provide an enhanced view of the physical space. Objects in the enhanced view can be tracked by a tracking system that detects and measures coordinate changes of a moving object. Machine learning techniques can also be used to track moving objects in AR and predict where the object might move throughout the AR scene. Summary of the Invention

[0003] A system of one or more computers can be configured to perform particular operations or actions by virtue of software, firmware, hardware, or a combination thereof installed on the system, the software, firmware, hardware, or a combination thereof causing the system to perform the actions in operation. One or more computer programs can be configured to perform particular operations or actions by virtue of instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.

[0004] In a first general aspect, a computer program product is described that includes: obtaining a first image at a processor from an image capture device on a computing device; using the processor and at least one sensor to detect a device orientation of the computing device associated with the capture of the first image; determining a rotation angle for rotating the first image based on the device orientation and a tracking stack associated with the computing device; rotating the first image to the rotation angle to generate a second image; and providing the second image as gravity-aligned content to at least one machine learning model associated with the computing device to trigger at least one augmented reality (AR) feature associated with the first image.

[0005] Particular embodiments of the computer program product can include any or all of the following features. For example, the at least one sensor can include or have access to a tracking stack corresponding to trackable features captured in the first image. In some embodiments, the at least one sensor is an inertial measurement unit (IMU) of the computing device, and the tracking stack is associated with changes detected at the computing device. In some embodiments, the second image is generated to match a capture orientation associated with previously captured training data. In some embodiments, the first image is a live camera image feed that generates a plurality of images, and the plurality of images are continuously aligned based on detected movement associated with the tracking stack. In some embodiments, the computer program product can include the steps of: using the plurality of images to generate an input for a neural network, the input including a generated lateral image based on a captured longitudinal orientation image.

[0006] In a second general aspect, a computer-implemented method is described. The method can include obtaining, at a processor, a first image from an image capture device included on a computing device; using the processor and at least one sensor to detect a device orientation of the computing device and associated with the capture of the first image; based on the orientation, determining a rotation angle for rotating the first image; rotating the first image to the rotation angle to generate a second image; and using the processor to provide the second image to at least one neural network to generate a lighting estimate of the first image based on the second image.

[0007] Particular implementations of the computer-implemented method can include any or all of the following features. For example, the detected device orientation can occur during an augmented reality (AR) session operating on the computing device. In some implementations, the lighting estimate is rotated in the reverse of the rotation angle. In some implementations, the first image is rendered in the AR session on the computing device and using the rotated lighting estimate. In some implementations, AR content is generated and rendered as an overlay on the first image using the rotated lighting estimate.

[0008] In some implementations, the second image is generated to match a capture orientation associated with previously captured training data, and wherein the second image is used to generate a horizontally oriented lighting estimate. In some implementations, the rotation angle is used to align the first image to generate a gravity-aligned second image. In some implementations, the at least one sensor includes a tracking stack associated with tracking features captured in a live camera image feed. In some implementations, the at least one sensor is an inertial measurement unit (IMU) of the computing device, and the movement change represents the tracking stack associated with the IMU and the computing device.

[0009] In some implementations, the first image is a live camera image feed that generates a plurality of images, and the plurality of images are continuously aligned based on the detected movement change associated with the computing device.

[0010] In a third general aspect, a system is described that includes an image capture device associated with a computing device, at least one processor, and a memory storing instructions that, when executed by the at least one processor, cause the system to: obtain, at the processor, a first image from the image capture device; use the processor and at least one sensor to detect a device orientation of the computing device and associated with the capture of the first image; use the processor and at least one sensor to detect a movement change associated with the computing device; based on the orientation and the movement change, determine a rotation angle for rotating the first image and rotate the first image to the rotation angle to generate a second image. The instructions can also generate a face tracking estimate of the first image based on the second image and according to the movement change.

[0011] Certain embodiments of the system may include any one or all of the following features. For example, the image capture device may be a front-facing image capture device or a rear-facing image capture device of a computing device. In some embodiments, a first image is captured using the front-facing image capture device, and the first image includes at least one face rotated at a rotation angle to generate a second image, and the second image is aligned with the eyes associated with the face, and the eyes are located above the mouth associated with the face. In some embodiments, the movement change is associated with an augmented reality (AR) session operating on the computing device, the face tracking estimates a reverse rotation at the rotation angle, and the first image is rendered in the AR session on the computing device, and the second image is provided as gravity-aligned content to at least one machine learning model associated with the computing device to trigger an augmented reality (AR) experience associated with the first image and the rotated face tracking estimate.

[0012] In some embodiments, the second image is used as an input to a neural network to generate laterally oriented content in which there is at least one gravity-aligned face. In some embodiments, the first image is a real-time camera image feed that generates multiple images, and the multiple images are continuously aligned based on the detected movement change associated with the computing device. In some embodiments, the second image is generated to match the capture orientation associated with previously captured training data, and the second image is used to generate a laterally oriented face tracking estimate. In some embodiments, at least one sensor is an inertial measurement unit (IMU) of the computing device, and the movement change represents a tracking stack associated with the IMU and the computing device.

[0013] Embodiments of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0014] Details of one or more embodiments are set forth in the accompanying drawings and the following description. Other features will be apparent from the specification, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Shows an example augmented reality (AR) scene captured using various lighting characteristics according to an example embodiment.

[0016] Figure 2 Is a block diagram of an example computing device having a framework for determining what data is used to estimate the screen orientation of an image presented in an AR experience according to an example embodiment.

[0017] Figures 3A to 3B Shows an example of generating laterally oriented image content from captured image content according to an example embodiment.

[0018] Figure 4 is an example showing an orientation diagram of a computing device and a transformation of such an orientation for performing gravity-based alignment according to an example implementation.

[0019] Figures 5A-5D shows an example of providing gravity-based alignment for face tracking in an AR experience according to an example embodiment.

[0020] Figure 6 is an example process for inferring gravity alignment of image content according to an example embodiment.

[0021] Figure 7 shows examples of a computer device and a mobile computer device that can be used with the techniques described herein.

[0022] The use of like or identical reference numerals in the various drawings is intended to indicate the presence of like or identical elements or features. Detailed Description

[0023] A machine learning model utilizing a neural network can receive an image as input in order to provide any number of output types. One such example output includes image classification, where the machine learning model is trained to indicate the category associated with an object in the image. Another example includes object detection, where the machine learning model is trained to output the specific location of an object in the image. Yet another example includes image-to-image class transformation, where the input is an image and the output is a stylized version of the original input image. Other examples can include, but are not limited to, face feature tracking for augmented reality (AR) (e.g., localizing 2D face features from an input image or video), face mesh generation for AR (e.g., inferring a 3D face mesh from an input image or video), hand, body, and / or pose tracking, and lighting estimation for AR (e.g., estimating scene lighting from an input image for realistically rendering virtual resources into an image or video feed).

[0024] Generally, a machine learning model described herein (e.g., using a neural network) can be configured to receive an image that is gravity - aligned (e.g., normalized to indicate upward alignment), where the image content expected to be near the top of the captured image is actually near the top of the image. For example, if a camera device is used in an upright - orientation configuration and / or to capture an image, it is typically expected that the sky or ceiling in the image is near the top of the image. Using a gravity - aligned image and / or object as the input to a neural network can ensure that directional content (e.g., heads, faces, sky, ceiling, etc.) is treated as being orthonormalized with respect to other images (e.g., orthogonal to the ground (floor)). If, for example, a particular element is corrected and / or confirmed as being upright (e.g., gravity - aligned in an upright position relative to the ground / bottom of the image) before being used in a neural network, any lighting and / or tracking associated with such gravity - based alignment can be correctly estimated using the neural network.

[0025] The systems and techniques described herein can provide the advantage of correcting image content that is in a non - gravity - aligned orientation. For example, the systems described herein can detect that a user is accessing an AR session on a computing device in a landscape orientation. As described in further detail below, such a detection can trigger a modification of specific image content to avoid providing inaccurate (e.g., incorrect) alignment content to a neural network configured to estimate the lighting (e.g., illumination) of the image content.

[0026] The techniques described herein can be used to align images for estimating and / or computing the lighting aspects of an AR scene in an AR session to ensure a realistic lighting estimate when rendering virtual content is composited into the scene. The techniques described herein can include using algorithms and neural networks that are used to determine when and how to gravity - align an image (e.g., tilt the image to ensure the correct orientation for tracking and / or lighting purposes) so that the gravity - aligned image is used when estimating the lighting and / or tracking for the image. The gravity - aligned image can be provided to a neural network, for example, to generate a lighting estimate and / or perform face tracking for a landscape - oriented AR session.

[0027] If the computing device is detected to be in an orientation other than the portrait orientation, the systems and techniques described herein can use a machine learning model (e.g., a neural network) based on the device orientation and device tracking metrics to generate realistic lighting for an AR session to infer specific aspects of a scene (e.g., an image). Additionally, the systems and techniques described herein can generate additional training data for the neural network. For example, the systems and techniques described herein can use one or more images captured in the portrait mode to generate data for estimating the lighting of an image displayed in the landscape mode. Additionally, when using such images as inputs to a neural network, gravity-based alignment can be used to generate missing and / or corrected image data to infer the upright orientation of a specific image. Gravity-based alignment can be used to ensure that a specific image is in an upright orientation when provided as an input to a neural network, e.g., regardless of the angle at which the computing device is held during image capture (using an on-board camera). For example, gravity-based alignment can be continuously performed on images (e.g., videos) from a stored or real-time camera image feed based on movement changes corresponding to the movement of the computing device.

[0028] The systems and techniques described herein can provide an improved lighting solution for AR, VR, and / or MR by computing the gravity-based alignment of the input image before providing such an image to one or more neural networks. The gravity-based alignment of the input image can ensure that directional content in the image is considered when generating realistic lighting for the image or object and / or the scene containing the image or object.

[0029] The systems and techniques described herein can provide the advantage of using sensor-measured image orientation, e.g., in a manner that a machine learning model is executed on a computing device to learn to modify image content. For example, a computing device that includes at least one image sensor (e.g., a camera, an inertial measurement unit (IMU), a tracking stack, etc.) and other sensors that are also used for tracking can be measured and provided to a machine learning model (e.g., a neural network) to produce a correctly oriented (e.g., upright) and realistic lighting for an image captured by the device. In addition to reducing the difficulty of each machine learning problem, determining how to rotate a specific image to achieve an upright input to the neural network can also enable the technique to simulate landscape-oriented training images when only portrait-oriented training images have been acquired for training.

[0030] In some embodiments, the systems and techniques described herein can incorporate orientation knowledge that indicates the upright orientation of a given input image for use in tracking a face captured in the image and tracking the content around the face. For example, for a face feature tracker, the technique can determine whether the face in the input image received at a model (e.g., a neural network model / machine learning model) is properly rotated such that the eyes are above the user's nose and mouth in the image. The model can learn the spatial relationships between different parts of the face and can be configured to be unlikely to provide a prediction where the eyes are below the mouth.

[0031] Similarly, in an AR lighting estimation example, if the input image to the model typically has a sky (e.g., in an outdoor image) or a ceiling (e.g., in an indoor image) in the upper portion of the input image, the model can be configured to be unlikely that sunlight is coming from a lower region (e.g., the lower hemisphere, the lower half, etc.), which can represent the natural occurrence and / or source of sunlight in the real world.

[0032] In some embodiments, the systems described herein can track approximate joint coordinates in a two-dimensional image (e.g., in a hand or body). The approximate joint coordinates can be used to properly align specific body parts in the image. For example, the systems described herein can ensure that the image provided to the neural network includes an upright image (e.g., the hand is above the foot, the shoulder is below the head, the knee is above the foot, etc.).

[0033] In some embodiments, the systems described herein can perform gravity alignment on images to ensure that each image is provided to the neural network in an upright orientation. Gravity alignment of imagery can assist computer vision tasks that utilize neural networks with a tracking stack associated with the computing device, such that the network is trained using upright images. Gravity alignment can be performed to benefit computer vision tasks of applications that benefit from receiving input images with upward, downward, or other directional evaluations of the images defined uniformly. For example, gravity alignment can provide computer vision benefits for any one or all of the following tasks: facial feature tracking (e.g., locating facial features in an image), face detection (e.g., locating a face in an image), body detection (e.g., locating a body in an image), body pose estimation (e.g., locating joint positions in an image), hand pose estimation and / or hand tracking (e.g., locating hand joint positions in an image), illumination estimation, surface normal estimation (e.g., estimating the surface normal for each point in an image), general object detection (i.e., an object specific to a location in an image, e.g., “find the chair”), object classification (e.g., determining if this object is a chair, etc.), semantic segmentation (e.g., determining if an image pixel represents part of a table, etc.), body segmentation (e.g., determining if an image pixel represents part of a person), head segmentation (e.g., determining if an image pixel represents part of a person's head), hand segmentation (e.g., determining if an image pixel represents part of a person's hand), monocular 3D depth estimation (e.g., determining the depth of a pixel without direct measurement, etc.).

[0034] Figure 1 FIG. 4 shows an example augmented reality (AR) scene 100 captured with various lighting characteristics according to an example embodiment. Scene 100 may be captured by a rear camera of a computing device 102 and provided by an AR application 104. In this example, user 106 may be accessing a camera mode that provides software and algorithms that enable user 106 to generate AR content and place it around the captured image (e.g., instantaneously and in real time). Computing device 102 may utilize a tracking system 108, training data 110, a neural network 112, a lighting engine 114, and facial tracking software 116 to access the AR environment and place AR content. For example, computing device 102 may use tracking system 108 to detect device orientation during capture of scene 100. For example, the detected device 102 orientation may be used to improve lighting estimation generated using lighting engine 114, training data 110, and neural network 112. Additionally, the detected device 102 orientation may be used to improve facial tracking estimation generated using facial tracking software 116.

[0035] As Figure 1As shown, user 106 of computing device 102 captures scene 100 while holding the device at a particular angle (rotation, orientation, etc.). The angle can be used to determine how to rotate a particular image when providing the particular image to neural network 112, e.g., to determine light estimation, motion estimation, face tracking estimation, etc. In this example, user 106 can capture content by twisting device 102 to the right (or left) from the vertical y-axis, as shown by arrow 120. The systems described herein can determine camera pose changes and / or device pose changes associated with the user and / or device 202 to correctly capture and render the user (in the front camera view) and any VR and / or AR content associated with the user in the camera feed (from the front camera view and / or the rear camera view). Similarly, the systems described herein can determine pose changes associated with the movement of the user or mobile device in a direction related to z-axis 122 and / or x-axis 124.

[0036] In some embodiments, e.g., the device orientation detected during the capture of scene 100 can be used with face tracking software 116 to detect an upright face for the front-facing camera on computing device 102. Tracking system 108 can determine the movement (i.e., position change) of device 102 to ascertain device orientation changes in order to determine the upright orientation (e.g., gravity-aligned) of the face depicted in the captured image.

[0037] In some embodiments, the systems and techniques described herein provide a solution for illuminating AR images and scenes using a motion tracking stack (e.g., representing changes in the movement of a device over time) and an inertial measurement unit (IMU) sensor of a computing device (e.g., a mobile phone) that is performing an AR session. For example, the systems and techniques can determine the true light of an image and scene by using such sensors and motion tracking to detect how (or whether) to rotate an image captured by one or more image sensing devices on the computing device before feeding the image to a neural network. For example, this can be performed on a live camera feed to ensure that the input image is gravity-aligned with the upward ceiling / sky in the detected image.

[0038] To train the neural network 112, multiple computing device (e.g., mobile phone) videos can be captured to provide training data. Typically, videos and / or images are captured in a portrait orientation, where the computing device is held with its base parallel to the ground, which is associated with the user capturing the video and / or image. When using the computing device to trigger real-time machine learning inference during an AR session for light estimation, the estimated light results may be inaccurate or untrue for the content in the scene when the computing device is held in a landscape orientation (e.g., the base of the computing device is perpendicular to the ground). The light engine 114, face tracking software 116, and tracking system 108 can correct such estimated light by inferring landscape-based imagery from portrait-based imagery. Additionally, the computing device 102 can trigger gravity-based alignment for the captured content to ensure that the content is correctly oriented before being provided to the neural network 112. For example, the computing device 102 can determine whether a rotation of the content is to be performed and, if so, to what extent, to ensure proper lighting and / or tracking when the content is rendered on the device 102.

[0039] Accordingly, the systems described herein can use such determination and rotation to ensure accurate training data 110 for training the neural network 112. In some embodiments, accurate training data 110 can be defined as data that matches a particular device orientation, sensor orientation, and / or image orientation. For example, the systems described herein can rely on portrait-based imagery being correctly indicated as upright. Additionally, the systems described herein can simulate landscape-based imagery and configure such imagery to be provided to the neural network 112 in the upright orientation expected by the system.

[0040] During an AR session, the AR sensing stack can be used to determine when and how to rotate the current image from the live camera feed to ensure that such image is provided to the neural network 112 as an upright image. The system can then apply an inverse rotation to the predicted light to output the estimated light aligned with the physical camera coordinates of the captured image.

[0041] Figure 2 is a block diagram of an example computing device 202 having a framework for determining what data to use to estimate the screen orientation of an image presented in an AR experience. In some embodiments, the framework can be used to determine what data to use to estimate the screen orientation for face detection and / or for generating a light estimation for an AR experience.

[0042] In operation, the systems and techniques described herein can provide a mechanism for using machine learning to estimate high dynamic range (HDR) omnidirectional (360-degree) lighting / illumination to illuminate and render virtual content into a real scene for use in an AR environment and / or other synthetic applications. The systems and techniques described herein can also determine a specific device orientation during the capture of an image and then generate aspects of lighting estimation and face tracking based on the determined device orientation to use the aspects of lighting estimation and face tracking to render the scene.

[0043] In some embodiments, system 200 can be used to generate lighting estimates for AR, VR, and / or MR environments. Generally, a computing device (e.g., a mobile device, a tablet, a laptop computer, an HMD device, AR glasses, a smartwatch, etc.) 202 can generate lighting conditions to illuminate an AR scene. Additionally, device 202 can generate an AR environment for a user of system 200 to trigger the rendering of an AR scene on device 202 or another device using the generated lighting conditions. In some embodiments, system 200 includes a computing device 202, a head-mounted display (HMD) device 204 (e.g., AR glasses, VR glasses, etc.), and an AR content source 206. A network 208 is also shown, and computing device 202 can communicate with AR content source 206 via network 208. In some embodiments, computing device 202 is a pair of AR glasses (or other HMD device).

[0044] Computing device 202 includes a memory 210, a processor component 212, a communication module 214, a sensor system 216, and a display device 218. Memory 210 can include an AR application 220, AR content 222, an image buffer 224, an image analyzer 226, a lighting engine 228, and a rendering engine 230. Computing device 202 can also include various user input devices 232, such as one or more controllers that communicate with computing device 202 using a wireless communication protocol. In some embodiments, input device 232 can include, for example, a touch input device that can receive tactile user input, a microphone that can receive audible user input, etc. Computing device 202 can also be one or more output devices 234. Output device 234 can include, for example, a display for visual output, a speaker for audio output, etc.

[0045] The computing device 202 may also include any number of sensors and / or devices in the sensor system 216. For example, the sensor system 216 may include a camera assembly 236 and a 3-DoF and / or 6-DoF tracking system 238. The tracking system 238 may include (or may have access to) for example, light sensors, IMU sensors 240, audio sensors 242, image sensors 244, distance / proximity sensors (not shown), location sensors (not shown), and / or different combinations of other sensors and / or sensors. Some of the sensors included in the sensor system 216 may provide positioning detection and tracking of the device 202. Some of the sensors in the system 216 may provide capture of images of the physical environment for display on components of the user interface of the rendered AR application 220.

[0046] The computing device 202 may also include a tracking stack 245. The tracking stack may represent the movement changes of the computing device and / or the AR session over time. In some embodiments, the tracking stack 245 may include IMU sensors 240 (e.g., gyroscopes, accelerometers, magnetometers). In some embodiments, the tracking stack 245 may perform image feature movement detection. For example, the tracking stack 245 may be used to detect motion by tracking features in an image. For example, the image may include or be associated with a plurality of trackable features, e.g., the trackable features may be tracked frame by frame in a video including the image. Camera calibration parameters (e.g., projection matrix) are generally known as part of the on-board device camera, and thus, the tracking stack 245 may use image feature movement together with other sensors to detect motion. The detected motion may be used to generate gravity-aligned images to be provided to the neural network 256, which may use such images to further learn and provide illumination, additional tracking, or other image variations.

[0047] The computing device 202 may also include face tracking software 260. The face tracking software 260 may include (or may have access to) one or more face cue detectors (not shown), smoothing algorithms, pose detection algorithms, and / or neural network 256. The face cue detector may operate on or with one or more camera assemblies 236 to determine the movement of specific facial features of the user or the positioning of the head. For example, the face tracking software 260 may detect or obtain an initial three-dimensional (3D) positioning of the computing device 202 with respect to facial features (e.g., image features) captured by one or more camera assemblies 236. For example, one or more camera assemblies 236 may be used with the software 260 to retrieve a specific positioning of the computing device 202 with respect to the facial features captured by the camera assemblies 236. In addition, the tracking system 238 may access the on-board IMU sensors 240 to detect or obtain an initial orientation associated with the computing device 202.

[0048] The face tracking software 260 can detect and / or estimate a particular computing device orientation 262 (e.g., screen orientation) of the device 202 during image capture, e.g., to detect an upright face in a scene. The computing device orientation 262 can be used to determine whether to rotate the captured images and by how much in the detected and / or estimated screen orientation.

[0049] In some embodiments, the sensor system 216 can detect the computing device 202 (e.g., a mobile phone device) orientation during the capture of an image. The detected computing device orientation 262 can be used as an input to modify the captured image content, e.g., to ensure that the captured image content is aligned upright (e.g., gravity-based alignment) before the captured image content is provided to the neural network 256 for illumination estimation (using the illumination reproduction software 250) and / or face tracking (using the face tracking software 260). The gravity alignment process can detect a particular degree of rotation of the computing device and can correct the image to provide portrait-based and landscape-based images with realistic and accurate illumination and tracking.

[0050] In some embodiments, the computing device orientation 262 can be used to generate the lateral training data 264 of the neural network 256 used by the illumination engine 228. The lateral training data 264 can be a cropped and padded version of the originally captured portrait-based image.

[0051] In some embodiments, the computing device 202 is a mobile computing device (e.g., a smart phone) that can be configured to provide or output AR content to a user via the HMD 204. For example, the computing device 202 and the HMD 204 can communicate via a wired connection (e.g., a Universal Serial Bus (USB) cable) or via a wireless communication protocol (e.g., any Wi-Fi protocol, any Bluetooth protocol, Zigbee, etc.). Additionally or alternatively, the computing device 202 is a component of the HMD 204 and can be included within the housing of the HMD 204.

[0052] The memory 210 can include one or more non-transitory computer-readable storage media. The memory 210 can store instructions and data that can be used to generate an AR environment for a user.

[0053] The processor component 212 includes one or more devices capable of executing instructions such as those stored by the memory 210 to perform various tasks associated with generating an AR, VR, and / or MR environment. For example, the processor component 212 can include a central processing unit (CPU) and / or a graphics processing unit (GPU). For example, if a GPU exists, certain image / video rendering tasks such as coloring the content based on the determined lighting parameters can be offloaded from the CPU to the GPU.

[0054] The communication module 214 includes one or more devices for communicating with other computing devices such as the AR content source 206. The communication module 214 may communicate via a wireless or wired network such as the network 208.

[0055] The IMU 240 detects the movement, motion, and / or acceleration of the computing device 202 and / or the HMD 204. The IMU 240 may include various different types of sensors such as, for example, accelerometers, gyroscopes, magnetometers, and other such sensors. The position and orientation of the HMD 204 may be detected and tracked based on data provided by the sensors included in the IMU 240. The detected position and orientation of the HMD 204 may allow the system to sequentially detect and track the user's gaze direction and head movement. Such tracking may be added to a tracking stack that may be polled by the lighting engine 228 to determine changes in the device and / or user movement and associate the associated time with such changes in movement. In some embodiments, the AR application 220 may use the sensor system 216 to determine the position and orientation of the user within the physical space and / or identify features or objects within the physical space.

[0056] The camera assembly 236 captures images and / or video of the physical space surrounding the computing device 202. The camera assembly 236 may include one or more cameras. The camera assembly 236 may also include an infrared camera.

[0057] The AR application 220 may present the AR content 222 to the user or provide the AR content 222 to the user via one or more output devices 234 of the HMD 204 and / or the computing device 202 such as a display device 218, a speaker (e.g., using the audio sensor 242), and / or other output devices (not shown). In some embodiments, the AR application 220 includes instructions stored in the memory 210 that, when executed by the processor component 212, cause the processor component 212 to perform the operations described herein. For example, the AR application 220 may generate and present an AR environment to the user based on, for example, AR content such as the AR content 222 and / or AR content received from the AR content source 206.

[0058] The AR content 222 may include AR, VR, and / or MR content, such as images or videos that may be displayed on a portion of the user's field of view in the HMD 204 or on the display 218 associated with the computing device 202 or on other display devices (not shown). For example, the AR content 222 may be generated using lighting (using the lighting engine 228) that substantially matches the physical space in which the user is located. The AR content 222 may include objects that overlay various portions of the physical space. The AR content 222 may be rendered as a flat image or a three-dimensional (3D) object. The 3D object may include one or more objects represented as a polygon mesh. The polygon mesh may be associated with various surface textures (such as colors and images). The polygon mesh may be shaded based on various lighting parameters generated by the AR content source 206 and / or the lighting engine 228.

[0059] The AR application 220 may use the image buffer 224, the image analyzer 226, the lighting engine 228, and the rendering engine 230 to generate an image for display via the HMD 204 based on the AR content 222. For example, one or more images captured by the camera component 236 may be stored in the image buffer 224. The AR application 220 may determine the location for inserting the content. For example, the AR application 220 may prompt the user to identify the location for inserting the content and may then receive user input indicating the location of the content on the screen. The AR application 220 may determine the location of the inserted content based on the user input. For example, the location of the content to be inserted may be the location indicated by the user accessing the AR experience. In some embodiments, the location is determined by mapping the location indicated by the user to a plane corresponding to a surface such as a floor or ground in the image (e.g., by finding a location on the plane below the location indicated by the user). The location may also be determined based on the location determined for the content in a previous image captured by the camera component (e.g., the AR application 220 may move the content on a surface identified within the physical space captured in the image).

[0060] The image analyzer 226 can then identify a region of the image stored in the image buffer 224 based on the determined position. The image analyzer 226 can determine one or more attributes of the region, such as luminance (or photometric value), hue, and saturation. In some embodiments, the image analyzer 226 filters the image to determine such characteristics. For example, the image analyzer 226 can apply a mipmap filter (e.g., a trilinear mipmap filter) to the image to generate a sequence of lower-resolution representations of the image. The image analyzer 226 can identify the lower-resolution representations of the image in which a single pixel or a small number of pixels correspond to the region. The characteristics of the region can then be determined from the single pixel or the small number of pixels. The lighting engine 228 can then generate one or more light sources or ambient light maps 254 based on the determined attributes. The rendering engine 230 can use the light source or ambient light map to render the inserted content or an enhanced image that includes the inserted content.

[0061] In some embodiments, the image buffer 224 is a region of the memory 210 configured to store one or more images. In some embodiments, the computing device 202 stores the images captured by the camera component 236 as textures within the image buffer 224. Alternatively or additionally, the image buffer 224 can also include a memory location integrated with the processor component 212, such as dedicated random access memory (RAM) on a GPU.

[0062] In some embodiments, the image analyzer 226, the lighting engine 228, and the rendering engine 230 can include instructions stored in the memory 210 that, when executed by the processor component 212, cause the processor component 212 to perform the operations described herein to generate an image or a series of images displayed to the user (e.g., via the HMD 204) and illuminated with lighting characteristics computed using the neural network 256 described herein.

[0063] The system 200 can include (or can access) one or more neural networks 256 (e.g., neural network 112). The neural network 256 can utilize an internal state (e.g., memory) to process an input sequence, such as a sequence of user movements and position changes when in an AR experience. In some embodiments, the neural network 256 can utilize the memory to process lighting aspects and generate a lighting estimate for the AR experience.

[0064] In some embodiments, neural network 256 can be a Recurrent Neural Network (RNN). In some embodiments, the RNN can be a deep RNN with multiple layers. For example, the RNN can include a Long Short-Term Memory (LSTM) architecture or a Gated Recurrent Unit (GRU) architecture. In some embodiments, system 200 can use the LSTM and GRU architectures based on determining which architecture reduces errors and / or latency. In some embodiments, neural network 256 can be a Convolutional Neural Network (CNN). In some embodiments, the neural network can be a deep neural network. As used herein, any number or type of neural network can be used to implement the specific illumination estimation and / or facial position of a scene.

[0065] Neural network 256 can include a detector that operates on an image to compute, for example, illumination estimation and / or facial position to model the predicted illumination and / or facial position as the face / user moves in world space. Additionally, neural network 256 can operate to compute illumination estimation and / or facial position for several future time steps. Neural network 256 can include a detector that operates on an image to compute, for example, device position and illumination variables to model the predicted illumination of a scene, for example, based on device orientation.

[0066] Neural network 256 can utilize the omnidirectional light or light probe image obtained from previous imaging, and can use such content to generate a specific ambient light map 254 (or other output images and illumination) from neural network 256.

[0067] In some embodiments, neural network 256 can directly predict a (cropped) light probe image in a two-step method where it can be a light estimation network (also known as a deep neural network, convolutional neural network, etc.) (the loss function can be the squared difference or absolute difference between the cropped input probe image and the net output), and then obtain the directional light values by solving a linear system using constrained least squares.

[0068] The captured images and associated illumination can be used to train neural network 256. The training data (e.g., captured images) can include Low Dynamic Range (LDR) images of one or more light detectors (not shown) with measured or known Bidirectional Reflectance Distribution Function (BRDF) under various (e.g., different) illumination conditions. The appearance of the gray sphere is a convolved version of the ambient illumination. By solving a linear system, the detector image can be further processed into High Dynamic Range (HDR) illumination coefficients. In some embodiments, the type of training data that can be used is a general LDR panorama, where even more is available.

[0069] In general, any number of lighting representations can be used for real-time graphics applications. In some embodiments, for example, ambient light can be used to evaluate and support ambient light estimation for AR development environments. In some embodiments, for example, directional light can be used to evaluate and work in conjunction with shadow mapping and approximations of primary and distant light sources (e.g., the sun). In some embodiments, for example, ambient light mapping can be used. It stores direct 360-degree lighting information. Several typical parameterizations include cube mapping, hexahedron, equirectangular mapping, or orthogonal projection can be used. In some embodiments, spherical harmonics can be used, for example, to model low-frequency lighting and serve as precomputed radiative transfer for fast integration.

[0070] Device 202 can use lighting engine 228 to generate one or more light sources for AR, VR, and / or MR environments. Lighting engine 228 includes lighting reproduction software 250 that can utilize and / or generate HDR lighting estimator 252, ambient light map 254, and neural network 256. Lighting reproduction software 250 can be executed locally on computing device 202, remotely on one or more remote computer systems (e.g., third-party provider server systems accessible via network 208), cloud networks, or a combination of one or more of each of the foregoing.

[0071] Lighting reproduction software 250 can present, for example, a user interface (UI) for displaying relevant information such as controls, calculations, and images on display device 218 of computing device 202. Lighting reproduction software 250 is configured to analyze, process, and manipulate data generated by the lighting estimation techniques described herein. Lighting reproduction software 250 can be implemented to automatically calculate, select, estimate, or control various aspects of the disclosed lighting estimation methods, such as functions for photographing color charts and / or disposing of or generating ambient light map 254.

[0072] Neural network 256 can represent a light estimation network that is trained to estimate HDR lighting from at least one LDR background image (not shown) using HDR lighting estimator 252. For example, the background image can be from the camera view of computing device 202. In some embodiments, as described in detail below, training examples can include background images, images of light detectors (e.g., spherical) in the same environment, and the bidirectional reflectance distribution function (BRDF) of the light detectors.

[0073] Figure 2The framework shown in FIG. supports training one or more of neural network 256 using multiple photodetectors (not shown) with different materials (e.g., bright, dim, etc. photodetector materials). The bright photodetector materials capture high-frequency information, which may include cropped pixel values in an image. The dimmer photodetector materials capture low information without any cropping. In some embodiments, these two datasets may be complementary to each other such that the neural network 256 can estimate HDR illumination without HDR training data.

[0074] The AR application 220 may update the AR environment based on inputs received from the camera assembly 236, the IMU 240, and / or other components of the sensor system 216. For example, the IMU 240 may detect the movement, motion, and / or acceleration of the computing device 202 and / or the HMD 204. The IMU 240 may include various different types of sensors, such as, for example, accelerometers, gyroscopes, magnetometers, and other such sensors. The positioning and orientation of the HMD 204 may be detected and tracked based on the data provided by the sensors included in the IMU 240. The detected positioning and orientation of the HMD 204 may allow the system to sequentially detect and track the positioning and orientation of the user within the physical space. Based on the detected positioning and orientation, the AR application 220 may update the AR environment to reflect the changed orientation and / or positioning of the user within the environment.

[0075] Although the computing device 202 and the HMD 204 are shown as separate devices in Figure 2 FIG., in some embodiments, the computing device 202 may include the HMD 204. In some embodiments, the computing device 202 communicates with the HMD 204 via a wired (e.g., cable) connection and / or via a wireless connection. For example, the computing device 202 may transmit video signals and / or audio signals to the HMD 204 for display to the user, and the HMD 204 may transmit motion, positioning, and / or orientation information to the computing device 202.

[0076] The AR content source 206 may generate and output AR content, which may be distributed or sent via the network 208 to one or more computing devices, such as the computing device 202. In some embodiments, the AR content 222 includes three-dimensional scenes and / or images. Additionally, the AR content 222 may include audio / video signals streamed or distributed to one or more computing devices. The AR content 222 may also include all or a portion of the AR application 220 that is executed on the computing device 202 to generate 3D scenes, audio signals, and / or video signals.

[0077] The network 208 can be the Internet, a local area network (LAN), a wireless local area network (WLAN), and / or any other network. For example, the computing device 202 can receive audio / video signals via the network 208, which can be provided as part of the AR content in the illustrative example embodiments.

[0078] The AR, VR, and / or MR systems described herein can include systems that insert computer-generated content into a user's perception of the physical space around the user. The computer-generated content can include tags, text information, images, sprites, and three-dimensional entities. In some embodiments, the content is inserted for entertainment, educational, or informational purposes.

[0079] Example AR, VR, and / or MR systems are portable electronic devices, such as smartphones, which include a camera and a display device. The portable electronic device can use the camera to capture an image and display the image on the display device, which includes computer-generated content superimposed on the image captured by the camera.

[0080] Another example AR, VR, and / or MR system includes a head-mounted display (HMD) worn by a user. The HMD includes a display device located in front of the user's eyes. For example, the HMD may block the user's entire field of view so that the user can only see the content displayed by the display device. In some examples, the display device is configured to display two different images, one for each of the user's eyes. For example, at least some of the content in one image may be slightly offset relative to the same content in the other image, thereby creating a perception of a three-dimensional scene due to parallax. In some embodiments, the HMD includes a chamber in which a portable electronic device, such as a smartphone, can be placed to allow viewing of the display device of the portable electronic device through the HMD.

[0081] Another example AR, VR, and / or MR system includes an HMD that allows the user to see the physical space while wearing the HMD. The HMD can include a microdisplay device that can display computer-generated content superimposed on the user's field of view. For example, the HMD can include goggles with a combiner that is at least partially transparent, which allows light from the physical space to reach the user's eyes while also reflecting the image displayed by the microdisplay device towards the user's eyes.

[0082] Although many of the examples described herein relate to AR systems that insert and / or synthesize visual content into an AR environment, the techniques described herein can also be used to insert content in other systems. For example, the techniques described herein can be used to insert content into an image or video.

[0083] Typically, systems and techniques can be hosted on a mobile electronic device such as computing device 202. However, an accommodation for one or more cameras and / or image sensors or other electronic devices associated with one or more cameras and / or image sensors can be used to perform the techniques described herein. In some embodiments, a tracking sensor and an associated tracking stack can also be used as an input to perform a light estimation technique.

[0084] Figures 3A to 3B An example of generating landscape image content from captured image content according to an example embodiment is shown. For example, the generated landscape image content can be provided as landscape training data 264 to neural network 256.

[0085] Figure 3A An image 302A captured in a portrait-based orientation is shown. Image 302A can be a single image / scene or can be a video. Multiple different reflective spheres 304, 306, and 308 can be used during the capture of image 302A. Image 302A can be processed by computing device 202 to determine device orientation, determine image orientation, and generate an output with adjustments for differences in such orientations. The resulting output can be provided to a neural network to generate a light estimation and / or face tracking task. In some embodiments, for example, such an output can be used as landscape training data 264 for neural network 256. Generally, the training data for neural network 256 can include captured content (in portrait mode) and landscape content generated using portrait-based captured content, as described in detail below.

[0086] Landscape training data 264 can include modified versions of videos and / or images captured using various reflective spheres (e.g., spheres 304, 306, and 308) placed within the camera's field of view. Such captured content can keep the background imagery unobscured while showing different lighting cues in a single exposure using materials with multiple reflective functions. Landscape training data 264 can be used to train a deep convolutional neural network (e.g., neural network 256) to regress from unoccluded portions of an LDR background image to HDR lighting by matching LDR ground truth sphere images with those rendered with predicted lighting using image-based relighting of landscape content.

[0087] For example, if an image 302A is to be illuminated and rendered for display to a user of the computing device 202, the system can use the captured content 310A to train the neural network 256 to generate lighting and / or other tracking for realistically lighting and rendering the content in the scene on the device 202. This method can be used if the system 216 detects that the computing device 202 is oriented in alignment with the portrait-based capture mode gravity used to capture the content 310A. In such an example, the system can access the image 302A and crop the image to remove the spheres 304, 306, and 308. The remaining content 310A can be used to generate a lighting estimate of the scene accessible to the user via the device 202.

[0088] Because the lighting engine 228 and the face tracking software 260 expect to receive gravity-aligned image content (e.g., an upright orientation where the sky and / or ceiling is in the upper half of the image content, or in the case of face tracking the eyes are above the lips in the image), the system can determine that the portrait-based capture is gravity-aligned, and thus, the content 310A can be used to generate lighting estimates or tracked facial features without rotational modification.

[0089] However, if the system determines that a particular device orientation does not match a particular image content orientation, the system can correct the mismatch. For example, if the system detects that the image content is being accessed for an AR session in a landscape mode (or within a threshold angle of landscape mode), the system can adjust the image content to ensure realistic rendering of lighting estimates, tracking, and content placement within the AR environment for the AR session (e.g., the images or scenes within the AR session).

[0090] For example, if the user uses the computing device in a landscape orientation to access an AR session, the sensor system 216 and the lighting engine 228 can work together to generate landscape-based content to appropriately illuminate the content accessed in the AR session. The device 202 can modify the portrait-based captured content to generate landscape-based content. For example, the computing device 202 can use the content 310A to generate landscape training data 264 by cropping out the same spheres 304, 306, and 308, but can additionally crop (e.g., and mask) the upper portion 312 of the content to generate the content 310B. The upper portion 312 can then be filled with white, gray, black, or other colored pixels. Such a mask can ensure that the image aspect ratio is maintained for both the captured portrait image and the generated landscape image.

[0091] During inference, device 202 can retrieve the landscape image 302B and generate a new image for training. For example, device 202 can retrieve a landscape (e.g., rotated from portrait) image 302B, crop the internal portion (e.g., 310B), fill portion 310B with pixels 312, and then can send the generated image (at the same resolution as the portrait image 310A) to neural network 256.

[0092] Additionally, tracking system 238 may have detected that device 202 is in a landscape orientation during inference, and thus, when predicting the light estimation from engine 228, device 202 can provide a light prediction that is rotated approximately 90 degrees from the position aligned with the actual image sensor 244 (e.g., camera sensor). Accordingly, the sensor outputs for pre-rotating the input image, cropping out the left and right portions of the image content 302B, and leaving the predicted light estimation un-rotated back to the orientation of image sensor 244 of device 202. In operation, device 202 can utilize a sensor stack that is part of the tracking stack such that device movement and user movement can also be considered when generating lights and / or tracking updates that can be represented in the rendered scene.

[0093] In some embodiments, an input image such as image 302A can be used to generate four cropped versions that represent four different rotations in which the camera can be moved to capture content using the front or rear camera. In such an example, the tracking stack may not be used.

[0094] In some embodiments, for example, if the system determines that the device has not moved beyond a threshold level, the rotations and corrections described herein may not be performed. Similarly, if device 202 determines that the previous state of the device and / or the image is sufficient, a particular image may not be modified. Accordingly, no changes to rotate the image, move the content, and / or update the light may be performed.

[0095] However, alternatively, if device 202 determines a movement from landscape mode to portrait mode or vice versa, device 202 can trigger a reset of the last state of the device to trigger new updates to tracking and / or light based on the detected change in device orientation or other movement.

[0096] Figure 4 is an example showing a device orientation diagram of a computing device according to an example embodiment and the transformation of such an orientation for performing gravity-based alignment. The device orientation diagram includes a detected phone (e.g., computing device 202) orientation column 402, a VGA image column 404, a gravity alignment angle column 406, a VGA image column 408 after rotation, and a display rotation column 410.

[0097] Typically, system 200 (e.g., on computing device 202) can use sensors (e.g., IMU 240, image sensor 244, camera assembly 236, etc.) and / or sensor system 216 to detect phone orientation. Device orientation can be detected during content capture. For example, during the capture of image content, tracking system 238 can detect 3-DoF and / or 6-DoF device pose. The device pose and / or camera pose can be used to trigger an orientation rotation to improve the output for rendering an image captured in the detected orientation of device 202. The improved output can involve improved accuracy and rendering of the lighting estimation of the captured content.

[0098] In operation, sensor system 216 can detect computing device (e.g., mobile phone device) orientation 262 during image capture. The detected computing device orientation 262 can be used to modify the captured image content to ensure that the captured image content is aligned upward (gravity-based alignment) before providing the captured image content to neural network 256 for, e.g., lighting estimation (using lighting re-software 250) and / or face tracking (using face tracking software 260).

[0099] Sensor system 216 can detect a specific rotation of degrees of computing device 202. For example, system 216 can detect whether computing device 202 is tilted and / or rotated at 0 degrees, 90 degrees, 180 degrees, and 270 degrees from the ground plane (e.g., parallel to z-axis 122). In some embodiments, system 216 can detect changes in degrees of tilt and / or rotation at approximately ten-degree increments from the x-axis, y-axis, or z-axis and from zero to 360 degrees around any one such axis. For example, system 216 can detect that device 202 is rotated at approximately positive or negative ten degrees from the ground plane (e.g., parallel to z-axis 122) at 0 degrees, 90 degrees, 180 degrees, and 270 degrees. In some embodiments, system 216 can also determine pitch, yaw, and roll to incorporate aspects of the tilt of device 202.

[0100] As Figure 4As shown, system 216 can detect a device orientation of zero degrees, as indicated by gravity alignment element 414, where device 202 is held in an upright vertical orientation with zero (or less than 10 degrees) tilt or rotation. Here, system 216 can determine that the content being captured can be used to generate landscape training data and / or landscape-based content. Since image sensor 244 and camera assembly 236 may not be able to recognize device orientation changes, IMU 240 and / or other tracking sensors can determine the device and / or camera pose to correct the image content to an upright and gravity-aligned orientation, as indicated by, for example, gravity alignment elements 414, 418, and 420. Thus, regardless of how computing device 202 is held, system 216 can determine an upright and gravity-aligned way to generate an image (e.g., a scene) for neural network 256 to ensure accurate execution of light estimation, which provides the advantage of generating and rendering realistic lighting when rendering an image (e.g., a scene).

[0101] Generally, if the device is held during capture such that the on-board camera 403 is positioned parallel to and above the bottom edge 205 of computing device 202, the device orientation can be at zero degrees. In some embodiments, within ten degrees in the clockwise direction (rotating clockwise from the normal to edge 205), counterclockwise direction (rotating counterclockwise from the normal to edge 205), forward (rotating from the normal to edge 205), or backward (rotating from the normal to edge 205) of such a position, the device orientation can still be considered zero degrees.

[0102] Similarly, if the device is held during capture such that the on-board camera 403 is positioned parallel to and to the left of the bottom edge 205 of computing device 202, the device orientation can be detected as approximately 270 degrees. In some embodiments, within ten degrees in the clockwise direction (rotating clockwise from the normal to edge 205), counterclockwise direction (rotating counterclockwise from the normal to edge 205), forward (rotating from the normal to edge 205), or backward (rotating from the normal to edge 205) of such a position, the device orientation can still be considered 270 degrees.

[0103] Similarly, if the device is held during capture such that the on-board camera 403 is positioned parallel to and below the bottom edge 205 of computing device 102, the device orientation can be detected as approximately 180 degrees. In some embodiments, within ten degrees in the clockwise direction (rotating clockwise from the normal to edge 205), counterclockwise direction (rotating counterclockwise from the normal to edge 205), forward (rotating from the normal to the edge of edge 205), or backward (rotating from the normal to edge 205) of such a position, the device orientation can still be considered 180 degrees.

[0104] Similarly, if the device is held during capture such that the on-board camera 403 is positioned parallel to the bottom edge 205 of the computing device 102 and to the right of that position, the device orientation can be detected as approximately 90 degrees. In some embodiments, within ten degrees in the clockwise direction (rotating clockwise from the normal to the edge 205), counterclockwise direction (rotating counterclockwise from the normal to the edge 205), forward (rotating from the normal to the edge 205), or backward (rotating from the normal to the edge 205) of such a position, the device orientation can still be considered to be 90 degrees.

[0105] If the system 216 determines that the device 202 is oriented in a portrait orientation, but determines that the captured content can be used to generate landscape-based training imagery, the system can determine that the captured image 401 is to be realigned (e.g., rotated counterclockwise) by approximately 90 degrees. For example, if the system 216 indicates that landscape-based image content can be generated, the system 216 can trigger the captured image 401 to be rotated counterclockwise by approximately 90 degrees, as indicated by the gravity alignment element 418, to generate the rotated image 416.

[0106] When providing the image content to the neural network 256 for illumination estimation, the image 416 can be provided according to the phone orientation, as indicated by the gravity alignment element 420. For example, the image 416 can be rotated counterclockwise by approximately 270 degrees, and / or cropped, and provided as an input to the neural network 256 for generating the illumination estimation. After the illumination estimation is completed, the illumination engine 228 can use the illumination estimation to trigger the rendering of the image content 420, and can trigger a realignment back to the physical camera space coordinates (as shown by the phone orientation 402). In such an example, when rendered for display, the output image from the neural network can be rotated in the reverse of the gravity alignment angle 406.

[0107] In another example, if system 216 determines that device 202 is oriented in a landscape orientation, as indicated by gravity alignment element 422, and the content being captured is landscape (e.g., zero degrees and not rotated from the camera sensor, as indicated by gravity alignment element 424), then system 216 can determine that the captured image 426 will not benefit from rotation / realignment. For example, if system 216 indicates that landscape-based image content is being captured and device 202 is in a landscape orientation where the sky / ceiling is indicated as upright in the capture, then the system can retain the original capture and alignment used during capture, as indicated by gravity alignment element 428. Instead, system 216 can crop (not shown) the central portion (e.g., region of interest) of the captured content, and can zero-pad (with black pixels, white pixels, gray pixels, etc.) the top landscape-based image to ensure that the image content is the same size as an image captured using a portrait-based orientation. As described above, when providing the image content to neural network 256 for illumination estimation, the image can be provided in the captured, cropped, and padded orientation.

[0108] If system 216 determines that the orientation of device 202 is inverted from upright (e.g., the camera is at the bottom of device 202), as indicated by gravity alignment element 430, but the content can be used to generate landscape training data, then system 216 can rotate the content 90 degrees, as indicated by image 426 and gravity alignment element 434. When providing the image content to neural network 256 for illumination estimation, image 426 can be provided according to the device orientation, as indicated by gravity alignment element 436, which indicates a 90-degree counterclockwise rotation from content 432 to content 433. After completing the illumination estimation, illumination engine 228 can use the illumination estimate to trigger the rendering of the image content and can trigger a realignment back to the physical camera space coordinates (as shown by phone orientation 402). In such an example, the output image from neural network 256 can be rotated in the reverse of gravity alignment angle 406.

[0109] In another example, if system 216 determines that the orientation of device 202 is a landscape orientation rotated clockwise from an upright portrait orientation, as indicated by gravity alignment element 382, and the content 440 being captured is landscape (e.g., 180 degrees and rotated from the camera sensor, as indicated by gravity alignment element 442), then the system can determine that the captured image 440 can be used to generate landscape-based training data for neural network 256. Thus, image 440 can be rotated (e.g., realigned) approximately 180 degrees to ensure that the ceiling or sky is facing up, e.g., as indicated by image 443 and gravity alignment element 444.

[0110] Additionally, system 216 can crop (not shown) the central portion (e.g., region of interest) of the captured content 440, and can zero-pad (with black pixels, white pixels, gray pixels, etc.) the top of the landscape-based image 440 to ensure that the image content is the same size as an image captured using a portrait-based orientation. As described above, when providing the image content to the neural network 256 for illumination estimation, the image can be provided in a rotated orientation and cropped and padded. When the image content is to be rendered, an inverse rotation (e.g., -180 degrees) of the gravity-aligned rotation (e.g., 180 degrees) can be performed to provide illumination and / or face tracking that is aligned with the physical camera space.

[0111] Although the rear camera 403 was described in the above example, the front camera 450 can be substituted and used in the systems and techniques described herein for content that includes a user captured using the front camera.

[0112] Figures 5A-5D An example of providing gravity-based alignment for face tracking in an AR experience is shown. Gravity-based alignment can ensure that the input image provided to the neural network used in face tracking is upright. For example, the face tracking software 260 can be used to determine which data to use to estimate the screen orientation and thus the device orientation in order to detect an upright face captured, for example, by the front camera on the device 202.

[0113] In some embodiments, for example, gravity-based alignment for face tracking can be performed to ensure that the image provided to the neural network 256 is provided in an upright manner in which the eyes are above the mouth. In such a determination, the screen orientation and / or the image content can be considered. For example, the tracking system 238 can determine the 3-DoF and / or 6-DoF device or camera pose. The pose can be used to estimate the current screen orientation in order to determine the orientation in which to provide the training image to the neural network 256.

[0114] For example, the sensor system 216 can determine the 3-DoF pose of the device 202. If the pose is within a threshold level of a predefined normal vector defined according to the surface, a new screen orientation may not be computable. For example, if the computing device is tilted up or down on the surface at an angle of approximately 8 degrees to approximately 12 degrees relative to the surface, the system 216 can maintain the current screen orientation. For example, if an additional tilt or an additional directional rotation is detected, the system 216 can generate a quantified rotation between the normal vector of the surface and the camera vector calculated from the 3-DoF pose.

[0115] Refer to Figure 5A, a user is shown in an image 502 captured by a front camera 504 on a computing device 506A. The device 506A is shown in a longitudinal orientation aligned about the y-z axis (shown by the y-axis 120 and the z-axis 122). The orientation of the device 506A can be determined by a tracking system 238. Facial tracking software 260 can use the computing device orientation 262 of the device 202 to ensure that the captured image 502 is upright (i.e., has an upright face where the eyes are above the mouth).

[0116] If the system 216 determines that the device 202 is oriented at zero degrees, but the content can be used to generate a landscape image, the system can determine that the image 502 will be realigned by approximately 270 degrees to ensure that a neural network using the image can appropriately track the content to be displayed along with the image in an AR session and provide the content to be displayed, for example, at a location associated with the user's upright face. The system can trigger a counterclockwise rotation of the captured image by approximately 270 degrees, thereby generating an image 508. When the image content is provided to the neural network 256 for illumination estimation and / or facial tracking, the rotated image 508 within the dashed lines may be cropped and rotated upright, as shown by the image 510 before the image 510 is provided to the neural network 256. After completing the facial tracking task, the facial tracking software 260 can use the facial tracking to trigger the rendering of the image content and can trigger a realignment back to the physical camera space coordinates (as shown by the device orientation of the device 506A). In such an example, the output image from the neural network can be rotated in reverse at a gravity alignment angle of 270 degrees.

[0117] As Figure 5B shown, the user may be accessing an AR experience on a device 506B and using the camera to capture image content 512. The system 216 can determine that the device 506B is capturing content in a landscape mode. In such an example, the content may not be rotated. Instead, the system 216 can crop a central portion 514 (e.g., the region of interest) of the captured content 512 and can zero-fill (with black pixels, white pixels, gray pixels, etc.) the top 516 of the landscape-based image to ensure that the image content is the same size as an image captured using a portrait-based orientation. As described above, when the image content is provided to the neural network 256 for facial tracking, the image 518 can be provided in the captured, cropped, and filled orientation.

[0118] As Figure 5CAs shown, the system 216 can determine that the device 506C is capturing content in a portrait mode, but the camera device is in an upright, inverted orientation. For example, the camera 504 is located at the lower part of the device 506C, rather than in the upright orientation as shown for the device 506A. If the system 216 is to generate landscape-oriented content using the content 520 captured with the device 506C in an inverted orientation, the system 216 may have to rotate the content 520 clockwise by 90 degrees, as shown by the content 522. When providing the image content to the neural network 256 for face tracking estimation, the image 522 can be cropped and padded, as shown by the region 524, and can be reoriented by rotating counterclockwise by 90 degrees, as shown by the image 526. After completing the face tracking task, the face tracking software 260 can use the face tracking to trigger the rendering of the image content 526 and can trigger a realignment back to the physical camera space coordinates shown for the device 506C. In such an example, the output image from the neural network can be rotated in the reverse of the gravity alignment angle shown by the rotated image 526.

[0119] As Figure 5D shown, if the system 216 determines that the device 506D is in a landscape orientation (rotated clockwise from the upright portrait orientation as indicated by the device 506A) and the content 530 being captured is in a landscape position (e.g., rotated 180 degrees from the camera sensor), the system 216 can determine to rotate (e.g., realign) the captured image 530 by approximately 180 degrees to ensure that the ceiling or sky is upward, e.g., as shown by the rotated image 532. To generate additional landscape training data using the image 532, the system 216 can crop the central portion 532 (e.g., the region of interest) of the captured content 530 and can zero-pad (with black pixels, white pixels, gray pixels, etc.) the side portions 534 and 536 of the landscape-based image 532 to ensure that the image content 538 is the same size as an image captured using a portrait-based orientation when provided to the neural network 256. As described above, when providing the image content to the neural network 256 for face tracking, the image 532 can be provided in a rotated orientation and cropped and padded to generate the image 538. When the image content is to be rendered, a reverse rotation of the gravity alignment rotation (e.g., 180 degrees) (e.g., -180 degrees from the normal to the ground plane) can be performed to provide face tracking aligned with the physical camera space.

[0120] Figure 6 is an example process 600 for inferring the gravity alignment of image content according to an example embodiment. With respect to Figure 2 the example embodiments of the electronic devices described in, it should be understood that the process 600 can be implemented by devices and systems with other configurations.

[0121] In short, the computing device 202 can incorporate determined orientation knowledge that indicates the upright orientation of a given input image for use in tracking the content captured in the image (e.g., face, object, etc.) and other content surrounding the face, object, etc. For example, for a face feature tracker, the device 202 can determine whether the face in the input image received at the model (e.g., neural network 256) is pre-rotated such that the eyes are above the user's mouth and nose, and the model can learn the spatial relationships between different parts of the face and can be configured to be less likely to provide a prediction where the eyes are below the mouth. Similarly, in an AR lighting estimation example, if the input image to the model typically has the sky (e.g., in an outdoor image) or ceiling (e.g., in an indoor image) in the upper portion of the input image, the model can be configured to be less likely to generate sunlight from a lower region (e.g., lower hemisphere, lower half, etc.), which may represent the source of sunlight occurring naturally and / or in the real world.

[0122] At block 602, process 600 can include obtaining a first image from an image capture device on the computing device at a processor. For example, the camera component 236 can use the image sensor 244 to capture any number of images for the computing device 202. Additionally, any number of processors from the processor component 212 or another on-board device in the computing device 202 can act as the processor throughout process 600. In some embodiments, at least one image includes a live camera image feed that functions to generate multiple images. In some embodiments, at least one sensor includes a tracking stack associated with the tracking features captured in the live camera image feed. For example, when the camera component 236 is capturing video, the sensor system 216 can capture and evaluate one or more images taking into account the tracking stack 245 in order to generate a gravity-aligned image to present to the neural network 256.

[0123] At block 604, process 600 includes detecting, using the processor and at least one sensor, the device orientation of the computing device performing an AR session and associated with the capture of the first image. For example, the IMU 240 can determine the device orientation, such as portrait, landscape, reverse portrait, or reverse landscape and / or another angle between portrait, landscape, reverse portrait, and / or reverse landscape. The orientation of the device 202 can be based on 3-DoF sensor data indicating the device pose or position in space. In some embodiments, the detected device orientation is obtained during operation of the AR session on the computing device 202. In some embodiments, the image can be generated and / or rotated to match the capture orientation associated with previously captured training data. For example, if the previously captured training data is detected as being in a portrait orientation, content oriented differently can be rotated to match the previously captured training data in the portrait orientation.

[0124] At block 606, process 600 includes detecting a movement change associated with the computing device using a processor and at least one sensor. For example, the movement change can be associated with a tracking stack generated by the device 202 using IMU measurements and / or tracked and / or measured by other sensor systems 216. In some embodiments, during image capture, the movement change is detected in real time when the user moves or rotates / tilts the device 202. In some embodiments, the first image includes (or represents) a real-time camera image feed that generates, for example, a plurality of images that will be gravity-aligned. For example, the plurality of images can be continuously aligned based on the detected movement change (in the tracking stack) associated with the computing device 202. For example, the at least one sensor can be the IMU sensor of the computing device 202, and the movement change can represent the tracking stack associated with the IMU 240 and the computing device 202.

[0125] In some embodiments, block 606 can be optional. For example, device movement (i.e., movement change of device 202) may not be detected, but device 202 can still determine orientation and evaluate or access the tracking stack 245 associated with device 202, as described above with respect to Figure 2 as described.

[0126] At block 608, process 600 includes determining a rotation angle for rotating the first image based on the orientation and the movement change. For example, the computing device 202 can determine how to rotate a particular image to achieve an upright input to the neural network 256 for performing face tracking estimation and / or illumination estimation and / or gravity-aligned content (e.g., objects, elements, etc.) for presentation in an AR environment. Additionally, when only portrait-oriented training images have been acquired for training, the rotation angle can be determined to enable simulation of landscape-oriented training images.

[0127] At block 610, process 600 includes rotating the first image to the rotation angle to generate a second image. For example, the determined orientation of the device 202 can be used to rotate, for example, the first image 310A to generate the second image 310B. In some embodiments, the rotation angle is used to align the first image to generate a gravity-aligned second image. For example, when the second image is provided to the neural network 256, the rotation angle can be selected such that the sky or ceiling is located in the upper portion of the second image. In some embodiments, the second image 310B is used by the neural network 256 as training data to generate a landscape-oriented illumination estimate, as Figures 3A-5D shown.

[0128] At block 612, process 600 includes using a processor to provide a second image to at least one neural network to generate a lighting estimate for a first image based on providing the second image to the neural network. For example, as described above, lighting estimation can be performed on the second image (or any number of images in the case of a live image feed), and such lighting estimation can be accurate because process 600 ensures that the image provided to neural network 256 is in an upright orientation, where, for example, image content based on the sky / ceiling is in the upper portion of the image.

[0129] In some embodiments, the second image is generated to match a capture orientation associated with previously captured training data. For example, to ensure that the image provided to neural network 256 is gravity aligned, process 600 can perform a rotation on an image determined to not match the orientation associated with the images used as training data. In some embodiments, the second image is used to generate a horizontally oriented lighting estimate that includes gravity aligned lighting. In some embodiments, the second image is used as an input to a neural network to generate horizontally oriented content having at least one gravity aligned face in the content.

[0130] In some embodiments, the first image is rendered in an AR session on computing device 202, and the second image is provided as gravity aligned content to at least one machine learning model (e.g., neural network 256 associated with computing device 202) to trigger an augmented reality (AR) experience and / or one or more AR features and lighting estimates and / or face tracking estimates associated with the first image. For example, the first image can be rendered as a live camera feed, and machine learning inference is performed on the second (gravity aligned / rotated) image, which can be used to add an overlay of AR content to the first image. In another example, a rotated face tracking estimate can be utilized to render the first image during an AR session. In some embodiments, the face tracking estimate may not be rendered, but instead can be used to render other content or implement a particular AR experience and / or feature. For example, the second image can be provided to neural network 256 to generate a face tracking estimate that can be used to render virtual cosmetics onto the tracked face and overlay the cosmetics (as AR content) onto the live camera feed of the tracked face. In another example, the neural network can use the first image, the second image, and a neural network that detects when a tracked facial feature forms a smile and trigger a still image capture during the exact time the smile is detected.

[0131] In some embodiments, the movement change is associated with an augmented reality (AR) session operating on computing device 202, and the light estimation rotates in the reverse of the rotation angle, and the first image is rendered in the AR session on the computing device using the rotated light estimation. In some embodiments, the rotated light estimation is used to generate AR content and render it as an overlay on the first image. For example, the first image may be rendered as a background image in the AR session on computing device 202, and the light estimation may be used to render specific content onto the image or around, partially overlay, or partially occlude the first image.

[0132] In some embodiments, process 600 may further include using multiple images to generate training data for neural network 256. The training data may include horizontally oriented images generated (e.g., created, produced, etc.) based on the captured vertically oriented images.

[0133] In some embodiments, device 202 may utilize processor component 212, computing device orientation 262, and face tracking software 260 to track a captured face (or head) in an image in order to generate, for example, content that is correctly tracked and placed in the AR environment according to face tracking estimates generated using neural network 256. For example, instead of using the rotation of the first image to the second image to provide the light estimation, computing device 202 may instead use the processor to provide the second image to at least one neural network to generate a face tracking estimate for the first image. The second image may be provided as gravity-aligned content to at least one machine learning model (e.g., neural network 256) to trigger at least one augmented reality (AR) feature associated with the first image. Such features may include audio content, visual content, and / or tactile content to provide an AR experience in the AR environment.

[0134] In some embodiments, the image capture device is a front-facing image capture device of computing device 202, where the user's face is captured along with background content. In some embodiments, the first image is captured using the front-facing image capture device, and the first image includes at least one face rotated by a rotation angle to generate a second image. The second image may be aligned with an eye associated with the face, which is located above the mouth associated with the face.

[0135] In some embodiments, the movement change is associated with an augmented reality (AR) session operating on a computing device, and the face tracking estimate rotates in the reverse of the rotation angle. For example, if the rotation angle indicates a 90-degree clockwise rotation, the reverse of the rotation angle will be a 90-degree counterclockwise rotation (e.g., -90 degrees). The rotated face tracking estimate may be used to render the first image in the AR session on computing device 202.

[0136] In some embodiments, a second image in a facial tracking example can be generated to match a capture orientation associated with previously captured training data. Such data can be used to generate laterally oriented content with a correctly aligned face. For example, computing device 202 can use facial tracking software 260, computing device orientation, and neural network 256 to generate laterally oriented content in which there is at least one gravity-aligned face. In some embodiments, the first image is rendered as a background in an AR session on the computing device, and facial tracking is used to render other content (e.g., audio content, VR content, video content, AR content, lighting, etc.) onto the first image.

[0137] Figure 7 FIG. shows example computer device 700 and example mobile computer device 750, which can be used with the techniques described herein. Generally, the devices described herein can generate and / or provide any one or all aspects in a virtual reality, augmented reality, or mixed reality environment. Features described with respect to computer device 700 and / or mobile computer device 750 can be included in the portable computing device 102 and / or 202 described above. Computing device 700 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 750 is intended to represent various forms of mobile devices, such as, personal digital assistants, cellular telephones, smart phones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are meant to be exemplary only and are not meant to limit embodiments of the systems and techniques claimed and / or described in this document.

[0138] Computing device 700 includes a processor 702, a memory 704, a storage device 706, a high-speed interface 708 connected to the memory 704 and a high-speed expansion port 710, and a low-speed interface 712 connected to a low-speed bus 714 and the storage device 706. Each of the components 702, 704, 706, 708, 710, and 712 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. Processor 702 can process instructions for execution within computing device 700, including instructions stored in memory 704 or storage device 706, to display graphical information of a GUI on an external input / output device, such as a display 716 coupled to the high-speed interface 708. In other embodiments, multiple processors and / or multiple buses can be used as appropriate, as well as multiple memories and memory types. Also, multiple computing devices 700 can be connected, each providing a portion of the necessary operations (e.g., as a server group, blade server group, or multi-processor system).

[0139] Memory 704 stores information within computing device 700. In one embodiment, memory 704 is one or more volatile storage units. In another embodiment, memory 704 is one or more non-volatile storage units. Memory 704 can also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0140] Storage device 706 can provide mass storage for computing device 700. In one embodiment, storage device 706 can be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, a tape device, a flash memory, or other similar solid-state storage devices or arrays of devices, including devices in a storage area network or other configurations. A computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as the methods described above. The information carrier is a computer or machine-readable medium, such as memory 704, storage device 706, or memory on processor 702.

[0141] High-speed controller 708 manages bandwidth-intensive operations of computing device 700, while low-speed controller 712 manages lower bandwidth-intensive operations. This functional assignment is merely exemplary. In one embodiment, high-speed controller 708 (e.g., via a graphics processor or accelerator) is coupled to memory 704, display 716, and to a high-speed expansion port 710 that can accept various expansion cards (not shown). In this embodiment, low-speed controller 712 is coupled to storage device 706 and low-speed expansion port 714. The low-speed expansion port (which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet)) can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a network device, such as a switch or a router, for example, via a network adapter.

[0142] Computing device 700 can be implemented in many different forms, as shown. For example, it can be implemented as a standard server 720, or implemented multiple times in a group of such servers. It can also be implemented as part of a rack server system 724. Additionally, it can be implemented in a personal computer, such as a laptop computer 722. Alternatively, components from computing device 700 can be combined with other components in a mobile device (not shown), such as device 750. Each such device can include one or more computing devices 700, 750, and the entire system can consist of multiple computing devices 700, 750 that communicate with each other.

[0143] Among other components, computing device 750 includes a processor 752, a memory 764, input / output devices such as a display 754, a communication interface 766, and a transceiver 768. Device 750 may also be equipped with a storage device, such as, a microdrive or other device, to provide additional storage. Each of the components 750, 752, 764, 754, 766, and 768 is interconnected using various buses, and several of the components may be mounted on a common motherboard or otherwise mounted as appropriate.

[0144] Processor 752 may execute instructions within computing device 750, including instructions stored in memory 764. The processor may be implemented as a chipset including individual as well as multiple analog and digital processors. The processor may provide, for example, coordination for other components of device 750 (such as, control of a user interface), applications run by device 750, and wireless communications performed by device 750.

[0145] Processor 752 may communicate with a user through a control interface 758 and a display interface 756 coupled to display 754. Display 754 may be, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display or other suitable display technology. Display interface 756 may include appropriate circuitry for driving display 754 to present graphics and other information to a user. Control interface 758 may receive commands from a user and convert them for submission to processor 752. Additionally, an external interface 762 may be provided in communication with processor 752 to enable near area communication of device 750 with other devices. External interface 762 may provide, for example, for wired communication in some embodiments, or for wireless communication in other embodiments, and may also use multiple interfaces.

[0146] Memory 764 stores information within computing device 750. Memory 764 may be implemented as a computer-readable medium, a volatile storage unit, or one or more non-volatile storage units. An extended memory 774 may also be provided and connected to device 750 through an expansion interface 772, which may include, for example, a SIMM (Single In-line Memory Module) card interface. Such extended memory 774 may provide additional storage space for device 750, or may also store applications or other information for device 750. Specifically, extended memory 774 may include instructions for performing or supplementing the above processes, and may also include security information. Thus, for example, extended memory 774 may be provided as a security module for device 750, and may be programmed with instructions that allow secure use of device 750. Additionally, security applications may be provided via the SIMM card as well as additional information (such as, placing identification information on the SIMM card in a non-intrusive manner).

[0147] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as the methods described above. The information carrier is a computer or machine-readable medium, such as the memory 764, the extended memory 774, or the memory on the processor 752, which may be received, for example, via the transceiver 768 or the external interface 762.

[0148] The device 750 may communicate wirelessly via the communication interface 766, which may include digital signal processing circuitry when necessary. The communication interface 766 may provide communication in various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, etc. Such communication may occur, for example, via the radio frequency transceiver 768. In addition, short-range communication may occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, the GPS (Global Positioning System) receiver module 770 may provide other wireless data related to navigation and location to the device 750, and the applications running on the device 750 may appropriately use the above data.

[0149] The device 750 may also communicate audibly using the audio codec 760, which may receive voice information from the user and convert it into usable digital information. The audio codec 760 may similarly generate audible sounds for the user, such as via the speaker in the earpiece of the device 750. Such sounds may include sounds from a voice telephone call, may include recorded sounds (e.g., voice messages, music files, etc.), and may also include sounds generated by applications running on the device 750.

[0150] As shown, the computing device 750 may be implemented in a variety of different forms. For example, it may be implemented as a cellular phone 780. It may also be implemented as part of a smart phone 782, a personal digital assistant, or other similar mobile devices.

[0151] Embodiments of the various techniques described herein can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or combinations thereof. Embodiments can be implemented as a computer program product, i.e., a computer program tangibly embodied in an information carrier, e.g., a machine-readable storage device or a propagated signal, for execution by, or to control the operation of, data processing apparatus, e.g., a programmable processor, a computer, or multiple computers. A computer program such as the computer program described above can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed to be executed on one computer or multiple computers at one site, or distributed across multiple sites and interconnected by a communication network.

[0152] Method steps can be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Method steps can also be performed by special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the apparatus can be implemented as the special purpose logic circuitry described above.

[0153] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. Elements of a computer may include at least one processor for executing instructions and one or more storage devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks, or the computer can be operatively coupled to receive data from, or transfer data to, one or more mass storage devices, or both. Information carriers suitable for embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0154] To provide interaction with a user, embodiments can be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor, an LED display) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input.

[0155] Embodiments can be implemented in a computing system including a backend component (e.g., a data server), or including a middleware component (e.g., an application server), or including a frontend component (e.g., a client computer having a graphical user interface or a web browser through which the user can interact with the embodiment), or including a combination of such backend, middleware, or frontend components. The components can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as, for example, the Internet.

[0156] A computing device based on the example embodiments described herein can be implemented using any suitable combination of hardware and / or software configured to interface with a user interface including a user device, a user interface (UI) device, a user terminal, a client device, or a customized device. The computing device can be implemented as a portable computing device, such as, for example, a laptop computer. The computing device can be implemented as some other type of portable computing device suitable for interfacing with a user interface, such as, for example, a PDA, a notebook computer, or a tablet computer. The computing device can be implemented as some other type of computing device suitable for interfacing with a user interface, such as, for example, a PC. The computing device can be implemented as a portable communication device (e.g., a mobile phone, a smart phone, a wireless cellular phone, etc.) suitable for interfacing with a user interface and suitable for wireless communication via a network including a mobile communication network.

[0157] A computer system (e.g., a computing device) can be configured to wirelessly communicate with a network server via a network through a communication link established with the network server using any known wireless communication technologies and protocols, including radio frequency (RF), microwave frequency (MWF), and / or infrared frequency (IRF) wireless communication technologies and protocols suitable for communicating over a network.

[0158] In accordance with various aspects of the present disclosure, implementations of the various techniques described herein can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or combinations thereof. Implementations can be implemented as a computer program product (e.g., a computer program tangibly embodied in an information carrier, a machine-readable storage device, a computer-readable medium, a tangible computer-readable medium) for processing by a data processing apparatus (e.g., a programmable processor, a computer, or multiple computers), or to control the operation of the above data processing apparatus. In some implementations, the tangible computer-readable storage medium can be configured to store instructions that, when executed, cause the processor to perform a process. A computer program such as the above computer program can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The computer program can be deployed to be processed on one computer or multiple computers at one site, or distributed across multiple sites and interconnected by a communication network.

[0159] The specific structural and functional details disclosed herein are merely representative for the purpose of describing example embodiments. However, example embodiments can be embodied in many alternative forms and should not be construed as limited to the embodiments set forth herein.

[0160] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "comprise", "comprising", "include" and / or "including" specify the presence of the stated features, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components and / or groups thereof.

[0161] It will be understood that when an element is referred to as being "coupled", "connected" or "responsive" to another element or "on" another element, it can be directly coupled, connected or responsive, or directly thereon, and other elements or intermediate elements can also be present. In contrast, when an element is referred to as being "directly coupled", "directly connected" or "directly responsive" to another element or "directly on" another element, no intermediate elements are present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0162] For ease of description, spatial relative terms, such as "below", "beneath", "lower", "above", "upper", etc., may be used herein to describe the relationship of one element or feature to another as shown in the figures. It will be understood that the spatial relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, an element described as "below" or "beneath" another element or feature will then be oriented "above" the other element or feature. Thus, the term "below" can encompass both an upper and a lower orientation. The device may be otherwise oriented (rotated 70 degrees or at other orientations), and the spatial relative descriptors used herein may be interpreted accordingly.

[0163] Exemplary embodiments of the concepts are described herein with reference to cross-sectional views, which are schematic illustrations of idealized embodiments (and intermediate structures) of the exemplary embodiments. As such, for example, variations in the illustrated shapes due to manufacturing techniques and / or tolerances are to be expected. Accordingly, the example embodiments of the concepts described herein should not be construed as limited to the particular shapes of regions shown herein, but should include, for example, shape deviations resulting from manufacturing. Thus, the regions shown in the figures are schematic in nature, and their shapes are not intended to show the actual shape of regions of the device nor are they intended to limit the scope of the example embodiments.

[0164] It will be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. Thus, a "first" element may be termed a "second" element without departing from the teachings of this embodiment.

[0165] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which these concepts belong. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or this specification, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0166] Although certain features of the described embodiments have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. Accordingly, it will be understood that the appended claims are intended to cover all such modifications and changes that fall within the scope of the embodiments. It should be understood that they are given by way of example only and not by way of limitation, and that various changes in form and detail may be made. Except for mutually exclusive combinations, any part of the apparatus and / or method described herein may be combined in any combination. The embodiments described herein may include various combinations and / or sub-combinations of the functions, components, and / or features of the different embodiments described.

Claims

1. A computer-implemented method, the method comprises: obtaining, at a processor, a first image from an image capture device included on a computing device; detecting, using the processor and at least one sensor, a device orientation of the computing device associated with the capture of the first image; determining, based on the device orientation, a rotation angle for rotating the first image; rotating the first image to the rotation angle to generate a second image; and using the processor to provide the second image to at least one model to generate a lighting estimate of the first image based on the second image, wherein the generated lighting estimate is rotated from the rotation angle based on the device orientation, and augmented reality (AR) content is generated using the rotated lighting estimate and rendered as an overlay on the first image.

2. The method according to claim 1, wherein: the detected device orientation occurs during an AR session operating on the computing device; and the first image is rendered in the AR session on the computing device.

3. The method according to claim 1, wherein, the second image is generated to match a capture orientation associated with previously captured training data, and wherein the second image is used to generate a horizontally oriented lighting estimate.

4. The method according to claim 1, wherein, the rotation angle is used to align the first image to generate a gravity-aligned second image.

5. The method according to claim 1, wherein: the first image is a real-time camera image feed that generates a plurality of images; and the plurality of images are continuously aligned based on detected movement changes associated with the computing device.

6. The method according to claim 5, wherein, the at least one sensor includes a tracking stack associated with tracking features captured in the real-time camera image feed.

7. The method according to claim 5, wherein, the at least one sensor is an inertial measurement unit (IMU) of the computing device, and the movement change represents a tracking stack associated with the IMU and the computing device.

8. A system for inferring gravity-aligned imagery, comprises: an image capture device associated with a computing device; at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to: obtain, at a processor, a first image from the image capture device; detect, using the processor and at least one sensor, a device orientation of the computing device associated with the capture of the first image; detect, using the processor and the at least one sensor, a movement change associated with the computing device, the movement change being associated with an augmented reality (AR) session operating on the computing device; determine, based on the orientation and the movement change, a rotation angle for rotating the first image; rotate the first image to the rotation angle to generate a second image; and Generate a facial tracking estimate of the first image based on the second image and according to the movement change, wherein the second image includes gravity-aligned content, and the gravity-aligned content is provided to at least one machine learning model associated with the computing device to trigger an AR experience in the AR session associated with the first image and the rotated facial tracking estimate.

9. The system according to claim 8, wherein: The image capture device is a front image capture device of the computing device; The first image is captured using the front image capture device; and The first image includes at least one face rotated by the rotation angle to generate the second image, and the second image is aligned with the eyes associated with the face, and the eyes are located above the mouth associated with the face.

10. The system according to claim 8, wherein: The facial tracking estimate is rotated in the reverse of the rotation angle; and The first image is rendered in the AR session on the computing device.

11. The system according to claim 8, wherein, The second image is used as an input to a model to generate laterally oriented content in which there is at least one gravity-aligned face.

12. The system according to claim 8, wherein: The first image is a real-time camera image feed that generates multiple images; and The multiple images are continuously aligned based on the detected movement change associated with the computing device.

13. The system according to claim 12, wherein: The second image is generated to match the capture orientation associated with previously captured training data, and The second image is used to generate a laterally oriented facial tracking estimate.

14. The system according to claim 8, wherein, The at least one sensor is an inertial measurement unit (IMU) of the computing device, and the movement change represents a tracking stack associated with the IMU and the computing device.

15. A computer program product, the computer program product being tangibly embodied on a non-transitory computer-readable medium and including instructions that, when executed, are configured to cause at least one processor to: At the processor, obtain a first image from an image capture device on a computing device; Using the processor and at least one sensor, detect the device orientation of the computing device associated with the capture of the first image; Based on the device orientation and a tracking stack associated with the computing device, determine a rotation angle for rotating the first image; Rotate the first image to the rotation angle to generate a second image; and Provide the second image as gravity-aligned content to at least one machine learning model associated with the computing device to trigger the generation of at least one augmented reality (AR) feature associated with the first image and a lighting estimate of the at least one feature, wherein the generated lighting estimate is rotated from the rotation angle based on the device orientation, and the AR feature is generated using the rotated lighting estimate and rendered as an overlay on the first image.

16. The computer program product according to claim 15, wherein: the at least one sensor includes the tracking stack corresponding to the trackable features captured in the first image.

17. The computer program product according to claim 15, wherein, the second image is generated to match the capture orientation associated with the previously captured training data.

18. The computer program product according to claim 15, wherein: the first image is a real-time camera image feed that generates multiple images; and the multiple images are continuously aligned based on the detected movement associated with the tracking stack.

19. The computer program product according to claim 18, further comprises: using the multiple images to generate an input to a model, the input including a generated lateral orientation image based on the captured longitudinal orientation image.

20. The computer program product according to claim 15, wherein, the at least one sensor is an inertial measurement unit (IMU) of the computing device, and the tracking stack is associated with the changes detected at the computing device.