Object depth estimation process within imaging device

By using the six-degree of freedom tracker and machine learning process to generate sparse depth maps in the imaging device, the problems of low efficiency and insufficient accuracy of depth estimation in traditional methods are solved, and more efficient and accurate depth map generation is achieved, reducing equipment cost and power consumption.

CN120266157APending Publication Date: 2025-07-04QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081176.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing imaging devices have problems with inefficient and insufficient accuracy in object depth estimation, especially in virtual reality and augmented reality devices, where traditional methods require depth sensors to increase the cost and power consumption of the device.

Method used

By using a six-degree of freedom tracker to acquire three-dimensional feature points, generate sparse depth maps, and use machine learning processes including image encoding, sparse encoding and decoding processes to generate more accurate depth values and reduce dependence on depth sensors.

Benefits of technology

More efficient and accurate depth map generation is achieved, reducing equipment cost and power consumption, while reducing equipment volume and weight.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266157A_ABST
    Figure CN120266157A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus are provided for determining object depth within a captured image. For example, an imaging device, such as a VR or AR device, captures an image. The imaging device applies a first encoding process to the image to generate a first set of features. The imaging device also generates a sparse depth map based on the image and applies a second encoding process to the sparse depth map to generate a second set of features. Further, the imaging device applies a decoding process to the first set of features and the second set of features to generate predicted depth values. In some examples, the decoding process receives a jump connection from a layer of the second encoding process as an input to a corresponding layer of the decoding process. The imaging device generates an output image, such as a 3D image, based on the predicted depth value.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art Technical Field

[0001] The present disclosure generally relates to imaging devices and, more particularly, to object depth estimation using machine learning processes in imaging devices.

[0002] Description of Related Technologies

[0003] Imaging devices (such as virtual reality devices, augmented reality devices, cellular devices, tablet computers, and smart devices) may use various signal processing techniques to render three-dimensional (3D) images. For example, an imaging device may capture an image and may apply conventional image processing techniques to the captured image to reconstruct a 3D image. In some examples, the imaging device may include a depth sensor to determine the depth of an object in the field of view of the imaging device's camera. In some examples, the imaging device may perform depth estimation algorithms on the captured images (e.g., left-eye image and right-eye image) to determine object depth. There is an opportunity to improve depth estimation within imaging devices. Summary of the Invention

[0004] According to one aspect, a method by an imaging device includes receiving three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker. The method further includes generating sparse depth values based on the three-dimensional feature points. Additionally, the method includes generating predicted depth values based on an image and the sparse depth values. The method further includes storing the predicted depth values in a data repository.

[0005] According to another aspect, an apparatus includes a non-transitory machine-readable storage medium storing instructions, and at least one processor coupled to the non-transitory machine-readable storage medium. The at least one processor is configured to execute the instructions to receive three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker. The at least one processor is further configured to execute the instructions to generate sparse depth values based on the three-dimensional feature points. Additionally, the at least one processor is configured to execute the instructions to generate predicted depth values based on an image and the sparse depth values. The at least one processor is further configured to execute the instructions to store the predicted depth values in a data repository.

[0006] According to another aspect, a non-transitory machine-readable storage medium stores instructions that, when executed by at least one processor, cause the at least one processor to perform operations including receiving three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker. The operations further include generating sparse depth values based on the three-dimensional feature points. Additionally, the operations include generating predicted depth values based on an image and the sparse depth values. The operations further include storing the predicted depth values in a data repository. Brief Description of the Drawings

[0007] Figure 1 is a block diagram of an exemplary imaging device according to some specific implementations;

[0008] Figure 2 and Figure 3 is a block diagram of an exemplary portion of an imaging device that illustrates according to some specific implementations Figure 1 thereof;

[0009] Figure 4 is a diagram that illustrates a machine learning model according to some specific implementations;

[0010] Figure 5 is a flowchart of an exemplary process for determining depth values for objects within an image according to some specific implementations;

[0011] Figure 6 is a flowchart of an exemplary process for rendering an image based on the determined depth values according to some specific implementations; and

[0012] Figure 7 is a flowchart of an exemplary process for training a machine learning process according to some specific implementations. DETAILED DESCRIPTION

[0013] Although the features, methods, devices, and systems described herein may be embodied in various forms, some exemplary and non - limiting implementations are shown in the drawings and described below. Some of the components described in this disclosure are optional, and some specific implementations may include additional components, different components, or fewer components compared to those explicitly described in this disclosure.

[0014] In some specific implementations, the imaging device projects three - dimensional feature points from tracking information onto two - dimensional feature points to generate a sparse depth map. The imaging device applies a machine learning process to the sparse depth map and the color image to generate a refined depth map. For example, an imaging device such as a VR or AR device captures an image. The imaging device applies a first encoding process to the image to generate a first set of features. The imaging device also generates a sparse depth map based on the image and applies a second encoding process to the sparse depth map to generate a second set of features. Additionally, the imaging device applies a decoding process to the first set of features and the second set of features to generate predicted depth values. In some examples, the decoding process receives skip connections from the layers of the second encoding process as inputs to the corresponding layers of the decoding process. The imaging device generates an output image, such as a 3D image, based on the predicted depth values.

[0015] In some specific implementations, the imaging device may include one or more cameras, one or more sensors (e.g., gyro sensors, accelerometers, etc.), an image encoder engine, a sparse encoder engine, and a decoder engine. In some examples, one or more of the image encoder engine, the sparse encoder engine, and the decoder engine may include instructions executed by one or more processors. Additionally, each camera may include, for example, one or more lenses and one or more imaging sensors. Each camera may also include one or more lens controllers that can adjust the position of the lens. The imaging device may capture image data from each of the cameras. For example, the imaging device may capture first image data from a first camera and may also capture second image data from a second camera. In some examples, the first camera and the second camera may jointly form a stereo camera (e.g., a left camera and a right camera).

[0016] When executed by one or more processors, the image encoder engine may receive image data representing an image captured by one or more of the cameras and may generate image encoder feature data (e.g., image feature values) representing the features of the image data. For example, the executed image encoder engine may receive the image data and apply an encoding process (e.g., a feature extraction process) to the image data to extract a set of image features. In some examples, the executed image encoder engine constructs a neural network encoder, such as an encoder of a convolutional neural network (CNN) (e.g., an Xception convolutional neural network, a CNN-based image feature extractor, or a deep neural network (DNN) encoder (e.g., a DNN- or CNN-based encoder)), and applies the constructed encoder to the image data to extract a set of image features.

[0017] When executed by one or more processors, the sparse encoder engine may receive sparse depth data (e.g., sparse depth values) representing a sparse depth map of an image (such as an image captured by a camera) and may generate sparse encoder feature data (e.g., sparse feature values) representing the sparse features of the image. For example, the executed sparse encoder engine may receive the sparse depth data and apply an encoding process to the sparse depth data to extract a set of sparse features. In some examples, the executed sparse encoder engine constructs a neural network encoder, such as an encoder of a CNN (e.g., a CNN-based sparse depth feature extractor or a deep neural network (DNN) encoder (e.g., a DNN- or CNN-based encoder)), and applies the constructed encoder to the sparse depth data to extract sparse features. In some instances, the executed sparse encoder engine provides skip connections (e.g., the outputs of one or more layers of the encoding process) to the executed decoder engine, as further described below.

[0018] The executed decoder engine, when executed by one or more processors, can receive a set of image features (e.g., such as those generated by an encoder engine) and a set of sparse features (e.g., such as those generated by a sparse encoder engine), and can generate depth map data (e.g., depth map values) representing a predicted depth map of an image (e.g., such as an image captured by a camera). For example, the executed decoder engine can receive image encoder feature data generated by the executed image encoder engine and sparse encoder feature data generated by the executed sparse encoder engine. Additionally, the executed decoder engine can apply a decoding process to the image encoder feature data and the sparse encoder feature data to generate a predicted depth map. In some examples, the executed decoder engine builds a decoder of a neural network, such as a decoder of a deep neural network built by the executed encoder engine (e.g., a CNN-based decoder or a DNN-based decoder), and applies the built decoder to the image encoder feature data and the sparse encoder feature data to generate a predicted depth map (e.g., to generate predicted depth values based on the image and sparse features). In some examples, the executed decoder engine receives skip connections from the executed sparse encoder and inputs the skip connections into corresponding layers of the decoding process.

[0019] In some examples, the imaging device includes one or more of a head tracker engine, a sparse point engine, and a rendering engine. In some examples, one or more of the head tracker engine, the sparse point engine, and the rendering engine may include instructions executable by one or more processors. When executed by one or more processors, the head tracker engine may receive image data representative of an image captured by one or more cameras in the camera, and may generate feature point data representative of features of the image. For example, the executed head tracker engine may apply one or more processes (e.g., trained machine learning processes) to the received image data, and in some instances to additional sensor data, to generate the feature point data. In some examples, the image captured by the camera is a monochrome image (e.g., a grayscale image, a grayscale image in each of three color channels), and thus the generated features are based on the monochrome image. In some examples, the executed head tracker engine may apply one or more processes to sensor data such as accelerometer and / or gyroscope data to generate the feature point data. In some instances, the executed head tracker engine may also generate pose data representative of the pose of a user (such as a user of the imaging device). Additionally, the determined pose may be temporally associated with the capture time of an image (such as an image captured by the camera). In some examples, the head tracker engine includes instructions that provide six degrees of freedom (6Dof) tracking functionality when executed by one or more processors. For example, the executed head tracker engine may detect movement of a user's head including yaw, pitch, roll, and movement within a space including left, right, backward, forward, up, and down based on, for example, the image data and / or the sensor data.

[0020] When executed by one or more processors, the sparse point engine may receive feature point data (e.g., key points) from the executed head tracker engine, and may generate sparse depth values representative of a sparse depth map based on the feature point data. For example, the executed sparse point engine may apply a depth estimation process to the feature point data to generate the sparse depth values representative of the sparse depth map. In some instances, the executed sparse encoder engine operates on the sparse depth values generated by the executed sparse point engine.

[0021] When executed by one or more processors, the rendering engine may receive depth map data (e.g., a predicted depth map) from the executed decoder engine, and may render an output image based on the depth map data. For example, the executed rendering engine may execute one or more mesh rendering processes to generate mesh data representative of a mesh of an image (such as an image captured by one of the cameras in the camera) based on the depth map data. In some examples, the executed rendering engine may execute one or more plane estimation processes to generate plane data representative of one or more planes based on the mesh data.

[0022] Among other advantages, the imaging device may not require a depth sensor, thereby reducing cost and power consumption, as well as reducing the size and weight of the imaging device. In addition, the imaging device may generate a depth map more accurately and efficiently compared to conventional depth estimation methods.

[0023] Figure 1 is a block diagram of an exemplary imaging device 100. The functions of the imaging device 100 may be implemented in one or more processors, one or more field programmable gate arrays (FPGAs), one or more application specific integrated circuits (ASICs), one or more state machines, digital circuits, any other suitable circuits, or any suitable hardware. The imaging device 100 may perform one or more of the example functions and processes described in this disclosure. Examples of the imaging device 100 include, but are not limited to, extended reality devices (e.g., virtual reality devices (e.g., virtual reality headsets), augmented reality devices (e.g., augmented reality glasses), mixed reality devices, etc.), cameras, video recording devices such as camcorders, mobile devices such as tablet computers, wireless communication devices (such as, for example, mobile phones, cellular phones, etc.), handheld devices such as portable video game devices or personal digital assistants (PDAs), or any device that may include one or more cameras.

[0024] As Figure 1 illustrated in the example of, the imaging device 100 may include one or more imaging sensors 112 (such as imaging sensor 112A), one or more lenses 113 (such as lens 113A), and one or more camera processors (such as camera processor 114). The camera processor 114 may also include a lens controller that may be operable to adjust the position of one or more lenses 113 (such as 113A). In some instances, the camera processor 114 may be an image signal processor (ISP) that employs various image processing algorithms to process image data (e.g., as captured by the corresponding lenses and sensors among these lenses and sensors). For example, the camera processor 114 may include an image front end (IFE) and / or an image processing engine (IPE) as part of the processing pipeline. In addition, the camera 115 may refer to a collective device that includes one or more imaging sensors 112, one or more lenses 113, and one or more camera processors 114.

[0025] In some examples, one or more image sensors in the image sensor 112 may be assigned to each lens in the lens 113. Additionally, in some examples, one or more image sensors in the image sensor 112 may be assigned to corresponding lenses in the lens 113 having respective and different lens types (e.g., wide-angle lens, ultra-wide-angle lens, telephoto lens, and / or periscope lens, etc.). For example, the lens 113 may include a wide-angle lens, and a corresponding image sensor in the image sensor 112 having a first size (e.g., 108MP) may be assigned to the wide-angle lens. In other instances, the lens 113 may include an ultra-wide-angle lens, and a corresponding image sensor in the image sensor 112 having a second and different size (e.g., 16MP) may be assigned to the ultra-wide-angle lens. In another instance, the lens 113 may include a telephoto lens, and a corresponding image sensor in the image sensor 112 having a third size (e.g., 12MP) may be assigned to the telephoto lens.

[0026] In an illustrative example, a single imaging device 100 may include two or more cameras (e.g., two or more cameras in the camera 115), and at least two of the cameras include image sensors (e.g., the image sensor 112) having the same size (e.g., two 12MP sensors, three 108MP sensors, three 12MP sensors, two 12MP sensors and one 108MP sensor, etc.). Additionally, in some examples, a single image sensor (e.g., the image sensor 112A) may be assigned to multiple lenses in the lens 113. Additionally or alternatively, each image sensor in the image sensor 112 may be assigned to a different lens in the lens 113, e.g., to provide multiple cameras to the imaging device 100.

[0027] In some examples, the imaging device 100 may include multiple cameras 115 (e.g., a VR device or an AR device having multiple cameras, a mobile phone having one or more front cameras and one or more rear cameras). For example, the imaging device 100 may be a VR headset that includes: a first camera (such as the camera 115) having a first field of view and located in a first part (e.g., a corner point) of the headset; a second camera having a second field of view and located in a second part of the headset; a third camera having a third field of view and located in the second part of the headset; and a fourth camera having a fourth field of view and located in a fourth part of the headset. Each camera 115 may include an image sensor 112A having a corresponding resolution (such as 12MP, 16MP, or 108MP).

[0028] In some examples, the imaging device 100 may include multiple cameras oriented in different directions. For example, the imaging device 100 may include dual "front-facing" cameras. Additionally, in some examples, the imaging device 100 may include a "front-facing" camera (such as camera 115) and a "rear-facing" camera. In other examples, the imaging device 100 may include dual "front-facing" cameras (which may include camera 115) and one or more "side-facing" cameras. In other examples, the imaging device 100 may include three "front-facing" cameras, such as camera 115. In yet other examples, the imaging device 100 may include three "front-facing" cameras and one, two, or three "rear-facing" cameras. Additionally, those skilled in the art will appreciate that the techniques of the present disclosure may be implemented for any type of camera and for any number of cameras of the imaging device 100.

[0029] Each imaging sensor in the imaging sensors 112 (including imaging sensor 112A) may represent an image sensor including processing circuitry, an array of pixel sensors (e.g., pixels) for capturing a representation of light, memory, an adjustable lens (such as lens 113), and an actuator for adjusting the lens. By way of example, imaging sensor 112A may be associated with a corresponding lens in lens 113 (such as lens 113A) and may capture an image through the corresponding lens. In other examples, an additional or alternative imaging sensor in the imaging sensors 112 may be associated with a corresponding additional lens in lens 113 and may capture an image through the corresponding additional lens.

[0030] In some instances, the imaging sensors 112 may include monochrome sensors (e.g., "clear" pixel sensors) and / or color sensors (e.g., Bayer sensors). For example, a monochrome pixel sensor may be established by disposing a monochrome filter over imaging sensor 112A. Additionally, in some examples, a color pixel sensor may be established by disposing a color filter (such as a Bayer filter) over imaging sensor 112A or by disposing a red filter, a green filter, or a blue filter over imaging sensor 112A. There are various other filter patterns, such as red, green, blue, white ("RGBW") filter arrays; cyan, magenta, yellow, white (CMYW) filter arrays; and / or variants thereof, including proprietary or non-proprietary filter patterns.

[0031] In addition, in some examples, multiple lenses in lens 113 may be associated with and disposed over corresponding subsets of imaging sensor 112. For example, a first subset of imaging sensor 112 may be assigned to a first lens in lens 113 (e.g., a wide-angle lens camera, an ultra-wide-angle lens camera, a telephoto lens camera, a periscope lens camera, etc.), and a second subset of imaging sensor 112 may be assigned to a second lens in lens 113 that is different from the first subset. In some instances, each lens in lens 113 may serve a corresponding function provided by various attributes of the camera (e.g., lens attributes, aperture attributes, viewing angle attributes, thermal imaging attributes, etc.), and a user of imaging device 100 may utilize the various attributes of each lens in lens 113 to capture one or more images or image sequences (e.g., as in video recording).

[0032] Imaging device 100 may also include a central processing unit (CPU) 116, one or more sensors 129, an encoder / decoder 117, a transceiver 119, a graphics processing unit (GPU) 118, local memory 120 of GPU 118, a user interface 122, a memory controller 124 that provides access to system memory 130 and instruction memory 132, and a display interface 126 that outputs a signal to display graphic data on display 128.

[0033] Sensor 129 may be, for example, a gyroscope sensor (e.g., a gyroscope), which is operable to measure the rotation of imaging device 100. In some instances, gyroscope sensors 129 may be distributed across imaging device 100 to measure the rotation of imaging device 100 about one or more axes of imaging device 100 (e.g., yaw, pitch, and roll). In addition, each gyroscope sensor 129 may generate gyroscope data representative of the measured rotation and may store the gyroscope data in a memory device (e.g., internal RAM, first-in-first-out (FIFO), system memory 130, etc.). For example, the gyroscope data may include one or more rotation values that identify the rotation of imaging device 100. CPU 116 and / or camera processor 114 may obtain (e.g., read) the generated gyroscope data from each gyroscope sensor.

[0034] As another example, sensor 129 may be an accelerometer, which is operable to measure the acceleration of imaging device 100. In some examples, imaging device 100 may include multiple accelerometers 129 to measure acceleration in multiple directions. For example, each accelerometer may generate acceleration data representative of the acceleration in one or more directions and may store the acceleration data in a memory device (such as internal memory or system memory 130).

[0035] Additionally, in some instances, the imaging device 100 may receive user input via the user interface 122, and in response to the received user input, the CPU 116 and / or the camera processor 114 may activate a corresponding lens or combination of lenses in the lens 113. For example, the received user input may correspond to a user selection of the lens 113A (e.g., a fish-eye lens), and based on the received user input, the CPU 116 may select an initial lens in the lens 113 for activation, and additionally or alternatively, may transition from the initially selected lens 113A to another lens in the lens 113.

[0036] In other examples, the CPU 116 and / or the camera processor 114 may detect operating conditions that meet certain lens selection criteria (e.g., a digital zoom level that meets a predefined camera transition threshold, a change in lighting conditions, an input from a user requiring a specific lens 13, etc.), and may select an initial lens (such as the lens 113A) in the lens 113 for activation based on the detected operating conditions. For example, the CPU 116 and / or the camera processor 114 may generate a lens adjustment command and provide the lens adjustment command to the lens controller 114A to adjust the position of the corresponding lens 113A. For example, the lens adjustment command may identify the position to which the lens 113A is to be adjusted, or for example, the amount by which to adjust the current lens position. In response, the lens controller 114A may adjust the position of the lens 113A according to the lens adjustment command. In some examples, the imaging device 100 may include multiple cameras in the camera 115, and the multiple cameras may jointly capture a composite image or a stream of composite images such that the camera processor 114 or the CPU 116 may process a composite image or a stream of composite images based on the image data captured from the imaging sensor 112. In some examples, the operating conditions detected by the CPU 116 and / or the camera processor 114 include rotation as determined based on rotation data or acceleration data received from the sensor 129.

[0037] In some examples, each of the lens 113 and the imaging sensor 112 may operate jointly to provide various optical zoom levels, angle of view (AOV), focal length, and FOV. Additionally, an optical waveguide may be used to direct incident light from the lens 113 to the corresponding imaging sensor in the imaging sensor 112, and examples of the optical waveguide may include, but are not limited to, a prism, a movable prism, or one or more mirrors. For example, light received from the lens 113A may be redirected from the imaging sensor 112A towards another imaging sensor in the imaging sensor 112. Additionally, in some instances, the camera processor 114 may perform an operation of moving the prism and redirecting the light incident on the lens 113A in order to effectively change the focal length of the received light.

[0038] In addition, as Figure 1As illustrated, a single camera processor, such as camera processor 114, may be assigned to and docked with all or a selected subset of the imaging sensor 112. In other instances, multiple camera processors may be assigned to and docked with all or a selected subset of the imaging sensor 112, and each of the camera processors may coordinate with one another to effectively allocate processing resources to all or a selected subset of the imaging sensor 112. For example, and by execution of stored instructions, the camera processor 114 may implement various processing algorithms in various situations to perform digital zoom operations or other image processing operations.

[0039] Although the various components of the imaging device 100 are illustrated as separate components, in some examples, these components may be combined to form a system-on-chip (SoC). As an example, the camera processor 114, CPU 116, GPU 118, and display interface 126 may be implemented on a common integrated circuit (IC) chip. In some examples, one or more of the camera processor 114, CPU 116, GPU 118, and display interface 126 may be implemented in separate IC chips. Various other arrangements and combinations are possible, and the techniques of the present disclosure should not be considered limited to Figure 1 the examples.

[0040] The system memory 130 may include one or more volatile or non-volatile memories or storage devices, such as, for example, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic data media, cloud-based storage media, or optical storage media.

[0041] The system memory 130 may store program modules and / or instructions and / or data that can be accessed by the camera processor 114, CPU 116, and GPU 118. For example, the system memory 130 may store user applications (e.g., instructions for a camera application) and resulting images from the camera processor 114. The system memory 130 may also store rendered images, such as three-dimensional (3D) images, rendered by one or more of the camera processor 114, CPU 116, and GPU 118. The system memory 130 may additionally store information for use and / or generation by other components of the imaging device 100. For example, the system memory 130 may serve as device memory for the camera processor 114.

[0042] Similarly, the GPU 118 can store data to and read data from the local memory 120. For example, the GPU 118 can store a working instruction set to the local memory 120, such as instructions loaded from the instruction memory 132. The GPU 118 can also use the local memory 120 to store dynamic data created during the operation of the imaging device 100. Examples of the local memory 120 include one or more volatile or non-volatile memories or storage devices, such as RAM, SRAM, DRAM, EPROM, EEPROM, flash memory, magnetic data media, cloud-based storage media, or optical storage media.

[0043] The instruction memory 132 can store instructions that can be accessed (e.g., read) and executed by one or more of the camera processor 114, the CPU 116, and the GPU 118. For example, the instruction memory 132 can store instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to perform one or more of the operations described herein. For example, the instruction memory 132 can include instructions 133 that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to establish one or more machine learning processes 133 to generate depth values representative of a predicted depth map.

[0044] For example, the encoder model data 132A can include instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to establish a first encoding process and apply the established first encoding process to an image (such as an image captured by the camera 115) to generate a set of image features.

[0045] The instruction memory 132 can also include sparse encoder model data 132B, which can include instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to establish a second encoding process and apply the established second encoding process to a sparse depth map to generate a set of sparse features.

[0046] In addition, the instruction memory 132 may include decoder model data 132C, which may include instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to establish a decoding process and apply the established decoding process to a set of image features and a set of sparse features to generate a predicted depth map.

[0047] The instruction memory 132 may further include head tracker model data 132D, which may include instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to establish a feature extraction process and apply the feature extraction process to an image (such as an image captured by the camera 115) to generate feature point data representing features of the image. The head tracker model data 132D may further include instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to generate pose data representing the pose of a user based on the feature point data. In some examples, the head tracker model data 132D includes instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, provide six degrees of freedom (6Dof) tracking functionality. For example, when the head tracker model data 132D is executed, one or more of the camera processor 114, the CPU 116, and the GPU 118 may detect the movement of a user's head including yaw, pitch, and roll, as well as movement within the space including left, right, backward, forward, up, and down.

[0048] Instruction memory 132 may also include rendering model data 132E, which may include instructions that, when executed by one or more of camera processor 114, CPU 116, and GPU 118, cause one or more of camera processor 114, CPU 116, and GPU 118 to render an output image based on a predicted depth map generated by a decoding process. In some examples, rendering model data 132E may include instructions that, when executed by one or more of camera processor 114, CPU 116, and GPU 118, cause one or more of camera processor 114, CPU 116, and GPU 118 to render an output image based on the predicted depth map and pose data. For example, in some instances, rendering model data 132E includes instructions that, when executed by one or more of camera processor 114, CPU 116, and GPU 118, cause one or more of camera processor 114, CPU 116, and GPU 118 to establish a mesh rendering process and apply the established mesh rendering process to the predicted depth map and pose data to generate one or more of mesh data representing a mesh of the image and plane data representing one or more planes of the image.

[0049] Instruction memory 132 may also store instructions that, when executed by one or more of camera processor 114, CPU 116, and GPU 118, cause one or more of camera processor 114, CPU 116, and GPU 118 to perform additional image processing operations on the captured image, such as one or more of automatic gain (AG), automatic white balance (AWB), color correction, or scaling operations.

[0050] As Figure 1 illustrated, the various components of imaging device 100 may be configured to communicate with each other across bus 135. Bus 135 may include any one of a variety of bus structures, such as a third-generation bus (e.g., HyperTransport bus or InfiniBand bus), a second-generation bus (e.g., Advanced Graphics Port bus, Peripheral Component Interconnect (PCI) Express bus, or Advanced eXtensible Interface (AXI) bus), or another type of bus or device interconnect. It will be appreciated that Figure 1 the specific configuration of the illustrated components and the communication interfaces between the different components are merely exemplary, and other configurations of the components and / or other image processing systems having the same or different components may be configured to implement the operations and processes of the present disclosure.

[0051] Memory controller 124 may be communicatively coupled to system memory 130 and instruction memory 132. Memory controller 124 may facilitate data transfer to and from system memory 130 and / or instruction memory 132. For example, memory controller 124 may receive memory read and write commands from, such as, camera processor 114, CPU 116, or GPU 118, and service such commands to provide memory services to system memory 130 and / or instruction memory 132. Although memory controller 124 is illustrated as separate from both CPU 116 and system memory 130 in the example of Figure 1 , in other examples, some or all of the functionality of memory controller 124 may be implemented on one or both of CPU 116 and system memory 130. Similarly, some or all of the functionality of memory controller 124 may be implemented on one or both of GPU 118 and instruction memory 132.

[0052] Camera processor 114 may also be configured by the instructions executed to analyze image pixel data and store the resulting image (e.g., the pixel value of each image pixel in the image pixels) in system memory 130 via memory controller 124. Each image in the image may be further processed to generate a final image for display. For example, GPU 118 or some other processing unit (including camera processor 114 itself) may execute any of the machine learning processes described herein as well as any color correction, white balance, blending, compositing, rotation, digital zooming, or any other operation to generate final image content for display (e.g., on display 128).

[0053] The CPU 116 may include a general-purpose or special-purpose processor that controls the operation of the imaging device 100. A user may provide input to the imaging device 100 to cause the CPU 116 to execute one or more software applications. Software applications executed by the CPU 116 may include, for example, VR applications, AR applications, camera applications, graphics editing applications, media player applications, video game applications, graphical user interface applications, or another program. For example, and when executed by the CPU 116, a camera application may allow control of various settings of the camera 115, for example, via input provided to the imaging device 100 through the user interface 122. Examples of the user interface 122 include, but are not limited to, a pressure-sensitive touchscreen unit, a keyboard, a mouse, or an audio input device (such as a microphone). For example, the user interface 122 may receive input from the user to select an object in the field of view of the camera 115 (e.g., for VR or AR game applications), adjust a desired zoom level (e.g., digital zoom level), change the aspect ratio of the image data, record a video, take a snapshot while recording a video, apply a filter when capturing an image, select a region of interest (ROI) for AF (e.g., PDAF), AE, AG, or AWB operations, record a slow-motion video or an ultra-slow-motion video, apply night-shot settings, and / or capture panoramic image data, among other examples.

[0054] In some examples, one or more of the CPU 116 and the GPU 118 cause output data (e.g., a focused image of an object, a captured image, etc.) to be displayed on the display 128. In some examples, the imaging device 100 sends output data to another computing device (such as a server (e.g., a cloud-based server) or the user's handheld device (e.g., a mobile phone)) via the transceiver 119. For example, the imaging device 100 may capture an image (e.g., using the camera 115) and may send the captured image to another computing device. The computing device that receives the captured image may apply any of the machine learning processes described herein to generate depth map values representing a predicted depth map of the captured image and may send the depth map values to the imaging device 100. The imaging device 100 may render an output image (e.g., a 3D image) based on the received depth map values and display the output image on the display 128 (e.g., via the display interface 126).

[0055] Figure 2 is illustrative Figure 1Diagram of an exemplary portion of the imaging device 100. In this example, the imaging device 100 includes an image encoder engine 202, a sparse encoder engine 206, a decoder engine 208, and a system memory 130. As described herein, in some examples, each of the image encoder engine 202, the sparse encoder engine 206, and the decoder engine 208 may include instructions that, when executed by one or more of the camera processor 114, the CPU 116, and the GPU 118, cause one or more of the camera processor 114, the CPU 116, and the GPU 118 to perform corresponding operations. For example, the image encoder engine 202 may include encoder model data 132A, the sparse encoder engine 206 may include sparse encoder model data 132B, and the decoder engine 208 may include decoder model data 132C. In some examples, one or more of the image encoder engine 202, the sparse encoder engine 206, and the decoder engine 208 may be implemented in hardware, such as within one or more FPGAs, ASICs, digital circuits, or any other suitable hardware, or a combination of hardware and software.

[0056] In this example, the image encoder engine 202 receives input image data 201. The input image data 201 may characterize an image captured by a camera, such as camera 115. In some examples, the input image data 201 characterizes a color image. For example, a color image may include a red channel, a green channel, and a blue channel, where each channel includes pixels of the image for the corresponding color. In some examples, the input image data 201 characterizes a monochrome image. For example, a monochrome image may include grayscale pixels for multiple color channels, such as grayscale pixel values for the corresponding red channel, green channel, and blue channel.

[0057] In addition, the image encoder engine 202 may apply an encoding process (e.g., a first encoding process) to the input image data 201 to generate encoder feature data 203 that characterizes a set of image features. In some examples, the encoder engine 202 establishes an encoder of a neural network (e.g., a trained neural network), such as a CNN-based encoder (e.g., Xception convolutional neural network) or a DNN-based encoder, and applies the established encoder to the input image data 201 to generate the encoder feature data 203. For example, the image encoder engine 202 may generate one or more image data vectors based on the pixel values of the input image data 201 and apply a deep neural network to the image data vectors to generate the encoder feature data 203 that characterizes the features of the image.

[0058] The sparse encoder engine 206 receives the sparse depth map 205. The sparse depth map 205 may include sparse depth values generated based on the image characterized by the input image data 201. For example, and as described herein, the executed sparse point engine may generate sparse depth values of the image based on feature point data (e.g., key points), where the feature point data is generated by the executed head tracker engine based on the image.

[0059] The sparse encoder engine 206 applies an encoding process (e.g., a second encoding process) to the sparse depth map 205 to generate sparse encoder feature data 207 that characterizes sparse features of the image. For example, the sparse encoder engine 206 may establish an encoder of a neural network (e.g., a trained deep neural network), such as a CNN-based encoder or a DNN-based encoder, and may apply the established encoder to the sparse depth map 205 to generate sparse encoder feature data 207 that characterizes sparse features of the image. For example, the sparse encoder engine 206 may generate one or more sparse data vectors based on the sparse values of the sparse depth map 205, and may apply a deep neural network to the sparse data vectors to generate sparse encoder feature data 207 that characterizes sparse features of the image. The sparse encoder feature data 207 may include, for example, sparse feature values that characterize one or more detected features.

[0060] The decoder engine 208 receives the encoder feature data 203 from the encoder engine 202 and the sparse encoder feature data 207 from the sparse encoder engine 206. Additionally, the decoder engine 208 may apply a decoding process to the encoder feature data 203 and the sparse encoder feature data 207 to generate a predicted depth map 209 that characterizes depth values of the image. For example, the decoder engine 208 may establish a decoder of a neural network (e.g., a trained neural network), such as a decoder corresponding to the encoder of the deep neural network established by the encoder engine 202 or the sparse encoder engine 206 (e.g., a CNN-based decoder or a DNN-based decoder), and may apply the established decoder to the encoder feature data 203 and the sparse encoder feature data 207 to generate the predicted depth map 209. The decoder engine 208 may store the predicted depth map 209 in a data repository (such as the system memory 130). An output image (such as a 3D image) may be rendered based on the predicted depth map 209.

[0061] In some instances, the sparse encoder engine 206 provides the skip connection 250 to the decoder engine 208. For example, the sparse encoder engine 206 may provide the output of one or more layers (e.g., convolutional layers) of the established encoder to the decoder engine 208, and the decoder engine 208 may provide one or more outputs as inputs to the corresponding layers (e.g., convolutional layers) of the established decoder. In this example, the decoder engine 208 may generate the predicted depth map 209 based on the encoder feature data 203, the sparse encoder feature data 207, and the skip connection 250.

[0062] For example, Figure 4 Illustrate a machine learning model 400 including an image encoder 402, a sparse encoder 404, and a decoder 406. The image encoder 402 includes a plurality of convolutional layers, such as convolutional layers 402A, 402B, 402C, and 402D. Although four convolutional layers are illustrated, in some embodiments, the number of convolutional layers may be greater than four or less than four. The image encoder 402 may receive the input image data 201 and may apply the first convolutional layer 402A to the input image data 201 to generate a first layer output. The image encoder 402 may then apply the second convolutional layer 402B to the first layer output to generate a second layer output. Similarly, the image encoder 402 may apply the third convolutional layer 402C to the second layer output to generate a third layer output, and apply the fourth convolutional layer 402D to the third layer output to generate the encoder feature data 203.

[0063] Although not illustrated for simplicity, the image encoder 402 may also include corresponding non-linear layers (e.g., sigmoid, rectified linear unit (ReLU), etc.) and pooling layers. For example, the fourth output may pass through a pooling layer to generate the encoder feature data 203.

[0064] The sparse encoder 404 also includes a plurality of convolutional layers, such as convolutional layers 404A, 404B, 404C, and 404D. Although four convolutional layers are illustrated, in some embodiments, the number of convolutional layers may be greater than four or less than four. The sparse encoder 404 may receive the sparse depth map 205 and may apply the first convolutional layer 404A to the sparse depth map 205 to generate a first layer output. The sparse encoder 404 may then apply the second convolutional layer 404B to the first layer output to generate a second layer output. Similarly, the sparse encoder 404 may apply the third convolutional layer 404C to the second layer output to generate a third layer output, and apply the fourth convolutional layer 404D to the third layer output to generate a fourth layer output, i.e., the sparse encoder feature data 207.

[0065] Although not illustrated for simplicity, the sparse encoder 404 may also include corresponding non - linear layers (e.g., sigmoid, rectified linear unit (ReLU), etc.) and pooling layers. For example, the output of the fourth layer may pass through a pooling layer to generate sparse encoder feature data 207.

[0066] In addition, the output of each of the convolutional layers 404A, 404B, 404C, and 404D is passed as a skip connection to the corresponding layer of the decoder 406. As shown, the decoder 406 includes a plurality of convolutional layers, including convolutional layers 406A, 406B, 406C, and 406D. Although four convolutional layers are illustrated, in some embodiments, the number of convolutional layers may be greater than four or less than four. The decoder 406 receives the encoder feature data 203 from the encoder 402 and the sparse encoder feature data 207 from the sparse encoder 404, and applies the first convolutional layer 406A to the encoder feature data 203 and the sparse encoder feature data 207 to generate a first - layer output.

[0067] In addition, the output of the first convolutional layer 404A of the sparse encoder 404 is provided as a skip - connection input 250A to the fourth convolutional layer 406D of the decoder 406. Similarly, the output of the second convolutional layer 404A of the sparse encoder 404 is provided as a skip - connection input 250B to the third convolutional layer 406C of the decoder 406. In addition, the output of the third convolutional layer 404C of the sparse encoder 404 is provided as a skip - connection input 250C to the second convolutional layer 406B of the decoder 406. Although three skip connections are illustrated, in some embodiments, the number of skip connections may be greater than or less than three. In addition, at least in some embodiments, the sparse encoder 404 and the decoder 406 include the same number of convolutional layers. In some embodiments, the sparse encoder 404 may include more or fewer convolutional layers compared to the decoder 406.

[0068] Thus, the first convolutional layer 406A generates a first - layer output based on the encoder feature data 203 and the sparse encoder feature data 207. The second convolutional layer 406B generates a second - layer output based on the first - layer output and the skip - connection input 250C. The third convolutional layer 406C generates a third - layer output based on the second - layer output and the skip - connection input 250B. Finally, the fourth convolutional layer 406D generates a fourth - layer output, i.e., the predicted depth map 209, based on the third - layer output and the skip - connection input 250A.

[0069] Although not illustrated for simplicity, decoder 406 may also include corresponding non-linear layers (e.g., sigmoid, rectified linear unit (ReLU), etc.) and upsampling layers, as well as a flattening layer, a fully-connected layer, and a softmax layer. For example, the fourth layer output of the fourth convolutional layer 406D may pass through a flattening layer, a fully-connected layer, and a softmax layer before being provided as the predicted depth map 209. In some examples, decoder 406 may receive skip connections from image encoder 402, in addition to or instead of skip connections 250A, 250B, 250C received from sparse encoder 404.

[0070] Figure 3 is an illustration Figure 1 of an exemplary portion of imaging device 100. In this example, imaging device 100 includes camera 115, encoder engine 202, sparse encoder engine 206, decoder engine 208, head tracker engine 220, sparse point engine 225, and rendering engine 230. As described herein, in some examples, each of image encoder engine 202, sparse encoder engine 206, decoder engine 208, head tracker engine 302, sparse point engine 304, and rendering engine 306 may include instructions that, when executed by one or more of camera processor 114, CPU 116, and GPU 118, cause one or more of camera processor 114, CPU 116, and GPU 118 to perform corresponding operations. For example, and as described herein, image encoder engine 202 may include encoder model data 132A, sparse encoder engine 206 may include sparse encoder model data 132B, and decoder engine 208 may include decoder model data 132C. Additionally, head tracker engine 302 may include head tracker model data 132D, and rendering engine 306 may include rendering model data 132E.

[0071] In some examples, one or more of image encoder engine 202, sparse encoder engine 206, decoder engine 208, head tracker engine 302, sparse point engine 304, and rendering engine 306 may be implemented in hardware, such as within one or more FPGAs, ASICs, digital circuits, or any other suitable hardware, or a combination of hardware and software.

[0072] In this example, camera 115 captures an image through corresponding lens 113A, such as an image of the field of view of one of sensors 112. Camera processor 114 may generate input image data 201 representative of the captured image and provide the input image data 201 to encoder engine 202 and head tracker engine 220. In some examples, the input image data 201 may represent a color image. For example, the image may include a red channel, a green channel, and a blue channel, where each channel includes pixels of the image for the corresponding color. In some examples, the input image data 201 represents a monochrome image. For example, a monochrome image may include grayscale pixel values of a single channel or grayscale pixel values of each of multiple channels, such as grayscale pixel values corresponding to a red channel, a green channel, and a blue channel.

[0073] As described herein, encoder engine 202 may receive the input image data 201 and may apply an established encoding process to the input image data 201 to generate encoder feature data 203 representative of a set of image features. Additionally, head tracker engine 302 may apply one or more processes to the input image data 201 and, in some examples, to sensor data 311 from one or more sensors 129 (e.g., accelerometer data, gyroscope data, etc.) to generate feature point data 301 representative of image features and may also generate pose data 303 representative of the pose of the user. For example, head tracker engine 302 may employ a Harris corner detector to generate the feature point data 301 representative of key points. In some instances, the feature point data 301 includes 6DoF tracking information (e.g., 6DoF tracking data) as described herein. In some examples, head tracker engine 302 applies one or more processes to the sensor data 311 to generate the feature point data 301. The feature point data 301 may be temporally associated with the time at which camera 115 captures an image. For example, camera 115 may have captured an image while sensor 129 generates the sensor data 311 (from which the feature point data 301 is generated).

[0074] Additionally, sparse point engine 304 may receive the feature point data 301 from head tracker engine 302 and may perform operations to generate sparse depth map 205. For example, sparse point engine 304 may perform operations to map the feature point data 301 (which may include 6DoF tracking information such as 3D depth information) to a two-dimensional space. The sparse depth map 205 may include sparse depth values of the captured image. In some examples, sparse point engine 304 projects 3D feature points from the 6DoF tracking information onto two dimensions to generate the sparse depth map 205.

[0075] Additionally, and as described herein, the sparse encoder engine 206 may receive the sparse depth map 205 from the sparse point engine 304, and may apply an encoding process to the sparse depth map 205 to generate sparse encoder feature data 207 that characterizes sparse features of the image. Additionally, the decoder engine 208 may receive the sparse encoder feature data 207 from the sparse encoder engine 206, and may apply a decoding process to the sparse encoder feature data 207 to generate a predicted depth map 209. For example, the decoder engine 208 may apply a trained decoder of a neural network (such as a CNN-based decoder or a DNN-based decoder) to the sparse encoder feature data 207 to generate a predicted depth map 209. In some instances, the sparse encoder engine 206 provides a skip connection 250 to the decoder engine 208. In these examples, the decoder engine 208 provides the skip connection 250 to the corresponding layer of the established decoding process, and generates a predicted depth map 209, as described herein. In some instances, the image encoder engine 202 provides a skip connection 253 to the decoder engine 208, as a supplement to or an alternative to the skip connection from the sparse encoder engine 206. In these examples, the decoder engine 208 provides the skip connection 250 to the corresponding layer of the established decoding process, and generates a predicted depth map 209.

[0076] The rendering engine 306 may receive the predicted depth map 209 from the decoder engine 208, and receive pose data 303 from the head tracker engine 302. The rendering engine 306 may apply a rendering process to the predicted depth map 209 and the pose data 303 to generate output image data 300 that characterizes an output image (such as a 3D image). For example, the rendering engine may apply a mesh rendering process to the predicted depth map 209 and the pose data 303 to generate mesh data that characterizes a mesh of the image, and may perform one or more plane estimation processes to generate plane data that characterizes one or more planes based on the mesh data. For example, the output image data 330 may include one or more of the mesh data and the plane data. The rendering engine 306 may store the output image data 330 in a data repository, such as within the system memory 130.

[0077] Figure 5 is a flowchart of an exemplary process 500 for determining depth values for objects within an image. For example, one or more computing devices (such as the imaging device 100) may perform one or more operations of the exemplary process 500, as described below with reference to Figure 5 as described.

[0078] See Figure 5, the imaging device 100 may perform any of the processes described herein at block 502 to receive an input image. For example, the camera 115 of the imaging device 100 may capture an image within its field of view. For example, the image may have the environment of a user of the imaging device 100 (e.g., a player wearing a VR headset). At block 504, the imaging device 100 may perform any of the processes described herein to apply a first encoding process to the input image to generate a first set of features. For example, the imaging device 100 may establish a trained CNN- or DNN-based encoder and may apply the trained encoder to the input image to generate image features.

[0079] In addition, at block 506, the imaging device 100 may perform any of the processes described herein to receive sparse depth values that characterize a sparse depth map associated with the input image in time. For example, the imaging device 100 may receive a sparse depth map that characterizes sparse depth values and is associated with the image captured by the camera 115 in time, such as the sparse depth map 205. At block 508, the imaging device 100 may perform any of the processes described herein to apply a second encoding process to the sparse depth values to generate a second set of features. For example, the imaging device 100 may establish a trained encoder of a neural network (e.g., such as a CNN- or DNN-based encoder) and may apply the trained encoder to the sparse depth values to generate sparse features.

[0080] Proceeding to block 510, the imaging device 100 may perform any of the processes described herein to apply a decoding process to the first set of features and the second set of features to generate predicted depth values. For example, as described herein, the imaging device 100 may establish a trained decoder of a neural network (e.g., such as a CNN-based decoder or a DNN-based decoder) and may apply the trained decoder to the first set of features and the second set of features to generate a predicted depth map that characterizes the predicted depth values of the image, such as the predicted depth map 209. In some instances, the second encoding process provides skip connections to the decoding process for determining the predicted depth values, as described herein.

[0081] At block 512, the imaging device 100 may perform any of the processes described herein to store the predicted depth values in a data repository. For example, the imaging device 100 may store the predicted depth values (e.g., the predicted depth map 209) in the system memory 130. As described herein, the imaging device 100 or another computing device may generate an output image (such as a 3D image) based on the predicted depth values and may provide the output image for display.

[0082] Figure 6is a flowchart of an exemplary process 600 for rendering an image based on the determined depth values. For example, one or more computing devices (such as imaging device 100) may perform one or more operations of the exemplary process 600 as described below with reference to Figure 6 as described.

[0083] Starting at block 602, imaging device 100 may perform any of the processes described herein to capture an image (e.g., using camera 115). At block 604, imaging device 100 may perform any of the processes described herein to generate sparse depth values representative of a sparse depth map based on the image. Additionally, and at block 606, imaging device 100 may perform any of the processes described herein to apply a first encoding process (e.g., an encoding process by encoder engine 202) to the image to generate a first set of features. At block 608, imaging device 100 may perform any of the processes described herein to apply a second encoding process (e.g., an encoding process by sparse encoder engine 206) to the sparse depth values to generate a second set of features.

[0084] Proceeding to block 610, imaging device 100 may perform any of the processes described herein to provide skip connection features (e.g., skip connection 250) from each of the multiple layers of the second encoding process to each of the corresponding multiple layers of a decoding process (e.g., a decoding process by decoder engine 208).

[0085] At block 612, imaging device 100 may perform any of the processes described herein to apply a decoding process to the skip connection features, the first set of features, and the second set of features to generate predicted depth values. For example, and as described herein, imaging device 100 may decode encoder feature data 203 and sparse encoder feature data 207 to generate predicted depth map 209.

[0086] Additionally, and at block 614, imaging device 100 may perform any of the processes described herein to render an output image based on the predicted depth values. For example, as described herein, rendering engine 306 may generate output image data 330 based on predicted depth map 209 and in some instances further based on pose data 303. At block 616, imaging device 100 may perform any of the processes described herein to provide the output image for display. For example, imaging device 100 may provide the output image to display interface 126 for display on display 128.

[0087] Figure 7is a flowchart of an exemplary process 600 for training a machine learning process. For example, one or more computing devices (such as imaging device 100) may perform one or more operations of exemplary process 700 as described below with reference to Figure 7 as described.

[0088] Starting at block 702, imaging device 100 may perform any of the processes described herein to apply a first encoding process to an input image to generate a first set of features. The input image may be a training image obtained from a training set of images stored in a data repository (such as system memory 130). At block 704, imaging device 100 may perform any of the processes described herein to apply a second encoding process to sparse depth values to generate a second set of features. The sparse depth values may correspond to a sparse depth map generated for and stored in system memory 130 for the input image. Additionally, and at block 706, imaging device 100 may perform any of the processes described herein to apply a decoding process to the first set of features and the second set of features to generate predicted depth values.

[0089] Proceeding to block 708, imaging device 100 may determine a loss value based on the predicted depth values and the corresponding ground truth. For example, imaging device 100 may compute the value of one or more of berHu, SSIM, Edge, MAE, the average variant with MAE, and the average variant with berHu, mean absolute relative error, root mean square error, mean absolute error, accuracy, recall, precision, F-score, or any other metric. Additionally, and at block 710, imaging device 100 may determine whether training is complete. For example, imaging device 100 may compare each computed loss value to a corresponding threshold to determine whether training is complete. For example, if each computed loss value indicates a loss greater than the corresponding threshold, training is not complete and the process returns to block 702. Otherwise, if each computed loss value indicates a loss not greater than the corresponding threshold, the process proceeds to block 712.

[0090] At block 712, imaging device 100 stores any configuration parameters, hyperparameters, and weights associated with the first encoding process, the second encoding process, and the decoding process in the data repository. For example, imaging device 100 may store any configuration parameters, hyperparameters, and weights associated with the first encoding process in encoder model data 132A within instruction memory 132. Similarly, imaging device 100 may store any configuration parameters, hyperparameters, and weights associated with the second encoding process in sparse encoder model data 132B within instruction memory 132. Imaging device 100 may also store any configuration parameters, hyperparameters, and weights associated with the second encoding process in decoder model data 132C within instruction memory 132.

[0091] Specific implementation examples are further described in the following numbered clauses:

[0092] 1. An apparatus, the apparatus comprising:

[0093] A non-transitory machine-readable storage medium storing instructions; and

[0094] At least one processor coupled to the non-transitory machine-readable storage medium, the at least one processor configured to execute the instructions to:

[0095] Receive three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker;

[0096] Generate sparse depth values based on the three-dimensional feature points;

[0097] Generate predicted depth values based on an image and the sparse depth values; and

[0098] Store the predicted depth values in a data repository.

[0099] 2. The apparatus according to clause 1, wherein the at least one processor is configured to execute the instructions to generate an output image based on the predicted depth values.

[0100] 3. The apparatus according to clause 2, wherein the at least one processor is configured to execute the instructions to generate pose data characterizing a user's pose and generate the output image based on the pose data.

[0101] 4. The apparatus according to any one of clauses 2 to 3, the apparatus comprising an extended reality environment, wherein the at least one processor is configured to execute the instructions to provide the output image for viewing in the extended reality environment.

[0102] 5. The apparatus according to any one of clauses 1 to 4, wherein the at least one processor is configured to execute the instructions to:

[0103] Apply a first encoding process to the image to generate a first set of features;

[0104] Apply a second encoding process to the sparse depth values to generate a second set of features; and

[0105] Apply a decoding process to the first set of features and the second set of features to generate the predicted depth values.

[0106] 6. The apparatus according to clause 5, wherein the at least one processor is further configured to execute the instructions to provide at least one skip connection from the second encoding process to the decoding process.

[0107] 7. The apparatus according to clause 6, wherein the at least one skip connection includes a first skip connection and a second skip connection, and the at least one processor is configured to execute the instructions to:

[0108] provide the first skip connection from the first layer of the second encoding process to the first layer of the decoding process; and

[0109] provide a second skip connection from the second layer of the second encoding process to the second layer of the decoding process.

[0110] 8. The apparatus according to any one of clauses 5 to 7, wherein the at least one processor is configured to execute the instructions to:

[0111] obtain a first parameter from the data repository and establish the first encoding process based on the first parameter;

[0112] obtain a second parameter from the data repository and establish the second encoding process based on the second parameter; and

[0113] obtain a third parameter from the data repository and establish the decoding process based on the third parameter.

[0114] 9. The apparatus according to any one of clauses 1 to 8, wherein the image is a monochrome image.

[0115] 10. The apparatus according to any one of clauses 1 to 9, the apparatus includes at least one camera, wherein the at least one camera is configured to capture the image.

[0116] 11. The apparatus according to any one of clauses 1 to 10, wherein the three-dimensional feature points are generated based on the image.

[0117] 12. A method for adjusting a lens of an imaging device, the method comprising:

[0118] receiving three-dimensional feature points from a six degrees of freedom (6Dof) tracker;

[0119] generating a sparse depth value based on the three-dimensional feature points;

[0120] generating a predicted depth value based on the image and the sparse depth value; and

[0121] storing the predicted depth value in a data repository.

[0122] 13. The method according to clause 12, the method comprising: generating an output image based on the predicted depth value.

[0123] 14. The method according to clause 13, the method comprising: generating an output image based on the predicted depth value.

[0124] 15. The method according to any one of clauses 13 to 14, the method comprising: providing the output image for viewing in an extended reality environment.

[0125] 16. The method according to any one of clauses 12 to 15, the method comprising:

[0126] applying a first encoding process to the image to generate a first set of features;

[0127] applying a second encoding process to the sparse depth value to generate a second set of features; and

[0128] applying a decoding process to the first set of features and the second set of features to generate the predicted depth value.

[0129] 17. The method according to clause 16, the method comprising: providing at least one skip connection from the second encoding process to the decoding process.

[0130] 18. The method according to clause 17, wherein the at least one skip connection includes a first skip connection and a second skip connection, the method comprising:

[0131] providing the first skip connection from the first layer of the second encoding process to the first layer of the decoding process; and

[0132] providing the second skip connection from the second layer of the second encoding process to the second layer of the decoding process.

[0133] 19. The method according to any one of clauses 17 to 18, the method comprising:

[0134] obtaining a first parameter from the data repository and establishing the first encoding process based on the first parameter;

[0135] obtaining a second parameter from the data repository and establishing the second encoding process based on the second parameter;

[0136] and

[0137] obtaining a third parameter from the data repository and establishing the decoding process based on the third parameter.

[0138] 20. The method according to any one of clauses 12 to 19, wherein the image is a monochromatic image.

[0139] 21. The method according to any one of clauses 12 to 20, the method comprising: causing at least one camera to capture the image.

[0140] 22. The method according to any one of clauses 12 to 21, wherein the three-dimensional feature points are generated based on the image.

[0141] 23. A non-transitory machine-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations comprising:

[0142] Receiving three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker;

[0143] Generating sparse depth values based on the three-dimensional feature points;

[0144] Generating predicted depth values based on the image and the sparse depth values; and

[0145] Storing the predicted depth values in a data repository.

[0146] 24. The non-transitory machine-readable storage medium according to clause 23, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising generating an output image based on the predicted depth values.

[0147] 25. The non-transitory machine-readable storage medium according to clause 24, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising generating an output image based on the predicted depth values.

[0148] 26. The non-transitory machine-readable storage medium according to any one of clauses 24 to 25, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising providing the output image for viewing in an extended reality environment.

[0149] 27. The non-transitory machine-readable storage medium according to any one of clauses 23 to 26, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

[0150] Applying a second encoding process to the sparse depth values to generate a second set of features; and

[0151] Apply the decoding process to the first set of features and the second set of features to generate the predicted depth value.

[0152] 28. The non-transitory machine-readable storage medium according to clause 27, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including providing at least one skip connection from the second encoding process to the decoding process.

[0153] 29. The non-transitory machine-readable storage medium according to clause 28, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including:

[0154] Provide the first skip connection from the first layer of the second encoding process to the first layer of the decoding process; and

[0155] Provide the second skip connection from the second layer of the second encoding process to the second layer of the decoding process.

[0156] 30. The non-transitory machine-readable storage medium according to any one of clauses 28 to 29, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including:

[0157] Obtain a first parameter from the data repository and establish the first encoding process based on the first parameter;

[0158] Obtain a second parameter from the data repository and establish the second encoding process based on the second parameter; and

[0159] Obtain a third parameter from the data repository and establish the decoding process based on the third parameter.

[0160] 31. The non-transitory machine-readable storage medium according to any one of clauses 23 to 30, wherein the image is a monochromatic image.

[0161] 32. The non-transitory machine-readable storage medium according to any one of clauses 23 to 31, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including causing at least one camera to capture the image.

[0162] 33. The non-transitory machine-readable storage medium according to any one of clauses 23 to 32, wherein the three-dimensional feature points are generated based on the image.

[0163] 34. An image capture device, the image capture device comprising:

[0164] A component for receiving three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker;

[0165] A component for generating sparse depth values based on the three-dimensional feature points;

[0166] A component for generating predicted depth values based on an image and the sparse depth values; and

[0167] A component for storing the predicted depth values in a data repository.

[0168] 35. The image capture device according to clause 34, the image capture device comprising: a component for generating an output image based on the predicted depth values.

[0169] 36. The image capture device according to clause 35, the image capture device comprising: a component for generating an output image based on the predicted depth values.

[0170] 37. The image capture device according to any one of clauses 35 to 36, the image capture device comprising: a component for generating an output image based on the predicted depth values.

[0171] 38. The image capture device according to any one of clauses 34 to 37, the image capture device comprising:

[0172] A component for applying a first encoding process to the image to generate a first set of features;

[0173] A component for applying a second encoding process to the sparse depth values to generate a second set of features; and

[0174] A component for applying a decoding process to the first set of features and the second set of features to generate the predicted depth values.

[0175] 39. The image capture device according to clause 38, the image capture device comprising: a component for providing at least one skip connection from the second encoding process to the decoding process.

[0176] 40. The image capture device according to clause 39, wherein the at least one skip connection includes a first skip connection and a second skip connection, the image capture device comprising:

[0177] A component for providing the first skip connection from the first layer of the second encoding process to the first layer of the decoding process; and

[0178] A component for providing a second skip connection from the second layer of the second encoding process to the second layer of the decoding process.

[0179] 41. The image capture device according to any one of clauses 39 to 40, the image capture device comprising:

[0180] a component for obtaining a first parameter from the data repository and establishing the first encoding process based on the first parameter;

[0181] a component for obtaining a second parameter from the data repository and establishing the second encoding process based on the second parameter; and

[0182] a component for obtaining a third parameter from the data repository and establishing the decoding process based on the third parameter.

[0183] 42. The image capture device according to any one of clauses 34 to 41, wherein the image is a monochrome image.

[0184] 43. The image capture device according to any one of clauses 34 to 42, the image capture device comprising: a component for causing at least one camera to capture the image.

[0185] 44. The image capture device according to any one of clauses 34 to 43, wherein the three-dimensional feature points are generated based on the image.

[0186] Although the methods described above refer to the illustrated flowcharts, many other ways of performing the actions associated with the methods can be used. For example, the order of some operations can be changed, and some embodiments can omit one or more of the described operations and / or include additional operations.

[0187] In addition, the methods and systems described herein can be embodied at least in part in the form of computer-implemented processes and apparatus for practicing those processes. The disclosed methods can also be embodied at least in part in the form of a tangible non-transitory machine-readable storage medium encoded with computer program code. For example, the method can be embodied in hardware, executable instructions executed by a processor (e.g., software), or a combination of both. The medium can include, for example, RAM, ROM, CD-ROM, DVD-ROM, BD-ROM, hard disk drive, flash memory, or any other non-transitory machine-readable storage medium. When the computer program code is loaded into and executed by a computer, the computer becomes an apparatus for practicing the method. The method can also be embodied at least in part in the form of a computer, with the computer program code loaded into or executed in the computer, such that the computer becomes a special-purpose computer for practicing the method. When implemented on a general-purpose processor, the computer program code segments configure the processor to create specific logic circuits. The method can alternatively be embodied at least in part in a special integrated circuit for performing the method.

[0188] The present subject matter has been described in terms of exemplary embodiments. Since they are merely examples, the claimed invention is not limited to these embodiments. Changes and modifications may be made without departing from the spirit of the claimed subject matter. The claims are intended to cover such changes and modifications.

Claims

1. A device, the device comprising: A non-transitory machine-readable storage medium storing instructions; And At least one processor coupled to the non-transitory machine-readable storage medium, the at least one processor configured to execute the instructions to: Receive three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker; Generate sparse depth values based on the three-dimensional feature points; Generate predicted depth values based on an image and the sparse depth values; And Store the predicted depth values in a data repository.

2. The device according to claim 1, wherein the at least one processor is configured to execute the instructions to generate an output image based on the predicted depth values.

3. The device according to claim 2, wherein the at least one processor is configured to execute the instructions to generate pose data characterizing a user's pose and generate the output image based on the pose data.

4. The device according to claim 2, the device comprising an extended reality environment, wherein the at least one processor is configured to execute the instructions to provide the output image for viewing in the extended reality environment.

5. The device according to claim 1, wherein the at least one processor is further configured to execute the instructions to: Apply a first encoding process to the image to generate a first set of features; Apply a second encoding process to the sparse depth values to generate a second set of features; and Apply a decoding process to the first set of features and the second set of features to generate the predicted depth values.

6. The device according to claim 5, wherein the at least one processor is further configured to execute the instructions to provide at least one skip connection from the second encoding process to the decoding process.

7. The device according to claim 6, wherein the at least one skip connection includes a first skip connection and a second skip connection, wherein the at least one processor is configured to execute the instructions to: Provide the first skip connection from the first layer of the second encoding process to the first layer of the decoding process; and Provide the second skip connection from the second layer of the second encoding process to the second layer of the decoding process.

8. The device according to claim 5, wherein the at least one processor is configured to execute the instructions to: Obtain a first parameter from the data repository and establish the first encoding process based on the first parameter; Obtain a second parameter from the data repository and establish the second encoding process based on the second parameter; And Obtain a third parameter from the data repository and establish the decoding process based on the third parameter.

9. The device according to claim 1, wherein the image is a monochromatic image.

10. The device according to claim 1, the device comprising at least one camera, wherein the at least one camera is configured to capture the image.

11. The device according to claim 1, wherein the three-dimensional feature points are generated based on the image.

12. A method for adjusting a lens of an imaging device, the method comprising: Receiving three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker; Generating sparse depth values based on the three-dimensional feature points; Generating predicted depth values based on an image and the sparse depth values; And Storing the predicted depth values in a data repository.

13. The method according to claim 12, the method comprising: Generating an output image based on the predicted depth values.

14. The method according to claim 13, the method comprising: Generating an output image based on the predicted depth values.

15. The method according to claim 13, the method comprising: Providing the output image for viewing in an extended reality environment.

16. The method according to claim 12, the method comprising: Applying a first encoding process to the image to generate a first set of features; Applying a second encoding process to the sparse depth values to generate a second set of features; And Applying a decoding process to the first set of features and the second set of features to generate the predicted depth values.

17. The method according to claim 16, the method comprising: Providing at least one skip connection from the second encoding process to the decoding process.

18. The method according to claim 17, wherein the at least one skip connection includes a first skip connection and a second skip connection, the method comprising: Providing the first skip connection from a first layer of the second encoding process to a first layer of the decoding process; And Providing a second skip connection from a second layer of the second encoding process to a second layer of the decoding process.

19. The method according to claim 17, the method comprising: Obtaining a first parameter from the data repository and establishing the first encoding process based on the first parameter; Obtaining a second parameter from the data repository and establishing the second encoding process based on the second parameter; And Obtaining a third parameter from the data repository and establishing the decoding process based on the third parameter.

20. A non-transitory machine-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations comprising: Receiving three-dimensional feature points from a six-degree-of-freedom (6Dof) tracker; Generating sparse depth values based on the three-dimensional feature points; Generating predicted depth values based on an image and the sparse depth values; And Storing the predicted depth values in a data repository.