Systems and methods for image reprojection - Patents.com

JP2025504307A5Pending Publication Date: 2025-12-12QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024539048
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2022-12-21
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing image processing techniques are difficult to effectively generate high-quality images from different perspectives, especially when changing the perspectives, there are problems of image distortion and excessive computational load.

Method used

By generating the first motion vector using depth data and image sensors, a second motion vector is generated using grid inversion techniques, thereby modifying the image data to reconstruct the image from different perspectives, including filling the vacant area using machine learning models.

Benefits of technology

It realizes stable reconstruction of images when viewing angle changes, reduces computational load and energy consumption, and improves image quality and efficiency, and is suitable for extending real-time image processing in real-life devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The imaging system receives depth data (corresponding to an environment) from the depth sensor and first image data (representation of the environment) from the image sensor. The imaging system generates a first motion vector corresponding to a change in viewpoint of the representation of the environment in the first image data based on the depth data. The imaging system generates a second motion vector indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint using grid inversion based on the first motion vector. The imaging system generates the second image data by modifying the first image data according to the second motion vector. The second image data includes a second representation of the environment from a different viewpoint than the first image data. Some image reprojection applications (e.g., frame interpolation) can be performed without depth data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application relates to image processing, and more particularly to a system and method for reprojecting a first image captured from a first viewpoint to generate a second image that appears to have been captured from a second viewpoint, for example using grid inversion. [Background technology]

[0002] A camera is a device that uses an image sensor to accept light and capture image frames, such as still images or video frames. Cameras capture images that depict an environment from a viewpoint that corresponds to the camera's field of view.

[0003] An extended reality (XR) device is a device that displays an environment to a user, for example through a head mounted display (HMD) or a mobile handset. The environment is at least partially different from the real-world environment in which the user is located. The user can generally interactively change the view of the environment, for example by tilting or moving the HMD or other device. Virtual reality (VR), augmented reality (AR), and mixed reality (MR) are examples of XR. An XR device can include sensors that capture information from the environment. Summary of the Invention

[0004] In some examples, systems and techniques for image processing are described. In some examples, an imaging system receives depth data (corresponding to an environment). The imaging system receives first image data (including a representation of the environment) captured by an image sensor. The imaging system generates a first motion vector corresponding to a change in viewpoint of the representation of the environment in the first image data based on the depth data. The imaging system generates a second motion vector indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint using grid inversion based on the first motion vector. The imaging system generates the second image data by modifying the first image data according to the first motion vector and / or the second motion vector. The second image data includes a second representation of the environment from a different viewpoint than the first image data. The imaging system outputs the second image data. Some image reprojection applications (e.g., frame interpolation) can be performed without depth data.

[0005] In one example, an apparatus for image processing is provided. The apparatus includes a memory and one or more processors (e.g., implemented in a circuit) coupled to the memory. The one or more processors are configured to receive depth data including depth information corresponding to an environment, receive first image data captured by an image sensor, the first image data including a representation of the environment, generate a first plurality of motion vectors corresponding to a change in viewpoint of the representation of the environment in the first image data based at least on the depth data, generate a second plurality of motion vectors indicative of a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint using grid inversion based on the first plurality of motion vectors, generate second image data by at least partially modifying the first image data corresponding to the second plurality of motion vectors, the second image data including a second representation of the environment from a different viewpoint than the first image data, and output the second image data.

[0006] In another example, a method of image processing is provided, the method including: receiving depth data including depth information corresponding to an environment, receiving first image data captured by an image sensor, the first image data including a representation of an environment, generating a first plurality of motion vectors corresponding to a change in viewpoint of the representation of the environment in the first image data based at least on the depth data, generating a second plurality of motion vectors indicative of a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint using grid inversion based on the first plurality of motion vectors, generating second image data by at least partially correcting the first image data corresponding to the second plurality of motion vectors, the second image data including a second representation of the environment from a different viewpoint than the first image data, and outputting the second image data.

[0007] In another example, a non-transitory computer-readable medium is provided that, when executed by one or more processors, causes the one or more processors to receive depth data including depth information corresponding to an environment; receive first image data captured by an image sensor, the first image data including a representation of the environment; generate a first plurality of motion vectors corresponding to a change in viewpoint of the representation of the environment in the first image data based at least on the depth data; generate a second plurality of motion vectors using grid inversion based on the first plurality of motion vectors indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint; generate second image data by at least partially modifying the first image data corresponding to the second plurality of motion vectors, the second image data including a second representation of the environment from a different viewpoint than the first image data; and output the second image data.

[0008] In another example, an apparatus for image processing is provided that includes: means for receiving first image data captured by an image sensor, the first image data including a representation of an environment, means for generating a first plurality of motion vectors corresponding to a change in viewpoint of the representation of the environment in the first image data based at least on the depth data, means for generating a second plurality of motion vectors using grid inversion based on the first plurality of motion vectors, the second plurality of motion vectors indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint, means for generating second image data by at least partially modifying the first image data corresponding to the second plurality of motion vectors, the second image data including a second representation of the environment from a different viewpoint than the first image data, and means for outputting the second image data.

[0009] In some aspects, the second image data includes an interpolated image configured to depict the environment at a second time between the first time and the third time, and the first image data includes at least one image depicting the environment at at least one of the first time or the third time.

[0010] In some aspects, the first image data includes a plurality of frames of video data including parallax motion, and the second image data includes a stabilized transformation of the plurality of frames of video data that reduces the parallax motion.

[0011] In some aspects, the first image data includes a person viewing the image sensor from a first angle, and the second image data includes a person viewing the image sensor from a second angle that is different from the first angle.

[0012] In some aspects, the change in viewpoint comprises a rotation of the viewpoint about an axis according to an angle. In some aspects, the change in viewpoint comprises a translation of the viewpoint according to a direction and a distance. In some aspects, the change in viewpoint comprises a translation. In some aspects, the change in viewpoint comprises a movement along an axis between an original viewpoint of the representation of the environment in the first image data and a position of an object in the environment, at least a portion of which is represented in the first image data.

[0013] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include identifying one or more gaps in the second image data based on the one or more gaps in the second plurality of motion vectors, and modifying the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0014] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors, and modifying the second image data by at least partially filling the one or more gaps in the second image data using inpainting before outputting the second image data.

[0015] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors, and modifying the second image data by at least partially filling the one or more gaps in the second image data using inpainting using the one or more trained machine learning models before outputting the second image data.

[0016] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include identifying one or more conflicts in the second image data based on one or more conflict values ​​from the first image data in the second plurality of motion vectors, and selecting one of the one or more conflict values ​​from the first image data based on the movement data associated with the second plurality of motion vectors.

[0017] In some aspects, the depth information includes a three-dimensional representation of the environment from a first perspective. In some aspects, the depth data is received from at least one depth sensor, the at least one depth sensor including at least one time-of-flight sensor.

[0018] In some aspects, outputting the second image data includes causing the second image data to be displayed using at least one display. In some aspects, outputting the second image data includes causing the second image data to be transmitted to at least a receiving device using at least the communication interface.

[0019] In some aspects, the representation of the environment in the first image data depicts the environment from a first perspective, and the change in perspective is a change between the first perspective and a different perspective that corresponds to a second representation of the environment in the second image data.

[0020] In some aspects, the change in viewpoint includes at least one of a parallax movement of the viewpoint or a rotation of the viewpoint about an axis, and further includes receiving, via the user interface, an indication of one of a distance of the parallax movement of the viewpoint, or an indication of an angle or axis of the rotation of the viewpoint.

[0021] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include identifying one or more gaps in the second plurality of motion vectors that cause one or more gaps in the second image data based on the one or more gaps at respective endpoints of the first plurality of motion vectors, and modifying the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0022] In some aspects, the device is, is a part of, and / or includes a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a head mounted display (HMD) device, a wireless communication device, a mobile device (e.g., a mobile phone and / or mobile handset and / or a so-called "smartphone" or other mobile device), a camera, a personal computer, a laptop computer, a server computer, a vehicle or a computing device or component of a vehicle, another device, or a combination thereof. In some aspects, the device includes a camera or multiple cameras for capturing one or more images. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the devices described above may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors.

[0023] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used independently to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0024] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0025] Exemplary embodiments of the present application are described in detail below with reference to the following drawings: [Brief description of the drawings]

[0026] [Figure 1] 1 is a block diagram illustrating an example architecture of an image capture and processing system, according to some examples. [Diagram 2] 1 is a block diagram illustrating an example architecture of an imaging system for performing reprojection operations for various applications, in accordance with some examples. [Figure 3A] FIG. 1 is a perspective view illustrating a head mounted display (HMD) used as an extended reality (XR) system, according to some examples. [Figure 3B] FIG. 3B is a perspective view illustrating the head mounted display (HMD) of FIG. 3A being worn by a user, according to some examples. [Figure 4A] FIG. 1 is a perspective view illustrating the front of a mobile handset that includes a forward-facing camera and can be used as an extended reality (XR) system, according to some examples. [Figure 4B] FIG. 1 is a perspective view illustrating the rear of a mobile handset that includes a rear-facing camera and can be used as an extended reality (XR) system, according to some examples. [Diagram 5] FIG. 1 is a block diagram illustrating an example of grid inversion, in accordance with some examples. [Figure 6] FIG. 1 is a conceptual diagram illustrating an example of depth-based reprojection, according to some examples. [Figure 7] 1 is a conceptual diagram illustrating an example of a time warp performed by a time warp engine, according to some examples. [Figure 8] FIG. 2 is a conceptual diagram illustrating an example of depth sensor support performed by a depth sensor support engine, according to some examples. [Figure 9] FIG. 2 is a conceptual diagram illustrating an example of 3D stabilization performed by a 3D stabilization engine, according to some examples. [Figure 10] 1 is a conceptual diagram illustrating an example of 3D zoom (or cinematic zoom) performed by a 3D zoom engine, according to some examples. [Figure 11] FIG. 2 is a conceptual diagram illustrating an example of reprojection performed by a reprojection SAT engine, according to some examples. [Figure 12] FIG. 2 is a conceptual diagram illustrating an example of head pose correction performed by a head pose correction engine, according to some examples. [Figure 13] FIG. 1 is a conceptual diagram illustrating an example of XR late-stage reprojection performed by an XR late-stage reprojection engine, according to some examples. [Figure 14] 4 is a conceptual diagram illustrating examples of special effects performed by a special effects engine, according to some examples. [Figure 15] FIG. 1 is a conceptual diagram illustrating an image reprojection transformation based on matrix operations, according to some examples. [Figure 16] 1 is a block diagram illustrating a grid inversion transformation and a 3D transformation based on depth data, in accordance with some examples. [Figure 17] FIG. 2 is a block diagram illustrating an image reprojection transformation based on motion vectors, according to some examples. [Figure 18] FIG. 1 is a conceptual diagram illustrating an example of inpainting to address occlusion, according to some examples. [Figure 19] FIG. 1 is a block diagram illustrating an architecture of a reprojection and grid inversion system, in accordance with some examples. [Figure 20] FIG. 13 is a conceptual diagram illustrating an example of a triangular walking motion, according to some examples. [Figure 21] FIG. 1 is a conceptual diagram illustrating an example of occlusion masking, according to some examples. [Figure 22] FIG. 13 is a conceptual diagram illustrating an example of hole filling, according to some examples. [Figure 23] 1 is a conceptual diagram illustrating additional examples of time warping performed by a time warp engine, according to some examples. [Figure 24] 1 is a block diagram illustrating an example architecture of a reprojection engine in some examples of a time warp engine, according to some examples. [Diagram 25] 1 is a block diagram illustrating an example architecture of a reprojection engine with temporal deblurring in some examples of a time warp engine with temporal deblurring, in accordance with some examples. [Figure 26] FIG. 2 is a block diagram illustrating an example architecture of a depth sensor support engine for a time-of-flight (ToF) sensor, in accordance with some examples. [Figure 27] FIG. 1 is a conceptual diagram illustrating an example of adding depth sensor support performed by a depth sensor support engine, according to some examples. [Figure 28] 1 is a block diagram illustrating an example architecture of an imaging system including an image reprojection engine and / or a 3D stabilization engine, in accordance with some examples. [Figure 29] 1 is a conceptual diagram illustrating an additional example of time warping performed using a time warp engine compared to an image without time warp engine processing, according to some examples. [Diagram 30] FIG. 13 is a conceptual diagram illustrating additional examples of 3D stabilization performed by a 3D stabilization engine, according to some examples. [Diagram 31] 1 is a conceptual diagram illustrating additional examples of 3D zoom (or cinematic zoom) performed by a 3D zoom engine, according to some examples. [Diagram 32]FIG. 13 is a conceptual diagram illustrating additional examples of reprojection performed by a reprojection SAT engine, according to some examples. [Diagram 33] FIG. 11 is a conceptual diagram illustrating additional examples of head pose correction performed by a head pose correction engine, according to some examples. [Diagram 34] FIG. 13 is a conceptual diagram illustrating additional examples of grid inversion, according to some examples. [Diagram 35] FIG. 1 is a conceptual diagram illustrating an example of the use of deep learning-based inpainting, according to some examples. [Diagram 36] A conceptual diagram showing an example of inpainting without deep learning, with some examples. [Figure 37] FIG. 1 is a conceptual diagram illustrating an example of the use of edge and depth filters on edges, in accordance with some examples. [Figure 38] FIG. 1 is a conceptual diagram illustrating an example of reprojection, according to some examples. [Figure 39] FIG. 2 is a block diagram illustrating an example of a neural network that can be used for media processing operations, according to some examples. [Diagram 40] 1 is a flow diagram illustrating a process for media processing, according to some examples. [Diagram 41] FIG. 1 illustrates an example of a computing system for implementing certain aspects described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0027] Specific aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments can be practiced without these specific details. The figures and descriptions are not intended to be limiting.

[0028] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0029] A camera is a device that uses an image sensor to accept light and capture image frames, such as still images or video frames. The terms "image," "image frame," and "frame" are used interchangeably herein. A camera may be configured with various image capture and image processing settings. Different settings result in images with different appearances. Some camera settings, such as ISO, exposure time, aperture size, F / stop, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters may be applied to an image sensor to capture one or more image frames. Other camera settings, such as contrast, brightness, saturation, sharpness, levels, curves, or color changes, may constitute post-processing of one or more image frames. For example, settings or parameters may be applied to a processor (e.g., an image signal processor or ISP) to process one or more image frames captured by the image sensor.

[0030] A depth sensor is a sensor that measures depth, range, or distance from the depth sensor to one or more portions of the environment in which the depth sensor resides. Examples of depth sensors include Light Detection and Ranging (LIDAR) sensors, Radio Detection and Ranging (RADAR) sensors, Acoustic Detection and Ranging (SODAR) sensors, Acoustic Navigation and Ranging (SONAR) sensors, Time of Flight (ToF) sensors, structured light sensors, or combinations thereof. Depth data captured by a depth sensor may include a point cloud, a 3D model, and / or a depth image.

[0031] An extended reality (XR) system or device can provide virtual content to a user and / or combine a real-world view of a physical environment (scene) with a virtual environment (including virtual content). The XR system facilitates user interaction with such a combined XR environment. The real-world view can include real-world objects (also called physical objects), such as people, vehicles, buildings, tables, chairs, and / or other real-world or physical objects. The XR system or device can facilitate interaction with different types of XR environments (e.g., a user can interact with an XR environment using an XR system or device). XR systems can include virtual reality (VR) systems that facilitate interaction with an AR environment, augmented reality (AR) systems that facilitate interaction with an MR environment, mixed reality (MR) systems that facilitate interaction with an MR environment, and / or other XR systems. Examples of XR systems or devices include head-mounted displays (HMDs), smart glasses, among others. In some cases, the XR system can track parts of the user (e.g., the user's hands and / or fingertips) to allow the user to interact with items of virtual content.

[0032] The imaging system can include a camera depth sensor and an image sensor. The depth sensor captures depth data including depth information corresponding to the environment, such as a point cloud, a 3D model, a depth image, a set of disparity values, and / or a 3D representation of the environment. The image sensor captures first image data including a 2D representation of the environment.

[0033] The imaging system uses the depth data to generate a first set of motion vectors, the first set of motion vectors corresponding to a change in viewpoint of a representation of the environment in the first image data from a first viewpoint to a second viewpoint.

[0034] The imaging system applies grid inversion to the first set of motion vectors to generate a second set of motion vectors, the second set of motion vectors indicating respective distances moved by respective pixels of the representation of the environment in the first image data for a change in viewpoint from the first viewpoint to the second viewpoint. In some cases, to apply the grid inversion, the imaging system resolves conflicts with the grid inversion by prioritizing larger motions over smaller motions and / or by prioritizing motion of objects closer in the environment over motion of objects more distant in the environment. In some cases, to apply the grid inversion, the imaging system uses interpolation to fill in missing areas.

[0035] The imaging system generates the second image data by modifying the image data according to the second set of motion vectors. For example, the imaging system can modify the image data according to the second set of motion vectors by moving pixel data for each pixel of the representation of the environment in the first image data by a respective distance indicated by the second set of motion vectors. The second image data includes a second representation of the environment from a different perspective than the first image data. The imaging system outputs the second image data, for example, by displaying the second image data or by transmitting the second image data to a receiving device.

[0036] The viewpoint change, performed by generating second image data through modification of the first image data based on a second set of motion vectors based on the grid inversion, has a variety of useful applications. For example, the viewpoint change can be used for 3D stabilization of video data to reduce or eliminate parallax shifts that may occur, for example, by a user's hands holding the camera unsteadily and / or by the user's steps. The viewpoint change can be used for frame interpolation to increase the effective frame rate of a video by generating intermediate frames between two existing frames. The viewpoint change can be used for a "3D zoom" effect to scale the foreground of an environment more quickly than the background of the environment to appear more similar to a true forward movement in the environment rather than upscaling. The viewpoint change can be used to accommodate an offset between two sensors (e.g., two cameras, a camera and a depth sensor, etc.). The viewpoint change can be used for head pose correction, for example, to make the camera appear to be at the same height as a person's head when in fact the camera is above or on top of the person, as is often the case in video conferencing. The viewpoint change can be used for XR to quickly simulate different viewpoints on an environment even if the different viewpoints have not finished rendering. The change in viewpoint can be used for a variety of special effects, such as an effect that simulates rotation around an object in the scene.

[0037] In some examples, systems and techniques for image processing are described. In some examples, an imaging system receives depth data (corresponding to an environment) captured by a depth sensor, and the imaging system receives first image data (representation of the environment) captured by the image sensor. The imaging system generates a first motion vector corresponding to a change in viewpoint of the representation of the environment in the first image data based on the depth data. The imaging system generates a second motion vector indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint using grid inversion based on the first motion vector. The imaging system generates the second image data by modifying the first image data according to the first motion vector and / or the second motion vector. The second image data includes a second representation of the environment from a different viewpoint than the first image data. The imaging system outputs the second image data.

[0038] The imaging systems and techniques described herein provide several technical improvements over conventional image processing systems. For example, the image processing systems and techniques described herein can provide reprojection to a different viewpoint for any translational and / or rotational movement of the viewpoint. The image processing systems and techniques described herein can use this reprojection and the grid inversion techniques that support it for a variety of applications, including using optical flow to improve video frame quality, aligning depth and image data to overcome offset distances between two sensors, 3D depth-based video stabilization, 3D depth-based zoom (also called cinematic zoom), aligning image data from two different cameras to overcome offset distances between two sensors, head pose compensation, late reprojection for extended reality (XR), special effects, or combinations thereof. The use of grid inversion provides increased efficiency, reduced computational load, reduced power usage, reduced heat generation, and reduced need for heat dissipating components.

[0039] Various aspects of the application are described with respect to the figures. FIG. 1 is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components used to capture and process images of one or more scenes (e.g., an image of a scene 110). The image capture and processing system 100 can capture a standalone image (or photo) and / or can capture a video including multiple images (or video frames) in a particular sequence. A lens 115 of the system 100 faces the scene 110 and accepts light from the scene 110. The lens 115 bends the light toward the image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130. In some examples, the scene 110 is a scene in an environment. In some examples, the scene 110 is a scene of at least a portion of a user. For example, the scene 110 can be a scene of one or both of the user's eyes and / or at least a portion of the user's face.

[0040] The one or more controls 120 may control exposure, focus, and / or zoom based on information from image sensor 130 and / or based on information from image processor 150. The one or more controls 120 may include multiple mechanisms and components. For example, the control 120 may include one or more exposure controls 125A, one or more focus controls 125B, and / or one or more zoom controls 125C. The one or more controls 120 may include additional controls beyond those shown, such as controls that control analog gain, flash, HDR, depth of field, and / or other image capture characteristics.

[0041] The focus control mechanism 125B of the control mechanism 120 can obtain the focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or servo, thereby adjusting the focus. In some cases, additional lenses, such as one or more microlenses above each photodiode of the image sensor 130, may be included in the system 100, each of which bends light received from the lens 115 toward a corresponding photodiode before the light reaches the photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus settings may be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus settings may be referred to as image capture settings and / or image processing settings.

[0042] The exposure control 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control 125A stores the exposure setting in a memory register. Based on this exposure setting, the exposure control 125A can control the size of the aperture (e.g., aperture size or F / stop), the duration the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.

[0043] The zoom control 125C of the control mechanism 120 can obtain the zoom setting. In some examples, the zoom control 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control 125C can control the focal length of the lens assembly by actuating one or more motors or servos to move one or more of the lenses relative to each other. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly may include a parfocal zoom lens or a variable focus zoom lens. In some examples, the lens assembly may include a focusing lens (which may be the lens 115 in some cases) that first accepts light from the scene 110, and then the light passes through an infinity focus zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before the light reaches the image sensor 130. In some cases, an afocal zoom system may include two positive (e.g., converging, convex) lenses of equal or similar focal lengths (e.g., within a threshold difference) with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control 125C moves one or more of the lenses in the afocal zoom system, such as the negative lens and one or both of the positive lenses.

[0044] The image sensor 130 includes one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures an amount of light that ultimately corresponds to a particular pixel in the image produced by the image sensor 130. In some cases, different photodiodes may be covered by different color filters, and thus may measure light that matches the color of the filter covering the photodiode. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, and each pixel of the image is generated based on red light data from at least one photodiode covered by a red color filter, blue light data from at least one photodiode covered by a blue color filter, and green light data from at least one photodiode covered by a green color filter. Other types of color filters may use yellow, magenta, and / or cyan (also called "emerald") color filters instead of or in addition to red, blue, and / or green filters. Some image sensors may be completely devoid of color filters, and instead use different photodiodes (possibly stacked vertically) throughout the pixel array. Different photodiodes across the pixel array can have different spectral sensitivity curves and therefore respond to different wavelengths of light. Monochrome image sensors may also lack color filters and therefore no color depth.

[0045] In some cases, image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching some photodiodes or portions of some photodiodes at some times and / or from some angles, which may be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier for amplifying an analog signal output by the photodiode and / or an analog-to-digital converter (ADC) for converting an analog signal output from the photodiode (and / or amplified by the analog gain amplifier) ​​to a digital signal. In some cases, instead or in addition, some components or functions discussed with respect to one or more of control mechanisms 120 may be included within image sensor 130. The image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active-pixel sensor (APS), a complimentary metal-oxide semiconductor (CMOS), an n-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0046] Image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other types of processors 4110 discussed with respect to computing system 4100. Host processor 152 may be a digital signal processor (DSP) and / or other types of processors. In some implementations, image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes host processor 152 and ISP 154. In some cases, the chip may include one or more input / output ports (e.g., input / output (I / O) ports 156), central processing units (CPUs), graphics processing units (GPUs), broadband modems (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth, Global Positioning System (GPS), etc.), any combination thereof, and / or other components.The I / O ports 156 may include any suitable input / output ports or interfaces according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a serial General Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (e.g., a MIPI CSI-2 physical (PHY) layer port or interface, etc.), an Advanced High-performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In one illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port and the ISP 154 may communicate with the image sensor 130 using a MIPI port.

[0047] Image processor 150 may perform several tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form HDR images, image recognition, object recognition, feature recognition, accepting input, managing output, managing memory, or some combination thereof. Image processor 150 may store image frames and / or processed images in random access memory (RAM) 140 and / or 4120, read-only memory (ROM) 145 and / or 4125, a cache, a memory unit, another storage device, or some combination thereof.

[0048] Various input / output (I / O) devices 160 may be connected to image processor 150. I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a track pad, a touch sensitive screen, a printer, any other output device 4135, any other input device 4145, or some combination thereof. In some cases, captions may be entered into image processing device 105B through a physical keyboard or keypad of I / O device 160 or through a virtual keyboard or keypad of a touch screen of I / O device 160. I / O 160 may include one or more ports, jacks, or other connectors that enable a wired connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. I / O 160 may include one or more wireless transceivers that enable a wireless connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. The peripheral devices may include any of the types of I / O devices 160 previously described, and may themselves be considered I / O devices 160 when coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector.

[0049] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to a camera). In some implementations, the image capture device 105A and the image processing device 105B may be coupled, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly via one or more wireless transceivers. In some implementations, the image capture device 105A and the image processing device 105B may be decoupled from each other.

[0050] As shown in Figure 1, a vertical dashed line divides the image capture and processing system 100 of Figure 1 into two portions representing image capture device 105A and image processing device 105B, respectively. Image capture device 105A includes lens 115, control mechanism 120, and image sensor 130. Image processing device 105B includes image processor 150 (including ISP 154 and host processor 152), RAM 140, ROM 145, and I / O 160. In some cases, some components shown in image capture device 105A, such as ISP 154 and / or host processor 152, may be included within image capture device 105A.

[0051] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smartphone, a mobile phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 wi-fi communication, wireless local area network (WLAN) communication, or any combination thereof. In some implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile handset, a desktop computer, or other computing device.

[0052] Although image capture and processing system 100 is shown as including several components, one skilled in the art will appreciate that image capture and processing system 100 may include many more components than those shown in FIG. 1. The components of image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of image capture and processing system 100 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing image capture and processing system 100.

[0053] 2 is a block diagram illustrating an example architecture of an imaging system 200 for performing reprojection operations for various applications. In some examples, the imaging system 200 includes at least one image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination or combinations thereof. In some examples, the imaging system 200 includes at least one computing system 4100. In some examples, the imaging system 200 includes at least one neural network 3900.

[0054] In some examples, the imaging system 200 includes one or more sensors 205. The sensors 205 capture sensor data that measures and / or tracks information about aspects of an environment the imaging system 200 and / or a user of the imaging system 200 are in. In some examples, the sensors 205 can capture sensor data that measures and / or tracks information about a user's body and / or behavior by the user. In some examples, the sensors 205 include one or more cameras facing at least a portion of the environment and / or the user. The one or more cameras can include one or more image sensors that capture images of at least a portion of the environment and / or the user. In some examples, the sensors 205 include one or more depth sensors facing at least a portion of the environment and / or the user. The one or more depth sensors can capture depth data (e.g., a depth image, a point cloud, a 3D model, a range between the depth sensor and a portion of the environment, a depth between the depth sensor and a portion of the environment, and / or a distance between the depth sensor and a portion of the environment) of at least a portion of the environment and / or the user. In some examples, depth data (such as any of the types of depth data enumerated above) can also be determined using image data from the stereo camera using stereo depth sensing. In some examples, depth data can be determined using image data from the stereo camera by inputting the image data into a trained machine learning model(s) trained based on training data. The training data includes other images captured by the stereo camera (or other cameras in a similar stereoscopic configuration) along with corresponding depth data. In some examples, the sensor 205 includes one or more other types of sensors, such as microphones, accelerometers, gyroscopes, positioning receivers, inertial measurement units (IMUs), biometric sensors, or combinations thereof. In FIG. 2, the one or more sensors 205 are depicted as a camera icon and a microphone icon.

[0055] The sensors 205 may include one or more cameras, image sensors, microphones, heart rate monitors, oximeters, biometric sensors, positioning transceivers, inertial measurement units (IMUs), accelerometers, gyroscopes, gyrometers, barometers, thermometers, altimeters, depth sensors, other sensors described herein, or combinations thereof. Examples of depth sensors include Light Detection and Ranging (LIDAR) sensors, Radio Detection and Ranging (RADAR) sensors, Acoustic Detection and Ranging (SODAR) sensors, Acoustic Navigation and Ranging (SONAR) sensors, Time of Flight (ToF) sensors, structured light sensors, or combinations thereof. Examples of positioning receivers include Global Navigation Satellite System (GNSS) receivers, Global Positioning System (GPS) receivers, cellular signal transceivers, Wi-Fi transceivers, Wireless Local Area Network (WLAN) transceivers, Bluetooth transceivers, beacon transceivers, Near Field Communication (NFC) transceivers, Personal Area Network (PAN) transceivers, Radio Frequency Identification (RFID) transceivers, communication interfaces 4140, or combinations thereof. In some examples, the one or more sensors 205 include at least one image capture and processing system 100, image capture device 105A, image processing device 105B, or combination(s) thereof. In some examples, the one or more sensors 205 include at least one input device 4145 of the computing system 4100. In some implementations, one or more of the sensor(s) 205 may supplement or refine sensor readings from other sensor(s) 205. For example, the application engine 210 and / or the image reprojection engine 215 may use sensor data from positioning receivers, inertial measurement units (IMUs), accelerometers, gyroscopes, and / or other sensors to refine and / or complement the image data and / or depth data.For example, the application engine 210 and / or the image reprojection engine 215 may use such sensor data to assist in determining the attitude (e.g., 3D position coordinates and / or orientation (e.g., pitch, yaw, and / or roll)) of the imaging system 200 within the environment during capture of image data and / or depth data and / or using image stabilization and / or motion compensation.

[0056] In some examples, the imaging system 200 includes a virtual content generator 207 that generates virtual content. The virtual content can include a two-dimensional (2D) shape, a three-dimensional (3D) shape, a 2D object, a 3D object, a 2D model, a 3D model, a 2D animation, a 3D animation, a 2D image, a 3D image, a texture, a portion of another image, a character, a string of characters, or a combination thereof. In some examples, the imaging system 200 can combine the virtual content generated by the virtual content generator 207 with sensor data from the sensor(s) 205 to form media data 285. In some examples, the imaging system 200 can combine the virtual content generated by the virtual content generator 207 with the media data 285. In FIG. 2, the virtual content generated by the virtual content generator 207 is shown as a tetrahedron. In some examples, the virtual content generator 207 includes one or more software elements, such as one or more sets of instructions corresponding to one or more programs, executed on one or more processors of the imaging system 200, such as the processor 4110 of the computing system 4100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the virtual content generator 207 includes one or more hardware elements. For example, the virtual content generator 207 can include a processor, such as the processor 4110 of the computing system 4100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the virtual content generator 207 includes a combination of one or more software elements and one or more hardware elements.

[0057] The imaging system 200 includes a set of application engines 210. The application engines 210 receive media data 285 from the sensor(s) 205. The media data 285 is captured by the sensor(s) 205. The media data 285 can include image data, including, for example, one or more images or portions thereof. The image data can include video data, including, for example, video frames of a video. The media data 285 can include depth data, including, for example, a depth image, a point cloud, a 3D model, a range between a depth sensor and a portion of the environment, a depth between a depth sensor and a portion of the environment, and / or a distance between a depth sensor and a portion of the environment, or a combination thereof. The media data 285 can include audio data, including, for example, audio recorded by one or more microphones of the sensor(s) 205. In some cases, the audio data can include an audio track corresponding to a video of the image data. In some cases, the audio data can be multi-channel audio from multiple microphones of the sensor(s) 205, allowing for separate audio tracks corresponding to audio arriving at the sensor(s) 205 from different directions in the environment, for example. The media data 285 can include attitude data including, for example, a position (e.g., latitude, longitude, and / or altitude) of the imaging system 200 in the environment, an orientation (e.g., pitch, yaw, and / or roll) of the imaging system 200, a rate of movement of the imaging system 200, an acceleration of the imaging system 200, a velocity of the imaging system 200, a momentum of the imaging system 200, a rotation of the imaging system 200, or a combination thereof. In some examples, the attitude data can be captured using positioning receivers, inertial measurement units (IMUs), accelerometers, and / or gyroscopes of the imaging system 200. In some examples, the imaging system 200 can infer aspects of the attitude data and refine the attitude data based on an attitude determination based on other types of media data 285, such as image data, depth data, and / or audio data.

[0058] The application engine 210 includes an image reprojection engine 215 having a motion vector engine 220 and a grid inversion engine 225. The motion vector engine 220 of the image reprojection engine 215 can determine and / or generate a first set of motion vectors corresponding to a movement from a first viewpoint of the environment to a second viewpoint of the environment. In some examples, the motion vector engine 220 can identify or generate a 3D representation of the environment based on depth data captured by a depth sensor of the sensor(s) 205 and / or image data captured by an image sensor of the sensor(s) 205. The motion vector engine 220 can rotate, translate, and / or transform the 3D representation of the environment from one representing the environment from a first viewpoint to one representing the environment from a second viewpoint. The motion vector engine 220 can determine a first set of motion vectors based on this change in viewpoint from the first viewpoint to the second viewpoint.

[0059] The motion vectors output by the motion vector engine 220 of the image reprojection engine 215 may be output to a grid inversion engine 225. The grid inversion engine 225 of the image reprojection engine 215 may perform grid inversion on the motion vectors to generate a second set of motion vectors. The image reprojection engine 215 may use the second set of motions to modify at least a subset of the media data 285 to generate the modified media data 290. For example, the image reprojection engine 215 may receive an image of the media data 285 depicting an environment from a third perspective and may apply the second set of motion vectors to the image to generate a modified image of the modified media data 290. The modified image may depict the environment from a fourth perspective. The change from the third perspective to the fourth perspective may be consistent with the change from the first perspective to the second perspective, for example applying a rotation, translation, and / or transformation of the same amount, distance(s), and / or angle(s). For example, in some examples, the change from a first perspective to a second perspective includes a rotation of the perspective according to an angle, and the change from a third perspective to a fourth perspective includes a rotation of the perspective according to the angle. In some examples, the change from a first perspective to a second perspective includes a translation of the perspective according to a direction and distance, and the change from a third perspective to a fourth perspective includes a translation of the perspective according to the direction and distance. In some examples, the change from a first perspective to a second perspective includes a transformation, and the change from a third perspective to a fourth perspective includes a translation of the perspective according to a transformation.

[0060] In some examples, the image reprojection engine 215 includes one or more software elements, such as one or more sets of instructions corresponding to one or more programs, executed on one or more processors of the imaging system 200, such as the processor 4110 of the computing system 4100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the image reprojection engine 215 includes one or more hardware elements. For example, the image reprojection engine 215 can include a processor, such as the processor 4110 of the computing system 4100, the image processor 150, the host processor 152, the ISP 154, and / or a combination thereof. In some examples, the image reprojection engine 215 includes a combination of one or more software elements and one or more hardware elements.

[0061] In some examples, the image reprojection engine 215 includes an ML system(s) and / or trained ML model(s) that receive as input the media data 285 from the sensor(s) 205 and / or the virtual content generator 207. The ML system(s) and / or trained ML model(s) output modified media data 290 based on the media data 285 and the virtual content. In some cases, the ML system(s) and / or trained ML model(s) can modify the media data 285 and / or the virtual content such that the modified media data 290 includes a depiction(s) and / or representation(s) of an environment from a perspective that differs from a perspective of the depiction(s) and / or representation(s) of the environment in the media data 285. In some examples, the ML system and / or trained ML model of the image reprojection engine 215 may include one or more neural networks (NNs) (e.g., neural network 3900), one or more convolutional neural networks (CNNs), one or more trained time-delay neural networks (TDNNs), one or more deep networks, one or more autoencoders, one or more deep belief nets (DBNs), one or more recurrent neural networks (RNNs), one or more generative adversarial networks (GANs), one or more other types of neural networks, one or more trained support vector machines (SVMs), one or more trained random forests (RFs), one or more computer vision systems, one or more deep learning systems, or combinations thereof.

[0062] The application engine 210 includes several engines that apply image reprojection by the image reprojection engine 215 (including, for example, a motion vector engine 220 and / or a grid inversion engine 225) in various ways for various applications. These engines of the application engine 210 include a time warp engine 230, a depth sensor support engine 235, a 3D stabilization engine 240, a 3D zoom engine 245, a reprojection SAT engine 250, a head pose correction engine 255, an extended reality (XR) late reprojection engine 260, and a special effects engine 265. The "SAT" in the reprojection SAT engine 250 can refer to a sensor alignment, a spatial alignment transform, or both. The reprojection SAT engine 250 may use a sensor alignment, a spatial alignment transform, or both. These engines of the application engine 210 modify at least a subset of the media data 285 to generate modified media data 290, for example, using image reprojection by the image reprojection engine 215 (e.g., including the motion vector engine 220 and / or the grid inversion engine 225) to do so.

[0063] In some examples, at least one of the application engines 210 includes an ML system(s) and / or trained ML model(s) that receive media data 285 as input from the sensor(s) 205 and / or virtual content generator 207. The ML system(s) and / or trained ML model(s) output modified media data 290 based on the media data 285 and the virtual content. In some cases, the ML system(s) and / or trained ML model(s) can modify the media data 285 and / or the virtual content such that the modified media data 290 includes a depiction(s) and / or representation(s) of an environment from a perspective that differs from the perspective of the depiction(s) and / or representation(s) of the environment in the media data 285. In some examples, at least one ML system and / or trained ML model of application engine 210 may include one or more NNs, one or more CNNs, one or more TDNNs, one or more deep networks, one or more autoencoders, one or more DBNs, one or more RNNs, one or more GANs, one or more trained SVMs, one or more trained RFs, one or more computer vision systems, one or more deep learning systems, or a combination thereof.

[0064] In some examples, the application engine 210 including the image reprojection engine 215 can analyze, process, and / or modify (e.g., to determine motion vectors) the media data 285 having virtual content generated by the virtual content generator 207 embedded in the media data 285. In some examples, the application engine 210 including the image reprojection engine 215 can analyze, process, and / or modify (e.g., to determine motion vectors) the media data 285 without the virtual content generated by the virtual content generator 207 embedded in the media data 285. In some examples, the modified media data 290 output by the application engine 210 including the image reprojection engine 215 can already include the virtual content generated by the virtual content generator 207, for example if the virtual content was embedded in the media data 285 input to the application engine 210. In some examples, the modified media data 290 output by the application engine 210, including the image reprojection engine 215, lacks the virtual content generated by the virtual content generator 207, for example if the virtual content was not incorporated into the media data 285 input to the application engine 210. In such examples, the virtual content generated by the virtual content generator 207 may be added to the modified media data 290 after the modified media data 290 is output by the application engine 210, but before the modified media data 290 is output using the output device(s) 270 and / or transceiver(s) 275.

[0065] In some examples, at least one of the application engines 210 includes one or more software elements, such as one or more instruction sets corresponding to one or more programs, executed on one or more processors of the imaging system 200, such as the processor 4110 of the computing system 4100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, at least one of the application engines 210 includes one or more hardware elements. For example, at least one of the application engines 210 can include a processor, such as the processor 4110 of the computing system 4100, the image processor 150, the host processor 152, the ISP 154, and / or a combination thereof. In some examples, at least one of the application engines 210 includes a combination of one or more software elements and one or more hardware elements.

[0066] In some examples, the imaging system 200 includes one or more output devices 270 configured and capable of outputting the modified media data 290. In some examples, the output device(s) 270 include a display(s) configured and capable of displaying visual media such as images and / or videos. In some examples, the output device(s) 270 include an audio output device(s) such as a loudspeaker or headphones or a connector configured to connect the imaging system 200 to a loudspeaker or headphones. The audio output device(s) are configured and capable of playing audio media such as music, sound effects, audio tracks corresponding to a video, audio recordings recorded by a microphone(s) (e.g., of the sensor(s) 205), or a combination thereof. The output device(s) 270 may output media including a representation of the environment (e.g., media data 285 captured by the sensor(s) 205), virtual content (e.g., generated by the virtual content generator 207), a combination of the representation of the environment and the virtual content, modification(s) to the representation(s) of the environment and / or the virtual content and / or the combination (e.g., modified by the application engine 210 and / or the image reprojection engine 215), or a combination thereof. In some examples, the output device(s) 270 may face a user of the imaging system 200. For example, the display(s) of the output device(s) 270 may face a user of the imaging system 200 and / or display visual media to (e.g., towards) a user of the imaging system 200. Similarly, the audio output device(s) of output device(s) 270 may face a user of the imaging system 200 and may play audio media to (e.g., towards) the user of the imaging system 200.In some examples, output device(s) 270 includes output device 4135. In some examples, output device(s) 4135 can include output device 270. In Figure 2, output device(s) 270 is shown as a display for displaying visual media data and corresponding loudspeakers for playing audio media data.

[0067] The imaging system 200 also includes one or more transceivers 275 that the imaging system 200 can use to output modified media data 290 generated by the application engine 210 (e.g., including the image reprojection engine 215), for example, by transmitting the media to a receiving device. The receiving device can output the media using its own output device(s), for example, by displaying visual media data of the media using a display(s) of the output device(s) and / or by playing audio media data of the media using an audio output device(s) of the output device(s). The transceiver(s) 275 may include wired or wireless transceiver(s), communication interface(s), antenna(s), connection, coupling, coupling system, or combinations thereof. In some examples, the transceiver(s) 275 may include a communication interface 4140 of the computing system 4100. In some examples, the communication interface 4140 of the computing system 4100 may include a transceiver(s) 275. In FIG. 2, the transceiver(s) 275 are illustrated as wireless transceiver(s) 275 that transmit media data.

[0068] In some examples, the imaging system 200 includes a feedback engine 280. The feedback engine 280 can detect feedback received from a user through a user interface of the imaging system. The feedback engine 280 can detect feedback regarding one engine of the imaging system 200 received from other engines of the imaging system 200, such as whether one engine decides to use data from the other engine. The feedback can be feedback regarding any of the application engines 210, such as the image reprojection engine 215, the motion vector engine 220, the grid inversion engine 225, the time warp engine 230, the depth sensor support engine 235, the 3D stabilization engine 240, the 3D zoom engine 245, the reprojection SAT engine 250, the head pose correction engine 255, the XR late stage reprojection engine 260, the special effects engine 265, or a combination thereof. The feedback received by the feedback engine 280 can be positive feedback or negative feedback. For example, if one engine of the imaging system 200 uses data from another engine of the imaging system 200, the feedback engine 280 may interpret this as positive feedback. If one engine of the imaging system 200 rejects data from another engine of the imaging system 200, the feedback engine 280 may interpret this as negative feedback. Positive feedback may also be based on attributes of sensor data from the sensor(s) 205 and / or input from a user interface, such as a user smiling, laughing, nodding, pressing a button associated with positive feedback, making a gesture associated with positive feedback (e.g., thumbs up), making an affirmative statement (e.g., "yes," "confirmed," "okay," "next"), or otherwise responding positively to the media.Negative feedback can also be based on attributes of sensor data from the sensor(s) 205 and / or input from a user interface, such as a user grimacing, crying, shaking their head (e.g., in a "no" motion), pressing a button associated with negative feedback, making a gesture associated with negative feedback (e.g., thumbs down), making a negative statement (e.g., "no," "negative," "bad," "not this"), or otherwise reacting negatively to the virtual content.

[0069] In some examples, the feedback engine 280 provides feedback as training data to one or more ML systems of the imaging system 200 to update the one or more ML systems of the imaging system 200. For example, the feedback engine 280 can provide feedback as training data to any ML system(s) and / or trained ML model(s) of the application engine 210, such as the image reprojection engine 215, the motion vector engine 220, the grid inversion engine 225, the time warp engine 230, the depth sensor support engine 235, the 3D stabilization engine 240, the 3D zoom engine 245, the reprojection SAT engine 250, the head pose correction engine 255, the XR late stage reprojection engine 260, the special effects engine 265, or a combination thereof. Positive feedback can be used to strengthen and / or reinforce weights associated with the output of the ML system(s) and / or trained ML model(s). Negative feedback can be used to weaken and / or remove weights associated with the outputs of the ML system(s) and / or trained ML model(s).

[0070] In some examples, the feedback engine 280 includes a software element, such as a set of instructions corresponding to a program executing on a processor, such as the processor 4110, the image processor 150, the host processor 152, the ISP 154, and / or a combination thereof, of the computing system 4100. In some examples, the feedback engine 280 includes one or more hardware elements. For example, the feedback engine 280 may include a processor, such as the processor 4110, the image processor 150, the host processor 152, the ISP 154, and / or a combination thereof, of the computing system 4100. In some examples, the feedback engine 280 includes a combination of one or more software elements and one or more hardware elements.

[0071] FIG. 3A is a perspective view 300 showing a head mounted display (HMD) 310 used as an extended reality (XR) system 200. The HMD 310 may be, for example, an augmented reality (AR) headset, a virtual reality (VR) headset, a mixed reality (MR) headset, an extended reality (XR) headset, or some combination thereof. The HMD 310 may be an example of the imaging system 200. The HMD 310 includes a first camera 330A and a second camera 330B along the front of the HMD 310. The first camera 330A and the second camera 330B may be examples of the sensor(s) 205 of the imaging system 200. The HMD 310 includes a third camera 330C and a fourth camera 330D that face the user's eye(s) when the user's eye(s) face the display(s) 340. The third camera 330C and the fourth camera 330D may be examples of sensors 205 of the imaging system 200. In some examples, the HMD 310 may have only a single camera with a single image sensor. In some examples, the HMD 310 may include one or more additional cameras in addition to the first camera 330A, the second camera 330B, the third camera 330C, and the fourth camera 330D. In some examples, the HMD 310 may include one or more additional sensors in addition to the first camera 330A, the second camera 330B, the third camera 330C, and the fourth camera 330D, which may also include other types of sensors 205 and / or sensor(s) 205 of the imaging system 200. In some examples, the first camera 330A, the second camera 330B, the third camera 330C, and / or the fourth camera 330D may be examples of the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.

[0072] The HMD 310 may include one or more displays 340 visible to a user 320 wearing the HMD 310 on the user's head. The one or more displays 340 of the HMD 310 may be examples of one or more displays of the output device(s) 270 of the imaging system 200. In some examples, the HMD 310 may include one display 340 and two viewfinders. The two viewfinders may include a left viewfinder for the left eye of the user 320 and a right viewfinder for the right eye of the user 320. The left viewfinder may be oriented so that the left eye of the user 320 sees the left side of the display. The right viewfinder may be oriented so that the left eye of the user 320 sees the right side of the display. In some examples, the HMD 310 may include two displays 340, including a left display that displays content to the left eye of the user 320 and a right display that displays content to the right eye of the user 320. The display(s) 340 of the HMD 310 may be a digital "pass-through" display or an optical "see-through" display.

[0073] The HMD 310 may include one or more earpieces 335 that may function as speakers and / or headphones to output audio to one or more ears of a user of the HMD 310. Although one earpiece 335 is shown in FIGS. 3A and 3B, it should be understood that the HMD 310 may include two earpieces, one for each ear (left and right) of the user. In some examples, the HMD 310 may also include one or more microphones (not shown). The one or more microphones may be examples of the sensor(s) 205 of the imaging system 200. The one or more earpieces may be examples of the output device(s) 270 of the imaging system 200. In some examples, the audio output by the HMD 310 through the one or more earpieces 335 to the user may include or be based on audio recorded using the one or more microphones.

[0074] FIG. 3B is a perspective view 350 showing the head mounted display (HMD) of FIG. 3A being worn by a user 320. The user 320 wears the HMD 310 on the user's 320 head over the user's 320 eyes. The HMD 310 can capture images using a first camera 330A and a second camera 330B. In some examples, the HMD 310 displays one or more output images toward the user's 320 eyes using a display(s) 340. In some examples, the output images can include virtual content generated by the virtual content generator 207, composited using a compositor, and / or displayed by the display(s) of the output device(s) 270. The output images can be based on images captured by the first camera 330A and the second camera 330B, for example, with the virtual content overlaid. The output images can provide a stereoscopic view of the environment, possibly with the virtual content overlaid and / or with other modifications. For example, the HMD 310 can display a first display image based on an image captured by the first camera 330A to the right eye of the user 320. The HMD 310 can display a second display image based on an image captured by the second camera 330B to the left eye of the user 320. For example, the HMD 310 can provide virtual content overlaid within the display image overlaid on top of the images captured by the first camera 330A and the second camera 330B. The third camera 330C and the fourth camera 330D can capture images of the eye before, during, and / or after the user views the display image displayed by the display(s) 340. In this manner, sensor data from the third camera 330C and / or the fourth camera 330D can capture the reaction of the user's eye (and / or other parts of the user) to the virtual content. The earpiece 335 of the HMD 310 is shown within the ear of the user 320.The HMD 310 may output audio to the user 320 through the earpiece 335 and / or through another earpiece (not shown) of the HMD 310 in the other ear (not shown) of the user 320.

[0075] 4A is a perspective view 400 showing the front of a mobile handset 410 that includes a forward-facing camera and can be used as an extended reality (XR) system 200. The mobile handset 410 can be an example of the imaging system 200. The mobile handset 410 can be, for example, a mobile phone, a satellite phone, a portable game console, a music player, a health tracking device, a wearable device, a wireless communication device, a laptop, a mobile device, any other type of computing device or computing system described herein, or a combination thereof.

[0076] The front surface 420 of the mobile handset 410 includes a display 440. The front surface 420 of the mobile handset 410 includes a first camera 430A and a second camera 430B. The first camera 430A and the second camera 430B may be examples of sensors 205 of the imaging system 200. The first camera 430A and the second camera 430B may face a user, including the user's eye(s), while content (e.g., modified media output by the media modification engine 235) is displayed on the display 440. The display 440 may be an example of a display(s) of the output device(s) 270 of the imaging system 200.

[0077] The first camera 430A and the second camera 430B are shown within a bezel around the display 440 on the front 420 of the mobile handset 410. In some examples, the first camera 430A and the second camera 430B can be located in a notch or cutout cut out of the display 440 on the front 420 of the mobile handset 410. In some examples, the first camera 430A and the second camera 430B can be under-display cameras located between the display 440 and the remainder of the mobile handset 410, so that light passes through a portion of the display 440 before reaching the first camera 430A and the second camera 430B. The first camera 430A and the second camera 430B in the perspective view 400 are forward-facing cameras. The first camera 430A and the second camera 430B face in a direction perpendicular to the plane of the front 420 of the mobile handset 410. The first camera 430A and the second camera 430B may be two of one or more cameras of the mobile handset 410. The first camera 430A and the second camera 430B may be first and second image sensors, respectively. In some examples, the front face 420 of the mobile handset 410 may have only a single camera.

[0078] In some examples, the front surface 420 of the mobile handset 410 may include one or more additional cameras in addition to the first camera 430A and the second camera 430B. The one or more additional cameras may also be examples of the sensors 205 of the imaging system 200. In some examples, the front surface 420 of the mobile handset 410 may include one or more additional sensors in addition to the first camera 430A and the second camera 430B. The one or more additional sensors may also be examples of the sensors 205 of the imaging system 200. In some cases, the front surface 420 of the mobile handset 410 includes two or more displays 440. The one or more displays 440 of the front surface 420 of the mobile handset 410 may be examples of the display(s) of the output device(s) 270 of the imaging system 200. For example, the one or more displays 440 may include one or more touch screen displays.

[0079] The mobile handset 410 may include one or more speakers 435A and / or other audio output devices (e.g., earphones or headphones or connectors thereto) that can output audio to one or more ears of a user of the mobile handset 410. While one speaker 435A is shown in FIG. 4A, it should be understood that the mobile handset 410 can include more than one speaker and / or other audio device. In some examples, the mobile handset 410 can also include one or more microphones (not shown). The one or more microphones can be examples of the sensor 205 and / or the sensor(s) 205 of the imaging system 200. In some examples, the mobile handset 410 can include one or more microphones along and / or adjacent to the front surface 420 of the mobile handset 410, which are examples of the sensors 205 of the imaging system 200. In some examples, audio output by the mobile handset 410 to the user through one or more speakers 435A and / or other audio output devices may include or be based on audio recorded using one or more microphones.

[0080] 4B is a perspective view 450 showing a back 460 of a mobile handset that includes a rear-facing camera and can be used as an extended reality (XR) system 200. The mobile handset 410 includes a third camera 430C and a fourth camera 430D on the back 460 of the mobile handset 410. The third camera 430C and the fourth camera 430D of the perspective view 450 are rear-facing. The third camera 430C and the fourth camera 430D may be examples of the sensor(s) 205 of the imaging system 200 of FIG. 2. The third camera 430C and the fourth camera 430D are oriented perpendicular to the plane of the back 460 of the mobile handset 410.

[0081] The third camera 430C and the fourth camera 430D may be two of the one or more cameras of the mobile handset 410. In some examples, the back surface 460 of the mobile handset 410 may have only a single camera. In some examples, the back surface 460 of the mobile handset 410 may include one or more additional cameras in addition to the third camera 430C and the fourth camera 430D. The one or more additional cameras may also be examples of the sensor(s) 205 of the imaging system 200. In some examples, the back surface 460 of the mobile handset 410 may include one or more additional sensors in addition to the third camera 430C and the fourth camera 430D. The one or more additional sensors may also be examples of the sensor(s) 205 of the imaging system 200. In some examples, the first camera 430A, the second camera 430B, the third camera 430C, and / or the fourth camera 430D may be examples of the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.

[0082] The mobile handset 410 may include one or more speakers 435B and / or other audio output devices (e.g., earphones or headphones or connectors thereto) that can output audio to one or more ears of a user of the mobile handset 410. The one or more speakers 435B may be examples of the output device(s) 270 of the imaging system 200. While one speaker 435B is shown in FIG. 4B, it should be understood that the mobile handset 410 may include more than one speaker and / or other audio device. In some examples, the mobile handset 410 may also include one or more microphones (not shown). The one or more microphones may be examples of the sensor 205 and / or the sensor(s) 205 of the imaging system 200. In some examples, the mobile handset 410 may include one or more microphones along and / or adjacent to the back surface 460 of the mobile handset 410, which are examples of the sensor(s) 205 of the imaging system 200. In some examples, audio output by the mobile handset 410 to the user through one or more speakers 435B and / or other audio output devices may include or be based on audio recorded using one or more microphones.

[0083] The mobile handset 410 can use the display 440 on the front face 420 as a pass-through display. For example, the display 440 can display an output image. The output image can be based on an image captured by the third camera 430C and / or the fourth camera 430D, e.g., with virtual content overlaid and / or with modification by the media modification engine 235 applied. The first camera 430A and / or the second camera 430B can capture images of the user's eye (and / or other parts of the user) before, during, and / or after the display of the output image with the virtual content on the display 440. In this manner, the sensor data from the first camera 430A and / or the second camera 430B can capture a reaction of the user's eye (and / or other parts of the user) to the virtual content.

[0084] FIG. 5 is a conceptual diagram illustrating an example of grid inversion. The input to the grid inversion includes a first set of motion vectors, shown using solid black arrows from a first image Img1 510 to a second image Img2 515 in FIG. 5 as a motion vector (MV) grid. The motion vector grid indicates, for each pixel (or group of pixels), how much that pixel (or group of pixels) is going to move between a first image Img1 510 (e.g., vision or depth) in an environment and a second image Img2 515 (e.g., vision or depth) in that environment using a motion vector in the motion vector (MV) grid 505. The motion vector grid 505 may be referred to as a motion vector map of the image. The motion vectors of the motion vector grid 505 may be determined using a motion vector engine 220, for example using optical flow.

[0085] 5. The grid inversion engine 225 may perform a grid inversion that changes a characteristic(s) (e.g., direction, origin, position, length, and / or size) of the motion vectors in the first group of motion vectors (motion vector grid 505) to generate a second set of motion vectors (inverse MV grid 520). Instead of showing how each pixel from Img1 510 moves to Img2 515 (as in MV grid 505), the motion vectors of the second set of motion vectors (inverse MV grid 520) show how each pixel from Img2 515 can move back to Img1 510. The motion vectors of the second set of motion vectors (inverse MV grid 520) are shown using black dashed arrows pointing from the second image Img2 515 to the first image Img1 510 in FIG. 5.

[0086] The various black icons in FIG. 5 represent various elements in the environment depicted in the two images Img1 510 and Img2 515. For example, the elements include a house, a bird, a person, a car, and a tree. According to the MV grid 505, the house and the tree have not moved from Img1 510 to Img2 515, as represented by a 0 in the MV grid 505. Similarly, in the inverse MV grid 520, the house and the tree have not moved from Img2 515 to Img1 510. The house is represented by a 0 in the MV grid 505 and the inverse MV grid 520, both in cell 0 where the house is located. The tree can be represented by a 0 in the inverse MV grid in cell 8 where the tree is located, but there is a conflict with the car, discussed below, and is represented by a black circle. The bird has moved one grid cell to the right from Img1 510 to Img2 515 (from cell 1 to cell 2), which is represented by a 1 in cell 1 in the MV grid 505. The bird moves one grid cell to the left from Img2 515 to Img1 510 (cell 2 to cell 1), which is represented by −1 in cell 2 in the inverse MV grid 520. Not only are values ​​inverted (multiplied by −1) from MV grid 505 to inverse MV grid 520, but they are also moved from the cell corresponding to the element's old position in Img1 510 to the cell corresponding to the element's new position in Img2 515. The black star in cell 1, where the bird was in Img1 510 but not in Img2 515, indicates that the area of ​​the image corresponding to cell 1 is missing and may need to be filled in (e.g., using interpolation and / or inpainting) in the inverse MV grid 520. The person moves two grid cells to the left from Img1 510 to Img2 515 (cell 6 to cell 4), which is represented by −2 in cell 6 in the MV grid 505. The person moves two grid cells to the right, from Img2 515 to Img1 510 (cell 4 to cell 6), which is represented by the 2 in cell 4 in the inverse MV grid 520. The black star in cell 6, where the person was in Img1 510 but no longer in Img2 515, indicates that the area of ​​the image corresponding to cell 1 in the inverse MV grid 520 is missing and may need to be filled in (e.g., using interpolation and / or inpainting).The car moves one grid cell to the right, from Img1 510 to Img2 515 (cell 7 to cell 8), which is represented by a 1 in MV grid 505. The car moves one grid cell to the left, from Img2 515 to Img1 510 (cell 8 to cell 7), which can be represented by a -1 in the inverse MV grid 520. However, the car and tree are in the same grid cell (cell 8) in Img2 515, so the red circle indicates a conflicting value in that cell of the inverse MV grid 520 (e.g., 0 for the tree and -1 for the car).

[0087] FIG. 6 is a conceptual diagram 600 illustrating an example of depth-based reprojection. Depth-based reprojection is performed by the image reprojection engine 215. This example shows a camera image 610 of an environment (referred to as world scene 605) that has a desk with a toolbox on it and several chairs around it. The image reprojection engine 215 uses depth data 620 of the environment (e.g., of the world scene 605) to reproject the camera image 610 to generate a reprojected image 615. The reprojected image 615 depicts the same environment as the camera image 610 (e.g., world scene 605), but reprojected as if the environment was captured from a different viewpoint or point in the reprojected image 615 compared to the camera image 610. In the example shown in FIG. 6, the reprojected image 615 appears to have been captured from a viewpoint or point in the environment that is translated to the left of the viewpoint or point in the environment depicted in the camera image 610. In some examples, the image reprojection engine 215 may perform image reprojection using an inverse MV grid (e.g., the inverse MV grid 520) generated by the grid inversion engine 225, for example based on the depth data 620.

[0088] 7 is a conceptual diagram 700 illustrating an example of a time warp 705 performed by the time warp engine 230. On the left, a large or dense motion vector map 720 is shown as a solid black arrow, illustrating how pixels move between image frame n and image frame n-4. Image frame n and image frame n-4 are shown as tall vertical lines. The time warp 705 uses grid inversion (using the grid inversion engine 225) on the large or dense motion vector map 720 to create smaller motion vector maps, shown as shorter vertical arrows, for example, from image frame n to image frame n-1, from image frame n-1 to image frame n-2, from image frame n-2 to image frame n-3, and from image frame n-3 to image frame n-4.

[0089] To create the smaller vector map, the time warp engine 230 uses resampling. For example, to generate the smaller vector map, the time warp engine 230 reduces the values ​​in the motion vector map (representing the distance of movement of the element between frame n and frame n-4), for example by multiplying the value by 1 / 4. In addition, the time warp engine 230 moves the values ​​to the new position of each element in the corresponding frame, similar to the movement of values ​​in the grid inversion of FIG.

[0090] Time warp 705 can be used to interpolate motion vector maps between existing motion vector maps, for example when optical flow is performed only every k frames. Optical flow is a computationally expensive operation that can use a lot of power to execute, but time warp 705 as demonstrated here is a lower cost, lower power operation. Thus, optical flow can be used sparingly to reduce computational cost and power usage, but time warp 705 can still enable imaging system 200 to obtain motion vectors for every frame transition between any two adjacent frames (and possibly between any two frames).

[0091] In some examples, the smaller motion vector map generated by Time Warp 705 can be used to interpolate additional frames between existing frames of the video, for example to increase the frame rate of the video from a first frame rate to a second frame rate that is higher than the first frame rate.

[0092] In some examples, the smaller motion vector map generated by the time warp 705 can be used to improve the quality of a particular frame of the video. For example, if a particular frame of the video is blurry, contains significant compression artifacts, contains compression artifacts that obscure the imaged scene from clarity, or otherwise suffers from poor quality, the time warp 705 can improve the quality of such frame of the video. The time warp 705 can be used to determine a motion vector map from one or more adjacent or nearby frames of the video, and the image data from these frames can be used to generate a modified image to replace the particular frame of interest so as to improve the image quality of the particular frame of interest. The conceptual diagram 700 shows two examples of images of a boy, a first image 710 on the left where the time warp 705 has not been applied, and a second image 715 on the right where the time warp 705 has been applied, improving the clarity of the depiction of the boy in the second image 715 compared to the first image 710. The image on the right 715, which has been enhanced using time warp 705, appears sharper and more distinct than image 710, particularly at and near the various edges of the depiction of the boy, as shown using solid lines to represent the various lines and edges of the depiction of the boy in image 715. Additionally, in some instances, patterns such as hair patterns, fabric patterns, other patterns, text, logos, and / or other designs may appear sharper and more distinct in an image with time warp 705 applied (e.g., image on the right 715) than in an image without time warp 705 applied (e.g., image on the left 710).

[0093] Further examples of time warp 705 and image enhancement using time warp 705 are shown in FIG. 23 and FIG.

[0094] FIG. 8 is a conceptual diagram 800 illustrating an example of depth sensor support 805 performed by depth sensor support engine 235. A cluster of sensors 205 on imaging system 200 is shown, including a set of image sensors 810 and a set of depth sensors 815, which may include time-of-flight (ToF) sensors. In some cases, in image processing, image data from image sensor 810 and depth data from depth sensor 815 may be useful to use together, for example, to generate bokeh, simulated depth-of-field blur, object recognition, etc. However, image sensor 810 and depth sensor 815 are not collocated. Instead, image sensor 810 and depth sensor 815 are offset from each other by offset 820. Thus, use of image data from image sensor 810 and depth data from depth sensor 815 may result in parallax issues due to slight mismatch of viewpoints caused by offset 820. Thus, the depth in the depth data may not match the object depicted in the image data. This discrepancy can be particularly noticeable for objects in the environment close to the sensor, which can appear in significantly different positions in the image data versus the depth data. More distant objects may appear more similar in the image data and depth data.

[0095] To correct for this mismatch, in some examples, the image reprojection engine 215 can reproject the depth data from the depth sensor 815 to appear to come from the perspective of the image sensor 810. In some examples, the image reprojection engine 215 can reproject the image data from the image sensor 810 to appear to come from the perspective of the depth sensor 815. Because depth data may be required for the image reprojection engine 215 to perform the reprojection, the image reprojection engine 215 can rely on an external calibration between the image sensor 810 and the depth sensor 815 for appropriate depth data.

[0096] FIG. 9 is a conceptual diagram 900 illustrating an example of 3D stabilization 905 performed by the 3D stabilization engine 240. Conventional stabilization techniques can compensate for rotational movements, but generally cannot compensate for translational (e.g., parallax) movements in the real world. Image reprojection using the image reprojection engine 215 based on depth data of the environment can provide true 3D stabilization 905 that corrects for parallax movements, including translational movements, rotational movements, or both. For each video frame of the video captured using the sensor(s) 205, including the four video frames labeled Original ("Original") in FIG. 9, reprojection is performed using the image reprojection engine 215 to generate a stabilized variant ("Stable") of the original video frame. The resulting reprojected video frames are reprojected such that their respective viewpoints all fall on a line representing a virtual stabilized movement path, without any parallax movements perpendicular to the line, or any rotations about the axis corresponding to the line (or any other axis). The lines may be curved to represent curved paths of movement, but do not have any jagged edges corresponding to such parallax movements or rotations.

[0097] For filmed 3D stabilization 905, the input video shown by the video frames is shaking in different directions, i.e. translationally up, translationally down, translationally left, translationally right, translationally forward, translationally navigational, and / or rotationally (e.g., pitch, yaw, and roll). All of these movements in shaking are stabilized by reprojection using the image reprojection engine 215, as the image reprojection engine 215 reprojects the image to change the viewpoint on the environment.

[0098] In some cases, blank areas may appear in the stabilized frames, for example at the edges of the frame and / or around people in the frame (e.g., to the right of the woman in the fourth stabilized frame at the bottom right of FIG. 9). These may represent occlusion areas where there is no corresponding data in the original image. These occlusion areas may be filled by the image reprojection engine 215, for example, using interpolation and / or inpainting (e.g., deep learning based inpainting). Additional examples 3205 of 3D stabilization 905 are shown in FIG. 30. In some examples, these blank areas may appear black. In some examples, these blank areas may appear white. In FIG. 9, these blank areas are shown in white.

[0099] In some examples, it may be useful for 3D stabilization, as well as certain other applications of the image reprojection engine 215, to treat distant pixels as if they were infinitely distant and to make the positions of such pixels unchanged under reprojection. In some examples, the image reprojection engine 215 can use translation falloff to smoothly transition the translation value toward a value that represents infinity to treat distant pixels as if they were infinitely distant.

[0100] FIG. 10 is a conceptual diagram 1000 illustrating an example of a 3D zoom 1005 (also referred to as a cinematic zoom) performed by the 3D zoom engine 245. The 3D zoom 1005 performed by the 3D zoom engine 245 may include zooming in on an image (e.g., magnifying a particular portion of an image while removing other portions of the image), moving a virtual camera in different directions (e.g., panning, rotating, etc.), and / or other types of zoom. In some cases, to perform a digital zoom on an image, the entire image is conventionally upscaled and cropped, as shown in the sequence of four images labeled as digital zoom ("dig.zm.") in FIG. 10. The image shows a skateboarder in front of a house. By performing a digital zoom (or an optical zoom, in some examples, using an optical zoom lens or a switch between the camera and / or lens), a significant portion of the view of the house is lost. However, when the camera is moved closer to the skateboarder, not as much view of the house is lost as would be lost using a digital zoom. This is because the skateboarder is closer to the camera than the house. In other words, the skateboarder is in the foreground and the house is in the background.

[0101] 3D zoom 1005, or depth-based zoom or cinematic zoom, uses image reprojection using the image reprojection engine 215 based on the depth data 1020 of the environment to simulate forward movement of the camera in the environment, in this case, toward the skateboarder. As shown in the sequence of four images labeled as depth-based zoom ("depth.zm.") in FIG. 10, the skateboarder increases in size to the same extent as with the digital zoom, but the depth of field of the house is not lost significantly. For example, at the end of the four images in the sequence, the extent of four windows of the house are at least partially in frame under the digital zoom, but the extent of six windows of the house are at least partially in frame under the 3D depth-based zoom (although one of these windows is completely behind the skateboarder). Thus, 3D depth-based zoom (or cinematic zoom) minimizes the loss of view, especially of background elements. 3D zoom 1005 (or depth-based zoom or cinematic zoom) is illustrated in FIG. 31.

[0102] FIG. 11 is a conceptual diagram 1100 illustrating an example of a reprojection 1105 performed by the reprojection SAT engine 250. A cluster of sensors 205 of the imaging system 200 is shown in FIG. 11 with a telephoto sensor 1110, a wide-angle sensor 1115, and other sensors 1125. In some cases, the imaging system 200 may switch between the telephoto sensor 1110 and the wide-angle sensor 1115, for example, to provide different levels of zoom to the image of the environment. However, similar to the scenario with the image sensor 810 and the depth sensor 815 of FIG. 8, the telephoto sensor 1110 and the wide-angle sensor 1115 are not collocated. Instead, there is an offset 1120 between the telephoto sensor 1110 and the wide-angle sensor 1115. Thus, switching between the telephoto sensor 1110 and the wide-angle sensor 1115 creates a parallax effect. For example, a telephoto image 1130 is taken (labeled "tele") captured using the telephoto sensor 1110, and a wide-angle image 1135 is taken (labeled "wide") captured using the wide-angle sensor 1115 and cropped to match the telephoto field of view, i.e., digitally zoomed before transitioning to the telephoto sensor. Both images depict a man in front of a distant background. In the telephoto image 1130, the man is slightly to the right of his position in the wide-angle image 1135.

[0103] Similar to the depth sensor support 805 of FIG. 8, the reprojection SAT engine 250 can perform a reprojection 1105 to correct the offset 1120 based on the depth data 1160. For example, the reprojection SAT engine 250 can perform a reprojection 1105 to correct the telephoto image to correct the viewpoint so that the modified telephoto image 1140 (labeled “modif.tele”) appears to have been captured from the viewpoint of the wide angle sensor 1115 (e.g., as in the wide angle image 1135) rather than the viewpoint of the telephoto sensor 1110 (e.g., as in the telephoto image 1130). In the modified telephoto image 1140, the man appears slightly to the left of his position in the unmodified telephoto image 1130. In the modified telephoto image 1140, the man appears in a position similar to the position of the man in the wide angle image 1135. A dark shadow caused by parallax shift of the image data depicting the man against the background appears to the right of the man in the modified telephoto image 1140. The black shadows represent "holes" that can be filled with image data, for example using interpolation and / or inpainting, as further described.

[0104] In some examples, the reprojection SAT engine 250 can instead perform reprojection 1105 based on depth data 1160 to correct the wide-angle image to correct the viewpoint so that a corrected wide-angle image (not shown) appears to have been captured from the viewpoint of the telephoto sensor 1110 rather than the viewpoint of the wide-angle sensor 1115. Unlike sensor-to-sensor transformation, where a set of digitally zoomed images from one sensor are warped based on an image estimation that matches the second sensor before the switch, the reprojection SAT engine 250 can correct the offset based on depth data to reduce parallax issues (e.g., parallax error), especially for closer objects (e.g., objects in the foreground and / or objects below a threshold depth). Additional examples of reprojection 1105 are shown in FIG. 32.

[0105] FIG. 12 is a conceptual diagram 1200 illustrating an example of head pose correction 1205 performed by the head pose correction engine 255. In some cases, a user's image may be captured from a suboptimal angle and / or a sub-flattering angle (e.g., an angle other than a vertical angle perpendicular to the user's face). For example, when a user captures a selfie of themselves or points a camera at themselves for a video conference, the angle at which the image is captured often does not align with the user's head pose, such that the user appears to be looking down, up, left, and / or right. In some cases, the user's hands may become tired and / or uncomfortable from holding the phone or other imaging system 200 for extended periods of time, which may exacerbate this problem when the user's hands drop or rest against a perceptual surface.

[0106] Head pose correction 1205 performed by head pose correction engine 255 can perform reprojection using image reprojection engine 215 to reproject the actual sensor position to match a more optimal and / or better looking viewpoint, such as a viewpoint from a vertical angle perpendicular to the user's face.

[0107] For example, the original head pose of the woman in input image 1210 was captured from a less flattering angle slightly below the height of the woman's head, emphasizing the woman's neck and chin area. Head pose correction 1205 uses image reprojection engine 215 based on input image 1210 and depth data 1220 to generate a reprojected image 1215 from a viewpoint from a vertical angle perpendicular to the user's face. The reprojected image 1215 appears to view the woman's face from a much more flattering vertical angle, emphasizing the woman's facial features rather than her neck and chin as in input image 1210. Additional examples of head pose correction 1205 are shown in FIG.

[0108] FIG. 13 is a conceptual diagram 1300 illustrating an example of XR late stage reprojection 1305 performed by the XR late stage reprojection engine 260. Some XR devices (e.g., HMD 1320) or other mobile devices capture sensor data (e.g., images, videos, depth images, and / or point clouds) using their sensors 205 at low frame rates to conserve battery power. Interpolation can be used to generate additional frames between frames of low frame rate sensor data to improve frame rates. A high frame rate can be important for XR applications because low frame rate XR can cause nausea to the user and / or make the XR appear jittery and unrealistic.

[0109] Interpolation techniques may not always be able to realistically represent all changes in the viewpoint of the XR device (e.g., HMD 1320). For example, the interpolation may use digital zoom to simulate a user moving closer or further away from an object, which may cause field of view mismatches similar to those described with respect to the 3D zoom 1005 of FIG. 10. Interpolation techniques may also have difficulty with parallax movements, for example caused by translational movements of the XR device (e.g., HMD 1320). Interpolation techniques may also have difficulty with rotational movements, for example caused by changes in the orientation (e.g., pitch, roll, and / or yaw) of the XR device (e.g., HMD 1320).

[0110] The XR late stage reprojection 1305 performed by the XR late stage reprojection engine 260 can perform image reprojection using the image reprojection engine 215 to reproject an image of the environment based on a change in the position of the XR device. The change in the position of the XR device (e.g., HMD 1320) can be determined based on sensor data from an orientation sensor of the XR device (e.g., HMD 1320), which may use less bandwidth and / or power than an image sensor or depth sensor. The change in the position of the XR device (e.g., HMD 1320) can be inferred based on image data, depth data, and / or audio data from an image sensor, depth sensor, and / or microphone of the sensor 205 of the XR device (e.g., HMD 1320).

[0111] For example, an input image 1310 is shown, based on which the XR late stage reprojection engine 260 generates a reprojected image 1315 using XR late stage reprojection 1305 based on the illustrated change in orientation of an HMD 1320, which is an example of an XR device.

[0112] FIG. 14 is a conceptual diagram 1400 illustrating an example of a special effect 1405 performed by the special effects engine 265. The special effects 1405 performed by the special effects engine 265 may perform image reprojection using the image reprojection engine 215 to reproject an input image 1410 to rotate around an object, pan along an object, rotate a viewpoint around an axis, move a viewpoint along a path, or some combination of these. In the example shown in FIG. 14, an input image 1410 of an environment is reprojected from a different viewpoint of the environment to form a reprojected image 1415. The viewpoint on the environment in the reprojected image 1415 is to the left of the viewpoint on the environment in the input image 1410, making, for example, a toolbox appear to rotate and / or tilt to the right in the reprojected image 1415 relative to the input image 1410.

[0113] FIG. 15 is a conceptual diagram 1500 illustrating an image reprojection transformation based on a matrix operation. The conceptual diagram 1500 illustrates how the image reprojection engine 215 can reproject a captured image 1510 of an environment to generate a reprojected image 1515 of the environment from a different perspective than the captured image 1510. The image reprojection engine 215 receives a captured image 1510 from the sensor(s) 205, specifically a camera. The captured image depicts the environment from a first perspective ("first persp."). An example of the captured image 1510 is shown in FIG. 15. For example, using a pinhole camera paradigm with focal length f and depth, the imaging system can determine where an object is in the environment relative to the camera. The image reprojection engine 215 can use an intrinsic matrix describing a first camera (also known as an original camera, source camera, or first viewpoint), a second intrinsic matrix describing a second camera or virtual camera (also known as a target camera, or second viewpoint) in the 3D world, and a 3D transformation matrix to move or reproject from a first camera to a second camera. In some examples, the image reprojection engine can perform depth reprojection to create a second depth map that describes the environment from the second viewpoint based on the same principles as the image reprojection described herein. Additionally, various transformation paradigms can be used for image and / or depth reprojection, such as transformation paradigms that take into account lens distortion (e.g., radial distortion).

[0114] The image reprojection engine 215 receives a depth map ("depth over image area") (e.g., depth data 620), for example, from a depth sensor and / or based on a depth determination using a camera (e.g., stereo depth sensing, ToF sensor, and / or structured light). Based on the depth map, the image reprojection engine 215 can determine the exact location in 3D coordinates (e.g., X, Y, and Z) of any given object in the captured image 1510, such as a chair, or a table, or a toolbox depicted in the captured image 1510. For example, the depth of the object, the camera (intrinsic cam ) and the coordinates of the object in the captured image 1510

number

number

[0115] Camera Intrinsic Matrix cam ) can be used to convert 3D camera coordinates to 2D image coordinates, and the focal length measurement(s) (f x and / or f y ) and / or principal point offset(s) (c x and / or c y ) can be based on

number

[0116] The 3D transformation can be based on an intrinsic matrix at the source camera position and the target camera position that corresponds to the reprojection, for example as shown below:

number

[0117] The image reprojection engine 215 receives and / or determines a reprojection matrix that indicates how the viewpoint should move within the reprojection environment (e.g., a simulated movement of the camera). The values ​​in the reprojection matrix depicted in FIG. 15 are labeled R11, R12, R13, Tx, R21, R22, R23, Ty, R31, R32, R33, and Tz. In another example, the image reprojection engine can obtain the transformation directly as a 3D transformation matrix (e.g., without performing at least some of the calculations set forth above). Once the image reprojection engine 215 knows how the viewpoint should move within the environment in the form of a reprojection matrix, the image reprojection engine 215 can calculate the transformation as follows: out , Y out , and Z out By determining (e.g., in the reprojected image 1515), the new 3D position of the object in the environment after the camera movement can be determined.

number

[0118] The image reprojection engine 215 includes:

number

number

number

[0119] The image reprojection engine 215 computes the coordinates of an object in the captured image 1510 to determine a motion vector of the object from the captured image 1510 to the reprojected image 1515.

number

number

number

[0120] The image reprojection engine 215 calculates a motion vector MV for any pixel of any object in the captured image 1510 to know where that pixel should fall in the reprojected image 1515. x and M.V. y may be used. In an illustrative example, a portion of the chair may move 4 pixels to the right from the captured image 1510 to the reprojected image 1515. On the other hand, because the toolbox is closer to the camera than the chair, a portion of the toolbox may move 10 pixels to the right from the captured image 1510 to the reprojected image 1515. Thus, for each object, the image reprojection engine 215 may calculate where the object should move in the reprojected image 1515 compared to the captured image 1510.

[0121] The motion vectors can represent pixel displacements of each pixel in the first image data to pixel locations in the second image data, where the displacements depend on the relative viewing points of the first and second viewpoints and the inverse of the depth. As described above, the motion vectors can be determined based on depth data (e.g., "Depth" in the above equation). For example, in some examples, the motion vectors can be determined based on the position(s) of the object(s) in the environment, such as 3D coordinates (e.g., X, Y, Z), which can be determined from the captured image data based on the depth data. In some examples, the motion vectors can be determined based on the output(s) (e.g., X, Y, Z) of a transformation (e.g., a 3D transformation) of the 3D coordinates (e.g., X, Y, Z) of the object(s). out , Y out , Z out ), may be determined based on the output(s) of the transformation of the position(s) of the object(s) in the environment.

[0122] In some examples, the focal length f of the camera may also be incorporated into some of the above equations. For example, determining the X and Y coordinates of an object in the environment may be calculated using the focal length f and determining the coordinates of the object in the reprojected image 1515, e.g., as shown below:

number

number

[0123] FIG. 16 is a block diagram 1600 illustrating a grid inversion transform based on depth data and a 3D transform. The grid inversion transform takes a 3D transform 1605 (e.g., in the form of a reprojection matrix) and a depth map 1610, and generates a motion vector (MV) 1620 that indicates the movement of an object in the environment from the captured image 1510 to the reprojected image 1515 using MV calculations 1615, as shown in FIG. 15. In some examples, the initial motion vector may also be referred to as an existing motion vector. The grid inversion transform performs a grid inversion 1625 on the existing MV 1620 to an inverse motion vector 1630. In some examples, the inverse motion vector may also be referred to as a required motion vector.

[0124] FIG. 17 is a block diagram 1700 illustrating an image reprojection transformation based on a motion vector. A warp engine 1705 is shown, which may be part of the image reprojection engine 215. The warp engine 1705 uses an inverse motion vector 1730 (e.g., inverse MV in FIG. 15-16) instead of the originally determined motion vector (MV in FIG. 15-16). This is because the inverse motion vector 1730 is an out-to-in motion vector, whereas the originally determined motion vector (MV) is an in-to-out motion vector. An out-to-in motion vector transformation is less computationally expensive than an in-to-out motion vector transformation. In particular, if the warp engine 1705 uses an out-to-in motion vector such as the inverse motion vector 1730 to generate the reprojection image 1715, the warp engine 1705 can generate the reprojection image 1715 pixel-by-pixel in the raster order (or inverse raster order, or any other order) of the reprojection image. For each pixel in the reprojected image 1715, the reverse out-to-in motion vector 1730 instructs the warp engine 1705 to pull pixel data from a particular location in the captured image 1710 and fill that pixel in the reprojected image 1715 with that pixel data from the captured image 1710. For example, for a particular pixel in the reprojected image 1715, the warp engine 1705 can read the reverse out-to-in motion vector 1730 to determine that the value for that pixel should be taken from the pixel four pixels to the left in the captured image 1710, and so on.

[0125] An in-to-out motion vector may refer to a motion vector indicating the movement of pixels from an initial image of a scene (from an initial viewpoint) to a target image of the scene (from a target viewpoint). The initially determined motion vector (e.g., MV in Figs. 15-16) may be an example of an in-to-out motion vector. An out-to-in motion vector may refer to a motion vector indicating the movement of pixels from a target image of a scene (from a target viewpoint) to an initial image of the scene (from an initial viewpoint). An inverse MV 1730 may be an example of an out-to-in motion vector.

[0126] When the warp engine 1705 performs a warp (e.g., from a captured image 1710 to a reprojected image 1715), the use of out-to-in motion vectors for warping (e.g., inverse motion vectors 1730) can provide a reduction in consumption of computational resources over the use of in-to-out motion vectors for warping (e.g., MVs in FIGS. 15-16). The in-to-out motion vectors (e.g., MVs in FIGS. 15-16) are organized based on the captured image 1710 rather than based on the reprojected image 1715. Meanwhile, the out-to-in motion vectors (e.g., inverse motion vectors 1730) are instead organized based on the reprojected image 1715. When the warp engine 1705 performs warping to generate the reprojected image 1715, it is optimal to generate the reprojected image 1715 according to a pixel order based on the reprojected image 1715 (e.g., in a raster order according to the reprojected image 1715) rather than generating the reprojected image 1715 according to a pixel order based on the captured image 1710 (e.g., in a raster order according to the captured image 1710). The use of out-to-in motion vectors (e.g., inverse motion vectors 1730) for warping can enable the warp engine 1705 to generate the reprojected image 1715 according to a pixel order based on the reprojected image 1715 (e.g., in a raster order according to the reprojected image 1715). For example, using the inverse motion vectors 1730, the warp engine 1705 can generate each pixel of the reprojected image 1715 with any conflicts or missing regions already resolved, as described with respect to FIG. On the other hand, in order for the warp engine 1705 to generate the reprojected image 1715 in raster order for pixels in the reprojected image 1715 using in-to-out motion vectors, the warp engine 1705 iteratively searches through the motion vectors through a pixel-by-pixel search of the captured image 1710 and the in-to-out motion vectors for each particular pixel of the reprojected image 1715 to find data that should end up within that particular pixel of the reprojected image 1715. The iterative search through the captured image 1710 and the in-to-out motion vectors is computationally expensive and uses significant power.In some cases, the warp engine 1705 may further need to resolve conflicts or fill missing regions if these searches brought up motion vectors in the wrong order, e.g., erroneously prioritizing distant objects over nearby objects rather than prioritizing nearby objects over distant objects. Thus, even though it may cost some computation to generate out-to-in motion vectors (e.g., reverse motion vectors 1730) from in-to-out motion vectors (e.g., the motion vectors of Figures 15-16), the end result of using out-to-in motion vectors (e.g., reverse motion vectors 1730) for warping is still a savings in computational resources and improved accuracy.

[0127] In some examples, since determining the in-to-out MVs can be costly, the in-to-out MVs (existing MVs) are determined at a lower resolution, e.g., 1 / 4 of the resolution of the captured image. Generating the out-to-in MVs (required MVs) by applying grid inversion to the in-to-out MVs is not computationally expensive. Furthermore, reprojection using the out-to-in MVs (required MVs) is not computationally expensive. The computationally inexpensive nature of these operations allows the grid inversion and / or reprojection using the out-to-in MVs (required MVs) to be efficiently performed even at a higher resolution, such as the full resolution of the captured image. Thus, the warp engine 1705 can generate the reprojected image to be a full reprojection of the captured image, despite determining the in-to-out MVs (existing MVs) at a lower resolution. This allows for further savings in computational resources and power.

[0128] The grid inversion engine 225 includes several mechanisms for handling missing data and / or conflicts in the inverted MV grid. As explained above, the grid inversion engine changes the positions of MVs to correlate the positions of pixels in the target image (e.g., reprojected image 1715). In some cases, there are pixels that are not pointed to by MVs in the input grid, and therefore no MVs will be placed at these positions using inversion alone. The grid inversion engine fills these cells in the inverted MV grid during its process by interpolation. Referring again to FIG. 5, the inverted MV grid 520 includes missing cells generated via grid inversion and marked using stars. For example, cell 1 in the inverted MV grid 520 does not have a corresponding motion vector from the MV grid 505 and is filled instead using inpainting. One option for interpolation is to interpolate the value of cell 1 using the values ​​of its neighboring cells 0 and 2. For example, the weights for the interpolation can be by distance, so the interpolated value for cell 1 can be -1 / 2 based on a value of 0 in cell 0 and a value of -1 in cell 2. A similar type of interpolation can be performed for cells 3, 5, 6, and 7.

[0129] The grid inversion engine 225 may also include a mechanism for handling conflicts in the inverted MV grid. In some cases, multiple MVs in the MV grid 505 may point to the same pixel in the second image (e.g., second image Img2 515, reprojected image 1715), thus creating conflicts of MVs in the inverted MV grid 520 and requiring the grid inversion engine to choose one of the conflicting values ​​for a given cell in the inverted MV grid 525. An example of such a conflict is shown in cell 8 of the inverted MV grid 520. Both the car in cell 7 of the first image Img1 510 and the tree in cell 8 of the first image Img1 510 end up in the same pixel corresponding to cell 8 in the second image Img2 515 for each motion vector extending from cells 7 and 8 in the MV grid 505. As a result, it may be unclear which value the grid inversion engine should choose to place in cell 8 of the inverted MV grid 520.

[0130] To resolve the conflict, the grid inversion engine 225 may select either value. In some examples, a weighted average of the conflicting values ​​may be used. If the grid inversion engine 225 has depth information corresponding to two objects (e.g., from the depth data 620), the grid inversion engine 225 may select the value corresponding to the object closer to the sensor 205. This is because the closer object often obscures, obstructs, or occludes the view of the more distant object. If the grid inversion engine 225 lacks depth information corresponding to two objects, the grid inversion engine 225 may select a value based on other heuristics or techniques, such as selecting a value corresponding to greater motion, or an object that appears larger. An object experiencing greater motion is more likely to be closer to the sensor 205, regardless of the size of the object, because a closer object appears to cover a greater amount of the field of view of the sensor 205 than a more distant object, even if the motion is the same speed. In some examples, an object that appears larger may also be closer to the sensor 205.

[0131] In some examples, referring to FIG. 5, a car moving from cell 7 in the first image Img1 510 to cell 8 in the second image Img2 515 is closer to the sensor 205 than the tree, in which case the grid inversion engine 225 may select the value of cell 8 in the inverse MV grid 520 to be −1 (as the inverse of the corresponding value of 1 in cell 7 in the MV grid 505). In some examples, in FIG. 5, a tree is closer to the sensor 205 than the car, in which case the grid inversion engine 225 may select the value of cell 8 in the inverse MV grid 520 to be 0 (based on the corresponding value of 0 in cell 8 in the MV grid 505). In some examples, the grid inversion engine 225 may lack information about the relative depth of the car compared to the tree. In such a case, the value of cell 8 in the inverse MV grid 520 is selected to be −1 because the car is experiencing greater motion (its value is 1 in the MV grid 505 compared to the tree's value of 0), and thus the car is likely closer to the sensor 205 than the tree. In some examples, if a car appears larger than a tree in the image(s), then the car is more likely to be closer to the sensor 205 than the tree, and so the value of cell 8 of the inverse MV grid 520 is selected to be −1. In some examples, the value of cell 8 of the inverse MV grid 520 is selected to be −½, as the average of the reciprocals of the values ​​of cells 7 and 8 of the MV grid 505.

[0132] Different types of interpolation may be performed, and in one example, the interpolation may weight values ​​based on the distance to neighboring cells. In another example, the interpolation may weight values ​​based on the depth of neighbors. Other methods may also be applied. For example, in larger gaps, such as cells 5, 6, and 7 of inverse MV grid 520, the interpolation may weight information from nearby cells higher than information from distant cells. For example, the value of cell 6 of inverse MV grid 520 may be the average between the value of cell 4 of inverse MV grid 520 (2) and the value of cell 8 of inverse MV grid 520. The value of cell 8 of inverse MV grid 520 may depend on how the conflict in cell 8 is resolved, as described above. Assuming that the value of cell 8 of inverse MV grid 520 is -1, then the value of cell 6 of inverse MV grid 520 may be "1 / 2. The value of cell 5 of the inverse MV grid may weight 520 the value of cell 4 (2) of inverse MV grid 520 more highly in its interpolation than the value of cell 8 of inverse MV grid 520, e.g., is the average of the value of cell 4 of inverse MV grid 520 and the interpolated value of cell 6 of inverse MV grid 520. Similarly, the value of cell 7 of inverse MV grid 520 may weight the value of cell 4 (2) of inverse MV grid 520 more highly in its interpolation than the value of cell 8 of inverse MV grid 520, e.g., is the average of the value of cell 8 of inverse MV grid 520 and the interpolated value of cell 6 of inverse MV grid 520. For example, assuming the value of cell 8 of the inverse MV grid is −1, the value of cell 5 of the inverse MV grid may be set to 1.25, while the value of cell 7 of the inverse MV grid may be set to −0.25.

[0133] FIG. 18 is a conceptual diagram 1800 illustrating an example of inpainting to address occlusion. Some regions in a particular reprojected image may not have the proper data from the input image and thus may represent gaps or occlusions in such reprojected image. In the reprojected image 1805, the occlusion regions appear as black regions. For example, occlusion regions are visible to the left of each of the chairs (especially the leftmost chair), to the left of the toolbox, and to the left of the table. These occlusion regions may occur when an object close to the sensor 205 is moved left or right. An occlusion map 1810 of the reprojected image 1805 shows the occlusion regions in white and all non-occluded regions in black. The imaging system 200 modifies the reprojected image 1805 to fill in the occlusion regions using inpainting to generate an inpainted image 1815. In some examples, deep learning based inpainting is used to intelligently inpaint to provide high quality inpainting based on training of a deep learning model used for deep learning based inpainting, which may have been trained based on training data including the original of the image and a second copy of the image with occlusions, as well as the occlusions shown in the reprojection image 1805 and the occlusion map 1810. An example of deep learning based inpainting is shown in the inpainting image 1815.

[0134] In some examples, a less computationally expensive form of inpainting can be used for the inpainting operation, such as interpolation or in-line or nearest value inpainting, based on the available computational bandwidth and / or power allowance of the imaging system 200. An example of interpolation-based inpainting, for example using interpolation and / or in-line or nearest value inpainting, is shown using 3D depth-based zoom at the bottom of FIG. 18. A 3D depth-based zoom image 1825 is shown in FIG. 18, where an occlusion region 1835 is visible between the skateboarder's legs at the previous position of the skateboard. An inpainting image 1830 is shown using interpolation-based inpainting, for example interpolation or in-line or nearest value inpainting, to inpaint this occlusion region 1835.

[0135] FIG. 19 is a block diagram 1900 illustrating the architecture of a reprojection and grid inversion system 1905. The reprojection and grid inversion system 1905 can read data in raster order. In some examples, the reprojection and grid inversion system 1905 reads the MV grid 1910 in raster order and / or reads depth data (e.g., from a depth sensor) in raster order (e.g., first option 1915) to obtain a 3D matrix. For each pixel in the input, for each motion vector and / or depth value in the input, the reprojection and grid inversion system 1905 places the pixel in the output at a location in the output. Each tile number represents a group of pixels in the output. Going in raster order, the pixel indicated by the arrow 1930 goes to tile 1, and the pixel indicated by the arrow 1935 goes to tile 2. Pixels that are not close to each other in the input grid can be closer in the output grid. Based on this, it may be useful to keep the tiles in a cache if the reprojection and grid inversion system 1905 needs to write more data to the tiles. For example, if the reprojection and grid inversion system 1905 starts with tile 1 and then moves to tile 2, the reprojection and grid inversion system 1905 may later need tile 1 again. Keeping the tile in cache (as long as the reprojection and grid inversion system 1905 can be based on a least recently used (LRU) caching system) allows the reprojection and grid inversion system 1905 to quickly modify the tile again and not have to read it from DRAM.

[0136] In some cases, using depth-based reprojection, closer objects can move more than distant objects. Thus, objects from different regions in the input image can appear in the same region in the reprojected image. Pixel / arrow 1930 and pixel / arrow 1940 are an example of this, originating from different locations in the input (e.g., MV grid 1910) but falling into the same region in the output, e.g., tile 1. Thus, the reprojection and grid inversion system 1905 can hold tile 1 in memory so that it can be modified (e.g., overwriting tile 1 with the value of the pixel indicated by arrow 1940). Because holding the entire output buffer in memory hardware can be excessive, the reprojection and grid inversion system 1905 can include a caching mechanism to hold the tiles in the memory hardware.

[0137] If the reprojection and grid inversion system 1905 starts at the beginning of the raster order and this is the first time the reprojection and grid inversion system 1905 wants to write to a tile (e.g., the value of the pixel indicated by arrow 1930 to tile 1), the reprojection and grid inversion system 1905 simply resets tile 1 and writes the value in question to tile 1 without having to read the tile from DRAM first. In some examples, the value from tile 1 may be moved from the cache to the DRAM. The reprojection and grid inversion system 1905 uses a cache so that it does not have to perform too many read / modify / write operations, but the reprojection and grid inversion system 1905 has the capability of read / modify / write operations when necessary. As long as the tiles are in the cache, the reprojection and grid inversion system 1905 can access them immediately. At some point, the cache may become full and the reprojection and grid inversion system 1905 may send tiles from the cache to the DRAM to make room for other tiles (based on LRU). At some other time, the reprojection and grid inversion system 1905 again needs the tile that was sent from the cache to the DRAM, and the reprojection and grid inversion system 1905 can then move the tile back from the DRAM to the cache to correct for this, and at some other time, the tile can be written to the DRAM.

[0138] In addition, the reprojection and grid inversion system 1905 has a prefetch mechanism that allows the reprojection and grid inversion system 1905 to bring up the necessary files in advance of processing to avoid latency issues from reading tiles from DRAM. The reprojection and grid inversion system 1905 operates in an ordered manner and the prefetch mechanism can ensure that the reprojection and grid inversion system 1905 always has what it needs in the cache. The reprojection and grid inversion system 1905 can switch between prefetching and processing in a regular, rather than random, manner to ensure that the reprojection and grid inversion system 1905 processes all of the data in an ordered manner and can have everything it needs to process in the cache.

[0139] The reprojection and grid inversion system 1905 can receive the depth data and a 3D matrix in a first option 1915. In some examples, the reprojection and grid inversion system 1905 can generate an MV grid 1910 from the depth data and the 3D matrix. The reprojection and grid inversion system 1905 can receive the depth data and an MV grid having a 2D matrix in a second option 1920. In some examples, the reprojection and grid inversion system 1905 can generate an MV grid 1910 from the MV grid having the depth data and a 2D matrix. When the reprojection and grid inversion system 1905 receives the depth and a 3D matrix (first option 1915) or when the reprojection and grid inversion system 1905 receives an MV grid and / or a 2D matrix (second option 1920), the reprojection and grid inversion system 1905 uses the coordinate calculation system to calculate the output coordinates (outCoord) and the output data (outData). In some examples, the output data can include an output motion vector (outMV) and an output depth (outDepth). The reprojection and grid inversion system 1905 can also output additional output data (as part of outData), such as confidence (outConf) and / or occlusion (outOcc) to determine where the occlusion regions are. The output from the reprojection and grid inversion system 1905 can be output as output data to one or more buffers, caches, or other memories. In one illustrative example, the output buffers (or caches or other memories) shown on the right side of FIG. 19 include an output buffer (or cache or other memory) for depth, an output buffer (or cache or other memory) for the MV grid (e.g., with depth and / or confidence), and an output buffer (or cache or other memory) for occlusion. These output buffers (or caches or other memories) can be output as multiple output images. The prefetching and caching mechanisms can process three buffers at a time.Since each output buffer can store a different amount of bits for each tile, the prefetching and caching mechanisms can handle synchronization between all the different levels of bits and different sized tiles at all stages.

[0140] In some examples, the reprojection and grid inversion system 1905 uses specialized hardware designed to be particularly efficient in motion vector manipulation, coordinate calculation, caching, pre-fetching, and generating output buffers. In some aspects, a processor, such as a CPU or GPU, may be used to perform certain operations.

[0141] In some examples, the output confidence (outConf) is not generated specifically for reprojection, but is a by-product of the depth measurement from the depth sensor. In some examples, the acquired depth may suffer from measurement inaccuracies and / or other issues that can be represented by the confidence map. It may be beneficial to improve the depth based on the confidence map and / or the visual (RGB) image. The reprojection and grid inversion system 1905 can reproject the depth and confidence so that it matches the visual (RGB) image and the confidence is available in the correct area in the reprojected image. Once the depth matches the RGB image, the reprojection and grid inversion system 1905 can use the confidence to improve the depth.

[0142] In some examples, the imaging system can use a "triangle walk" operation to determine where a given pixel from an input image (e.g., first image Img1 510, captured image 1710) should move to in the reprojected image (e.g., second image Img2 515, reprojected image 1715).

[0143] FIG. 20 is a conceptual diagram 2000 illustrating an example of a triangle walking operation. In some examples, different pixels from an input image can be moved to different locations in the reprojected image. The system can process X inputs at a time, where X is equal to any integer value (e.g., 3, 4, 5, 6, 10, etc.). The system can generate Y output triangles (for each set of inputs), where Y is equal to any integer value (e.g., 6, 7, 8, 9, 10, 15, etc.). The pixels in the inputs include pixel a, pixel b, pixel c, etc. In some examples, pixel data from pixel a in the input image can be moved to a first one of the locations in the reprojected image, pixel data from pixel b in the input image can be moved to a second one of the locations in the reprojected image, pixel data from pixel c in the input image can be moved to a third one of the locations in the reprojected image, and so on. Through a map (e.g., MV grid 505 or inverse MV grid 520), the system finds where each pixel in the input image should go in the reprojected image. Thus, in an illustrative example, pixel a of the input image ends up at pixel 2010 in the output, pixel b of the input ends up at pixel 2015 in the output, pixel 1 of the input ends up at pixel 2020 in the output, and so on. For each input pixel, the imaging system calculates where the value of the input pixel is configured to end up in the output. In areas between certain pixels in the output (e.g., the shaded triangular area between pixels 2010, 2015, and 2020), the imaging system uses interpolation to fill the area. To perform the interpolation, the imaging system can have a processor (e.g., a GPU or other processor) look at each of the triangles separately and interpolate one by one for each output pixel individually.

[0144] However, to improve efficiency, the imaging system can consolidate the triangles to form a larger polygon at the output side of FIG. 20, i.e., a polygon made from all combinations of triangles (including the triangle with pixels 2010, 2015, and 2020). The imaging system can have a dedicated hardware processor specifically designed to be effective at interpolation, or can have another processor perform the interpolation (e.g., a GPU or other processor). Because many of these triangles are close to each other and contain similar image data, it can be inefficient for the imaging system to use a processor (e.g., a GPU) to look at each of the triangles separately and interpolate for each output pixel individually. To improve efficiency, the imaging system can consolidate the triangles into polygons and have the processor (e.g., a GPU) look at the entire polygon at once and perform the interpolation across the pixels of the entire polygon.

[0145] The imaging system includes a main walk engine 2025, N triangle control engines 2030 (N can be equal to any integer value such as 6, 8, 10, or other values), and M pixel interpolation engines 2035 (M can be equal to any integer value such as 6, 8, 10, or other values, and may be equal to N in some implementations). The main walk engine 2025, shown as a white dashed box, walks the entire polygon at once. The N triangle control engines 2030, two of which are shown as dashed and light shaded boxes, each responsible for one of the triangles. The main walk engine 2025 traverses the entire polygon, effectively pre-scanning output locations and / or regions used by the imaging system for image reprojection, caching the data, and thereby causing the imaging system to pre-fetch and / or retrieve data (e.g., tiles) from DRAM to reduce or eliminate delays (e.g., filling, interpolation, or other image processing operations) that might otherwise be caused by retrieving data from DRAM.

[0146] FIG. 21 is a conceptual diagram 2100 illustrating an example of occlusion masking. An occlusion region is a region of a reprojected image for which the image reprojection engine 215 does not have image data available. As previously mentioned, the image reprojection engine 215 performs an interpolation for regions that do not have a particular value in the originally captured image. Even for occlusion regions, this interpolation is still performed, for example to avoid these regions being filled with unreliable data (e.g., no matter what happens to the DRAM). Image 2110 may be an example of filling using such unreliable data. To perform the reprojection, certain objects, such as a toolbox, may be stretched slightly in a particular direction (e.g., horizontally), but this stretching is generally not significant enough to have adverse effects and in some cases may improve the appearance of a new perspective in the reprojected image. However, in certain regions, holes or gaps exceed a threshold size where the interpolation becomes unreliable, which may be determined by the image reprojection engine 215 to be an occlusion region.

[0147] In some examples, the image reprojection engine 215 can determine that an occlusion region exists based on the corner depth. For example, the image reprojection engine 215 can determine that an occlusion region exists in a region (such as a triangle or other shape in FIG. 20) if the difference between the depths at the corners of the region exceeds a threshold difference. The threshold difference can vary based on a minimum value of the depth.

[0148] When the image reprojection engine 215 determines that an occlusion region exists (e.g., based on a difference between depths at the corners of the region exceeding a threshold difference), the image reprojection engine 215 can perform inpainting to fill the occlusion region(s) in the reprojected image with image data. The "unreliable leftovers" in the image 2110 can represent one form of inpainting, using a portion of the toolbox image data in the occlusion region. In some cases, this type of inpainting can work well even if it appears abnormal in the image 2110. In some examples, the occlusion may be performed using deep learning, e.g., using one or more trained ML models.

[0149] FIG. 22 is a conceptual diagram 2200 illustrating an example of hole filling. Hole filling refers to interpolation in gaps where no motion vector data exists. Flow 2220 shows that with hole filling turned off, the reprojected image has many visual artifacts, e.g., black and white dots in visual artifact patterns that are particularly noticeable on the toolbox and other objects near the camera. When hole filling is turned on, the holes in the reprojected image are filled using interpolation and the image looks clean without such visual artifacts or visual artifact patterns. In some examples, hole filling can use inpainting, such as deep learning-based inpainting, instead of or in addition to interpolation.

[0150] FIG. 23 is a conceptual diagram 2300 illustrating an additional example of time warping 705 performed by the time warp engine 230. The time warp engine 230 here calculates dense optical flows between frames n+1 and n, and between frames n and n-1, respectively. The input frame rate (in frames per second (FPS)) is equal to Fin, which may be 30 FPS, 60 FPS, 120 FPS, 240 FPS, or other frame rates. The output frame rate is equal to Fout, which may be 60 FPS, 120 FPS, 240 FPS, 480 FPS, or other frame rates. These dense optical flows are calculated with high quality, but may be computationally expensive and / or use a large amount of power. The time warp engine 230 splits the dense optical flow to generate smaller partial optical flows between other frames, such as between frames n-1 and n, or between frames n and n+1, similar to the time warp 705 of FIG. 7. For example, the time warp engine 230 splits the dense optical flow to generate smaller partial optical flows for frames n+3 / 4, n+1 / 2, n+1 / 4, n-1 / 4, n-1 / 2, and n-3 / 4. These partial optical flows can serve as substitutions to the optical flow as if each of the partial optical flows was directly calculated using the optical flow calculation. These partial optical flows can be decomposed into quarters, as in this example, or into other similar fractions. These partial optical flows, if present, can be used to improve existing frames at frames n+3 / 4, n+1 / 2, n+1 / 4, n-1 / 4, n-1 / 2, and n-3 / 4. These partial optical flows can be used to generate new interpolated frames at frames n+3 / 4, n+1 / 2, n+1 / 4, n-1 / 4, n-1 / 2, and n-3 / 4.In some examples, Time Warp 705 can be used to generate optical flow for videos at high frame rates (e.g., 90, 120, 240, 480, or 960 fps) by first generating dense optical flow for videos at a lower frame rate (e.g., 30 or 60 fps) and then using Time Warp 705 to split the computed dense optical flow into optical flows for the frames in between.

[0151] In some examples, the time warp engine 230 can obtain motion vectors for optical flow, combine the motion vectors with a global matrix, and split the result into partial optical flow or motion vectors, such as in time warp 705 after the combination.

[0152] Further examples of the benefits of image sharpening are shown for images without and using time warp 705. Details are restored using time warp 705, as shown in the areas indicated by the arrows, e.g., the boy's hair, ears, and T-shirt in the center image, and the markings in the right-hand image. In particular, edges and / or areas that appear blurry are represented using dashed lines, and edges and / or areas that appear clear and sharp are represented using solid lines.

[0153] FIG. 24 is a block diagram 2400 illustrating an example architecture of a reprojection engine 24341 in some examples of the time warp engine 230. The optical flow engine 2420 receives frame n and frame nM from a camera 2405 having an image sensor 2410 and a dynamic random access memory (DRAM) 2415. The optical flow engine 2420 generates motion information. In some examples, the motion information includes two types of motion information, including global motion and local motion. For example, a matrix (e.g., a global matrix) may represent global motion in some cases. The optical flow engine may generate a dense grid of motion vectors to indicate local motion and 3D motion. In another example, the dense grid of motion vectors may also indicate global motion and / or a combination of local motion, 3D motion, and global motion.

[0154] The grid inversion engine 2425 receives motion information from the optical flow engine 2420 (e.g., a dense grid of motion vectors and possibly a matrix representing global motion). The grid inversion engine 2425 runs multiple times (M times), with each run splitting the motion vector and outputting a different portion of the motion vector. The grid inversion engine 2425 outputs M motion vectors. In some cases, the motion vectors may be multiplied by a factor. The motion vectors may be downscaled using a warp engine 2430 to provide different resolutions. The warp engine 2430 may receive the motion vectors from the dense grid and perform some warping, scaling, and / or other operations on the dense motion grid. In some examples, the warp engine 2430 may also obtain a transformation matrix and warp the dense grid based on this. In another example, the warp engine 2430 may obtain a transformation matrix and combine it with the dense grid. The inverse motion vectors output by the grid inversion engine 2425 and / or the warp engine 2430 are output to an image processing engine 2440 for generating a reprojected image based on the inverse motion vectors.

[0155] FIG. 25 is a block diagram 2500 illustrating an example architecture of a reprojection engine with temporal deblurring 2535 in some examples of the time warp engine with temporal deblurring 230. The architecture of FIG. 25 is similar to that of FIG. 24, but the system's temporal deblurring engine 2505 determines which M frames are blurred (e.g., based on motion detection and / or image analysis) and uses the partial motion vectors generated by the grid inversion engine 2425 to deblur and / or sharpen the blurred frames. In some examples, the temporal deep learning algorithm of the reprojection engine 2535 analyzes the pose sensor data to see how much movement (and how much blur) there was during the capture of each frame. In some examples, the original motion vectors are provided to the image processing engine 2440 from the optical flow engine 2420, possibly after further transformation 2520 (e.g., shrinking).

[0156] FIG. 26 is a block diagram 2600 illustrating an example architecture of the depth sensor support engine 235. Although a time-of-flight (ToF) sensor is an example of a depth sensor, the depth sensor support engine 235 may use different types of depth sensors as described herein in some examples. Post-processing may be applied to clean up the depth values ​​from the depth sensor to provide higher quality depth values, for example by filtering outliers and / or normalizing noise. In some cases, the post-processing may also receive a confidence map along with the depth, which may then clean the confidence map and / or use the confidence map to assist the depth processing. The depth, and possibly the confidence, are sent to a reprojection engine, which may reproject the depth image and confidence map based on a 3D transformation, for example to align with the image sensor (e.g., wide-angle or telephoto). The reprojection engine may generate reprojected depth and confidence values, which may be run through depth post-processing again to clean up the depth and confidence values. The depth post-processing may also accept images from the wide-angle and telephoto sensors, and / or secondary depth sensor data (e.g., DFS depth) from a secondary depth sensor, and the depth post-processing can adjust the depth to further improve it and correct for inaccuracies resulting from the original depth. The 3D transformation can be based on 3D calibration between the image sensor and the depth sensor. If the depth sensor and the image sensor move relative to each other (e.g., focus change, zoom, OIS, and / or other), the 3D calibration can take this into account and update the 3D transformation. It should be understood that the secondary depth flow (i.e., DFS with wide-angle and telephoto images) at the bottom of FIG. 26 is an illustrative example. In another example, the secondary depth can be obtained from other depth sensors, deep learning depth engines, and / or any other depth source. In some examples, the depth post-processing does not have a secondary depth. In some examples, the depth post-processing can have more than two depth sources.

[0157] FIG. 27 is a conceptual diagram 2700 illustrating additional examples of depth sensor support 805 performed by the depth sensor support engine 235. In these additional examples, the main image sensor (e.g., RGB3) and depth sensor (e.g., TOF system) are shown on a circuit board. Both depth maps and images are shown. In the example on the left (projection alignment 2705), some elements are aligned, but other objects at different distances from the camera, such as the teddy bear or head in the figure, are not aligned between the image data and the depth data. For example, the bear's depth data (e.g., shown using a dashed line) is to the right (parallax shift) compared to the bear's image data. Similarly, the figure's depth data (e.g., shown using a dashed line) is to the right (parallax shift) compared to the image data in the figure. Meanwhile, in the example on the right (depth-based alignment 2710), the parallax is fixed and the depth and image data of each object are aligned.

[0158] FIG. 28 is a block diagram 2800 illustrating an example architecture of an imaging system including an image reprojection engine 215 and a 3D stabilization engine 240. The imaging system takes an input and reprojects the viewpoint to a new location in the environment. For 3D stabilization, it can be performed to reduce or eliminate camera shake and / or to simulate a state in which the camera is stable and / or stabilized so that any movement involves no (or little) shaking or trembling. For example, the 3D stabilization engine 240 of the imaging system can create a virtual path as if the video was captured along the virtual path that involves little or no shaking and / or trembling. The imaging system can also be used for at least some of the other applications of image reprojection described herein, such as time warp, head pose correction, sensor support, etc. The imaging system receives image data and / or depth data as input, stabilizes or otherwise corrects any distortions in the data, and then provides the data to the reprojection engine. For 3D stabilization, the 3D stabilization engine 240 of the imaging system can create a matrix that indicates a stable and smooth virtual path. The imaging system can create a 3D transformation to change the viewpoint of the image. For example, in 3D stabilization, a 3D transformation can change the viewpoint of each of a series of images so that each viewpoint of the images has an origin along a virtual path (e.g., a stable smooth virtual path). The 3D transformation, and possibly the virtual path, can be fed to a reprojection engine. The reprojection engine can generate motion vectors (MV grid) to warp the images to the identified viewpoint (e.g., so that the capture viewpoint is along the virtual path). In some examples, the imaging system can use another motion vector grid to perform lens distortion correction (LDC) and / or rolling shutter correction (RSC) on the images to reduce any distortion from the lens and / or rolling shutter. In another example, the motion vectors and / or matrices can also be used to correct other distortions and / or translation errors.As shown in FIG. 30, in some examples, the 3D stabilization and grids for LDC and RSC are combined together by combining motion vectors from both and warped together. A new set of MVs can perform both 3D stabilization and LDC and RSC. In some examples, the LDC and RSC MV grids can be sparser than the 3D stabilization MV grid, in which case the LDC and RSC MV grids can be upscaled before combining. In some examples, the 3D stabilization MV grids can be sparser than the LDC and RSC MV grids, in which case the 3D stabilization MV grids can be upscaled before combining. The combined MV grid can be sent to a warp engine that performs the warp. The resulting image with 3D stabilization (via reprojection), LDC, and RSC applied is shown.

[0159] Occlusion regions may still remain in the resulting image due to the use of reprojection for 3D stabilization. Depth reprojection, occlusion map, a low-resolution copy of the image (e.g., with full field of view (FoV)), and / or Q high-resolution patches from the image (e.g., 500 patches of size 64×64, or any other number of patches having any suitable size) may be sent to the deep learning engine (NSP) to perform inpainting. For example, the 3D stabilization engine 240 may take a patch from one region but not need to read other regions. The 3D stabilization engine 240 knows which region to focus on in the high-resolution patch due to the occlusion map. In some examples, the patches and occlusion map are small (e.g., the occlusion map is binary or may contain a small number of bits, such as 3 bits, 4 bits, 6 bits, etc.), making the patch a low-cost input to the deep learning engine (NSP) to perform inpainting. Depth reprojection can help ensure that the correct type of material is used for inpainting. For example, the Deep Learning Engine (NSP) does not use nearby objects like the toolbox to inpaint background regions, the only thing it uses for inpainting background regions is image data from background regions at similar depths. This smart inpainting is efficient and uses less power.

[0160] In some examples, inpainting can use temporal filtering, such as using a previous image in a video to capture image content of a particular region. For example, if a previous image has clear image content in an area of ​​a scene depicted in an occlusion region in a current image frame, the image data from the previous image can be used for inpainting and / or 3D stabilization to mitigate shaking. The patches can be aligned with the compression tiles so that the inpainting patches output by the deep learning engine (NSP) can be moved into memory (e.g., directly into DRAM) for the relevant portion of the resulting image.

[0161] 29 is a conceptual diagram 2900 illustrating an additional example of time warping 705 performed using time warp engine 230 compared to an image without time warp engine 230 processing. The example using time warp engine 230 appears clearer and sharper than the image without time warp engine 230, especially at and around edges and corners in the image. For example, edges that appear blurry are reprojected using the dashed lines in FIG. 29, while edges that appear shared and sharp are reprojected using the solid lines in FIG.

[0162] 30 is a conceptual diagram 3000 illustrating a further example 3005 of 3D stabilization 905 performed by the 3D stabilization engine 240. The further example 3005 includes four video frames of a video shown in both original (unstabilized) and stabilized form. Reprojection is used to remove shaking and / or parallax shifting, as previously described.

[0163] 31 is a conceptual diagram 3100 of an additional example of 3D zoom 1005 performed by the 3D zoom engine 245. The digital zoom 3105 crops and upscales as shown using the dashed boxes and lines on the left side of the diagram. A depth image of the skateboarder is shown alongside the 3D depth-based zoom. The 3D depth-based zoom uses reprojection based on the depth image to simulate moving the camera closer to the skateboarder as shown in the illustration 3110 of moving the phone closer to the man.

[0164] 32 is a conceptual diagram 3200 illustrating an additional example of reprojection 1105 performed by reprojection SAT engine 250. Reprojection 1105 uses reprojection from one sensor's viewpoint to a different sensor's viewpoint to shift the viewpoint by an offset.

[0165] FIG. 33 is a conceptual diagram 3300 of an additional example of head pose correction 1205 performed by the head pose correction engine 255. A depth image 3515 of a woman's head is shown, which is the basis for the reprojection. An occlusion map 3320 of the reprojected image 1215 is also shown. A depiction of the person's relative position to the camera is shown below the input image 1210, showing the camera taking the picture from slightly below the user's face, angled slightly upwards. A depiction of the person's simulated relative position to the camera is shown below the reprojected image 1215, showing the simulated camera position taking the picture from an elevation or height that matches the elevation or height of the user's face, an offset distance 3305 away from the position at which the input image 1210 was captured, and an offset angle 3310 away from the angle at which the input image 1210 was captured. The capture angle of the reprojected image 1215 is perpendicular to the person's face, body, and / or gravity.

[0166] FIG. 34 is a conceptual diagram 3400 illustrating additional examples of grid inversion. The original MV grid and the inverse MV grid are shown for a target image with a sun and clouds. Examples where missing content will be filled (via interpolation and / or inpainting) are shown using stars, e.g., part of the sun was occluded by a cloud in the input image but not in the reprojected image. Examples of competing values ​​are shown using circles, e.g., there is data for both clouds and sun, and the clouds are in front of the sun, so the cloud data ultimately wins.

[0167] 35 is a conceptual diagram 3500 illustrating an example of the use of deep learning based inpainting. A set of images is shown, each of which includes an occlusion region 3505 in one of the images of the set. The occlusion region is shown as blank before being filled in using a trained deep learning inpainting engine such as neural network 3900.

[0168] FIG. 36 is a conceptual diagram 3600 illustrating an example of the use of inpainting without deep learning. A set of images are shown arranged in vertical columns. The first column includes images output by the grid inversion engine (RGE) that include occlusion regions 3605, which are shown as blank. The second column includes images output by the grid inversion engine (RGE) where inpainting has been performed to fill the occlusion regions 3605. For example, the inpainting in FIG. 36 can use interpolation and / or inline or nearest value inpainting. Patches for inpainting can be selected based on similarity and / or priority as shown. The third column includes images output by the grid inversion engine (RGE) without the occlusion regions 3605. The third row of images contains blurring or visual "smearing" around some of the edges where the occlusion regions 3605 are in the first row of images, which may appear similar to motion blur and may be caused by other positions and / or depictions of objects from the originally captured images that have been transformed using a Grid Inversion Engine (RGE).

[0169] FIG. 37 is a conceptual diagram 3700 illustrating an example of the use of edge and depth filters on edges. Edge filters can be used in some examples to smooth blocky edges in depth and / or image data, reducing visual artifacts of image reprojection. Although the filters are shown to have a size of 3×3, in some cases the filters can be larger (e.g., 4×4, 6×6, etc.). Edge filters can detect edges in the depth map. Depth filters on edges can reduce interpolated depth values ​​that do not belong to any object.

[0170] Figure 38 is a conceptual diagram 3800 illustrating an example of reprojection. The sensor 205 includes a camera cam1 that captures image(s) and depth data (cam1 depth) of a 3D scene. Inter-camera 3D translation is used to reproject the depicted 3D scene into image(s) in 3D space to use perspective camera cam2. Forward mapping (e.g., motion vector grid) is shown using dashed lines. Reverse mapping (e.g., reverse motion vector grid) is shown using solid lines from cam2 back to cam1.

[0171] 39 is a block diagram illustrating an example of a neural network (NN) 3900 that may be used for media processing operations. The neural network 3900 may include any type of deep network, such as a convolutional neural network (CNN), an autoencoder, a deep belief net (DBN), a Recurrent Neural Network (RNN), a Generative Adversarial Network (GAN), and / or other types of neural networks. Neural network 3900 may be an example of one of the one or more trained neural networks of imaging system 200, such as a neural network of any of application engines 210, such as the image reprojection engine 215, motion vector engine 220, grid inversion engine 225, time warp engine 230, depth sensor support engine 235, 3D stabilization engine 240, 3D zoom engine 245, reprojection SAT engine 250, head pose correction engine 255, XR late stage reprojection engine 260, special effects engine 265, or a combination thereof.

[0172] The input layer 3910 of the neural network 3900 includes input data. The input data of the input layer 3910 may include data representing pixels of one or more input image frames, such as media data 285, sensor data from the sensor(s) 205, virtual content from the virtual content generator 207, or a combination thereof. The input data of the input layer 3910 may include depth data from a depth sensor(s). The input data of the input layer 3910 may include motion vectors and / or optical flow. The input data of the input layer 3910 may include matrices. The input data of the input layer 3910 may include an occlusion map.

[0173] The image may include image data from an image sensor including raw pixel data (e.g., including a single color per pixel based on a Bayer filter) or processed pixel values ​​(e.g., RGB pixels for an RGB image). The neural network 3900 includes multiple hidden layers 3912A, 3912B to 3912N. The hidden layers 3912A, 3912B to 3912N include "N" hidden layers, where "N" is an integer equal to or greater than 1. The number of hidden layers may be adapted to include as many layers as required for a given application. The neural network 3900 further includes an output layer 3914 that provides output resulting from the processing performed by the hidden layers 3912A, 3912B to 3912N.

[0174] In some examples, the output layer 3914 may provide an output image or a portion thereof, such as the modified media data 290, any reprojected image described herein, any reprojected depth data described herein, any motion vectors or optical flow described herein, any inpainting image data described herein, or a combination thereof.

[0175] Neural network 3900 is a multi-layered neural network of interconnected filters. Each filter can be trained to learn features that represent the input data. Information related to the filters is shared between different layers, with each layer retaining the information as it is processed. In some cases, neural network 3900 can include a feed-forward network, in which there are no feedback connections where the output of the network is fed back to itself. In some cases, network 3900 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading the input.

[0176] In some cases, information can be exchanged between layers through interconnections of nodes and nodes between various layers. In some cases, the network can include a convolutional neural network, which may not connect every node in one layer to every other node in the next layer. In a network in which information is exchanged between layers, the nodes of the input layer 3910 can activate a set of nodes in the first hidden layer 3912A. For example, as shown, each of the input nodes of the input layer 3910 can be connected to each of the nodes of the first hidden layer 3912A. The nodes of the hidden layer can transform the information of each input node by applying an activation function (e.g., a filter) to this information. The information derived from the transformation can then be passed to the nodes of the next hidden layer 3912B to activate those nodes, which can perform their own designated functions. Exemplary functions include convolution functions, downsampling, upscaling, data transformation, and / or any other suitable functions. The output of hidden layer 3912B may then activate a node in the next hidden layer, and so on. The output of the final hidden layer 3912N may activate one or more nodes in output layer 3914, which provides a processed output image. In some cases, a node in neural network 3900 (e.g., node 3916) is shown as having multiple output lines, however, the node has a single output and all lines shown as outputting from the node represent the same output value.

[0177] In some cases, each node or interconnection between nodes can have a weight, which is a set of parameters derived from training of the neural network 3900. For example, the interconnections between nodes can represent information learned about the interconnected nodes. The interconnections can have adjustable numerical weights that can be adjusted (e.g., based on a training data set), allowing the neural network 3900 to be adaptive to the input and to learn as more and more data is processed.

[0178] The neural network 3900 is pre-trained to process features from the data in the input layer 3910 using different hidden layers 3912A, 3912B through 3912N to provide output through the output layer 3914.

[0179] 40 is a flow diagram illustrating a process for media processing. The process 4000 may be implemented by a media processing system. In some examples, the media processing system may include, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the imaging system 200, the HMD 310, the mobile handset 410, the reprojection and grid inversion system 2490, the system of FIG. 25, the system of FIG. 26, the system of FIG. 27, the system of FIG. 28, the neural network 3900, the computing system 4100, the processor 4110, or a combination thereof.

[0180] At operation 4005, the media processing system is configured to receive and may receive depth data including depth information corresponding to the environment. In some examples, the depth information may include depth measurements of a representation of the environment from a first perspective. In some examples, the depth information includes a point cloud corresponding to the environment. In some examples, the depth data may be captured using one or more depth sensors, such as one or more Light Detection and Ranging (LIDAR) sensors, Radio Detection and Ranging (RADAR) sensors, Acoustic Detection and Ranging (SODAR) sensors, Acoustic Navigation and Ranging (SONAR) sensors, Time of Flight (ToF) sensors, structured light sensors, or combinations thereof. In some examples, the depth data may be captured using one or more cameras and / or image sensors based on stereoscopic depth sensing, for example using a stereo camera configuration. In some examples, depth data may be captured using image capture and processing system 100, sensor 205, cameras 330A-330B, cameras 430A-430D, image sensor 810, depth sensor 815, telephoto sensor 1110, wide-angle sensor 1115, sensor 1125, image sensor 2610, cam1 of FIG. 38, cam2 of FIG. 38, any other sensor described herein, or a combination thereof. Examples of depth data include media data 285, depth data 620, depth data 1020, depth data 1160, depth data 1220, the depth data of FIG. 15, depth map 1610, depth data associated with the first option 1915, depth input 2402, the depth of FIG. 26, the depth data of FIG. 27, the depth data of FIG. 28, depth data 3315, depth image 3410, the depth map of FIG. 37, the Cam1 depth of FIG. 38, any other depth data described herein, or combinations thereof.

[0181] At operation 4010, the media processing system is configured to receive and may receive first image data captured by an image sensor, the first image data including a depiction of an environment. In some examples, the first image data may be captured using image capture and processing system 100, sensor 205, cameras 330A-330B, cameras 430A-430D, image sensor 810, depth sensor 815, telephoto sensor 1110, wide angle sensor 1115, sensor 1125, image sensor 2610, cam1 of FIG. 38, cam2 of FIG. 38, any other sensor described herein, or a combination thereof. Examples of the first image data include media data 285, first image Img1 510, camera image 610, image 710, the “original” image of FIG. 9, the original non-zoomed image (before zoom) of FIG. 10, telephoto image 1130, input image 1210, input image 1310, input image 1410, capture image 1510, capture image 1710, input image 1 of flow 2310 flow 2320, the input image without time warp 705 of FIG. 25, frames n and nM of FIG. 24-FIG. 25, the m blurred frames of FIG. 25, the wide angle and telephoto images of FIG. 26, the input image of FIG. 27, the “original” image of FIG. 30, the non-zoomed input image of FIG. 31, the input image of FIG. 34, the input image of FIG. 35, the input image of FIG. 36, the original pixels of FIG. 38, the image(s) provided to the input layer 3910, other image data described herein, or combinations thereof.

[0182] At operation 4015, the media processing system is configured to and can generate, based at least on the depth data, a first plurality of motion vectors corresponding to a change in viewpoint of a depiction of the environment in the first image data. Examples of the first plurality of motion vectors include the motion vectors in the MV grid 505, the motion vectors of FIG. 15 (e.g., the MV in , M.V. x , M.V. y), the MV 1620, the dense MV of FIG. 23, the motion vectors associated with the optical flow engine 2420, the MV grid of FIG. 28, the original MV and MV grid of FIG. 34, the forward mapping of FIG. 38, other motion vectors described herein, or a combination thereof.

[0183] At operation 4020, the media processing system is configured to and can generate, using grid inversion based on the first plurality of motion vectors, a second plurality of motion vectors indicative of respective distances moved by respective pixels of the representation of the environment in the first image data for a change in viewpoint. Examples of the second plurality of motion vectors include motion vectors in the inverse MV grid 520, the inverse MV 1630, the inverse MV 1730, the inverse motion vectors associated with the grid inversion engine 2425, the MV grid of Figure 28, the inverse MV and MV grid of Figure 24, the backward mapping of Figure 38, other inverse motion vectors described herein, or combinations thereof.

[0184] In operation 4025, the media processing system is configured to and may generate second image data by at least partially rectifying the first image data according to a second plurality of motion vectors, the second image data including a second representation of the environment from a different perspective than the first image data. Examples of the second image data include rectified media data 290, second image Img2 515, reprojected image 615, image 715, the "stable" image of FIG. 9, the 3D zoom image of FIG. 10, rectified telephoto image 1140, reprojected image 1215, input image 1315, reprojected image 1415, reprojected image 1515, reprojected image 1715, reprojected image 1805, inpainting image 1815, reprojected image 2110, reprojected image 2115, reprojected images of flow 2210, reprojected images of flow 2 ... The projection images include images output using the image processing engine 2440, the depth-based aligned 2710 image of FIG. 27, the time warped image of FIG. 29, the "stable" image of FIG. 30, the depth-based 3D zoom image of FIG. 31, the output image of FIG. 34, the output image of FIG. 35, the output image of FIG. 36, the reprojected pixels of FIG. 38, the image(s) output using the output layer 3914, other image data described herein, or combinations thereof.

[0185] In some examples, the second image data includes an interpolated image configured to depict the environment at a second time between the first time and the third time. In such examples, the first image data includes at least one image depicting the environment at at least one of the first time or the third time. Examples of such image interpolation may be performed using time warp 705 as in FIG. 7 and / or FIG. 23. In some examples, the imaging system may generate the interpolated image without using depth data.

[0186] In some examples, the first image data includes a plurality of frames of video data exhibiting parallax movement, and the second image data includes a stabilized transformation of the plurality of frames of video data that reduces the parallax movement. For example, the 3D stabilization 905 can stabilize, reduce, and / or eliminate parallax movement, rotation, or a combination thereof, as in FIG. 9 and / or FIG. 30.

[0187] In some examples, the first image data includes a person viewing the image sensor from a first angle, and the second image data includes a person viewing the image sensor from a second angle different than the first angle. This example includes head pose correction 1205, as in FIG. 12 and / or FIG. 33.

[0188] In some examples, the change in viewpoint includes a rotation of the viewpoint about an axis according to an angle. In some examples, the change in viewpoint includes a translation of the viewpoint according to a direction and a distance. In some examples, the change in viewpoint includes a transformation. In some examples, the change in viewpoint includes a movement along an axis between an original viewpoint of the representation of the environment in the first image data and a position of an object in the environment, at least a portion of which is depicted in the first image data. In some examples, the rotation, translation, transformation, and / or movement can be identified based on what is needed to perform any of the types of reprojection and / or warping described herein, for example, in any of FIGS. 7-14. In some examples, the rotation, translation, transformation, and / or movement can be identified using a user interface. In some examples, the change in viewpoint includes at least one of a parallax movement of the viewpoint or a rotation of the viewpoint about an axis, further including receiving, via the user interface, an indication of one of a distance of the parallax movement of the viewpoint, or an indication of an angle or axis of the rotation of the viewpoint.

[0189] At operation 4030, the media processing system is configured to and may output the second image data (e.g., using output device(s) 270). For example, the media processing system may display the second image data, output the second image data for further processing, store the second image data, any combination of these, and / or output the second image data in other manners.

[0190] In some examples, outputting the second image data includes causing the second image data to be displayed using at least one display. In some examples, outputting the second image data includes causing the second image data to be transmitted to at least a receiving device using at least the communication interface.

[0191] In some examples, the media processing system is configured to and can modify the second image data by identifying one or more gaps in the second image data based on one or more gaps in the second plurality of motion vectors and at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data. In some examples, the media processing system is configured to and can modify the second image data by identifying one or more gaps in the second plurality of motion vectors that cause one or more gaps in the second image data based on one or more gaps at respective endpoints of the first plurality of motion vectors and at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data. Examples of gaps include gaps in the inverted MV grid 520 (and / or the second image Img2 515) indicated by stars in FIG. 5.

[0192] In some examples, the media processing system is configured to and may: identify one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors; and modify the second image data by at least partially filling the one or more gaps in the second image data using inpainting before outputting the second image data. The inpainting may use interpolation, machine learning, neural networks, or combinations thereof. Examples of inpainting are shown in FIG. 18, FIG. 21, FIG. 22, FIG. 28, FIG. 33, FIG. 34, FIG. 35, FIG. 36, and / or FIG. 37.

[0193] In some examples, the media processing system is configured to and may modify the second image data by identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors and at least partially filling the one or more gaps in the second image data using inpainting using one or more trained machine learning models before outputting the second image data. The inpainting may use interpolation, machine learning, neural networks, or combinations thereof. Examples of inpainting are shown in Figures 18, 21, 22, 28, 33, 34, 35, 36, and / or 37.

[0194] In some examples, the media processing system is configured to and may: identify one or more conflicts in the second image data based on one or more conflict values ​​from the first image data in the second plurality of motion vectors, and select one of the one or more conflict values ​​from the first image data based on the motion data associated with the second plurality of motion vectors. An example of the one or more conflicts includes a conflict in cell 8 of inverted MV grid 520.

[0195] In some examples, the representation of the environment in the first image data depicts the environment from a first viewpoint and the change in viewpoint is a change between the first viewpoint and a different viewpoint corresponding to a second representation of the environment in the second image data, in some examples, the first plurality of motion vectors point from the first viewpoint to a different viewpoint and the second plurality of motion vectors point from the different viewpoint to the first viewpoint.

[0196] In some examples, the processes described herein (e.g., process 4000 and / or other processes described herein) may be performed by a computing device or apparatus. In some examples, the processes described herein may be performed by image capture and processing system 100, image capture device 105A, image processing device 105B, image processor 150, ISP 154, host processor 152, imaging system 200, HMD 310, mobile handset 410, reprojection and grid inversion system 2490, the system of FIG. 23, the system of FIG. 24, the system of FIG. 25, the system of FIG. 26, the system of FIG. 28, the system of FIG. 29, neural network 3900, computing system 4100, processor 4110, or a combination thereof.

[0197] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device, having resource capabilities to perform the processes described herein. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) configured to perform steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other component(s). The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0198] Components of a computing device may be implemented in circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0199] The processes described herein are illustrated as logic flow diagrams, block diagrams, or conceptual diagrams, whose operations represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement a process.

[0200] In addition, the processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0201] Fig. 41 is a diagram illustrating an example of a system for implementing certain body elements of the present technology. In particular, Fig. 41 illustrates an example of a computing system 4100, which may be, for example, an internal computing system, a remote computing system, a camera, or any computing device constituting any of these components, in which the components of the system communicate with each other using a connection 4105. The connection 4105 may be a physical connection using a bus, or a direct connection to a processor 4110, such as in a chipset architecture. The connection 4105 may also be a virtual connection, a network connection, or a logical connection.

[0202] In some embodiments, the computing system 4100 is a distributed system in which the functionality described in this disclosure may be distributed across one data center, multiple data centers, a peer network, etc. In some embodiments, one or more of the system components described represent many components, each performing some or all of the functionality that is the subject of the component description. In some embodiments, the components may be physical or virtual devices.

[0203] The exemplary system 4100 includes at least one processing unit (CPU or processor) 4110 and connections 4105 coupling various system components to the processor 4110, including system memory 4115, such as read only memory (ROM) 4120 and random access memory (RAM) 4125. The computing system 4100 may include a cache 4112 of high speed memory, either directly connected to the processor 4110, in close proximity to the processor 4110, or integrated as part of the processor 4110.

[0204] The processor 4110 may include any general-purpose processor, as well as hardware or software services, such as services 4132, 4134, and 4136, stored in storage device 4130 and configured to control the processor 4110, as well as special-purpose processors whose software instructions are built into the actual processor design. The processor 4110 may essentially be a completely self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0205] To enable user interaction, the computing system 4100 includes input devices 4145, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. The computing system 4100 may also include output devices 4135, which may be one or more of several output mechanisms. In some cases, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 4100. The computing system 4100 may generally include a communication interface 4140, which may govern and manage user input and system output.The communications interface may be any of the following: audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, Apple® Lightning® port / plug, Ethernet port / plug, fiber optic port / plug, proprietary wired port / plug, BLUETOOTH® wireless signal transmission, BLUETOOTH® low energy (BLE) wireless signal transmission, IBEACON® wireless signal transmission, radio-frequency identification (RFID) wireless signal transmission, near-field communications (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WLAN), and Bluetooth® wireless signal transmission. The wireless communication device may perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including those utilizing WiMAX (Wireless Access), infrared (IR) communications wireless signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or any combination thereof.The communication interface 4140 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine the position of the computing system 4100 based on reception of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite system (BDS), and the European Galileo GNSS. There is no constraint to operate with any particular hardware arrangement, and therefore the basic features herein may be easily substituted for improved hardware or firmware arrangements as they are developed.

[0206] The storage device 4130 can be a non-volatile and / or non-transitory and / or computer readable memory device, and can be a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, a flash memory, a memristor memory, any other solid state memory, a compact disc read only memory (CD-ROM) optical disk, a rewritable compact disc (CD) optical disk, a digital video disk (DVD) optical disk, a blu-ray disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) circuit, IC chips / cards, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASH EPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random-access memory (RRAM),The memory may be a hard disk or other type of computer readable medium capable of storing data that is accessible by a computer, such as a memory, RRAM / ReRAM, phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0207] The storage device 4130 may include software services, servers, services, etc. that cause the system to perform functions when code defining such software is executed by the processor 4110. In some embodiments, hardware services that perform particular functions may include software components stored in a computer-readable medium in conjunction with the necessary hardware components, such as the processor 4110, connections 4105, output devices 4135, etc., to perform the functions.

[0208] The term "computer-readable medium" as used herein includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, storing, or transporting instruction(s) and / or data. Computer-readable media may also include non-transitory media on which data is stored and does not include carrier waves and / or transitory electronic signals propagating wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memories, or memory devices. A computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted using any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0209] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referring to non-transitory computer-readable storage media, media such as energy, carrier signals, electromagnetic waves, and the signals themselves are expressly excluded.

[0210] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the embodiments may be practiced without these specific details. For ease of explanation, in some cases, the present technology may be presented as including individual functional blocks, including devices, device components, steps or routines in a method embodied in software, or functional blocks comprising a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.

[0211] Individual embodiments may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although the flowcharts may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of steps may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.

[0212] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored on or available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions, or in some cases configure a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, etc.

[0213] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) to perform the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor or processors may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small-footprint personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, and the like. The functionality described herein may also be embodied in a peripheral device or an add-in card. Such functionality may also be implemented on a circuit board, between different chips on a single device, or between different processes running thereon, as further examples.

[0214] The instructions, media for carrying such instructions, computing resources for executing them, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.

[0215] In the above description, aspects of the present application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while exemplary embodiments of the present application have been described in detail herein, it should be understood that the inventive concepts may be embodied and employed in various other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. The various features and aspects of the present application described above may be used individually or jointly. Moreover, the embodiments may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present specification. Thus, the present specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, the methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in an order different from that described.

[0216] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of this description.

[0217] When a component is described as being "configured to" perform a particular operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming a programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof.

[0218] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or that is in communication, either directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).

[0219] Claim language or other language stating "at least one" of a set and / or "one or more" of a set indicates that one member of a set or multiple members of a set (in any combination) satisfy the claim. For example, a claim language stating "at least one of A and B" means A, B, or A and B. In another example, a claim language stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one" of a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, a claim language stating "at least one of A and B" can mean A, B, or A and B, and can further include items not listed in the set of A and B.

[0220] The various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, the various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0221] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together as an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium, such as random access memory (RAM), such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage medium, etc. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer.

[0222] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuit configurations. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor, or alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or may be incorporated within a combined video encoder-decoder (CODEC).

[0223] Exemplary aspects of the present disclosure include the following.

[0224] Aspect 1A. An apparatus for image processing, the apparatus comprising at least one memory and at least one processor coupled to the at least one memory, the at least one processor configured to:

[0225] Aspect 2A. The apparatus of Aspect 1A, wherein the second image data includes an interpolated image configured to depict the environment at a second time between the first time and the third time, and the first image data includes at least one image depicting the environment at at least one of the first time or the third time.

[0226] Aspect 3A. The apparatus of any one of aspects 1A to 2A, wherein the first image data includes a plurality of frames of video data including parallax movement, and the second image data includes a stabilized transformation of the plurality of frames of video data that reduces the parallax movement.

[0227] Aspect 4A. The apparatus of any one of Aspects 1A to 3A, wherein the first image data includes a person viewing the image sensor from a first angle, and the second image data includes a person viewing the image sensor from a second angle that is different from the first angle.

[0228] Embodiment 5A. The apparatus of any one of embodiments 1A to 4A, wherein the change in viewpoint includes a rotation of the viewpoint around an axis according to an angle.

[0229] Embodiment 6A. The apparatus of any one of embodiments 1A to 5A, wherein the change in viewpoint includes a translation of the viewpoint according to direction and distance.

[0230] Embodiment 7A. The apparatus of any one of embodiments 1A to 6A, wherein the change in viewpoint includes a translation.

[0231] Aspect 8A. An apparatus described in any one of aspects 1A to 7A, wherein the change in viewpoint includes movement along an axis between an original viewpoint of the representation of the environment in the first image data and a position of an object in the environment, at least a portion of which is depicted in the first image data.

[0232] Aspect 9A. The apparatus of any one of aspects 1A to 8A, wherein at least one processor is configured to identify one or more gaps in the second image data based on one or more gaps in the second plurality of motion vectors, and to modify the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0233] Aspect 10A. The apparatus of any one of aspects 1A to 9A, wherein at least one processor is configured to identify one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors, and to modify the second image data by at least partially filling the one or more gaps in the second image data using inpainting before outputting the second image data.

[0234] Aspect 11A. The apparatus of any one of aspects 1A to 10A, wherein at least one processor is configured to identify one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors, and modify the second image data by at least partially filling the one or more gaps in the second image data using inpainting using one or more trained machine learning models before outputting the second image data.

[0235] Aspect 12A. The apparatus of any one of aspects 1A to 11A, wherein at least one processor is configured to identify one or more conflicts in the second image data based on one or more conflict values ​​from the first image data in the second plurality of motion vectors, and select one of the one or more conflict values ​​from the first image data based on the movement data associated with the second plurality of motion vectors.

[0236] Aspect 13A. The apparatus of any one of aspects 1A to 12A, wherein the depth information includes a three-dimensional representation of the environment from a first perspective.

[0237] Embodiment 14A. The device of any one of embodiments 1A to 13A, wherein the depth data is received from at least one depth sensor.

[0238] Embodiment 15A. The device of any one of embodiments 1A to 14A, further comprising a display, and wherein to output the second image data, the at least one processor is configured to display the second image data using at least the display.

[0239] Aspect 16A. The apparatus of any one of Aspects 1A to 15A, further comprising a communications interface, and wherein to output the second image data, the at least one processor is configured to transmit at least the second image data to at least a receiving device using at least the communications interface.

[0240] Aspect 17A. The apparatus of any one of Aspects 1A to 16A, wherein the apparatus includes at least one of a head mounted display (HMD), a mobile handset, or a wireless communication device.

[0241] Aspect 18A. An apparatus described in any one of aspects 1A to 17A, wherein the representation of the environment in the first image data depicts the environment from a first viewpoint, and the change in viewpoint is a change between the first viewpoint and a different viewpoint corresponding to a second representation of the environment in the second image data.

[0242] Aspect 19A. The apparatus of any one of aspects 1A to 18A, wherein the change in viewpoint includes at least one of a parallax movement of the viewpoint or a rotation of the viewpoint around an axis, and wherein at least one processor is configured to receive, via a user interface, an indication of one of a distance of the parallax movement of the viewpoint or an indication of an angle or axis of the rotation of the viewpoint.

[0243] Aspect 20A. The apparatus of any one of aspects 1A to 19, wherein at least one processor is configured to identify one or more gaps in the second plurality of motion vectors that cause one or more gaps in the second image data based on one or more gaps at respective endpoints of the first plurality of motion vectors, and to modify the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0244] Aspect 21A. A method for image processing, the method including: receiving depth data including depth information corresponding to an environment; receiving first image data captured by an image sensor, the first image data including a representation of the environment; generating a first plurality of motion vectors corresponding to a change in viewpoint of the representation of the environment in the first image data based at least on the depth data; generating a second plurality of motion vectors indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint using grid inversion based on the first plurality of motion vectors; generating second image data by at least partially correcting the first image data corresponding to the second plurality of motion vectors, the second image data including a second representation of the environment from a different viewpoint than the first image data; and outputting the second image data.

[0245] Aspect 22A. The method of aspect 21A, wherein the second image data includes an interpolated image configured to depict the environment at a second time between the first time and the third time, and the first image data includes at least one image depicting the environment at at least one of the first time or the third time.

[0246] Aspect 23A. The method of any one of aspects 21A to 22A, wherein the first image data includes a plurality of frames of video data including parallax movement, and the second image data includes a stabilized transformation of the plurality of frames of video data that reduces the parallax movement.

[0247] Aspect 24A. The method of any one of aspects 21A to 23A, wherein the first image data includes a person viewing the image sensor from a first angle, and the second image data includes a person viewing the image sensor from a second angle different from the first angle.

[0248] Embodiment 25A. The method of any one of embodiments 21A to 24A, wherein the change in viewpoint includes a rotation of the viewpoint around an axis according to an angle.

[0249] Embodiment 26A. The method of any one of embodiments 21A to 25A, wherein the change in viewpoint includes a translation of the viewpoint according to direction and distance.

[0250] Embodiment 27A. The method of any one of embodiments 21A to 26A, wherein the change in viewpoint includes a transformation.

[0251] Aspect 28A. A method according to any one of aspects 21A to 27A, wherein the change in viewpoint includes a movement along an axis between an original viewpoint of the representation of the environment in the first image data and a position of an object in the environment, at least a portion of which is depicted in the first image data.

[0252] Embodiment 29A. The method of any one of embodiments 21A to 28A, further comprising: identifying one or more gaps in the second image data based on one or more gaps in the second plurality of motion vectors; and modifying the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0253] Embodiment 30A. The method of any one of embodiments 21A to 29A, further comprising: identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors; and modifying the second image data by at least partially filling the one or more gaps in the second image data using inpainting before outputting the second image data.

[0254] Aspect 31A. The method of any one of aspects 21A to 30A, further comprising: identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors; and modifying the second image data by at least partially filling the one or more gaps in the second image data using inpainting using one or more trained machine learning models before outputting the second image data.

[0255] Aspect 32A. The method of any one of aspects 21A to 31A, further comprising: identifying one or more conflicts in the second image data based on one or more conflict values ​​from the first image data in the second plurality of motion vectors; and selecting one of the one or more conflict values ​​from the first image data based on the movement data associated with the second plurality of motion vectors.

[0256] Embodiment 33A. The method of any one of embodiments 21A to 32A, wherein the depth information includes a three-dimensional representation of the environment from a first viewpoint.

[0257] Embodiment 34A. The method of any one of embodiments 21A to 33A, wherein the depth data is received from at least one depth sensor.

[0258] Embodiment 35A. The method of any one of embodiments 21A to 34A, wherein outputting the second image data includes displaying the second image data using at least one display.

[0259] Aspect 36A. The method of any one of aspects 21A to 35A, wherein outputting the second image data includes causing the second image data to be transmitted to at least the receiving device using at least the communication interface.

[0260] Aspect 37A. The method of any one of aspects 21A to 36A, wherein the method is performed using an apparatus including at least one of a head mounted display (HMD), a mobile handset, or a wireless communication device.

[0261] Aspect 38A. A method according to any one of aspects 21A to 37A, wherein the representation of the environment in the first image data depicts the environment from a first viewpoint, and the change in viewpoint is a change between the first viewpoint and a different viewpoint corresponding to a second representation of the environment in the second image data.

[0262] Aspect 39A. The method of any one of aspects 21A to 38A, wherein the change in viewpoint includes at least one of a parallax shift of the viewpoint or a rotation of the viewpoint around an axis, and further includes receiving, via a user interface, an indication of one of a distance of the parallax shift of the viewpoint or an indication of an angle or axis of the rotation of the viewpoint.

[0263] Embodiment 40A. The method of any one of embodiments 21A to 39A, further comprising: identifying one or more gaps in the second plurality of motion vectors that cause one or more gaps in the second image data based on one or more gaps at respective endpoints of the first plurality of motion vectors; and modifying the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0264] Aspect 41A. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to receive depth data including depth information corresponding to an environment; receive first image data captured by an image sensor, the first image data including a representation of the environment; generate a first plurality of motion vectors corresponding to a change in viewpoint of the representation of the environment in the first image data based at least on the depth data; generate a second plurality of motion vectors using grid inversion based on the first plurality of motion vectors indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint; generate second image data by at least partially modifying the first image data corresponding to the second plurality of motion vectors, the second image data including a second representation of the environment from a different viewpoint than the first image data; and output the second image data.

[0265] Embodiment 42A. The non-transitory computer-readable medium of embodiment 41A, further comprising the operations of any one of embodiments 2A to 20A and / or any one of embodiments 22A to 40A.

[0266] Aspect 43A. An apparatus for image processing, the apparatus comprising: means for receiving first image data captured by an image sensor, the first image data including a representation of an environment; means for generating a first plurality of motion vectors corresponding to a change in viewpoint of the representation of the environment in the first image data based at least on the depth data; means for generating a second plurality of motion vectors using grid inversion based on the first plurality of motion vectors, the second plurality of motion vectors indicating a respective distance moved by each pixel of the representation of the environment in the first image data for the change in viewpoint; means for generating second image data by at least partially modifying the first image data corresponding to the second plurality of motion vectors, the second image data including a second representation of the environment from a different viewpoint than the first image data; and means for outputting the second image data.

[0267] Embodiment 44A. The apparatus of embodiment 43A, further comprising the operation of any one of embodiments 2A to 20A and / or any one of embodiments 22A to 40A.

[0268] Aspect 1B. An apparatus for image processing, the apparatus comprising at least one memory and one or more processors coupled to the at least one memory, the one or more processors configured to: receive depth data captured by a depth sensor, the depth data including a three-dimensional representation of a representation of an environment from a first perspective, determine a first plurality of motion vectors corresponding to a change from the first perspective to a second perspective based at least on the depth data, receive first image data captured by an image sensor, the first image data depicting an environment from a third perspective, determine a second plurality of motion vectors corresponding to a change from the third perspective to a fourth perspective using grid inversion based on the first plurality of motion vectors, generate second image data by at least partially modifying the first image data corresponding to the second plurality of motion vectors, the second image data depicting the environment from the fourth perspective, and output the second image data.

[0269] Aspect 2B. The apparatus of Aspect 1B, wherein the second image data includes an interpolated image configured to depict the environment at a second time between the first time and the third time, and the first image data includes a first image depicting the environment at the first time and a second image depicting the environment at the third time.

[0270] Aspect 3B. The apparatus of any one of aspects 1B to 2B, wherein the first image data includes video data including parallax motion and the second image data includes a stabilized version of the video data without parallax motion.

[0271] Aspect 4B. The apparatus of any one of aspects 1B to 3B, wherein the first image data includes and depicts a person viewing the image sensor from a first angle, and the second image data includes and depicts a person viewing the image sensor from a second angle that is different from the first angle.

[0272] Embodiment 5B. The apparatus of any one of embodiments 1B to 4B, wherein the fourth viewpoint is the first viewpoint.

[0273] Embodiment 6B. The apparatus of any one of embodiments 1B to 5B, wherein the fourth viewpoint is the second viewpoint.

[0274] Embodiment 7B. The apparatus of any one of embodiments 1B to 6B, wherein the change from the first viewpoint to the second viewpoint includes a rotation of the viewpoint according to an angle, and the change from the third viewpoint to the fourth viewpoint includes a rotation of the viewpoint according to an angle.

[0275] Embodiment 8B. The apparatus of any one of embodiments 1B to 7B, wherein the change from the first viewpoint to the second viewpoint includes a translation of the viewpoint according to a direction and distance, and the change from the third viewpoint to the fourth viewpoint includes a translation of the viewpoint according to the direction and distance.

[0276] Embodiment 9B. The apparatus of any one of embodiments 1B to 8B, wherein the change from the first viewpoint to the second viewpoint comprises a transformation, and the change from the third viewpoint to the fourth viewpoint comprises a transformation.

[0277] Embodiment 10B. The apparatus of any one of embodiments 1B to 9B, wherein the one or more processors are configured to identify one or more gaps in the second image data based on one or more gaps in the second plurality of motion vectors, and modify the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0278] Aspect 11B. The apparatus of any one of aspects 1B to 10B, wherein the one or more processors are configured to identify one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors, and modify the second image data by at least partially filling the one or more gaps in the second image data using inpainting before outputting the second image data.

[0279] Aspect 12B. A method for image processing, the method including: receiving depth data captured by a depth sensor, the depth data including a three-dimensional representation of an environment from a first viewpoint; determining a first plurality of motion vectors corresponding to a change from the first viewpoint to a second viewpoint based at least on the depth data; receiving first image data captured by an image sensor, the first image data depicting an environment from a third viewpoint; determining a second plurality of motion vectors corresponding to a change from the third viewpoint to a fourth viewpoint using grid inversion based on the first plurality of motion vectors; generating second image data by at least partially correcting the first image data corresponding to the second plurality of motion vectors, the second image data depicting the environment from the fourth viewpoint; and outputting the second image data.

[0280] Aspect 13B. The method of aspect 12B, wherein the second image data includes an interpolated image configured to depict the environment at a second time between the first time and the third time, and the first image data includes a first image depicting the environment at the first time and a second image depicting the environment at the third time.

[0281] Aspect 14B. The method of any one of aspects 12B to 13B, wherein the first image data includes video data including parallax motion, and the second image data includes a stabilized version of the video data without parallax motion.

[0282] Embodiment 15B. The method of any one of embodiments 12B to 14B, wherein the first image data includes and depicts a person viewing the image sensor from a first angle, and the second image data includes and depicts a person viewing the image sensor from a second angle that is different from the first angle.

[0283] Embodiment 16B. The method of any one of embodiments 12B to 15BB, wherein the fourth aspect is the first aspect.

[0284] Embodiment 17B. The method of any one of embodiments 12B to 16B, wherein the fourth aspect is the second aspect.

[0285] Embodiment 18B. The method of any one of embodiments 12B to 17B, wherein the change from the first viewpoint to the second viewpoint includes a rotation of the viewpoint according to an angle, and the change from the third viewpoint to the fourth viewpoint includes a rotation of the viewpoint according to an angle.

[0286] Embodiment 19B. A method according to any one of embodiments 12B to 18B, wherein the change from the first viewpoint to the second viewpoint includes a translation of the viewpoint according to a direction and distance, and the change from the third viewpoint to the fourth viewpoint includes a translation of the viewpoint according to the direction and distance.

[0287] Embodiment 20B. The method of any one of embodiments 12B to 19B, wherein the change from the first perspective to the second perspective includes a transformation, and the change from the third perspective to the fourth perspective includes a transformation.

[0288] Embodiment 21B. The method of any one of embodiments 12B to 20B, further comprising: identifying one or more gaps in the second image data based on one or more gaps in the second plurality of motion vectors; and modifying the second image data by at least partially filling the one or more gaps in the second image data using interpolation before outputting the second image data.

[0289] Embodiment 22B. The method of any one of embodiments 12B to 21B, further comprising: identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors; and modifying the second image data by at least partially filling the one or more gaps in the second image data using inpainting before outputting the second image data.

[0290] Aspect 23B. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform the operations recited in any one of aspects 1B to 22B.

[0291] Aspect 24B. An apparatus for image processing, comprising one or more means for performing the operations of any one of aspects 1B to 22B.

Claims

1. 1. An apparatus for image processing, comprising: at least one memory; at least one processor coupled to the at least one memory, the at least one processor comprising: receiving depth data including depth information corresponding to an environment; receiving first image data captured by an image sensor, the first image data including a representation of the environment; generating, for each pixel or group of pixels of the first image data based on at least the depth data, a motion vector grid including a first plurality of motion vectors indicative of a distance between the first image data and a change in viewpoint of the representation of the environment; using grid inversion performed on the motion vector grid to generate, for each pixel or group of pixels, an inverse motion vector grid comprising a second plurality of motion vectors indicative of a distance between the representation of the environment and the first image data with respect to the change in viewpoint; generating second image data by modifying the first image data at least in part according to the second plurality of motion vectors, the second image data including a second representation of the environment from a different perspective than the first image data; outputting the second image data; An apparatus configured to:

2. 2. The apparatus of claim 1, wherein the second image data includes an interpolated image configured to depict the environment at a second time between a first time and a third time, and the first image data includes at least one image depicting the environment at at least one of the first time or the third time.

3. 2. The apparatus of claim 1, wherein the first image data comprises a plurality of frames of video data that include parallax motion, and the second image data comprises a stabilized transformation of the plurality of frames of video data that reduces the parallax motion.

4. 2. The apparatus of claim 1, wherein the first image data includes a person viewing the image sensor from a first angle, and the second image data includes the person viewing the image sensor from a second angle different from the first angle.

5. The change in viewpoint is Rotation of the viewpoint around an axis according to an angle, or A translation of the viewpoint according to direction and distance, or conversion, or 2. The apparatus of claim 1, further comprising: a movement along an axis between an original viewpoint of the representation of the environment in the first image data and a position of an object in the environment, wherein at least a portion of the object is depicted in the first image data.

6. the at least one processor: identifying one or more gaps in the second image data based on one or more gaps in the second plurality of motion vectors; modifying the second image data at least in part by filling the one or more gaps in the second image data using interpolation before outputting the second image data; The apparatus of claim 1 configured to:

7. the at least one processor: identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors; modifying the second image data, at least in part, by filling the one or more gaps in the second image data using inpainting before outputting the second image data; The apparatus of claim 1 configured to:

8. the at least one processor: identifying one or more occlusion regions in the second image data based on one or more gaps in the second plurality of motion vectors; modifying the second image data at least in part by filling the one or more gaps in the second image data using inpainting using one or more trained machine learning models before outputting the second image data; The apparatus of claim 1 configured to:

9. the at least one processor: identifying a conflict in the second image data based on a plurality of motion vectors in the motion vector grid that point to a same pixel in the second image data, the plurality of motion vectors that point to the same pixel in the second image data being associated with a plurality of conflict values ​​from the first image data; selecting one of the plurality of competing values ​​from the first image data based on motion data associated with the second plurality of motion vectors; The apparatus of claim 1 configured to:

10. The device of claim 1 , wherein the depth information comprises a three-dimensional representation of an environment from a first viewpoint.

11. The device of claim 1 , wherein the depth data is received from at least one depth sensor.

12. a display, wherein the at least one processor is configured to display the second image data using at least the display to output the second image data; or a communications interface, wherein the at least one processor is configured to transmit at least the second image data to at least a receiving device using the communications interface to output the second image data. The apparatus of claim 1 further comprising:

13. The apparatus of claim 1 , wherein the apparatus comprises at least one of a head-mounted display (HMD), a mobile handset, or a wireless communication device.

14. 1. A method for image processing, comprising: receiving depth data including depth information corresponding to an environment; receiving first image data captured by an image sensor, the first image data including a representation of the environment; generating, for each pixel or group of pixels of the first image data based on at least the depth data, a motion vector grid comprising a first plurality of motion vectors indicative of a distance between the first image data and a change in viewpoint of the representation of the environment; using grid inversion performed on the motion vector grid to generate, for each pixel or group of pixels, an inverse motion vector grid comprising a second plurality of motion vectors indicative of a distance between the representation of the environment and the first image data with respect to the change in viewpoint; generating second image data by modifying the first image data at least in part according to the second plurality of motion vectors, the second image data including a second representation of the environment from a different perspective than the first image data; outputting the second image data; A method comprising:

15. A computer-readable storage medium having stored thereon instructions, the instructions, when executed by one or more processors, causing the one or more processors to: receiving depth data including depth information corresponding to an environment; receiving first image data captured by an image sensor, the first image data including a representation of the environment; generating, for each pixel or group of pixels of the first image data based on at least the depth data, a motion vector grid including a first plurality of motion vectors indicative of a distance between the first image data and a change in viewpoint of the representation of the environment; using grid inversion performed on the motion vector grid to generate, for each pixel or group of pixels, an inverse motion vector grid comprising a second plurality of motion vectors indicative of a distance between the representation of the environment and the first image data with respect to the change in viewpoint; generating second image data by modifying the first image data at least in part according to the second plurality of motion vectors, the second image data including a second representation of the environment from a different perspective than the first image data; outputting the second image data; A computer-readable storage medium that causes the