Single view image to stereoscopic image conversion and adjustment for viewing on spatial computer
Through the deformation process based on viewpoint and depth, combined with boundary adjustment and comfort parameters, the single-view image is successfully converted into a stereo image pair, solving the problem of poor image conversion effect in the prior art and achieving efficient and comfortable stereo image generation.
Patent Information
- Application Number
- CN202510060817.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-10
- Filing Date
- 2025-01-15
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art cannot effectively convert single-view images into stereo images to provide an efficient, desired viewing experience, especially in head-mounted devices, resulting in visual discomfort and insufficient immersion.
Using a deformation process based on viewpoint, depth and boundary adjustment, a single-view image is converted into a stereoscopic image pair. By generating a left-eye view and a right-eye view, image details are preserved using sparse depth information and a cutout network, and adjustments are combined with comfort parameter to meet the viewing preferences of different users.
It realizes the generation of clear, immersive and comfortable stereo pairs of images while maintaining image resolution and details, adapting to the visual comfort needs of different users and reducing visual discomfort.
Smart Images

Figure CN120390075A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to systems, methods, and devices for converting monoscopic image content into a stereoscopic image pair using a viewpoint, depth, or boundary adjustment process plus comfort parameter adjustment. Background Art
[0002] The prior art for viewing two-dimensional (2D) images is not sufficient to advantageously enhance such images to have an effect of improving the realism or other aspects of the image to provide an efficient, desired, and enhanced viewing experience. Summary of the Invention
[0003] Various embodiments disclosed herein include devices, systems, and methods for converting a monoscopic image into a stereoscopic image pair using a viewpoint-based warping process, a depth-based warping process, and / or a boundary adjustment process. The monoscopic image can be converted into a stereoscopic image pair in real time for display via (including but not limited to (inter alia)) a head-mounted device (HMD), etc.
[0004] In some embodiments, a viewpoint-based warping process is implemented to generate a right-eye view-based image and a left-eye view-based image from an input image associated with a central viewpoint. In some embodiments, the input image can correspond to the appearance of a scene relative to the central viewpoint of a user and can include any type of image, such as (including but not limited to) a photograph including appearance values (e.g., color values) at the pixel locations of the input image. In some embodiments, a depth image including depth values at the original pixel locations mapped to the pixel locations of the input image can be used to generate a left-eye output image and a right-eye output image relative to a left-eye viewpoint and a right-eye viewpoint different from the central viewpoint of the input image. In some embodiments, the left-eye output image and the right-eye output image can be used combinatorially to form a stereoscopic output image pair depicting a scene for viewing on a stereoscopic display of a head-mounted device (HMD).
[0005] In some embodiments, a depth-based warping process is implemented to maintain the resolution of the input image (e.g., high resolution, such as (including but not limited to) 20 megapixels or greater, etc.). The depth-based warping process can be configured to warp the input image to generate one viewpoint or two viewpoints. For example, a left-eye view can be generated from a right-eye view, and / or a right-eye view can be generated from a left-eye view. In some embodiments, it may be more preferable to generate both a right-eye view and a left-eye view from the central viewpoint.
[0006] In some specific implementations, the depth-based deformation process may utilize sparse depth information to deform the input image to a single viewpoint or multiple (e.g., two) viewpoints while maintaining the resolution of the input image. For example, the sparse depth information may include (including but not limited to) a low-resolution depth map that includes a lower resolution (e.g., 2 megapixels, etc.) than the resolution of the original input image (e.g., 20 megapixels, etc.). The sparse depth information is used to determine how to deform the coordinate image. For example, the sparse depth information may include a low-resolution image that provides a mapping of the pixel locations of the low-resolution image to the associated pixel locations in the high-resolution input image. Subsequently, the coordinate image may be upsampled. For example, the coordinate image may be upsampled by interpolating between pixels to identify the intermediate pixel mapping values of the intermediate pixels. The upsampled coordinate image may include the same or a similar resolution as the original input or image and may be used to extract red, green, and blue (RGB) values from the pixels of the original input image. Therefore, using the upsampled coordinate image as a mapping structure enables the details of the original input image to be retained in the output of the deformation process, such that the input image can be used as a lookup table for filling the output image.
[0007] In some specific implementations, a boundary adjustment process may be used to convert a single-view image into a stereoscopic image pair, and the boundary adjustment process may retain details that may otherwise appear to be blended in the output (stereoscopic) image from the regions of the input image. For example, the foreground portion of the image (e.g., a person's hair or facial features) may be alpha-blended with the background portion of the image (e.g., the walls in a room) instead of performing a process for blurring the foreground and background portions. In some specific implementations, the input image and the estimated depth may be used to classify the pixels within the boundary region between the local foreground and background regions. The local foreground and background regions may be extended and blended, for example, by using a matting network to determine blending weights, alpha blending values, etc. For example, in the local boundary region, the first part / pixels may include all local foreground regions, such as (including but not limited to) hair / hairs. Similarly, the second part / pixels may include all local background regions, such as (including but not limited to) walls. Therefore, a third intermediate transition region (i.e., the region between the foreground region and the background region) may be blended. For example, the hair / hairs in the local foreground portion may be presented in a partially transparent layer located above the top of the wall (in the background portion), and the wall is presented on the opaque layer below.
[0008] In some specific implementations, a comfort-based three-dimensional (3D) style preset can be implemented for adjusting comfort parameters for viewing stereoscopic 3D content (such as a stereoscopic image pair generated from 2D content) via a device (such as (including but not limited to) an HMD). Typical stereoscopic 3D image and video playback specific implementations may cause visual discomfort due to vergence accommodation conflict, and thus a comfort-based 3D style preset can be used to meet different content viewing preferences of different users to deliver different levels of stereoscopic visual comfort and different levels of immersion. For example, the comfort parameters adjusted via the 3D style preset can include (including but not limited to) a maximum parallax parameter, a parallax adjustment parameter, a motion parameter, a binocular rivalry parameter, a vertical parallax parameter, a bad image quality parameter, a low light parameter, a cardboard effect parameter, a puppet theater effect parameter, a color / brightness / saturation mismatch parameter, etc.
[0009] In some specific implementations, the maximum parallax parameter can be adjusted via parallax map adjustment. In some specific implementations, the parallax adjustment parameter can be adjusted relative to a target real-world parallax. In some specific implementations, the cardboard effect parameter can include a flat depth plane. In some specific implementations, the puppet theater effect parameter can be associated with unnatural object sizes and shapes.
[0010] In some specific implementations, an electronic device has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method performs one or more steps or processes. In some specific implementations, the electronic device obtains an input image including appearance values at pixel locations, the input image corresponding to the appearance of a scene from a first viewpoint. Some specific implementations determine a depth image including depth values at original pixel locations that map to at least a subset of the pixel locations in the input image. Coordinate mapping can be used to map the original pixel locations to corresponding pixel locations in the input image. Some specific implementations generate a first output image corresponding to a second viewpoint of the scene different from the first viewpoint. The first output image is generated by determining a first set of changed pixel locations for the depth values and identifying the appearance values of the first set of changed pixel locations based on the coordinate mapping and the input image. Some specific implementations generate a second output image corresponding to a third viewpoint of the scene different from the second viewpoint. The second output image is generated by determining a second set of changed pixel locations for the depth values and identifying the appearance values of the second set of changed pixel locations based on the coordinate mapping and the input image.
[0011] In some specific implementations, an electronic device has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method performs one or more steps or processes. In some specific implementations, the electronic device obtains an input image depicting a scene. The input image may include pixels and have a first resolution. In some specific implementations, a depth image may be determined. The depth image may correspond to a subset of the pixels of the input image from a first viewpoint. The depth image may have a second resolution that is less than the first resolution. In some specific implementations, a coordinate map may be generated for mapping positions in the depth image and positions in the input image. Some specific implementations may perform a first adjustment to the coordinate map to change the coordinate map to correspond to a second viewpoint different from the first viewpoint. Some specific implementations may perform a second adjustment to the coordinate map to increase the resolution of the coordinate map, and may provide an output image corresponding to a view of the scene from the second viewpoint. The output image may be provided based on the input image and the coordinate map.
[0012] In some specific implementations, an electronic device has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method performs one or more steps or processes. In some specific implementations, the electronic device obtains an input image depicting a scene from a first viewpoint. In response, an output image may be generated based on the input image. The output image may depict the scene from a second viewpoint different from the first viewpoint. In some specific implementations, a boundary region of the output image is identified based on depth information. The boundary region includes a first portion associated only with a relatively close part of the scene, a second portion associated only with a relatively distant part of the scene, and a third portion associated with both the relatively close part and the relatively distant part of the scene. In some specific implementations, extended foreground content may be generated by extending foreground content in the first portion into the third portion, and extended background content may be generated by extending background content in the second portion into the third portion. In some specific implementations, the boundary region of the output image may be updated by providing mixed content for the third portion using the extended foreground content and the extended background content.
[0013] In some embodiments, an electronic device has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method performs one or more steps or processes. In some embodiments, the electronic device obtains an image depicting two-dimensional (2D) content. In some embodiments, an adjustment of 3D tuning parameters is performed. The 3D tuning parameters may be associated with a 3D content viewing style. In some embodiments, the 3D tuning parameters are used to generate a 3D stereoscopic image pair corresponding to the image, and in response, a view of a 3D environment including the 3D stereoscopic image pair is presented.
[0014] According to some embodiments, a device includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, instructions are stored in a non-transitory computer-readable storage medium that, when executed by one or more processors of a device, cause the device to perform or cause to perform any of the methods described herein. According to some embodiments, a device includes: one or more processors, non-transitory memory, and components for performing or causing to perform any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Thus, the present disclosure may be understood by those of ordinary skill in the art, and a more detailed description may be made with reference to aspects of some illustrative embodiments, some of which are shown in the drawings.
[0016] Figure 1 An exemplary electronic device operating in a physical environment according to some embodiments is illustrated.
[0017] Figure 2 An example representing a view-point based warping process for converting a single-view image to a stereoscopic image pair according to some embodiments is illustrated.
[0018] Figure 3 A process for converting a single-view image to a stereoscopic image pair using a depth-based warping process according to some embodiments is illustrated.
[0019] Figure 4 A series of images associated with a warping process according to some embodiments is illustrated.
[0020] Figure 5 A local determination process associated with determining the foreground and background of an input image according to some embodiments is illustrated.
[0021] Figure 6 Illustrates a process associated with determining foreground and background classifications according to some specific implementations.
[0022] Figure 7A Illustrates a diagram showing the foreground and background of an image portion representing an image according to some specific implementations.
[0023] Figure 7B Illustrates according to some specific implementations with respect to Figure 7A A diagram modified from the illustrated diagram.
[0024] Figure 8 Illustrates a process associated with generating a tripartite graph according to some specific implementations.
[0025] Figure 9 Illustrates a process associated with generating a predicted alpha matte according to some specific implementations.
[0026] Figure 10 Is a flowchart representation of an exemplary method for dynamically converting a single-view image into a stereoscopic image pair by generating two views from an input image associated with a central viewpoint according to some specific implementations.
[0027] Figure 11 Is a flowchart representation of an exemplary method for dynamically converting a single-view image into a stereoscopic image pair using a depth-based deformation process according to some specific implementations.
[0028] Figure 12 Is a flowchart representation of an exemplary method for dynamically converting a single-view image into a stereoscopic image pair using a boundary adjustment process according to some specific implementations.
[0029] Figure 13 Is a workflow representation that enables modification of the disparity map of a stereoscopic image pair with respect to a maximum disparity parameter according to some specific implementations.
[0030] Figure 14 Is a flowchart representation of an exemplary method for implementing a comfort-based 3D style preset for adjusting comfort parameters for viewing stereoscopic 3D content via a device (such as an HMD) according to some specific implementations.
[0031] Figure 15 Is a block diagram of an electronic device according to some specific implementations.
[0032] According to common practice, the various features shown in the drawings may not be drawn to scale. Therefore, for clarity, the dimensions of the various feature portions may be arbitrarily enlarged or reduced. Additionally, some of the drawings may not depict all the components of a given system, method, or device. Finally, throughout the specification and the drawings, like reference numerals may be used to represent like feature portions. Detailed Implementation Modes
[0033] Numerous details are described to provide a thorough understanding of the exemplary specific implementations shown in the accompanying drawings. However, the accompanying drawings only show some exemplary aspects of the present disclosure and should not be considered restrictive. Those of ordinary skill in the art will understand that other valid aspects and / or variations do not include all the specific details described herein. In addition, well-known systems, methods, components, devices, and circuits have not been described in detail so as not to obscure more relevant aspects of the exemplary specific implementations described herein.
[0034] Figure 1 An exemplary electronic device 105 operating in a physical environment 100 is shown. In Figure 1 the example, the physical environment 100 is a room. The electronic device 105 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about the physical environment 100 and the objects therein, as well as information about the user 102 of the electronic device 105. Information about the physical environment 100 and / or the user 102 can be used to provide visual and audio content, and / or identify the current location of the physical environment 100 and / or the location of the user within the physical environment 100.
[0035] In some specific implementations, a view of an extended reality (XR) environment can be provided to one or more participants (e.g., the user 102 and / or other participants not shown) via the electronic device 105 (e.g., a wearable device such as an HMD). Such an XR environment can include a view of a 3D environment generated based on camera images and / or depth camera images of the physical environment 100 and a representation of the user 102 based on camera images and / or depth camera images of the user 102. Such an XR environment can include virtual content positioned at 3D positions relative to a 3D coordinate system associated with the XR environment, and the 3D coordinate system can correspond to the 3D coordinate system of the physical environment 100.
[0036] In some specific implementations, an HMD (e.g., device 105), a server communicatively coupled thereto, or other external devices can be configured to convert a single-view image (e.g., a photo in a photo library, a frame of a video, etc.) into a stereoscopic image pair (e.g., in real time) viewable via a head-mounted device (such as (including but not limited to) an HMD). A viewpoint-based warping process, a depth-based warping process, and / or a boundary adjustment process can be used to convert the single-view image into a stereoscopic image pair.
[0037] The view-point based warping process may include: obtaining an input image, and generating a left-eye view image and a right-eye view image from the input image associated with a center view-point of a user. The input image may correspond to an appearance of a scene relative to the center view-point of the user, and may include any type of image, such as (including but not limited to) a photograph including color values at pixel locations of the input image. In some embodiments, a depth image may be determined. The depth image may include depth values at original pixel locations of a subset of pixel locations of the input image mapped to the pixel locations of the input image. The depth image may be used to generate a left-eye output image and a right-eye output image relative to eye view-points (e.g., left and right eye view-points) different from the center view-point of the input image. In some embodiments, the left-eye output image and the right-eye output image combinatorially form a stereoscopic output image pair depicting the scene for viewing on a stereoscopic display of an HMD.
[0038] A depth-based warping process may be implemented to maintain the resolution of the input image (e.g., the original single-view image) relative to the output image (e.g., the resulting stereoscopic image pair). A depth-based warping process may be implemented to warp the input image to produce view-points (e.g., one view-point, two view-points, etc.). For example, a left-eye view may be generated from a right-eye view, and / or a right-eye view may be generated from a left-eye view.
[0039] In some embodiments, a depth-based warping process may utilize depth information to warp the input image to at least one view-point. For example, the depth information may include (including but not limited to) a low-resolution (e.g., 2 megapixels, etc.) depth map that includes a resolution smaller than the resolution of the original input image (e.g., 20 megapixels, etc.). The depth information is used to determine a process for warping a coordinate image. For example, the depth information may include a low-resolution image that provides a mapping of pixel locations of the low-resolution image relative to associated pixel locations in the high-resolution input image. Subsequently, the coordinate image is upsampled by interpolating between pixels to identify intermediate pixel mapping values of intermediate pixels. The upsampled coordinate image may include a resolution similar to the original input image, and may be used to extract RGB values from pixels of the original input image. Utilizing the upsampled coordinate image as a mapping structure enables details of the original input image to be retained within the output of the warping process, such that the input image can be used as a look-up table for filling outputs for the output image.
[0040] In some specific implementations, a boundary adjustment process can be used to convert a single-view image into a stereoscopic image pair. This boundary adjustment process can preserve details from regions of the input image that might otherwise appear blended in the output (stereoscopic) image. For example, the foreground portion of an image (e.g., the fur of an animal such as a dog) can be alpha-blended with the background portion of the image (e.g., a wall), rather than blurring the background and foreground portions. In some specific implementations, the input image and the estimated depth can be used to classify pixels within the boundary region between local foreground and background regions. The local foreground and background regions can be extended and blended, for example, by using a matting network to determine blending weights, alpha blending values, etc. For example, in a local boundary region, a first portion / pixel can include all local foreground portions, such as (including but not limited to) hair / fur. Similarly, a second portion / pixel can include all local background portions, such as (including but not limited to) walls, ceilings, etc. Thus, a third / intermediate portion can be blended. For example, the hair / fur in the local foreground portion can be presented in a partially transparent layer located above the top of a wall (in the background portion), and the wall is presented on an opaque layer below.
[0041] In some specific implementations, stereoscopic 3D image and video playback may cause visual discomfort due to vergence-accommodation conflicts, and thus 3D tuning parameter adjustments can be performed relative to an image or video (e.g., frames of a video) to meet the different content viewing preferences of different users, such that different levels of stereoscopic visual comfort and associated levels of immersion can be achieved for different users. For example, an image depicting 2D content (e.g., a photo or video) can be obtained (e.g., via an HMD), and adjustments to 3D tuning parameters associated with a 3D content viewing style can be performed such that the 3D tuning parameters are used to generate a 3D stereoscopic image pair corresponding to the image, thereby enabling a view of a 3D environment including the 3D stereoscopic image pair (i.e., a customized version) to be presented to the user.
[0042] In some specific implementations, the adjustment of 3D tuning parameters can include modifying a disparity map based on a maximum disparity parameter. The modified disparity map can be used to perform adjustments to control the amount of perceived depth within the view.
[0043] In some specific implementations, the adjustment of 3D tuning parameters can include performing disparity adjustment by modifying the disparity map to match a target real-world disparity. Disparity adjustment can be performed when the maximum disparity parameter exceeds a threshold level.
[0044] In some specific implementations, the adjustment of 3D tuning parameters can include modifying the scene depth representation format, which is different from disparity modification.
[0045] In some specific implementations, the adjustment of 3D tuning parameters may include: activating preset 3D tuning parameters, variably adjusting 3D tuning parameters, and so on.
[0046] In some specific implementations, the adjustment of 3D tuning parameters may be enabled in response to user input.
[0047] In some specific implementations, the adjustment of 3D tuning parameters may include modifying motion parameters within a view.
[0048] Figure 2 An example of a view - point - based deformation process 200 according to some specific implementations is illustrated. The view - point - based deformation process converts a single - view image into a stereoscopic image pair by generating a left - eye view (output) image 202b and a right - eye view (output) image 202c from a (single - view) input image 202 associated with a central view - point 203 of a device 205 of a user 201 relative to a display input image 202. The input image 202 may include (including but not limited to) a 2D photograph (e.g., from a photo library) or a screenshot (e.g., from a video game) representing the appearance of a scene including a person 208 in the foreground and a mountain 204 in the background. The input image 202 may include appearance values, such as color values located at pixel positions.
[0049] The view - point - based deformation process 200 may include determining a depth image 202a (e.g., a low - resolution three - dimensional (3D) model illustrating the person 208 in the foreground and the mountain 204 in the background), which includes depth values at original pixel positions mapped to a subset of pixel positions of the input image 202. The depth image 202a includes a coordinate mapping for mapping the original pixel positions to corresponding pixel positions in the input image 202.
[0050] The left - eye view image 202b corresponds to the left - eye view - point of the scene relative to the input image 202, and may be generated by: determining a first set of changed pixel positions for the (left - eye - view - point) depth values, and identifying the appearance (e.g., color) values of the first set of changed pixel positions based on the (depth - image 202a's) coordinate mapping and the input image 202. The left - eye view image 202b represents a deformed view 208b of the person 208 located at a first position, which is different from the original position 207 of the person 208 in the original input image 202 (e.g., horizontally offset in direction 212a).
[0051] The right-eye view image 202c corresponds to the right-eye view point of the scene relative to the input image 202, and can be generated by: determining a second set of modified pixel locations for depth values (e.g., for the right-eye view point), and identifying the appearance (e.g., color) values of the second set of modified pixel locations based on the coordinate mapping (of the depth image 202a) and the input image 202. The right-eye view image 202c represents a deformed view 208c of the person 208 located at the second location, which is different from the original location 207 of the user 208 in the original input image 202 (e.g., horizontally offset in the direction 212b). The first location represents the user 208 at a position in the left-eye image version 202a that is different from the second location in the right-eye image version 202b.
[0052] Thus, when viewed via the HMD, the combination of the left-eye image version 202a and the right-eye image version 202b forms a stereoscopic output image pair 218 depicting the scene for viewing on the stereoscopic display of the device 205 (e.g., HMD).
[0053] Figure 3 Illustrated is a process 300 for using a depth-based warping process to convert a single-view image into a stereoscopic image pair according to some specific implementations. The depth-based warping process preserves details such as the resolution of the original input image 302. Process 300 includes warping the image content based on the depth map 318, and subsequently using coordinate adjustment during an upsampling process to maintain the high-frequency details (resolution) from the original image 302. For example, process 300 may include a view-point based warping process that converts a single-view image (e.g., the original input image 302) into a stereoscopic image pair (output image 320) by generating a left-eye view image 320a and a right-eye view image 320b associated with the view point of the user relative to the device displaying the original input image.
[0054] Process 300 is initiated in response to performing a downsampling process 304 on the original input image 302 (e.g., a high-resolution image such as, for example, a 20 megapixel (MP) image) to generate a downsampled input image 306 (e.g., including a low resolution such as, for example, 2MP, etc.). The downsampling process 304 can be performed such that a low-resolution image (e.g., the input image 306) is generated to provide a mapping of the pixel locations in the low-resolution image to the locations in the high-resolution input image (e.g., the original input image 302), such that when the low-resolution image is upsampled (e.g., by interpolating between pixels to identify intermediate pixel mapping values), the upsampled image can be used to obtain RGB values from the pixels of the original input image so that the details of the input image can be retained in the output of process 300.
[0055] Subsequently, the deep network 316 is enabled to predict / generate a low-resolution depth map 318, such as, for example, 2MP. The depth map 318 is used to determine how to warp the coordinate image 308 associated with the downsampled input image 306 (via the forward warping module 310). The coordinate image 308 includes a low-resolution image that provides a mapping of pixel locations in the low-resolution image to pixel locations in the original high-resolution input image 302. Subsequently, the coordinate image 308 is warped (via the warping module 310) into a new perspective coordinate image 314 (e.g., including a left-eye view image and a right-eye view image). This warping process includes transforming the location of each pixel (of the coordinate image 308) based on the depth information of the depth map 318 to create the new perspective coordinate image 314.
[0056] The new perspective coordinate image 314 is then upsampled (via the upsampling process 315) by interpolating values of intermediate pixels between adjacent pixels to identify intermediate pixel mapping values, thereby producing an upsampled coordinate image 317 having the same resolution as the original input image 302. The upsampled coordinate image 317 is used to obtain RGB values from the pixels of the original input image 302 (via the backward warping module 319) to populate the output image 320, which may include a left-eye view image 320a from a first perspective and a right-eye view image 320b from a second different perspective. Thus, using the upsampled coordinate image as a map enables the details of the original input image 302 to be retained in the output of the warping process, such that the original input image 302 can be used as a look-up table for populating the output image 320.
[0057] Figure 4 Illustrated is a series of images 400 associated with a warping process (such as a depth-based warping process as described with respect to Figure 3 The images 400 include a first image 404 of a person 401, a second image 406 of the person 401, and a third image 408 of the person 401. The first image 404 represents the original input image for single-view image to stereoscopic image processing via the warping process. The second image 406 represents the result of using the upsampled coordinate image as a map to retain details (such as the resolution of the first image 404) via a depth-based warping process (described above with respect to Figure 3The output image (e.g., a stereoscopic image pair) generated from the first image 404 (e.g., from different viewpoints) as illustrated and described. Thus, the aforementioned single-view image to stereoscopic image processing generates a second image 406 including high resolution for providing a true and accurate representation of the person 401 without generating an untrue, smooth, and blurred representation of the face of the person 401 as illustrated with respect to image 308. For example, the third image 408 represents an output image generated from the first image 404 via a warping process that includes generating a downsampled image, warping the downsampled image, and performing an upsampling process using conventional upsampling techniques, resulting in the loss of some details of the original input image (i.e., the first image 404) (e.g., hair 410, eyes 412a and 412b, and nose 414, etc.). For example, image 408 includes an untrue, smooth, and blurred representation of the face of the person 401.
[0058] Figure 5 Illustrates a local determination process 500 according to some embodiments associated with determining the foreground and background of an input image 502 to convert a single-view image to a stereoscopic image pair as described with respect to Figure 4 The local determination process 500 enables the determination of the classification of the local foreground and background from the input image 502 and the estimated depth map 504. The estimated depth map 504 is used to segment the input image 502 into an opaque foreground layer 506 for rendering and a transparent foreground 508 and an opaque background 510 for placement on top of the opaque foreground 506. For example, the input image 502 uses the estimated depth map 504 to classify the pixels in the boundary regions (e.g., at the hairline, on the face, etc.) as local foreground (e.g., hair / fur) and background (e.g., environment) regions. The local foreground and background regions can be extended and blended, for example, by using a matting network (e.g., the matting network 610 as described below with respect to Figure 6 to determine blending weights, alpha blending values, etc. For example, in a local boundary region, a first portion / pixels can be all local foreground (e.g., hair / fur), and another portion / pixels can be all local background (e.g., a wall), and the background and foreground can be rendered separately such that they can move independently of each other to enable the generation of a clear, focused, and realistic single-view image to stereoscopic image pair for presentation to the user.
[0059] Figure 6 Illustrates according to some embodiments associated with determining in relation to as described with respect to Figure 5Process 600 associated with the foreground / background classification associated with the described local determination process 500. Process 600 is associated with rendering transparent regions that are partially visible in both the foreground and background of an (image) to avoid visual artifacts in the generated stereoscopic image pair for presentation to the user. Thus, the process includes classifying the boundaries, foreground regions, and background regions of the input image 602 using the depth map 604 (via module 606). Subsequently, two layers of the image 602 are locally generated by using the trimap 608 (i.e., a three-channel image / map representing the absolute background, foreground, and unknown regions of the input image 602) and a matting network 610 (e.g., a pre-trained model) to create a soft alpha matte for the boundary regions that will be used to blend the foreground and background with each other to generate a predicted alpha matte 612, which represents the transparency level of each associated pixel and indicates whether each associated pixel belongs to the foreground or background of the input image 602. The matting network 610 is configured to accurately estimate the transparency of pixels to allow for smooth blending between the foreground region and the background region, thereby reducing visual artifacts that may occur at hard cutoffs, and thus providing a visually accurate and appealing stereoscopic image pair for presentation to the user.
[0060] Figure 7A Illustrates a diagram 700a of the foreground 707 and background 706 of an image portion 702a of an image 702 (such as (including but not limited to) a photograph as described with respect to Figure 5 and Figure 6 ). The foreground 707 includes the fur of a dog, and the background 706 includes the portion of the image that is behind the fur. Diagram 700 illustrates a depth prediction representation 708 that represents the average between the foreground representation 705 of (the foreground 707) and the background representation 704 of (the background 706). The depth prediction representation 708 is used to classify between the foreground representation 705 and the background representation 704, but may predict an incorrect boundary 712 between the foreground representation 705 of (the foreground 707) and the background representation 704 of (the background 706). Thus, the predicted incorrect boundary 712 can be removed as described below with respect to Figure 7B .
[0061] Figure 7B Illustrates a diagram 700b modified from diagram 700a according to some specific implementations with respect to Figure 7A . Diagram 700b illustrates the foreground representation 705 of (the foreground 707) and the background representation 704 of (the background 706), where ( Figure 7AThe depth prediction representation 708 is removed, such that there is a missing region 712 between the foreground representation 705 and the background representation 704. Similarly, illustration 700b shows a foreground representation 705 extended (via an extended foreground representation portion 718) and a background representation 704 extended (via an extended background representation portion 715), such that the foreground representation 705 and the background representation 704 currently overlap with the missing content therebetween. Subsequently, a matting network can be implemented to determine which portion (of image portion 702a) belongs to the foreground 707 and which portion belongs to the background 706, to create a true and accurate stereoscopic image pair for presentation to the user.
[0062] Figure 8 Illustrates a process 800 according to some embodiments associated with generating a tri-map 808 to identify boundary regions (e.g., local regions of 1+ pixels) of an output image (such as, Figure 3 output image 320) including a stereoscopic image pair based on depth information. Process 800 receives an input image 802 and associated depth 804, and applies a blurred disparity image operator 806 to generate a tri-map 808, which is used to classify the input image 802 into a foreground portion and a background portion for input into a matting network (e.g., Figure 6 matting network 610) to predict a fine boundary. The tri-map 808 represents a local foreground region, a local background region, and a region between the local foreground region and the local background region, for updating a boundary region of the output image by providing blended content for the intermediate region using the foreground content and the background content. (e.g., using a combination of multiple layers, opaque / transparent features in multiple layers, alpha values, etc.). For example, tri-map portion 808a represents a magnified view of a portion of the tri-map 808, and illustrates a local foreground region 814 and a local background region 816, where region 810 is a transition between the local foreground region 814 and the local background region 816. Thus, the matting network performs a query to determine a soft boundary between the local foreground region 814 and the local background region 816, thereby implementing a process for creating a stereoscopic image pair for presentation to the user.
[0063] Figure 9 Illustrates a process 900 according to some embodiments associated with generating a predicted alpha matte 908 (such as, Figure 6 matting network 610). Process 900 obtains an input image 902 (which includes a foreground image portion 903 placed above a background image portion 904), and generates / utilizes a tri-map 906 (e.g., Figure 8A trimap 808 is used to generate a predicted alpha matte 908 (i.e., a grayscale image where the luminance of each pixel represents its transparency): An image segmentation algorithm or a deep learning-based segmentation model is applied and the trimap 906 is used to classify pixels into foreground, background, and unknown regions. The alpha matte 908 includes an image with an additional channel (i.e., the alpha channel) representing the transparency or opacity of each pixel. The alpha channel is configured to define how much of the associated pixel is opaque (visible) and how much of the associated pixel is transparent, and can indicate whether the associated pixel belongs to the foreground or the background, thereby determining which parts of the input image 902 should be visible and which parts of the input image 902 should be transparent in order to create a stereoscopic image pair for viewing.
[0064] Figure 10 is a flowchart representation of an exemplary method 1000 for dynamically converting a single-view image into a stereoscopic image pair by generating two views from an input image associated with a central viewpoint according to some specific implementations. In some specific implementations, the method 1000 is executed by a device (such as a mobile device, a desktop computer, a laptop computer, an HMD, or a server device). In some specific implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images, such as an HMD (an HMD, such as for example Figure 1 device 105). In some specific implementations, the method 1000 is executed by processing logic (including hardware, firmware, software, or a combination thereof). In some specific implementations, the method 1000 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., a memory). Each block in the method 1000 can be enabled and executed in any order.
[0065] At block 1002, the method 1000 obtains an input image (e.g., a photo) including appearance (e.g., color) values at pixel locations. The input image corresponds to the appearance of a scene from a first viewpoint (such as (including but not limited to) the central viewpoint 203 of the user 201 relative to the device 205 displaying the input image 202, as described above with respect to Figure 2 ). The input image can include (including but not limited to) photos, etc. The appearance values can include (including but not limited to) color values, etc.
[0066] At block 1004, the method 1000 determines a depth image including depth values at original pixel locations that map to at least a subset of the pixel locations of the input image, such that the coordinate mapping maps the original pixel locations to the input image (such as, as described above with respect to Figure 2The corresponding pixel localization in the described depth image 202a). In some embodiments, the depth image may be generated based on evaluating an input image using a neural network. In some embodiments, the coordinate mapping may be a coordinate image. In some embodiments, the depth image may be generated by using predefined rules and / or algorithms according to a rule-based / deterministic approach to manipulate depth data from the image. The rule-based / deterministic approach may include techniques such as (including but not limited to) depth thresholding, edge detection, histogram analysis, depth filtering, and the like.
[0067] At block 1006, method 1000 generates a first output image corresponding to a second view point of the scene that is different from the first view point. The first output image may be generated by determining a first set of modified pixel localizations for the depth values and identifying appearance values for the first set of modified pixel localizations based on the coordinate mapping and the input image. For example, the first output image may be a left-eye output image, such as the left-eye view (output) image 202b described above with respect to Figure 2 The left-eye view (output) image 202b described above.
[0068] In some embodiments, generating the first output image may include identifying appearance values for additional pixel localizations in addition to a second set of modified pixel localizations based on the coordinate mapping and the input image. In some embodiments, identifying appearance values for the additional pixel localizations may include: identifying an intermediate pixel localization between two adjacent pixel localizations in the second set of modified pixel localizations; and identifying the appearance value of the intermediate pixel localization by: according to the coordinate mapping, identifying the pixel localization in the input image between the pixel localizations in the input image corresponding to the two adjacent pixel localizations.
[0069] At block 1008, method 1000 generates a second output image corresponding to a third view point of the scene that is different from the second view point. The second output image may be generated by determining a second set of modified pixel localizations for the depth values and identifying appearance values for the second set of modified pixel localizations based on the coordinate mapping and the input image. For example, the first output image may be a right-eye output image, such as the right-eye view (output) image 202c described above with respect to Figure 2 The right-eye view (output) image 202c described above.
[0070] In some embodiments, the first view point may correspond to a center view point, the second view point may correspond to a left-eye view point, and the third view point may correspond to a right-eye view point, as with respect to Figure 2As described. In some specific implementations, the first output image can be a left-eye image generated based on the input image and the first coordinate image, where the first coordinate image (a) is determined based on the depth image and (b) is distorted for the left-eye view point. In some specific implementations, the second output image can be a right-eye image generated based on the input image and the second coordinate image, where the second coordinate image (a) is determined based on the depth image and (b) is distorted for the right-eye view point. Generating two different views (i.e., the right-eye view point and the left-eye view point) enables the 2D image to be viewed in 3D. In some specific implementations, using the central input image to create two view-point images (for viewing in 3D) can enable the process to perform fewer adjustments relative to the view point, thereby creating smaller holes to be filled or hidden using alpha blending. Similarly, using the central input image to create two view-point images enables the distance of the view point relative to the object to be accurately determined.
[0071] Some specific implementations also provide the first output image and the second output image to form a stereoscopic output image pair depicting a scene for viewing on the stereoscopic display of the HMD. For example, as described with respect to Figure 2 the output image pair 218.
[0072] Figure 11 is a flowchart representation of an exemplary method 1100 for dynamically converting a single-view image into a stereoscopic image pair using a depth-based distortion process according to some specific implementations. In some specific implementations, method 1100 is executed by a device such as a mobile device, a desktop computer, a laptop computer, an HMD, or a server device. In some specific implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images, such as an HMD (such as, for example, Figure 1 device 105). In some specific implementations, method 1100 is executed by processing logic components including hardware, firmware, software, or a combination thereof. In some specific implementations, method 1100 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., a memory). Each block in method 1100 can be enabled and executed in any order.
[0073] At block 1102, method 1100 obtains an input image depicting a scene (e.g., a photograph such as the original input image 302 as described with respect to Figure 3 ). The input image can include pixels and have a first resolution.
[0074] At block 1104, method 1100 determines a depth image corresponding to a subset of the pixels of the input image from the first view point (e.g., as described with respect to Figure 3The described depth map 318). The depth image may have a second resolution less than the first resolution. The depth image may be generated based on evaluating an input image using a neural network. Alternatively, the depth image may be generated according to a rule-based / deterministic approach using predefined rules and / or algorithms to manipulate depth data from the image. The rule-based / deterministic approach may include techniques such as (including but not limited to) depth thresholding, edge detection, histogram analysis, depth filtering, etc.
[0075] At block 1106, method 1100 generates a coordinate map (e.g., a coordinate image such as with respect to Figure 3 the described coordinate image 308) that maps locations in the depth image and locations in the input image.
[0076] At block 1108, method 1100 performs a first adjustment to the coordinate map (e.g., via a warping process performed by a forward warping module 310 such as with respect to Figure 3 the described one) to change the coordinate map to correspond to a second viewpoint different from the first viewpoint. The first adjustment may include warping the coordinate image and may be determined according to disparity information determined based on the depth image. In some embodiments, during the first adjustment, the input image (i.e., the low-resolution image) is warped to one or two viewpoints, such as, for example, generating a left-eye view from a right-eye view, generating a right-eye view from a left-eye view, generating both eye views from a center viewpoint, etc. In some embodiments, generating the left-eye view and the right-eye view may be performed continuously in sequence. In some embodiments, generating the left-eye view and the right-eye view may be performed simultaneously in parallel.
[0077] At block 1110, method 1100 performs a second adjustment to the coordinate map (e.g., upsampling, such as with respect to Figure 3 the described upsampling process 319) to increase the resolution of the coordinate map. The second adjustment may include upsampling the coordinate map. Additionally or alternatively, the second adjustment may include upsampling the coordinate map from a second resolution to a first resolution. In some embodiments, upsampling may include interpolating between pixel locations for intermediate pixels of the coordinate map.
[0078] At block 1112, method 1100 provides an output image (e.g., such as with respect to Figure 3 the described output image 317) corresponding to a view of the scene from the second viewpoint. The output image may be provided based on the input image and the coordinate map.
[0079] In some embodiments, the input image and the output image together provide a stereoscopic image pair depicting a scene. In some embodiments, the input image may correspond to a center view point, and the output image may correspond to a left-eye image or a right-eye image of a stereoscopic image pair depicting the scene. In some embodiments, the output images may be generated continuously in sequence. In some embodiments, the output images may be generated simultaneously in parallel.
[0080] In some embodiments, a left-eye image may be generated based on the input image and a first coordinate image that (a) is determined based on a depth image, (b) is distorted for a left-eye view point; and (c) has an increased resolution. Similarly, a right-eye image may be generated based on the input image and a second coordinate image that (a) is determined based on a depth image, (b) is distorted for a right-eye view point; and (c) has an increased resolution as described with respect to Figure 3 what is described.
[0081] In some embodiments, providing the output image may include using a pixel value of the input image at a pixel position mapped in the output image based on coordinates. In some embodiments, the output image may be provided as part of a stereoscopic image pair depicting a scene for viewing on a stereoscopic display of an HMD.
[0082] Figure 12 is a flowchart representation of an exemplary method 1200 for dynamically converting a single-view image into a stereoscopic image pair using a boundary adjustment process according to some embodiments. In some embodiments, method 1200 is performed by a device such as a mobile device, a desktop computer, a laptop computer, an HMD, or a server device. In some embodiments, the device has a screen for displaying images and / or a screen for viewing stereoscopic images, such as an HMD (an HMD, such as, for example Figure 1 device 105). In some embodiments, method 1200 is performed by processing logic including hardware, firmware, software, or a combination thereof. In some embodiments, method 1200 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). Each block in method 1200 may be enabled and executed in any order.
[0083] At block 1202, method 1200 obtains an input image (e.g., image 702 as described with respect to Figure 2 and Figure 7A and Figure 7B what is described) depicting a scene from a first view point (such as, for example, the center view point 203 as described with respect to
[0084] At block 1204, method 1200 generates an output image based on the input image. The output image may depict the scene from a second viewpoint different from the first viewpoint.
[0085] At block 1206, method 1200 identifies a boundary region of the output image based on depth information such as the depth prediction representation 708 described relative to Figure 7A The boundary region may include a first part associated only with a relatively proximal portion of the scene (e.g., the foreground 707 as described relative to Figure 7A ), a second part associated only with a relatively distal portion of the scene (e.g., the background 706 as described relative to Figure 7A ), and a third part associated with both the relatively proximal portion and the relatively distal portion of the scene (e.g., the trimap 808 as described relative to Figure 8 ).
[0086] At block 1208, method 1200 generates extended foreground content by extending foreground content in the first part into the third part. For example, the extended foreground representation part 718 as corresponding to Figure 7B described.
[0087] At block 1210, method 1200 generates extended background content by extending background content in the second part into the third part. For example, the extended background representation part 715 as described relative to Figure 7B described.
[0088] At block 1212, method 1200 updates the boundary region of the output image by providing blended content for the third part using the extended foreground content and the extended background content. For example, updating the boundary region may include (including but not limited to) using combinations of multiple layers, opaque / transparent features in multiple layers, alpha values, etc. The blending process may utilize a matting neural network. For example, the matting network 610 as described relative to Figure 6 described.
[0089] In some embodiments, the input image and the output image together provide a stereoscopic image pair depicting a scene. For example, the stereoscopic output image pair 218 as described relative to Figure 2 described. In some embodiments, the input image corresponds to a central viewpoint, and the output image corresponds to a left-eye image or a right-eye image of a stereoscopic image pair depicting the scene as described relative to Figure 2 described. Similarly, the left-eye image may be generated by deforming the input image for the left-eye viewpoint, and the right-eye image may be generated by deforming the input image for the right-eye viewpoint. In some embodiments, providing the output image includes providing a stereoscopic image pair depicting a scene for viewing on a stereoscopic display of an HMD.
[0090] In some specific implementations, the output image can be updated in the following ways: different foreground content thresholds and background content thresholds are used for the boundary regions to update multiple boundary regions. In some specific implementations, updating the boundary regions can include using extended foreground content and extended background content in different display layers. In some specific implementations, updating the boundary regions can include using extended foreground content in a first display layer, and this first display layer is displayed on top of a second display that displays the extended background content. In some specific implementations, providing the mixed content can include displaying the extended foreground content on a partially transparent layer. In some specific implementations, providing the mixed content can include displaying the extended foreground content with an alpha value to be mixed with the extended background content. In some specific implementations, providing the mixed content can include inputting the extended foreground content and the extended background content into a matting network.
[0091] Figure 13 is a workflow representation 1300 that enables modification of the disparity map of a stereoscopic image pair relative to a maximum disparity parameter according to some specific implementations.
[0092] In some cases, stereoscopic 3D image and video playback may cause visual discomfort (for users of, for example, HMDs) due to vergence-accommodation conflicts, such as (including but not limited to) in a scene (e.g., near the viewpoint) within a view generated via, for example, an HMD where there are objects with excessive disparity or negative disparity effects, which may cause focusing difficulties during viewing and may result in excessive retinal disparity and visual discomfort. Similarly, the same 3D image or video rendered stereoscopically may cause different levels of discomfort for different users viewing the 3D content. In addition, for 3D content associated with large disparities or disparity effects, some users may prefer a higher level of immersion and depth experience. Therefore, a comfort-based 3D style preset for stereoscopic 3D content playback and rendering can be implemented to address the aforementioned visual discomfort problem and the problem of different levels of immersion for different user preferences. The comfort-based 3D style preset can be configured to meet the different 3D content viewing preferences of different users to deliver different levels of stereoscopic visual comfort and different levels of immersion.
[0093] In some specific implementations, 3D style presets can be implemented, which involve (including but not limited to) high comfort level presets, medium comfort level presets, and low comfort level presets. For example, the high comfort level preset can provide a conservative tuning property for the parallax parameters of 3D space images and videos to provide a comfortable setting for users (viewers) who may be sensitive to high parallax effects and depth attributes. Similarly, the low comfort level preset can provide a loose tuning of the parallax parameters of 3D space images and videos to provide a comfortable setting for users (viewers) who may prefer to be exposed to high parallax effects and depth. Additionally, the medium comfort level preset can provide a fine tuning of the parallax parameters to achieve an intermediate level of stereoscopic visual comfort and an intermediate level of immersion, thus providing a preset level between the high comfort level preset and the low comfort level preset.
[0094] In some specific implementations, 3D style presets can be managed by selecting two parallax parameters to achieve different levels of stereoscopic visual comfort and different levels of immersion. For example, the first parallax parameter can include a maximum parallax parameter, and the second parallax parameter can include a parallax adjustment parameter.
[0095] The maximum parallax parameter (e.g., the maximum negative parallax effect) represents the maximum near-field depth that will be perceived by a user viewing a stereoscopic 3D image or video. Similarly, the maximum parallax parameter can be defined as the maximum amount of negative horizontal parallax that exists between the left-eye view and the right-eye view of a 3D space image or video (e.g., across the entire duration of the video).
[0096] In some specific implementations, a parallax map (as illustrated in block 1302) can be used for a pair of synthetic stereoscopic images, and the maximum value of the parallax map can be constrained by the maximum parallax parameter, thereby enabling the control of the perceived depth generated by the stereoscopic playback of a 3D space image or video to match the target comfort level for the associated 3D style preset.
[0097] In some specific implementations, the workflow representation 1300 represents a process for enabling modification (e.g., scaling) of a disparity map (M) associated with a stereo image pair (as illustrated in block 1302) to generate a modified disparity map (M') (as illustrated in block 1306) when the maximum map disparity (Max M) is greater than the maximum disparity parameter (max_disparity) (as illustrated in block 1304). Similarly, when the maximum map disparity (Max M) is not greater than the maximum disparity parameter (max_disparity), then it can be determined that the disparity map (M) is equivalent to the modified disparity map (M'), as illustrated in block 1308. In some specific implementations, the modified disparity map (M') can be determined via the following equation illustrated in block 1306: M' = M * (max_disparity / max(M)). Subsequently, a stereo image pair can be synthesized from the modified disparity map (M') to control the amount of depth perceived by a viewer.
[0098] In some specific implementations, the maximum disparity parameter (max_disparity) can be applied to the disparity map (M) corresponding to a stereo composite image pair having a specified reference resolution. Similarly, comfort tuning of the disparity preset for a defined 3D style can be performed relative to a set of target real-world disparities (i.e., one target real-world disparity is preset for each defined 3D style), since the perceived depth can be determined by the real-world disparity associated with the comfort tuning preset for the defined 3D style.
[0099] In some specific implementations, the relationship between the real-world disparity and the maximum disparity for a given reference horizontal resolution can be as follows:
[0100] real_world_disparity = viewing_distance * 2 * tan(\frac{hFOV}{2}) * \frac
[0101] {max_disparity}{ref_resolution}
[0102] In the foregoing relationship, hFOV is the horizontal field of view occupied by the rendered stereo image, viewing_distance represents the distance between the viewer and the screen (e.g., of an HMD), max_disparity represents the maximum disparity for a given reference horizontal resolution, and ref_resolution is the reference horizontal resolution of the composite image.
[0103] In some specific implementations, the maximum disparity parameter (max_disparity) can be the maximum allowable disparity that is generally set for all types of stereoscopic 3D images or videos. In some specific implementations, as a statistic from the reference disparity map, the maximum disparity parameter (max_disparity) can be adaptive relative to each piece of material. For example, the maximum disparity parameter (max_disparity) can be selected as the maximum value of the reference disparity map of a piece of material. Alternatively, the maximum disparity parameter (max_disparity) can be selected based on a given percentile distribution from the reference disparity map.
[0104] In some specific implementations, the stereoscopic playback and rendering of 3D spatial images and videos can be associated with different viewing configurations and different screen sizes (e.g., with respect to width and height) and distances. Changes in the viewing configurations with different screen distances and sizes can result in changes in the real-world disparity for rendering a given 3D spatial image or video, thus affecting the perceived depth and stereoscopic vision comfort. Therefore, each 3D style preset can have different target real-world disparities for different viewing configurations to maintain the same target level of stereoscopic vision comfort. For example, the disparity map can be scaled according to the maximum disparity parameter of a given 3D style preset and a given viewing configuration (e.g., a reference viewing configuration with a reference screen distance and reference screen size (such as width and height)) relative to the aspect ratio of the 3D spatial image or video.
[0105] In some specific implementations, for an alternative viewing configuration, the necessary modification of the disparity map to match the target real-world disparity of a given 3D style preset may require applying disparity_adjustment to the disparity map (M) as follows:
[0106] disparity_adjustment=\frac{max_disparity}{ref_resolution}-
[0107] \frac{real_world_disparity_mode}{width_in_meters}
[0108] In the above disparity adjustment, max_disparity represents the maximum disparity of a given reference viewing mode, ref_resolution is the reference horizontal resolution of the composite image, width_in_meters represents the width in meters of the rendered image for a given viewing configuration, and real_world_disparity_mode represents the target real_world_disparity for a given 3D style preset and a given viewing configuration.
[0109] In some specific implementations, the post - processing method can be combined with disparity_adjustment to further constrain the range. For example, disparity_adjustment can be set based on max(thr, disparity_adjustment) to allow adjustment only when it exceeds a predefined threshold (thr). Similarly, disparity_adjustment can be determined according to the function LUT(disparity_adjustment), where the disparity_adjustment value is obtained from a pre - configured look - up table LUT.
[0110] In some specific implementations, the comfort - based 3D style presets for stereoscopic 3D content playback and rendering can include adjustments or presets associated with (including but not limited to) motion parameters, binocular rivalry parameters, vertical disparity parameters, poor image quality parameters, low - light parameters, cardboard effect parameters (e.g., flat depth planes), puppet theater effect parameters (e.g., unnatural object sizes and shapes), color / brightness / saturation mismatch parameters, etc.
[0111] In some specific implementations, motion (e.g., camera motion within a captured video or content motion within the video and the resulting jitter and stutter artifacts) can cause different levels of visual discomfort for different users. Therefore, various objective metrics for quantifying motion (such as pixel - difference metrics or optical - flow - based metrics) can be used to define various 3D style presets based on the motion comfort level, thus providing different video experiences for different 3D style presets, e.g., by differently adjusting the screen size for different presets to reduce the discomfort impact of motion.
[0112] Vertical disparity caused by epipolar misalignment between stereoscopic image pairs due to calibration errors can cause different levels of visual discomfort for different users. Similarly, binocular rivalry triggered by the lack of stereoscopic correspondence (e.g., artifacts caused by flaws in occluded - area repair) can cause different levels of visual discomfort for different users. Therefore, various metrics can be evaluated to, for example, determine the amount of vertical disparity as a percentage of the image width or to detect and quantify the size of the occluded area. The evaluated metrics are used to define 3D style presets and map these 3D style presets to different types of stereoscopic content experiences.
[0113] Figure 14is a flowchart representation of an exemplary method 1400 for implementing a comfort-based 3D style preset for adjusting comfort parameters for viewing stereoscopic 3D content via a device such as an HMD. In some implementations, method 1400 is performed by a device such as a mobile device, a desktop computer, a laptop computer, an HMD, or a server device. In some implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images, such as an HMD (e.g., an HMD such as Figure 1 device 105). In some implementations, method 1400 is performed by processing logic including hardware, firmware, software, or a combination thereof. In some implementations, method 1400 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). Each block in method 1400 can be enabled and executed in any order.
[0114] At block 1402, method 1400 obtains an image depicting 2D content (e.g., input image 202 as described with respect to Figure 2 ).
[0115] At block 1404, an adjustment is performed with respect to 3D tuning parameters associated with a 3D viewing style (such as, including but not limited to, a disparity parameter preset associated with the level of disparity existing between the left-eye view and the right-eye view of a 3D spatial image or video as described with respect to Figure 13 ).
[0116] In some implementations, performing the adjustment of the 3D tuning parameters may include modifying a disparity map (e.g., disparity map (M) as illustrated in block 1302 of Figure 13 ) based on a maximum disparity parameter (max_disparity) as illustrated in block 1304 of Figure 13 . The modified disparity map can be used to perform the adjustment to control the amount of perceived depth within subsequent views of the 3D environment. In some implementations, the maximum disparity parameter can be determined as a function of viewing distance, horizontal field of view, reference resolution, and target real-world disparity, as described with respect to Figure 13 .
[0117] In some implementations, performing the adjustment of the 3D tuning parameters may include performing a disparity adjustment by modifying a disparity map (e.g., disparity map (M) of Figure 13 ) to match a target real-world disparity, as described in Figure 13 . In some implementations, the disparity adjustment parameter is determined as a function of the maximum disparity for a given reference viewing mode, reference resolution, width of the rendered image, and target real-world disparity, as described with respect to Figure 13 .
[0118] In some specific implementations, when the maximum disparity parameter (e.g., the maximum disparity (max_disparity) as described and illustrated in block 1304) exceeds a threshold level, disparity adjustment is performed. Figure 13 In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the scene depth representation format, which is different from the disparity modification.
[0119] In some specific implementations, performing an adjustment to the 3D tuning parameters may include activating preset 3D tuning parameters.
[0120] In some specific implementations, performing an adjustment to the 3D tuning parameters may include variably adjusting the 3D tuning parameters.
[0121] In some specific implementations, performing an adjustment to the 3D tuning parameters may be performed in response to user input.
[0122] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the motion parameters in subsequent views of the 3D environment.
[0123] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the binocular rivalry parameters in subsequent views of the 3D environment.
[0124] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the vertical display parameters in subsequent views of the 3D environment.
[0125] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the poor image quality parameters in subsequent views of the 3D environment.
[0126] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the low light parameters in subsequent views of the 3D environment.
[0127] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the cardboard effect (e.g., flat depth planes) parameters in subsequent views of the 3D environment.
[0128] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the puppet theater effect (e.g., unnatural object sizes and shapes) parameters in subsequent views of the 3D environment.
[0129] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the color, brightness, or sharpness mismatch parameters in subsequent views of the 3D environment.
[0130] In some specific implementations, performing an adjustment to the 3D tuning parameters may include modifying the color, brightness, or sharpness mismatch parameters in subsequent views of the 3D environment.
[0131] At block 1406, the 3D tuning parameters are used to generate a 3D stereoscopic image pair corresponding to the image, as described with respect to Figure 13 as described.
[0132] At block 1408, a view of the 3D environment including the 3D stereoscopic image pair is presented to the user via, for example, an HMD, as described with respect to Figure 13 as described.
[0133] Figure 15 is a block diagram of an example device 1500. Device 1500 illustrates an exemplary device configuration of an electronic device 105 of Figure 1 Although certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to not obscure more relevant aspects of the specific implementations disclosed herein. To that end, as a non-limiting example, in some specific implementations, device 1500 includes one or more processing units 1502 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 1504, one or more communication interfaces 1508 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.14x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, SPI, I2C, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 1510, output devices (e.g., one or more displays) 1512, one or more internal and / or external-facing image sensor systems 1514, a memory 1520, and one or more communication buses 1504 for interconnecting these and various other components.
[0134] In some specific implementations, one or more communication buses 1504 include circuitry that interconnects system components and controls communication between system components. In some specific implementations, one or more I / O devices and sensors 1506 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a tactile engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.), one or more cameras (e.g., an HMD's inward-facing camera and outward-facing camera), one or more infrared sensors, one or more thermogram sensors, and / or the like.
[0135] In some specific implementations, one or more displays 1512 are configured to present a view of a physical environment, a graphical environment, an extended reality environment, etc. to a user. In some specific implementations, one or more displays 1512 are configured to present content (determined based on a determined user / object location of the user within the physical environment) to a user. In some specific implementations, one or more displays 1512 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some specific implementations, one or more displays 1512 correspond to diffractive, reflective, polarization, holographic, etc. waveguide displays. In one example, device 1500 includes a single display. In another example, device 1500 includes a display for each eye of the user.
[0136] In some specific implementations, one or more image sensor systems 1514 are configured to obtain image data corresponding to at least a portion of the physical environment 100. For example, one or more image sensor systems 1514 include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. In various specific implementations, one or more image sensor systems 1514 further include an illumination source that emits light, such as a flash. In various specific implementations, one or more image sensor systems 1514 further include an on-camera image signal processor (ISP) that is configured to perform multiple processing operations on the image data.
[0137] In some specific implementations, sensor data may be obtained by a device (e.g., Figure 1acquired during a scan of a room in a physical environment. The sensor data can include a 3D point cloud and a sequence of 2D images corresponding to views of the room captured during the scan of the room. In some embodiments, the sensor data includes image data (e.g., from an RGB camera), depth data (e.g., a depth image from a depth camera), ambient light sensor data (e.g., from an ambient light sensor), and / or motion data from one or more motion sensors (e.g., an accelerometer, a gyroscope, an IMU, etc.). In some embodiments, the sensor data includes visual inertial odometry (VIO) data determined based on the image data. The 3D point cloud can provide semantic information about one or more elements of the room. The 3D point cloud can provide information about the location and appearance of surface portions within the physical environment. In some embodiments, the 3D point cloud is acquired over time (e.g., during a scan of the room), and the 3D point cloud can be updated, and updated versions of the 3D point cloud are acquired over time. For example, when the 3D representation is updated / adjusted over time (e.g., when a user scans a room), the 3D representation can be acquired (and analyzed / processed).
[0138] In some embodiments, the sensor data can be positioning information, and some embodiments include VIO to use sequence camera images (e.g., light intensity image data) and motion data (e.g., obtained from an IMU / motion sensor) to determine equivalent odometry information to estimate travel distance. Alternatively, some embodiments of the present disclosure can include a simultaneous localization and mapping (SLAM) system (e.g., a positioning sensor). The SLAM system can include a multi-dimensional (e.g., 3D) laser scanning and range measurement system that is independent of GPS and provides real-time simultaneous localization and mapping. The SLAM system can generate and manage data of a very accurate point cloud generated by reflections of laser scans from objects in the environment. Over time, the movement of any point in the point cloud is accurately tracked such that the SLAM system can use the points in the point cloud as reference points for position and maintain an accurate understanding of its position and orientation as it travels through the environment.
[0139] In some particular implementations, the device 1500 includes an eye tracking system for detecting eye position and eye movement (e.g., eye gaze detection). For example, the eye tracking system can include one or more infrared (IR) light emitting diodes (LEDs), an eye tracking camera (e.g., a near infrared (NIR) camera), and an illumination source (e.g., a NIR light source) that emits light (e.g., NIR light) towards the user's eyes. Additionally, the illumination source of the device 1500 can emit NIR light to illuminate the user's eyes, and the NIR camera can capture images of the user's eyes. In some particular implementations, the images captured by the eye tracking system can be analyzed to detect the positioning and movement of the user's eyes, or to detect other information about the eyes such as pupil dilation or pupil diameter. Additionally, the gaze point estimated from the eye tracking images can enable gaze-based interaction with the content shown on the near-eye display of the device 1500.
[0140] The memory 1520 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some particular implementations, the memory 1520 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 1520 optionally includes one or more storage devices that are remotely located from one or more processing units 1502. The memory 1520 includes non-transitory computer-readable storage media.
[0141] In some particular implementations, the memory 1520 or the non-transitory computer-readable storage media of the memory 1520 stores an optional operating system 1530 and one or more instruction sets 1540. The operating system 1530 includes procedures for handling various basic system services and for performing hardware-related tasks. In some particular implementations, the instruction set 1540 includes executable software defined by binary information stored in the form of charge. In some particular implementations, the instruction set 1540 is software that can be executed by one or more processing units 1502 to implement one or more of the techniques described herein.
[0142] The instruction set 1540 includes an input image instruction set 1542 and an output image conversion instruction set 1544. The instruction set 1540 can be embodied as a single software executable file or multiple software executable files.
[0143] The input image instruction set 1542 is configured with instructions executable by a processor to determine, receive, and process a single-view input image for conversion into a stereoscopic image pair.
[0144] The output image conversion instruction set 1544 is configured with instructions that can be executed by a processor to convert single-view image content into a stereoscopic image pair using a viewpoint, depth, or boundary adjustment process plus comfort parameter adjustment.
[0145] Although instruction set 1540 is shown as residing on a single device, it should be understood that in other embodiments, any combination of the elements may be located in separate computing devices. Additionally, Figure 15 It is more used as a functional description of the various features present in a particular embodiment, which is different from the structural schematic of the embodiments described herein. As will be recognized by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. The actual number of instruction sets and how the features are allocated therein will vary depending on the embodiment and may depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.
[0146] Those of ordinary skill in the art will understand that well-known systems, methods, components, devices, and circuits are described so as not to obscure more relevant aspects of the example embodiments described herein. Additionally, other effective aspects and / or variations do not include all of the details of the specific details described herein. Thus, several details are described in order to provide a thorough understanding of the example aspects shown in the figures. Additionally, the figures only show some example embodiments of the present disclosure and should not be considered limiting.
[0147] Although this specification contains many specific implementation details, these specific implementation details should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of different embodiments in this specification may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although certain features may be described as acting in certain combinations and even initially claimed as such, one or more features of the claimed combination may in some cases be removed from the combination, and the claimed combination may relate to a sub-combination or a variation of a sub-combination.
[0148] Similarly, although operations are shown in the figures in a particular order, this should not be construed as requiring that such operations be performed in a sequential order or the particular order shown, or that all of the illustrated operations be performed to achieve the desired result. In some circumstances, multitasking and parallel processing may be advantageous. In addition, the partitioning of various system components in the above-described embodiments should not be construed as requiring such partitioning in all embodiments, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged into multiple software products.
[0149] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order or sequence shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.
[0150] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more modules of computer program instructions encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium can be, or include, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, although a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be, or include, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0151] The term "data processing apparatus" encompasses all kinds of devices, equipment, and machines for processing data, such as including programmable processors, computers, systems on a chip, or multiple or combinations of the foregoing. The apparatus may include dedicated logic circuitry (e.g., FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit)). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program under consideration, such as code that constitutes processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations of one or more of them. The apparatus and the execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. Unless otherwise specifically stated, it should be understood that throughout the specification, discussions using terms such as "processing", "computing", "computing out", "determining", and "identifying" refer to actions or processes of a computing device, such as one or more computers or similar electronic computing devices, which manipulate or transform data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.
[0152] One or more of the systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditional on one or more inputs. Suitable computing devices include computer systems based on general-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device into a dedicated computing device that implements one or more specific implementations of the subject matter of the present invention. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in the software for programming or configuring a computing device.
[0153] Specific implementations of the methods disclosed herein may be performed in the operation of such computing devices. The order of the boxes presented in the examples above may vary; for example, the boxes may be reordered, combined, and / or divided into sub-blocks. Some boxes or processes may be performed in parallel. The operations described in this specification may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0154] The use herein of "suitable for" or "configured to" means open and inclusive language that does not exclude devices suitable for or configured to perform additional tasks or steps. Additionally, the use of "based on" means open and inclusive because a process, step, computation, or other action "based on" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbers included herein are for ease of explanation only and are not intended to be restrictive.
[0155] It will also be understood that, although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first node may be referred to as a second node, and similarly, a second node may be referred to as a first node, which changes the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. The first node and the second node are both nodes, but they are not the same node.
[0156] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the present embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the term "comprising", when used in this specification, specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0157] As used herein, the term "if" can be interpreted to mean "when the precondition is true" or "while the precondition is true" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the precondition is true, depending on the context. Similarly, the phrases "if it is determined [that the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" are interpreted to mean "when it is determined that the precondition is true" or "in response to determining" or "in accordance with determining" that the precondition is true or "when it is detected that the precondition is true" or "in response to detecting" that the precondition is true, depending on the context.
Claims
1. A method comprising: At an electronic device having a processor: obtaining an input image comprising appearance values at pixel locations, the input image corresponding to the appearance of a scene from a first viewpoint; determining a depth image comprising depth values at original pixel locations mapped to at least a subset of the pixel locations of the input image, wherein a coordinate mapping maps the original pixel locations to corresponding pixel locations in the input image; generating a first output image corresponding to a second viewpoint of the scene different from the first viewpoint, the first output image generated by determining a first set of altered pixel locations for the depth values and identifying appearance values for the first set of altered pixel locations based on the coordinate map and the input image; as well as A second output image is generated corresponding to a third viewpoint of the scene that is different from the second viewpoint, the second output image being generated by determining a second set of altered pixel locations for the depth values and identifying appearance values for the second set of altered pixel locations based on the coordinate map and the input image. 2 . The method of claim 1 , wherein the first viewpoint corresponds to a center viewpoint, the second viewpoint corresponds to a left-eye viewpoint, and the third viewpoint corresponds to a right-eye viewpoint.
3. The method according to claim 2, wherein: The first output image is a left-eye image generated based on the input image and a first coordinate image, wherein the first coordinate image is (a) determined based on the depth image and (b) warped for the left-eye viewpoint; and The second output image is a right-eye image generated based on the input image and a second coordinate image, wherein the second coordinate image is (a) determined based on the depth image and (b) warped for the right-eye viewpoint.
4. The method of claim 1, wherein the depth image is generated based on evaluating the input image using a neural network. The method of claim 1 , wherein the coordinate map is a coordinate image.
6. The method of claim 1, further comprising providing the first output image and the second output image to form a stereoscopic output image pair depicting the scene for viewing on a stereoscopic display of a head mounted device (HMD).
7. The method of claim 1 , wherein generating the output input image further comprises: Appearance values for additional pixel locations other than the second set of altered pixel locations are identified based on the coordinate map and the input image.
8. The method of claim 7, wherein identifying the appearance value of the additional pixel location comprises: identifying an intermediate pixel location between two adjacent pixel locations in the second set of altered pixel locations; as well as The appearance value of the intermediate pixel location is identified by identifying, based on the coordinate mapping, the pixel location in the input image between pixel locations in the input image that correspond to the two adjacent pixel locations.
9. A system comprising: processor; A computer-readable medium storing instructions that, when executed by the processor, cause the processor to perform operations including the following: Obtain an input image including appearance values at pixel locations, the input image corresponding to the appearance of a scene from a first viewpoint; Determine a depth image including depth values at original pixel locations that map to at least a subset of the pixel locations in the input image, where a coordinate mapping maps the original pixel locations to corresponding pixel locations in the input image; Generate a first output image corresponding to a second viewpoint of the scene different from the first viewpoint, the first output image being generated by determining a first set of modified pixel locations for the depth values and identifying appearance values of the first set of modified pixel locations based on the coordinate mapping and the input image; And Generate a second output image corresponding to a third viewpoint of the scene different from the second viewpoint, the second output image being generated by determining a second set of modified pixel locations for the depth values and identifying appearance values of the second set of modified pixel locations based on the coordinate mapping and the input image.
10. The system according to claim 9, wherein the first viewpoint corresponds to a center viewpoint, the second viewpoint corresponds to a left-eye viewpoint, and the third viewpoint corresponds to a right-eye viewpoint.
11. The system according to claim 10, wherein: The first output image is a left-eye image generated based on the input image and a first coordinate image, the first coordinate image being (a) determined based on the depth image and (b) distorted for the left-eye viewpoint; and The second output image is a right-eye image generated based on the input image and a second coordinate image, the second coordinate image being (a) determined based on the depth image and (b) distorted for the right-eye viewpoint.
12. The system according to claim 10, wherein the depth image is generated based on evaluating the input image using a neural network.
13. The system according to claim 9, wherein the coordinate mapping is a coordinate image.
14. The system according to claim 9, wherein the operations further include providing the first output image and the second output image to form a stereoscopic output image pair depicting the scene for viewing on a stereoscopic display of a head-mounted device (HMD).
15. The system according to claim 9, wherein generating the output input image further includes: Identifying appearance values of additional pixel locations other than the second set of modified pixel locations based on the coordinate mapping and the input image.
16. The system according to claim 15, wherein identifying the appearance values of the additional pixel locations includes: Identifying intermediate pixel locations between two adjacent pixel locations in the second set of modified pixel locations; And The appearance value of the intermediate pixel location is identified by identifying, based on the coordinate mapping, the pixel location in the input image between pixel locations in the input image that correspond to the two adjacent pixel locations.
17. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform operations comprising: obtaining an input image comprising appearance values at pixel locations, the input image corresponding to the appearance of a scene from a first viewpoint; determining a depth image comprising depth values at original pixel locations mapped to at least a subset of the pixel locations of the input image, wherein a coordinate mapping maps the original pixel locations to corresponding pixel locations in the input image; generating a first output image corresponding to a second viewpoint of the scene different from the first viewpoint, the first output image generated by determining a first set of altered pixel locations for the depth values and identifying appearance values for the first set of altered pixel locations based on the coordinate map and the input image; as well as A second output image is generated corresponding to a third viewpoint of the scene that is different from the second viewpoint, the second output image being generated by determining a second set of altered pixel locations for the depth values and identifying appearance values for the second set of altered pixel locations based on the coordinate map and the input image. 18 . The non-transitory computer-readable medium of claim 17 , wherein the first viewpoint corresponds to a center viewpoint, the second viewpoint corresponds to a left-eye viewpoint, and the third viewpoint corresponds to a right-eye viewpoint.
19. The non-transitory computer readable medium of claim 18, wherein: The first output image is a left-eye image generated based on the input image and a first coordinate image, wherein the first coordinate image is (a) determined based on the depth image and (b) warped for the left-eye viewpoint; and The second output image is a right-eye image generated based on the input image and a second coordinate image, wherein the second coordinate image is (a) determined based on the depth image and (b) warped for the right-eye viewpoint.
20. The non-transitory computer-readable medium of claim 17, further comprising providing the first output image and the second output image to form a stereoscopic output image pair depicting the scene for viewing on a stereoscopic display of a head-mounted device (HMD).
Citation Information
Patent Citations
Synthetic stereoscopic content capture
CN110546951A
Runtime optimized artificial vision
CN114930392A
Image processing device, image processing method, image capturing device, and mobile vehicle
JP2023029064A
Apparatus and method for generating depth image that have same viewpoint and same resolution with color image
KR1020120018915A