Neural Blending for Novel View Synthesis

Deep learning methods with neural networks and witness cameras enhance object rendering by reducing artifacts and improving geometric accuracy in novel view synthesis, addressing computational inefficiencies and image noise in traditional rendering.

JP7804725B2Active Publication Date: 2026-01-22GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024109644
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2026-01-22
Estimated Expiration
2041-04-08

AI Technical Summary

Technical Problem

Traditional object rendering in motion involves significant computational effort and often results in noisy images with geometric artifacts due to inaccurate blending of image contributions from different views.

Method used

The use of deep learning methods employing neural networks to blend image content for novel view synthesis, incorporating learned blending weights and multi-resolution techniques to reduce artifacts, and employing witness cameras for ground truth data to improve accuracy.

Benefits of technology

Generates high-quality novel views with reduced artifacts and improved geometric accuracy by leveraging learned blending weights and ground truth data, enhancing image-based rendering processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007804725000010
    Figure 0007804725000010
  • Figure 0007804725000011
    Figure 0007804725000011
  • Figure 0007804725000012
    Figure 0007804725000012
Patent Text Reader

Abstract

To provide a method, an apparatus, and an algorithm for generating a virtual view of a target object.SOLUTION: A method comprises: receiving a plurality of input images, a plurality of depth images, and a plurality of view parameters; generating a plurality of warped images based on the plurality of input images, the plurality of view parameters, and one of the plurality of depth images; receiving blending weights from a neural network for assigning colors to pixels of a virtual view of a target object in response to providing the plurality of depth images, the plurality of view parameters, and the plurality of warped images to the neural network; and generating a synthetic image according to the view parameters based on the blending weights and the virtual view.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Technical Field This description relates generally to methods, apparatus and algorithms used in three-dimensional (3D) content synthesis. [Background technology]

[0002] background Traditional object rendering typically involves a significant amount of computation to generate a realistic image. When an object is in motion, additional computational effort may be used to generate a realistic image of the object. Such rendering may use neural networks to model the appearance of the object. However, such models may produce images with excessive noise and geometric artifacts. Summary of the Invention

[0003] overview The systems and methods described herein may perform image-based rendering using input images and predefined view parameters to generate (e.g., synthesize) novel (e.g., unseen) views of video and / or images based on the input images. Image-based rendering of unseen views may utilize a warping process on the received input images. Generally, warping processes may introduce geometric inaccuracies and view- and / or image-dependent effects that may generate artifacts when contributions from different input views are blended. The systems and methods described herein use deep learning methods that employ neural networks (NNs) to blend image content for image-based rendering of novel views. Specific blending weights are learned and used to combine the contributions of the input images to the final synthesized view. The blending weights are generated to provide the advantage of generating a synthesized image that exhibits reduced view and / or image-dependent effects and a fewer number of image artifacts.

[0004] A technical challenge that can arise when using NNs, warping, and / or blending weights is the lack of sufficiently accurate geometry so that the NN (e.g., a convolutional neural network) can select appropriate blending weights to avoid image artifacts. The systems and methods described herein may solve this technical challenge by using learned blends of color and depth views of the input image and / or employing multi-resolution blending techniques to select pixel colors that provide an accurate image with reduced image artifacts. For example, blending weights may be applied to heavily weight projected (e.g., probabilistically provided) pixel colors that are likely to be appropriate and accurate for a ground truth image, while de-weighting projected pixel colors that are unlikely to be appropriate and / or accurate for a given ground truth image.

[0005] To employ such blending techniques, the systems and methods described herein may utilize one or more witness cameras in addition to specific on-board system cameras (e.g., color cameras, infrared cameras, etc.). The witness camera(s) may monitor the content used to generate the novel view. For example, the witness camera(s) may be high-resolution cameras that may function to provide ground truth data. The novel view generated may be based on the images captured by the witness camera(s). The new view is compared to the received (e.g., captured) ground truth data. In some implementations, the image details of the new view can be scored based on the image details captured by the witness camera(s) in generating the new view.

[0006] In some implementations, the systems and methods described herein take training loss into account. For example, the system may generate training data with a variety of captured scenes to minimize loss in order to provide high-quality novel view synthesis while reducing temporal flickering artifacts in the synthesized view. Also, in some implementations, the systems and methods described herein may employ occlusion inference to correct artifacts in the synthesized novel view.

[0007] A system of one or more computers can be configured to perform particular operations or actions by having installed on the system software, firmware, hardware, or a combination thereof that is capable of causing the system to perform actions during operation. One or more computer programs can be configured to perform particular operations or actions by containing instructions that, when executed by a data processing device, cause the device to perform the actions.

[0008] In one general aspect, systems and methods are described for receiving a plurality of input images, receiving a plurality of depth images associated with a target object in at least one of the plurality of input images, receiving a plurality of view parameters, and generating a plurality of warped images based on the plurality of input images, the plurality of view parameters, and at least one of the plurality of depth images to generate a virtual view of the target object. In response to providing the plurality of depth images, the plurality of view parameters, and the plurality of warped images to a neural network, the systems and methods may receive blending weights from the neural network for assigning colors to pixels of a virtual view of the target object. The systems and methods may generate a composite image in accordance with the view parameters based on the blending weights and the virtual view.

[0009] These and other aspects may include one or more of the following, alone or in combination: In some implementations, the systems and methods may comprise reconstructing a consensus surface using a geometric fusion process on the plurality of depth images to generate a geometrically fused model, and generating a plurality of reprojection images based on the plurality of input images and the consensus surface, wherein the systems and methods may receive additional blending weights from the neural network in response to providing the plurality of depth images, the plurality of view parameters, and the plurality of reprojection images to the neural network for assigning colors to pixels in the composite image.

[0010] In some implementations, the system and method may further comprise providing a difference between the depth of the geometrically fused model and the depth observed in the plurality of depth images to the neural network, and the method further comprises correcting the detected occlusion in the composite image based on the depth difference. In some implementations, the plurality of input images are color images captured according to predefined view parameters associated with at least one camera that captured the plurality of input images, and / or the plurality of depth images each include a depth map associated with at least one camera that captured at least one of the plurality of input images, at least one occlusion map, and / or a depth map associated with a ground truth image captured by at least one witness camera at a time corresponding to the capture of at least one of the plurality of input images. In some implementations, the blending weights assign a blended color to each pixel of the composite image. It is configured to

[0011] In some implementations, the neural network is trained based on minimizing an occlusion loss function between a synthetic image generated by the neural network and a ground truth image captured by at least one witness camera, hi some implementations, the synthetic image is an uncaptured view of a target subject generated for three-dimensional video conferencing.

[0012] In some implementations, generating a plurality of warped images based on a plurality of input images, a plurality of view parameters, and at least one of a plurality of depth images includes using at least one of the plurality of depth images to determine candidate projections of colors associated with the plurality of input images onto an uncaptured view that includes at least a portion of image features of at least one of the plurality of input images.

[0013] In another general aspect, an image processing system is described that is particularly adapted to perform the method of any one of the preceding claims. The image processing system may include at least one processor and a memory storing instructions that, when executed, cause the system to perform operations including receiving a plurality of input images captured by the image processing system, receiving a plurality of depth images captured by the image processing system, receiving a plurality of view parameters associated with an uncaptured view associated with at least one of the plurality of input images, and generating a plurality of warped images based on the plurality of input images, the plurality of view parameters, and at least one of the plurality of depth images. In response to providing the plurality of depth images, the plurality of view parameters, and the plurality of warped images to a neural network, the system may include receiving blending weights from the neural network for assigning colors to pixels of the uncaptured view. The system may further include generating a composite image according to the blending weights, the composite image corresponding to the uncaptured view.

[0014] These and other aspects may include one or more of the following, alone or in combination: In some implementations, the plurality of input images are color images captured by the image processing system according to predefined view parameters associated with the image processing system, and / or the plurality of depth images include a depth map associated with at least one camera that captured at least one of the plurality of input images, at least one occlusion map, and / or a depth map associated with a witness camera of the image processing system.

[0015] In some implementations, the blending weights are configured to assign a blended color to each pixel of the composite image. In some implementations, the neural network is trained based on minimizing an occlusion loss function between the composite image generated by the neural network and a ground truth image captured by at least one witness camera. In some implementations, the composite image is a novel view generated for three-dimensional video conferencing.

[0016] In another general aspect, a non-transitory machine-readable medium is described as storing instructions that, when executed by a processor, cause a computing device to receive a plurality of input images, receive a plurality of depth images associated with a target object in at least one of the plurality of input images, and receive a plurality of view parameters to generate a virtual view of the target object. The non-transitory machine-readable medium also describes a method for reconstructing a consensus surface using a geometric fusion process on the plurality of depth images to generate a geometrically fused model of the target object, and a method for reconstructing a consensus surface using a geometric fusion process on the plurality of inputs and the plurality of view parameters. and generating a plurality of reprojection images based on the meter and the consensus surface. The machine-readable medium may be configured to, in response to providing the plurality of depth images, the plurality of view parameters, and the plurality of reprojection images to the neural network, receive blending weights from the neural network for assigning colors to pixels of a virtual view of the target subject, and generate a composite image according to the view parameters based on the blending weights and the virtual view.

[0017] These and other aspects may include one or more of the following, alone or in combination: In some implementations, the machine-readable medium further includes providing a difference between the depth of the geometrically fused model and the depth observed in the plurality of depth images to a neural network, and correcting the detected occlusion in the composite image based on the difference in depth. In some implementations, the plurality of input images are color images captured according to predefined view parameters associated with at least one camera that captured the plurality of input images, and / or the plurality of depth images include a depth map associated with at least one camera that captured at least one of the plurality of input images, at least one occlusion map, and / or a depth map associated with a ground truth image captured by at least one witness camera when corresponding to the capture of at least one of the plurality of input images.

[0018] In some implementations, the blending weights are configured to assign blended colors to each pixel of the composite image. In some implementations, the neural network is trained based on minimizing an occlusion loss function between the synthetic image generated by the neural network and a ground truth image captured by at least one witness camera. In some implementations, the synthetic image is a novel view for 3D video conferencing. In some implementations, the neural network is further configured to perform multi-resolution blending to assign pixel colors to pixels in the composite image, the multi-resolution blending triggering providing an image pyramid as input to the neural network and receiving from the neural network multi-resolution blending weights for multiple scales and opacity values ​​associated with each scale.

[0019] These and other aspects may include one or more of the following, alone or in combination: According to some aspects, the methods, systems, and computer-readable media claimed herein may include one or more (e.g., all) of the following features (or any combination thereof).

[0020] Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium. Details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 is a block diagram illustrating an example of a 3D content system for displaying composite content on a display device, according to implementations described throughout this disclosure. [Figure 2] FIG. 1 is a block diagram illustrating an example system for compositing content for rendering on a display, according to implementations described throughout this disclosure. [Figure 3] FIG. 2 is a block diagram illustrating an example of reprojection of an input image onto a target camera viewpoint, according to implementations described throughout this disclosure. [Figure 4] FIG. 10 is a block diagram illustrating an example of a flow diagram for using neural blending to generate synthetic content for rendering on a display, according to implementations described throughout this disclosure. [Figure 5] FIG. 10 is a block diagram illustrating an example of a flow diagram for generating blending weights, according to implementations described throughout this disclosure. [Figure 6] 1 is a flowchart illustrating an example process for generating synthetic content using neural blending, according to implementations described throughout this disclosure. [Figure 7] 1A-1C illustrate examples of computing devices and mobile computing devices that can be used with the techniques described herein. DETAILED DESCRIPTION OF THE INVENTION

[0022] Like symbols in the various drawings indicate like elements. Detailed Description Examples are described herein related to generating novel (e.g., unseen) views of image content. Examples described herein may synthesize (e.g., generate) real-time novel views based on captured video content and / or image content. For example, image-based rendering techniques may be used to synthesize novel views of video content (e.g., objects, users, scene content, image frames, etc.) using learned blending of color and depth views.

[0023] The systems and methods described herein may generate novel color images with fewer artifacts than conventional systems. For example, the systems and methods described herein may correct for certain image noise and loss function analysis to generate novel images with less inaccurate depth and occlusion. Correction may be performed by employing a neural network (NN) to learn to detect and correct image regions containing visibility errors. Furthermore, the NN can learn and predict color values ​​for novel views using a blending algorithm that constrains output values ​​to be a linear combination of reprojected input colors derived from the color input image.

[0024] During operation, the process may retrieve (e.g., capture, acquire, receive, etc.) multiple input images and data (e.g., target view parameters) to predict a new view (e.g., an unseen color image) by combining color image streams from input images (e.g., views) of the same scene (e.g., image content within the scene). The color image streams may be provided to the NN to employ neural rendering techniques to enhance low-quality output from a real-time image capture system (e.g., a 3D video conferencing system such as a telepresence system). For example, the new view may be a predicted color image generated by the systems and techniques described herein. The predicted image may be generated by providing the NN with input images and combined color image streams (e.g., and / or reprojections or re-representations of such input images) to learn specific blending weights for assigning pixel colors to the predicted color image. The learned blending weights may be applied to generate pixel colors of the new color image. The learned blending weights may also be used to generate other new views of the image content shown in one or more provided input images.

[0025] In some implementations, the NNs described herein may model view-dependent effects that predict future user motion (e.g., movement) to mitigate misprojection artifacts caused by the noisiness of the particular geometry information used to generate the user's image, and / or the geometry information received from the camera capturing the user, and / or the information received from image processing performed on the user's image.

[0026] In some implementations, the systems and methods described herein may include, for example, a separate witness camera's viewpoint, which may be used to provide surveillance to the output color image. One or more NNs (e.g., convolutional NNs such as U-net) can be trained to predict the resulting image. The witness camera may serve as a ground truth camera for the image capture and / or processing system described herein. In some implementations, two or more witness cameras may be used as training data for the NN. The two or more witness cameras may represent one or more pairs of witness cameras.

[0027] In some implementations, the systems and methods may utilize a captured input image, predefined parameters associated with a desired new output view, and / or an occlusion map and a depth map including a depth difference. The depth difference may be generated using a view from a color camera between the surface closest to the new view and the surface of the camera view. The depth difference may be used for occlusion inference to correct for occluded views and / or other errors in the generated image. In some implementations, the depth map may include a depth map from a view captured by a witness camera.

[0028] In some implementations, the systems and methods described herein may reconstruct a consensus surface (e.g., a geometric surface) by geometric fusion of input depth images. In some implementations, the systems and methods described herein may use depth information, such as separately captured depth images and / or a consensus surface, to determine the projection of input colors onto a new view.

[0029] In some implementations, the systems and methods described herein may generate a color image of a new view (e.g., a color image) by assigning a blended color to each pixel in the new view. The blended color may be determined using the color input image and blending weights determined by the NN described herein. In some implementations, the blending weights are normalized through a loss function. In some implementations, the new view is a weighted combination of one or more pixel color values ​​of the image projected from the original input image onto the new view.

[0030] As used herein, a novel (e.g., unseen) view may include image content and / or video content that is interpreted (e.g., synthesized, interpolated, modeled, etc.) based on one or more frames of image content and / or video content captured by a camera. Interpretation of camera-captured image content and / or video content may be used in combination with techniques described herein to, for example, create unseen versions and views (e.g., poses, expressions, angles, etc.) of the captured image content and / or video content.

[0031] In some implementations, the techniques described herein can be used to synthesize accurate and realistic-looking images for display on the screen of a 2D or 3D display used, for example, in multi-directional 2D or 3D video (e.g., telepresence) conferencing. The techniques described herein can be used to generate and display accurate and realistic views (e.g., image content, video content) of users in a video conference. Traditionally, views include invisible views that can be difficult to depict in 3D without significant image artifacts.

[0032] The systems and methods described herein provide the advantage of generating novel views without significant image artifacts by using one or more witness cameras and a neural network to learn blending weights based on multi-view color input images and noise occlusion cues. The learned blending weights are used to generate the resulting It can be ensured that occlusion and color artifacts are corrected in the output image. Furthermore, the learned blending weights and one or more witness cameras can be used by the system described herein to ensure that image content not captured in the input image can be used to accurately predict new views associated with the image content in the input image. For example, because the blending weights are learned and evaluated with respect to witness camera images, accurate predictions can be made for image portions of the scene that were not captured or represented in the original input image.

[0033] In some implementations, the techniques described herein may be used for entertainment purposes in movies, videos, short films, gaming content, virtual and / or augmented reality content, or other formats involving user images that can benefit from the predictive techniques described herein. For example, the techniques described herein may be used to generate novel views for moving characters rendered in images and / or video content.

[0034] In some implementations, the techniques described herein can be used by a virtual assistant device or other intelligent agent that can use the techniques described herein to recognize objects, reproduce objects, and / or perform image processing to generate synthetic images from such objects.

[0035] FIG. 1 is a block diagram illustrating an example of a 3D content system 100 for displaying content on a stereoscopic display device, according to implementations described throughout this disclosure. The 3D content system 100 can be used by multiple users, for example, to conduct videoconference communications in 3D (e.g., telepresence sessions) and / or access augmented reality and / or virtual reality content. In general, the system of FIG. 1 can be used to capture video and / or images of a user and / or scene during a 2D or 3D videoconference and generate a new view based on the captured content using the systems and techniques described herein to render accurate images depicting the new view within the videoconference session. System 100 can benefit from the use of the techniques described herein, since such techniques can generate and display real-time new views that accurately represent the user, for example, within a videoconference. The new view may be provided for display to another user in 2D and / or 3D via system 100, for example.

[0036] 1, a 3D content system 100 is accessed by a first user 102 and a second user 104. For example, users 102 and 104 may access the 3D content system 100 to participate in a 3D telepresence session. In such an example, the 3D content system 100 may enable each of users 102 and 104 to view highly realistic, visually consistent representations of each other, facilitating the users to interact with each other in a manner similar to being physically present.

[0037] Each user 102, 104 may conduct a 3D telepresence session using a corresponding 3D system, where user 102 accesses 3D system 106 and user 104 accesses 3D system 108. 3D systems 106, 108 may provide functionality related to 3D content, including, but not limited to, capturing images for 3D display, processing and presenting image information, and processing and presenting audio information. 3D system 106 and / or 3D system 108 may comprise a collection of sensing devices integrated as a single unit. 3D system 106 and and / or the 3D system 108 may include some or all of the components described with reference to FIGS.

[0038] The 3D content system 100 may include one or more 2D or 3D displays. Here, 3D display 110 is depicted for 3D system 106, and 3D display 112 is depicted for 3D system 108. The 3D displays 110, 112 can use any of several types of 3D display technologies to provide stereoscopic viewing for the respective viewer (e.g., user 102 or user 104). In some implementations, the 3D displays 110, 112 may be standalone units (e.g., self-supporting or wall-mounted). In some implementations, the 3D displays 110, 112 may include or have access to wearable technology (e.g., a controller, a head-mounted display, AR glasses, etc.). In some implementations, the displays 110, 112 may be 2D displays.

[0039] In general, displays 110, 112 can provide images that approximate the 3D optical properties of real-world physical objects without the use of a head-mounted display (HMD) device. The displays described herein may include flat-panel displays containing lenticular lenses (e.g., microlens arrays) and / or parallax barriers that redirect images to multiple different viewing areas associated with the displays.

[0040] In some implementations, displays 110, 112 may include high-resolution, glasses-free lenticular 3D displays. For example, displays 110, 112 may include a microlens array (not shown) including multiple lenses (e.g., microlenses) with glass spacers coupled (e.g., glued) to the microlenses of the display. The microlenses may be designed such that, from a selected viewing position, the left eye of a user of the display can see a first set of pixels and the right eye of the user can see a second set of pixels (e.g., the second set of pixels is mutually exclusive to the first set of pixels).

[0041] In some display examples, there may be one location that provides a 3D view of the image content (e.g., a user, an object, etc.) provided by such a display. The user may sit in one location to experience proper parallax, minimal distortion, and realistic 3D imagery. As the user moves to a different physical location (or changes head position or gaze position), the image content (e.g., the user, objects worn by the user, and / or other objects) may begin to appear less realistic, 2D, and / or distorted. The systems and techniques described herein may reconfigure the image content projected from the display so that the user can move around and still reliably experience proper parallax, low distortion, and realistic 3D imagery in real time. Thus, the systems and techniques described herein provide the advantage of maintaining and providing 3D image content and objects for display to the user regardless of any user movement that occurs while the user is viewing the 3D display.

[0042] As shown in Figure 1, 3D content system 100 can be connected to one or more networks. Here, network 114 is connected to 3D system 106 and 3D system 108. Network 114 can be a publicly available network (e.g., the Internet) or a private network, to name just two examples. Network 114 can be wired, or wireless, or a combination of the two. Network 114 may include or utilize one or more other devices or systems, including, but not limited to, one or more servers (not shown).

[0043] The 3D systems 106, 108 may include multiple components related to capturing, processing, transmitting or receiving 3D information, and / or presenting 3D content. The 3D systems 106, 108 may include one or more cameras for capturing image content and / or video (e.g., visible and infrared image data) for images included in the 3D representation. In the illustrated example, the 3D system 106 includes cameras 116 and 118. For example, the cameras 116 and / or 118 can essentially be disposed within the housing of the 3D system 106 such that the objective or lens of each camera 116 and / or 118 captures image content through one or more openings in the housing. In some implementations, the cameras 116 and / or 118 can be separate from the housing, such as in the form of a standalone device (e.g., with a wired and / or wireless connection to the 3D system 106). The cameras 116 and 118 can be positioned and / or oriented to capture a sufficiently representative view of the user (e.g., the user 102).

[0044] Although cameras 116 and 118 generally do not obstruct the view of 3D display 110 for user 102, the placement of cameras 116 and 118 can be selected arbitrarily. For example, one of cameras 116, 118 can be positioned somewhere above user 102's face, and the other can be positioned somewhere below. For example, one of cameras 116, 118 can be positioned somewhere on the right side of user 102's face, and the other can be positioned somewhere on the left side of the face. 3D system 108 can include, for example, cameras 120 and 122 in a similar manner. Additional cameras are also possible. For example, a third camera can be positioned near or behind display 110.

[0045] In some implementations, the 3D systems 106, 108 may include one or more witness cameras 119, 121. The witness cameras 119, 121 may be used to capture high-quality images (e.g., witness camera image 132) that may represent ground truth images. Images captured by the witness cameras 119 and / or 121 may be used with the techniques described herein to generate new views and to be used as comparisons when calculating losses and corrections for such losses. Generally, an image captured by the witness cameras 119, 121 may be captured at substantially the same moment as a corresponding one of other images (e.g., frames) captured by cameras 116, 118, 120, 122, 124, and / or 126, as well as combinations of such cameras and / or camera pods. In some implementations, the witness camera image 134 may be captured and used as training data for one or more neural networks to generate new views.

[0046] In some implementations, 3D systems 106, 108 may include one or more depth sensors to capture depth data used in the 3D representation. Such depth sensors may be considered part of a depth capture component within 3D content system 100 used to characterize the scene captured by 3D systems 106 and / or 108 to properly represent the scene on a 3D display. Additionally, the system may track the viewer's head position and orientation so that the 3D representation may be rendered with an appearance corresponding to the viewer's current viewpoint. Here, 3D system 106 includes depth sensor 124, which may also represent an infrared camera. In a similar manner, 3D system 108 may include depth sensor 126. Any of several types of depth perception or depth capture may be used to generate the depth data.

[0047] In some implementations, each of cameras 116, 118, 119, and 124 may represent multiple cameras in a pod. For example, depth sensor 124 may be housed together with camera 116 and / or camera 118 in a camera pod. In some implementations, three or more camera pods may be positioned around and / or behind display 110, and each pod may include camera 124 (e.g., a depth sensor / camera) and one or more cameras 116, 118. Similarly, three or more camera pods may be positioned around and / or behind display 112, and each pod may include camera 126 (e.g., a depth sensor / camera) and one or more cameras 120, 122.

[0048] During operation of the system 106, assisted stereo depth capture may be performed. The scene may be illuminated with dots of light, and stereo matching may be performed, for example, between two respective cameras. Such illumination may be performed using waves of a selected wavelength or wavelength range. For example, infrared (IR) light may be used. The depth data may include or be based on any information about the scene that reflects the distance between a depth sensor (e.g., depth sensor 124) and an object in the scene. For content in the image that corresponds to an object in the scene, the depth data reflects the distance (or depth) to the object. For example, the spatial relationship between the camera(s) and the depth sensor may be known and may be used to correlate images from the camera(s) with signals from the depth sensor to generate depth data for the image.

[0049] Images captured by the 3D content system 100 can be processed and then displayed as a 3D representation. As illustrated in the example of FIG. 1 , a 3D image of a user 104 is presented on a 3D display 110. In this manner, a user 102 can perceive a 3D image 104' (e.g., of the user) as a 3D representation of the user 104, who may be located at a distance from the user 102. Similarly, a 3D image 102' is presented on a 3D display 112. In this manner, the user 104 can perceive the 3D image 102' as a 3D representation of the user 102.

[0050] The 3D content system 100 may enable participants (e.g., users 102, 104) to engage in audio communications with each other and / or others. In some implementations, the 3D system 106 includes a speaker and a microphone (not shown). For example, the 3D system 108 may similarly include a speaker and a microphone. In this manner, the 3D content system 100 may enable users 102 and 104 to participate in 3D telepresence sessions with each other and / or others. In general, the systems and techniques described herein may work in conjunction with the system 100 to generate image and / or video content for display among users of the system 100.

[0051] During operation of the system 100, a set of input images 132 may be captured by cameras 116, 118, 119, 124 and / or 120, 121, 122, and 126. The input images may include, for example, a witness camera image 134 and an RGB color image 136. In some implementations, the system 100 may also generate and / or otherwise acquire a depth image 138. In one example, the depth image 138 may be generated by performing one or more stereo calculations from a pair of IR images acquired from an IR camera, as described above. The input images 132 may be used as a basis for predicting an output image that is a linear combination of reprojected colors from the input image(s). In some implementations, the input images 132 represent reprojected color images (e.g., red-green-blue (RGB)) captured with known (e.g., predetermined, predefined) view parameters. The input images 132 may include two or more color images. In some implementations, the input images 132 also include one or more depth images 138 calculated (e.g., generated) with known view parameters. The input images 132 may be used in combination with specific camera parameters, view parameters, and / or NN blending algorithms 140 to generate new views for display on the displays 110 and / or 112.

[0052] 2 is a block diagram illustrating an example system for compositing content for rendering on a display, according to implementations described throughout this disclosure. System 200 can function as or be included in one or more implementations described herein and / or can be used to perform the operation(s) of one or more examples of image content compositing, processing, modeling, or representation described herein. Overall system 200 and / or one or more of its individual components can be implemented in accordance with one or more examples described herein.

[0053] System 200 may include one or more 3D systems 202. In the illustrated example, 3D systems 202A, 202B, through 202N are shown, where the index N denotes any number. 3D system 202 can provide visual and audio information capture for 2D or 3D representations and can transfer 2D or 3D information for processing. Such information can include images of a scene, depth data about the scene, parameters associated with image capture, and / or audio from the scene. 2D / 3D system 202 can function as or be included within systems 106 and 108 and 2D / 3D displays 110 and 112 ( FIG. 1 ). Systems 202B and 202N do not depict the same modules as those depicted in system 202A, although each module in system 202A can also be present in systems 202B and 202N.

[0054] System 200 may include multiple cameras, as shown by camera 204. Any type of light-sensing technology can be used to capture images, such as image sensors of the type used in common digital cameras. Cameras 204 can be the same type or different types. Camera locations can be located anywhere on a 3D system, such as system 106. In some implementations, each system 202A, 202B, and 202N includes three or more camera pods, each including a depth camera (e.g., a depth sensor 206 and / or one or more pairs of IR cameras whose content is analyzed using a stereo algorithm to infer a depth image) and one or more color cameras. In some implementations, systems 202A, 202B, and 202N also include one or more witness cameras (not shown), which may capture images used as ground truth images when generating new views and / or for training a neural network, for example.

[0055] System 202A includes a depth sensor 206. In some implementations, depth sensor 206 operates by propagating an IR signal into a scene and detecting a responsive signal. For example, depth sensor 206 can generate and / or detect beams 128A and / or 128B and / or 130A and / or 130B. In some implementations, depth sensor 206 can be used to calculate an occlusion map. System 202A also includes at least one microphone 208 and speaker 210. In some implementations, microphone 208 and speaker 210 may be part of system 106.

[0056] The system 202 further includes a 3D display 212 capable of presenting 3D images. In some implementations, the 3D display 212 may be a standalone display. In some implementations, the 3D display 212 may be integrated into AR glasses, head-mounted display devices, and the like. In some implementations, the 3D display 212 operates using parallax barrier technology. For example, the parallax barrier may include parallel vertical stripes of an essentially non-transparent material (e.g., an opaque film) disposed between the screen and the viewer. Due to the parallax between each of the viewer's eyes, different portions of the screen (e.g., different pixels) are seen by the left and right eyes, respectively. In some implementations, the 3D display 212 operates using lenticular lenses. For example, alternating rows of lenses may be disposed in front of the screen, with the rows directing light from the screen toward the viewer's left and right eyes, respectively.

[0057] System 200 may include a computing system 214 capable of performing specific tasks of data processing, data modeling, data conditioning, and / or data transmission. In some implementations, computing system 214 may also generate images, blend weights, and perform neural processing tasks. In some implementations, computing system 214 is an image processing system. Computing system 214 and / or its components may include some or all of the components described with reference to FIG. 8.

[0058] The computing system 214 includes an image processor 216 that may generate 2D and / or 3D information. For example, the image processor 216 may receive (e.g., obtain) one or more input images 132 and / or view parameters 218 and may generate image content for further processing by the image warp engine 220, the blend weight generator 222, and / or the NN 224. The input images 132 may include captured color (e.g., RGB, YUV, CMYK, CIE, RYB) images.

[0059] The view parameters 218 may include camera parameters associated with capturing a particular input image 132 and / or associated with capturing an image to be generated (e.g., synthesized). In general, the view parameters 218 may represent a camera model approximation. The view parameters 218 may include any or all of the view direction, pose, camera viewpoint, lens distortion, and / or intrinsic and extrinsic parameters of the camera.

[0060] The image processor 216 also includes (and / or generates and / or receives) an occlusion map 226, a depth map 228, a UV map 230, target view parameters 232, a loss function 234, and a mesh proxy geometry 236.

[0061] The occlusion map 226 may encode a signed distance between the surface point determined to be closest to the target viewpoint and the camera capturing the surface. A positive value may indicate that the point is occluded from the view. Thus, the system 200 may configure the blending weight generator 222 (and the NN 224) to avoid using positive distances when determining the blending weights 242 because such occluded image content would not provide accurate reconstruction data when generating new or novel views based on the captured image content. In some implementations, the occlusion map 226 may be used to assess the difference between the depth observed in a particular view and the depth of a geometrically fused model associated with that view.

[0062] The depth map 228 represents one or more images containing information related to the distance of the surface of a particular scene object from a selected viewpoint. In some implementations, the depth map 228 corresponds to each of the three color camera images and / or the depth from the target viewpoint to the nearest surface point determined for each output pixel in the combined (e.g., novel) view.

[0063] The UV map 230 may be generated from visible content in the input image 132. In particular, the UV map 230 represents the projection of the 2D image onto the surface of the 3D model to perform texture mapping to generate features that may be used to generate a composite image (e.g., a novel view).

[0064] The target view parameters 232 represent view parameters for the new synthetic image (i.e., view parameters for generating a virtual view of the target subject). The target view parameters 232 may include image parameters associated with the image to be generated (e.g., synthesized) and / or camera parameters. The target view parameters 232 may include view direction, pose, camera viewpoint, etc.

[0065] The loss function 234 may evaluate the difference between the ground truth image and the predicted image, where the predicted image is predicted based on a combination of both the visible light information captured for the frame, the IR light captured for the frame, and blending weights associated with color and / or depth. The loss function 234 may include functions that describe any or all image errors, image holes, and image misprojection artifacts, etc.

[0066] In some implementations, the loss function 234 may include a reconstruction loss based on a reconstruction difference between a segmented ground truth image mapped to activations of layers in the NN and a segmented predicted image mapped to activations of layers in the NN. The segmented ground truth image may be segmented by a ground truth mask to remove background pixels, and the segmented predicted image may be segmented by a prediction mask to remove background pixels. The prediction mask may be predicted based on a combination of both visible light information captured for the frame and infrared light captured for the frame.

[0067] The mesh proxy geometry 236 is a set of K proxies {P i ,1,···,P i , K 236). For example, a 2D image may be projected onto a 3D proxy model surface to generate a mesh proxy geometry 236. The proxy may function to represent a version of the actual geometry of specific image content. In operation, the system 200 uses the principles of proxy geometry to encode geometric structure using a set of coarse proxy surfaces (e.g., mesh proxy geometry 236) in addition to shape, albedo, and view-dependent effects.

[0068] The image warp engine 220 may be configured to receive one or more input images (e.g., frames, streams) and / or other capture / feature parameter data and generate one or more output images (e.g., frames, streams) that retain the features. The image warp engine 220 may utilize the capture / feature parameter data to reconstruct the input images in some manner. For example, the image warp engine 220 may generate reconstructed candidate color images from the input images, where each pixel in the reconstructed image is a candidate pixel for a new composite image that corresponds to one or more of the input images.

[0069] In some implementations, the image warp engine 220 may perform functions on the input image at the pixel level to preserve small-scale image features. In some implementations, the image warp engine 220 may use nonlinear or linear functions to generate the reconstructed image.

[0070] The blending weight generator 222 combines the blending algorithm 238 and the visibility score 24 0. The blending algorithm 238 may be used to generate blend weights 242. In particular, the blending algorithm 238 may be accessed via the NN 224 to generate the blend weights 242. The blend weights 242 represent values ​​for particular pixels of the images that may be used to contribute aspects of the pixel in a resulting (e.g., final, new) image. The blending algorithm 238 includes a heuristic-based algorithm for calculating blend weights for shading a particular set of depth images and / or fused geometry representing the depth images. The blending algorithm receives the multi-view color images and noisy occlusion cues as inputs to learn output blend weights for the new view (e.g., new composite image). In some implementations, texture and visibility scores 240 (e.g., received from the camera pod(s)) for the target view and input images may also be provided as inputs to the blending algorithm 238.

[0071] The visibility scores 240 may represent the visibility of particular pixels or features of a captured object in an image. Each visibility score 240 may represent a single scalar value to indicate which portion of an image (e.g., pixel, feature, etc.) is visible in a particular view of the input image. For example, if the left-most side of a user's face is not visible in that user's input image, the visibility scores 240 of pixels representing the left-most side of the user's face may be weighted low, while other regions that are visible in the input image and / or are well-captured may be weighted high. The visibility scores may be taken into account when generating blending weights 242 for a new view (e.g., image).

[0072] The neural network 224 includes an embedder network 244 and a generator network 246. The embedder network 244 includes one or more convolutional layers and a downsampling layer. The generator network 246 includes one or more convolutional layers and an upsampling layer.

[0073] The inpainter 254 may generate content (e.g., pixels, regions, etc.) that may be missing from a particular texture or image based on the local neighborhood of pixels surrounding the particular missing content portion. In some implementations, the inpainter 254 may utilize blend weights 242 to determine how to inpaint for a particular pixel, region, etc. The inpainter 254 may utilize output from the NN 224 to predict a particular background / foreground matte for rendering. In some implementations, the inpainter 254 may work in conjunction with the image correction engine 252 to perform pull-push hole-filling. This may be performed on images with regions / pixels of missing depth information that may not result in the output color predicted by the NN 224. The image correction engine 252 may trigger the inpainter to color specific regions / pixels in the image.

[0074] Once the blending weights 242 are determined, the system 214 may provide the weights to the neural renderer 248. The neural renderer 248 may generate an intermediate representation of the object (e.g., a user) and / or scene, for example, utilizing the NN 224 (or another NN). The neural renderer 248 may incorporate view-dependent effects, for example, by modeling the difference between the true appearance (e.g., ground truth) and the diffuse reprojection using an object-specific convolutional network.

[0075]

number

[0076]

number

[0077] In some implementations, the system 214 may perform multi-resolution blending using a multi-resolution blending engine 256. The multi-resolution blending engine 256 may take an image pyramid as input to a convolutional neural network (e.g., NN224 / 414), which generates blend weights at multiple scales with an opacity value associated with each scale. In operation, the multi-resolution blending engine 256 may employ a two-stage, end-to-end trained convolutional network process. The engine 256 may utilize multiple source cameras.

[0078] The synthetic view 250 represents a 3D stereoscopic image of content (e.g., a VR / AR object, a user, a scene, etc.) with appropriate disparity and viewing configuration for both eyes associated with a user accessing a display (e.g., display 212) based at least in part on the calculated blending weights 242 as described herein. At least a portion of the synthetic view 250 may be determined using system 214 based on output from a neural network (e.g., NN 224) each time the user moves their head position while looking at the display and / or each time a particular image changes on the display. In some implementations, the synthetic view 250 represents the user's face and other features of the user surrounding the user's face within a view capturing the user's face. In some implementations, the synthetic view 250 represents the entire field of view captured by one or more cameras associated with telepresence system 202A, for example.

[0079] In some implementations, the processors (not shown) of systems 202 and 214 may include (or be in communication with) a graphics processing unit (GPU). During operation, the processors may communicate with memory, storage, and other processors (e.g., CPUs). ), to facilitate graphics and image generation, the processor may communicate with a GPU to display images on a display device (e.g., display device 212). The CPU and GPU may be connected via a high-speed bus such as PCI, AGP, or PCI-Express. The GPU may be connected to a display via another high-speed interface such as HDMI, DVI, or Display Port. Generally, a GPU may render image content in pixel form. The display device 212 may receive image content from the GPU and display the image content on a display screen.

[0080] 2, additional maps, such as feature maps, may be provided to one or more NNs 224 to generate image content. Feature maps may be generated by analyzing an image to generate features for each pixel of the image. Such features may be used to generate feature and texture maps, which may be provided to blending weight generator 222 and / or NNs 224 to assist in generating blending weights 242.

[0081] 3 is a block diagram illustrating an example of reprojection of an input image to a target camera viewpoint, according to implementations described throughout this disclosure. System 200 may be used, for example, to generate a reprojection of an image to be used as an input image to a neural network. Warping the image may include reprojecting the captured input image 132 to a target camera viewpoint using fused depth (from the depth image). In some implementations, the input image 132 is already in the form of a reprojected image. In some implementations, an image warp engine 220 performs the warping.

[0082] For example, image warp engine 220 may backproject target image point x 302 onto the ray. Then, image warp engine 220 may find point X 304 at distance d from target camera 308. Then, image warp engine 220 may project X to pod image point x' 306, which is distance d' from pod camera 310. Equations [1]-[3] below describe this calculation.

[0083]

number

[0084] The image warp engine 220 may then bilinearly sample the texture camera image at x′ as shown in equations [4] and [5] below.

[0085]

number

[0086] 4 is a block diagram illustrating an example of a flow diagram 400 for using neural blending to generate composite content for rendering on a display, according to implementations described throughout this disclosure. Diagram 400 may generate data (e.g., multi-view color images, noisy occlusion cues, depth data, etc.) that is provided to the blending algorithm via a neural network. The neural network may then learn output blend weights.

[0087] In this example, multiple input images 402 may be acquired (e.g., received). For example, system 202A may capture multiple input images 402 (e.g., image frames, video). The input images 402 may be color images. The input images 402 may also be associated with depth images captured substantially simultaneously with the input images. The depth images may be captured, for example, by an infrared camera.

[0088] The computing system 214 may warp (e.g., reproject) the input image 402 into a reprojected image 404 using the input image color and depth images. For example, the warp engine 220 may reproject the input image 402 into an output view representing a desired new view. In particular, the warp engine 220 may take the color from the input image 402 and warp the color to the output view using the depth view associated with the input image. In general, each input image may be warped to a single reprojected view. Thus, if four input images are taken, the warp engine 220 may generate four reprojected views, each associated with a single input image. The reprojected image 404 serves as candidate colors that may be selected for pixels in the new, synthesized output image. Depth views captured substantially simultaneously with the input image 402 may be used to generate a depth map 406 and an occlusion map 408 (similar to the depth map 228 and the occlusion map 226).

[0089] The reprojected image 404 may be used to generate a weighted sum image 410, which represents a weighted combination of colors for the pixels. The weighted sum image 410 may also take into account a ground truth image 412. The ground truth image 412 may be captured by one or more witness cameras.

[0090] The reprojected image 404, depth map 406, and occlusion map 408 may be provided to NN 414, which is a convolutional neural network having the U-Net geometry shown in Figure 4. Of course, other NNs are possible. In one non-limiting example, the input of NN 414 may include three color RGB images, an occlusion map, and a target view depth map, and may utilize approximately 14 channels.

[0091] Also, in some implementations, a number of view parameters 415 may be provided to the NN 414. The view parameters 415 may relate to a desired new view (e.g., an image). The view parameters 415 may include any or all of the following: view direction, pose, camera viewpoint, lens distortion, and / or intrinsic and extrinsic parameters of the camera (virtual or real).

[0092] The NN 414 may generate blending weights 416 for each reprojected image 404 to determine how to combine the colors of the reprojected image 404 to generate an accurate new output image. The reprojected image 404 may be calculated, for example, by warping the input image 402 into a new view according to the depth image 406. The NN 414 may use the blending weights 416 and the reprojected images 404 to generate a blended texture image 418, for example, by blending at least a portion of the reprojected images 404 together using the blending weights 416. The blended texture image 418 may be used to generate an image associated with each camera pod associated with the input image 402 and the reprojected image 404. In this example, three camera pods were used to capture three color images (e.g., input image 402) and three depth images (e.g., represented by depth map 406). Thus, three corresponding image views are output, as shown by image 420. This allows images 418 and 420 to be used to synthesize a new view, as shown in composite image 422. can.

[0093] In operation, the NN 414 may use blending weights 416 to determine how to combine the reprojected colors associated with the reprojected image 404 to generate an accurate composite image 422. The NN 414 may determine the blending weights by training on a predefined space of output views.

[0094] The network architecture of NN414 may be a deep neural network, a U-Net-shaped network in which all convolutional layers use the same padding value and rectified linear unit activation function. The output may include three reprojected images 404, blending weights 416 for each channel of the camera pod, where the output weights are generated according to Equation [6].

[0095]

number

[0096] Diagram 400 may be implemented taking into account training losses, for example, reconstruction loss, perceptual loss and integrity loss for the blended color image, which may be determined and used to improve the resulting composite image 422.

[0097] In operation, the system 200 may utilize several aspects to generate pixel-by-pixel loss values. For example, given a new view image I of texture camera i, N and the neural blending weights W i can be expressed as shown in equation [7].

[0098]

number

[0099] An invalid target depth mask where none of the inputs have RGB values ​​is I Mask It can be expressed as: In particular, an example of the loss function can be expressed as the following equation [8]:

[0100]

number

[0101] The integrity loss for the network output blending weights for each x,y pixel coordinate can be expressed as shown in Equation

[10] .

[0102]

number

[0103] The occlusion loss on the network can be expressed as shown in Equation

[11] .

[0104]

number

[0105] In some implementations, the NN 414 may be trained based on minimizing an occlusion loss function (i.e., Equation [8]) between the synthetic image 422 generated by the NN 414 and the ground truth image 412 captured by at least one witness camera.

[0106] 5 is a block diagram illustrating an example flow diagram for generating blending weights according to implementations described throughout this disclosure. This example may employ, for example, a convolutional neural network (e.g., a convolutional U-Net) to process pixels of each input view. A multi-layer perceptron (MLP) may be used to generate blending weights to assign each pixel of the proposed synthesized view. The blending weights generated by the MLP can be used to combine features from the input image(s) / view(s).

[0107] In some implementations, generating the blending weights may include using a multi-resolution blending technique, which employs a two-stage, end-to-end trained convolutional network process. This technique utilizes multiple source cameras. For example, system 202A may capture one or more input images (e.g., RGB color images) from each of first camera pod 502, second camera pod 504, and third camera pod 506. Similarly, substantially simultaneously, each of pods 502-504 can capture (or calculate) a depth image corresponding to a particular input image.

[0108] At least three color source input images and at least three source depth images may be provided to convolutional network(s) 508A, 508B, and 508C (e.g., Convolutional U-Net) to generate feature maps incorporating view-dependent information. For example, one or more feature maps (not shown) may represent features of the input images in a feature space. In particular, for each of the input / depth images 502-504, a feature map (e.g., feature maps 510A, 510B, and 510C) may be generated using extracted features of the image. In some implementations, the input image may include two color source images and one depth image. In such an example, the system 500 may reproject each of the two color input images to an output view using a single depth image.

[0109] The feature maps 510A-510C may be used to generate UV maps 512A, 512B, and 512C. For example, the UV maps 512A-C may be generated from visible content in the input images 502-504 using the feature maps 510A-510C. The UV maps 512A-512C represent the projection of the 2D images onto the 3D model surface to perform texture mapping to generate features that can be used to generate a composite image (e.g., a novel view). The output neural texture remains in source camera image coordinates.

[0110] Each of the feature maps 510A-510C may be sampled with a respective UV map 512A-512C and witness camera parameters 514. For example, the system 500 may use a witness camera as a target camera for generating a synthesized new image. The witness (e.g., target) camera parameters 514 may be predefined. Each of the sampled feature maps 510A-510C and UV maps 512A-C may be used with the parameters 514 and sampled with an occlusion map and a depth map 516. The sampling may be performed using pre-computed fusion geometry (e.g., mesh proxy geometry 236). The neural texture may include a differentiable sampling layer that warps each neural texture using the UV maps 512A-512C.

[0111] The sampled content may be used by a pixel-wise multi-layer perceptron (MLP) NN 518 to generate an occlusion map, depth map, etc. of sampled features from all source camera views. From the maps, the MLP 518 may generate a set of blending weights 520. For example, a pixel-wise MLP 518 map may include sampled features from any number of source camera views that can be used to generate a set of blending weights 520. Such blending weights 520 may be used to generate a composite image.

[0112] In some implementations, the processes described herein may incorporate multi-resolution blending, which may be performed by the multi-resolution blending engine 256, for example, and may take an image pyramid as input to a convolutional neural network (e.g., NN224 / 414), which generates blending weights at multiple scales with an opacity value associated with each scale.

[0113] The output blending weights at each scale are used to construct an output color image using the input reprojected color image at that scale, forming an output image pyramid. Each level of this pyramid is then weighted by its associated opacity value and upsampled to the original scale. The resulting set of images is then summed to construct the final output image. This is advantageous because if the input reprojected image has small holes (due to missing geometry), the downscaling and upscaling processes can fill the missing areas with neighboring pixel values. This procedure may also produce softer silhouettes that are more visually appealing than traditional blending methods.

[0114] In some implementations, the input pyramid can be constructed by downsampling the bilinear reprojected colors of the reprojected image, un-pre-multiplying them by a downsampled effective depth mask (e.g., a map), upsampling them back to a predefined (e.g., original) resolution, and un-pre-multiplying them by the upsampled effective depth mask. For each layer, the flow diagram may add an output layer decoder (for blending weights and alpha), upsample to a predefined (e.g., original) resolution, adjust the additional background alpha at the highest resolution, normalize the alpha with a softmax function, and blend it with the reprojected colors and background.

[0115] The multi-resolution blending method employs two stages of trained end-to-end convolutional network processing. For each stage, the multi-resolution blending method may add an output layer decoder (e.g., for blending weights and alpha loss). In this technique, an RGB image may be calculated, a loss may be added, alpha may be multiplied, and concatenated to determine a candidate RGB image. The candidate RGB image may be upsampled. An output image (e.g., a new view / sync image) may be generated using the upsampled candidate image that takes the loss into account.

[0116] In operation, the technology utilizes multiple source cameras. For example, system 202A may capture one or more input images (e.g., RGB color images) from each of first camera pod 502, second camera pod 504, and third camera pod 506. Similarly, substantially simultaneously, pods 502-504 may each capture a depth image corresponding to a particular input image.

[0117] Multi-resolution blending can use the same 3D point on the scene map for the same point location in the feature map regardless of how the output viewpoint moves, ensuring that the output contains the same blend weights for that point location since no 2D convolutions are performed and the input features are fixed.

[0118] 6 is a flowchart illustrating an example of a process 600 for generating synthetic content using neural blending, according to implementations described throughout this disclosure. While process 600 is described with respect to implementations of systems 100 and / or 200 and systems 500 and / or 800 of FIGS. 1 and 2, it will be understood that the method may be implemented by systems having other configurations. In general, one or more processors and memory on system 202 and / or computing system 214 may be used to implement process 600.

[0119] At a high level, process 600 may utilize a color input image, a depth image corresponding to the input image, and view parameters associated with a desired new view corresponding to at least a portion of the content in the input image. Process 600 may provide the above elements or versions of the above elements to a neural network to receive blending weights for determining the specific pixel colors and depths of the desired new view. The view may be used together with the blending weights to generate a new output image.

[0120] At block 602, process 600 may include receiving multiple input images. For example, system 202A (or other image processing system) may capture input images from two or more camera pods using a camera (e.g., camera 204). Typically, the multiple input images are color images captured according to predefined view parameters. However, in some implementations, the multiple input images may be gradient images of a single color (e.g., sepia, grayscale, or other gradient colors). The predefined view parameters may include camera parameters associated with the capture of a particular input image 132 (e.g., input image 402) and / or camera parameters associated with the capture of an image to be generated (e.g., synthesized). In some implementations, the view parameters may include view direction, pose, camera viewpoint, lens distortion, and / or any or all of the camera's intrinsic and extrinsic parameters. In some implementations, the multiple input images may include multiple target subjects captured within the frames of the images. The target subjects may include a user, a background, a foreground, a physical object, a virtual object, a gesture, a hairstyle, a wearable device, etc.

[0121] At block 604, process 600 may include receiving multiple depth images associated with a target object in at least one of the multiple input images. For example, substantially simultaneously with the capture of an input image (e.g., RGB color image 136), system 202A may capture depth image 138. The depth image may capture a target object also captured in one or more of the multiple input images. Each depth image may include a depth map (e.g., map 228) associated with at least one camera 204 that captured at least one of the multiple input images 132, at least one occlusion map 226, and a depth map associated (e.g., via target view parameters 232) with a ground truth image captured by at least one witness camera at a time corresponding to the capture of at least one of the multiple input images. In short, system 200 may consider the depth of the input image and the depth of the desired target view (or other determined target view) of the witness camera when generating blending weights 242 for the target view.

[0122] At block 606, process 600 may include receiving a plurality of view parameters for generating a virtual view of the target subject. For example, the view parameters may relate to a desired new view (e.g., a new synthetic image for a new (e.g., virtual) view not previously captured by a camera). The view parameters may include, for example, target parameters of a witness camera capturing content substantially simultaneously with color image 136 and depth image 138. The view parameters may include predefined lens parameters, gaze direction, pose, and certain intrinsic and / or extrinsic parameters of a camera configured to capture the new view.

[0123] At block 608, process 600 may include generating multiple warped images based on the multiple input images, the multiple view parameters, and at least one of the multiple depth images. For example, image warp engine 220 may generate the warped image using input image 132 by reprojecting input image 132 onto a reprojected version of input image 132. Warping may be performed using depth information (e.g., either individual depth images or a geometric consensus surface) to determine the projection of input colors of input image 132 onto a new view. Warping may generate a reprojected image (e.g., image 404) by taking colors from one or more original input views and manipulating the colors of the new view (e.g., image) using depth images (e.g., depth map 406 and occlusion map 408). Each input image may be used to generate a separate reprojection. The reprojected image (e.g., image 404) may represent pixels of candidate colors that may be used in the new composite image.

[0124] In some implementations, the process 600 may also include generating multiple warped images based on the multiple input images, the multiple view parameters, and at least one of the multiple depth images by determining candidate projections of colors associated with the multiple input images 402 onto uncaptured views (i.e., new views / images, virtual views / images) using at least one of the multiple depth images (e.g., the depth map 406 and the occlusion map 408). The uncaptured views may include at least a portion of image features of at least one of the multiple input images. For example, if the input image includes an object, the uncaptured views may take into account at least a portion, color, pixels, etc. of the object.

[0125] At block 610, process 600 may include receiving blending weights 416 from a neural network (e.g., NN 224, NN 414, NN 508A-C) for assigning colors to pixels of a virtual view (e.g., invisible image / uncaptured view) of a target subject (e.g., user 104′). In some implementations, the target subject may include or be based on at least one element captured in at least one frame of multiple input images 402. The blending weights 416 may be received in response to providing multiple depth images (e.g., depth image 138 and / or depth map 406 and / or occlusion map 408), multiple view parameters 415, and multiple warped images (e.g., reprojected images 404) to the NN 414. The NN 414 may generate the blending weights 416 to indicate a probabilistic manner in which to combine colors in the reprojected images 404 to provide a realistic output image that is likely to realistically represent the target subject. In some implementations, blending weights 416 are configured to assign a blending color to each pixel of a virtual view (i.e., a new and / or unseen and / or previously uncaptured view) that results in the assignment of such blending color to an output composite image (e.g., composite image 422). For example, blending weights 416 are used to blend at least some of reprojected images 404 together.

[0126] At block 612, process 600 may include generating a composite image according to the view parameters based on the blending weights and the virtual view. The composite image 422 may represent an image captured using parameters for an uncaptured view (e.g., not captured by a physical camera, generated as a virtual view from a virtual camera or a physical camera, etc.), or it may represent an unseen view (not captured by any camera in the imaging system but instead synthesized). The composite image 422 may be generated for and / or during a three-dimensional (e.g., telepresence) video conference. For example, the composite image 422 may be generated in real time during the video conference to provide an error-corrected, accurate image of a user or content being captured by a camera associated with the video conference. In some implementations, the composite image 422 represents a new view generated for the three-dimensional video conference. In some implementations, the composite image represents an uncaptured view of a target subject generated for the three-dimensional video conference.

[0127] In operation, blending weights are applied to pixels in the virtual view according to the view parameters. The resulting virtual view may include pixel colors generated using the blended weights for the target object. The colorized image of the virtual view may be used to generate a synthetic view according to the view parameters associated with the virtual camera, for example.

[0128] In some implementations, process 600 may additionally perform a geometric fusion operation. In some implementations, process 600 may perform a geometric fusion operation instead of providing input images to individual depth images. For example, process 600 may reconstruct a consensus surface (e.g., a geometric proxy) using a geometric fusion operation on multiple depth images to generate a geometric fusion model.

[0129] The geometrically fused model may be used to replace multiple views of depth image data (e.g., captured depth views of the image content) with updated (e.g., calculated) views of the depth image data. The updated depth views may be generated as views of the image content that include depth data from the captured depth views and further include image and / or depth information from each of the other available captured depth views of the image content. One or more of the updated depth views may be used by the NN 414 to synthesize additional (and new) views of the object using additional (and new) blending weights, for example, by utilizing the geometrically fused depth image data and image and / or depth information associated with multiple other views of the object. The depth image data may be fused using any number of algorithms to replace each (input) depth view with a new depth view that incorporates depth data information from multiple other depth views. In some implementations, the geometrically fused model can be used by system 200 to generate depth data (e.g., a depth map) that can be used to infer occlusion to correct for such occlusion losses.

[0130] Process 600 may then generate a plurality of reprojected images based on the plurality of input images and the consensus surface used to generate the geometrically fused depth image data, and provide the geometrically fused depth image data (along with the plurality of view parameters 415 and the plurality of reprojected images 404) to NN 414. In response, process 600 may include receiving from NN 414 blending weights 416 and / or additional blending weights generated using the consensus surface depth image data for assigning colors to pixels in composite image 422.

[0131] In some implementations, the process 600 further includes a step of configuring the NN 414 to perform geometrically fused The process 400 may include providing a difference between the depth of the model and the depth observed in the multiple depth images. The depth difference may be used, for example, to correct for detected occlusion in the synthetic image 422. In some implementations, the NN 414 may be trained based on minimizing an occlusion loss function between the synthetic image generated by the NN 414 and ground truth images 412 captured by at least one witness camera (e.g., associated with system 202A), as described in detail with respect to FIG. 4. In some implementations, the process 400 may be performed using a single depth image rather than multiple depth images.

[0132] In some implementations, NN 414 is further configured to perform multi-resolution blending to assign pixel colors to pixels in the composite image. In operation, multi-resolution blending may trigger the provision of an image pyramid as input to NN 414 to receive from NN 414 multi-resolution blend weights for multiple scales (e.g., additional blend weights 520) and further receive an opacity value associated with each scale.

[0133] FIG. 7 illustrates an example of a computing device 700 and a mobile computer device 750 usable with the described techniques. Computing device 700 includes a processor 702, memory 704, a storage device 706, a high-speed interface 708 connecting memory 704 and a high-speed expansion port 710, and a low-speed interface 712 connecting a low-speed bus 714 and storage device 706. Components 702, 704, 706, 708, 710, and 712 are interconnected using various buses and may be mounted on a common motherboard or in other suitable manners. Processor 702 is capable of processing instructions for execution within computing device 700, including instructions stored in memory 704 or storage device 706, for displaying graphical information for a GUI on an external input / output device, such as a display 716 coupled to high-speed interface 708. In some embodiments, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Additionally, multiple computing devices 700 may be connected with each device providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multiprocessor system).

[0134] Memory 704 stores information within computing device 700. In some embodiments, memory 704 is one or more volatile memory units. In other embodiments, memory 704 is one or more non-volatile memory units. Memory 704 may also be any form of computer-readable medium, such as a magnetic or optical disk.

[0135] The storage device 706 can provide mass storage for the computing device 700. In one embodiment, the storage device 706 can be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or devices in a storage area network or other configuration. A computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described herein. The information carrier is a computer-readable or machine-readable medium, such as memory 704, the storage device 706, or memory on the processor 702.

[0136] The high-speed controller 708 provides bandwidth-intensive The high-speed controller 708 manages operations, and the low-speed controller 712 manages less bandwidth-intensive operations. This allocation of functionality is merely exemplary. In one embodiment, the high-speed controller 708 is coupled to the memory 704, the display 716 (e.g., via a graphics processor or accelerator), and to a high-speed expansion port 710 that may accept various expansion cards (not shown). The low-speed controller 712 may be coupled to the storage device 706 and the low-speed expansion port 714. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or to a networking device, such as a switch or router, for example, via a network adapter.

[0137] Computing device 700 can be implemented in many different forms as shown. For example, it can be implemented as a standard server 720, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 724. It can also be implemented in a personal computer, such as a laptop computer 722. Alternatively, components from computing device 700 can be combined with other components in a mobile device (not shown), such as device 750. Each such device may include one or more of computing devices 700, 750, and an entire system may be formed of multiple computing devices 700, 750 in communication with each other.

[0138] Computing device 750 includes, among other components, a processor 752, memory 764, input / output devices such as a display 754, a communications interface 766, and a transceiver 768. Device 750 may also be provided with a storage device such as a microdrive or other device to provide additional storage. Each of components 750, 752, 764, 754, 766, and 768 are interconnected using various buses, and multiple of these components may be mounted on a common motherboard or otherwise as appropriate.

[0139] The processor 752 can execute instructions within the computing device 750, including instructions stored in the memory 764. The processor may be implemented as a chipset of chips including separate analog and digital processors. The processor may provide, for example, control of a user interface, applications executed by the device 750, and coordination of other components of the device 750, such as wireless communication by the device 750.

[0140] Processor 752 may communicate with a user via a control interface 758 and a display interface 756 coupled to a display 754. Display 754 may be, for example, a TFT LCD (thin film transistor liquid crystal display) or an OLED (organic light-emitting diode) display, or any other display technology. Display interface 756 may include suitable circuitry for driving display 754 to present graphical and other information to the user. Control interface 758 may receive commands from the user and translate these commands for transmission to processor 752. Additionally, an external interface 762 may communicate with processor 752 to enable near-field communication of device 750 with other devices. External interface 762 may provide, for example, wired or wireless communication, and multiple interfaces may be used in some embodiments.

[0141] The memory 764 stores information within the computing device 750. 4 may be embodied as one or more of one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Expansion memory 784 may also be provided in device 750 and connected via expansion interface 782, which may include, for example, a SIMM (Single In-Line Memory Module) guard interface. Such expansion memory 784 may provide additional storage space for device 750 or may also store applications or other information for device 750. Specifically, expansion memory 784 may include instructions for performing or supplementing the processes described above, and may also include secure information. Thus, expansion memory 784 may, for example, be a security module for device 750, programmable with instructions that enable secure use of device 750. Additionally, secure applications may be provided via a SIMM card along with additional information, such as by placing identifying information on the SIMM card in an unhackable manner.

[0142] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 764, expansion memory 784, or memory on processor 752, which may be received, for example, via transceiver 768 or external interface 762.

[0143] Device 750 can communicate wirelessly via a communication interface 766, which may include digital signal processing circuitry as needed. Communication interface 766 can provide communications under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communications may occur, for example, via a radio frequency transceiver 768. Additionally, short-range communications may occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, a GPS (Global Positioning System) receiver module 770 can provide additional wireless data related to navigation and location to device 750, which can be used as appropriate by applications executing on device 750.

[0144] Device 750 can also communicate audibly using a voice codec 760, which can receive spoken information from a user and convert it into usable digital information. Voice codec 760 may also generate audible sounds for the user, such as via a speaker in a handset of device 750. Such sounds may include sounds from a voice telephone call, may include recorded sounds (e.g., voice messages, music files, etc.), and may include sounds generated by applications running on device 750.

[0145] The computing device 750 can take many different forms, as shown in the figure, such as a mobile phone 780, or as part of a smartphone 783, personal digital assistant, or other similar mobile device.

[0146] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuits, specially designed application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from the storage system and to transmit data and instructions to the storage system, at least one input device, and may include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one output device.

[0147] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with a user. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0149] The systems and techniques described herein may be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a client computer having a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.

[0150] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0151] In some embodiments, the computing device illustrated in Figure 7 may include sensors that interface with a virtual reality headset (VR headset / HMD device 790). For example, one or more sensors included in computing device 750 or other computing devices illustrated in Figure 7 may provide input to the VR headset 790, or generally to the VR space. The sensors may include, but are not limited to, a touchscreen, an accelerometer, a gyroscope, a pressure sensor, a biometric sensor, a temperature sensor, a humidity sensor, and an ambient light sensor. Using sensors, computing device 750 can determine an absolute position and / or detected rotation of the computing device in VR space, which can then be used as input to the VR space. For example, computing device 750 may be incorporated into the VR space as a virtual object such as a controller, laser pointer, keyboard, weapon, etc. The user's positioning of the computing device / virtual object when incorporated into the VR space allows the user to position the computing device to view the virtual object in a particular manner in the VR space.

[0152] In some embodiments, one or more input devices included in or connected to computing device 750 can be used as input to the VR space. Input devices may include, but are not limited to, a touchscreen, a keyboard, one or more buttons, a trackpad, a touchpad, a pointing device, a mouse, a trackball, a joystick, a camera, a microphone, earphones or buds with input capabilities, a game controller, or other connectable input devices. When the computing device is integrated into the VR space, a user interacting with the input devices included in computing device 750 can cause certain actions to occur in the VR space.

[0153] In some embodiments, one or more output devices included in computing device 750 can provide output and / or feedback to a user of VR headset 790 in the VR space. The output and feedback can be visual, tactical, or audio. The output and / or feedback can include, but is not limited to, rendering the VR space or virtual environment, vibration, turning one or more lights or strobes on and off or blinking and / or flashing, sounding an alarm, playing a chime, playing a song, and playing an audio file. Output devices can include, but are not limited to, vibration motors, vibration coils, piezoelectric devices, electrostatic devices, light-emitting diodes (LEDs), strobes, and speakers.

[0154] In some embodiments, the computing device 750 can be placed within a VR headset 790 to create a VR system. The VR headset 790 can include one or more positioning elements that allow the computing device 750, such as the smartphone 783, to be placed in an appropriate position within the VR headset 790. In such embodiments, the display of the smartphone 783 can render stereoscopic images that represent the VR space or virtual environment.

[0155] In some embodiments, the computing device 750 may appear as another object in the computer-generated 3D environment. A user's interactions with the computing device 750 (e.g., rotating, shaking, touching the touchscreen, swiping a finger across the touchscreen) can be interpreted as interactions with objects in the VR space. By way of example only, the computing device may be a laser pointer. In such an example, the computing device 750 appears as a virtual laser pointer in the computer-generated 3D environment. As the user manipulates the computing device 750, the user in the VR space can see the movement of the laser pointer. The user receives feedback on the computing device 750 or on the VR headset 790 from their interactions with the computing device 750 in the VR environment.

[0156] In some embodiments, computing device 750 may include a touchscreen. For example, a user may want to make things happen in the VR space that happen on the touchscreen. In a virtual world, a user can interact with the touchscreen in certain ways that can mimic the experience of a virtual world. For example, a user may use a pinch-type motion to zoom content displayed on the touchscreen. This pinch-type motion on the touchscreen can zoom information provided in the VR space. In another example, a computing device may be rendered as a virtual book in a computer-generated 3D environment. In the VR space, pages of the book are displayed, and swiping of the user's finger across the touchscreen can be interpreted as turning / flipping pages of the virtual book. In addition to seeing the content of the pages change as each page is turned / flipped, the user may be provided with audio feedback, such as the sound of pages turning in a book.

[0157] In some embodiments, one or more input devices (e.g., a mouse, a keyboard) in addition to a computing device can be rendered in the computer-generated 3D environment, and the rendered input devices (e.g., a rendered mouse, a rendered keyboard) can be used as rendered in the VR space to control objects in the VR space.

[0158] Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 750 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the disclosed embodiments.

[0159] Additionally, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. Furthermore, other steps may be provided or removed from the described flows, and other components may be added or removed from the described systems. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. 1. A computer-implemented method comprising: generating blending weights for assigning colors to pixels of a virtual view of a target subject by providing an input image and view parameters to a neural network, the virtual view including at least one of a view not visible in the input image and an uncaptured view, and a depth image including an occlusion map, the neural network not using occluded image content when determining the blending weights, the computer-implemented method further comprising: generating a composite image based on the blending weights, the virtual view, and a plurality of reprojected images based on the input image and the view parameters.

2. 2. The computer-implemented method of claim 1, wherein the neural network is trained to generate the blending weights to minimize an occlusion loss function between the synthetic image and a ground truth image captured by at least one witness camera.

3. 3. The computer-implemented method of claim 1, wherein the composite image is an uncaptured view of the target subject generated for three-dimensional video conferencing.

4. The computer-implemented method of claim 3 , wherein the blending weights are used to generate pixel colors of the composite image.

5. The computer-implemented method of any one of claims 1 to 4, wherein generating the plurality of reprojected images comprises reprojecting the input image onto a target viewpoint.

6. The computer-implemented method of any one of claims 1 to 5, wherein the view parameters represent a camera model approximation.

7. The computer-implemented method of any one of claims 1 to 6, wherein the view parameters include at least one of a view direction, a pose, and a camera viewpoint associated with a target viewpoint.

8. The computer-implemented method of any one of claims 1 to 7, wherein the blending weights are used to combine features from the input images.

9. 9. The computer-implemented method of claim 1, wherein the neural network is a two-stage end-to-end convolutional neural network.

10. 1. An image processing system, comprising: at least one processing device; a memory storing instructions that, when executed, cause the image processing system to perform operations, the operations including: generating blending weights for assigning colors to pixels of a virtual view of a target object by providing an input image and view parameters to a neural network, the virtual view including at least one of a view not visible in the input image and an uncaptured view, a depth image including an occlusion map, the neural network not using occluded image content when determining the blending weights, the operations further comprising: generating a composite image based on the blending weights, the virtual view, and a plurality of reprojected images based on the input image and the view parameters.

11. 11. The image processing system of claim 10, wherein the neural network is trained to generate the blending weights to minimize an occlusion loss function between the synthetic image and a ground truth image captured by at least one witness camera.

12. 12. The image processing system of claim 10 or 11, wherein the composite image is an uncaptured view of the target subject generated for 3D video conferencing.

13. The image processing system of claim 12 , wherein the blending weights are used to generate pixel colors of the composite image.

14. The image processing system of any one of claims 10 to 13, wherein generating the plurality of reprojected images comprises reprojecting the input image onto a target viewpoint.

15. The image processing system of any one of claims 10 to 14, wherein the view parameters represent a camera model approximation.

16. The image processing system of any one of claims 10 to 15, wherein the view parameters include at least one of a view direction, a pose, and a camera viewpoint associated with a target viewpoint.

17. The image processing system of any one of claims 10 to 16, wherein the blending weights are used to combine features from the input images.

18. The image processing system of any one of claims 10 to 17, wherein the neural network is a two-stage end-to-end convolutional neural network.

19. A program having instructions that, when executed by a processor, cause a computing device to: and causing the computing device to generate blending weights for assigning colors to pixels of a virtual view of a target subject based on a depth image and view parameters associated with the input image, the virtual view including at least one of a view not visible in the input image and an uncaptured view, the depth image including an occlusion map, the neural network not using occluded image content when determining the blending weights; and wherein the instructions, when executed by the processor, cause the computing device to further: generating a composite image based on the blending weights, the virtual view, and a plurality of reprojected images based on the input image and the view parameters.

20. 20. The program of claim 19, wherein the neural network is trained to generate the blending weights to minimize an occlusion loss function between the synthetic image and a ground truth image captured by at least one witness camera.

21. A program that causes a computer to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Intermediate image synthesis and multiview data signal extraction

    JP2012504805A

  • New view composition using deep-layer convolutional neural network

    JP2018081672A

  • System and method for depth-guided filtering in a video conference environment

    US20140160239A1

  • Free-viewpoint photorealistic view synthesis from casually captured video

    US20200228774A1

  • Virtual visual point image generating method and 3-d image display method and device

    WO2004114224A1