Single image 3D photography with soft layering and depth-aware repair
By using soft layering and depth-aware inpainting techniques, the scene is decomposed into foreground and background layers, generating soft foreground visibility maps and background demasking masks. This solves the problem of inaccurate 3D effects in existing technologies and achieves accurate modeling of thin objects and good 3D effect generation.
Patent Information
- Application Number
- CN202510830627.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-05
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies struggle to generate high-quality 3D effects, especially inaccurate representations of thin objects such as hair and fur. Furthermore, the difficulty in obtaining multi-view datasets leads to poor performance of methods in scenes outside the training data distribution.
By using soft layering and depth-aware inpainting techniques, the scene is decomposed into a foreground layer and a background layer. The depth image is used to generate a soft foreground visibility map and a background demasking mask, which repairs the demasking areas of the monocular image and the depth image, generating a modified 3D image.
It achieves accurate modeling of thin objects, generates visually consistent 3D effects, and performs well in scenes outside the training data distribution.
Smart Images

Figure CN120894263A_ABST
Abstract
Description
[0001] Divisional Statement
[0002] This application is a divisional application of Chinese Patent Application No. 202180063773.2, filed on August 5, 2021, which has a filing date of August 5, 2021. TECHNICAL FIELD
[0003] The present disclosure relates to single image 3D photography with soft delamination and depth-aware inpainting. BACKGROUND
[0004] Machine learning models can be used to process various types of data including images to generate various desired outputs. Improvements in machine learning models allow the models to perform processing of data more quickly, to utilize fewer computational resources for processing, and / or to generate outputs with relatively higher quality. SUMMARY
[0005] A three-dimensional (3D) photograph system can be configured to generate, based on a monocular image, a 3D viewing effect / experience that simulates a different viewpoint of a scene represented by the monocular image. Specifically, a depth image can be generated based on the monocular image. A soft foreground visibility map can be generated based on the depth image and can indicate a transparency of different portions of the monocular image to a background layer. The monocular image can be considered to form a foreground layer. A soft background occlusion mask can be generated based on the depth image and can indicate a likelihood of different background regions of the monocular image being occluded due to a change in viewpoint. An inpainting model can use the soft background occlusion mask to inpaint the monocular image and the depth image for the occluded background regions, thereby generating the background layer. The foreground layer and the background layer can be represented in 3D, and these 3D representations can be projected from the new viewpoint to generate a new foreground image and a new background image. The new foreground image and the new background image can be combined according to the soft foreground visibility map as re-projected into the new viewpoint, thereby generating a modified image with the new viewpoint of the scene.
[0006] In a first example embodiment, a method includes obtaining a monocular image having an initial viewpoint, and determining a depth image including a plurality of pixels based on the monocular image. Each respective pixel of the depth image can have a corresponding depth value. The method further includes determining, for each respective pixel of the depth image, a corresponding depth gradient associated with the respective pixel of the depth image, and determining a foreground visibility map including, for each respective pixel of the depth image, a visibility value inversely proportional to the corresponding depth gradient. The method additionally includes determining, based on the depth image, a background removal occlusion mask including, for each respective pixel of the depth image, a removal occlusion value indicating a likelihood that a corresponding pixel of the monocular image will be removed occluded by a change in the initial viewpoint. The method still additionally includes (i) generating a repaired image by repairing portions of the monocular image according to the background removal occlusion mask using a repair model, and (ii) generating a repaired depth image by repairing portions of the depth image according to the background removal occlusion mask using the repair model. The method further includes (i) generating a first three-dimensional (3D) representation of the monocular image based on the depth image, and (ii) generating a second 3D representation of the repaired image based on the repaired depth image. The method still further includes generating a modified image having an adjusted viewpoint different from the initial viewpoint by combining the first 3D representation with the second 3D representation according to the foreground visibility map.
[0007] In a second example embodiment, a system can include a processor and a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations according to the first example embodiment.
[0008] In a third example embodiment, a non-transitory computer-readable medium can have stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations according to the first example embodiment.
[0009] In a fourth example embodiment, a system can include various means for performing each of the operations of the first example embodiment.
[0010] These and other embodiments, aspects, advantages, and alternatives will become apparent to those of ordinary skill in the art after reviewing this description. Further, the Summary and the other descriptions and drawings provided herein are intended to be illustrative only and not restrictive. Numerous variations are possible that will fall within the scope of the disclosure as defined by the following claims. For example, elements and / or process steps can be rearranged, combined, distributed, eliminated, or altered, entirely or in part, without rewarding from the scope of the disclosure as claimed. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 A computing device is illustrated in accordance with examples described herein.
[0012] Figure 2 A computing system is illustrated in accordance with examples described herein.
[0013] Figure 3A A 3D photo system is illustrated in accordance with examples described herein.
[0014] Figure 3B An image and a depth image are illustrated in accordance with examples described herein.
[0015] Figure 3C A visibility map and a background removed occlusion mask are illustrated in accordance with examples described herein.
[0016] Figure 3D A repaired image, a repaired depth image, and a modified image are illustrated in accordance with examples described herein.
[0017] Figure 4A A soft foreground visibility model is illustrated in accordance with examples described herein.
[0018] Figure 4B An image and a depth image are illustrated in accordance with examples described herein.
[0019] Figure 4C A depth based foreground visibility map and a foreground alpha mask are illustrated in accordance with examples described herein.
[0020] Figure 4D A mask based foreground visibility map and a foreground visibility map are illustrated in accordance with examples described herein.
[0021] Figure 5A A training system is illustrated in accordance with examples described herein.
[0022] Figure 5B A background occlusion mask is illustrated in accordance with examples described herein.
[0023] Figure 6 A flowchart including examples described herein.
[0024] Figure 7A , Figure 7B and Figure 7C A table of performance metrics including examples described herein. DETAILED DESCRIPTION
[0025] Example methods, devices, and systems are described herein. It should be understood that the words “example” and “exemplary” are used herein to mean “serving as an example, instance, or illustration.” Any implementation described herein as “example,” “exemplary,” and / or “illustrative” is not necessarily to be construed as preferred or advantageous over other implementations. Thus, the examples are presented for illustrative purposes and do not have to accomplish any particular advantage over other implementations.
[0026] Accordingly, the example embodiments described herein are not meant to be limiting. It will be readily understood that the aspects of the present disclosure as generally described herein and illustrated in the figures can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0027] Moreover, the features illustrated in each figure are for the purpose of example and can be combined with each other where appropriate. One of ordinary skill in the art will further appreciate that the figures are not necessarily drawn to scale and that, where appropriate, elements can have been exaggerated, simplified, removed or otherwise not drawn necessarily to scale.
[0028] Additionally, any listing of elements, blocks, or steps in a method or claim is for clarity only. Thus, such listing should not be interpreted as requiring or implying that those elements, blocks, or steps be performed in the particular order or sequence in which they are recited. The drawings are not necessarily to scale.
[0029] SUMMARY
[0030] Monocular images representing a scene from an initial viewpoint can be used to generate modified images representing the scene from an adjusted viewpoint. By sequentially viewing the monocular images and / or the modified images, a 3D effect can be simulated, creating a richer and / or more interactive viewing experience. In particular, the viewing experience can be created using a single monocular image that is generated without additional optical hardware beyond the optical hardware of a monoscopic camera.
[0031] Some techniques for generating images with modified viewpoints can rely on decomposing a scene into two or more layers based on hard discontinuities. A hard discontinuity can define a sharp and / or discrete (e.g., binary) transition between two layers, such as a foreground and a background. Hard discontinuities can not allow for accurate modeling of some appearance effects, such as very thin objects (e.g., hair), resulting in visually flawed or deficient 3D effects. Other techniques for generating images with modified viewpoints can rely on a training dataset that includes multi-view images that ground-truth different viewpoints. However, it can be difficult to obtain multi-view datasets that represent a wide range of scenes, and these approaches can therefore perform poorly for out-of-distribution scenes that are not adequately represented by the training data distribution.
[0032] Accordingly, provided herein is a 3D photo system that relies on soft layering formation and depth-aware inpainting to decompose a given scene into a foreground layer and a background layer. These layers can be combined to generate a modified image from a new viewpoint. In particular, a depth model can be used to generate a depth image corresponding to a monocular image, and the depth image can be used to generate a soft foreground visibility map and / or a soft background disocclusion mask.
[0033] The soft foreground visibility map can be generated based on depth gradients of the depth image and can indicate how different portions of the foreground layer are transparent and thus allow corresponding regions of the background layer to be seen through the foreground layer. The foreground visibility map can be soft in that it is generated by a soft foreground visibility function that is continuous and smooth at least along a range of input values of the depth gradients. The foreground layer can be defined by the monocular image and the foreground visibility map. In some implementations, the soft foreground visibility map can also be based on a foreground alpha matte, which can improve the representation of thin and / or high frequency features such as hair and / or fur.
[0034] The soft background disocclusion mask can be generated based on the depth image to quantify how different pixels of the monocular image are likely to be disoccluded due to a change in viewpoint. The background disocclusion mask can be soft in that it is generated by a soft background disocclusion function that is continuous and smooth at least along a range of input values of the depth image. The terms “map” and “mask” can be used interchangeably to refer to grayscale and / or binary images.
[0035] The inpainting model can use the soft background disocclusion mask to inpaint portions of the monocular image and the depth image that are likely to be disoccluded, thereby generating the background layer. In particular, the same inpainting model can be trained to inpaint the monocular image and the corresponding depth image such that the inpainted pixel values are based on both the color / intensity values of the monocular image and the depth values of the depth image. The inpainting model can be trained on both image and depth data such that it is trained to perform depth-aware inpainting. For example, the inpainting model can be trained to inpaint the monocular image and the corresponding depth image concurrently. Thus, the inpainting model can be configured to inpaint disoccluded background regions by borrowing information from other background regions rather than from foreground regions, thereby resulting in a visually consistent background layer. Additionally, the inpainting model can be configured to inpaint the monocular image and the depth image in a single iteration (e.g., a single pass through) of the inpainting model rather than using multiple iterations to pass through.
[0036] The respective 3D representations of the foreground layer and the background layer can be generated based on the depth image and the inpainted depth image, respectively. The viewpoint from which these 3D representations are observed can be adjusted, and the 3D representations can be projected to generate foreground and background images having the adjusted viewpoint. The foreground and background images can then be combined according to the modified foreground visibility map, which also has the adjusted viewpoint. Thus, the occlusion-removed regions in the foreground layer can be filled based on corresponding regions of the background layer, where the transition between the foreground and the background includes a blend of both the foreground layer and the background layer.
[0037] Example computing devices and systems
[0038] Figure 1 An example computing device 100 is illustrated. The computing device 100 is shown in the form factor of a mobile phone. However, the computing device 100 can alternatively be implemented as a laptop computer, a tablet computer, and / or a wearable computing device, among other possibilities. The computing device 100 can include various elements, such as a body 102, a display 106, and buttons 108 and 110. The computing device 100 can further include one or more cameras, such as a front-facing camera 104 and a rear-facing camera 112.
[0039] The front-facing camera 104 can be positioned on a side of the body 102 that is generally facing the user when in operation (e.g., on the same side as the display 106). The rear-facing camera 112 can be positioned on a side of the body 102 opposite the front-facing camera 104. The designation of the cameras as front-facing and rear-facing is arbitrary, and the computing device 100 can include multiple cameras positioned on various sides of the body 102.
[0040] The display 106 can represent a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light emitting diode (OLED) display, or any other type of display known in the art. In some examples, the display 106 can display a digital representation of a current image captured by the front-facing camera 104 and / or the rear-facing camera 112, an image that can be captured by one or more of these cameras, an image that was recently captured by one or more of these cameras, and / or a modified version of one or more of these images. Thus, the display 106 can function as a viewfinder for the cameras. The display 106 can also support touchscreen functionality that is capable of adjusting settings and / or configurations of one or more aspects of the computing device 100.
[0041] The front-facing camera 104 can include an image sensor and associated optical elements such as a lens. The front-facing camera 104 can provide zoom capabilities or can have a fixed focal length. In other examples, interchangeable lenses can be used with the front-facing camera 104. The front-facing camera 104 can have a variable mechanical aperture and mechanical and / or electronic shutter. The front-facing camera 104 can also be configured to capture still images, video images, or both. In addition, the front-facing camera 104 can represent, for example, a single view, stereo view, or multi-view camera. The rear-facing camera 112 can be similarly or differently arranged. Additionally, one or more of the front-facing camera 104 and / or the rear-facing camera 112 can be an array of one or more cameras.
[0042] One or more of the front-facing camera 104 and / or the rear-facing camera 112 can include or be associated with an illumination assembly that provides a light field to illuminate a target object. For example, the illumination assembly can provide a flash or constant illumination of the target object. The illumination assembly can also be configured to provide a light field that includes one or more of structured light, polarized light, and light having a particular spectral content. Within the context of the examples herein, other types of light fields known and used to recover three-dimensional (3D) models from objects are possible.
[0043] The computing device 100 can also include an ambient light sensor that can continuously or from time to time determine the ambient brightness of a scene that the cameras 104 and / or 112 can capture. In some implementations, the ambient light sensor can be used to adjust the display brightness of the display 106. Additionally, the ambient light sensor can be used to determine, or assist in the determination of, the exposure length of one or more of the cameras 104 or 112.
[0044] The computing device 100 can be configured to use the display 106 and the front-facing camera 104 and / or the rear-facing camera 112 to capture images of a target object. The captured images can be a plurality of still images or a video stream. The image capture can be triggered by activating the button 108, pressing a soft key on the display 106, or by some other mechanism. Depending on the implementation, the images can be automatically captured at specific time intervals, for example, upon pressing the button 108, upon suitable lighting conditions of the target object, upon moving the computing device 100 a predetermined distance, or according to a predetermined capture schedule.
[0045] Figure 2is a simplified block diagram illustrating some components of an example computing system 200. By way of example and not limitation, the computing system 200 can be a cellular mobile phone (e.g., a smartphone), a computer such as a desktop computer, a notebook computer, a tablet computer, or a handheld computer, a home automation component, a digital video recorder (DVR), a digital television, a remote controller, a wearable computing device, a gaming console, a robotic device, a vehicle, or some other type of device. The computing system 200 can represent aspects of the computing device 100, for example.
[0046] As shown in Figure 2 , the computing system 200 can include a communication interface 202, a user interface 204, a processor 206, a data store 208, and a camera component 224, all of which can be communicatively linked together by a system bus, network, or other connection mechanism 210. The computing system 200 can be equipped with at least some image capture and / or image processing capabilities. It will be appreciated that the computing system 200 can represent a physical image processing system, a particular physical hardware platform on which an image sensing and / or processing application operates in software, or other combinations of hardware and software configured to perform image capture and / or processing functions.
[0047] The communication interface 202 can allow the computing system 200 to communicate with other devices, access networks, and / or transport networks using analog or digital modulation. Thus, the communication interface 202 can facilitate circuit- switched and / or packet-switched communications, such as plain old telephone service (POTS) communications and / or Internet Protocol (IP) or other packetized communications. For example, the communication interface 202 can include a chipset and antenna arrangement for wireless communication with a radio access network or access point. Moreover, the communication interface 202 can take or include a wired interface such as an Ethernet, Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI) port. The communication interface 202 can also take or include a wireless interface such as a Wi-Fi, Bluetooth interface, and a wide-area wireless interface (e.g., WiMAX or 3GPP Long-Term Evolution (LTE)). However, other forms of physical layer interface and other types of standard or proprietary communication protocols can be used on the communication interface 202. Moreover, the communication interface 202 can include multiple physical communication interfaces (e.g., a Wi-Fi interface, a Bluetooth
[0048] The user interface 204 can be used to allow the computing system 200 to interact with a human or non-human user, such as receiving input from and providing output to a user. Thus, the user interface 204 can include input components such as a keypad, keyboard, touch panel, computer mouse, trackball, joystick, microphone, etc. The user interface 204 can also include one or more output components, such as a display screen, which may, for example, be combined with a touch panel. The display screen can be based on CRT, LCD, and / or LED technology, or other technologies now known or later developed. The user interface 204 can also be configured to generate audible output via a speaker, speaker jack, audio output port, audio output device, headphones, and / or other similar devices. The user interface 204 can also be configured to receive and / or capture audible utterances, noises, and / or signals by way of a microphone and / or other similar devices.
[0049] In some examples, the user interface 204 can include a display that functions as a viewfinder for still and / or video camera functionality supported by the computing system 200. Additionally, the user interface 204 can include one or more buttons, switches, knobs, and / or dials that facilitate configuration and focusing of the camera functionality, as well as capture of images. Some or all of these buttons, switches, knobs, and / or dials can be implemented by way of a touch panel.
[0050] The processor 206 can include one or more general-purpose processors (e.g., microprocessors) and / or one or more special-purpose processors (e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating point units (FPUs), network processors, or application-specific integrated circuits (ASICs)). In some cases, the special-purpose processor can be capable of image processing, image alignment, and merging images, among other possibilities. The data storage 208 can include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic memory devices, and can be integrated in whole or part with the processor 206. The data storage 208 can include removable and / or non-removable components.
[0051] The processor 206 can be capable of executing program instructions 218 (e.g., compiled or interpreted program logic and / or machine code) stored in the data storage 208 to perform various functions described herein. Thus, the data storage 208 can include a non-transitory computer-readable medium having program instructions stored therein that, when executed by the computing system 200, cause the computing system 200 to perform any of the methods, processes, or operations disclosed in the specification and / or drawings. Execution of the program instructions 218 by the processor 206 can result in the processor 206 using the data 212.
[0052] As an example, program instructions 218 can include an operating system 222 (e.g., an operating system kernel, device drivers, and / or other modules) and one or more application programs 220 installed on computing system 200 (e.g., a camera function, an address book, an email, a Web browser, a social network, an audio-to-text function, a text translation function, and / or a gaming application). Similarly, data 212 can include operating system data 216 and application data 214. Operating system data 216 can be primarily accessible by operating system 222, and application data 214 can be primarily accessible by one or more application programs 220. Application data 214 can be arranged in a file system that is visible to or hidden from a user of computing system 200.
[0053] Application programs 220 can communicate with operating system 222 through one or more application programming interfaces (APIs). These APIs can facilitate, for example, application programs 220 reading and / or writing application data 214, transmitting or receiving information via communication interface 202, receiving and / or displaying information on user interface 204, and so on.
[0054] In some cases, application programs 220 can be referred to simply as "apps." Additionally, application programs 220 can be downloadable to computing system 200 through one or more online application stores or application markets. However, application programs can also be installed on computing system 200 in other ways, such as via a Web browser or through a physical interface on computing system 200 (e.g., a USB port).
[0055] Camera assembly 224 can include, without limitation, an aperture, a shutter, a recording surface (e.g., photographic film and / or an image sensor), a lens, a shutter button, an infrared projector, and / or a visible light projector. Camera assembly 224 can include components configured to capture images in the visible spectrum (e.g., electromagnetic radiation having wavelengths of 380 nanometers to 700 nanometers) and / or components configured to capture images in the infrared spectrum (e.g., electromagnetic radiation having wavelengths of 701 nanometers to 1 millimeter), among other possibilities. Camera assembly 224 can be controlled at least in part by software executed by processor 206.
[0056] Example 3D photo system
[0057] Figure 3A An example system for generating 3D photos based on monocular images is illustrated. Specifically, 3D photo system 300 can be configured to generate modified image 370 based on image 302. 3D photo system 300 can include depth model 304, soft foreground visibility model 308, soft background removal occlusion function 312, inpainting model 316, pixel removal projector 322, and pixel projector 330.
[0058] Image 302 can be a monocular / monoscopic image comprising a plurality of pixels, each of which can be associated with one or more color values (e.g., a red-green-blue color image) and / or intensity values (e.g., a grayscale image). Image 302 can have an initial viewpoint from which the camera has captured the image. The initial viewpoint can be represented by initial camera parameters that indicate a spatial relationship between the camera and a world frame of reference of the scene represented by image 302. For example, the world frame of reference can be defined such that it is initially aligned with the camera frame of reference.
[0059] Modified image 370 can represent the same scene as image 302 from an adjusted viewpoint that is different from the initial viewpoint. The adjusted viewpoint can be represented by adjusted camera parameters that are different from the initial camera parameters and that indicate an adjusted spatial relationship between the camera and the world frame of reference in the scene. Specifically, in the adjusted spatial relationship, the camera frame of reference can be rotated and / or translated relative to the world frame of reference. Thus, by generating one or more instances of modified image 370 and viewing them in sequence, a 3D photo effect can be achieved due to the change in viewpoint from which the scene represented by image 302 is observed. This can allow image 302 to appear visually richer and / or more interactive by simulating motion of the camera relative to the scene.
[0060] Depth model 304 can be configured to generate a depth image 306 based on image 302. Image 302 can be represented as I e R nx3 . Image 302 can thus have n pixels, each of which can be associated with 3 color values (e.g., red, green, and blue). Depth image 306 can be represented as D = Φ D (I), where D e R nx1 and Φ D represents depth model 304. Depth image 306 can thus have n pixels, each of which can be associated with a corresponding depth value. The depth values can be scaled to a predetermined range, such as 0 to 1 (i.e., D e [0, 1] n). In some implementations, the depth values can represent parallax rather than metric depth, and can be converted to metric depth by an appropriate inversion of parallax. The depth model 304 can be a machine learning model that has been trained to generate depth images based on monocular images. For example, the depth model 304 can represent the MiDaS model as discussed in the article authored by Ranft et al. and published as arXiv: 1907.01341v3, entitled “Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer.” Additionally or alternatively, the depth model 304 can be configured to generate the depth image 306 using other techniques and / or algorithms, some of which can be based on additional image data corresponding to the image 302.
[0061] The soft foreground visibility model 308 can be configured to generate a foreground visibility map 310 based on the depth image 306. The foreground visibility map 310 can alternatively be referred to as a soft foreground visibility map 310. The foreground visibility map 310 can be represented as A e R n . The foreground visibility map 310 can thus have n pixels, each of which can be associated with a corresponding visibility value (which can be referred to as a foreground visibility value). The visibility values can be selected from and / or scaled to a predetermined range, such as 0 to 1 (i.e., A e [0, 1] n ). In some cases, the visibility values can alternatively be represented as opacity values, as the visibility of a foreground and the opacity of a foreground are complementary to each other. That is, the nth opacity value can be represented as Because It will thus be appreciated that whenever discussed herein, the visibility values can be explicitly represented in terms of visibility, or equivalently, implicitly represented in terms of opacity.
[0062] Within examples, a function can be considered to be “soft” in that along at least a particular input interval, the function is continuous and / or smooth. A function can be considered to be “hard” in that along at least a particular input interval, the function is not continuous and / or smooth. A map or mask can be considered to be “soft” when it is generated using a soft function rather than a hard function. In particular, the map or mask itself can not be continuous and / or smooth, but can still be considered to be “soft” if it has been generated by a soft function.
[0063] The soft foreground visibility model 308 can be configured to generate a foreground visibility map 310 based on the depth gradient of the depth image 306 and possibly also on the foreground alpha mask of the image 302. The visibility value of the foreground visibility map 310 can be relatively low at pixel locations representing the transition from foreground features to background features of the image 302 (and correspondingly, the transparency value can therefore be relatively high). The visibility value of the foreground visibility map 310 can be relatively high at pixel locations not representing the transition from foreground features to background features (and correspondingly, the transparency value can therefore be relatively low). Therefore, the visibility value can allow features that have been unmasked due to the change of viewpoint to be seen at least partially through the content of the foreground. Additional details of the soft foreground visibility model 308 are provided in... Figure 4A , Figure 4B , Figure 4C and Figure 4D The above is shown and discussed with reference to it.
[0064] The soft background demasking function 312 can be configured to generate a background demasking mask 314 based on the depth image 306. The background demasking mask 314 can alternatively be referred to as the soft background demasking mask 314. The background demasking mask 314 can be represented as S∈R n The background demasking mask 314 can therefore have n pixels, each pixel being associated with a corresponding demasking value (which may be referred to as the background demasking value). The demasking value can be selected from and / or scaled to a predetermined range, such as 0 to 1 (i.e., S∈[0,1]). n ).
[0065] Each corresponding demasking value of the background demasking mask 314 can indicate the probability that the corresponding pixel of image 302 will be demasked (i.e., become visible) due to a change in the initial viewpoint of image 302. Specifically, if there is an adjacent reference pixel (x i y j ), such that (i) the depth value D(x, y) associated with a given pixel and (ii) the depth value D(x, y) associated with a neighboring reference pixel. i y j The depth difference between these pixels is greater than the scaled distance between their corresponding locations. Then a given pixel associated with pixel position (x, y) can be unmasked due to a change in the initial viewpoint. That is, if Then a given pixel at position (x, y) can be demasked, where ρ is a scalar parameter that can be used when the depth image 306 represents disparity in order to scale the disparity to a metric depth, where the value of ρ can be chosen based on the maximum assumed camera movement between the initial viewpoint and the modified viewpoint, and where Thus, if the corresponding depth is relatively low (or the corresponding disparity is relatively high) compared to surrounding pixels, a given pixel can be more likely to be unoccluded by a change in the initial viewpoint, as features closer to the camera will appear to move more within the image 302 when the viewpoint is changed.
[0066] Thus, the soft background unocclusion function 312 can be expressed as where γ is an adjustable parameter that controls the steepness of the hyperbolic tangent function, i and j are each iterated through a corresponding range to compare the given pixel (x, y) with a predetermined number of neighboring pixels (x i , y j ) and a rectified linear unit (ReLU) or leaky ReLU can be applied to the output of the hyperbolic tangent function to make S (x,y) positive. Specifically, (x i , y j ) e N (x,y) , where N (x,y) represents a fixed neighborhood of m pixels around the given pixel at position (x, y). The value of m and the shape of N (x,y) may be selected based on, for example, a determined computational time allocated to the background unocclusion mask 314. For example, N (x,y) may define a square, a rectangle, an approximately circular region, or two or more scan lines. The two or more scan lines can include a vertical scan line that includes a predetermined number of pixels above and below the given pixel (x, y) and / or a horizontal scan line that includes a predetermined number of pixels to the right and left of the given pixel (x, y).
[0067] Thus, the soft background unocclusion function 312 can be configured to determine, for each respective pixel of the depth image 306, a first plurality of difference values, where each respective difference value of the first plurality of difference values is determined by subtracting from a corresponding depth value D(x i ,y j ) of the respective pixel (i) a corresponding depth value D(x of a corresponding reference pixel located within a predetermined pixel distance of the respective pixel and (ii) a scaled number of pixels separating the respective pixel and the corresponding reference pixel Further, the soft background unocclusion function 312 can be configured to determine, for each respective pixel of the depth image 306, a first maximum difference value of the first plurality of difference values (i.e., and determine the unocclusion value S (x,y) by applying a hyperbolic tangent function to the first maximum difference value.
[0068] The inpainting model 316 can be configured to generate an inpainted image 318 and an inpainted depth image 320 based on the image 302, the depth image 306, and the background-removed occlusion mask 314. The inpainted image 318 can be represented as and the inpainted depth image 320 can be represented as The inpainted image 318 and the inpainted depth image 320 can thus each have n pixels. The inpainting model 316 can implement a function wherein represents a depth-wise concatenation of the inpainted image 318 and the inpainted depth image 320, and Θ(·) represents the inpainting model 316.
[0069] In particular, the inpainted image 318 can represent the image 302 inpainted according to the background-removed occlusion mask 314, while the inpainted depth image 320 can represent the depth image 306 inpainted according to the background-removed occlusion mask 314. Changes to the initial viewpoint can cause apparent movement of one or more foreground regions, causing corresponding background regions (indicated by the background-removed occlusion mask 314) to become visible. Since the image 302 and the depth image 306 do not contain pixel values for these background regions, the inpainting model 316 can be used to generate these pixel values. The inpainted image 318 and the inpainted depth image 320 can thus be considered to form a background layer, while the image 302 and the depth image 306 form a foreground layer. These layers can be combined as part of generating the modified image 370.
[0070] The inpainting model 316 can be trained to inpaint background regions with values that match other background regions in the image 302, rather than values that match foreground regions of the image 302. Additionally or alternatively, the inpainting model 316 can be configured to generate the inpainted image 318 and the inpainted depth image 320 in parallel and / or concurrently. For example, both the inpainted image 318 and the inpainted depth image 320 can be generated by a single pass / iteration of the inpainting model 316. The output of the inpainting model 316 can thus be a depth-wise concatenation of the inpainted image 318 and the inpainted depth image 320, and each of these inpainted images can thus be represented by a corresponding one or more channels of the output.
[0071] Both the inpainted image 318 and the inpainted depth image 320 can be based on the image 302 and the depth image 306, such that the pixel values of the inpainted portions of the inpainted image 318 are consistent with the depth values of the corresponding inpainted portions of the inpainted depth image 320. That is, the inpainting model 316 can be configured to consider both depth information and color / intensity information in generating values for the occluded background regions, resulting in inpainted regions that exhibit consistency in depth and color / intensity. The depth information can allow the inpainting model 316 to distinguish between foreground regions and background regions of the image 302, and thus select appropriate portions from which to borrow values for inpainting. The inpainting model 316 can thus be considered depth-aware. Additional details of the inpainting model 316 are shown in and discussed with respect to Figure 5A
[0072] The pixel removal projector 322 can be configured to generate (i) a foreground visibility map 3D representation 324 based on the foreground visibility map 310 and the depth image 306, (ii) a foreground 3D representation 326 based on the image 302 and the depth image 306, and (iii) a background 3D representation 328 based on the inpainted image 318 and the inpainted depth image 320. In particular, the pixel removal projector 322 can be configured to back-project each respective pixel of the foreground visibility map 310, the image 302, and the inpainted image 318 to generate a corresponding 3D point to represent the respective pixel within a 3D camera reference frame. Thus, the 3D points of the foreground visibility map 3D representation can be represented as The 3D points of the foreground 3D representation 326 can be represented as The 3D points of the background 3D representation 326 can be represented as where p denotes the respective pixel, denotes a homography augmentation of the coordinates of the respective pixel, M denotes an intrinsic camera matrix with the principal point at the center of the image 302, D(p) denotes a depth value (rather than an inverse depth value) corresponding to the respective pixel as represented by the depth image 306, denotes a depth value (rather than an inverse depth value) corresponding to the respective pixel as represented by the inpainted depth image 320, and the camera reference frame is assumed to be aligned with the world reference frame. Since V(p) = F(p), a single set of 3D points can be determined for and shared by the 3D representations 324 and 326.
[0073] Based on the respective 3D points determined to represent the pixels of the foreground visibility map 310, the image 302, and the inpainted image 318, the pixel removal projector 322 can be configured to connect 3D points corresponding to adjacent pixels, thereby generating a respective polygonal mesh (e.g., a triangular mesh) for each of the foreground visibility map 310, the image 302, and the inpainted image 318. Specifically, if a pixel corresponding to a given 3D point and a pixel corresponding to one or more other 3D points are adjacent to each other within the corresponding image, the given 3D point can be connected to the one or more other 3D points.
[0074] The pixel removal projector 322 can also be configured to texture each respective polygonal mesh based on the pixel values of the corresponding image or map. Specifically, the polygonal mesh of the foreground visibility map 3D representation 324 can be textured based on the corresponding values of the foreground visibility map 310, the polygonal mesh of the foreground 3D representation 326 can be textured based on the corresponding values of the image 302, and the polygonal mesh of the background 3D representation 328 can be textured based on the corresponding values of the inpainted image 318. When a single set of 3D points is determined for and shared by the 3D representations 324 and 326, the 3D representations 324 and 326 can also represent using a single polygonal mesh that includes multiple channels, where one or more first channels represent the values of the foreground visibility map 310 and one or more second channels represent the values of the image 302. Additionally or alternatively, the 3D representations 324, 326, and 328 can include point clouds, layered representations (e.g., layered depth images), multiplane images, and / or implicit representations (e.g., Neural Radiance Fields (NeRF)), among other possibilities.
[0075] The pixel projector 330 can be configured to generate a modified image 370 based on the 3D representations 324, 326, and 328. Specifically, the pixel projector 330 can be configured to perform a rigid transformation T that represents a translation and / or a rotation of the camera relative to the camera pose associated with the initial viewpoint. Thus, the rigid transformation T can represent adjusted camera parameters that define an adjusted viewpoint for the modified image 370. For example, the rigid transformation T can be represented relative to a world reference frame associated with the initial viewpoint of the image 302. The value of T can be adjustable and / or iteratable to generate multiple different instances of the modified image 370, each representing the image 302 from a corresponding adjusted viewpoint.
[0076] Specifically, based on the rigid transformation T, the pixel projector 330 can be configured to (i) project the foreground visibility map 3D representation 324 to generate a modified foreground visibility map A T , (ii) project the foreground 3D representation 326 to generate a foreground image IT and (iii) projecting the background 3D representation 328 to generate a background image Thus, the modified foreground visibility map A T may represent the foreground visibility map 310 from the modified viewpoint, the foreground image I T may represent the image 302 from the modified viewpoint, and the background image may represent the inpainted image 318 from the modified viewpoint.
[0077] The pixel projector 330 can additionally be configured to generate the modified image 370 according to the modified foreground visibility map A T combining the foreground image I T with the background image . For example, the pixel projector 330 can be configured to generate the modified image 370 according to wherein represents the modified image 370. Thus, the modified image 370 can be based primarily on the image 302 where the values of the foreground visibility map 310 are relatively high, and can be based primarily on the inpainted depth image 320 where the values of the foreground visibility map 310 are relatively low. The visibility values of the modified visibility map A T may be selected from and / or scaled to the same predetermined range as the values of the foreground visibility map 310. For example, the range of values of the modified foreground visibility map can be from zero to a predetermined value (e.g., A T ∈ [0, 1] n ). Thus, the modified image 370 can be based on the sum of (i) a first product of the foreground image and the modified foreground visibility map (i.e., A T I T ) and (ii) a second product of the background image and the difference between the predetermined value and the modified foreground visibility map (i.e., ).
[0078] In other implementations, the pixel projector 330 can be configured to generate the modified image 370 according to wherein θ(·) represents a machine learning model that has been trained to combine the foreground image I T with the background image T according to A . For example, the machine learning model can be trained using a plurality of training image sets, each training image set including a plurality of different views of a scene and thus representing ground-truth pixel values with occluded regions removed.
[0079] Figure 3B , Figure 3C and Figure 3D contain examples of images and maps generated by the 3D photo system 300. Specifically,Figure 3B includes image 380 (example of image 302) and corresponding depth image 382 (example of depth image 306). Figure 3C includes foreground visibility map 384 (example of foreground visibility map 310) and background removal occlusion mask 386 (example of background removal occlusion mask 314), each of which corresponds to image 380 and depth image 382. Figure 3D includes inpainted image 388 (example of inpainted image 318), inpainted depth image 390 (example of inpainted depth image 320), and modified image 392 (example of modified image 370), each of which corresponds to image 380 and depth image 382.
[0080] In particular, image 380 contains a giraffe (foreground feature / object) relative to a background (background features / objects) of grass and trees. Depth image 382 indicates that the giraffe is closer to the camera than most of the background. In foreground visibility map 384, the maximum visibility value is shown in white, the minimum visibility value is shown in black, and values in between are shown using grayscale levels. Thus, as described above, the area around the giraffe, the top of the grass, and the top of the tree line are shown in black and / or dark grayscale levels, indicating that the depth gradient is relatively high in these areas, and that these areas are thus relatively more transparent to the inpainting of background pixel values after projection in three dimensions. In background removal occlusion mask 386, the maximum removal occlusion value is shown in white, the minimum removal occlusion value is shown in black, and values in between are shown using grayscale levels. Thus, the area behind the giraffe and the area behind the top of the grass are shown in white and / or light grayscale levels, indicating that these areas are likely to be removed occluded (i.e., become visible) by a change in viewpoint from which image 380 is rendered.
[0081] In inpainted image 388 and inpainted depth image 390, the portions of image 380 and depth image 382 that previously represented the giraffe and the top of the grass are inpainted using new values that correspond to the background portions of image 380 and depth image 382, rather than using values that correspond to the foreground portions of image 380 and depth image 382. Thus, images 388 and 390 represent image 380 and 382, respectively, with at least some of the foreground content replaced by newly synthesized background content that is contextually coherent (in terms of color / intensity and depth) with the background portions of image 380 and 382.
[0082] During rendering of the modified image 392, the content of the image 380 (foreground layer) can be combined with the inpainted content of the inpainted image 388 (background layer) to fill the occlusion-removed regions revealed by the modification of the initial viewpoint of the image 380. After re-projecting the images 380 and 388 and the map 384 into the modified viewpoint, the blending ratio between the images 380 and 388 is specified by the foreground visibility map 384. Specifically, the blending ratio (after re-projecting into the modified viewpoint) is specified by the corresponding pixel of the foreground visibility map 384, where the corresponding pixel is determined after removing the projected image 380, the inpainted image 388, and the foreground visibility map 384 to generate the respective 3D representations and the projection of these 3D representations into the modified viewpoint. Thus, the modified image 392 represents the content of the image 380 from a modified viewpoint that is different from the initial viewpoint of the image 380. The modified image 392 as shown is not an exact rectangle because the images 380 and 388 can not include content outside the outer bounds of the image 380. In some cases, the modified image 392 and / or one or more of the images on which it is based can undergo further inpainting so that the modified image 392 is rectangular even after the viewpoint modification.
[0083] Example soft foreground visibility model
[0084] Figure 4A Aspects of the soft foreground visibility model 308 are illustrated, which can be configured to generate the foreground visibility map 310 based on the depth image 306 and the image 302. Specifically, the soft foreground visibility model 308 can include a soft background occlusion function 332, an inverse function 336, a soft foreground visibility function 340, a foreground segmentation model 350, a mask model 354, a dilation function 358, a difference operator 362, and a multiplication operator 344.
[0085] The soft foreground visibility function 340 can be configured to generate a depth-based foreground visibility map 342 based on the depth image 306. Specifically, the soft foreground visibility function can be represented as where A DEPTH represents the depth-based foreground visibility map 342, represents a gradient operator (e.g., a Sobel gradient operator), and β is an adjustable scalar parameter. Thus, the foreground visibility value for a given pixel p of the depth image 306 can be based on an exponential of a depth gradient determined by applying the gradient operator to the given pixel p and one or more surrounding pixels.
[0086] In some implementations, the depth-based foreground visibility map 342 can be equal to the foreground visibility map 310. That is, the soft background occlusion function 332, the inverse function 336, the foreground segmentation model 350, the mask model 354, the dilation function 358, the difference operator 362, and / or the multiplication operator 344 can be omitted from the soft foreground visibility model 308. However, in some cases, the depth image 306 can inaccurately represent some thin and / or high-frequency features (e.g., due to the properties of the depth model 304) such as hair or fur. To improve the representation of these thin and / or high-frequency features in the modified image 370, the depth-based foreground visibility map 342 can be combined with the inverse background occlusion mask 338 and the mask-based foreground visibility map 364 to generate the foreground visibility map 310. Because the depth-based foreground visibility map 342 is soft, the depth-based foreground visibility map 342 can be combined with the mask-based foreground segmentation techniques.
[0087] The foreground segmentation model 350 can be configured to generate a foreground segmentation mask 352 based on the image 302. The foreground segmentation mask 352 can distinguish foreground features of the image 302 (e.g., represented using a binary pixel value of 1) from background features of the image 302 (e.g., represented using a binary pixel value of 0). The foreground segmentation mask 352 can distinguish foreground from background, but can not accurately represent thin and / or high-frequency features of foreground objects represented in the image 302. The foreground segmentation model 350 can be implemented, for example, using the model discussed in the article titled “U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection” authored by Qin et al. and published as arXiv:2005.09007.
[0088] The mask model 354 can be configured to generate a foreground alpha mask 356 based on the image 302 and the foreground segmentation mask 352. The foreground alpha mask 356 can distinguish foreground features of the image 302 from background features of the image 302, and can do so using a binary value per pixel or a grayscale intensity value per pixel. Unlike the foreground segmentation mask 352, the foreground alpha mask 356 can accurately represent thin and / or high-frequency features of foreground objects represented in the image 302 (e.g., hair or fur). The mask model 354 can be implemented, for example, using the model discussed in the article titled “F, B, Alpha Matting” authored by Forte et al. and published as arXiv:2003.07711.
[0089] To incorporate the foreground alpha mask 356 into the foreground visibility map 310, the foreground alpha mask 356 can undergo further processing. Specifically, the foreground visibility map 310 is expected to include low visibility values around the foreground object boundary, but is not expected to include low visibility values in areas that are not near the foreground object boundary. However, the foreground alpha mask 356 includes low visibility values for all background regions. Thus, a dilation function 358 can be configured to generate a dilated foreground alpha mask 360 based on the foreground alpha mask 356, which can then be subtracted from the foreground alpha mask 356 as indicated by a difference operator 362 to generate a mask-based foreground visibility map 364. The mask-based foreground visibility map 364 can thus include low visibility values around the foreground object boundary but not in background regions that are not near the foreground object boundary.
[0090] Additionally, the soft background occlusion function 332 can be configured to generate a background occlusion mask 334 based on the depth image 306. The background occlusion mask 334 can alternatively be referred to as a soft background occlusion mask 334. The background occlusion mask 334 can be represented as and thus can have n pixels, each of which can be associated with a corresponding occlusion value (which can be referred to as a background occlusion value). The occlusion values can be selected from and / or scaled to a predetermined range, such as 0 to 1 (i.e., ).
[0091] Each respective occlusion value of the background occlusion mask 334 can indicate a likelihood that a corresponding pixel of the image 302 will be occluded (i.e., completely covered) by a change in the initial viewpoint of the image 302. The computation of the background occlusion mask 334 can be similar to the computation of the background removal occlusion mask 314. Specifically, a given pixel associated with pixel location (x i , y j ) can be occluded due to a change in the initial viewpoint if (i) a depth value D(x i , y j ) associated with the given pixel and (ii) a depth value D(x associated with a neighboring reference pixel are less than a scaled distance between their respective locations . That is, if where (x i , y j ) ∈ N (x,y)such that i and j are each iterated through the corresponding ranges to compare the given pixel (x, y) to a predetermined number of neighboring pixels (x i , y j ) and ReLU or leaky ReLU can be applied to the output of the hyperbolic tangent function to make positive.
[0092] Accordingly, the soft background occlusion function 332 can be configured to determine, for each respective pixel of the depth image 306, a second plurality of difference values, where each respective difference value of the second plurality of difference values is determined by subtracting (i) the corresponding depth value D(x, y) of the respective pixel from (ii) the corresponding depth value D(x i , y j ) of a corresponding reference pixel that is separated from the respective pixel by a scaled number of pixels Further, the soft background removal occlusion function 312 can be configured to determine, for each respective pixel of the depth image 306, a second maximum difference value of the second plurality of difference values (i.e., and determine a removal occlusion value
[0093] The inverse function 336 can be configured to generate an inverse background occlusion mask 338 based on the background occlusion mask 334. For example, when the inverse function 336 can be represented as
[0094] The foreground visibility map 310 can be determined based on a product of (i) the depth-based foreground visibility map 342, (ii) the mask-based foreground visibility map 364, and (iii) the inverse background occlusion mask 338, as indicated by the multiplication operator 344. Determining the foreground visibility map 310 in this way can allow for representation of thin and / or high frequency features of the foreground object while accounting for depth discontinuities and reducing and / or avoiding “leakage” of the mask-based foreground visibility map 364 onto too much background. Accordingly, the net effect of this multiplication is to incorporate thin and / or high frequency features represented by the mask-based foreground visibility map 364 around the foreground object’s boundaries into the depth-based foreground visibility map 342.
[0095] Figure 4B , Figure 4C and Figure 4D include examples of images and maps used and / or generated by the soft foreground visibility model 308. Specifically, Figure 4B includes an image 402 (an example of the image 302) and a corresponding depth image 406 (an example of the depth image 306). Figure 4Cincluding a depth-based foreground visibility map 442 (example of depth-based foreground visibility map 342) and a foreground alpha mask 456 (example of foreground alpha mask 356), each of which correspond to image 402 and depth image 406. Figure 4D including a mask-based foreground visibility map 462 (example of mask-based foreground visibility map 364) and a foreground visibility map 410 (example of foreground visibility map 310), each of which correspond to image 402 and depth image 406.
[0096] In particular, image 402 contains a lion (foreground object) against a blurred background. Depth image 406 indicates that the lion is closer to the camera than the background, with the lion's snout being closer than other parts of the lion. Depth image 402 does not accurately represent the depth of some of the hairs in the lion's mane. In depth-based foreground visibility map 442 and foreground visibility map 410, the maximum visibility value is shown in white, the minimum visibility value is shown in black, and values between them are shown using shades of gray. In foreground alpha mask 456 and mask-based foreground visibility map 462, the maximum alpha value is shown in white, the minimum alpha value is shown in black, and values between them are shown using shades of gray. The maximum alpha value corresponds to the maximum visibility value, as both indicate that the foreground content is completely opaque.
[0097] Thus, in depth-based foreground visibility map 442, the outline regions representing the transitions between the lion and the background and between the lion's snout and other parts of the lion are shown in darker shades of gray, indicating that the depth gradient is relatively higher in these outline regions, and these outline regions are relatively more transparent for repairing background pixel values, while other regions are shown in lighter shades of gray. In foreground alpha mask 456, the lion is shown in lighter shades of gray, while the background is shown in darker shades of gray. Unlike depth-based foreground visibility map 442, foreground alpha mask 456 details the hairs of the lion's mane. In mask-based foreground visibility map 462, the lion is shown in lighter shades of gray, the outline region around the lion is shown in darker shades of gray, and the background portion outside the outline region is shown in lighter shades of gray. The size of this outline region can be controlled by controlling the number of pixels that dilation function 358 dilates foreground alpha mask 356.
[0098] In foreground visibility map 410, most of the lion is shown with lighter gray levels, the contour region around the lion and around the nose-mouth of the lion is shown with darker gray levels, and the background portion outside of the contour region is shown with lighter gray levels. The contour region of the lion in foreground visibility map 410 is smaller than the contour region of the lion in mask-based foreground visibility map 462, and includes a detailed representation of the fur of the lion's mane. In addition, foreground visibility map 410 includes a contour region around the nose-mouth of the lion, which is present in depth-based foreground visibility map 442 but not in mask-based foreground visibility map 462. Foreground visibility map 410 thus represents both depth-based discontinuities of image 402 and mask-based thin and / or high frequency features of the lion's mane.
[0099] Example training system and operations
[0100] Figure 5A An example training system 500 that can be used to train one or more components of 3D photo system 300 is illustrated. In particular, training system 500 can include depth model 304, soft background masking function 332, inpainting model 316, discriminator model 512, adversarial loss function 522, reconstruction loss function 508, and model parameter adjuster 526. Training system 500 can be configured to determine updated model parameters 528 based on training images 502. Training images 502 and training depth images 506 can be similar to images 302 and depth images 306, respectively, but can be determined at training time rather than at inference time. Depth model 304 and soft background masking function 332 have been discussed with reference to Figure 3A and Figure 4A .
[0101] Inpainting model 316 can be trained to inpaint regions of image 302 that can be unmasked by a change in viewpoint as indicated by background unmasking mask 314. However, in some cases, ground truth pixel values for these unmasked regions can not be available (e.g., multi-views of training images 502 can not be available). Thus, rather than training inpainting model 316 using a training background unmasking mask corresponding to training images 502, training inpainting model 316 can be performed using training background masking mask 534. Training background masking mask 534 can be similar to background masking mask 334, but can be determined at training time rather than at inference time. Because ground truth pixel values for regions that can be masked by foreground features due to a change in initial viewpoint are available in training images 502, this technique can allow inpainting model 316 to be trained without relying on a multi-view training image dataset. In cases where ground truth data for unmasked regions is available, inpainting model can additionally or alternatively be trained using an unmasking mask.
[0102] In particular, the inpainting model 316 can be configured to generate, based on the training images 502, the training depth images 506, and the training background occlusion masks 534, inpainted training images 518 and inpainted training depth images 520, which can be analogous to the inpainted images 318 and the inpainted depth images 320, respectively. In some cases, the inpainting model 316 can also be trained using stroke-shaped inpainting masks (not shown), thereby allowing the inpainting model 316 to learn to inpaint thin or small objects.
[0103] The quality with which the inpainting model 316 inpaints the missing regions of the images 502 and 506 can be quantified using a reconstruction loss function 508 and / or an adversarial loss function. The reconstruction loss function 508 can be configured to generate a reconstruction loss value 510 by comparing the training images 502 and 506 to the inpainted training images 518 and 520, respectively. For example, the reconstruction loss function 508 can be configured to determine, at least with respect to the regions that have been inpainted (as indicated by the training background occlusion masks 534 and / or the stroke-shaped inpainting masks), (i) a first L-1 distance between the training images 502 and the inpainted training images 518 and (ii) a second L-1 distance between the training depth images 506 and the inpainted training depth images 520. The reconstruction loss function 508 can thus incentivize the inpainting model 316 to generate inpainting intensity values and inpainting depth values that are (i) consistent with one another and (ii) based on background features (rather than foreground features) of the training images 502.
[0104] The discriminator model 512 can be configured to generate, based on the inpainted training images 518 and the inpainted training depth images 520, discriminator outputs 514. In particular, the discriminator outputs 514 can indicate whether the discriminator model 512 estimates that the inpainted training images 518 and the inpainted training depth images 520 were generated by the inpainting model 316 or as ground-truth images that have not yet been generated by the inpainting model 316. Thus, the inpainting model 316 and the discriminator model 512 can implement an adversarial training architecture. Accordingly, the adversarial loss function 522 can include, for example, a hinge adversarial loss, and can be configured to generate an adversarial loss value 524 based on the discriminator outputs 514. The adversarial loss function 522 can thus incentivize the inpainting model 316 to generate inpainting intensity values and inpainting depth values that look realistic and thus accurately mimic natural scenes.
[0105] The model parameter adjuster 526 can be configured to determine updated model parameters 528 based on the reconstruction loss value 510 and the adversarial loss value 524 (and any other loss values that can be determined by the training system 500). The model parameter adjuster 526 can be configured to determine a total loss value based on a weighted sum of these loss values, where the relative weights of the corresponding loss values can be adjustable training parameters. The updated model parameters 528 can include one or more updated parameters of the inpainting model 316, which can be denoted as ΔΘ. In some implementations, the updated model parameters 528 can additionally include one or more updated parameters for other components of the 3D photo system 300, such as the depth model 304 or the soft foreground visibility model 308, among other possibilities, as well as for the discriminator model 512.
[0106] The model parameter adjuster 526 can be configured to determine the updated model parameters 528 by, for example, determining the gradient of the total loss function. Based on the gradient and the total loss value, the model parameter adjuster 526 can be configured to select updated model parameters 528 that are expected to reduce the total loss value, and thus improve the performance of the 3D photo system 300. After applying the updated model parameters 528 to at least the inpainting model 316, the operations discussed above can be repeated to compute another instance of the total loss value, and based on this, another instance of the updated model parameters 528 can be determined and applied to at least the inpainting model 316 to further improve its performance. This training of the components of the 3D photo system 300 can be repeated until, for example, the total loss value is reduced below a target threshold loss value.
[0107] Figure 5B includes a background occlusion mask 544 (examples of the background occlusion mask 334 and the training background occlusion mask 534) that corresponds to a portion of the image 380 and the depth image 382. In the background occlusion mask 544, the maximum occlusion value is shown in white, the minimum occlusion value is shown in black, and values in between are shown using grayscale levels. Thus, the areas around the giraffe are shown in white or light grayscale levels, indicating that these areas are likely to be occluded by a change in the viewpoint from which the image 380 is rendered. Figure 3B
[0108] Additional example operations
[0109] Figure 6 A flowchart illustrating operations related to generating 3D effects based on monocular images. The operations can be performed by various computing devices 100, computing systems 200, 3D photo systems 300, and / or training systems 500, among other possibilities. Figure 6 Embodiments can be simplified in implementation by the removal of any one or more of the features described therein. Additionally, one or more features of the embodiments can be combined with features, aspects and / or implementations of any of the previous figures or features, aspects and / or implementations described herein.
[0110] Block 600 can involve obtaining a monocular image having an initial viewpoint.
[0111] Block 602 can involve determining, based on the monocular image, a depth image comprising a plurality of pixels. Each respective pixel of the depth image can have a corresponding depth value.
[0112] Block 604 can involve determining, for each respective pixel of the depth image, a corresponding depth gradient associated with the respective pixel of the depth image.
[0113] Block 606 can involve determining a foreground visibility map comprising, for each respective pixel of the depth image, a visibility value inversely proportional to the corresponding depth gradient.
[0114] Block 608 can involve determining, based on the depth image, a background removal occlusion mask comprising, for each respective pixel of the depth image, a removal occlusion value indicative of a likelihood that the corresponding pixel of the monocular image will be removal occluded by a change in the initial viewpoint.
[0115] Block 610 can involve (i) generating a repaired image by repairing portions of the monocular image according to the background removal occlusion mask using a repair model, and (ii) generating a repaired depth image by repairing portions of the depth image according to the background removal occlusion mask using the repair model.
[0116] Block 612 can involve (i) generating a first 3D representation of the monocular image based on the depth image, and (ii) generating a second 3D representation of the repaired image based on the repaired depth image.
[0117] Block 614 can involve generating a modified image having an adjusted viewpoint different from the initial viewpoint by combining the first 3D representation with the second 3D representation according to the foreground visibility map.
[0118] In some embodiments, the foreground visibility map can be determined by a soft foreground visibility function that is continuous and smooth along at least the first interval.
[0119] In some embodiments, determining the foreground visibility map comprises, for each respective pixel of the depth image, determining the corresponding depth gradient associated with the respective pixel by applying a gradient operator to the depth image, and determining, for each respective pixel of the depth image, the visibility value based on an exponent of the corresponding depth gradient associated with the respective pixel.
[0120] In some embodiments, the background removal occlusion mask can be determined by a soft background removal occlusion function that is continuous and smooth along at least the second interval.
[0121] In some embodiments, determining the background removal occlusion mask can include determining, for each respective pixel of the depth image, a first plurality of difference values. Each respective difference value of the first plurality of difference values can be determined by subtracting, from a corresponding depth value of the respective pixel, (i) a corresponding depth value of a corresponding reference pixel that is within a predetermined pixel distance of the respective pixel and (ii) a scaled number of pixels separating the respective pixel and the corresponding reference pixel. Determining the background removal occlusion mask can further include determining, for each respective pixel of the depth image, a removal occlusion value based on the first plurality of difference values.
[0122] In some embodiments, determining the removal occlusion value can include determining, for each respective pixel of the depth image, a first maximum difference value of the first plurality of difference values, and determining, for each respective pixel of the depth image, the removal occlusion value by applying a hyperbolic tangent function to the first maximum difference value.
[0123] In some embodiments, the corresponding reference pixel of a respective difference value can be selected from (i) a vertical scan line comprising a predetermined number of pixels above and below the respective pixel, or (ii) a horizontal scan line comprising a predetermined number of pixels to the right and left of the respective pixel.
[0124] In some embodiments, determining the foreground visibility map can include determining a depth-based foreground visibility map comprising, for each respective pixel of the depth image, a visibility value that is inversely proportional to a corresponding depth gradient associated with the respective pixel of the depth image, and determining a foreground alpha mask based on and corresponding to the monocular image. Determining the foreground visibility map can further include determining a mask-based foreground visibility map based on a difference between (i) the foreground alpha mask and (ii) a dilation of the foreground alpha mask, and determining the foreground visibility map based on a product of (i) the depth-based foreground visibility map and (ii) the mask-based foreground visibility map.
[0125] In some embodiments, determining the foreground alpha mask can include determining, by a foreground segmentation model, a foreground segmentation based on the monocular image, and determining, by a mask model and based on the foreground segmentation and the monocular image, the foreground alpha mask.
[0126] In some embodiments, determining the foreground visibility map can include determining, based on the depth image, a background occlusion mask including, for each respective pixel of the depth image, an occlusion value indicating a likelihood that a corresponding pixel of the monocular image would be occluded by a change in the initial viewpoint. Determining the foreground visibility map can further include determining the foreground visibility map based on a product of (i) the depth-based foreground visibility map, (ii) the mask-based foreground visibility map, and (iii) an inverse of the background occlusion mask.
[0127] In some embodiments, determining the background occlusion mask can include determining, for each respective pixel of the depth image, a second plurality of difference values. Each respective difference value of the second plurality of difference values can be determined by subtracting (i) a corresponding depth value of the respective pixel from (ii) a corresponding depth value of a corresponding reference pixel and a scaled number of pixels separating the respective pixel from the corresponding reference pixel within a predetermined pixel distance of the respective pixel. Determining the background occlusion mask can further include determining, for each respective pixel of the depth image, the occlusion value based on the second plurality of difference values.
[0128] In some embodiments, determining the occlusion value can include determining, for each respective pixel of the depth image, a second maximum difference value of the second plurality of difference values, and determining, for each respective pixel of the depth image, the occlusion value by applying a hyperbolic tangent function to the second maximum difference value.
[0129] In some embodiments, the monocular image can include foreground features and background features. The inpainting model can have been trained to inpaint (i) a de-occluded background region of the monocular image having matching background features and independent of intensity values of the foreground features, and (ii) a corresponding de-occluded background region of the depth image having matching background features and independent of depth values of the foreground features.
[0130] In some embodiments, the inpainting model can have been trained such that intensity values of the de-occluded region of the inpainted image are contextually consistent with corresponding depth values of the corresponding de-occluded background region of the inpainted depth image.
[0131] In some embodiments, the inpainting model can have been trained through a training process that includes obtaining a training monocular image and determining a training depth image based on the training monocular image. The training process can also include determining a training background occlusion mask based on the training depth image, the training background occlusion mask including, for each respective pixel of the training depth image, an occlusion value indicating a likelihood that a corresponding pixel of the training monocular image would be occluded by a change in a viewpoint of the training monocular image. The training process can additionally include: (i) generating an inpainted training image by inpainting portions of the training monocular image from the training background occlusion mask using the inpainting model, and (ii) generating an inpainted training depth image by inpainting portions of the training depth image from the training background occlusion mask using the inpainting model. The training process can further include determining a loss value by applying a loss function to the inpainted training image and the inpainted training depth image, and adjusting one or more parameters of the inpainting model based on the loss value.
[0132] In some embodiments, the loss value can be based on one or more of: (i) an adversarial loss value determined based on processing the inpainted training image and the inpainted training depth image through a discriminator model, or (ii) a reconstruction loss value determined based on comparing the inpainted training image to the training monocular image and comparing the inpainted training depth image to the training depth image.
[0133] In some embodiments, generating the modified image can include generating a 3D foreground visibility map corresponding to the first 3D representation based on the depth image. Generating the modified image can also include: (i) generating a foreground image by projecting the first 3D representation based on the adjusted viewpoint, (ii) generating a background image by projecting the second 3D representation based on the adjusted viewpoint, and (iii) generating a modified foreground visibility map by projecting the 3D foreground visibility map based on the adjusted viewpoint. Generating the modified image can further include combining the foreground image with the background image according to the modified foreground visibility map.
[0134] In some embodiments, values of the modified foreground visibility map can have a range of zero to a predetermined value. Combining the foreground image with the background image can include determining a sum of (i) a first product of the foreground image and the modified foreground visibility map and (ii) a second product of the background image and a difference between the predetermined value and the modified foreground visibility map.
[0135] In some embodiments, generating the first 3D representation and the second 3D representation can include (i) generating a first plurality of 3D points by removing projections of pixels of the monocular image based on the depth image pair, and (ii) generating a second plurality of 3D points by removing projections of pixels of the inpainted image based on the inpainted depth image pair, and (i) generating a first polygonal mesh by interconnecting respective subsets of the first plurality of 3D points corresponding to contiguous pixels of the monocular image, and (ii) generating a second polygonal mesh by interconnecting respective subsets of the second plurality of 3D points corresponding to contiguous pixels of the inpainted image. Generating the first 3D representation and the second 3D representation can further include (i) applying one or more first textures to the first polygonal mesh based on the monocular image, and (ii) applying one or more second textures to the second polygonal mesh based on the inpainted image.
[0136] Example performance metrics
[0137] Figure 7A 、 Figure 7B and Figure 7C each include respective tables containing results of testing various 3D photo models against corresponding data sets. Specifically, Figure 7A 、 Figure 7B and Figure 7C each compare the performance of a SynSin model (described in the article titled “SynSin: End-to-end View Synthesis from a Single Image” authored by Wiles et al. and published as arXiv: 1912.08804), an SMPI model (described in the article titled “Single-View View Synthesis with Multiplane Images” authored by Tucker et al. and published as arXiv: 2004.11364), a 3D photo model (described in the article titled “3D Photography using Context-aware Layered Depth Inpainting” authored by Shih et al. and published as arXiv: 2004.04727), and the 3D photo system 300 as described herein.
[0138] Performance is quantified by learned perceptual image patch similarity (LPIPS) metric (lower values indicate better performance), peak signal-to-noise ratio (PSNR) (higher values indicate better performance), and structural similarity index measure (SSIM) (higher values indicate better performance). Figure 7A 、 Figure 7B andFigure 7C The results in Table 1 correspond to the RealEstate10k image dataset (discussed in the article authored by Zhou et al. and published as arXiv: 1805.09817, entitled "Stereo Magnification: Learning View Synthesis using Multiplane Images"), the Dual-Pixels image dataset (discussed in the article authored by Garg et al. and published as arXiv: 1904.05822, entitled "Learning Single Camera Depth Estimation using Dual-Pixels"), and the Human3.6M image dataset (discussed in the article authored by Li et al. and published as arXiv: 1904.11111, entitled "Learning the Depths of Moving People by Watching Frozen People"), respectively.
[0139] In Table 1, Figure 7A In Table 1, Figure 7B In Table 1, Figure 7C In Table 1, Figure 7A Figure 7B Figure 7C As indicated by the shaded pattern in the respective bottom-most row of the table in each of Tables 2,
[0140] CONCLUSION
[0141] The present disclosure should not be limited to the particular embodiments described in this application, as these are intended to illustrate various aspects. Many modifications and variations will be apparent to those skilled in the art without departing from their scope. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those described herein, will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.
[0142] The above detailed description describes various features and operations of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the drawings, similar symbols typically identify corresponding or similar components throughout the several views, unless context dictates otherwise. The example embodiments described herein and in the drawings are not intended to be limiting. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0143] With respect to any or all of the message flow diagrams, scenarios, and flowcharts in the drawings and as discussed herein, each step, block, and / or communication can represent a processing of information and / or a transfer of information in accordance with example embodiments. Alternative embodiments are included within the scope of the example embodiments. In these alternative embodiments, for example, the operations described as steps, blocks, transfers, communications, requests, responses, and / or messages can not be performed in the order indicated or discussed, including substantially concurrently or in reverse order, depending on the functionality involved. Moreover, more or fewer blocks and / or operations can be utilized in any of the message flow diagrams, scenarios, and flowcharts discussed herein, and these message flow diagrams, scenarios, and flowcharts can be combined or split into further message flow diagrams, scenarios, and flowcharts as needed.
[0144] Steps or blocks representing processing of information can correspond to a circuit that can be configured to perform a specific logical function or a portion of a method or technique described herein. Alternatively or additionally, a block representing processing of information can correspond to a portion of a module, segment, or program code (including related data). Program code can include one or more instructions executable by a processor for implementing a specific logical operation or action in a method or technique. Program code and / or related data can be stored on any type of computer readable medium, such as a storage device including a random access memory (RAM), a magnetic disk drive, a solid state drive, or another storage medium.
[0145] Computer-readable media can also include non-transitory computer-readable media, such as computer-readable media that store data for short periods of time like register memory, processor cache and RAM. Computer-readable media also can include non-transitory computer-readable media that store programs of instructions or data for periods of time, such as secondary or persistent long term storage, like ROM, optical or magnetic disks, solid state drives, or flash memory. Therefore, computer-readable media can include both volatile and non-volatile media, removable and non-removable media, and storage media parted and non-parted. Computer-readable media can be considered computer-readable storage media, for example, or tangible storage devices.
[0146] Also, the steps or blocks of the steps representing one or more information transmissions can correspond to information transmissions between software and / or hardware modules in the same physical device. However, other information transmissions can be between software modules and / or hardware modules in different physical devices.
[0147] The particular arrangement of elements shown in the figures should not be interpreted as limiting. It will be understood that other embodiments can include more or fewer of each element shown in the given figure. Furthermore, various elements of some embodiments can be combined with each other and / or omitted. Moreover, example embodiments can also include additional elements that are not shown in the figures.
[0148] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the appended claims.
Claims
1. A computer-implemented method, comprising: For each corresponding pixel among a plurality of pixels in a depth image, a corresponding depth gradient associated with the corresponding pixel in the depth image is determined, wherein the depth image corresponds to an input image having an initial viewpoint, and wherein each corresponding pixel among the plurality of pixels in the depth image has a corresponding depth value; Determine a foreground visibility map, the foreground visibility map comprising a visibility value for each of the plurality of pixels of the depth image, the sum of which is inversely proportional to the corresponding depth gradient associated with the corresponding pixel of the depth image; (i) generating a repaired image by repairing portions of the input image that are expected to be unoccluded due to the change in the initial viewpoint; and (ii) generating a repaired depth image by repairing portions of the depth image that are expected to be unoccluded due to the change in the initial viewpoint; and By combining the visual information of the input image and the repaired image based on the depth image, the repaired depth image, and the foreground visibility map, a modified image with an adjusted viewpoint different from the initial viewpoint is generated.
2. The computer-implemented method according to claim 1, wherein, The foreground visibility mapping is determined using a soft foreground visibility function that is continuous and smooth along at least one interval.
3. The computer-implemented method according to claim 1, wherein: Determining the corresponding depth gradient includes: for each corresponding pixel among the plurality of pixels in the depth image, determining the corresponding depth gradient associated with the corresponding pixel by applying a gradient operator to the depth image; and Determining the foreground visibility map includes: for each of the plurality of pixels in the depth image, determining the visibility value based on the exponent of the corresponding depth gradient associated with the corresponding pixel.
4. The computer-implemented method according to claim 1, wherein, Determining the foreground visibility mapping includes: Determine a depth-based foreground visibility map, the depth-based foreground visibility map comprising: a visibility value for each corresponding pixel in the plurality of pixels of the depth image, the sum of which is inversely proportional to the corresponding depth gradient associated with the corresponding pixel in the depth image; Based on the foreground alpha mask corresponding to the input image, determine a mask-based foreground visibility mapping; and The foreground visibility map is determined by combining (i) the depth-based foreground visibility map and (ii) the mask-based foreground visibility map.
5. The computer-implemented method according to claim 4, wherein, Determining the foreground visibility mapping includes: Based on the depth image, a background occlusion mask is determined, wherein for each corresponding pixel of the plurality of pixels in the depth image, the background occlusion mask indicates the probability that the corresponding pixel in the input image will be occluded due to a change in the initial viewpoint; and The foreground visibility map is determined based on the product of: (i) the depth-based foreground visibility map, (ii) the mask-based foreground visibility map, and (iii) the inverse of the background occlusion mask.
6. The computer-implemented method according to claim 5, wherein, Determining the background occlusion mask includes: For each corresponding pixel among the plurality of pixels of the depth image, a plurality of differences are determined, wherein each of the plurality of differences is determined by subtracting (i) the corresponding depth value of the corresponding pixel and a scaled number of pixels separating the corresponding pixel from the corresponding reference pixel located within a predetermined pixel distance of the corresponding pixel from the corresponding depth value of (ii) the corresponding reference pixel; and For each corresponding pixel among the plurality of pixels in the depth image, the occlusion probability is determined based on the plurality of differences.
7. The computer-implemented method according to claim 6, wherein, Determining the occlusion probability includes: For each corresponding pixel among the plurality of pixels in the depth image, determine the maximum difference among the plurality of differences; and For each corresponding pixel among the plurality of pixels in the depth image, the occlusion probability is determined by applying the hyperbolic tangent function to the maximum difference.
8. The computer-implemented method according to claim 1, further comprising: Determine a background demasking mask, for each corresponding pixel among the plurality of pixels in the depth image, the background demasking mask indicating the occlusion probability that the corresponding pixel in the input image will be demasked by the change of the initial viewpoint, wherein generating the restored image and the restored depth image includes: (i) generating a repaired image using the input image based on the background occlusion mask repair portion; and (ii) generating a repaired depth image using the depth image based on the background occlusion mask repair portion.
9. The computer-implemented method according to claim 8, wherein, The background masking is determined using a soft background removal masking function that is continuous and smooth along at least one interval.
10. The computer-implemented method according to claim 8, wherein, The background removal mask is determined in the following way: For each corresponding pixel of the plurality of pixels in the depth image, a plurality of differences are determined, wherein each of the plurality of differences is determined by subtracting (i) the corresponding depth value of a corresponding reference pixel located within a predetermined pixel distance of the corresponding pixel and (ii) the scaled number of pixels separating the corresponding pixel from the corresponding depth value of the corresponding pixel; and For each corresponding pixel of the depth image, the likelihood of removing the occlusion is determined based on the plurality of differences.
11. The computer-implemented method according to claim 10, wherein, Determining the likelihood of removing the occlusion includes: For each corresponding pixel among the plurality of pixels in the depth image, determine the maximum difference among the plurality of differences; and For each corresponding pixel among the plurality of pixels in the depth image, the likelihood of removing the occlusion is determined by applying a hyperbolic tangent function to the maximum difference.
12. The computer-implemented method according to claim 10, wherein, The corresponding reference pixel for the corresponding difference is selected from: (i) a vertical scan line including a predetermined number of pixels above and below the corresponding pixel, or (ii) a horizontal scan line including the predetermined number of pixels to the right and left of the corresponding pixel.
13. The computer-implemented method according to claim 1, wherein, Generating the modified image includes: (i) generating a first three-dimensional (3D) representation of the input image based on the depth image; and (ii) generating a second 3D representation of the restored image based on the restored depth image; and The modified image is generated by combining the first 3D representation with the second 3D representation according to the foreground visibility mapping.
14. The computer-implemented method according to claim 13, wherein, Generating the modified image includes: Based on the depth image, a 3D foreground visibility map corresponding to the first 3D representation is generated; (i) generating a foreground image by projecting the first 3D representation based on the adjusted viewpoint, (ii) generating a background image by projecting the second 3D representation based on the adjusted viewpoint, and (iii) generating a modified foreground visibility map by projecting the 3D foreground visibility map based on the adjusted viewpoint; and The foreground image is combined with the background image according to the modified foreground visibility mapping.
15. The computer-implemented method according to claim 13, wherein, Generating the first 3D representation and the third includes: (i) generating a first plurality of 3D points by removing the pixel projection of the input image based on the depth image, and (ii) generating a second plurality of 3D points by removing the pixel projection of the repaired image based on the repaired depth image; (i) generating a first polygonal mesh by interconnecting corresponding subsets of a first plurality of 3D points corresponding to adjacent pixels of the input image, and (ii) generating a second polygonal mesh by interconnecting corresponding subsets of a second plurality of 3D points corresponding to adjacent pixels of the repaired image; and (i) applying one or more first textures to the first polygonal mesh based on the input image, and (ii) applying one or more second textures to the second polygonal mesh based on the repaired image.
16. The computer-implemented method according to claim 1, wherein, The restored image and the restored depth image are generated by the restoration model.
17. The computer-implemented method according to claim 16, wherein, The training of the repair model includes: Obtain a training input image with the original viewpoint and a training depth image corresponding to the training input image; A training background occlusion mask is determined based on the training depth image. The training background occlusion mask includes a training occlusion value for each corresponding pixel among a plurality of pixels in the training depth image. The training occlusion value indicates the probability that the corresponding pixel in the training input image will be occluded by a change in the original viewpoint. (i) generating a repair training image by using the repair model and the training input image of the training background occlusion mask repair portion, and (ii) generating a repair training depth image by using the repair model and the training depth image of the training background occlusion mask repair portion; The loss value is determined by applying a loss function to the repaired training image and the repaired training depth image; and One or more parameters of the repair model are adjusted based on the loss value.
18. The computer-implemented method according to claim 1, wherein, The input image is a monocular image.
19. A system comprising a processor configured to perform operations of the method according to any one of claims 1 to 18.
20. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by a computing device, cause the computing device to perform the operation of the method according to any one of claims 1 to 18.