Method and system for single-image 3D photography with soft stratification and depth perception repair
Through soft layering formation and depth perception repair technology, the problem of low 3D photo generation quality in the prior art is solved, and high-quality 3D photo generation and visual effects are improved.
Patent Information
- Application Number
- CN202180063773.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-05
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-08-05
AI Technical Summary
The prior art is difficult to generate high-quality 3D photos, especially when dealing with complex scenes and thin features, resulting in visually flawed or defective 3D effects.
Through soft layering formation and depth perception repair techniques, a given scene is decomposed into foreground layer and background layer, and depth model is used to generate depth images, soft foreground visibility maps and soft background removal occlusion masks are generated based on depth images, and the removal occlusion areas of monocular images and depth images are repaired.
The generation of high-quality 3D photos is achieved, enabling the generation of modified images from new perspectives, improving the richness and interactivity of 3D effects, especially when dealing with thin features and complex scenes.
Smart Images

Figure CN116250002B_ABST
Abstract
Description
Background Art
[0001] Machine learning models can be used to process various types of data including images to generate various desired outputs. Improvements to machine learning models allow the models to perform data processing faster, utilize fewer computing resources for processing, and / or generate outputs of relatively higher quality. Summary of the Invention
[0002] A three-dimensional (3D) photo system can be configured to generate a 3D viewing effect / experience that simulates different viewpoints of a scene represented by a monocular image. Specifically, a depth image can be generated based on the monocular image. A soft foreground visibility map can be generated based on the depth image and can indicate the transparency of different parts of the monocular image with respect to the background layer. The monocular image can be considered to form part of the foreground layer. A soft background removal occlusion mask can be generated based on the depth image and can indicate the likelihood of different background regions of the monocular image being unoccluded due to a change in viewpoint. A inpainting model can use the soft background removal occlusion mask to inpaint the unoccluded background regions of the monocular image and the depth image, thereby generating the background layer. The foreground layer and the background layer can be represented in 3D, and these 3D representations can be projected from the new viewpoint to generate a new foreground image and a new background image. The new foreground image and the new background image can be combined according to the soft foreground visibility map reprojected into the new viewpoint, thereby generating a modified image with the new viewpoint of the scene.
[0003] In a first exemplary embodiment, a method includes: obtaining a monocular image having an initial viewpoint, and determining a depth image including a plurality of pixels based on the monocular image. Each respective pixel of the depth image can have a corresponding depth value. The method further includes: determining, for each respective pixel of the depth image, a corresponding depth gradient associated with the respective pixel of the depth image, and determining a foreground visibility map including visibility values that are inversely proportional to the corresponding depth gradients for each respective pixel of the depth image. The method additionally includes determining a background removal occlusion mask based on the depth image, the background removal occlusion mask including removal occlusion values for each respective pixel of the depth image, the removal occlusion values indicating the likelihood that the corresponding pixel of the monocular image will be unoccluded by a change in the initial viewpoint. The method further additionally includes: (i) generating a repaired image by using a repair model to repair a portion of the monocular image according to the background removal occlusion mask, and (ii) generating a repaired depth image by using a repair model to repair a portion of the depth image according to the background removal occlusion mask. The method further includes: (i) generating a first three-dimensional (3D) representation of the monocular image based on the depth image, and (ii) generating a second 3D representation of the repaired image based on the repaired depth image. The method further includes generating a modified image having an adjusted viewpoint different from the initial viewpoint by combining the first 3D representation with the second 3D representation according to the foreground visibility map.
[0004] In a second exemplary embodiment, a system can include a processor and a non-transitory computer-readable medium having instructions stored thereon that, when executed by the processor, cause the processor to perform the operations according to the first exemplary embodiment.
[0005] In a third exemplary embodiment, a non-transitory computer-readable medium can have instructions stored thereon that, when executed by a computing device, cause the computing device to perform the operations according to the first exemplary embodiment.
[0006] In a fourth exemplary embodiment, a system can include various means for performing each of the operations of the first exemplary embodiment.
[0007] These and other embodiments, aspects, advantages, and alternatives will become apparent to those of ordinary skill in the art by reading the following detailed description with reference to the accompanying drawings where appropriate. Additionally, the present disclosure and the other descriptions and drawings provided herein are intended to illustrate embodiments by way of example only, and thus, various modifications are possible. For example, structural elements and process steps can be rearranged, combined, distributed, eliminated, or otherwise changed while remaining within the scope of the claimed embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 Illustrated is a computing device according to examples described herein.
[0009] Figure 2 A computing system according to examples described herein is illustrated.
[0010] Figure 3A Illustrated is a 3D photo system according to examples described herein.
[0011] Figure 3B Illustrated are images and depth images according to examples described herein.
[0012] Figure 3C Illustration of a visibility map and background subtraction occlusion mask according to examples described herein.
[0013] Figure 3D Illustrated is an inpainted image, an inpainted depth image, and a modified image according to examples described herein.
[0014] Figure 4A Illustrated is a soft foreground visibility model according to examples described herein.
[0015] Figure 4B Illustrated are images and depth images according to examples described herein.
[0016] Figure 4C Illustrated is a depth-based foreground visibility map and foreground alpha mask according to examples described herein.
[0017] Figure 4D Illustrated is a mask-based foreground visibility map and a foreground visibility map according to examples described herein.
[0018] Figure 5A Illustrated is a training system according to examples described herein.
[0019] Figure 5B Illustration of a background occlusion mask according to examples described herein.
[0020] Figure 6 Included are flow charts according to examples described herein.
[0021] Figure 7A , Figure 7B and Figure 7C Included is a table of performance metrics according to examples described herein. DETAILED DESCRIPTION
[0022] This document describes example methods, devices, and systems. It should be understood that the terms "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration". Any embodiment or feature described herein as "example", "exemplary", and / or "illustrative" need not be construed as preferred or advantageous over other embodiments or features, unless so stated. Thus, other embodiments can be utilized and other changes can be made without departing from the scope of the subject matter presented herein.
[0023] Accordingly, the example embodiments described herein are not meant to be limiting. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0024] Furthermore, unless the context otherwise suggests, the features shown in each figure can be used in combination with one another. Thus, the figures should generally be regarded as constituent aspects of one or more overall embodiments, and it should be understood that not all of the features shown are necessary for each embodiment.
[0025] Additionally, any listing of elements, boxes, or steps in this specification or claims is for the purpose of clarity. Thus, such listing should not be construed as requiring or implying that these elements, boxes, or steps follow a particular arrangement or are to be performed in a particular order. Unless otherwise noted, the figures are not drawn to scale.
[0026] Overview
[0027] A monocular image representing a scene from an initial viewpoint can be used to generate a modified image representing the scene from an adjusted viewpoint. By sequentially viewing the monocular image and / or the modified image, a 3D effect can be simulated, thereby creating a richer and / or more interactive viewing experience. In particular, a single monocular image can be used to create this viewing experience, which is generated without additional optical hardware beyond the optical hardware of a single field-of-view camera.
[0028] Some techniques for generating an image with a modified viewpoint can rely on decomposing a scene into two or more layers based on hard discontinuities. Hard discontinuities can define a sharp and / or discrete (e.g., binary) transition between two layers (e.g., foreground and background). Hard discontinuities may not allow for accurate modeling of some appearance effects such as very thin objects (e.g., hair), resulting in a visually flawed or defective 3D effect. Other techniques for generating an image with a modified viewpoint can rely on a training dataset that includes multi-view images providing ground truth for different viewpoints. However, it may be difficult to obtain a multi-view dataset representing a wide range of scenes, and these methods may therefore perform poorly on out-of-distribution scenes that are not well represented by the training data distribution.
[0029] Accordingly, the present disclosure provides a 3D photo system that relies on soft layering formation and depth-aware inpainting to decompose a given scene into a foreground layer and a background layer. These layers can be combined to generate a modified image from a new viewpoint. Specifically, a depth model can be used to generate a depth image corresponding to a monocular image, and the depth image can be used to generate a soft foreground visibility map and / or a soft background occlusion mask.
[0030] The soft foreground visibility map can be generated based on the depth gradient of the depth image and can indicate the degree to which different parts of the foreground layer are transparent and thus allow the corresponding regions of the background layer to be seen through the foreground layer. The foreground visibility map can be soft because it is generated by a soft foreground visibility function that is continuous and smooth at least along the range of input values of the depth gradient. The foreground layer can be defined by the monocular image and the foreground visibility map. In some embodiments, the soft foreground visibility map can also be based on a foreground alpha mask, which can improve the representation of thin features and / or high-frequency features such as hair and / or fur.
[0031] The soft background occlusion mask can be generated based on the depth image to quantify the likelihood that different pixels of the monocular image will be unoccluded due to a change in viewpoint. The background occlusion mask can be soft because it is generated by a soft background occlusion function that is continuous and smooth at least along the range of input values of the depth image. The terms "map" and "mask" can be used interchangeably to refer to grayscale images and / or binary images.
[0032] An inpainting model can use the soft background occlusion mask to inpaint the possibly occluded parts of the monocular image and the depth image, thereby generating the background layer. Specifically, the same inpainting model can be trained to inpaint the monocular image and the corresponding depth image such that the inpainted pixel values are based on both the color / intensity values of the monocular image and the depth values of the depth image. The inpainting model can be trained on both the image and the depth data such that it is trained to perform depth-aware inpainting. For example, the inpainting model can be trained to inpaint the monocular image and the corresponding depth image concurrently. Thus, the inpainting model can be configured to inpaint the occluded background regions by borrowing information from other background regions rather than from foreground regions, thereby producing a visually consistent background layer. Additionally, the inpainting model can be configured to inpaint the monocular image and the depth image in a single pass (e.g., a single processing pass) of the inpainting model rather than using multiple passes.
[0033] The corresponding 3D representations of the foreground layer and the background layer can be generated based on a depth image and a refined depth image, respectively. The viewpoints from which these 3D representations are observed can be adjusted, and the 3D representations can be projected to generate a foreground image and a background image with the adjusted viewpoints. Then, the foreground image and the background image can be combined according to a modified foreground visibility map that also has the adjusted viewpoint. Thus, the occluded regions removed in the foreground layer can be filled based on the corresponding regions of the background layer, where the transition between the foreground and the background includes the fusion of both the foreground layer and the background layer.
[0034] Example Computing Devices and Systems
[0035] Figure 1 FIG. illustrates an example computing device 100. Computing device 100 is shown in the form factor of a mobile phone. However, computing device 100 may alternatively be implemented as a laptop computer, a tablet computer, and / or a wearable computing device, among other possibilities. Computing device 100 may include various elements such as a body 102, a display 106, and buttons 108 and 110. Computing device 100 may further include one or more cameras, such as a front camera 104 and a rear camera 112.
[0036] The front camera 104 may be positioned on a side of the body 102 that generally faces the user during operation (e.g., on the same side as the display 106). The rear camera 112 may be positioned on the side of the body 102 opposite the front camera 104. Referring to the cameras as front and rear is arbitrary, and computing device 100 may include multiple cameras positioned on various sides of the body 102.
[0037] The display 106 can represent a cathode ray tube (CRT) display, a light-emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light-emitting diode (OLED) display, or any other type of display known in the art. In some examples, the display 106 can display a digital representation of a current image captured by the front camera 104 and / or the rear camera 112, an image that can be captured by one or more of these cameras, an image most recently captured by one or more of these cameras, and / or a modified version of one or more of these images. Thus, the display 106 can be used as a viewfinder for the cameras. The display 106 can also support a touchscreen function that can adjust the settings and / or configuration of one or more aspects of the computing device 100.
[0038] The front camera 104 may include an image sensor and associated optical elements such as a lens. The front camera 104 may provide zoom capabilities or may have a fixed focal length. In other examples, interchangeable lenses may be used with the front camera 104. The front camera 104 may have a variable mechanical aperture and mechanical and / or electronic shutters. The front camera 104 is also capable of being configured to capture still images, video images, or both. Additionally, the front camera 104 can represent, for example, a single field of view, a stereoscopic field of view, or a multi-field of view camera. The rear camera 112 may be arranged similarly or differently. Additionally, one or more of the front camera 104 and / or the rear camera 112 may be an array of one or more cameras.
[0039] One or more of the front camera 104 and / or the rear camera 112 may include or be associated with an illumination component that provides a light field to illuminate a target object. For example, the illumination component can provide a flash or constant illumination of the target object. The illumination component can also be configured to provide a light field that includes one or more of structured light, polarized light, and light having a particular spectral content. Other types of light fields that are known and used to recover a three-dimensional (3D) model from an object are possible within the context of the examples herein.
[0040] The computing device 100 may also include an ambient light sensor that can continuously or periodically determine the ambient brightness of the scene that the cameras 104 and / or 112 can capture. In some embodiments, the ambient light sensor can be used to adjust the display brightness of the display 106. Additionally, the ambient light sensor can be used to determine or assist in determining the exposure length of one or more of the cameras 104 or 112.
[0041] The computing device 100 can be configured to use the display 106 and the front camera 104 and / or the rear camera 112 to capture images of a target object. The captured images can be multiple still images or a video stream. Image capture can be triggered by activating the button 108, pressing a soft key on the display 106, or by some other mechanism. Depending on the embodiment, these images can be automatically captured at specific time intervals, for example, when the button 108 is pressed, when the target object is under appropriate lighting conditions, when the computing device 100 is moved a predetermined distance, or according to a predetermined capture schedule.
[0042] Figure 2FIG. 0 is a simplified block diagram showing some components of an example computing system 200. By way of example and not limitation, computing system 200 can be a cellular mobile phone (e.g., a smart phone), a computer (such as a desktop computer, laptop computer, tablet computer, or handheld computer), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a game console, a robotic device, a vehicle, or some other type of device. Computing system 200 can represent aspects of, for example, computing device 100.
[0043] As Figure 2 shown, computing system 200 can include a communication interface 202, a user interface 204, a processor 206, a data store 208, and a camera component 224, all of which can be communicatively linked together via a system bus, network, or other connection mechanism 210. Computing system 200 can be equipped with at least some image capture and / or image processing capabilities. It should be understood that computing system 200 can represent a physical image processing system, a particular physical hardware platform on which image sensing and / or processing applications operate in software, or some other combination of hardware and software configured to perform image capture and / or processing functions.
[0044] Communication interface 202 can allow computing system 200 to communicate with other devices, access networks, and / or transport networks using analog or digital modulation. Thus, communication interface 202 can facilitate circuit-switched and / or packet-switched communications, such as plain old telephone service (POTS) communications and / or Internet protocol (IP) or other packetized communications. For example, communication interface 202 can include a chipset and antenna arranged for wireless communication with a radio access network or access point. Additionally, communication interface 202 can take the form of or include a wired interface, such as Ethernet, universal serial bus (USB), or high-definition multimedia interface (HDMI) port. Communication interface 202 can also take the form of or include a wireless interface, such as Wi-Fi, Global Positioning System (GPS), or wide area wireless interface (e.g., WiMAX or 3GPP Long Term Evolution (LTE)). However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols can be used on communication interface 202. Additionally, communication interface 202 can include multiple physical communication interfaces (e.g., a Wi-Fi interface, interface, and a wide area wireless interface).
[0045] The user interface 204 can be used to allow the computing system 200 to interact with human or non - human users, such as receiving input from the user and providing output to the user. Thus, the user interface 204 can include input components such as a keypad, keyboard, touch - sensitive panel, computer mouse, trackball, joystick, microphone, etc. The user interface 204 can also include one or more output components such as a display screen, which can be combined with a touch - sensitive panel, for example. The display screen can be based on CRT, LCD, and / or LED technology, or other technologies now known or developed in the future. The user interface 204 can also be configured to generate audible output via a speaker, speaker jack, audio output port, audio output device, headphones, and / or other similar devices. The user interface 204 can also be configured to receive and / or capture audible speech, noise, and / or signals via a microphone and / or other similar devices.
[0046] In some examples, the user interface 204 can include a display that serves as a viewfinder for the still - camera and / or video - camera functions supported by the computing system 200. Additionally, the user interface 204 can include one or more buttons, switches, knobs, and / or dials that facilitate the configuration and focusing of the camera functions and the capture of images. Some or all of these buttons, switches, knobs, and / or dials can be implemented via the touch - sensitive panel.
[0047] The processor 206 can include one or more general - purpose processors (e.g., microprocessors) and / or one or more special - purpose processors (e.g., digital signal processor (DSP), graphics processing unit (GPU), floating - point unit (FPU), network processor, or application - specific integrated circuit (ASIC)). In some cases, the special - purpose processor can be capable of image processing, image alignment and merging images, and other possibilities. The data storage 208 can include one or more volatile and / or non - volatile storage components such as magnetic, optical, flash, or organic storage devices, and can be integrated with the processor 206 either wholly or in part. The data storage 208 can include removable components and / or non - removable components.
[0048] The processor 206 can be capable of executing program instructions 218 (e.g., compiled or uncompiled program logic and / or machine code) stored in the data storage 208 to perform the various functions described herein. Thus, the data storage 208 can include non - transitory computer - readable media having program instructions stored thereon that, when executed by the computing system 200, cause the computing system 200 to perform any of the methods, processes, or operations disclosed in this specification and / or the accompanying drawings. The execution of the program instructions 218 by the processor 206 can cause the processor 206 to use the data 212.
[0049] As an example, the program instructions 218 can include an operating system 222 (e.g., an operating system kernel, device drivers, and / or other modules) and one or more applications 220 installed on the computing system 200 (e.g., a camera function, an address book, email, web browsing, social networking, an audio-to-text function, a text translation function, and / or a game application). Similarly, the data 212 can include operating system data 216 and application data 214. The operating system data 216 can be primarily accessible by the operating system 222, and the application data 214 can be primarily accessible by one or more applications 220. The application data 214 can be arranged in a file system that is visible or hidden to the user of the computing system 200.
[0050] The applications 220 can communicate with the operating system 222 through one or more application programming interfaces (APIs). These APIs can facilitate, for example, the applications 220 reading and / or writing the application data 214, transmitting or receiving information via the communication interface 202, receiving and / or displaying information on the user interface 204, etc.
[0051] In some cases, the applications 220 can be simply referred to as "apps". Additionally, the applications 220 can be downloadable to the computing system 200 through one or more online app stores or app markets. However, the applications can also be installed on the computing system 200 in other ways (such as via a web browser or through a physical interface on the computing system 200 (e.g., a USB port)).
[0052] The camera component 224 can include, but is not limited to, an aperture, a shutter, a recording surface (e.g., photographic film and / or an image sensor), a lens, a shutter button, an infrared projector, and / or a visible light projector. The camera component 224 can include components configured to capture images in the visible spectrum (e.g., electromagnetic radiation having a wavelength of 380 nanometers to 700 nanometers) and / or components configured to capture images in the infrared spectrum (e.g., electromagnetic radiation having a wavelength of 701 nanometers to 1 millimeter), among other possibilities. The camera component 224 can be at least partially controlled by software executed by the processor 206.
[0053] Example 3D Photo System
[0054] Figure 3A An example system for generating a 3D photo based on a monocular image is illustrated. Specifically, the 3D photo system 300 can be configured to generate a modified image 370 based on the image 302. The 3D photo system 300 can include a depth model 304, a soft foreground visibility model 308, a soft background removal occlusion function 312, a refinement model 316, a pixel removal projector 322, and a pixel projector 330.
[0055] The image 302 can be a monocular / single field-of-view image including a plurality of pixels, and each pixel can be associated with one or more color values (e.g., a red-green-blue color image) and / or intensity values (e.g., a grayscale image). The image 302 can have an initial viewpoint from which the camera has captured the image. The initial viewpoint can be represented by initial camera parameters that indicate the spatial relationship between the camera and the world reference frame of the scene represented by the image 302. For example, a world reference frame can be defined such that it is initially aligned with the camera reference frame.
[0056] The modified image 370 can represent the same scene as the image 302 from an adjusted viewpoint that is different from the initial viewpoint. The adjusted viewpoint can be represented by adjusted camera parameters that are different from the initial camera parameters and indicate the adjusted spatial relationship between the camera and the world reference frame in the scene. Specifically, in the adjusted spatial relationship, the camera reference frame can be rotated and / or translated relative to the world reference frame. Thus, by generating one or more instances of the modified image 370 and viewing them in sequence, a 3D photo effect can be achieved due to the change in the viewpoint from which the scene represented by the image 302 is observed. This can allow the image 302 to appear more visually rich and / or more interactive by simulating the movement of the camera relative to the scene.
[0057] The depth model 304 can be configured to generate a depth image 306 based on the image 302. The image 302 can be represented as I ∈ R nx3 . The image 302 can thus have n pixels, and each pixel can be associated with 3 color values (e.g., red, green, and blue). The depth image 306 can be represented as D = Φ D (I), where D ∈ R nx1 and Φ D represents the depth model 304. The depth image 306 can thus have n pixels, and each pixel can be associated with a corresponding depth value. The depth values can be scaled to a predetermined range, such as 0 to 1 (i.e., D ∈ [0, 1] n)。In some embodiments, the depth value may represent parallax rather than metric depth and may be converted to metric depth by appropriate inversion of the parallax. The depth model 304 may be a machine learning model that has been trained to generate a depth image based on a monocular image. For example, the depth model 304 may represent the MiDaS model, as discussed in the article titled "Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer" by Ranflt et al. and published as arXiv:1907.01341 v3. Additionally or alternatively, the depth model 304 may be configured to use other techniques and / or algorithms to generate the depth image 306, some of which may be based on additional image data corresponding to the image 302.
[0058] The soft foreground visibility model 308 may be configured to generate a foreground visibility map 310 based on the depth image 306. The foreground visibility map 310 may alternatively be referred to as the soft foreground visibility map 310. The foreground visibility map 310 may be represented as A ∈ R n . The foreground visibility map 310 may thus have n pixels, each of which may be associated with a corresponding visibility value (which may be referred to as a foreground visibility value). The visibility values may be selected from and / or scaled to a predetermined range, such as 0 to 1 (i.e., A ∈ [0, 1] n ). In some cases, the visibility values may alternatively be represented as transparency values, since the visibility of the foreground and the transparency of the foreground are complementary to each other. That is, the nth transparency value may be represented as since Therefore, it should be understood that whenever discussed herein, the visibility values may be explicitly represented according to visibility or, equivalently, implicitly represented according to transparency.
[0059] In an example, a function may be considered "soft" because it is continuous and / or smooth along at least a particular input interval. A function may be considered "hard" because it is not continuous and / or smooth along at least a particular input interval. When a soft function is used instead of a hard function to generate a map or mask, the map or mask may be considered "soft". In particular, the map or mask itself may not be continuous and / or smooth, but may still be considered "soft" if it has been generated by a soft function.
[0060] The soft foreground visibility model 308 can be configured to generate a foreground visibility map 310 based on the depth gradients of the depth image 306 and possibly also based on the foreground alpha mask of the image 302. The visibility values of the foreground visibility map 310 can be relatively low at pixel positions representing the transition from the foreground features of the image 302 to the background features of the image 302 (and the corresponding transparency values can thus be relatively high). The visibility values of the foreground visibility map 310 can be relatively high at pixel positions that do not represent the transition from foreground features to background features (and the corresponding transparency values can thus be relatively low). Thus, the visibility values can allow features that are unoccluded due to a change in the viewing point to be seen at least partially through the content of the foreground. Additional details of the soft foreground visibility model 308 are shown and discussed with reference to Figure 4A 、 Figure 4B 、 Figure 4C and Figure 4D .
[0061] The soft background unoccluding function 312 can be configured to generate a background unoccluding mask 314 based on the depth image 306. The background unoccluding mask 314 can alternatively be referred to as the soft background unoccluding mask 314. The background unoccluding mask 314 can be represented as S ∈ R n . The background unoccluding mask 314 can thus have n pixels, each of which can be associated with a corresponding unoccluding value (which can be referred to as a background unoccluding value). The unoccluding values can be selected from a predetermined range and / or scaled to a predetermined range, such as 0 to 1 (i.e., S ∈ [0, 1] n ).
[0062] Each respective unoccluding value of the background unoccluding mask 314 can indicate the likelihood that the corresponding pixel of the image 302 will be unoccluded (i.e., become visible) due to a change in the initial viewing point of the image 302. Specifically, if there is an adjacent reference pixel (x i , y j ), such that (i) the depth value D(x, y) associated with the given pixel and (ii) the depth value D(x i , y j ) associated with the adjacent reference pixel, the depth difference between them is greater than the scaled distance between the respective positions of these pixels then the given pixel associated with the pixel position (x, y) can be unoccluded due to a change in the initial viewing point. That is, if then the given pixel at the position (x, y) can be unoccluded, where ρ is a scalar parameter that can be used when the depth image 306 represents disparity to scale the disparity to metric depth, where the value of ρ can be selected based on the maximum assumed camera movement between the initial viewing point and the modified viewing point, and where Thus, if the corresponding depth is relatively low (or the corresponding disparity is relatively high) compared to surrounding pixels, a given pixel can be more likely to be unoccluded by a change in the initial viewpoint because, when the viewpoint changes, features closer to the camera will appear to move more within image 302 than features further away.
[0063] Thus, the soft background unocclusion function 312 can be expressed as where γ is an adjustable parameter that controls the steepness of the hyperbolic tangent function, and i and j are each iterated through their respective ranges to compare a given pixel (x, y) with a predetermined number of neighboring pixels (x i , y j ), and a rectified linear unit (ReLU) or leaky ReLU can be applied to the output of the hyperbolic tangent function to make S (x,y) positive. Specifically, (x i , y j ) ∈ N (x,y) , where N (x,y) represents a fixed neighborhood of m pixels around the given pixel at position (x, y). The value of m and the shape of N (x,y) can be selected based on, for example, a determined computation time allocated to the background unocclusion mask 314. For example, N (x,y) can define a square, rectangle, approximately circular region, or two or more scan lines. The two or more scan lines can include vertical scan lines and / or horizontal scan lines, where the vertical scan lines include a predetermined number of pixels above and below the given pixel (x, y), and the horizontal scan lines include a predetermined number of pixels to the right and left of the given pixel (x, y).
[0064] Thus, the soft background unocclusion function 312 can be configured to determine a first plurality of differences for each respective pixel of the depth image 306, where each respective difference of the first plurality of differences is determined by subtracting (i) the corresponding depth value D(x, y) of a corresponding reference pixel located within a predetermined pixel distance of the respective pixel from the corresponding depth value D(x i , y j ) and (ii) a scaled number of pixels separating the respective pixel from the corresponding reference pixel . Additionally, the soft background unocclusion function 312 can be configured to determine a first maximum difference (i.e., of the first plurality of differences for each respective pixel of the depth image 306 and to determine an unocclusion value S (x,y) by applying the hyperbolic tangent function to the first maximum difference.
[0065] The inpainting model 316 can be configured to generate an inpainted image 318 and an inpainted depth image 320 based on the image 302, the depth image 306, and the background removal occlusion mask 314. The inpainted image 318 can be represented as and the inpainted depth image 320 can be represented as The inpainted image 318 and the inpainted depth image 320 can thus each have n pixels. The inpainting model 316 can implement a function where, represents the depth-wise concatenation of the inpainted image 318 and the inpainted depth image 320, and Θ(·) represents the inpainting model 316.
[0066] Specifically, the inpainted image 318 can represent the image 302 inpainted according to the background removal occlusion mask 314, and the inpainted depth image 320 can represent the depth image 306 inpainted according to the background removal occlusion mask 314. A change in the initial viewpoint can cause a significant movement of one or more foreground regions, resulting in the corresponding background regions (indicated by the background removal occlusion mask 314) becoming visible. Since the image 302 and the depth image 306 do not contain the pixel values of these background regions, the inpainting model 316 can be used to generate these pixel values. Thus, the inpainted image 318 and the inpainted depth image 320 can be considered to form a background layer, while the image 302 and the depth image 306 form a foreground layer. These layers can be combined as part of generating the modified image 370.
[0067] The inpainting model 316 can be trained to inpaint the background regions by using the values of other background regions in the matching image 302 rather than the values of the foreground regions of the matching image 302. Additionally or alternatively, the inpainting model 316 can be configured to generate the inpainted image 318 and the inpainted depth image 320 in parallel and / or concurrently. For example, both the inpainted image 318 and the inpainted depth image 320 can be generated by a single pass / iteration of the inpainting model 316. Thus, the output of the inpainting model 316 can be the depth-wise concatenation of the inpainted image 318 and the inpainted depth image 320, and each of these inpainted images can thus be represented by the corresponding one or more channels of the output.
[0068] Both the inpainted image 318 and the inpainted depth image 320 can be based on the image 302 and the depth image 306 such that the pixel values of the inpainted portion of the inpainted image 318 are consistent with the depth values of the corresponding inpainted portion of the inpainted depth image 320. That is, the inpainting model 316 can be configured to consider both depth information and color / intensity information when generating values for the occluded background regions, resulting in inpainted regions that exhibit consistency in depth and color / intensity. The depth information can allow the inpainting model 316 to distinguish between the foreground and background regions of the image 302 and thus select the appropriate portions from which to borrow values for inpainting. Thus, the inpainting model 316 can be considered depth-aware. Additional details of the inpainting model 316 are shown in Figure 5A and discussed with reference thereto.
[0069] The pixel de-projector 322 can be configured to generate (i) a 3D representation 324 of the foreground visibility map based on the foreground visibility map 310 and the depth image 306, (ii) a 3D representation 326 of the foreground based on the image 302 and the depth image 306, and (iii) a 3D representation 328 of the background based on the inpainted image 318 and the inpainted depth image 320. Specifically, the pixel de-projector 322 can be configured to de-project each respective pixel of the foreground visibility map 310, the image 302, and the inpainted image 318 to generate corresponding 3D points representing the respective pixels within the 3D camera reference frame. Thus, the 3D points of the 3D representation of the foreground visibility map can be represented as The 3D points of the 3D representation 326 of the foreground can be represented as The 3D points of the 3D representation 326 of the background can be represented as where p represents the respective pixel, represents the homogenous augmentation of the coordinates of the respective pixel, M represents the intrinsic camera matrix with the principal point at the center of the image 302 D(p) represents the depth value (as opposed to the inverse depth value) corresponding to the respective pixel as represented in the depth image 306, represents the depth value (as opposed to the inverse depth value) corresponding to the respective pixel as represented by the inpainted depth image 320, and the camera reference frame is assumed to be aligned with the world reference frame. Since V(p) = F(p), a single set of 3D points can be determined for and shared by the 3D representations 324 and 326.
[0070] Based on the corresponding 3D points of the pixels determined to represent the foreground visibility map 310, the image 302, and the inpainted image 318, the pixel remover projector 322 can be configured to connect the 3D points corresponding to adjacent pixels, thereby generating a corresponding polygon mesh (e.g., a triangular mesh) for each of the foreground visibility map 310, the image 302, and the inpainted image 318. Specifically, if the pixel corresponding to a given 3D point and the pixels corresponding to one or more other 3D points are adjacent to each other within the corresponding image, the given 3D point can be connected to the one or more other 3D points.
[0071] The pixel remover projector 322 can also be configured to texture each corresponding polygon mesh based on the pixel values of the corresponding image or map. Specifically, the polygon mesh of the foreground visibility map 3D representation 324 can be textured based on the corresponding values of the foreground visibility map 310, the polygon mesh of the foreground 3D representation 326 can be textured based on the corresponding values of the image 302, and the polygon mesh of the background 3D representation 328 can be textured based on the corresponding values of the inpainted image 318. When a single set of 3D points is determined for and shared by the 3D representations 324 and 326, the 3D representations 324 and 326 can also be represented using a single polygon mesh including multiple channels, where one or more first channels represent the values of the foreground visibility map 310 and one or more second channels represent the values of the image 302. Additionally or alternatively, the 3D representations 324, 326, and 328 can include point clouds, hierarchical representations (e.g., hierarchical depth images), multi-plane images, and / or implicit representations (e.g., neural radiance fields (NeRF)), among other possibilities.
[0072] The pixel projector 330 can be configured to generate a modified image 370 based on the 3D representations 324, 326, and 328. Specifically, the pixel projector 330 can be configured to perform a rigid transformation T that represents the translation and / or rotation of the camera relative to the camera pose associated with the initial viewpoint. Thus, the rigid transformation T can represent the adjusted camera parameters that define the adjusted viewpoint for the modified image 370. For example, the rigid transformation T can be represented relative to the world reference frame associated with the initial viewpoint of the image 302. The value of T can be adjustable and / or iterable to generate multiple different instances of the modified image 370, each instance representing the image 302 from a corresponding adjusted viewpoint.
[0073] Specifically, based on the rigid transformation T, the pixel projector 330 can be configured to (i) project the foreground visibility map 3D representation 324 to generate a modified foreground visibility map A T , (ii) project the foreground 3D representation 326 to generate a foreground image IT , and (iii) projecting the background 3D representation 328 to generate a background image Thus, the modified foreground visibility map A T can represent the foreground visibility map 310 from the modified viewing point, the foreground image I T can represent the image 302 from the modified viewing point, and the background image can represent the inpainted image 318 from the modified viewing point.
[0074] The pixel projector 330 can be further configured to, according to the modified foreground visibility map A T combine the foreground image I T with the background image . For example, the pixel projector 330 can be configured to generate a modified image 370 according to , where represents the modified image 370. Thus, the modified image 370 can be mainly based on the image 302 where the value of the foreground visibility map 310 is relatively high, and can be mainly based on the inpainted depth image 320 where the value of the foreground visibility map 310 is relatively low. The visibility values of the modified visibility map A T can be selected from and / or scaled to the same predetermined range as the values of the foreground visibility map 310. For example, the range of the values of the modified foreground visibility map can be from zero to a predetermined value (e.g., A T ∈ [0, 1] n ). Thus, the modified image 370 can be based on the sum of (i) the first product of the foreground image and the modified foreground visibility map (i.e., A T I T ) and (ii) the second product of the background image and the difference between the predetermined value and the modified foreground visibility map (i.e., ).
[0075] In other embodiments, the pixel projector 330 can be configured to generate a modified image 370 according to , where θ(·) represents a machine learning model that has been trained to combine the foreground image I T with the background image T according to A . For example, the machine learning model can be trained using multiple sets of training images, each set of training images including multiple different views of a scene and thus representing the ground truth pixel values of the occluded regions removed.
[0076] Figure 3B , Figure 3C and Figure 3D contain examples of images and maps generated by the 3D photo system 300. Specifically,Figure 3B Includes an image 380 (an example of the image 302) and a corresponding depth image 382 (an example of the depth image 306). Figure 3C Includes a foreground visibility map 384 (an example of the foreground visibility map 310) and a background removal occlusion mask 386 (an example of the background removal occlusion mask 314), each of which corresponds to the image 380 and the depth image 382. Figure 3D Includes an inpainted image 388 (an example of the inpainted image 318), an inpainted depth image 390 (an example of the inpainted depth image 320), and a modified image 392 (an example of the modified image 370), each of which corresponds to the image 380 and the depth image 382.
[0077] Specifically, the image 380 contains a giraffe (foreground feature / object) relative to a background (background feature / object) of grass and trees. The depth image 382 indicates that the giraffe is closer to the camera than most of the background. In the foreground visibility map 384, the maximum visibility value is shown in white, the minimum visibility value is shown in black, and the values in between are shown using gray levels. Thus, as described above, the areas around the giraffe, the top of the grass, and the top of the tree line are shown in black and / or dark gray levels, indicating that the depth gradient is relatively high in these areas and these areas are thus relatively more transparent for inpainting background pixel values after projection in three dimensions. In the background removal occlusion mask 386, the maximum occlusion removal value is shown in white, the minimum occlusion removal value is shown in black, and the values in between are shown using gray levels. Thus, the areas behind the giraffe and behind the top of the grass are shown in white and / or light gray levels, indicating that these areas may be unoccluded (i.e., become visible) by a change in the viewpoint from which the image 380 is rendered.
[0078] In the inpainted image 388 and the inpainted depth image 390, new values corresponding to the background portions of the image 380 and the depth image 382 are used instead of the values corresponding to the foreground portions of the image 380 and the depth image 382 to inpaint the portions of the image 380 and the depth image 382 that previously represented the giraffe and the top of the grass. Thus, the images 388 and 390 respectively represent the images 380 and 382, where at least some of the foreground content is replaced with newly synthesized background content that is contextually coherent (in terms of color / intensity and depth) with the background portions of the image 380 and 382.
[0079] During the rendering of the modified image 392, the content of the image 380 (foreground layer) can be combined with the inpainting content of the inpainting image 388 (background layer) to fill the occluded regions revealed by the modification of the initial viewpoint of the image 380. After re-projecting the images 380 and 388 and the mapping 384 into the modified viewpoint, the blending ratio between the images 380 and 388 is specified by the foreground visibility mapping 384. Specifically, the blending ratio (after re-projecting into the modified viewpoint) is specified by the corresponding pixels of the foreground visibility mapping 384, where the corresponding pixels are determined after de-projecting the image 380, the inpainting image 388, and the foreground visibility mapping 384 to generate the corresponding 3D representations and the projection of these 3D representations into the modified viewpoint. Thus, the modified image 392 represents the content of the image 380 from a modified viewpoint different from the initial viewpoint of the image 380. The modified image 392 shown is not an exact rectangle because the images 380 and 388 may not include content outside the outer boundaries of the image 380. In some cases, the modified image 392 and / or one or more of the images it is based on can undergo further inpainting so that the modified image 392 is rectangular even after the viewpoint modification.
[0080] Example soft foreground visibility model
[0081] Figure 4A Illustrating aspects of the soft foreground visibility model 308, which can be configured to generate a foreground visibility mapping 310 based on a depth image 306 and an image 302. Specifically, the soft foreground visibility model 308 can include a soft background occlusion function 332, an inverse function 336, a soft foreground visibility function 340, a foreground segmentation model 350, a mask model 354, a dilation function 358, a difference operator 362, and a multiplication operator 344.
[0082] The soft foreground visibility function 340 can be configured to generate a depth-based foreground visibility mapping 342 based on the depth image 306. Specifically, the soft foreground visibility function can be expressed as where A DEPTH represents the depth-based foreground visibility mapping 342, represents a gradient operator (e.g., Sobel gradient operator), and β is an adjustable scalar parameter. Thus, the foreground visibility value of a given pixel p in the depth image 306 can be an exponent based on the depth gradient determined by applying the gradient operator to the given pixel p and one or more surrounding pixels.
[0083] In some embodiments, the depth-based foreground visibility map 342 can be equal to the foreground visibility map 310. That is, the soft background occlusion function 332, the inverse function 336, the foreground segmentation model 350, the masking model 354, the dilation function 358, the difference operator 362, and / or the multiplication operator 344 can be omitted from the soft foreground visibility model 308. However, in some cases, the depth image 306 may not accurately represent some thin and / or high-frequency features such as hair or fur (e.g., due to the properties of the depth model 304). To improve the representation of these thin and / or high-frequency features in the modified image 370, the depth-based foreground visibility map 342 can be combined with the inverse background occlusion mask 338 and the mask-based foreground visibility map 364 to generate the foreground visibility map 310. Since the depth-based foreground visibility map 342 is soft, the depth-based foreground visibility map 342 can be combined with occlusion-based foreground segmentation techniques.
[0084] The foreground segmentation model 350 can be configured to generate a foreground segmentation mask 352 based on the image 302. The foreground segmentation mask 352 can distinguish the foreground features of the image 302 (e.g., represented using a binary pixel value of 1) from the background features of the image 302 (e.g., represented using a binary pixel value of 0). The foreground segmentation mask 352 can distinguish the foreground from the background, but may not accurately represent the thin and / or high-frequency features of the foreground objects represented in the image 302. The foreground segmentation model 350 can be implemented, for example, using the model discussed in the article titled "U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection" by Qin et al. and published as arXiv:2005.09007.
[0085] The masking model 354 can be configured to generate a foreground alpha mask 356 based on the image 302 and the foreground segmentation mask 352. The foreground alpha mask 356 can distinguish the foreground features of the image 302 from the background features of the image 302, and can do so using a per-pixel binary value or a per-pixel grayscale intensity value. Different from the foreground segmentation mask 352, the foreground alpha mask 356 can accurately represent the thin and / or high-frequency features (e.g., hair or fur) of the foreground objects represented in the image 302. For example, the masking model 354 can be implemented using the model discussed in the article titled "F, B, Alpha Matting" by Forte et al. and published as arXiv:2003.07711.
[0086] To incorporate the foreground alpha mask 356 into the foreground visibility map 310, the foreground alpha mask 356 may undergo further processing. Specifically, the foreground visibility map 310 is expected to include low visibility values around the foreground object boundary, but is not expected to include low visibility values in regions not near the foreground object boundary. However, the foreground alpha mask 356 includes low visibility values for all background regions. Thus, the dilation function 358 may be configured to generate a dilated foreground alpha mask 360 based on the foreground alpha mask 356, and then the dilated foreground alpha mask 360 may be subtracted from the foreground alpha mask 356 as indicated by the differential operator 362 to generate a mask-based foreground visibility map 364. The mask-based foreground visibility map 364 may thus include low visibility values around the foreground object boundary but not in background regions not near the foreground object boundary.
[0087] Additionally, the soft background occlusion function 332 may be configured to generate a background occlusion mask 334 based on the depth image 306. The background occlusion mask 334 may alternatively be referred to as the soft background occlusion mask 334. The background occlusion mask 334 may be represented as and may thus have n pixels, each of which may be associated with a corresponding occlusion value (which may be referred to as a background occlusion value). The occlusion values may be selected from and / or scaled to a predetermined range, such as 0 to 1 (i.e.,
[0088] Each respective occlusion value of the background occlusion mask 334 may indicate the likelihood that the corresponding pixel of the image 302 will be occluded (i.e., completely covered) by a change in the initial viewpoint of the image 302. The calculation of the background occlusion mask 334 may be similar to the calculation of the background removal occlusion mask 314. Specifically, if there exists an adjacent reference pixel (x i , y j ) such that (i) the depth value D(x, y) associated with the given pixel and (ii) the depth value D(x i , y j ) associated with the adjacent reference pixel have a depth difference less than the scaled distance between the respective positions of these pixels then the given pixel associated with the pixel position (x, y) may be occluded due to a change in the initial viewpoint. That is, if then the given pixel at the position (x, y) may be occluded. Thus, the soft background occlusion function 332 may be represented as where, (x i , y j ) ∈ N (x,y)Cause i and j to each iterate through their respective ranges to compare a given pixel (x, y) with a predetermined number of adjacent pixels (x i , y j ), and the ReLU or leaky ReLU can be applied to the output of the hyperbolic tangent function to make positive.
[0089] Thus, the soft background occlusion function 332 can be configured to determine a second plurality of differences for each respective pixel of the depth image 306, where each respective difference in the second plurality of differences is obtained by subtracting (i) the corresponding depth value D(x i , y j ) of the corresponding pixel from (ii) the corresponding depth value D(x, y) of the respective pixel and the scaled number of pixels separating the respective pixel from the corresponding reference pixel within a predetermined pixel distance of the respective pixel is determined. In addition, the soft background removal occlusion function 312 can be configured to: determine the second largest difference (i.e., ) among the second plurality of differences for each respective pixel of the depth image 306, and determine the removal occlusion value
[0090] The inverse function 336 can be configured to generate an inverse background occlusion mask 338 based on the background occlusion mask 334. For example, when , the inverse function 336 can be expressed as
[0091] The foreground visibility map 310 can be determined based on the product of (i) the depth-based foreground visibility map 342, (ii) the mask-based foreground visibility map 364, and (iii) the inverse background occlusion mask 338, as indicated by the multiplication operator 344. Determining the foreground visibility map 310 in this way can allow for the representation of thin and / or high-frequency features of the foreground object while taking into account depth discontinuities and reducing and / or avoiding the "leakage" of the mask-based foreground visibility map 364 onto too much of the background. Thus, the net effect of this multiplication is to incorporate the thin and / or high-frequency features around the foreground object boundary represented by the mask-based foreground visibility map 364 into the depth-based foreground visibility map 342.
[0092] Figure 4B , Figure 4C and Figure 4D include examples of images and maps used and / or generated by the soft foreground visibility model 308. Specifically, Figure 4B includes the image 402 (an example of the image 302) and the corresponding depth image 406 (an example of the depth image 306). Figure 4Cincluding a depth-based foreground visibility map 442 (an example of the depth-based foreground visibility map 342) and a foreground alpha mask 456 (an example of the foreground alpha mask 356), each corresponding to the image 402 and the depth image 406. Figure 4D including a mask-based foreground visibility map 462 (an example of the mask-based foreground visibility map 364) and a foreground visibility map 410 (an example of the foreground visibility map 310), each corresponding to the image 402 and the depth image 406.
[0093] Specifically, the image 402 includes a lion (foreground object) relative to a blurred background. The depth image 406 indicates that the lion is closer to the camera than the background, with the lion's muzzle being closer than other parts of the lion. The depth image 402 does not accurately represent the depth of some of the hairs in the lion's mane. In the depth-based foreground visibility map 442 and the foreground visibility map 410, the maximum visibility value is shown in white, the minimum visibility value is shown in black, and the values between them are shown using a gray scale. In the foreground alpha mask 456 and the mask-based foreground visibility map 462, the maximum alpha value is shown in white, the minimum alpha value is shown in black, and the values between them are shown using a gray scale. The maximum alpha value corresponds to the maximum visibility value because both indicate that the foreground content is completely opaque.
[0094] Thus, in the depth-based foreground visibility map 442, the contour regions representing the transitions between the lion and the background and between the lion's muzzle and other parts of the lion are shown in a darker gray scale, indicating that the depth gradient is relatively high in these contour regions and that these contour regions are relatively more transparent for restoring background pixel values, while other regions are shown in a lighter gray scale. In the foreground alpha mask 456, the lion is shown in a lighter gray scale while the background is shown in a darker gray scale. Different from the depth-based foreground visibility map 442, the foreground alpha mask 456 details the hairs of the lion's mane. In the mask-based foreground visibility map 462, the lion is shown in a lighter gray scale, the contour regions around the lion are shown in a darker gray scale, and the background portion outside the contour regions is shown in a lighter gray scale. The size of the contour region can be controlled by controlling the number of pixels by which the dilation function 358 dilates the foreground alpha mask 356.
[0095] In the foreground visibility map 410, most of the lion is shown in a lighter gray level, the contour regions around and around the muzzle of the lion are shown in a darker gray level, and the background portion outside the contour regions is shown in a lighter gray level. The contour region of the lion in the foreground visibility map 410 is smaller than the contour region of the lion in the mask-based foreground visibility map 462 and includes a detailed representation of the hair of the lion's mane. In addition, the foreground visibility map 410 includes a contour region around the muzzle of the lion that is present in the depth-based foreground visibility map 442 but not in the mask-based foreground visibility map 462. The foreground visibility map 410 thus represents the depth-based discontinuities of the image 402 and the mask-based thin and / or high-frequency features of the lion's mane.
[0096] Example Training Systems and Operations
[0097] Figure 5A FIG. illustrates an example training system 500 that can be used to train one or more components of the 3D photo system 300. Specifically, the training system 500 can include a depth model 304, a soft background masking function 332, an inpainting model 316, a discriminator model 512, an adversarial loss function 522, a reconstruction loss function 508, and a model parameter adjuster 526. The training system 500 can be configured to determine updated model parameters 528 based on a training image 502. The training image 502 and the training depth image 506 can be similar to the image 302 and the depth image 306, respectively, but can be determined at training time rather than at inference time. The depth model 304 and the soft background masking function 332 have been discussed with reference to Figure 3A and Figure 4A respectively.
[0098] The inpainting model 316 can be trained to inpaint regions of the image 302 that can be unoccluded by a change in viewpoint as indicated by the background removal occlusion mask 314. However, in some cases, the ground truth pixel values of these unoccluded regions may not be available (e.g., the multi-views of the training image 502 may not be available). Thus, instead of using the training background removal occlusion mask corresponding to the training image 502 to train the inpainting model 316, the training background masking mask 534 can be used to train the inpainting model 316. The training background masking mask 534 can be similar to the background masking mask 334 but can be determined at training time rather than at inference time. Since the ground truth pixel values of the regions that may be occluded by foreground features due to a change in the initial viewpoint are available in the training image 502, this technique can allow the inpainting model 316 to be trained without relying on a multi-view training image dataset. In cases where the ground truth data of the unoccluded regions is available, the removal occlusion mask can be used to additionally or alternatively train the inpainting model.
[0099] Specifically, the inpainting model 316 can be configured to generate inpainted training images 518 and inpainted training depth images 520 based on training images 502, training depth images 506, and training background occlusion masks 534. The inpainted training images 518 and inpainted training depth images 520 can be similar to the inpainted images 318 and inpainted depth images 320, respectively. In some cases, a stroke-shaped inpainting mask (not shown) can also be used to train the inpainting model 316, thereby allowing the inpainting model 316 to learn to inpaint thin or small objects.
[0100] The quality of the inpainting model 316 in inpainting the missing regions of images 502 and 506 can be quantified using a reconstruction loss function 508 and / or an adversarial loss function. The reconstruction loss function 508 can be configured to generate reconstruction loss values 510 by comparing the training images 502 and 506 with the inpainted images 518 and 520, respectively. For example, the reconstruction loss function 508 can be configured to determine (i) a first L-1 distance between the training image 502 and the inpainted training image 518 and (ii) a second L-1 distance between the training depth image 506 and the inpainted training depth image 520, at least with respect to the regions that have been inpainted (as indicated by the training background occlusion mask 534 and / or the stroke-shaped inpainting mask). The reconstruction loss function 508 can thus encourage the inpainting model 316 to generate (i) inpainting intensity values and inpainting depth values that are consistent with each other and (ii) based on the background features (rather than the foreground features) of the training image 502.
[0101] The discriminator model 512 can be configured to generate a discriminator output 514 based on the inpainted training images 518 and inpainted training depth images 520. Specifically, the discriminator output 514 can indicate whether the discriminator model 512 estimates that the inpainted training images 518 and inpainted training depth images 520 are generated by the inpainting model 316 or as ground truth images that have not yet been generated by the inpainting model 316. Thus, the inpainting model 316 and the discriminator model 512 can implement an adversarial training architecture. Therefore, the adversarial loss function 522 can include, for example, a hinge adversarial loss and can be configured to generate an adversarial loss value 524 based on the discriminator output 514. The adversarial loss function 522 can thus encourage the inpainting model 316 to generate inpainting intensity values and inpainting depth values that look realistic and thus accurately mimic natural scenes.
[0102] The model parameter adjuster 526 can be configured to determine updated model parameters 528 based on the reconstruction loss value 510 and the adversarial loss value 524 (and any other loss values that can be determined by the training system 500). The model parameter adjuster 526 can be configured to determine a total loss value based on a weighted sum of these loss values, where the relative weights of the corresponding loss values can be adjustable training parameters. The updated model parameters 528 can include one or more updated parameters of the inpainting model 316, which can be represented as ΔΘ. In some embodiments, the updated model parameters 528 can additionally include one or more other components for the 3D photo system 300 (such as other possibilities like the depth model 304 or the soft foreground visibility model 308) and one or more updated parameters for the discriminator model 512.
[0103] The model parameter adjuster 526 can be configured to determine the updated model parameters 528 by, for example, determining the gradient of the total loss function. Based on this gradient and the total loss value, the model parameter adjuster 526 can be configured to select updated model parameters 528 that are expected to decrease the total loss value and thus improve the performance of the 3D photo system 300. After applying the updated model parameters 528 to at least the inpainting model 316, the operations discussed above can be repeated to calculate another instance of the total loss value, and based on this, another instance of the updated model parameters 528 can be determined and applied to at least the inpainting model 316 to further improve its performance. This training of the components of the 3D photo system 300 can be repeated until, for example, the total loss value is reduced below a target threshold loss value.
[0104] Figure 5B including a background occlusion mask 544 (an example of the background occlusion mask 334 and the training background occlusion mask 534), which corresponds to Figure 3B a portion of the image 380 and the depth image 382. In the background occlusion mask 544, the maximum occlusion value is shown in white, the minimum occlusion value is shown in black, and the values between them are shown using gray levels. Thus, the areas around the giraffe are shown in white or light gray levels, indicating that these areas may be occluded by a change in the viewpoint from which the image 380 is rendered.
[0105] Additional example operations
[0106] Figure 6 A flowchart illustrating operations related to generating 3D effects based on monocular images. The operations can be performed by various computing devices 100, computing systems 200, 3D photo systems 300, and / or training systems 500, among other possibilities. Figure 6Embodiments can be simplified by removing any one or more of the features shown therein. Additionally, these embodiments can be combined with features, aspects, and / or embodiments of any one of the previous figures or features, aspects, and / or embodiments described elsewhere herein.
[0107] Block 600 can involve obtaining a monocular image with an initial viewpoint.
[0108] Block 602 can involve determining a depth image including a plurality of pixels based on the monocular image. Each respective pixel of the depth image can have a corresponding depth value.
[0109] Block 604 can involve determining, for each respective pixel of the depth image, a corresponding depth gradient associated with the respective pixel of the depth image.
[0110] Block 606 can involve determining a foreground visibility map that includes visibility values inversely proportional to the corresponding depth gradients for each respective pixel of the depth image.
[0111] Block 608 can involve determining a background removal occlusion mask based on the depth image, the background removal occlusion mask including occlusion values for each respective pixel of the depth image that indicate the likelihood that the corresponding pixel of the monocular image will be unoccluded by a change in the initial viewpoint.
[0112] Block 610 can involve (i) generating a repaired image by using a inpainting model to inpaint a portion of the monocular image according to the background removal occlusion mask, and (ii) generating a repaired depth image by using the inpainting model to inpaint a portion of the depth image according to the background removal occlusion mask.
[0113] Block 612 can involve (i) generating a first 3D representation of the monocular image based on the depth image, and (ii) generating a second 3D representation of the repaired image based on the repaired depth image.
[0114] Block 614 can involve generating a modified image with an adjusted viewpoint different from the initial viewpoint by combining the first 3D representation with the second 3D representation according to the foreground visibility map.
[0115] In some embodiments, the foreground visibility map can be determined by a soft foreground visibility function that is continuous and smooth along at least a first interval.
[0116] In some embodiments, determining the foreground visibility map includes: determining, for each respective pixel of the depth image, a corresponding depth gradient associated with the respective pixel by applying a gradient operator to the depth image, and determining a visibility value for each respective pixel of the depth image based on an exponent of the corresponding depth gradient associated with the respective pixel.
[0117] In some embodiments, a background removal occlusion mask may be determined by a soft background removal occlusion function that is continuous and smooth along at least a second interval.
[0118] In some embodiments, determining a background removal occlusion mask may include: determining a first plurality of differences for each respective pixel of a depth image. Each respective difference of the first plurality of differences may be determined by subtracting (i) the corresponding depth value of a corresponding reference pixel located within a predetermined pixel distance of the respective pixel and (ii) a scaled number of pixels separating the respective pixel from the corresponding reference pixel from the corresponding depth value of the respective pixel. Determining the background removal occlusion mask may further include: determining an occlusion value for each respective pixel of the depth image based on the first plurality of differences.
[0119] In some embodiments, determining the occlusion value may include: determining a first maximum difference among the first plurality of differences for each respective pixel of the depth image; and determining the occlusion value for each respective pixel of the depth image by applying a hyperbolic tangent function to the first maximum difference.
[0120] In some embodiments, the corresponding reference pixel of a respective difference may be selected from: (i) a vertical scan line including a predetermined number of pixels above and below the respective pixel, or (ii) a horizontal scan line including a predetermined number of pixels to the right and left of the respective pixel.
[0121] In some embodiments, determining a foreground visibility map may include: determining a depth-based foreground visibility map that includes a visibility value for each respective pixel of the depth image, the visibility value being inversely proportional to a corresponding depth gradient associated with the respective pixel of the depth image; and determining a foreground alpha mask based on and corresponding to a monocular image. Determining the foreground visibility map may further include: determining a mask-based foreground visibility map based on a difference between (i) the foreground alpha mask and (ii) a dilation of the foreground alpha mask, and determining the foreground visibility map based on a product of (i) the depth-based foreground visibility map and (ii) the mask-based foreground visibility map.
[0122] In some embodiments, determining the foreground alpha mask may include: determining a foreground segmentation based on the monocular image by a foreground segmentation model, and determining the foreground alpha mask by a mask model based on the foreground segmentation and the monocular image.
[0123] In some embodiments, determining the foreground visibility map may include determining a background occlusion mask based on a depth image, the background occlusion mask including an occlusion value for each respective pixel of the depth image, the occlusion value indicating the likelihood that the corresponding pixel of the monocular image will be occluded by a change in the initial viewpoint. Determining the foreground visibility map may also include determining the foreground visibility map based on the product of (i) a depth-based foreground visibility map, (ii) a mask-based foreground visibility map, and (iii) the inverse of the background occlusion mask.
[0124] In some embodiments, determining the background occlusion mask may include determining a second plurality of differences for each respective pixel of the depth image. Each respective difference of the second plurality of differences may be determined by subtracting (i) the corresponding depth value of the respective pixel from (ii) the corresponding depth value of a corresponding reference pixel and a scaled number of pixels separating the respective pixel from the corresponding reference pixel within a predetermined pixel distance of the respective pixel. Determining the background occlusion mask may also include determining an occlusion value for each respective pixel of the depth image based on the second plurality of differences.
[0125] In some embodiments, determining the occlusion value may include: determining a second maximum difference of the second plurality of differences for each respective pixel of the depth image, and determining the occlusion value for each respective pixel of the depth image by applying the hyperbolic tangent function to the second maximum difference.
[0126] In some embodiments, the monocular image may include foreground features and background features. The inpainting model may have been trained to inpaint (i) an occluded background region of the monocular image having intensity values that match the background features and are independent of the foreground features, and (ii) a corresponding occluded background region of the depth image having depth values that match the background features and are independent of the foreground features.
[0127] In some embodiments, the inpainting model may have been trained such that the intensity values of the occluded regions of the inpainted image are contextually consistent with the corresponding depth values of the corresponding occluded background regions of the inpainted depth image.
[0128] In some embodiments, the inpainting model may have been trained via a training process that includes obtaining training monocular images and determining training depth images based on the training monocular images. The training process may further include determining a training background occlusion mask based on the training depth images, the training background occlusion mask including an occlusion value for each respective pixel of the training depth images, the occlusion value indicating the likelihood that the corresponding pixel of the training monocular image will be occluded by a change in the original viewpoint of the training monocular image. The training process may additionally include: (i) generating inpainted training images by using the inpainting model to inpaint portions of the training monocular images according to the training background occlusion mask; and (ii) generating inpainted training depth images by using the inpainting model to inpaint portions of the training depth images according to the training background occlusion mask. The training process may further include: determining a loss value by applying a loss function to the inpainted training images and the inpainted training depth images; and adjusting one or more parameters of the inpainting model based on the loss value.
[0129] In some embodiments, the loss value may be based on one or more of the following: (i) an adversarial loss value determined based on processing the inpainted training images and the inpainted training depth images by a discriminator model, or (ii) a reconstruction loss value determined based on comparing the inpainted training images with the training monocular images and comparing the inpainted training depth images with the training depth images.
[0130] In some embodiments, generating a modified image may include generating a 3D foreground visibility map corresponding to a first 3D representation based on a depth image. Generating the modified image may further include: (i) generating a foreground image by projecting the first 3D representation based on an adjusted viewpoint, (ii) generating a background image by projecting a second 3D representation based on the adjusted viewpoint, and (iii) generating a modified foreground visibility map by projecting the 3D foreground visibility map based on the adjusted viewpoint. Generating the modified image may further include combining the foreground image and the background image according to the modified foreground visibility map.
[0131] In some embodiments, the values of the modified foreground visibility map may range from zero to a predetermined value. Combining the foreground image and the background image may include determining the sum of (i) a first product of the foreground image and the modified foreground visibility map and (ii) a second product of the background image and the difference between the predetermined value and the modified foreground visibility map.
[0132] In some embodiments, generating the first 3D representation and the second 3D representation may include (i) generating a first plurality of 3D points by unprojecting pixels of a monocular image based on a depth image, and (ii) generating a second plurality of 3D points by unprojecting pixels of a inpainted image based on an inpainted depth image, and (i) generating a first polygon mesh by interconnecting corresponding subsets of the first plurality of 3D points corresponding to adjacent pixels of the monocular image, and (ii) generating a second polygon mesh by interconnecting corresponding subsets of the second plurality of 3D points corresponding to adjacent pixels of the inpainted image. Generating the first 3D representation and the second 3D representation may further include: (i) applying one or more first textures to the first polygon mesh based on the monocular image, and (ii) applying one or more second textures to the second polygon mesh based on the inpainted image.
[0133] Example performance metrics
[0134] Figure 7A 、 Figure 7B and Figure 7C each include respective tables containing the results of testing various 3D photo models against corresponding datasets. Specifically, Figure 7A 、 Figure 7B and Figure 7C each compare the performance of the SynSin model (described in the article titled "SynSin: End-to-end View Synthesis from a Single Image" by Wiles et al. and published as "arXiv:1912.08804"), the SMPI model (described in the article titled "Single-View View Synthesis with Multiplane Images" by Tucker et al. and published as arXiv:2004.11364), the 3D photo model (described in the article titled "3D Photography using Context-aware Layered Depth Inpainting" by Shih et al. and published as arXiv:2004.04727), and the 3D photo system 300 as described herein.
[0135] Performance is quantified by the learned perceptual image patch similarity (LPIPS) metric (lower values indicate better performance), the peak signal-to-noise ratio (PSNR) (higher values indicate better performance), and the structural similarity index measure (SSIM) (higher values indicate better performance). Figure 7A 、 Figure 7B andFigure 7C The results in correspond to the RealEvent10k image dataset (discussed in the article titled "Stereo Magnification: Learning View Synthesis using Multiplane Images" created by Zhou et al. and published as arXiv: 1805.09817), the Dual-Pixel image dataset (discussed in the article titled "Learning Single Camera Depth Estimation using Dual-Pixels" created by Garg et al. and published as arXiv: 1904.05822), and the Human Model Challenger image dataset (discussed in the article titled "Learning the Depths of Moving People by Watching Frozen People" created by Li et al. and published as arXiv: 1904.11111).
[0136] In Figure 7A "T-5" and "T = 10" indicate that the corresponding metrics are calculated for the test image with respect to the modified viewpoints corresponding to the images that are 5 and 10 time steps away from the test image, respectively. In Figure 7B the corresponding metrics are calculated for the test image with respect to the modified viewpoints corresponding to the four images captured simultaneously with the test image. In Figure 7C the corresponding metrics are calculated for the test image with respect to the modified viewpoints corresponding to the four images that are after the test image and consecutive with the test image. As indicated by the shaded patterns in the respective bottommost rows of the tables in each of Figure 7A , Figure 7B and Figure 7C the system 300 outperforms other models on all metrics of the Dual-Pixel and Human Model Challenger image datasets, and outperforms or ties with other systems on all metrics of the RealEstate10k dataset. Thus, the 3D photo system 300 matches or exceeds the performance of other 3D photo systems.
[0137] Conclusion
[0138] The present disclosure should not be limited to the specific embodiments described in this application, which are intended to illustrate various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from its scope. Functionally equivalent methods and apparatuses within the scope of the present disclosure, other than those described herein, will be apparent to those skilled in the art from the foregoing description. These modifications and variations are intended to fall within the scope of the appended claims.
[0139] The above detailed description has described various features and operations of the disclosed systems, apparatuses, and methods with reference to the accompanying drawings. In the drawings, like symbols generally identify like components unless the context dictates otherwise. The exemplary embodiments described herein and in the drawings are not intended to be limiting. Other embodiments can be utilized and other changes can be made without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0140] Regarding any and all message flowcharts, scenarios, and flow diagrams in the drawings and as discussed herein, each step, block, and / or communication can represent the processing of information and / or the transmission of information according to an example embodiment. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages can be performed not in the order shown or discussed, including substantially simultaneously or in the reverse order, depending on the functions involved. Additionally, more or fewer blocks and / or operations can be used with any one of the message flowcharts, scenarios, and flow diagrams discussed herein, and these message flowcharts, scenarios, and flow diagrams can be combined with each other, in whole or in part.
[0141] A step or block representing the processing of information can correspond to circuitry that can be configured to perform a specific logical function described herein for a method or technique. Alternatively or additionally, a block representing the processing of information can correspond to a portion of a module, segment, or program code (including associated data). The program code can include one or more instructions executable by a processor to implement a specific logical operation or action in a method or technique. The program code and / or associated data can be stored on any type of computer-readable medium, such as a storage device including random access memory (RAM), disk drive, solid state drive, or another storage medium.
[0142] The computer-readable medium may also include non-transitory computer-readable media such as computer-readable media that store data for a short period of time, such as register memory, processor cache, and RAM. The computer-readable medium may also include non-transitory computer-readable media that store program code and / or data for a longer period of time. Thus, the computer-readable medium may include secondary or persistent long-term storage, such as read-only memory (ROM), optical discs or magnetic disks, solid state drives, compact disc read-only memory (CD-ROM). The computer-readable medium may also be any other volatile or non-volatile storage system. The computer-readable medium may be considered, for example, a computer-readable storage medium or a tangible storage device.
[0143] In addition, steps or blocks representing the transmission of one or more information may correspond to the transmission of information between software and / or hardware modules in the same physical device. However, other information transmissions may be between software modules and / or hardware modules in different physical devices.
[0144] The particular arrangements shown in the figures should not be regarded as limiting. It should be understood that other embodiments can include more or fewer of each element shown in a given figure. In addition, some of the elements shown can be combined or omitted. In addition, example embodiments can include elements not shown in the figures.
[0145] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and are not intended to be limiting, and the true scope is indicated by the appended claims.
Claims
1. A computer-implemented method, comprising: Obtaining a monocular image having an initial viewpoint; Determining a depth image including a plurality of pixels based on the monocular image, wherein each respective pixel of the depth image has a corresponding depth value; For each respective pixel of the depth image, determining a corresponding depth gradient associated with the respective pixel of the depth image; Determining a foreground visibility map, the foreground visibility map including visibility values that are inversely proportional to the corresponding depth gradients for each respective pixel of the depth image; Determining a background removal occlusion mask based on the depth image, the background removal occlusion mask including occlusion removal values for each respective pixel of the depth image, the occlusion removal values indicating the likelihood that the corresponding pixel of the monocular image will be unoccluded by a change in the initial viewpoint; (i) Generating a repaired image by using a repair model to repair a portion of the monocular image according to the background removal occlusion mask, and (ii) generating a repaired depth image by using the repair model to repair a portion of the depth image according to the background removal occlusion mask; (i) Generating a first 3D representation of the monocular image based on the depth image, and (ii) generating a second 3D representation of the repaired image based on the repaired depth image; and Generating a modified image by combining the first 3D representation and the second 3D representation according to the foreground visibility map, the modified image having an adjusted viewpoint different from the initial viewpoint.
2. The computer-implemented method according to claim 1, wherein, Determining the foreground visibility map by a soft foreground visibility function that is continuous and smooth along at least a first interval.
3. The computer-implemented method according to claim 1, wherein, Determining the foreground visibility map includes: For each respective pixel of the depth image, determining the corresponding depth gradient associated with the respective pixel by applying a gradient operator to the depth image; and For each respective pixel of the depth image, determining the visibility value based on an exponent of the corresponding depth gradient associated with the respective pixel.
4. The computer-implemented method according to claim 1, wherein, Determining the background removal occlusion mask by a soft background removal occlusion function that is continuous and smooth along at least a second interval.
5. The computer-implemented method according to claim 1, wherein, Determining the background removal occlusion mask includes: Determining a first plurality of differences for each respective pixel of the depth image, wherein each respective difference of the first plurality of differences is determined by subtracting (i) the corresponding depth value of a corresponding reference pixel located within a predetermined pixel distance of the respective pixel and (ii) a scaled number of pixels separating the respective pixel from the corresponding reference pixel from the corresponding depth value of the respective pixel; and For each respective pixel of the depth image, determining the occlusion removal value based on the first plurality of differences.
6. The computer-implemented method according to claim 5, wherein, Determining the occlusion removal value includes: For each respective pixel of the depth image, determining a first maximum difference among the first plurality of differences; and For each respective pixel of the depth image, determining the occlusion removal value by applying a hyperbolic tangent function to the first maximum difference.
7. The computer-implemented method according to claim 5, wherein, The corresponding reference pixels of the corresponding difference are selected from: (i) a vertical scan line including a predetermined number of pixels above and below the corresponding pixel, or (ii) a horizontal scan line including the predetermined number of pixels to the right and left of the corresponding pixel.
8. The computer-implemented method according to claim 1, wherein, Determining the foreground visibility map includes: Determining a depth-based foreground visibility map, the depth-based foreground visibility map including a visibility value for each corresponding pixel of the depth image, the visibility value being inversely proportional to the corresponding depth gradient associated with the corresponding pixel of the depth image; Determining a foreground alpha mask based on and corresponding to the monocular image; Determining a mask-based foreground visibility map based on the difference between (i) the foreground alpha mask and (ii) the dilation of the foreground alpha mask; and Determining the foreground visibility map based on the product of: (i) the depth-based foreground visibility map and (ii) the mask-based foreground visibility map.
9. The computer-implemented method according to claim 8, wherein, Determining the foreground alpha mask includes: Determining foreground segmentation based on the monocular image through a foreground segmentation model; and Determining the foreground alpha mask through a mask model and based on the foreground segmentation and the monocular image.
10. The computer-implemented method according to claim 8, wherein, Determining the foreground visibility map includes: Determining a background occlusion mask based on the depth image, the background occlusion mask including an occlusion value for each corresponding pixel of the depth image, the occlusion value indicating the likelihood that the corresponding pixel of the monocular image will be occluded by a change in the initial viewing point; and Determining the foreground visibility map based on the product of: (i) the depth-based foreground visibility map, (ii) the mask-based foreground visibility map, and (iii) the inverse of the background occlusion mask.
11. The computer-implemented method according to claim 10, wherein, Determining the background occlusion mask includes: Determining a second plurality of differences for each corresponding pixel of the depth image, wherein each corresponding difference of the second plurality of differences is determined by subtracting (i) the corresponding depth value of the corresponding pixel from (ii) the corresponding depth value of the corresponding reference pixel and the scaled number of pixels separating the corresponding pixel from the corresponding reference pixel within a predetermined pixel distance of the corresponding pixel; and For each corresponding pixel of the depth image, determining the occlusion value based on the second plurality of differences.
12. The computer-implemented method according to claim 11, wherein, Determining the occlusion value includes: For each corresponding pixel of the depth image, determining the second largest difference among the second plurality of differences; and For each corresponding pixel of the depth image, determining the occlusion value by applying the hyperbolic tangent function to the second largest difference.
13. The computer-implemented method according to claim 1, wherein, The monocular image includes foreground features and background features, wherein the inpainting model has been trained to inpaint (i) a background occluded region of the monocular image having an intensity value that matches the background features and is independent of the foreground features, and (ii) a corresponding background occluded region of the depth image having a depth value that matches the background features and is independent of the foreground features, and wherein the intensity value of the occluded region of the inpainted image is contextually consistent with the corresponding depth value of the corresponding background occluded region of the inpainted depth image.
14. The computer-implemented method according to claim 1, wherein, The inpainting model has been trained by a training process that includes: obtaining training monocular images; determining a training depth image based on the training monocular images; determining a training background occlusion mask based on the training depth image, the training background occlusion mask including an occlusion value for each respective pixel of the training depth image, the occlusion value indicating the likelihood that the corresponding pixel of the training monocular image will be occluded by a change in the original viewpoint of the training monocular image; (i) generating an inpainted training image by using the inpainting model to inpaint a portion of the training monocular image according to the training background occlusion mask, and (ii) generating an inpainted training depth image by using the inpainting model to inpaint a portion of the training depth image according to the training background occlusion mask; determining a loss value by applying a loss function to the inpainted training image and the inpainted training depth image; and adjusting one or more parameters of the inpainting model based on the loss value.
15. The computer-implemented method according to claim 14, wherein, The loss value is based on one or more of the following: (i) an adversarial loss value determined based on processing the inpainted training image and the inpainted training depth image by a discriminator model, and (ii) a reconstruction loss value determined based on comparing the inpainted training image with the training monocular image and comparing the inpainted training depth image with the training depth image.
16. The computer-implemented method according to claim 1, wherein, Generating the modified image includes: generating a 3D foreground visibility map corresponding to the first 3D representation based on the depth image; (i) generating a foreground image by projecting the first 3D representation based on the adjusted viewpoint, (ii) generating a background image by projecting the second 3D representation based on the adjusted viewpoint, and (iii) generating a modified foreground visibility map by projecting the 3D foreground visibility map based on the adjusted viewpoint; and combining the foreground image and the background image according to the modified foreground visibility map.
17. The computer-implemented method according to claim 16, wherein, The value of the modified foreground visibility map has a range from zero to a predetermined value, and wherein combining the foreground image and the background image includes: determining the sum of (i) a first product of the foreground image and the modified foreground visibility map and (ii) a second product of the background image and the difference between the predetermined value and the modified foreground visibility map.
18. The computer-implemented method according to any one of claims 1-17, wherein, Generating the first 3D representation and the second 3D representation includes: (i) Generating a first plurality of 3D points by de-projecting pixels of the monocular image based on the depth image, and (ii) generating a second plurality of 3D points by de-projecting pixels of the inpainted image based on the inpainted depth image; (i) Generating a first polygon mesh by interconnecting corresponding subsets of the first plurality of 3D points corresponding to adjacent pixels of the monocular image, and (ii) generating a second polygon mesh by interconnecting corresponding subsets of the second plurality of 3D points corresponding to adjacent pixels of the inpainted image; and (i) Applying one or more first textures to the first polygon mesh based on the monocular image, and (ii) applying one or more second textures to the second polygon mesh based on the inpainted image.
19. A system comprising: A processor; And A non-transitory computer-readable medium having stored instructions that, when executed by the processor, cause the processor to perform the operations of the method according to any one of claims 1-18.
20. A non-transitory computer-readable medium having stored instructions that, when executed by a computing device, cause the computing device to perform the operations of the method according to any one of claims 1-18.
Citation Information
Patent Citations
Depth image processing method and device, electronic equipment and storage medium
CN112541875A
System and process for generating a two-layer, 3D representation of a scene
CN1716311A