Method and apparatus for removing and rendering image

By identifying viewpoint correlation in the image and using multiple loss functions to train NeRF, the view inconsistency and blurring problems in NeRF editing are solved, and the 3D scenes are clearly rendered from any viewpoint are achieved.

CN120476430APending Publication Date: 2025-08-12SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480006898.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-13
Filing Date
2024-02-08
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing NeRF technology lacks interpretability when editing 3D scenes and is difficult to maintain multi-view consistency, resulting in image blurring repair.

Method used

By identifying the association between the first image of the multiple images and the first viewpoint, unwanted objects are removed, the repaired 3D scene is rendered using neural radiation field (NeRF), and training is combined with multiple loss functions (L_unmasked, L_depth, L_substituted, L_occluded) to ensure view consistency and deocclusion effect.

Benefits of technology

It realizes the rendering of repaired 3D scenes from any viewpoint, maintaining the consistency and clarity of images, and improving the interpretability and repair effect of NeRF editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476430A_ABST
    Figure CN120476430A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for training a neural radiation field and generating a rendering of a 3D scene with view-related effects from a new viewpoint. A neural radiation field is initially trained using a first loss associated with a plurality of unmasked regions associated with a reference image and a plurality of target images. The training may also be updated using a second loss associated with a depth estimate of a masked region in the reference image. A third loss associated with a view replacement image associated with a corresponding target image may also be used to further update the training. The view replacement image is rendered from a reference viewpoint across a volume of pixels having a view replacement target color. In an embodiment, the neural radiation field is additionally trained using a fourth loss. The fourth loss is associated with a de-occluded pixel in the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to synthesizing views of a 3D scene from new viewpoints. Background Art

[0002] The popularity of Neural Radiance Fields (NeRF) for view synthesis has led to the demand for NeRF editing tools.

[0003] Using existing NeRF techniques to provide scene representations presents technical challenges. First, the black-box nature of the implicit neural representation makes it infeasible to simply edit the underlying data structure based on geometric understanding. In NeRF neural networks, there is no interpretability at the internal node level. Second, because NeRF is trained from images, special considerations are required to maintain multi-view consistency. Using a 2D inpainter to independently inpaint an image of a scene produces viewpoint-inconsistent images. Training standard NeRF to reconstruct these 3D-inconsistent images will result in blurry inpaintings. Summary of the Invention

[0004] Technical Solution

[0005] According to an embodiment of the present disclosure, a method may include: obtaining a plurality of images from a user, wherein the plurality of images are acquired by an electronic device viewing a first scene, and each of the plurality of images is associated with a corresponding viewpoint of the first scene. The method may include: obtaining a first indication of identifying a first image from the plurality of images, wherein the first image is associated with a first viewpoint of the first scene. The method may include: obtaining a second indication of a first object to be removed from the first image. The method may include: removing the first object from the first image to obtain a reference image. The method may include: obtaining a third indication of a second viewpoint from the user, wherein the second viewpoint is different from each of the viewpoints of the plurality of images. The method may include: rendering a second image corresponding to a 3D scene seen from the second viewpoint using a neural radiance field (NeRF), wherein the 3D scene has been repaired into the NeRF. The method may include: displaying the second image on a display of the electronic device.

[0006] According to an embodiment of the present disclosure, a device may include one or more processors. The device may include one or more memories, wherein the one or more memories store instructions configured to cause the device to at least receive a plurality of images from a user, wherein the plurality of images are acquired by a device viewing a first scene, and each of the plurality of images is associated with a corresponding viewpoint of the first scene. The device may include one or more memories, wherein the one or more memories store instructions configured to cause the device to at least obtain a first indication of identifying a first image in the plurality of images, wherein the first image is associated with a first viewpoint of the first scene. The device may include one or more memories, wherein the one or more memories store instructions configured to cause the device to at least obtain a second indication of a first object to be removed from the first image. The device may include one or more memories, wherein the one or more memories store instructions configured to cause the device to at least obtain a third indication of a second viewpoint from the user, wherein the second viewpoint does not correspond to any of the plurality of images. The device may include: one or more memories storing instructions configured to cause the device to at least render a second image corresponding to a 3D scene viewed from a second viewpoint using Neural Radiance Fields (NeRF), wherein the 3D scene has been inpainted into NeRF. The device may include: one or more memories storing instructions configured to cause the device to at least display the second image on a display of the device.

[0007] According to one aspect of the present disclosure, a computer-readable storage medium storing instructions is provided. The instructions, when executed by at least one processor, may cause the at least one processor to obtain a plurality of images from a user, wherein the plurality of images are acquired by an electronic device viewing a first scene, and each of the plurality of images is associated with a corresponding viewpoint of the first scene. The instructions, when executed by the at least one processor, may cause the at least one processor to obtain a first indication of identifying a first image from the plurality of images, wherein the first image is associated with a first viewpoint of the first scene. The instructions, when executed by the at least one processor, may cause the at least one processor to obtain a second indication of a first object to be removed from the first image. The instructions, when executed by the at least one processor, may cause the at least one processor to remove the first object from the first image to obtain a reference image. The instructions, when executed by the at least one processor, may cause the at least one processor to obtain a third indication of a second viewpoint from the user, wherein the second viewpoint is different from each of the viewpoints of the plurality of images. The instructions, when executed by the at least one processor, may cause the at least one processor to render a second image corresponding to a 3D scene viewed from the second viewpoint using Neural Radiance Fields (NeRF), wherein the 3D scene has been inpainted into NeRF. The instructions, when executed by the at least one processor, may cause the at least one processor to display the second image on a display of the electronic device. A method is provided herein, comprising: receiving a plurality of images from a user, wherein the plurality of images are acquired by an electronic device viewing a first scene, and each of the plurality of images is associated with a corresponding viewpoint of the first scene; receiving a first indication identifying a first image from the plurality of images, wherein the first image is associated with a first viewpoint of the first scene; receiving a second indication of a first object to be removed from the first image; removing the first object from the first image to obtain a reference image; obtaining a third indication of a second viewpoint from the user, wherein the second viewpoint does not correspond to any of the plurality of images; rendering a second image corresponding to a 3D scene seen from the second viewpoint using a neural radiance field (NeRF), wherein the 3D scene has been restored into the NeRF; and displaying the second image on a display of the electronic device. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The text and drawings are provided as examples only to help the reader understand the present invention. They are not intended to, and should not be construed as, limiting the scope of the present invention in any way. Although specific embodiments and examples have been provided, it will be apparent to those skilled in the art based on the disclosure herein that the illustrated embodiments and examples may be modified without departing from the scope of the embodiments provided herein.

[0009] Figure 1A An example of logic for rendering an inpainted 3D scene from a new viewpoint is shown, in accordance with some embodiment embodiments.

[0010] Figure 1B An example of adding a selected object to a 3D scene is shown according to some embodiments.

[0011] Figure 1C Rendering from a new viewpoint according to some embodiments is shown. Figure 1B Example of a repaired 3D scene.

[0012] Figure 2A An example of a system for providing a new view in accordance with some embodiments is shown.

[0013] Figure 2B Shown are examples of logic for training NeRF to represent an inpainted 3D scene and using NeRF to obtain a new view, in accordance with some embodiments.

[0014] Figure 3 An example of training NeRF to represent an inpainted 3D scene is shown in accordance with some embodiments.

[0015] Figure 4 shows training according to some embodiments Figure 3 An example of using NeRF to represent further details of the inpainted 3D scene.

[0016] Figure 5 Examples of geometry relevant to view replacement techniques are shown.

[0017] Figure 6 Represents an example of an input image.

[0018] Figure 7 Shown from Figure 6 An example of removing the backpack from an input image and receiving a text command to repair the red fence and repairing the red fence.

[0019] Figure 8 Shown from Figure 6 An example of removing the backpack from the input image and receiving a text command to repair the rubber duck and repairing the rubber duck.

[0020] Figure 9 Shown from Figure 6 An example of removing the backpack from the input image and receiving a text command to repair the flower pot and repairing the flower pot.

[0021] Figure 10 An example of an input image with a backpack as the object to be removed is shown.

[0022] Figure 11 An example of using 2D inpainting to replace a backpack with a red fence is shown.

[0023] and Figure 10 Related Figure 12An example of replacing a backpack by pasting an image of a mailbox is shown.

[0024] and Figure 10 Related Figure 13 An example of replacing a backpack by repairing the red fence and then manually pasting in the bushes is shown.

[0025] Figure 14 An example of a reference image from which an object has been removed is shown.

[0026] Figure 15 An example of an input image set is shown.

[0027] Figure 16 Shown with Figure 15 An example of a set of masks corresponding to the input image .

[0028] Figure 17 An example of an initial target view with distortion in the area corresponding to the inpainting in the reference view is shown.

[0029] Figure 18 Shows relative to Figure 17 Example of residual error of the target view.

[0030] Figure 19 Shown based on Figure 18 An example of an updated rendering of the residual target view.

[0031] Figure 20 An example of improving the de-occlusion processing of an inpainted 3D scene represented by NeRF for views other than a reference view is shown.

[0032] Figure 21 Exemplary hardware of a computing device for implementing the systems and algorithms described by the figures is shown, according to some exemplary embodiments. DETAILED DESCRIPTION

[0033] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configuration in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some cases, well-known structures and components are shown in block diagram form to avoid obscuring these concepts. In the following description, identical components are marked with the same reference numerals throughout the specification and the accompanying drawings.

[0034] The following description provides examples and does not limit the scope, applicability, or embodiments set forth in the claims. The function and / or arrangement of the elements discussed may be changed without departing from the scope of this disclosure. Various examples may appropriately omit, replace, or add various processes or components. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, and / or combined. In embodiments, features described with reference to some examples may be combined in other examples.

[0035] Various aspects and / or features may be presented in terms of systems that may include multiple devices, components, modules, etc. It should be understood and appreciated that the various systems may include additional devices, components, modules, etc. and / or may not include the devices, components, modules, etc. discussed in conjunction with the figures. Combinations of these methods may also be used.

[0036] As a general introduction to the subject matter described in more detail below, the present disclosure provides methods, apparatus, and computer-readable media for repairing unwanted objects in one of a plurality of 2D images forming a complete 3D scene representation. The unwanted objects are removed from any viewpoint within the 3D scene image. Obtaining may include receiving, accessing, acquiring, and the like.

[0037] NeRF technology can be used to inpaint unwanted regions in a view-consistent manner, allowing users to exercise control over the generated scene through a single inpainted image.

[0038] NeRF is an implicit neural field representation (e.g., coordinate mapping) for 3D scenes and objects, generally suitable for multi-view image sets. The basic components are (i) field f θ : (x, d) → (c, σ), which transforms the 3D coordinate x∈R via the learnable parameters θ 3 and viewing direction d∈S 2 Mapping to color c∈R 3 and density σ∈R + , and (ii) rendering operators that produce the color and depth of a given view pixel. The field f can be constructed in various ways θ The rendering operator is implemented as a classical volume rendering integral via an orthogonal approximation, where the ray r is split into t n With t f N parts between (near boundary and far boundary), where t is sampled from the i-th part i The estimated color is then given by Equation 1.

[0039]

[0040] Among them, T i is the transmittance, δ i =t i+1 -ti And c i and σ i It is t i Instead, use t in Equation 1. i Replace c i Estimated Depth and parallax (inverse depth)

[0041] The input is n input images Their camera transformation matrices and a corresponding mask depicting the unwanted area The input also includes a single repair reference view I ref , where ref∈{1, 2, ..., K}, which provides information for embodiment mapping or extrapolation into 3D restoration of the scene represented by NeRF.

[0042] Example 1 ref , not only repairs NeRF, but also generates 3D details and VDE from other viewpoints.

[0043] Next, we discuss the following topics: i) Based on the reference image I ref The depth of the image is obtained by using a monocular depth estimator to guide the geometry of the repaired area (see Figure 4 Items A2-1, A2-2, and A2-3), ii) in conjunction with the view replacement technique of an embodiment, using a bilateral solver to add a VDE to a view other than the reference view (see Figure 5 and Figures 14-19 , for depicting geometry supervision and VDE processing), and iii) since not all masked target pixels are visible in the reference, embodiments provide supervision for such deoccluded pixels via additional inpainting during training (see Figure 20 ). However, the present disclosure is not limited in this regard, and other attachment methods may be utilized without departing from the scope of the present disclosure.

[0044] Training may include experience with a task and attempts to improve performance with respect to performing the task at a future time after the training.

[0045] In an embodiment, the training is based on the following four losses: i) L_unmasked, ii) L_depth, iii) L_substituted, and iv) L_occluded, which represent the unmasked appearance loss, the masked geometry loss, the view-dependent masked color loss, and the deocclusion loss, respectively.

[0046] The overall objective of NeRF fitting for repair is given by Equation 2 (including the weights γ for the last three terms).

[0047] L=L)unmasked+γ depth L_depth+γ sub L_substituted+γ occluded L_occluded (Equation 2)

[0048] Supervision is calculated modulo the iteration count. For example, every N unmasked 、N depth 、N sub and N occluded Iterates times, computing supervision for each added term in Equation 2. No specific loss is used until an appropriate number of iterations have passed.

[0049] In the first stage of training, f is supervised on unmasked pixels via the NeRF reconstruction loss shown in Equation 3. θ Perform N unmasked Iterations.

[0050]

[0051] In Equation 3, R unmasked (with R masked contrast) is the set of rays corresponding to pixels in the unmasked portion of the image (the portion not affected by the mask), and C GT (r) is the true (GT) color for ray r.

[0052] The depth-based loss for the masked part is derived by Equation 4, Equation 5, and Equation 6.

[0053]

[0054] In the above, the scalars h and w are the height and width of the input image.

[0055] Regarding matrices H and V, for position (p x , p y ) pixel p, H(p) = p x And V=p y .

[0056] According to the disparity, the monocular depth estimate of the masked region from the reference image is The disparity from the NeRF model is

[0057] The coefficients a0, a1, a2, a3 in Equation 4 are found by optimization, where F is the target (Equation 5).

[0058] In Equation 4, J is an all-one matrix.

[0059] The inverse of the distance between p and the mask is w(p).

[0060] In Equation 6, the expectation exceeds r′∈R masked .

[0061] Furthermore, in Equation 6, By optimizing D r The variable obtained by improving the smoothness around the mask. The example smoothing technique makes D around the mask boundary r Minimize total variation.

[0062] The loss for obtaining the VDE for the masked portion is given by Equations 7, 8, and 9.

[0063]

[0064] res t =B(I r ,I r -I r,t ,(1-M r )×c max ) (Equation 8)

[0065]

[0066] The expectation in Equation 9 exceeds

[0067] In the above, x i is the ray emitted from the reference camera (direction ) on the shadow point position, is the target image from the camera (at o t (at) and x i The corresponding light direction of the intersection. res t is the repaired residual, I r is the reference view, I r,t is the view replacing the image, is the target color, M r is the mask and B is the bilateral solver.

[0068] The loss for accounting for occluded regions in the reference image, which are however visible in the non-reference image, is given by Equation 10.

[0069]

[0070] In Equation 10, it is expected that the value of (t~T, r~R do,t ), η do >0, and the color and parallax are and

[0071] The above equations are discussed with reference to the accompanying drawings.Before discussing the drawings, a partial list of annotated identifiers is provided here.

[0072] L_unmasked: This is the NeRF reconstruction loss on the unmasked regions of the K input images. See Equation 3.

[0073] L_depth: This loss is based on monocular depth estimation to predict the uncalibrated disparity of the reference image and guide the geometry. See Equation 6.

[0074] L_substituted: This loss takes into account view-dependent effects (VDEs) such as specular reflections and surfaces that are not rough (do not deflect light in every direction). See Equation 9.

[0075] L_occluded: The entire algorithm focuses on the reference view, and pixels that are visible in the target view but not in the reference view are called deoccluded pixels (they are occluded in the reference view and become deoccluded when viewing the scene from other viewpoints). This loss supervises NeRF training so that NeRF produces reasonable results for these deoccluded pixels. See Equation 10.

[0076] I in : The input image chosen as the basis for the reference image.

[0077] I target : Input image set, excluding I in .

[0078] I ref : Reference image, repaired by I in part of the structure.

[0079] I novel : Image of 3D scene restored to NeRF; I novel From the user's request point of view,

[0080] And I novel Produced by NeRF.

[0081] I ref,target : A view replacement image generated by NeRF and associated with one of the target viewpoints.

[0082] After using Equation 8 with target The VDE view replaces the image.

[0083] Γ target : Confidence used by the bilateral solver in the deocclusion process.

[0084] Π target : Target view, showing deoccluded pixels.

[0085] Disparity image generated by NeRF during the deocclusion process.

[0086] An inpainted version of the target view showing deoccluded pixels.

[0087] Use Applies to The disparity image obtained by bilateral guidance.

[0088] res target =Δ target =I ref -I ref,target : The residual used to obtain the VDE for one of the target viewpoints.

[0089] Obtaining new views from 3D restoration into NeRF is now described with reference to the accompanying figures.

[0090] Figure 1A An example of a flow chart of a method L1 for rendering an inpainted 3D scene from a new viewpoint is shown.

[0091] For example, the user has a camera. In operation S1-1, the device may include the user capturing a number of pictures, which may be a video sequence.

[0092] In operation S1-2, the method may include selecting one of the images as an input image by a user.

[0093] In operation S1 - 3 , the method may include selecting an undesired object to be removed from the input image.

[0094] In an embodiment, the electronic device may perform the selection by recommending an object to be erased by the device. The electronic device may select a portion of the image having features such as, but not limited to, many light reflections, a blurred portion, or a portion recognized by the electronic device as a background object.

[0095] In an embodiment, the method may include performing the selection by a user.The method may be performed by the user selecting an area around the object, the electronic device analyzing the identified area, and the electronic device selecting an area around an outline of the object.

[0096] Embodiments may include additional selections by the electronic device based on the information selected by the user.In embodiments, the electronic device may analyze the selected object and recommend whether other objects of a similar type to the object selected by the user should also be selected and erased from the image.

[0097] The method may include removing undesired objects from the image using a mask.

[0098] In an embodiment, in operation S1-4, the device may obtain information from the user regarding the object to be inpainted in the image. The device may be an electronic device. The method of obtaining information (such as, but not limited to, via text, voice, click, or touch) may vary, and an image or video corresponding to the text and an image or video corresponding to the voice may be displayed, and those images and videos may be inserted into the desired input location. As an example, the device used by the user (perhaps a mobile terminal including a camera) may determine the identity of the desired object from the user. Identification may be performed through various methods, such as, but not limited to, voice commands, text commands, touch commands, click commands, or from an image or video submitted to the device. The desired object is inpainted into the reference image. As an example, in an embodiment, the method includes allowing the user to select not only to transmit the new object via text or voice communication, but also to provide an image of the desired object, such as performing manual insertion of the image. However, the present disclosure is not limited to this aspect, and other communication transmission methods may be utilized without departing from the scope of the present disclosure. In an embodiment, the inserted image is downloaded from the internet (something the user finds appealing), or the inserted image is from the user's photo library or another photo library.

[0099] In an embodiment, when a user enters text, there may be multiple images corresponding to the text. The embodiment may be configured to allow the user to select from the multiple images indicated in a list displayed at the bottom of the electronic device user interface display or on the side of the electronic device user interface display. Once an image corresponding to the multiple texts is selected, the embodiment also allows the device to move the image portion as needed, and the image portion enters the repair area.

[0100] In operation S1-5, the device may remove undesired objects and fill the blanks in the 3D scene with desired objects using NeRF, which creates an I ref Methods for performing such repairs are known to practitioners working in this field. ref is the reference view of the restoration, providing the image that the user expects to be extrapolated to {I i}The information in the 3D restoration of the scene of the subject.

[0101] At operation S1-6, the method may include training a neural radiance field (NeRF) to represent the inpainted 3D scene. Figure 3-Figure 4 and Equation 1-Equation 10.

[0102] In operation S1-7, the method may include providing a user with a viewpoint from which to view the 3D scene.

[0103] In operation S1-8, the device may render the new viewpoint and display it to the user. Figure 1C .

[0104] In operation S1-9, the method may include selecting, by the user, another object to repair or viewing the 3D scene from another viewpoint.

[0105] Figure 1B Adding a selected object to a 3D scene according to an embodiment is shown.

[0106] Figure 1B An example of a mobile device displaying a first image Iin is shown. Examples of mobile devices may be a smartphone with a camera, a tablet PC with a camera, etc. The mobile device is an example, and the embodiments are not limited to mobile devices. The embodiments are applicable to electronic devices such as, but not limited to, AR headsets, smart glasses, and smartphones. The method may include selecting an object (e.g., Figure 1B and add it to the first image I in To obtain the restored version of the first image, reference image I ref .

[0107] Figure 1C Shows rendering from the new viewpoint Figure 1B The repaired 3D scene is obtained to obtain image I novel .

[0108] Figure 2A An example of the entire system for providing a new view is shown. K views, K masks, a reference view with additional objects to be inpainted, and a request to render from a new viewpoint are provided to NeRF. The training of NeRF can take place at a mobile terminal, a server, etc. NeRF can provide a new view of the inpainted 3D scene. novel .

[0109] Figure 2B shows the method for training NeRF to represent the inpainted 3D scene and using NeRF to obtain the new view I novel In an embodiment, the method may include taking as input K views of a scene. The i-th view may be denoted as image Ii.

[0110] In operation S2-1, the method may include segmenting out undesired objects to remove them from the scene in each view. This may result in a mask for each scene. The i-th mask may be denoted as M i .

[0111] In operation S2-2, the method may include selecting from the set {I i} selects one of the images as an input image, wherein a reference image I is created from the input image refAt operation S2-3, the method may include training an inpainted neural radiance field to represent the inpainted 3D scene. NeRF may be a scene-specific neural network.

[0112] In operation S2-4, the method may include rendering the repaired 3D scene from the new viewpoint using NeRF to obtain I novel .

[0113] Figure 3 An example of method L3 for training NeRF to represent inpainted 3D scenes according to four training epochs is shown.

[0114] Each training stage in the figure has a predefined number of iterations within it. Each training iteration in NeRF training randomly samples rays from the input view in the scene, renders them using the current NeRF network, and updates the NeRF parameters by minimizing the corresponding loss.

[0115] In operation A1, the loss L_unmasked can be used. See Equation 3. unmasked Operation A1 is performed once per iteration. In operation A1, an input view and camera parameters, a mask, and a (painted) reference view may be used as input. In operation A1, the method may include training NeRF using a loss L_unmasked for the unmasked portion. In operation A1, the loss may be cumulative. The method may include training using the available loss.

[0116] In operation A2, the losses L_depth and L_unmasked can be used. See Equation 3 and Equation 6. depth Operation A2 is performed once for each iteration. In operation A2, the method may include estimating the depth of the masked portion. The depth estimation of the masked portion may include training using L_depth and L_unmasked. In operation A1, the loss may be cumulative. The method may include training using the available loss.

[0117] In operation A3, the losses L_substituted, L_depth, and L_unmasked can be used. See Equation 3, Equation 6, and Equation 9. substituted Operation A3 is performed once per iteration. The K-1 target views, the restored reference views, and the result of operation A2 may be used as inputs to operation A3. In operation A3, the method may include view substitution and training using L_substituted, L_depth, and L_unmasked. In operation A3, the loss may be accumulated. The method may include training using the available loss.

[0118] In operation A4, the losses L_occluded, L_substituted, L_depth, and L_unmasked may be used. See Equation 3, Equation 6, Equation 9, and Equation 10. occluded The method may include performing operation A4 once for each iteration. The method may include training the unoccluded pixels in the target view using L_occluded, L_substituted, L_depth, and L_unmasked. At operation A4, the loss may be cumulative. The method may include training using the available loss. Operation A4 may output a trained NeRF representing the inpainted 3D scene.

[0119] In an embodiment, one or more of A2, A3, and A4 may not be used at all to train NeRF.

[0120] Figure 4 Shows the training Figure 3 An example of NeRF to represent further details of the inpainted 3D scene.

[0121] In operation A1-1, the image {I i In operation A1-1, training can be performed using L_unmasked.

[0122] In operation A2-1, the depth of the masked portion in the reference image may be obtained. In operation A2-2, parallax alignment and smoothing may be performed. In operation A2-3, training may be performed using L_unmasked and L_depth.

[0123] In operation A3-1, the color along the ray from the reference camera can be obtained, but with the view direction from the target camera. This is called view replacement. In operation A3-2, the reference view can be used to replace the color along the ray from the reference camera. ref with I ref,target Compare the two to get the residual res target In operation A3-3, a view-dependent effect (VDE) may be obtained by using a bilateral solver. The bilateral solver may be used to obtain the view-dependent effect (VDE). ref is considered as a reference input. In operation A3-3, the confidence level may be zero within the mask. See Equation 8. In operation A3-4, a target color including the VDE for the view may be obtained. In operation A3-5, training may be performed using L_unmasked, L_depth, and L_substitute. See Equation 9.

[0124] In operation A4-1, deoccluded pixels may be determined by reprojecting all pixels from the reference view into the target view. In operation A4-2, the deoccluded pixels may be inpainted for view t using the leftmost, rightmost, and topmost target images. In operation A4-3, a bilateral solver may be used to inpaint the parallax versions of the deoccluded pixels. In operation A4-4, NeRF training may be performed using L_unmasked, L_depth, L_substitute, and L_occluded. See Equations 3, 6, 9, and 10. Operation A4-4 may output a trained NeRF representing the inpainted 3D scene.

[0125] Figure 5 An example of geometry related to the view replacement technique is shown. The view replacement technique disclosed herein enables rendering from a reference viewpoint but with view-dependent effects of a target viewpoint by replacing the directional input into a per-shade neural color field. Figure 5 The upper portion 510 of the diagram shows a given ray emitted from a reference camera (direction ) on the shadow point position x i , the embodiment can obtain the target image from the camera (at o t (at) and x i The corresponding light direction of the intersection See Equation 7. Figure 5 The lower portions 520 and 530 of 520 show that NeRF can be queried using standard input to obtain the shaded point x i Color Alternatively, Figure 5 530 shows that the NeRF query can be obtained by replacing the input with a view. As color.

[0126] NeRF (e.g., in Figure 5 The output from the NeRF network can be 3D.

[0127] Figures 6-11 Some example results on the image variation level are presented. Figure 6 Represents the input image I in . Figure 7 Shown from Figure 6 An example of removing the knapsack from an input image and obtaining text commands for repairing the red fence and repairing the red fence. Figure 8 Shown from Figure 6 An example of removing the backpack from an input image and obtaining text commands for repairing the rubber duck and repairing the rubber duck. Figure 9 Shown from Figure 6An example of removing the backpack from the input image and obtaining text commands for repairing the flower pot and repairing the flower pot. Figure 10 An example of an input image with a backpack as the object to be removed is shown. Figure 11 An example of using 2D inpainting to replace a backpack with a red fence in the inpainting area is shown. Text commands are examples, and embodiments are not limited to text commands. Embodiments can obtain information about the object to be inpainted in the image in various forms (such as, but not limited to, text, voice, click, or touch), display an image or video corresponding to the text and an image or video corresponding to the voice, and insert those images and videos into the desired input location.

[0128] In an embodiment, the device may obtain (e.g., receive, capture, download) multiple images or short videos while moving the camera around the scene. The device may then interactively segment objects of interest from the scene using well-known techniques (e.g., SPIn-NeRF).

[0129] In an embodiment, reference-guided controllable 3D scene inpainting is performed. The method may include selecting a view and inpainting an object using a controllable 2D inpainting method. For one example, the controllable inpainting method may be stable diffusion inpainting guided by text input. Alternatively, the method may include creating an inpainted image by first inpainting the inpainted image with a background using any 2D inpainting method, and then manually overlaying the object of interest in the inpainted area. The inpainting NeRF may then be trained using guidance from a single inpainted view. The inpainted NeRF may be used to render the inpainted 3D scene from any view.

[0130] For example, with Figure 10 Related Figure 12 shows an example of replacing a backpack by pasting an image of a mailbox. Figure 10 Related Figure 13 An example of replacing a backpack by inpainting a red fence and then manually pasting in a bush is shown. In an embodiment, the method may include obtaining an indication of a selection of an object to inpaint in the first image.

[0131] However, the present disclosure is not limited in this regard, and other methods or examples may be utilized without departing from the scope of the present disclosure.

[0132] Figures 14 to 19 An example of an image describing a method for training NeRF using view replacement is shown. Figure 14 shows a reference image I in which the object has been removed ref . Figure 15 An example of an input image set is shown. in Images other than {I i} is called the target image. Figure 15 The undesired object UO in is the sheet music on the piano stand. Figure 16 Shown with Figure 15 The mask set M corresponding to the input image i . Figure 17 shows the initial target view I with distortion in the area corresponding to the inpainting in the reference view ref,target . Figure 18 Shows relative to Figure 17 The residual res of the target view target .

[0133] Figure 19 Shown based on Figure 18 Update rendering of the residual target view .

[0134] Embodiments may provide view-dependent effects as follows. For each target t, the scene may be rendered from a reference camera with the target color to obtain a view replacement image I ref,target ( Figure 17 ). The bilateral solver can repair the residual between the reference view and the view replacement image, see Equation 8, and obtain the repaired residual res target ( Figure 18 ), subtract the repair residual res from the reference view target To get the target color ( Figure 19 ). The difference between the target color and the view replacement image provides supervision on the masked region.

[0135] Replace the image in the view After (at least N substitute After iterations), the training can be supervised to see the masked appearance of the target image. Each such image The scene can be viewed via a reference source camera (e.g., with I ref image structure), but may have I target ) (particularly, VDE). Embodiments may use those colors obtained by the bilateral solver of Equation 8 to supervise the mask (i.e., under R mask Embodiments may render each view replacement image within the mask (obtaining the target view appearance as shown in FIG. Figure 17 I in ref,target ), and by combining it with the output of the bilateral repair The comparison is performed to calculate the reconstruction loss, as shown in Equation 9.

[0136] Figure 20 An example of improving the de-occlusion processing of an inpainted 3D scene represented by NeRF for views other than a reference view is shown.

[0137] While single-reference inpainting may prevent issues caused by view-inconsistent inpainting, multi-view information is missing in the inpainted area. For example, when inserting a duck into a scene (see Figure 20 ), due to the removal of occlusion (see Figure 20 The second image from the left is labeled Γ target Viewing the scene from another perspective naturally reveals new details on and around the duck. Embodiments can construct these missing details.

[0138] Embodiments can identify target view Π target (also known as I target ) to construct the deocclusion mask Γ target Then, the embodiment can be obtained from target Repair Γ target Masking color, see Figure 20 The upper right image in This is followed by padding the parallax rendered image, using bilateral guidance to ensure consistency. Figure 20 The upper right image in and Figure 20 Parallax image They are the parameters of the term L_occluded in Equation 10. Finally, these restored deoccluded values can be used for supervision. Figure 3 A4.

[0139] A quantitative full-reference (FR) evaluation of 3D inpainting techniques on inpainted regions of preserved views from the SPIn-NeRF dataset is shown in Table 1. The columns show the difference from known ground-truth images of the scene (without the target object) based on learned perceptual patch similarity (LPIPS) and feature-based statistical distance (FID).

[0140] The embodiment with stable diffusion (SD) performs best on both metrics.

[0141]

Table 1

[0142] Quantitative Full Reference (FR) Evaluation of 3D Restoration Technologies

[0143]

[0144]

[0145] As shown in Table 1, the embodiments provide the best performance on both FR metrics. Object-NeRF and Masked-NeRF methods, which perform object removal without changing the newly displayed area, perform the worst. Masked-NeRF combined with DreamFusion performs slightly better. This shows some utility of the diffusion prior; however, while DreamFusion can generate impressive 3D entities on its own, it cannot produce sufficiently realistic output to inpaint real scenes. SPIn-NeRF-SD obtains similarly poor LPIPS but has better FID. SPIn-NeRF-SD cannot handle the larger mismatch of SD's generated results. NeRF-In outperforms the above models. Nevertheless, the use of pixel-level loss leads to blurry output. Finally, our model significantly outperforms the second best model (SPIn-NeRF-LaMa) in terms of FID, reducing it by about 25%. The embodiments are also applied to videos. Table 2 provides an indication of the technical improvement. SD and LaMa are known inpainters.

[0146]

Table 2

[0147] Quantitative Full-Reference (FR) Evaluation of 3D Restoration Techniques for Videos

[0148] method Clarity MUSIQ SPIn-NeRF-LaMa 354.31 58.10 Examples of using LaMa 394.55 62.0 Example using SD 398.56 61.47

[0149] The FR metric is limited by its use of a single ground-truth target image. Therefore, we also examined NR performance, demonstrating superiority over SPIn-NeRF in terms of sharpness (11.2% improvement) and MUSIQ (5.8% improvement); see Table 2. Table 2 indicates that the embodiment provides numerically sharper and more realistic new views. Figure 21 An exemplary device 21 - 1 for implementing the embodiments disclosed herein is shown. Figure 21 Hardware for performing the provided embodiments is shown. For example, the apparatus 21-1 may be a server, a computer, a laptop, a handheld device, or a tablet device.

[0150] As an example, execute Figure 2B Method L2 Figure 2A The NeRF is located on the electronic device, and the method L2 may process information obtained from an input unit of the electronic device.

[0151] As an example, execute Figure 2B Method L2 Figure 2A The NeRF is located on the server, and the image is a server image. Input values of the server image (region selection, information about the object to be repaired, content obtained from text, voice, etc.) can be obtained from the communication unit of the server and used Figure 2B The method L2 applies the input value of the server image.

[0152] The device 21-1 may include one or more hardware processors 21-9. The one or more hardware processors 21-9 may include an ASIC (Application Specific Integrated Circuit), a CPU (e.g., a CISC or RISC device), and / or custom hardware. The embodiments may be deployed on various GPUs. For example, the provider of the GPU is Nvidia. TM , Santa Clara, California. For example, an embodiment may have been deployed on an Nvidia TM On the A6000 GPU.

[0153] The embodiments can be deployed on various computers, servers or workstations. TM is a workstation company based in San Francisco, California. TM Experiments using the embodiments were conducted on a vector workstation.

[0154] The device 21-1 may also include a user interface 21-5 (e.g., a display screen and / or a keyboard and / or a pointing device such as a mouse). The device 21-1 may include one or more volatile memories 21-2. The device 21-1 may include one or more non-volatile memories 21-3. The one or more non-volatile memories 21-3 may include a computer-readable medium storing instructions for execution by one or more hardware processors 21-9 to cause the device 21-1 to perform any of the methods of the embodiments disclosed herein.

[0155] Device 21-1 may include a wired and / or wireless interface 21-4. The wired and / or wireless interface 21-4 may include a receiver component, a transmitter component, and / or a transceiver component. The wired and / or wireless interface 21-4 may enable device 21-1 to establish a connection with another device (e.g., a server, another device) and / or transmit communications. Communications may be effected via a wired connection, a wireless connection, or a combination of wired and wireless connections. The wired and / or wireless interface 21-4 may allow device 21-1 to receive information from another device and / or provide information to another device. In embodiments, the wired and / or wireless interface 21-4 may provide communication with another device via a network (such as, but not limited to, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic-based network, a cellular network (e.g., a fifth generation (5G) network, a long term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a telephone network (e.g., a public switched telephone network (PSTN)), etc. and / or a combination of these or other types of networks). In embodiments, the wired and / or wireless interface 21-4 may provide communication with another device via a device-to-device (D2D) communication link (such as, but not limited to, FlashLinQ, WiMedia, Bluetooth Bluetooth The wired and / or wireless interface 21-4 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a USB interface, an IEEE 1094 (FireWire) interface, or the like.

[0156] Device 21-1 may include a display device 21-6. Device 21-1 may include a display device 21-6. Display device 21-6 may include one or more components that may allow for presentation of information from a component set of device 21-1. For example, bus 21-7 may be a computer monitor, a smartphone screen, a television (TV), a tablet screen, a digital watch, an AR headset, etc. The present disclosure is not limited in this respect.

[0157] Device 21-1 may include a bus 21-7. The components of device 21-1 may be communicatively coupled via bus 21-7. Bus 21-7 may include one or more components that may allow communication between the components of device 21-1. For example, bus 21-7 may be a communication bus, a crossbar, a network, etc. Although bus 21-7 may be Figure 21Although depicted as a single line in FIG. 2 , bus 21 - 7 may be implemented using multiple (eg, two or more) connections between the set of components of device 21 - 1 . The present disclosure is not limited in this respect.

[0158] Embodiments provide a method for inpainting NeRF via a single inpainted reference image. Embodiments may use a monocular depth estimator, aligning its output with the coordinate system of the inpainted NeRF, to back-project the inpainted material from the reference view into 3D space. Embodiments use a bilateral solver to add VDEs to the inpainted regions and a 2D inpainter to fill in the deoccluded regions. Tables 1 and 2, using multiple evaluation metrics, illustrate the superiority of embodiments over existing 3D inpainting methods.

[0159] Finally, embodiments include controllability advantages, enabling the user to control the ref ) to easily change the generated 3D scene. However, the present disclosure is not limited in this regard, and advantages may be utilized without departing from the scope of the present disclosure.

[0160] Embodiments of the present disclosure may solve one or more technical problems.

[0161] Embodiments may use a single reference for inpainting, thereby avoiding view inconsistencies. To geometrically supervise the inpainted region, embodiments may use an optimization-based formulation utilizing monocular depth estimation. Embodiments may obtain view-dependent effects (VDEs) of non-reference views from a reference viewpoint. This may enable guided inpainting methods that propagate non-reference colors (with VDEs) into masked regions of a 3D scene represented by NeRF. Embodiments may also inpaint the de-occluded appearance and geometry in a consistent manner.

[0162] Embodiments may be provided for inpainting regions in a view-consistent and controllable manner. In addition to the typical NeRF input and a mask marking unwanted regions in each view, embodiments may require only a single inpainted view of the scene (e.g., a reference view). Embodiments may use a monocular depth estimator to back-project the inpainted view to the correct 3D position. Then, via new rendering techniques, the bilateral solver of the embodiments may construct view-dependent effects in non-reference views so that the inpainted region appears consistent from any view. For non-reference de-occluded regions that cannot be supervised by a single reference view, embodiments may provide an image inpainter-based approach to guide geometry and appearance. Embodiments may demonstrate performance superior to a NeRF inpainting baseline, with the added advantage that the user can control the generated scene via a single inpainted image.

[0163] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the embodiments.

[0164] As used herein, the terms "component," "module," "system," and the like are intended to include computer-related entities such as, but not limited to, hardware, firmware, a combination of hardware and software, software, or software in operation. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As an illustration, both an application running on a computing device and a computing device can be components. One or more components can reside within a process and / or execution thread, and a component can be located on one computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate by means of local and / or remote processes, such as by signals having one or more data packets, such as data from one component interacting with another component in a local system, a distributed system, and / or interacting with other systems across a network (such as the Internet) by means of signals.

[0165] Embodiments may relate to systems, methods, and / or computer-readable media at any possible level of integrated technical detail. A computer-readable medium may include a computer-readable storage medium (or medium) having computer-readable program instructions thereon for causing a processor to perform operations. A computer-readable medium may not include transient signals.

[0166] A computer-readable storage medium can be a tangible device that can retain and store instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, RAM, ROM, an erasable programmable read-only memory (EEPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a DVD, a memory stick, a floppy disk, a mechanical encoding device (such as a punched card or a raised structure in a groove on which instructions are recorded), and any suitable combination of the foregoing. The computer-readable storage medium used herein should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0167] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0168] The computer readable program code / instruction for performing an operation can be an assembly instruction, an instruction set architecture (ISA) instruction, a machine instruction, a machine dependent instruction, a microcode, a firmware instruction, a state setting data, a configuration data for an integrated circuit, or a source code or an object code written in any combination of one or more programming languages, wherein the programming language includes an object-oriented programming language (such as Smalltalk, C++, etc.) and a procedural programming language (such as "C" programming language or similar programming language). The computer readable program code / instruction can be run completely on the user's computer, partly on the user's computer, run as an independent software package, partly on the user's computer and partly on a remote computer, or run completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including LAN or WAN), or can be connected to an external computer (for example, using an Internet service provider (ISP) to connect through the Internet). In an embodiment, an electronic circuit (including, for example, a programmable logic circuit, FPGA, or programmable logic array (PLA)) can execute the computer readable program instruction with personalized electronic circuit by utilizing the state information of the computer readable program instruction, so as to perform various aspects or operations.

[0169] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create a method for implementing the functions / actions specified in the flowchart and / or block diagram blocks. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may instruct the computer, programmable data processing device, and / or other apparatus to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in the flowchart and / or block diagram blocks.

[0170] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing device, or other apparatus to cause a series of operational steps to be performed on the computer, other programmable device, or other apparatus to produce a computer-implemented process, so that the instructions running on the computer, other programmable device, or other apparatus implement the functions / actions specified in the flowchart and / or block diagram.

[0171] According to example embodiments, at least one of the components, elements, modules, or units (collectively referred to as "components" in this paragraph) represented by the blocks in the accompanying drawings may be embodied as various numbers of hardware, software, and / or firmware structures that perform the corresponding functions described above. According to example embodiments, at least one of these components may use a direct circuit structure (such as a memory, a processor, a logic circuit, a lookup table, etc.) that can perform the corresponding function under the control of one or more microprocessors or other control devices. In addition, at least one of these components may be specifically implemented as part of a module, program, or code, wherein the module, program, or code contains one or more executable instructions for performing a specified logical function and is executed by one or more microprocessors or other control devices. In addition, at least one of these components may include a processor (such as a CPU), a microprocessor, etc. that performs the corresponding function, or may be implemented by a processor (such as a CPU), a microprocessor, etc. that performs the corresponding function. Two or more of these components may be combined into a single component that performs all the operations or functions of the combined two or more components. In addition, at least a portion of the functions of at least one of these components may be performed by another of these components. The functional aspects of the above example embodiments may be implemented in an algorithm executed on one or more processors. Furthermore, components represented by blocks or processing steps may employ any number of related techniques for electronics configuration, signal processing and / or control, data processing and the like.

[0172] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementations of the architecture, functions, and operations of the systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of an instruction that includes one or more executable instructions for implementing a specified logical function. The method, computer system, and computer-readable medium may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those depicted in the accompanying drawings. In some optional embodiments, the functions mentioned in the blocks may not occur in the order mentioned in the accompanying drawings. For example, two blocks shown in succession may actually be executed simultaneously or substantially simultaneously, or the blocks may sometimes be executed in the opposite order, depending on the functions involved. It may also be noted that each block shown in the block diagram and / or flowchart, as well as the combination of blocks shown in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs a specified function or action or a combination of dedicated hardware and computer instructions.

[0173] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code - it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0174] Elements, actions or instructions used herein should not be interpreted as critical or necessary unless clearly described as such. In addition, as used herein, the singular form is intended to include one or more items and can be used interchangeably with "one or more". In addition, as used herein, the term "set" is intended to include one or more items (for example, related items, unrelated items, combinations of related items and unrelated items, etc.), and can be used interchangeably with "one or more". In the case of only one item being intended, the term "one" or similar language is used. In addition, as used herein, the terms "having", "including" etc. are intended to be open terms. In addition, unless otherwise expressly stated, the phrase "based on" is intended to represent "based at least in part". In addition, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0175] References throughout this specification to "one embodiment," "an embodiment," or similar language indicate that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of the present solution. Thus, the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment. As used herein, terms such as "1st" and "2nd" or "first" and "second" may be used solely to distinguish a respective component from another component and do not limit the components in other respects (e.g., importance or order). It should be understood that if an element (e.g., a first element) is referred to as "operably coupled to" another element (e.g., a second element) with or without the term "operably coupled to" or "in communication with," the references to "first" and "second" may be used solely to distinguish the respective component from one another and do not limit the components in other respects (e.g., importance or order). It should be understood that if an element (e.g., a first element) is referred to as "coupled to" another element (e.g., a second element) with or without the term "operably coupled to" or "in communication with," the references to "first" and "second" may be used solely to distinguish the respective component from one another and do not limit the components in other respects (e.g., importance or order).

[0176] When “coupled to,” “coupled to,” “connected with,” or “connected to” another element (e.g., a second element) it means that the element may be coupled to the other element directly (e.g., by wire), wirelessly, or via a third element.

[0177] It will be understood that when an element or layer is referred to as being “on,” “under,” “connected to,” or “coupled to” another element or layer, it can be directly on, under, connected to, or coupled to the other element or layer, or intervening elements or layers may be present. In contrast, when an element is referred to as being “directly on,” “directly under,” “directly connected to,” or “directly coupled to” another element or layer, there are no intervening elements or layers present.

[0178] The descriptions of the various aspects and embodiments have been presented for purposes of illustration and are not intended to be exhaustive or limited to the disclosed embodiments. Even though combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may be directly dependent on only one claim, the disclosure of possible implementations includes each dependent claim in combination with every other claim in the claim set. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.

[0179] It should be understood that the specific order or hierarchy of blocks in the disclosed processes / flowcharts is illustrative of exemplary approaches. It should be understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged based on design preferences. Furthermore, some blocks may be combined or omitted. The appended claims present elements of the various blocks in a sample order and are not meant to be limited to the specific order or hierarchy presented.

[0180] Furthermore, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. Based on the description herein, one skilled in the relevant art will recognize that the present disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other cases, additional features and advantages may be recognized in a particular embodiment that may not be present in all embodiments of the present disclosure.

[0181] According to an aspect of the present disclosure, a method may include obtaining a plurality of images from a user, wherein the plurality of images are acquired by an electronic device viewing a first scene, and each of the plurality of images is associated with a corresponding viewpoint of the first scene. The method may include obtaining a first indication of identifying a first image from the plurality of images, wherein the first image is associated with a first viewpoint of the first scene. The method may include obtaining a second indication of a first object to be removed from the first image. The method may include removing the first object from the first image to obtain a reference image. The method may include obtaining a third indication of a second viewpoint from the user, wherein the second viewpoint is different from each of the viewpoints of the plurality of images. The method may include rendering a second image corresponding to a 3D scene seen from the second viewpoint using a neural radiance field (NeRF), wherein the 3D scene has been repaired into the NeRF. The method may include displaying the second image on a display of the electronic device.

[0182] According to an embodiment of the present disclosure, removing the first object may include performing a first restoration on the first image by applying a mask to the first object to obtain a reference image. The method may also include restoring the 3D scene into NeRF, in part by adjusting a first size of the mask based on a second size of the first object appearing in the second image and applying the mask having the adjusted size to the second image. The method may include inputting the reference image and information of the second viewpoint into NeRF based on user input requesting an image of the 3D scene as seen from a second viewpoint to provide a second image corresponding to the 3D scene as seen from the second viewpoint.

[0183] According to an embodiment of the present disclosure, the method may include: obtaining a fourth indication from a user, wherein the fourth indication is associated with the second object to be restored into the first image.

[0184] The method may include updating the first image to remove the first object from the first image by using a mask before training the NeRF, wherein the first image includes an unmasked portion and a masked portion.

[0185] The method may include training NeRF after removing the first object from the first image, wherein the NeRF is trained to output an inpainted 3D scene from an unobserved viewpoint by accepting as input a reference inpainting view image obtained by selecting one of a plurality of views of the scene and applying a mask to inpaint the object into the reference view image.

[0186] According to an embodiment of the present disclosure, training may be performed at an electronic device. According to an embodiment of the present disclosure, training may be performed at a server.

[0187] According to an embodiment of the present disclosure, the method may include: obtaining a fifth indication from a user after the display, wherein the fifth indication is a selection of a second object to be restored into the first image. The method may include: obtaining a second representative image by restoring the second object into the first image. The method may include: updating the training of NeRF based on the second representative image. The method may include: rendering a third image using NeRF. The method may include: displaying the third image on a display of an electronic device. The operation of training NeRF may include training NeRF using a first loss associated with an unmasked portion based on a reference image and a plurality of images.

[0188] According to an embodiment of the present disclosure, the operation of training NeRF may include: training NeRF using a second loss based on the masked portion and an estimated depth, wherein the estimated depth is associated with a first geometry of the first scene in the masked portion.

[0189] According to an embodiment of the present disclosure, the operation of training NeRF may include: identifying a plurality of deoccluded pixels, wherein the plurality of deoccluded pixels are present in a target image and are associated with a second viewpoint. The method may include: determining a fourth loss, wherein the fourth loss is associated with a second restoration of the plurality of deoccluded pixels of the target image. The method may include: training NeRF using the fourth loss.

[0190] According to an embodiment of the present disclosure, the method may include: when removing a first object from a reference image using a first mask, and if a second size of the first object in other images is different from the first size in the reference image, proportionally adjusting mask sizes of respective masks in the other images to be proportional to respective object sizes of the first object in the other images.

[0191] According to an embodiment of the present disclosure, a method for training a neural radiance field may include: initially training the neural radiance field using a first loss associated with a plurality of unmasked regions, the plurality of unmasked regions being associated with a reference image and a plurality of target images, respectively, wherein the reference image is associated with a reference viewpoint and each target in the plurality of target images is associated with a corresponding target viewpoint. The method may include: updating the training of the neural radiance field using a second loss associated with a depth estimate of a masked region in the reference image. The method may include: updating the training of the neural radiance field using a third loss associated with a plurality of view replacement images, wherein each view replacement image in the plurality of view replacement images is associated with a corresponding target view in the plurality of target images, each view replacement image is rendered from the reference viewpoint across a volume of pixels having a view replacement target color, and the third loss is based on the plurality of view replacement images.

[0192] According to an embodiment of the present disclosure, the method may include: additionally updating the training of the neural radiance field using a fourth loss, wherein the fourth loss is associated with the de-occluded pixels in each of the plurality of target images.

[0193] According to an embodiment of the present disclosure, a method for rendering an image using depth information may include obtaining image data including multiple images showing a first scene from different viewpoints. The method for rendering an image using depth information may include, based on a first user input for identifying a target object from one of the multiple images, performing a first restoration on the one of the multiple images by applying a mask to the target object to obtain a reference image. The method for rendering an image using depth information may include, based on the reference image, adjusting a first size of a mask according to a second size of the target object in each of the remaining images except the one of the multiple images to obtain multiple adjusted masks, and applying the multiple adjusted masks to each of the remaining images, thereby restoring a 3D scene into a neural radiance field (NeRF). The method for rendering an image using depth information may include, based on a second user input for requesting a first image of the 3D scene seen from a requested viewpoint, inputting the reference image and the requested viewpoint into a neural radiance field (NeRF) model to provide a first image, wherein the first image corresponds to the 3D scene seen from the requested viewpoint.

[0194] According to an embodiment of the present disclosure, a device may include one or more processors. The device may include one or more memories, the one or more memories storing instructions configured to cause the device to obtain a plurality of images from a user, wherein the plurality of images are acquired by a device viewing a first scene, and each of the plurality of images is associated with a corresponding viewpoint of the first scene. The device may include one or more memories, the one or more memories storing instructions configured to cause the device to obtain a first indication of identifying a first image in the plurality of images, wherein the first image is associated with a first viewpoint of the first scene. The device may include one or more memories, the one or more memories storing instructions configured to cause the device to obtain a second indication of a first object to be removed from the first image. The device may include one or more memories, the one or more memories storing instructions configured to cause the device to remove the first object from the first image to obtain a reference image. The device may include one or more memories, the one or more memories storing instructions configured to cause the device to obtain a third indication of a second viewpoint from the user, wherein the second viewpoint does not correspond to any of the plurality of images. The device may include: one or more memories storing instructions configured to cause the device to render a second image corresponding to a 3D scene viewed from a second viewpoint using Neural Radiance Fields (NeRF), wherein the 3D scene has been inpainted into NeRF. The device may include: one or more memories storing instructions configured to cause the device to display the second image on a display of the device.

[0195] According to an embodiment of the present disclosure, the device may include instructions configured to cause the device to perform a first restoration on a first image by applying a mask to the first object to remove the first object to obtain a reference image. The device may include instructions configured to cause the device to restore the 3D scene into NeRF in part by adjusting a first size of the mask according to a second size of the first object appearing in a second image and applying the mask having the adjusted size to the second image. The device may include instructions configured to cause the device to input the reference image and information of the second viewpoint into NeRF based on user input requesting an image of the 3D scene seen from a second viewpoint to provide a second image corresponding to the 3D scene seen from the second viewpoint.

[0196] According to an embodiment of the present disclosure, the device may include instructions configured to cause the device to obtain a fourth indication from a user, wherein the fourth indication is associated with a second object to be restored into the first image. The device may also include instructions configured to cause the device to update the first image before training NeRF to remove the first object from the first image using a mask, wherein the first image includes an unmasked portion and a masked portion. The device may also include instructions configured to cause the device to train NeRF after removing the first object from the first image.

[0197] According to an embodiment of the present disclosure, the apparatus may be a mobile device.

[0198] According to an embodiment of the present disclosure, the device may include: instructions configured to enable the device to obtain NeRF from a server after training NeRF, wherein the NeRF has been trained at the server.

[0199] According to one aspect of the present disclosure, a computer-readable storage medium storing instructions is provided. The instructions, when executed by at least one processor, may cause the at least one processor to obtain a plurality of images from a user, wherein the plurality of images are acquired by an electronic device viewing a first scene, and each of the plurality of images is associated with a corresponding viewpoint of the first scene. The instructions, when executed by the at least one processor, may cause the at least one processor to obtain a first indication of identifying a first image from the plurality of images, wherein the first image is associated with a first viewpoint of the first scene. The instructions, when executed by the at least one processor, may cause the at least one processor to obtain a second indication of a first object to be removed from the first image. The instructions, when executed by the at least one processor, may cause the at least one processor to remove the first object from the first image to obtain a reference image. The instructions, when executed by the at least one processor, may cause the at least one processor to obtain a third indication of a second viewpoint from the user, wherein the second viewpoint is different from each of the viewpoints of the plurality of images. The instructions, when executed by the at least one processor, may cause the at least one processor to render a second image corresponding to a 3D scene viewed from the second viewpoint using Neural Radiance Fields (NeRF), wherein the 3D scene has been inpainted into NeRF. The instructions, when executed by the at least one processor, may cause the at least one processor to display the second image on a display of the electronic device.

Claims

1. A method comprising: Obtaining a plurality of images from a user, wherein the plurality of images are acquired by an electronic device viewing a first scene, and each image in the plurality of images is associated with a corresponding viewpoint of the first scene; obtaining a first indication identifying a first image in the plurality of images, wherein the first image is associated with a first viewpoint of the first scene; obtaining a second indication of a first object to be removed from the first image; removing the first object from the first image to obtain a reference image; obtaining a third indication of a second viewpoint from the user, wherein the second viewpoint is different from each of the viewpoints of the plurality of images; rendering a second image corresponding to the 3D scene seen from the second viewpoint using a neural radiance field (NeRF), wherein the 3D scene has been inpainted into the NeRF; and The second image is displayed on a display of the electronic device.

2. The method according to claim 1, wherein The operation of removing the first object includes performing a first restoration on the first image by applying a mask to the first object to obtain the reference image, wherein the method further includes: inpainting the 3D scene into the NeRF in part by adjusting a first size of the mask according to a second size of the first object appearing in the second image and applying the mask having the adjusted size to the second image; and Based on a user input requesting an image of the 3D scene viewed from the second viewpoint, the reference image and information of the second viewpoint are input into the NeRF to provide the second image corresponding to the 3D scene viewed from the second viewpoint.

3. The method according to any one of claims 1 to 2, wherein: The method further comprises: obtaining a fourth indication from the user, wherein the fourth indication is associated with a second object to be restored into the first image; updating the first image to remove the first object from the first image by using a mask before training the NeRF, wherein the first image includes an unmasked portion and a masked portion; and The NeRF is trained after removing the first object from the first image, wherein the NeRF is trained to output an inpainted 3D scene from an unobserved viewpoint by accepting as input a reference inpainting view image obtained by selecting one of a plurality of views of a scene and applying a mask to inpaint the object into the reference view image.

4. The method according to any one of claims 1 to 3, further comprising: obtaining a fifth indication from the user after the displaying, wherein the fifth indication is a selection of a second object to be inpainted into the first image; obtaining a second representative image by inpainting the second object into the first image; updating the training of the NeRF based on the second representative image; Rendering a third image using the NeRF; and The third image is displayed on a display of the electronic device.

5. The method according to any one of claims 1 to 4, wherein: Training the NeRF is performed at the server, and The operation of training the NeRF includes training the NeRF using a first loss associated with the unmasked portion based on the reference image and the multiple images.

6. The method according to any one of claims 1 to 5, wherein: The operation of training the NeRF further includes training the NeRF using a second loss based on the masked portion and an estimated depth associated with a first geometry of the first scene in the masked portion.

7. The method according to any one of claims 1 to 6, wherein: The operation of training the NeRF further includes: performing view replacement of a target image to obtain a view replacement image, wherein the view replacement image comprises: view dependent effects (VDEs) from a third viewpoint different from the first viewpoint associated with the first image, thereby obtaining a view replacement color associated with the third viewpoint, wherein a second geometry of the first scene underlying the view replacement image is a geometry of the reference image, wherein the plurality of images includes the target image and the target image is not the first image; and The NeRF is trained using a third loss based on the view replacement color.

8. The method according to any one of claims 1 to 7, wherein: The operation of training the NeRF further includes: identifying a plurality of deocclusion pixels, wherein the plurality of deocclusion pixels are present in the target image and are associated with the second viewpoint; determining a fourth loss, wherein the fourth loss is associated with a second inpainting of the plurality of deoccluded pixels of the target image; as well as The NeRF is trained using the fourth loss.

9. The method according to any one of claims 1 to 8, further comprising: When removing the first object from the reference image using the first mask, and if a second size of the first object in other images is different from the first size in the reference image, proportionally adjusting mask sizes of respective masks in the other images to be proportional to respective object sizes of the first object in the other images.

10. The method according to any one of claims 1 to 9, wherein: The operations of rendering images using depth information also include: obtaining image data comprising a plurality of images showing a first scene from different viewpoints; performing a first restoration on one of the plurality of images by applying a mask to the target object to obtain a reference image based on a first user input identifying a target object from the one of the plurality of images; Based on the reference image, inpainting the 3D scene into a neural radiation field (NeRF) by adjusting a first size of the mask according to a second size of the target object in each of the remaining images except the one of the plurality of images to obtain a plurality of adjusted masks, and applying the plurality of adjusted masks to each of the remaining images; and Based on a second user input requesting a first image of the 3D scene seen from a requested viewpoint, the reference image and the requested viewpoint are input into a neural radiance field (NeRF) model to provide the first image, wherein the first image corresponds to the 3D scene seen from the requested viewpoint.

11. A device comprising: one or more processors; as well as One or more memories storing instructions, wherein the instructions are configured to cause the device to at least perform the following operations: obtaining a plurality of images from a user, wherein the plurality of images are acquired by the device viewing a first scene, and each image in the plurality of images is associated with a corresponding viewpoint of the first scene; obtaining a first indication identifying a first image in the plurality of images, wherein the first image is associated with a first viewpoint of the first scene; obtaining a second indication of a first object to be removed from the first image; removing the first object from the first image to obtain a reference image; obtaining a third indication of a second viewpoint from the user, wherein the second viewpoint does not correspond to any of the plurality of images; rendering a second image corresponding to the 3D scene seen from the second viewpoint using a neural radiance field (NeRF), wherein the 3D scene has been inpainted into the NeRF; and The second image is displayed on a display of the device.

12. The apparatus of claim 11, wherein: The instructions are further configured to cause the device to remove the first object by performing a first restoration on the first image by applying a mask to the first object to obtain the reference image, and wherein the instructions are further configured to cause the device to: inpainting the 3D scene into the NeRF in part by adjusting a first size of the mask according to a second size of the first object appearing in the second image and applying the mask having the adjusted size to the second image; and Based on a user input requesting an image of the 3D scene viewed from the second viewpoint, the reference image and information of the second viewpoint are input into the NeRF to provide the second image corresponding to the 3D scene viewed from the second viewpoint.

13. The apparatus according to any one of claims 11 to 12, wherein: The instructions are further configured to cause the device to perform the following operations: obtaining a fourth indication from the user, wherein the fourth indication is associated with a second object to be restored into the first image; updating the first image to remove the first object from the first image by using a mask before training the NeRF, wherein the first image includes an unmasked portion and a masked portion; and The NeRF is trained after removing the first object from the first image.

14. The apparatus according to any one of claims 11 to 13, wherein The instructions are further configured to cause the device to perform the following operations: The NeRF is obtained from a server after the NeRF is trained, wherein the NeRF has been trained at the server.

15. A computer-readable medium storing instructions, wherein: When the instructions are executed by at least one processor, the at least one processor is caused to perform the method according to any one of claims 1 to 10.