Object removal method and device based on three-dimensional scene, equipment and medium

By locating the mask of the object to be removed in a multi-view image using image mapping relationship, the problem of inefficient removal of 3D objects in the prior art is solved, and a more efficient object removal process is achieved.

CN120147585APending Publication Date: 2025-06-13DUKE KUNSHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224283.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art relies on depth information during the 3D object removal process, resulting in excessive consumption of time and computing resources and inefficient.

Method used

By using image mapping relationships, mask positioning of objects to be removed in multi-view images is achieved, which simplifies the 3D object removal process and improves efficiency.

Benefits of technology

This greatly simplifies the 3D object removal process, improves removal efficiency, and reduces dependence on depth information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147585A_ABST
    Figure CN120147585A_ABST
Patent Text Reader

Abstract

The invention discloses an object removal method and device based on a three-dimensional scene, equipment and a medium. The method comprises the following steps: in response to a three-dimensional scene object removal request, determining a target view angle image and a corresponding to-be-removed object mask from multi-view angle images of a to-be-processed three-dimensional scene; determining a to-be-removed object mask of the next view angle image according to the to-be-removed object mask of the target view angle image and a mapping relation between the target view angle image and the next view angle image; taking the next view angle image as a new target view angle image, repeatedly executing the process, and determining the to-be-removed object masks of the remaining view angle images; and determining a repair image corresponding to each view angle image according to each view angle image in the multi-view angle image and the corresponding to-be-removed object mask, and determining a target three-dimensional scene corresponding to the to-be-processed three-dimensional scene according to the repair image. According to the scheme, the mask positioning of the to-be-removed object in the multi-view image is realized by using the image mapping relation, the 3D object removal process is simplified, and the 3D object removal efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of scene reconstruction, and particularly to a method, device, equipment and medium for object removal based on a three-dimensional scene. Background Art

[0002] Scene reconstruction refers to the reproduction of the depth, color, translucency, etc. of the scene contained in the pictures through multiple pictures. However, since the content of the reconstructed scene depends on the pictures themselves, the reconstructed scene may contain unnecessary or incorrect objects. Therefore, the 3D (Three-Dimensional) object removal task is proposed. 3D object removal refers to the process of eliminating a specified entity from a 3D scene and repairing the visual gap generated after the removal. 3D object removal technology has a wide range of applications and helps to enhance scene understanding and data diversity.

[0003] In the related art, it is necessary to rely on the depth information of the pictures to assist in the mask positioning of the object to be removed, so as to achieve 3D object removal based on the located mask. Among them, the depth information can be obtained through NeRF (Neural Radiance Fields), but it requires a large amount of time and computing resources, resulting in a complex 3D object removal process, and thus the 3D object removal efficiency is low. Summary of the Invention

[0004] The present invention provides a method, device, equipment and medium for object removal based on a three-dimensional scene, which can utilize the image mapping relationship to achieve the mask positioning of the object to be removed in multi-view images, greatly simplify the 3D object removal process, and improve the 3D object removal efficiency.

[0005] According to an aspect of the present invention, there is provided a method for object removal based on a three-dimensional scene, the method comprising:

[0006] In response to a three-dimensional scene object removal request, determining a target view image from multi-view images of the three-dimensional scene to be processed, and determining a mask of the object to be removed in the target view image;

[0007] According to the mask of the object to be removed in the target view image and the mapping relationship between the target view image and the next view image, determining the mask of the object to be removed in the next view image;

[0008] Taking the next view image as the new target view image, repeating the above process of determining the mask of the object to be removed based on the mapping relationship, and determining the masks of the objects to be removed in the remaining view images by traversing the multi-view images;

[0009] Determining a repaired image corresponding to each view image according to each view image and the corresponding mask of the object to be removed in the multi-view images;

[0010] Determine the target 3D scene corresponding to the 3D scene to be processed according to the repaired image corresponding to each perspective image; wherein, the target 3D scene does not include the object to be removed.

[0011] According to another aspect of the present invention, there is provided an object removal device based on a 3D scene, the device comprising:

[0012] A target perspective mask determination module, configured to, in response to a 3D scene object removal request, determine a target perspective image from multi-perspective images of the 3D scene to be processed, and determine a mask of the object to be removed for the target perspective image;

[0013] A next perspective mask determination module, configured to determine a mask of the object to be removed for the next perspective image according to the mask of the object to be removed for the target perspective image and the mapping relationship between the target perspective image and the next perspective image;

[0014] A remaining perspective mask determination module, configured to use the next perspective image as a new target perspective image, repeat the process of determining the mask of the object to be removed based on the mapping relationship, and determine the mask of the object to be removed for the remaining perspective images by traversing the multi-perspective images;

[0015] A repaired image determination module, configured to determine the repaired image corresponding to each perspective image according to each perspective image and the corresponding mask of the object to be removed in the multi-perspective images;

[0016] A target 3D scene determination module, configured to determine the target 3D scene corresponding to the 3D scene to be processed according to the repaired image corresponding to each perspective image; wherein, the target 3D scene does not include the object to be removed.

[0017] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0018] At least one processor; and,

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the object removal method based on a 3D scene according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the object removal method based on a three-dimensional scene according to any embodiment of the present invention when executed.

[0022] In the technical solution of the embodiment of the present invention, in response to a three-dimensional scene object removal request, a target perspective image is determined from multi-perspective images of a to-be-processed three-dimensional scene, and an object mask to be removed from the target perspective image is determined; according to the object mask to be removed from the target perspective image and the mapping relationship between the target perspective image and the next perspective image, an object mask to be removed from the next perspective image is determined; the next perspective image is used as a new target perspective image, and the above process of determining the object mask to be removed based on the mapping relationship is repeatedly executed to determine the object masks to be removed from the remaining perspective images by traversing the multi-perspective images; a repaired image corresponding to each perspective image is determined according to each perspective image and the corresponding object mask to be removed in the multi-perspective images; a target three-dimensional scene corresponding to the to-be-processed three-dimensional scene is determined according to the repaired images corresponding to each perspective image; wherein, the target three-dimensional scene does not include the object to be removed. This technical solution can use the image mapping relationship to realize the mask positioning of the object to be removed in the multi-perspective images, greatly simplifying the 3D object removal process and improving the 3D object removal efficiency.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a flowchart of a method for removing an object based on a three-dimensional scene according to Embodiment 1 of the present invention;

[0026] Figure 2 is a flowchart of a method for removing an object based on a three-dimensional scene according to Embodiment 2 of the present invention;

[0027] Figure 3 is a schematic flowchart of a method for removing an object based on a three-dimensional scene according to Embodiment 2 of the present invention;

[0028] Figure 4It is a schematic structural diagram of an object removal device based on a three-dimensional scene according to Embodiment 3 of the present invention;

[0029] Figure 5 It is a schematic structural diagram of an electronic device for implementing a method for removing an object based on a three-dimensional scene according to an embodiment of the present invention. Detailed implementation manners

[0030] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] It should be noted that the terms "first", "second", "target", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment 1

[0033] Figure 1 It is a flowchart of a method for removing an object based on a three-dimensional scene provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of efficiently removing an object in a three-dimensional scene. This method can be executed by an object removal device based on a three-dimensional scene. The object removal device based on a three-dimensional scene can be implemented in the form of hardware and / or software, and the object removal device based on a three-dimensional scene can be configured in an electronic device with data processing capabilities. As Figure 1 shown, the method includes:

[0034] S110, in response to a three-dimensional scene object removal request, determine a target perspective image from multi-perspective images of the three-dimensional scene to be processed, and determine a mask of the object to be removed in the target perspective image.

[0035] Among them, the three-dimensional scene object removal request can be used to request the removal of objects in the three-dimensional scene. Exemplarily, the three-dimensional scene object removal request can include the three-dimensional scene to be processed and the object to be removed in the three-dimensional scene to be processed. Among them, the three-dimensional scene to be processed can refer to the three-dimensional scene that needs to have objects removed. The object to be removed in the three-dimensional scene to be processed can refer to any object in the three-dimensional scene to be processed that needs to be removed, where the object to be removed can be one or more. The multi-view images can refer to two-dimensional images of the three-dimensional scene to be processed from multiple different perspectives, with one perspective corresponding to one image. Exemplarily, the multi-views can include at least two of a bird's-eye view, a front view, a top view, a bottom view, and an oblique view (such as a 45-degree top view). The target view image can refer to an image from the multi-view images that includes the object to be removed at any selected perspective. The object-to-be-removed mask can refer to a binary mask of the object to be removed in the two-dimensional image.

[0036] In this embodiment, when a three-dimensional scene object removal request is detected, first, the three-dimensional scene to be processed is obtained from the three-dimensional scene object removal request, and the three-dimensional scene to be processed is projected onto multiple different perspectives to obtain multi-view images. Then, an image including the object to be removed is randomly selected from the multi-view images as the target view image. Furthermore, the target view image and the prompt information of the object to be removed in the target view image are input into an image segmentation model, and the object-to-be-removed mask of the target view image is determined according to the output result of the image segmentation model. Exemplarily, the image segmentation model can adopt SAM (Segment Anything Model). As a deep learning-based image segmentation model, SAM aims to achieve interactive segmentation of images, that is, to divide different objects or regions in the image according to the prompt information provided by the interaction. It can segment any image without any annotation and output a segmentation mask. Among them, the prompt information can be represented in forms such as spatial prompts (such as point prompts) or semantic prompts (such as text prompts).

[0037] In this embodiment, optionally, the three-dimensional scene object removal request includes a target perspective image, a target area in the target perspective image, and objects to be retained in the target area, where the target area includes objects to be removed; correspondingly, determining the mask of the objects to be removed in the target perspective image includes: determining the mask of the objects to be retained in the target perspective image according to the target perspective image and the objects to be retained in the target area; taking the intersection of the mask of the objects to be retained in the target perspective image and the target area to obtain the mask of the objects to be retained in the target area; determining the mask of the objects to be removed in the target perspective image according to the difference between the target area and the mask of the objects to be retained in the target area; or, determining the mask of the objects to be retained in the target area according to the target area and the objects to be retained in the target area; determining the mask of the objects to be removed in the target perspective image according to the difference between the target area and the mask of the objects to be retained in the target area.

[0038] Among them, the target area may refer to the image area in the target perspective image that contains the complete object to be removed. The object to be retained may refer to the object that needs to be retained in the target area. It can be understood that in addition to the complete object to be removed, the target area usually also includes the object to be retained, where the object to be retained may be complete or incomplete. It should be noted that the size and shape of the target area in this embodiment are not specifically limited and can be flexibly set according to actual needs. For example, the target area can adopt regular shapes such as rectangles, squares, circles, ellipses, etc., or irregular shapes.

[0039] In this embodiment, if the mask of the objects to be removed in the target perspective image is determined by using the point hint information of the objects to be removed in the target perspective image, it is necessary to provide representative points of the objects to be removed as the point hint information of the objects to be removed. When there are multiple objects to be removed, the determination of the point hint information takes a lot of time, and due to the limitation of the number of points, it is difficult to ensure the accuracy of the provided point hint information, thereby reducing the determination efficiency and accuracy of the mask of the objects to be removed in the target perspective image.

[0040] To address the above problems, in this embodiment, the objects to be retained in the target area are used as the starting point. By determining the mask of the objects to be retained in the target area and using the difference between the target area and the mask of the objects to be retained in the target area, the mask of the objects to be removed in the target perspective image is determined. It should be noted that usually, in order to reduce the computational pressure, most of the area in the determined target area corresponds to the objects to be removed, and only a small amount of area corresponds to the objects to be retained. In this case, only a small amount of point hint information of the objects to be retained needs to be provided to quickly and accurately determine the mask of the objects to be retained in the target area. Compared with the mask of the objects to be removed, the determination process of the mask of the objects to be retained is simpler, which helps to improve the determination efficiency and accuracy of the mask of the objects to be removed.

[0041] Specifically, when determining the object mask to be removed from the target perspective image, it can be achieved in two ways. One way is as follows: First, input the target perspective image and the point prompt information of the object to be retained in the target area into an image segmentation model (such as SAM). According to the output result of the image segmentation model, determine the object mask to be retained in the target perspective image. Then, take the intersection of the object mask to be retained in the target perspective image and the target area to obtain the object mask to be retained in the target area. Furthermore, calculate the difference between the target area and the object mask to be retained in the target area as the object mask to be removed from the target perspective image. This method is more applicable to the situation where the object to be retained in the target area is incomplete, which helps to improve the determination accuracy of the object mask to be removed from the target perspective image.

[0042] Another way is as follows: First, input the target area and the point prompt information of the object to be retained in the target area into an image segmentation model (such as SAM). According to the output result of the image segmentation model, determine the object mask to be retained in the target area. Then, calculate the difference between the target area and the object mask to be retained in the target area as the object mask to be removed from the target perspective image. The implementation process of this method is relatively simplified, which helps to improve the determination efficiency of the object mask to be removed from the target perspective image.

[0043] It can be understood that when the object to be retained in the target area is incomplete, if only the object to be retained in the target area is segmented, it may lead to a deterioration in the segmentation effect of the object to be retained in the target area, thereby affecting the determination accuracy of the object mask to be removed from the target perspective image. However, due to the relatively simple processing method, it helps to improve the determination efficiency of the object mask to be removed from the target perspective image. Therefore, in practical applications, one of the above methods can be selected according to the requirements to determine the object mask to be removed from the target perspective image.

[0044] S120. Determine the object mask to be removed from the next perspective image according to the object mask to be removed from the target perspective image and the mapping relationship between the target perspective image and the next perspective image.

[0045] Among them, the next perspective can refer to the perspective with the smallest parallax among multiple perspectives in the multi-perspective. Specifically, the target perspective and the next perspective can be regarded as adjacent perspectives. Correspondingly, the target perspective image and the next perspective image can be regarded as adjacent perspective images. The mapping relationship can be used to describe the image correspondence relationship that needs to be satisfied when mapping one image to another. Exemplarily, the mapping relationship can be represented in the form of a homography matrix. Among them, the homography matrix can refer to the projection matrix from one plane to another plane.

[0046] In this embodiment, it is also necessary to determine the mapping relationship between the target perspective image and the next perspective image. Optionally, the method further includes: for each perspective image in the multi-perspective images, performing image feature matching on each perspective image and its next perspective image to obtain a plurality of key point matching pairs; determining a homography matrix between each perspective image and its next perspective image according to the plurality of key point matching pairs; and determining the homography matrix as the mapping relationship between each perspective image and its next perspective image.

[0047] Exemplarily, an image feature matching model can be used to perform image feature matching on each perspective image and its next perspective image, and a plurality of key point matching pairs can be obtained according to the output result of the image feature matching model. For example, the image feature matching model can adopt the LoFTR model. Among them, the LoFTR model is a local image feature matching model based on Transformer. Specifically, for each perspective image in the multi-perspective images, each perspective image and its next perspective image can be input into the image feature matching model, and the key points in each perspective image and its next perspective image can be extracted through the image feature matching model, and the corresponding relationship between the key points in the two images can be determined, whereby a plurality of key point matching pairs can be obtained.

[0048] Exemplarily, the RANSAC (Random Sample Consensus) framework can be used to calculate the homography matrix between each perspective image and its next perspective image by minimizing the reprojection error ε. Among them, p i and p′ i respectively represent the key points belonging to the current perspective image and the next perspective image in the key point matching pair, H represents the homography matrix, and n represents the number of key point matching pairs. Among them, the RANSAC framework is an iterative method for estimating the parameters of a mathematical model from a set of observed data containing outliers; its basic idea is to randomly select sample data to fit the model, calculate the errors between all other data in the dataset and the fitted model, divide the data into inliers and outliers (i.e., outliers) according to the error size, and find the best model containing the most inliers through continuous iteration. Among them, inliers refer to data points with errors less than a certain set threshold, which are considered to conform to the model, while outliers are considered to be data points that do not conform to the model. After obtaining the homography matrix between each perspective image and its next perspective image, the homography matrix can be determined as the mapping relationship between each perspective image and its next perspective image.

[0049] In this embodiment, after obtaining the mask of the object to be removed in the target perspective image, the mask of the object to be removed in the target perspective image can be mapped to the next perspective by using the mapping relationship between the target perspective image and the next perspective image, so that the mask of the object to be removed in the next perspective image can be obtained. Exemplarily, taking the homography matrix as the mapping relationship as an example, the mask of the object to be removed in the next perspective image can be determined by the following formula: where In and In+1 represent the target perspective and the next perspective respectively, and M In and M In+1 represent the masks of the objects to be removed in the target perspective image and the next perspective image respectively, and H In,In+1 represents the homography matrix between the target perspective image and the next perspective image.

[0050] S130. Take the next perspective image as the new target perspective image, and repeat the above process of determining the mask of the object to be removed based on the mapping relationship, and determine the masks of the objects to be removed in the remaining perspective images by traversing the multi-perspective images.

[0051] In this embodiment, after determining the mask of the object to be removed in the next perspective image of the target perspective image, the next perspective image can be taken as the new target perspective image, and the above process of determining the mask of the object to be removed in the next perspective image according to the mask of the object to be removed in the target perspective image and the mapping relationship between the target perspective image and the next perspective image in S120 can be repeated, and the masks of the objects to be removed in the remaining perspective images can be determined by traversing the multi-perspective images, so that the masks of the objects to be removed in each perspective image in the multi-perspective images can be determined. Among them, the remaining perspective images can refer to the other perspective images in the multi-perspective images except the initial target perspective image and its next perspective image.

[0052] S140. Determine the repaired image corresponding to each perspective image according to each perspective image and the corresponding mask of the object to be removed in the multi-perspective images.

[0053] Among them, the repaired image can refer to an image in which the object to be removed in each perspective image is removed, and the image area corresponding to the object to be removed in each perspective image is repaired and completed. Therefore, the repaired image does not include the object to be removed. Specifically, after determining the mask of the object to be removed in each perspective image of the multi-perspective image, each perspective image and the corresponding mask of the object to be removed can be input into an image repair model. The image repair model is used to remove the object to be removed in each perspective image and repair and complete the image area corresponding to the object to be removed. According to the output result of the image repair model, the repaired image corresponding to each perspective image can be determined. Among them, the image repair model can be used to remove the object in the image and repair and complete the image area where the object is removed. Exemplarily, the image repair model can adopt the LaMA (Large Mask Inpainting with Fourier Convolutions) model. Among them, the LaMA model is a high-resolution image repair model based on Fourier convolution.

[0054] S150. Determine the target 3D scene corresponding to the 3D scene to be processed according to the repaired image corresponding to each perspective image; among them, the target 3D scene does not include the object to be removed.

[0055] Among them, the target 3D scene can refer to the 3D scene after removing the object to be removed in the 3D scene to be processed. In this embodiment, after obtaining the repaired image corresponding to each perspective image, the trained NeRF (Neural Radiance Field) can be used to determine the target 3D scene according to the repaired image corresponding to each perspective image, so as to complete the 3D object removal task of the 3D scene to be processed. Specifically, the repaired image corresponding to each perspective image and the camera parameters corresponding to each perspective image are input into the trained NeRF. According to the output of the NeRF, a 3D reconstruction model can be generated, and this 3D reconstruction model is the target 3D scene corresponding to the 3D scene to be processed. Exemplarily, the camera parameters can include the camera position and direction.

[0056] Among them, NeRF refers to a deep learning model for three-dimensional implicit space modeling. Its core lies in mapping each point in the three-dimensional space to a vector representing the radiation (color, density, etc.) of this point. This mapping process is realized by a multi-layer perceptron (MLP) model, which can learn the mapping relationship from the input perspective and position to the output density and color. Specifically, the input of NeRF is a set of two-dimensional images and the corresponding camera parameters, and the output is a function representing the color and density of each point in the 3D scene (i.e., the 3D reconstruction model).

[0057] In the technical solution of the embodiment of the present invention, in response to a three-dimensional scene object removal request, a target perspective image is determined from multi-perspective images of a to-be-processed three-dimensional scene, and an object mask to be removed for the target perspective image is determined; according to the object mask to be removed for the target perspective image and the mapping relationship between the target perspective image and the next perspective image, an object mask to be removed for the next perspective image is determined; the next perspective image is used as a new target perspective image, and the above process of determining the object mask to be removed based on the mapping relationship is repeatedly executed, and the object masks to be removed for the remaining perspective images are determined by traversing the multi-perspective images; repair images corresponding to each perspective image are determined according to each perspective image and the corresponding object mask to be removed in the multi-perspective images; a target three-dimensional scene corresponding to the to-be-processed three-dimensional scene is determined according to the repair images corresponding to each perspective image; wherein, the target three-dimensional scene does not include the object to be removed. This technical solution can use the image mapping relationship to realize the mask positioning of the object to be removed in the multi-perspective images, greatly simplifying the 3D object removal process and improving the 3D object removal efficiency.

[0058] In this embodiment, optionally, determining the repair image corresponding to each perspective image according to each perspective image and the corresponding object mask to be removed in the multi-perspective images includes: determining the repair image corresponding to the target perspective image according to the target perspective image and the corresponding object mask to be removed; determining the repair image corresponding to the next perspective image according to the repair image corresponding to the target perspective image and the mapping relationship between the target perspective image and the next perspective image; using the next perspective image as a new target perspective image, and repeatedly executing the above process of determining the repair image based on the mapping relationship, and determining the repair images corresponding to the remaining perspective images by traversing the multi-perspective images.

[0059] It should be noted that if an image repair model is used to remove objects and repair images for each perspective image respectively, it will consume a large amount of time and computing resources, thus affecting the 3D object removal efficiency. To solve the above problem, in this embodiment, the mapping relationship between adjacent perspective images is used to realize the fast mapping between the repair images of adjacent perspective images. Specifically, first, the target perspective image and its corresponding object mask to be removed are input into the image repair model, and the image repair model is used to remove the object to be removed in the target perspective image and repair and complete the image area corresponding to the object to be removed, so as to obtain the repair image corresponding to the target perspective image. Then, using the mapping relationship between the target perspective image and the next perspective image, the repair image corresponding to the target perspective image is directly mapped to the next perspective to obtain the repair image corresponding to the next perspective image. Exemplarily, taking the homography matrix as the mapping relationship as an example, the repair image corresponding to the next perspective image can be determined by the following formula: where, I In and I In+1They respectively represent the restored images corresponding to the target perspective image and the next perspective image. Furthermore, the next perspective image is used as the new target perspective image, and the above process of determining the restored image based on the mapping relationship is repeatedly executed. By traversing multi-perspective images, the restored images corresponding to the remaining perspective images are determined. Thus, the restored image corresponding to each perspective image can be obtained.

[0060] With such a setting in this solution, by utilizing the mapping relationship between adjacent perspective images, fast mapping between the restored images of adjacent perspective images is achieved, which can significantly reduce the image restoration cost, shorten the image restoration time, contribute to improving the 3D object removal efficiency, and also improve the multi-perspective consistency.

[0061] Embodiment 2

[0062] Figure 2 The figure is a flowchart of a method for removing an object based on a three-dimensional scene provided in Embodiment 2 of the present invention. This embodiment is optimized based on the above embodiment. Specifically, the optimization is as follows: After determining the object mask to be removed from the next perspective image according to the object mask to be removed from the target perspective image and the mapping relationship between the target perspective image and the next perspective image, it further includes: determining a mask offset region according to the centroid and the initial radius of the object mask to be removed from the next perspective image, and selecting multiple sampling points on the boundary of the mask offset region as anchor points; determining an improved mask according to the next perspective image and the anchor points, and determining a loss score according to the improved mask and the object mask to be removed from the next perspective image; using a gradient to optimize the initial radius, and repeatedly executing the above process of determining the improved mask and the loss score during the optimization process, and judging whether the optimization end condition is satisfied based on the loss score; if it is satisfied, the optimization process is ended, and the improved mask at the end of the optimization is determined as the reference mask, and the object mask to be removed from the next perspective image is updated using the reference mask.

[0063] As Figure 2 shown, the method of this embodiment specifically includes the following steps:

[0064] S210, in response to a three-dimensional scene object removal request, determine a target perspective image from multi-perspective images of a to-be-processed three-dimensional scene, and determine an object mask to be removed from the target perspective image.

[0065] S220, according to the object mask to be removed from the target perspective image and the mapping relationship between the target perspective image and the next perspective image, determine the object mask to be removed from the next perspective image.

[0066] S230, determine a mask offset region according to the centroid and the initial radius of the object mask to be removed from the next perspective image, and select multiple sampling points on the boundary of the mask offset region as anchor points.

[0067] It should be noted that when there may be line-of-sight occlusion of the object to be removed in different perspective images or the mapping relationship between the target perspective image and the next perspective image is calculated inaccurately, it will lead to the problem that the mask of the object to be removed in the next perspective image mapped from the mask of the object to be removed in the target perspective image is inaccurate, thus affecting the accuracy of 3D object removal.

[0068] To solve the above problems, in this embodiment, after determining the mask of the object to be removed in the next perspective image, an anchor circle centered on the mask of the object to be removed in the next perspective image is constructed, and the mask of the object to be removed in the next perspective image is further optimized through anchor self-adaptation adjustment, so as to improve the accuracy of the mask of the object to be removed in the next perspective image.

[0069] Specifically, after determining the mask of the object to be removed in the next perspective image, the centroid of the mask of the object to be removed in the next perspective image can be used as the center of the circle, and a circular area is determined as the mask offset area according to the initial radius, and multiple sampling points are selected as anchor points at equal intervals or unequal intervals on the boundary of the mask offset area. The initial radius can refer to the radius length preset according to historical experience. The mask offset area can refer to the deviation area corresponding to the mask of the object to be removed in the next perspective image.

[0070] S240, determine the improved mask according to the next perspective image and the anchor points, and determine the loss score according to the improved mask and the mask of the object to be removed in the next perspective image.

[0071] Among them, the improved mask can refer to the mask obtained after preliminary optimization of the mask of the object to be removed in the next perspective image. The loss score can be used to characterize the deviation degree between the mask of the object to be removed in the next perspective image and the corresponding improved mask.

[0072] In this embodiment, after obtaining multiple anchor points, the next perspective image and the multiple anchor points can be input into an image segmentation model (such as SAM), and the improved mask corresponding to the next perspective image can be obtained according to the output result of the image segmentation model. After determining the improved mask, the loss score can be calculated by using a preset loss function according to the improved mask and the mask of the object to be removed in the next perspective image. The preset loss function can refer to the loss function preset according to actual needs. Exemplarily, the preset loss function can be constructed based on the weighted sum of the intersection over union and the shape similarity. The intersection over union is used to represent the ratio of the intersection of two objects to the union of these two objects. The shape similarity is used to represent the similarity degree of the shapes of two objects. Exemplarily, the shape similarity of two objects can be characterized by calculating the distance between the shape information of the two objects. Specifically, substituting the improved mask and the mask of the object to be removed in the next perspective image into the preset loss function for calculation, the corresponding loss score can be obtained.

[0073] In this embodiment, optionally, determining the loss score according to the improved mask and the mask of the object to be removed in the next perspective image includes: determining the intersection over union and the shape similarity between the improved mask and the mask of the object to be removed in the next perspective image; performing a weighted sum of the intersection over union and the shape similarity to obtain the loss score.

[0074] Specifically, when determining the loss score, first calculate the intersection over union and the shape similarity between the improved mask and the mask of the object to be removed in the next perspective image, and then perform a weighted sum of the calculated intersection over union and shape similarity according to a preset weight to obtain the loss score corresponding to the improved mask and the mask of the object to be removed in the next perspective image.

[0075] S250, Use the gradient to optimize the initial radius, and repeat the above process of determining the improved mask and the loss score during the optimization process, and determine whether the optimization end condition is satisfied based on the loss score.

[0076] Among them, the optimization end condition is used to describe the condition for ending the optimization of the initial radius. Exemplarily, the optimization end condition can be set to that the loss score reaches a convergent state. In this embodiment, the initial radius is optimized by using the gradient. For each new initial radius obtained, the process of S230 - S240 needs to be executed again to determine the new improved mask and the new loss score, and determine whether the loss score reaches a convergent state based on the change of the loss score. If the loss score reaches a convergent state, it can be determined that the optimization end condition is satisfied.

[0077] S260, If satisfied, end the optimization process, and determine the improved mask at the end of the optimization as the reference mask, and use the reference mask to update the mask of the object to be removed in the next perspective image.

[0078] In this embodiment, if it is determined that the optimization end condition is satisfied based on the change of the loss score, the optimization process is ended. At this time, the obtained improved mask is the mask that most conforms to the mask of the object to be removed in the next perspective image. Therefore, the improved mask at the end of the optimization can be determined as the reference mask, and the mask of the object to be removed in the next perspective image is updated by using the reference mask to realize the optimization of the mask of the object to be removed in the next perspective image.

[0079] S270, Take the next perspective image as the new target perspective image, and repeat the above process of determining the mask of the object to be removed based on the mapping relationship and updating the mask of the object to be removed in the next perspective image, and determine the masks of the objects to be removed in the remaining perspective images by traversing the multi-perspective images.

[0080] S280, Determine the repaired image corresponding to each perspective image according to each perspective image and the corresponding mask of the object to be removed in the multi-perspective image.

[0081] S290. Determine the target three-dimensional scene corresponding to the three-dimensional scene to be processed according to the repaired images corresponding to each perspective image; wherein, the target three-dimensional scene does not include the object to be removed.

[0082] In the technical solution of the embodiment of the present invention, after determining the object mask to be removed from the next perspective image according to the object mask to be removed from the target perspective image and the mapping relationship between the target perspective image and the next perspective image, determine the mask offset area according to the centroid and the initial radius of the object mask to be removed from the next perspective image, and select multiple sampling points on the boundary of the mask offset area as anchor points; determine the improved mask according to the next perspective image and the anchor points, and determine the loss score according to the improved mask and the object mask to be removed from the next perspective image; use the gradient to optimize the initial radius, and repeat the above process of determining the improved mask and the loss score during the optimization process, and judge whether the optimization end condition is met based on the loss score; if it is met, end the optimization process, and determine the improved mask at the end of the optimization as the reference mask, and use the reference mask to update the object mask to be removed from the next perspective image, so as to realize the optimization of the object mask to be removed from the next perspective image. This technical solution can accurately construct the index of the object masks to be removed corresponding to multiple perspectives through the mask mapping and the anchor point adaptive adjustment algorithm, improve the accuracy of the mask mapping, help improve the accuracy of 3D object removal, and ensure the consistency of 3D multi-perspective object removal.

[0083] In this embodiment, optionally, determining the repaired image corresponding to each perspective image according to each perspective image and the corresponding object mask to be removed in the multi-perspective image includes: determining the repaired image corresponding to the target perspective image according to the target perspective image and the corresponding object mask to be removed; determining the initial repaired image corresponding to the next perspective image according to the repaired image corresponding to the target perspective image and the mapping relationship between the target perspective image and the next perspective image; taking the intersection of the updated object mask to be removed from the next perspective image and the initial repaired image corresponding to the next perspective image to obtain the first mask; determining the second mask according to the difference between the updated object mask to be removed from the next perspective image and the first mask; determining the final repaired image according to the next perspective image and the second mask, and taking the final repaired image as the repaired image corresponding to the next perspective image; taking the next perspective image as the new target perspective image, and repeating the above process of determining the initial repaired image, the first mask, the second mask and the final repaired image to determine the repaired images corresponding to the remaining perspective images by traversing the multi-perspective images.

[0084] Specifically, first, the target perspective image and its corresponding object mask to be removed are input into an image inpainting model. The image inpainting model is used to remove the object to be removed from the target perspective image and repair and complete the image area corresponding to the object to be removed, obtaining a repaired image corresponding to the target perspective image. Then, using the mapping relationship between the target perspective image and the next perspective image, the repaired image corresponding to the target perspective image is directly mapped to the next perspective, obtaining an initial repaired image corresponding to the next perspective image. Next, the intersection of the updated object mask of the next perspective image and the initial repaired image corresponding to the next perspective image is taken to obtain a first mask, and the difference between the updated object mask of the next perspective image and the first mask is calculated as a second mask. Among them, the second mask can be used to represent the mask area not covered by the preliminary repair. Furthermore, the next perspective image and the second mask are input into the image inpainting model, and the final repaired image is determined according to the output result of the image inpainting model, and the final repaired image is used as the repaired image corresponding to the next perspective image. Finally, the next perspective image is used as the new target perspective image, and the above processes of determining the initial repaired image, the first mask, the second mask, and the final repaired image are repeatedly executed. By traversing multiple perspective images, the repaired images corresponding to the remaining perspective images can be determined, thereby optimizing the repaired images corresponding to each perspective image.

[0085] With such a setting, this solution optimizes the repaired image corresponding to each perspective image using the updated object mask to be removed, improving the accuracy of the repaired image and contributing to improving the accuracy of 3D object removal.

[0086] Figure 3 It is a schematic flowchart of a method for object removal based on a three-dimensional scene provided in the second embodiment of the present invention. As Figure 3 shown, a 3D scene can be rendered through multi-perspective projection to obtain multi-perspective images. The adjacent perspective images in the multi-perspective images are respectively input into the LoFTR model to obtain the key point matching pairs corresponding to the adjacent perspective images, and then the homography matrix between the adjacent perspective images is calculated using the key point matching pairs.

[0087] As Figure 3As shown, the 3D scene can be projected and rendered through the target perspective to obtain the target perspective image. The target region in the target perspective image that contains the complete object to be removed and the point hint information of the object to be retained in the target region are input into SAM to obtain the mask of the object to be removed in the target perspective image. Using the mapping relationship between the target perspective image and the next perspective image, the mask of the object to be removed in the target perspective image can be mapped to the next perspective to obtain the mask of the object to be removed in the next perspective image. At the same time, the target perspective image and the corresponding mask of the object to be removed are input into the LaMa model to obtain the restored image corresponding to the target perspective image, and using the mapping relationship between the target perspective image and the next perspective image, the restored image corresponding to the target perspective image can be mapped to the next perspective to obtain the initial restored image corresponding to the next perspective image. Initialize the anchor points for the mask of the object to be removed in the next perspective image, use SAM to determine the improved mask corresponding to the next perspective image, and then optimize the improved mask through anchor point adaptive optimization to obtain the reference mask, and use the reference mask as the updated mask of the object to be removed in the next perspective image. Determine the final restored image of the next perspective image according to the initial restored image corresponding to the next perspective image and the updated mask of the object to be removed in the next perspective image, and obtain the final restored image corresponding to each perspective image by iterating through all perspectives. Input the final restored image corresponding to each perspective image into NeRF to obtain the 3D scene after removing the object.

[0088] Exemplarily, in a typical indoor environment, the corresponding 3D scene contains multiple objects, such as a table, a dog, a wallet, and a dinosaur model. Now it is necessary to remove the dinosaur model and the wallet from the 3D scene. The following are the specific implementation steps:

[0089] Step1: Project the three-dimensional scene containing the dinosaur model and the wallet onto multiple preset perspectives to obtain multi-perspective images, and select one of the perspective images containing the dinosaur model and the wallet as the target perspective image. Exemplarily, the resolution of the projected multi-perspective images is 1920×1080.

[0090] Step2: Input the adjacent perspective images in the multi-perspective images into LoFTR to obtain the corresponding key point matching pairs of the adjacent perspective images, and use the key point matching pairs to calculate the homography matrix between the adjacent perspective images.

[0091] Step 3: Outline the area containing the dinosaur model and the wallet in the target perspective image as the target area. Assume that there is no dog in the target area. Just click on the table within the target area to keep it. Input the target area and the point prompt information corresponding to the table into SAM to obtain the mask of the object to be removed in the target perspective image. Input the target perspective image and the corresponding mask of the object to be removed into LaMA to obtain the restored image corresponding to the target perspective image. In the restored image, both the dinosaur model and the wallet are removed, and the background of the removed area is filled with the table.

[0092] Step 4: Map the dinosaur mask and the wallet mask of the target perspective image to the next perspective through the calculated homography matrix to obtain the preliminary dinosaur mask and the preliminary wallet mask of the next perspective image.

[0093] Step 5: Map the dinosaur restored image and the wallet restored image corresponding to the target perspective image to the next perspective through the calculated homography matrix to obtain the preliminary dinosaur restored image and the preliminary wallet restored image corresponding to the next perspective image.

[0094] Step 6: Select the centroid of the preliminary dinosaur mask of the next perspective image as the center, and initialize a sampling anchor point with a radius of r = 10 pixels. Input the next perspective image and the sampling anchor point into SAM to obtain the preliminarily improved dinosaur mask. Then select the centroid of the preliminary wallet mask of the next perspective image as the center, and initialize a sampling anchor point with a radius of r = 10 pixels. Input the next perspective image and the sampling anchor point into SAM to obtain the preliminarily improved wallet mask.

[0095] Step 7: Calculate the loss score using the intersection over union and shape similarity, and optimize the anchor point radius using the gradient to find the improved mask that best matches the mapped mask as the final dinosaur mask (r = 60 pixels) and the final wallet mask (r = 100 pixels) of the next perspective image.

[0096] Step 8: Keep the intersection part of the final dinosaur mask of the next perspective image and the preliminary dinosaur restored image corresponding to the next perspective image. At the same time, keep the intersection part of the final wallet mask of the next perspective image and the preliminary wallet restored image corresponding to the next perspective image. Input the preliminary restored area not covered by the preliminary restoration into LaMa to obtain the final restored image corresponding to the next perspective image. By iterating through all perspectives, the final restored image corresponding to each perspective image is obtained.

[0097] Step 9: Input the final restored image corresponding to each perspective image into NeRF to obtain the 3D scene after removing the dinosaur model and the wallet.

[0098] Example 3

[0099] Figure 4The following is a schematic structural diagram of an object removal device based on a three-dimensional scene provided in Embodiment 3 of the present invention. This device can execute the object removal method based on a three-dimensional scene provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. As Figure 4 shown, the device includes:

[0100] A target perspective mask determination module 310, configured to determine a target perspective image from multi-perspective images of a three-dimensional scene to be processed in response to a three-dimensional scene object removal request, and determine a mask of the object to be removed in the target perspective image;

[0101] A next perspective mask determination module 320, configured to determine a mask of the object to be removed in the next perspective image according to the mask of the object to be removed in the target perspective image and the mapping relationship between the target perspective image and the next perspective image;

[0102] A remaining perspective mask determination module 330, configured to use the next perspective image as a new target perspective image, repeat the process of determining the mask of the object to be removed based on the mapping relationship, and determine the masks of the objects to be removed in the remaining perspective images by traversing the multi-perspective images;

[0103] A repaired image determination module 340, configured to determine a repaired image corresponding to each perspective image according to each perspective image and the corresponding mask of the object to be removed in the multi-perspective images;

[0104] A target three-dimensional scene determination module 350, configured to determine a target three-dimensional scene corresponding to the three-dimensional scene to be processed according to the repaired image corresponding to each perspective image; wherein, the target three-dimensional scene does not include the object to be removed.

[0105] Optionally, the three-dimensional scene object removal request includes the target perspective image, the target area in the target perspective image, and the object to be retained in the target area, and the object to be removed is included in the target area;

[0106] Correspondingly, the target perspective mask determination module 310 is configured to:

[0107] Determine a mask of the object to be retained in the target perspective image according to the target perspective image and the object to be retained in the target area;

[0108] Take the intersection of the mask of the object to be retained in the target perspective image and the target area to obtain the mask of the object to be retained in the target area;

[0109] Determine the mask of the object to be removed in the target perspective image according to the difference between the target area and the mask of the object to be retained in the target area; or,

[0110] Determine the mask of the object to be retained in the target area according to the target area and the object to be retained in the target area;

[0111] Determine the mask of the object to be removed in the target perspective image according to the difference between the target area and the mask of the object to be retained in the target area.

[0112] Optionally, the device further includes: an image mapping relationship determination module, configured to:

[0113] For each perspective image in the multi-perspective images, perform image feature matching on the each perspective image and its next perspective image to obtain a plurality of key point matching pairs;

[0114] Determine the homography matrix between the each perspective image and the next perspective image according to the plurality of key point matching pairs;

[0115] Determine the homography matrix as the mapping relationship between the each perspective image and the next perspective image.

[0116] Optionally, the repaired image determination module 340 is configured to:

[0117] Determine the repaired image corresponding to the target perspective image according to the target perspective image and the corresponding mask of the object to be removed;

[0118] Determine the repaired image corresponding to the next perspective image according to the repaired image corresponding to the target perspective image and the mapping relationship between the target perspective image and the next perspective image;

[0119] Take the next perspective image as the new target perspective image, and repeat the process of determining the repaired image based on the mapping relationship, and determine the repaired images corresponding to the remaining perspective images by traversing the multi-perspective images.

[0120] Optionally, the device further includes: a next perspective mask optimization module, configured to:

[0121] After determining the mask of the object to be removed in the next perspective image according to the mask of the object to be removed in the target perspective image and the mapping relationship between the target perspective image and the next perspective image, determine the mask offset area according to the centroid and the initial radius of the mask of the object to be removed in the next perspective image, and select a plurality of sampling points on the boundary of the mask offset area as anchor points;

[0122] Determine an improved mask according to the next perspective image and the anchor points, and determine a loss score according to the improved mask and the mask of the object to be removed in the next perspective image;

[0123] Optimize the initial radius using a gradient, and repeatedly execute the above-mentioned determination process of the improved mask and the loss score during the optimization process, and determine whether the optimization end condition is satisfied based on the loss score;

[0124] If it is satisfied, end the optimization process, determine the improved mask at the end of the optimization as the reference mask, and update the mask of the object to be removed in the next perspective image using the reference mask.

[0125] Optionally, the next perspective mask optimization module is further configured to:

[0126] Determine the intersection over union (IoU) and the shape similarity between the improved mask and the mask of the object to be removed in the next perspective image;

[0127] Perform a weighted sum of the IoU and the shape similarity to obtain a loss score.

[0128] Optionally, the repaired image determination module 340 is further configured to:

[0129] Determine the repaired image corresponding to the target perspective image according to the target perspective image and the corresponding mask of the object to be removed;

[0130] Determine the initial repaired image corresponding to the next perspective image according to the repaired image corresponding to the target perspective image and the mapping relationship between the target perspective image and the next perspective image;

[0131] Take the intersection of the updated mask of the object to be removed in the next perspective image and the initial repaired image corresponding to the next perspective image to obtain a first mask;

[0132] Determine a second mask according to the difference between the updated mask of the object to be removed in the next perspective image and the first mask;

[0133] Determine the final repaired image according to the next perspective image and the second mask, and use the final repaired image as the repaired image corresponding to the next perspective image;

[0134] Take the next perspective image as the new target perspective image, and repeatedly execute the above-mentioned determination process of the initial repaired image, the first mask, the second mask, and the final repaired image, and determine the repaired images corresponding to the remaining perspective images by traversing the multi-perspective images.

[0135] The object removal device based on a three-dimensional scene provided by an embodiment of the present invention can execute the object removal method based on a three-dimensional scene provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0136] Embodiment IV

[0137] Figure 5 FIG. 1 shows a schematic structural diagram of an electronic device 10 that can be used to implement embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0138] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0139] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0140] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the object removal method based on a three-dimensional scene.

[0141] In some embodiments, the object removal method based on a three-dimensional scene can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the object removal method based on a three-dimensional scene described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the object removal method based on a three-dimensional scene by any other suitable means (e.g., by means of firmware).

[0142] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0146] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0147] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0148] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0149] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for removing objects based on a three-dimensional scene, characterized in that: The method comprises: In response to a three-dimensional scene object removal request, determining a target view image from the multi-view images of the three-dimensional scene to be processed, and determining a mask of the object to be removed of the target view image; Determining the mask of the object to be removed of the next perspective image according to the mask of the object to be removed of the target perspective image and the mapping relationship between the target perspective image and the next perspective image; The next viewing angle image is used as a new target viewing angle image, and the process of determining the mask of the object to be removed based on the mapping relationship is repeated, and the mask of the object to be removed of the remaining viewing angle images is determined by traversing the multi-view images; Determine, according to each perspective image in the multi-perspective images and the corresponding mask of the object to be removed, a repair image corresponding to each perspective image; A target three-dimensional scene corresponding to the three-dimensional scene to be processed is determined according to the repaired image corresponding to each perspective image; wherein the target three-dimensional scene does not contain the object to be removed.

2. The method according to claim 1, characterized in that The three-dimensional scene object removal request includes the target perspective image, a target area in the target perspective image, and an object to be retained in the target area, wherein the target area includes the object to be removed; Accordingly, determining the mask of the object to be removed of the target perspective image includes: Determining a mask of the object to be retained of the target perspective image according to the target perspective image and the object to be retained in the target area; Taking the intersection of the to-be-retained object mask of the target perspective image and the target area to obtain the to-be-retained object mask of the target area; Determine the mask of the object to be removed of the target perspective image according to the difference between the target area and the mask of the object to be retained in the target area; or, Determining a mask of the object to be retained in the target area according to the target area and the object to be retained in the target area; The object mask to be removed of the target viewing angle image is determined according to a difference between the target area and the object mask to be retained in the target area.

3. The method according to claim 1 or 2, characterized in that: The method further comprises: For each perspective image in the multi-perspective images, performing image feature matching between each perspective image and its next perspective image to obtain a plurality of key point matching pairs; Determine a homography matrix between each perspective image and the next perspective image according to the plurality of key point matching pairs; The homography matrix is ​​determined as a mapping relationship between each perspective image and the next perspective image.

4. The method according to claim 1, characterized in that Determining a repaired image corresponding to each perspective image in the multi-perspective images and a corresponding mask of the object to be removed includes: Determine a repaired image corresponding to the target perspective image according to the target perspective image and the corresponding mask of the object to be removed; Determining a restoration image corresponding to the next perspective image according to the restoration image corresponding to the target perspective image and a mapping relationship between the target perspective image and the next perspective image; The next viewing angle image is used as a new target viewing angle image, and the above process of determining the restoration image based on the mapping relationship is repeated, and the restoration images corresponding to the remaining viewing angle images are determined by traversing the multi-view images.

5. The method according to claim 1, characterized in that After determining the object mask to be removed of the next perspective image according to the object mask to be removed of the target perspective image and the mapping relationship between the target perspective image and the next perspective image, the method further includes: Determine a mask offset region according to the centroid and initial radius of the mask of the object to be removed in the next viewing angle image, and select multiple sampling points on the boundary of the mask offset region as anchor points; Determine an improved mask according to the next-view image and the anchor point, and determine a loss score according to the improved mask and a mask of the object to be removed of the next-view image; Optimizing the initial radius using a gradient, and repeatedly performing the above-mentioned determination process of the improved mask and the loss score during the optimization process, and determining whether the optimization end condition is met based on the loss score; If satisfied, the optimization process ends, and the improved mask at the end of the optimization is determined as a reference mask, and the mask of the object to be removed of the next viewing angle image is updated using the reference mask.

6. The method according to claim 5, characterized in that Determining a loss score according to the improved mask and the mask of the object to be removed of the next view image includes: Determining an intersection-over-union ratio and a shape similarity between the improved mask and the mask of the object to be removed in the next viewing angle image; The loss score is obtained by performing a weighted summation of the intersection-over-union ratio and the shape similarity.

7. The method according to claim 5, characterized in that Determining a repaired image corresponding to each perspective image in the multi-perspective images and a corresponding mask of the object to be removed includes: Determine a repaired image corresponding to the target perspective image according to the target perspective image and the corresponding mask of the object to be removed; Determining an initial restored image corresponding to the next perspective image according to the restored image corresponding to the target perspective image and a mapping relationship between the target perspective image and the next perspective image; Taking the intersection of the updated mask of the object to be removed of the next-view image and the initial repair image corresponding to the next-view image to obtain a first mask; Determining a second mask according to a difference between the updated mask of the object to be removed of the next viewing angle image and the first mask; Determine a final restoration image according to the next viewing angle image and the second mask, and use the final restoration image as the restoration image corresponding to the next viewing angle image; The next viewing angle image is used as a new target viewing angle image, and the above-mentioned process of determining the initial restoration image, the first mask, the second mask and the final restoration image is repeated, and the restoration images corresponding to the remaining viewing angle images are determined by traversing the multi-view images.

8. A device for removing objects based on a three-dimensional scene, characterized in that: The device comprises: A target view mask determination module is used to determine a target view image from the multi-view images of the 3D scene to be processed in response to a 3D scene object removal request, and determine a mask of the object to be removed of the target view image; A next viewing angle mask determination module, configured to determine the object mask to be removed of the next viewing angle image according to the object mask to be removed of the target viewing angle image and a mapping relationship between the target viewing angle image and the next viewing angle image; A remaining view mask determination module, configured to use the next view image as a new target view image, repeatedly perform the above process of determining the mask of the object to be removed based on the mapping relationship, and determine the mask of the object to be removed of the remaining view images by traversing the multi-view images; A repair image determination module, used to determine a repair image corresponding to each view image in the multi-view images according to each view image and a corresponding mask of the object to be removed; The target three-dimensional scene determination module is used to determine the target three-dimensional scene corresponding to the three-dimensional scene to be processed according to the repaired image corresponding to each perspective image; wherein the target three-dimensional scene does not contain the object to be removed.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the three-dimensional scene-based object removal method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the three-dimensional scene-based object removal method according to any one of claims 1 to 7 when executed.