Shielding removal method and device based on nerve radiation field, electronic equipment and medium
Through multi-view data processing and deep learning technology, a de-occlusion three-dimensional model is generated, which solves the problem of reconstruction stability and poor quality in complex occlusion scenarios, and achieves efficient and accurate occlusion removal and target object reconstruction, improving the quality and practicality of 3D scene reconstruction.
Patent Information
- Application Number
- CN202510485893.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-25
AI Technical Summary
When handling complex occlusion scenarios, it is difficult for the prior art to effectively utilize multi-view cues, resulting in poor stability and quality of the three-dimensional scene reconstruction results. In addition, when the occlusion area is large or the occlusion complexity is high, it is difficult for traditional methods to effectively reconstruct the target object, resulting in poor rendering quality and unnatural artifacts.
By obtaining scene data from multiple perspectives, generating depth images and predicting color map collections, using sparse annotation and segmentation models to generate masks, combining multi-view data and depth information, a de-occlusion three-dimensional model is trained to perform de-occlusion operations of the target perspective, and using pseudo-view poses and data enhancement to improve the generalization ability of the model.
Efficient occlusion removal and target object reconstruction in complex occlusion scenarios are achieved, and the stability and accuracy of three-dimensional scene reconstruction is improved. The generated deocclusion scene images have a sense of reality and credibility, which enhances the robustness and adaptability of the model.
Smart Images

Figure CN120374861A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and a method, apparatus, electronic device, and medium for occlusion removal based on neural radiance fields. Background Art
[0002] The method for occlusion removal based on neural radiance fields is an important technical means in the field of 3D scene reconstruction and occlusion processing. It processes and analyzes multi-view images, combines sparse annotations and depth information, and realizes the removal of occluded areas and the reconstruction of target objects. Currently, occlusion removal mainly relies on foreground-background separation techniques and image inpainting techniques based on external 2D vision priors. The above foreground-background separation techniques improve the visibility of target objects by separating sparse mesh foreground occluders (such as fences, water droplets, etc.). The above image inpainting techniques based on external 2D vision priors remove occluders and fill in missing areas by combining global information.
[0003] However, when using the above methods, there are often the following technical problems: The above foreground-background separation techniques are difficult to effectively reconstruct target objects in scenes with too large occlusion areas or high-complexity occluders, resulting in poor rendering quality. The above image inpainting techniques based on external 2D vision priors are difficult to provide convincing filled content and consistent 3D geometric information when dealing with severe occlusions, which may lead to unnatural artifacts and reduce the realism and credibility of the reconstruction results to a certain extent. In addition, existing methods are often difficult to effectively utilize multi-view cues when dealing with complex occlusion scenes, resulting in poor stability and quality of the reconstruction results.
[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the concept of the present disclosure, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0005] This summary of the present disclosure is used to introduce concepts in a brief form, and these concepts will be described in detail in the following detailed implementation section. This summary of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose a method, apparatus, electronic device, and computer-readable medium for occlusion removal based on neural radiance fields to address one or more of the technical problems mentioned in the above background art section.
[0007] In a first aspect, some embodiments of the present disclosure propose an occlusion removal method based on a neural radiance field. The method includes: obtaining scene data from each perspective to obtain a set of scene images, where the above-mentioned each perspective includes a first perspective and each second perspective; generating a set of depth images and a set of predicted color maps according to the above-mentioned set of scene images; generating a first perspective mask according to the scene image corresponding to the above-mentioned first perspective; generating each second perspective mask based on the above-mentioned first perspective mask and the above-mentioned set of depth images; generating non-occluded region supervision information based on multi-perspective data and the above-mentioned set of predicted color maps, where the above-mentioned multi-perspective data includes the above-mentioned set of scene images and each perspective mask; generating occluded region supervision information based on the above-mentioned multi-perspective data and the above-mentioned set of depth images; training to obtain a de-occluded three-dimensional model according to the above-mentioned non-occluded region supervision information and the above-mentioned occluded region supervision information; performing an occlusion removal operation on the scene image corresponding to the target perspective in the above-mentioned set of scene images based on the above-mentioned de-occluded three-dimensional model to obtain a de-occluded scene image of the target perspective.
[0008] In a second aspect, some embodiments of the present disclosure propose an occlusion removal device based on a neural radiance field. The device includes: an obtaining unit configured to obtain scene data from each perspective to obtain a set of scene images, where the above-mentioned each perspective includes a first perspective and each second perspective; a prediction unit configured to generate a set of depth images and a set of predicted color maps according to the above-mentioned set of scene images; a first generating unit configured to generate a first perspective mask according to the scene image corresponding to the above-mentioned first perspective; a second generating unit configured to generate each second perspective mask based on the above-mentioned first perspective mask and the above-mentioned set of depth images; a first supervision unit configured to generate non-occluded region supervision information based on multi-perspective data and the above-mentioned set of predicted color maps, where the above-mentioned multi-perspective data includes the above-mentioned set of scene images and each perspective mask; a second supervision unit configured to generate occluded region supervision information based on the above-mentioned multi-perspective data and the above-mentioned set of depth images; a training unit configured to train to obtain a de-occluded three-dimensional model according to the above-mentioned non-occluded region supervision information and the above-mentioned occluded region supervision information; and an occlusion removal unit configured to perform an occlusion removal operation on the scene image corresponding to the target perspective in the above-mentioned set of scene images based on the above-mentioned de-occluded three-dimensional model to obtain a de-occluded scene image of the target perspective.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the above first aspect.
[0010] Fourthly, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.
[0011] Fifthly, some embodiments of the present disclosure provide a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.
[0012] The above embodiments of the present disclosure have the following beneficial effects: Through the occlusion removal method based on neural radiance fields in some embodiments of the present disclosure, occlusion removal and target object reconstruction can be achieved, and the stability of the reconstruction effect, as well as the efficiency and accuracy of occlusion removal, are improved in different scenarios. Thereby, the quality and practicality of 3D scene reconstruction are enhanced. Specifically, traditional occlusion removal methods such as foreground-background separation techniques and image inpainting techniques based on external 2D vision priors often have the following technical problems: The above foreground-background separation techniques are difficult to effectively reconstruct the target object in scenarios with a large occlusion area or a high complexity of the occluder, resulting in poor rendering quality. The above image inpainting techniques based on external 2D vision priors are difficult to provide convincing filled content and consistent 3D geometric information when dealing with severe occlusions, which may lead to unnatural artifacts and to a certain extent reduce the realism and credibility of the reconstruction results. In addition, existing methods often have difficulty in effectively utilizing multi-view cues when dealing with complex occlusion scenarios, resulting in poor stability and quality of the reconstruction results. Based on this, the occlusion removal method based on neural radiance fields in some embodiments of the present disclosure, firstly, obtains the scene data of each view to obtain a set of scene images, wherein the above each view includes a first view and each second view. Thus, a multi-view comprehensive data basis is provided for subsequent occlusion removal, ensuring the diversity and geometric consistency of the data. Then, according to the above set of scene images, a set of depth images and a set of predicted color maps are generated. Thus, by combining depth information and color information, rich feature support is provided for the identification and completion of occlusion regions, improving the accuracy of reconstruction. Next, according to the scene image corresponding to the first view, a first-view mask is generated. Thus, by combining a sparse annotation and segmentation model, the target object and the occlusion region are located, providing a reliable basis for the generation of multi-view masks. Then, based on the first-view mask and the set of depth images, each second-view mask is generated. Thus, through the mapping of sampling points and depth screening, the annotation information of the first view is effectively propagated to other views, ensuring the consistency and accuracy of multi-view data. Then, based on the multi-view data and the set of predicted color maps, non-occluded region supervision information is generated. Thus, by combining a global difference map and a non-occluded region mask, the reconstruction quality of the non-occluded region is evaluated, providing a reliable supervision signal for model optimization. At the same time, based on the multi-view data and the set of depth images, occlusion region supervision information is generated. Thus, by combining the generation of the filled color image of the source view and the depth constraint error, it is ensured that the completion of the occlusion region conforms to both color consistency and geometric depth logic, improving the reconstruction effect of the occlusion region. Finally, according to the non-occluded region supervision information and the occlusion region supervision information, a de-occluded 3D model is trained, and based on the above de-occluded 3D model, a de-occlusion operation is performed on the target view to obtain a de-occluded scene image of the target view.Thus, through the comprehensive utilization of multi-view clues and the dynamic optimization of the model, the removal of occluded regions and the reconstruction of target objects are achieved. The generated occluded-free scene images are realistic and reliable, and can, to a certain extent, handle complex occluded scenes. In addition, through the generation of pseudo-view poses and data augmentation, the generalization ability and robustness of the model are further improved, thereby enhancing the stability of the reconstruction effect in different scenes. Moreover, by combining multi-view clues with deep learning techniques, the efficiency and accuracy of occlusion removal are increased, providing a more reliable and practical solution for fields such as 3D scene reconstruction, virtual reality, and augmented reality. Brief Description of the Drawings
[0013] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of a method for occlusion removal based on neural radiance fields according to the present disclosure;
[0015] Figure 2 is a schematic structural diagram of some embodiments of an apparatus for occlusion removal based on neural radiance fields according to the present disclosure;
[0016] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Description of the Embodiments
[0017] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the scope of protection of the present disclosure.
[0018] It should also be noted that, for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules, or units.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0023] Figure 1 Flow 100 of some embodiments of the occlusion removal method based on neural radiance fields according to the present disclosure is shown. The occlusion removal method based on neural radiance fields includes the following steps:
[0024] Step 101, obtain scene data from each perspective to obtain a set of scene images.
[0025] In some embodiments, an execution entity (such as a computing device) of the occlusion removal method based on a neural radiance field may obtain scene data from each perspective to obtain a set of scene images. Among them, the scene data from each of the above perspectives includes the scene images corresponding to each perspective and the parameter data corresponding to each scene image. Each of the above perspectives corresponds to each of the above scene images. The scene images corresponding to each of the above perspectives may be relatively clear images including a target object (the occluded object) and an occluder (the object occluding the target object) captured at each preset fixed perspective in the same environment by the same camera or a tool with a camera. The parameter data corresponding to each of the above scene images may include camera intrinsics (such as the focal lengths of the horizontal and vertical coordinates, the position of the principal point, etc.) and the camera extrinsics corresponding to each scene image (such as the rotation matrix and the translation vector). Since the scene images corresponding to each of the above perspectives are captured by the same camera under the same settings, the camera intrinsics corresponding to each scene image are the same. The camera extrinsics corresponding to each scene image may also be called the camera poses corresponding to each perspective. The set of scene images includes the scene images corresponding to each perspective. Each of the above perspectives may include a first perspective and each second perspective. The first perspective may be a preset perspective as the starting point of the research. The second perspective may be each perspective other than the first perspective among all perspectives. In practice, the execution entity may use a structured reconstruction method to perform sparse reconstruction on the scene images corresponding to each of the above perspectives to obtain the camera intrinsics (usually represented in matrix format) and the camera extrinsics corresponding to each scene image. Then, the scene images corresponding to each of the above perspectives are determined as the set of scene images. For example, the execution entity may use the COLMAP algorithm to obtain the intrinsics of the camera and the camera extrinsics corresponding to each scene image through feature extraction, feature matching, and geometric calculation. Then, the scene images corresponding to each of the above perspectives are integrated into a set of scene images.
[0026] Step 102: Generate a set of depth images and a set of predicted color maps according to the set of scene images.
[0027] In some embodiments, the execution entity may generate a set of depth images and a set of predicted color maps according to the set of scene images.
[0028] In some optional implementation manners of some embodiments, the execution entity may generate a set of depth images and a set of predicted color maps according to the set of scene images through the following steps:
[0029] First step: For each scene image in the set of scene images, perform the following steps:
[0030] Sub-step 1: Obtain the camera parameters corresponding to the above-mentioned scene images according to the parameter data corresponding to each scene image. Among them, the camera parameters corresponding to the above-mentioned scene images include the above-mentioned camera internal parameters and the camera external parameters corresponding to the above-mentioned scene images. In practice, the above-mentioned execution entity can retrieve the camera parameters corresponding to the above-mentioned scene images from the scene data from each perspective.
[0031] Sub-step 2: For each pixel point in the above scene image, generate a ray corresponding to each pixel point according to the camera parameters corresponding to the above scene image. Among them, the rays corresponding to the above pixel points can establish a corresponding relationship between the pixel information in the 2D image and the points in the 3D scene (usually represented in the form of an equation), that is, a straight line starting from the optical center of the camera (or the center point of the sensor), passing through the pixel point on the image plane, and extending into the 3D scene can be used to map the pixel point in the image to the 3D space. The color and depth of each of the above pixel points are determined by the attributes (color and density) of the 3D points distributed along the ray. In practice, for each pixel point in the above scene image, the above execution entity usually takes the upper left corner of the above scene image as the origin, the right as the positive direction of the abscissa, and the down as the positive direction of the ordinate to generate the coordinates of the pixel point. Then, the coordinates of the pixel point are decentralized (that is, subtract the principal point offset, subtract the principal point abscissa from the abscissa of the pixel point, and subtract the principal point ordinate from the ordinate of the pixel point) to obtain the decentralized coordinates. Next, the decentralized coordinates are normalized by the focal length (that is, converted to the physical scale, the abscissa is divided by the abscissa focal length, and the ordinate is divided by the ordinate focal length) to obtain the normalized plane coordinates. Combining the above normalized plane coordinates with the imaging plane position of the camera (the position where the depth is 1 in the optical axis direction of the camera can be taken), the coordinates of the 3D point corresponding to the normalized coordinates (that is, the 3D point in the camera coordinate system) are obtained. Since the ray points from the optical center to the 3D point corresponding to the above normalized coordinates, the coordinates of the 3D point corresponding to the above normalized coordinates can be converted into the camera coordinate system direction vector (multiplied by the imaging plane position of the camera). Since the rotation matrix describes the rotation from the world coordinate system to the camera coordinate system, the above camera coordinate system direction vector can be combined with the inverse matrix of the rotation matrix to generate the world coordinate system direction vector. Since the starting point of the ray is the position of the camera optical center in the world coordinate system, that is, the translation vector, the ray equation of this pixel point can be expressed as the translation vector corresponding to this pixel point plus the above world coordinate system direction vector multiplied by the distance along the ray (that is, the ray parameter). Similarly, the rays corresponding to each pixel point can be obtained. Among them, the value range of the above ray parameter can be the effective range of the ray in the scene. The effective range in the above scene is usually a closed interval with the near-end boundary and the far-end boundary as the minimum and maximum values. The above near-end boundary is usually the minimum distance preset from the camera optical center to the nearest effective scene surface. The above far-end boundary is usually the distance from the camera optical center to the farthest visible scene point preset (for example, the near-end boundary can be taken as 0.1m, and the far-end boundary can be taken as 10m). For example, for a scene image with a resolution of 1280×720, the coordinates of any pixel point are selected as (740, 460). The abscissa and ordinate focal lengths can be extracted from the camera internal parameter matrix as both 800, and the principal point position coordinates are (640, 360).Next, the coordinates of this pixel point can be subtracted from the coordinates of the principal point position to obtain a decentered coordinate of (100, 100). Then, the decentered coordinates are divided by the focal lengths of the horizontal and vertical coordinates respectively to obtain a normalized plane coordinate of (0.125, 0.125). Since in the camera coordinate system, the imaging plane is located at a depth of 1 (the straight-line distance from the object to the camera optical center is the position of the camera's imaging plane), therefore, the coordinates of the 3D point corresponding to the normalized coordinates are (0.125, 0.125, 1), and the direction vector of the camera coordinate system obtained by conversion is (0.125, 0.125, 1). If the rotation matrix is a matrix that rotates 30 degrees around the optical axis direction of the camera, the direction vector of the camera coordinate system can be multiplied by the inverse matrix of the rotation matrix to obtain a direction vector of the world coordinate system that can be (0.168, 0.045, 1). If the translation vector of the scene image (i.e., the world coordinates of the optical center) is (2, 0, 5), the ray equation corresponding to this pixel can be generated, that is, (2, 0, 5) + ray parameter × (0.168, 0.045, 1).
[0032] Sub-step three, uniformly sample on the rays corresponding to each of the above pixel points to obtain each sampling point and the position and ray direction corresponding to each sampling point. Among them, the sampling range can be the effective range in the above scene. The positions of each of the above sampling points can be obtained through the ray equations of the corresponding pixel points (i.e., the three-dimensional coordinates obtained by sampling on the rays). The ray directions of each of the above sampling points are represented by the direction vectors of the world coordinate system of the rays corresponding to each sampling point. In practice, the above execution entity can uniformly sample on the rays corresponding to each of the above pixel points and within the above sampling range to obtain each sampling point. Then, since both the sampling frequency and distance are preset in advance, the positions and ray directions corresponding to each sampling point can be obtained through the ray equations of the pixel points corresponding to each sampling point and in combination with the direction vectors of the world coordinate system of the ray equations of the pixel points.
[0033] Sub-step 4: Input the positions and ray directions of the above-mentioned respective sampling points into the constructed three-dimensional reconstruction model to obtain the colors and densities of the respective sampling points. Among them, the above-mentioned constructed three-dimensional reconstruction model can be a neural radiance field model implemented by a multi-layer perceptron. The input of the above-mentioned constructed three-dimensional reconstruction model is the positions and ray directions of the respective sampling points, and the output is the colors and densities of the respective sampling points. The above-mentioned constructed three-dimensional reconstruction model includes an encoding layer, a backbone network, and a color generation branch. The above-mentioned encoding layer can convert the input raw data (such as position coordinates and ray directions) into a higher-dimensional encoding so that the model can better capture details and complex patterns. The above-mentioned backbone network can be a multi-layer perceptron (MLP), which can extract features from the encoded data. The above-mentioned color generation branch can be a sub-network that generates the color of the sampling point according to the feature vector and direction encoding. In practice, for any one of the respective sampling points, the above-mentioned execution entity can use multi-frequency sine and cosine functions in the above-mentioned encoding layer to encode the position coordinates (such as (0.5, 1.2, 3.0)) (with a frequency parameter of 10) to obtain the axis-independent encoding of this sampling point (with a total dimension of 60). At the same time, the ray direction of this sampling point (such as (0°, 30°)) can be low-frequency encoded (with a frequency parameter of 4) to obtain the direction component-independent encoding of this sampling point (with a total dimension of 24). Then, feature extraction can be performed on the axis-independent encoding of this sampling point in the 8-layer fully connected layer (including ReLU activation, with 256 neurons in each layer) of the above-mentioned backbone network to obtain a feature vector (256-dimensional). Then, the above-mentioned feature vector can be mapped to a scalar, and ReLU can be used to ensure non-negativity to obtain the density of this sampling point (such as 0.8). Then, in the above-mentioned color branch, first, the feature vector and the direction component-independent encoding are concatenated to obtain an input vector (with a total dimension of 280). Then, the input vector is mapped to 128 dimensions (ReLU activation) in the fully connected layer of the above-mentioned color branch (including the fully connected layer and the color output layer) to obtain a dimension-reduced vector. Finally, the 128-dimensional dimension-reduced vector is mapped to 3 dimensions in the color output layer of the above-mentioned color branch (using Sigmoid to be constrained to [0, 1], such as (0.7, 0.5, 0.3)) to obtain the color of this sampling point. Similarly, the colors and densities of the respective sampling points can be obtained.
[0034] Sub-step five: Generate the predicted color of each pixel point based on the colors and densities of the above-mentioned sampling points. Among them, the predicted color of each pixel point can be obtained by calculating the colors and densities of the sampling points on the ray corresponding to each pixel point. In practice, for any pixel point among each pixel point, the execution entity can obtain the ray equation of this pixel point, so as to obtain the colors and densities of the corresponding sampling points on this ray equation. Then, for any sampling point, the transmittance is the probability that the light is not scattered or absorbed before reaching this sampling point. When calculating, it is necessary to combine the densities and distances of all the previous sampling points (that is, the distance from the sampling point to the optical center. Since it is uniform sampling, it is a predefined value). Starting from the first sampling point to the sampling point before this sampling point, multiply the density of each sampling point by the distance from each sampling point to the next sampling point in turn, then add up these products, and then take the opposite of the exponential function of this cumulative value to obtain the transmittance of the current sampling point. Then, multiply the transmittance by the density of the current sampling point and then by the distance from the current sampling point to the next sampling point to obtain the weight of the current sampling point. According to the above operations, the weights corresponding to each sampling point on this pixel point can be obtained. Then, multiply the colors of each sampling point by the corresponding weights respectively, and then add up these products to obtain the predicted color of this pixel point. Similarly, the predicted colors of each pixel point can be obtained.
[0035] Second step: Optimize the above-mentioned constructed 3D reconstruction model based on the predicted colors of each pixel point on each scene image to obtain a pre-trained 3D reconstruction model. In practice, the execution entity can obtain the predicted colors of each pixel point on each scene image through the methods of sub-step one to sub-step five in the first step above, calculate the mean square error between the predicted color and the real color of each pixel point on each scene image, perform weighted processing on each mean square error to obtain a loss function, and update the parameters of the constructed 3D reconstruction model according to the loss function using the Adam optimizer to obtain a pre-trained 3D reconstruction model.
[0036] Third step: For each scene image in the above-mentioned scene image set, perform the following steps:
[0037] Sub-step one: Perform secondary sampling on the rays corresponding to each pixel point in the above-mentioned scene image to obtain each secondary sampling point and the positions and ray directions corresponding to each secondary sampling point. Among them, the above-mentioned secondary sampling can be high-frequency uniform sampling. In practice, the execution entity can perform high-frequency uniform sampling on the rays corresponding to each pixel point in the above-mentioned scene image to obtain each secondary sampling point and the positions and ray directions corresponding to each secondary sampling point. Among them, the generation method of the positions and ray directions corresponding to each secondary sampling point can refer to the generation method of the positions and ray directions corresponding to each sampling point above, which will not be elaborated here.
[0038] Sub-step 2: Input the positions and ray directions corresponding to each of the above secondary sampling points into the pre-trained 3D reconstruction model to obtain the colors and densities of each secondary sampling point. In practice, the above-mentioned execution entity can input the positions and ray directions corresponding to each of the above secondary sampling points into the pre-trained 3D reconstruction model. Finally, the colors and densities of each secondary sampling point are obtained. Among them, the generation methods of the colors and densities of each secondary sampling point can refer to the methods for generating the colors and densities of each of the above sampling points, which will not be elaborated here.
[0039] Sub-step 3: Generate a depth image set and a predicted color map set according to the colors and densities of each secondary sampling point. Among them, the above-mentioned depth image set includes each depth image corresponding to each scene image. The above-mentioned predicted color map set includes each predicted color map corresponding to each scene image. In practice, for any scene image and any pixel point on the scene image, the above-mentioned execution entity can, according to the method in sub-step 5 of the first step above, obtain the secondary predicted color of the above pixel point, the transmittance of each secondary sampling point on the ray equation corresponding to the above pixel point, and the weights corresponding to each secondary sampling point. Then, the distances of each secondary sampling point (that is, the distance from the secondary sampling point to the optical center, which is a predefined value because of uniform sampling) can be extracted. Next, multiply the distances of each secondary sampling point on the ray equation corresponding to the above pixel point by the weights corresponding to each secondary sampling point respectively, then sum up these products, and then divide by the sum of the weights corresponding to each secondary sampling point to obtain the depth value of the above pixel point. Thus, in this way, the depth values and secondary predicted colors of each pixel point in the scene image can be obtained. Then, normalizing and visualizing the depth values and secondary predicted colors of each pixel point in the scene image can obtain the depth image and predicted color map corresponding to the scene image. Similarly, the depth images and predicted color maps corresponding to each scene image can be obtained, and the depth images corresponding to each scene image are integrated into a depth image set, and the predicted color maps corresponding to each scene image are integrated into a predicted color map set.
[0040] The above first to third steps are an inventive point of the embodiments of the present disclosure, which solve the problems of "when existing 3D reconstruction methods process complex scene images, there are problems of insufficient sampling accuracy, incomplete capture of subtle scene features, and inability to efficiently generate depth images and predict color maps". In the prior art, 3D reconstruction methods have deficiencies in ray generation, sampling strategies, and model optimization, and it is difficult to effectively adapt to image data of different scenes and different resolutions. At the same time, there are defects in model robustness and computational efficiency, resulting in low reconstruction accuracy and inability to fully meet the requirements of high-precision 3D reconstruction. To solve these problems, the present disclosure adopts a 3D reconstruction method based on deep learning. By optimizing ray generation and sampling strategies, it can more accurately capture subtle features in the scene; through multi-scale feature fusion and model optimization, it improves the adaptability of the model to different scenes and data; by optimizing sampling and model structures, it significantly improves computational efficiency and meets real-time requirements. The specific implementation process includes: First, extract camera parameters from the scene image and generate ray equations; Second, perform uniform sampling on the rays to obtain sampling points and their positions and ray directions; Third, input the positions and ray directions of the sampling points into the 3D reconstruction model to obtain the colors and densities of the sampling points; Fourth, generate the predicted color of the pixel points through the color integration formula; calculate the mean square error between the predicted color and the real color as the loss function, and use the Adam optimizer to update the model parameters to obtain a pre-trained 3D reconstruction model; Fifth, perform high-frequency uniform sampling on the rays to obtain secondary sampling points and their positions and ray directions. Then, input the secondary sampling points into the pre-trained 3D reconstruction model to generate the colors and densities of each secondary sampling point. Finally, combine the colors and densities of each secondary sampling point and calculate through the depth integration formula and the color integration formula to generate the depth and predicted color of each pixel point, and organize them into corresponding depth images and predicted color maps. Thus, the depth images and predicted color maps corresponding to all scene images can be integrated into a depth image set and a predicted color map set. Thereby, end-to-end optimization from scene images to depth images and predicted color maps can be achieved, which can effectively extract scene features, flexibly adapt to different scenes and data, optimize all links of 3D reconstruction, improve reconstruction accuracy and model robustness, and ultimately achieve high-precision 3D reconstruction.
[0041] Step 103: Generate a first perspective mask according to the scene image corresponding to the first perspective.
[0042] In some embodiments, the above execution subject may generate a first perspective mask according to the scene image corresponding to the first perspective.
[0043] In some optional implementation manners of some embodiments, the above execution subject may generate a first perspective mask according to the scene image corresponding to the first perspective through the following steps:
[0044] Step 1: Annotate the target object and the occlusion area on the scene image corresponding to the above first perspective to obtain each sparse annotation point on the scene image corresponding to the above first perspective. Among them, the above scene image includes the target object and the occlusion area. The above target object is the area where the occluded object (i.e., the target object) in the scene is displayed on the scene image. The above occlusion area is the area where the object (i.e., the occluder) that occludes the target object in the scene is displayed on the scene image. Each sparse annotation point on the scene image corresponding to the above first perspective includes the annotation points of the target object and the annotation points of the occlusion area on the scene image corresponding to the first perspective. In practice, the above execution entity can retrieve the scene image corresponding to the first perspective from the scene image set. Then, by identifying that the user uses an interactive annotation tool (such as Label Studio, CVAT, or a custom annotation interface) to manually click or box the sparse annotation points of the target object and the occlusion area on the scene image corresponding to the first perspective (such as 3 - 10 points on the target object or the occlusion area), each sparse annotation point on the scene image corresponding to the first perspective is obtained.
[0045] Step 2: Generate a first - perspective mask based on the above - mentioned sparse annotation points and the segmentation model. Among them, the above first - perspective mask includes the target - object mask and the occlusion - area mask in the scene image corresponding to the above first perspective. The above target - object mask can be used to represent the area where the target object is located in the image. The above occlusion - area mask can be used to represent the area where the occluder is located in the image. The above segmentation model can be a SAM model that supports sample segmentation based on point prompts. The above segmentation model can be input with each sampling point and the image of any image, and can output the mask of the image. In practice, the above execution entity can input the scene image corresponding to the first perspective and each sparse annotation point on the scene image corresponding to the first perspective into the above segmentation model to obtain the target - object mask and the occlusion - area mask in the scene image corresponding to the first perspective, that is, the first - perspective mask.
[0046] Step 104: Generate each second - perspective mask based on the first - perspective mask and the depth - image set.
[0047] In some embodiments, the above execution entity can generate each second - perspective mask based on the above first - perspective mask and the depth - image set.
[0048] In some optional implementation manners of some embodiments, the above execution entity can generate each second - perspective mask based on the above first - perspective mask and the depth - image set through the following steps:
[0049] Step 1: Sample the above first - perspective mask to obtain respective first - perspective sampling points. Among them, each first - perspective sampling point includes sampling points of each first - perspective target object and sampling points of each first - perspective occlusion area. In practice, the above - mentioned execution entity can sample respectively in the target object mask and the occlusion area mask in the scene image corresponding to the first perspective (for example, sample 1 point per 100 pixel points) to obtain sampling points of each first - perspective target object and sampling points of each first - perspective occlusion area, that is, each first - perspective sampling point.
[0050] Step 2: Map the above - mentioned respective first - perspective sampling points to each second perspective to obtain respective mapping points of each second perspective. In practice, the above - mentioned execution entity can construct a propagation tree. This propagation tree takes the first perspective as the root node, selects each second perspective as a child node, and recursively propagates the sampling points in the pre - order traversal order (from the root to the left child to the right child). First, use the depth image corresponding to the first perspective in the above - mentioned depth image set to convert the 2D pixel coordinates of each first - perspective sampling point into 3D spatial coordinates. Then, extract the parameter data corresponding to the first perspective and the parameter data corresponding to any second - perspective child node from the parameter data corresponding to the above - mentioned respective scene images. Next, calculate the change in the external camera parameters corresponding to the first perspective and the external camera parameters corresponding to any second - perspective child node to obtain a pose transformation matrix. Map the 3D spatial coordinates of any sampling point in the first perspective to the camera coordinate system of this second perspective through this pose transformation matrix. Then, combined with the parameter data corresponding to this second perspective, project the point mapped to this second perspective onto the 2D pixel coordinate system of this second perspective to obtain the mapping point of this second perspective for this sampling point in the first perspective. Similarly, the mapping points of each second perspective can be obtained.
[0051] Step 3: Based on the above - mentioned depth image set, filter the mapping points of the above - mentioned respective second perspectives to obtain respective sampling points of each second perspective. In practice, for the respective mapping points of any second perspective, the above - mentioned execution entity can retrieve the depth image corresponding to this second perspective in the depth image set. Then, according to the 2D coordinates of the respective mapping points of this second perspective, compare the depth values of the respective mapping points of this second perspective and the pixel points at the corresponding coordinate positions on the depth image corresponding to this second perspective, and filter out the mapping points that exceed the image boundary. The mapping points with a depth - value change less than the tolerance threshold (such as 0.05 of the depth value) and within the image range are used as the sampling points of this second perspective. Similarly, the sampling points of each second perspective can be obtained.
[0052] Step 4: Generate respective second - perspective masks based on the sampling points of the above - mentioned respective second perspectives. Among them, the above - mentioned respective second - perspective masks include the target - object masks and occlusion - area masks in the scene images corresponding to the respective second perspectives. In practice, the above - mentioned execution entity can extract the scene images corresponding to the respective second perspectives from the scene - image set. Then, input the scene image corresponding to any second perspective and the sampling points of this second perspective into the above - mentioned segmentation model to obtain the target - object mask and occlusion - area mask in the scene image corresponding to this second perspective, that is, this second - perspective mask. Similarly, respective second - perspective masks can be obtained. Among them, the generation method of the respective second - perspective masks can refer to the above - mentioned first - perspective mask generation method, which will not be elaborated here.
[0053] Step 105: Generate non - occlusion - area supervision information based on multi - perspective data and the set of predicted color maps.
[0054] In some embodiments, the above - mentioned execution entity can generate non - occlusion - area supervision information based on multi - perspective data and the above - mentioned set of predicted color maps. Among them, the above - mentioned multi - perspective data includes the above - mentioned scene - image set and respective perspective masks. The above - mentioned respective perspective masks include the first - perspective mask and respective second - perspective masks.
[0055] In some optional implementation manners of some embodiments, the above - mentioned execution entity can generate non - occlusion - area supervision information based on multi - perspective data and the above - mentioned set of predicted color maps through the following steps:
[0056] Step 1: Execute the following steps on the scene image corresponding to each perspective in the above - mentioned scene - image set:
[0057] Sub - step 1: Obtain the perspective mask corresponding to the above - mentioned scene image according to the multi - perspective data. Among them, the perspective mask corresponding to the above - mentioned scene image is the mask corresponding to the perspective of this scene image in the above - mentioned respective perspective masks. In practice, the above - mentioned execution entity can extract the perspective corresponding to this scene image from the above - mentioned respective perspective masks. Then, extract the target - object mask and occlusion - area mask of this perspective in the above - mentioned respective perspective masks as the mask of this scene image.
[0058] Sub - step 2: Generate the non - occlusion - area mask of the above - mentioned scene image according to the perspective mask corresponding to the above - mentioned scene image. Among them, the above - mentioned non - occlusion - area mask includes the target object and the background area (i.e., the area not occluded by an object). The above - mentioned non - occlusion - area mask is used to represent the area in the scene image except the occluder. The mask of the above - mentioned scene image includes the target - object mask and occlusion - area mask in the perspective mask corresponding to this scene image. In practice, the above - mentioned execution entity can take the inverse of the occlusion - area mask in the mask of the above - mentioned scene image and take the union with the target - object mask of the above - mentioned scene image to obtain the non - occlusion - area mask.
[0059] Sub-step 3: Generate a global difference map of the above scene image according to the above predicted color map set. Among them, each element in the above global difference map represents the difference value of the corresponding pixel point. In practice, the above execution entity can retrieve the predicted color map corresponding to the above scene image from the above predicted color map set. Then, the predicted color map of the above scene image and the above scene image can be compared pixel by pixel by calculating the color difference value (for example, calculating the mean square error for each RGB channel of each pixel). Then, the differences of each pixel point in the two images are integrated into a difference matrix with the same size as the scene image according to the position of each pixel point. Finally, the above difference matrix is used as the global difference map.
[0060] Sub-step 4: Generate non-occluded region supervision information according to the non-occluded region mask of the above scene image and the global difference map of the above scene image. In practice, the above execution entity can use the region where the above non-occluded region mask is located to screen out the difference values of each pixel point in the non-occluded region from the above global difference map. Then, the above difference values of each pixel point are averaged to obtain the non-occluded region loss of the above scene image. Similarly, the non-occluded region losses of the scene images corresponding to each perspective in the above scene image set can be obtained. The non-occluded region losses of the scene images corresponding to each perspective are averaged to obtain the non-occluded region supervision information.
[0061] Step 106: Generate occluded region supervision information based on multi-perspective data and a depth image set.
[0062] In some embodiments, the above execution entity can generate occluded region supervision information based on the above multi-perspective data and the above depth image set.
[0063] In some optional implementation manners of some embodiments, the above execution entity can generate occluded region supervision information based on the above multi-perspective data and the above depth image set through the following steps:
[0064] Step 1: Extract a target perspective and each source perspective from each perspective according to the scene data of each perspective. Among them, the above target perspective is a predefined perspective to be processed. The above source perspectives are each perspective adjacent to the target perspective based on geometric relationships (distance and viewing angle). The above geometric relationships include a distance condition (for example, the Euclidean distance between the optical centers of the source perspective and the target perspective is less than a threshold (such as 25 cm)) and an angle condition (such as the angle between the optical center axes of the two perspectives is less than 30 degrees). In practice, the above execution entity can calculate the distance between the optical centers of each perspective and the target perspective and the angle between the axes where the optical centers are located according to the parameter data corresponding to each perspective. Then, each perspective that satisfies the distance condition and the angle condition is used as each source perspective.
[0065] Step 2: Generate a color image for filling the occluded area in the target view based on the scene images corresponding to each source view and the above-mentioned view masks. Among them, the above-mentioned color image for filling can be an image generated by mapping the areas missing in the target view due to the target object being occluded by the occluder using the scene images of each source view and the view masks of each source view. The above-mentioned view masks include the above-mentioned source view masks. The above-mentioned source view masks include the target object masks and occluded area masks of each source view. In practice, for any source view, the execution entity can extract the above-mentioned source view mask from the above-mentioned view masks. Then, map the colors of each pixel of the target object mask in the above-mentioned source view mask and the target object in the scene image corresponding to the above-mentioned source view to the occluded area of the target view through homography transformation. Then, screen the colors of each pixel of the target object mask in the above-mentioned source view mask and the target object in the scene image corresponding to the above-mentioned source view that are mapped to the occluded area of the target view, so as to remove the parts with the same masks in the target object masks corresponding to the source view and the target view and the parts with the same colors in each pixel of the target object corresponding to the source view and the target view, and obtain the mapping result of the above-mentioned source view to the above-mentioned target view. Similarly, the mapping results of each source view to the above-mentioned target view can be obtained. Finally, perform weighted fusion on the mapping results of each source view to the above-mentioned target view (the smaller the angle between the source view and the target view, the higher the weight; the source view with a more complete observation of the occluded area has a higher weight, such as the occluder is not occluded in the source view; the matching degree between the color of the pixel points after mapping and the geometric structure (such as depth, surface normal) of the target view can be evaluated through the reprojection error, and the higher the matching degree, the higher the weight), and obtain the color image for filling the occluded area in the above-mentioned target view.
[0066] Step 3: Generate a loss value for filling the occluded area in the target view based on the above-mentioned color image for filling the occluded area in the target view. In practice, the execution entity can calculate the mean squared error of the pixel colors to compare the pixel points of the color image for filling the occluded area in the target view with the pixel points corresponding in position in the occluded area of the predicted color map corresponding to the target view, calculate the mean squared error of each pixel point in the occluded area of the target view, and finally average the mean squared errors of each pixel point to obtain the loss value for filling the occluded area.
[0067] Step 4: Based on the above-mentioned perspective masks and the depth image corresponding to the target perspective in the above-mentioned depth image set, generate the pixel depth of the occluded area and the pixel depth of the target object in the depth image corresponding to the target perspective. Among them, the above-mentioned target perspective mask includes the target object mask and the occluded area mask corresponding to the target perspective. The above-mentioned pixel depth can be the average depth value of each pixel point. In practice, the above-mentioned execution entity can extract each pixel point of the occluded area and each pixel point of the target object in the scene image corresponding to the target perspective according to the area represented by the target perspective mask. Then, on the depth image corresponding to the target perspective in the above-mentioned depth image set, the depth values of each pixel point of the occluded area in the depth image corresponding to the target perspective and the depth values of each pixel point of the target object in the depth image corresponding to the target perspective are extracted. Taking the average of the depth values of each pixel point of the occluded area in the depth image corresponding to the target perspective, the pixel depth of the occluded area in the depth image corresponding to the target perspective is obtained. Similarly, the pixel depth of the target object in the depth image corresponding to the target perspective can be obtained.
[0068] Step 5: Based on the pixel depth of the occluded area and the pixel depth of the target object in the depth image corresponding to the target perspective, generate a depth constraint error. Among them, the above-mentioned depth error can reflect the pixel depth difference between the occluded area and the target object in the depth image corresponding to the target perspective. In practice, when the pixel depth of the above-mentioned occluded area is greater than the pixel depth of the above-mentioned target object, the generated depth constraint error is the difference between the pixel depth of the above-mentioned occluded area and the pixel depth of the above-mentioned target object. When the pixel depth of the above-mentioned occluded area is less than or equal to the pixel depth of the above-mentioned target object, the generated depth constraint error is 0.
[0069] Step 6: Integrate the above-mentioned target perspective occluded area completion loss value and the above-mentioned depth constraint error into occluded area supervision information. In practice, the above-mentioned execution entity can perform weighted summation on the above-mentioned target perspective occluded area completion loss value (such as the weight is 1) and the above-mentioned depth constraint error (such as the weight is 0.5) to obtain the occluded area supervision information.
[0070] Step 107: Train a de-occluded three-dimensional model according to the non-occluded area supervision information and the occluded area supervision information.
[0071] In some embodiments, the above-mentioned execution entity can train a de-occluded three-dimensional model based on the above-mentioned non-occluded area supervision information and the above-mentioned occluded area supervision information. Among them, the above-mentioned de-occluded three-dimensional model can be a model obtained by introducing a total loss function into the loss function of the above-mentioned pre-trained three-dimensional reconstruction model and training. The above-mentioned total loss function includes the above-mentioned non-occluded area supervision information and the above-mentioned occluded area supervision information. In practice, the above-mentioned execution entity can perform weighted fusion on the above-mentioned non-occluded area supervision information and the above-mentioned occluded area supervision information according to a preset initial weight of the supervision information to obtain a total loss function (the model parameters can be updated by backpropagating the total loss function). Among them, the above-mentioned weighted fusion is a method of combining different data or information. When training a de-occluded three-dimensional model, the non-occluded area supervision information and the occluded area supervision information respectively represent the performance of the model in the non-occluded area and the occluded area. By assigning different weights to these two pieces of information, their contributions to model training can be comprehensively considered. The assignment of weights can be determined according to their importance or their impact on model performance (for example, if the supervision information in the non-occluded area has a greater impact on the accuracy of the model, a higher weight can be assigned to it). Then, the above-mentioned total loss function is introduced into an adaptive optimizer (such as Adam) to dynamically adjust the learning rate, and the model parameters are updated by the method of backpropagation to minimize the total loss function. When it is recognized that the loss in the training process no longer significantly decreases or reaches a preset number of iterations, the training is stopped to obtain a de-occluded three-dimensional model.
[0072] Optionally, the above-mentioned execution entity can also perform the following steps:
[0073] Step 1: Based on the scene data from the above-mentioned perspectives, generate respective pseudo-perspective poses between adjacent perspectives. Among them, the above-mentioned pseudo-perspectives are virtual perspectives generated through an interpolation algorithm, which are used to enhance the generalization ability of 3D scene modeling. The above-mentioned pseudo-perspective poses include quaternions and translation vectors. In practice, the above-mentioned execution entity can convert the rotation matrix in the external camera parameters corresponding to each perspective into the quaternion corresponding to each perspective. Then, perform spherical linear interpolation among the quaternions corresponding to each adjacent perspective and according to the preset interpolation parameter (generally between 0 and 1). Among them, the above-mentioned interpolation parameter can be the scalar value used in spherical linear interpolation, which is used to control the progress of interpolation. It defines the position of interpolation between two quaternions. For example, when the difference parameter is 0, the interpolation result is equal to the previous quaternion; when the difference parameter is 1, the interpolation result is equal to the subsequent quaternion; when the difference parameter is between 0 and 1, the interpolation result is a new quaternion (a certain intermediate rotation state between the previous and subsequent quaternions). Obtain the quaternions corresponding to each pseudo-perspective. Then, perform linear interpolation on the translation vectors in the external camera parameters corresponding to each adjacent perspective to obtain the translation vectors corresponding to each pseudo-perspective. Integrate the quaternions corresponding to each pseudo-perspective and the translation vectors corresponding to each pseudo-perspective into each pseudo-perspective pose.
[0074] Step 2: Based on the above-mentioned respective pseudo-perspective poses, generate a set of depth images of pseudo-perspectives and a set of predicted color maps of pseudo-perspectives. In practice, the above-mentioned execution entity can convert each pseudo-perspective pose into the external camera parameters corresponding to each pseudo-perspective. Then, use the above-mentioned camera intrinsics and the above-mentioned pre-trained 3D reconstruction model to obtain a set of depth images of pseudo-perspectives and a set of predicted color maps of pseudo-perspectives. The specific implementation method can refer to the generation methods of the above-mentioned set of depth images and the above-mentioned set of predicted color maps, which will not be elaborated here.
[0075] Step 3: Add the above-mentioned set of depth images of pseudo-perspectives and the above-mentioned set of predicted color maps of pseudo-perspectives to the training set. Among them, the above-mentioned training set can be the data set used to train the above-mentioned occluded-free 3D model. The above-mentioned training set includes the set of depth images corresponding to each of the above-mentioned scene images and the set of predicted color maps corresponding to each of the above-mentioned scene images. In practice, the above-mentioned execution entity can add the above-mentioned set of depth images of pseudo-perspectives and the above-mentioned set of predicted color maps of pseudo-perspectives to the above-mentioned training set to expand the training set and enhance robustness.
[0076] Step 108: Based on the occluded-free 3D model, perform an occlusion removal operation on the scene image corresponding to the target perspective in the scene image set to obtain the occluded-free scene image of the target perspective.
[0077] In some embodiments, the execution subject may perform a deocclusion operation on the scene image corresponding to the target perspective in the scene image set based on the deocclusion three-dimensional model to obtain the deocclusion scene image of the target perspective. In practice, first, the execution subject renders the scene image set using the deocclusion three-dimensional reconstruction model to generate a depth image set containing geometric structures and a predicted color map set containing geometric structures. The specific implementation method can refer to the generation method of the depth image set and the predicted color map set, which will not be repeated here. And the occlusion area is accurately located in combination with each perspective mask (such as the part of the car body blocked by the tree trunk in the outdoor scene). Then, the source perspective with the best visibility of the occlusion area is selected from the adjacent perspectives of the target perspective (such as the left perspective can observe the blocked car door, and the top perspective can see the roof). The specific implementation method can refer to the above-mentioned method for extracting the target perspective and each source perspective, which will not be repeated here. Afterwards, each pixel in the occlusion area of the target perspective is projected to each source perspective image plane through three-dimensional coordinate transformation, and the color value of the pixel that is not blocked at the corresponding position is extracted (such as the red door texture in the left perspective). Then, according to the angle of view, depth matching (obtained from the above-mentioned depth image set containing geometric structures, such as the depth difference of the door from the source view to the target view is ≤0.2m) and color similarity (obtained from the above-mentioned predicted color map set containing geometric structures, such as the color difference value with the unobstructed area around the target is less than 5), weights are dynamically assigned. For example, the left view gets 70% weight due to high geometric consistency, and the top view only accounts for 30% due to large depth deviation. Then, the candidate pixel values of all source views are weighted fused to generate a preliminary restoration result (such as the mixed texture of the door with the main color of the left view red and the details of the top view after fusion). Finally, using the geometric constraints of the depth map of the target view (such as the depth mutation boundary between the trunk and the car body), the Poisson fusion algorithm is used to optimize the color transition and eliminate the seam artifacts, and finally, the de-occluded scene image of the target view is obtained. For example, in street scenes, the de-occlusion operation can make the car body blocked by tree trunks appear completely, and the door outline and headlight structure are consistent with the real scene, without color discontinuity or geometric distortion. The real scene is restored to a high degree in both geometric structure (door outline error <0.5 pixel) and visual perception (PSNR>32dB).
[0078] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the occlusion removal method based on neural radiance fields in some embodiments of the present disclosure, occlusion removal and target object reconstruction can be achieved, and the stability of the reconstruction effect, as well as the efficiency and accuracy of occlusion removal, are improved in different scenarios. Thereby, the quality and practicality of 3D scene reconstruction are enhanced. Specifically, traditional occlusion removal methods such as foreground-background separation techniques and image inpainting techniques based on external 2D vision priors often have the following technical problems: The above-mentioned foreground-background separation techniques are difficult to effectively reconstruct the target object in scenarios with a large occlusion area or a high complexity of the occluder, resulting in poor rendering quality. The above-mentioned image inpainting techniques based on external 2D vision priors are difficult to provide convincing filled content and consistent 3D geometric information when dealing with severe occlusions, which may lead to unnatural artifacts and reduce the realism and credibility of the reconstruction results to a certain extent. In addition, existing methods often have difficulty in effectively utilizing multi-view cues when dealing with complex occlusion scenarios, resulting in poor stability and quality of the reconstruction results. Based on this, the occlusion removal method based on neural radiance fields in some embodiments of the present disclosure, first, obtains the scene data of each view to obtain a set of scene images, where the above-mentioned each view includes a first view and each second view. Thereby, a comprehensive multi-view data basis is provided for subsequent occlusion removal, ensuring the diversity and geometric consistency of the data. Then, according to the above set of scene images, a set of depth images and a set of predicted color maps are generated. Thereby, by combining depth information and color information, rich feature support is provided for the identification and completion of occlusion regions, improving the accuracy of reconstruction. Next, according to the scene image corresponding to the first view, a first view mask is generated. Thereby, by combining a sparse annotation and segmentation model, the target object and the occlusion region are located, providing a reliable basis for the generation of subsequent multi-view masks. Then, based on the first view mask and the set of depth images, each second view mask is generated. Thereby, through the mapping of sampling points and depth screening, the annotation information of the first view is effectively propagated to other views, ensuring the consistency and accuracy of multi-view data. Then, based on the multi-view data and the set of predicted color maps, non-occluded region supervision information is generated. Thereby, by combining a global difference map and a non-occluded region mask, the reconstruction quality of the non-occluded region is evaluated, providing a reliable supervision signal for model optimization. At the same time, based on the multi-view data and the set of depth images, occlusion region supervision information is generated. Thereby, by combining the generation of the filled color image of the source view and the depth constraint error, it is ensured that the completion of the occlusion region conforms to both color consistency and geometric depth logic, improving the reconstruction effect of the occlusion region. Finally, according to the non-occluded region supervision information and the occlusion region supervision information, a de-occluded 3D model is trained, and based on the above de-occluded 3D model, a de-occlusion operation is performed on the target view to obtain a de-occluded scene image of the target view.Thus, through the comprehensive utilization of multi-view clues and the dynamic optimization of the model, the removal of occluded regions and the reconstruction of target objects are achieved. The generated occluded-removed scene images have a sense of reality and credibility, and can, to a certain extent, handle complex occluded scenes. In addition, through the generation of pseudo-view poses and data augmentation, the generalization ability and robustness of the model are further improved, thereby enhancing the stability of the reconstruction effect in different scenes. Moreover, by combining multi-view clues with deep learning techniques, the efficiency and accuracy of occlusion removal are improved, providing a more reliable and practical solution for fields such as 3D scene reconstruction, virtual reality, and augmented reality.
[0079] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an occlusion removal device based on neural radiance fields. These device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.
[0080] As Figure 2 shown, some embodiments of the occlusion removal device 200 based on neural radiance fields include: an acquisition unit 201 configured to acquire scene data from each perspective to obtain a set of scene images, where the above perspectives include a first perspective and each second perspective; a prediction unit 202 configured to generate a set of depth images and a set of predicted color maps according to the above set of scene images; a first generation unit 203 configured to generate a first perspective mask according to the scene image corresponding to the first perspective; a second generation unit 204 configured to generate each second perspective mask based on the first perspective mask and the above set of depth images; a first supervision unit 205 configured to generate non-occluded region supervision information based on multi-view data and the above set of predicted color maps, where the above multi-view data includes the above set of scene images and each perspective mask; a second supervision unit 206 configured to generate occluded region supervision information based on the above multi-view data and the above set of depth images; a training unit 207 configured to train and obtain a de-occluded three-dimensional model according to the above non-occluded region supervision information and the above occluded region supervision information; and a de-occlusion unit 208 configured to perform a de-occlusion operation on the scene image corresponding to the target perspective in the above set of scene images based on the above de-occluded three-dimensional model to obtain a de-occluded scene image of the target perspective.
[0081] It can be understood that the units described in the device 200 correspond to the respective steps in the method described with reference to Figure 1 . Thus, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units included therein, and will not be elaborated herein.
[0082] Next, referring toFigure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0083] As Figure 3 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage device 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0084] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 3 Each block shown in
[0085] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are executed.
[0086] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0087] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0088] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain scene data of each perspective to obtain a set of scene images, where the above each perspective includes a first perspective and each second perspective; generate a set of depth images and a set of predicted color maps according to the above set of scene images; generate a first perspective mask according to the scene image corresponding to the above first perspective; generate each second perspective mask based on the above first perspective mask and the above set of depth images; generate non-occluded region supervision information based on multi-perspective data and the above set of predicted color maps, where the above multi-perspective data includes the above set of scene images and each perspective mask; generate occluded region supervision information based on the above multi-perspective data and the above set of depth images; train to obtain a de-occluded three-dimensional model according to the above non-occluded region supervision information and the above occluded region supervision information; perform a de-occlusion operation on the scene image corresponding to the target perspective in the above set of scene images based on the above de-occluded three-dimensional model to obtain a de-occluded scene image of the target perspective.
[0089] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0091] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a prediction unit, a first generation unit, a second generation unit, a first supervision unit, a second supervision unit, a training unit, and an occlusion removal unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit can also be described as "the unit that acquires scene data from various perspectives to obtain a set of scene images".
[0092] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0093] Some embodiments of the present disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any one of the above-described occlusion removal methods based on neural radiance fields.
[0094] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.
Claims
1. A method for occlusion removal based on neural radiance fields, comprising: Obtaining scene data from each perspective to obtain a set of scene images, where the each perspective includes a first perspective and each second perspective; Generating a set of depth images and a set of predicted color maps according to the set of scene images; Generating a first perspective mask according to the scene image corresponding to the first perspective; Generating each second perspective mask based on the first perspective mask and the set of depth images; Generating non-occluded region supervision information based on multi-perspective data and the set of predicted color maps, where the multi-perspective data includes the set of scene images and each perspective mask; Generating occluded region supervision information based on the multi-perspective data and the set of depth images; Training a de-occluded three-dimensional model according to the non-occluded region supervision information and the occluded region supervision information; Performing an occlusion removal operation on the scene image corresponding to the target perspective in the set of scene images based on the de-occluded three-dimensional model to obtain a de-occluded scene image of the target perspective.
2. The method according to claim 1, wherein The method further includes: Generating each pseudo-perspective pose between adjacent perspectives according to the scene data of each perspective; Generating a set of depth images of the pseudo-perspective and a set of predicted color maps of the pseudo-perspective according to each pseudo-perspective pose; Adding the set of depth images of the pseudo-perspective and the set of predicted color maps of the pseudo-perspective to the training set.
3. The method according to claim 1, wherein, The generating a first perspective mask according to the scene image corresponding to the first perspective includes: Annotating the target object and the occluded region on the scene image corresponding to the first perspective to obtain each sparse annotation point on the scene image corresponding to the first perspective; Generating a first perspective mask based on each sparse annotation point and a segmentation model, where the first perspective mask includes a target object mask and an occluded region mask in the scene image corresponding to the first perspective.
4. The method according to claim 1, wherein, The generating each second perspective mask based on the first perspective mask and the set of depth images includes: Sampling the first perspective mask to obtain each first perspective sampling point; Mapping each first perspective sampling point to each second perspective to obtain each mapped point of the second perspective; Filtering each mapped point of the second perspective based on the set of depth images to obtain each sampling point of the second perspective; Generating each second perspective mask based on each sampling point of the second perspective.
5. The method according to claim 1, wherein, The generating non-occluded region supervision information based on multi-perspective data and the set of predicted color maps includes: Performing the following steps on the scene image corresponding to each perspective in the set of scene images: Obtaining the perspective mask corresponding to the scene image according to the multi-perspective data; Generating a non-occluded region mask of the scene image according to the perspective mask corresponding to the scene image; Generating a global difference map of the scene image according to the set of predicted color maps; Generating non-occluded region supervision information according to the non-occluded region mask of the scene image and the global difference map of the scene image.
6. The method according to claim 1, wherein The generating occluded region supervision information based on the multi-perspective data and the set of depth images includes: Extract a target view and each source view from each of the perspectives according to the scene data of each perspective; Generate a color image for filling in the occluded area in the target view according to the scene images corresponding to each source view and the perspective masks; Generate a loss value for filling in the occluded area in the target view according to the color image for filling in the occluded area in the target view; Generate the pixel depth of the occluded area and the pixel depth of the target object in the depth image corresponding to the target view according to the perspective masks and the depth image corresponding to the target view in the depth image set; Generate a depth constraint error based on the pixel depth of the occluded area and the pixel depth of the target object in the depth image corresponding to the target view; Integrate the loss value for filling in the occluded area in the target view and the depth constraint error into occlusion area supervision information.
7. An occlusion removal device based on a neural radiance field, comprising: An acquisition unit configured to acquire scene data of each perspective to obtain a set of scene images, where each of the perspectives includes a first perspective and each second perspective; A prediction unit configured to generate a set of depth images and a set of predicted color maps according to the set of scene images; A first generation unit configured to generate a first perspective mask according to the scene image corresponding to the first perspective; A second generation unit configured to generate each second perspective mask based on the first perspective mask and the set of depth images; A first supervision unit configured to generate non-occluded area supervision information based on multi-perspective data and the set of predicted color maps, where the multi-perspective data includes the set of scene images and each perspective mask; A second supervision unit configured to generate occlusion area supervision information based on the multi-perspective data and the set of depth images; A training unit configured to train a de-occluded three-dimensional model according to the non-occluded area supervision information and the occlusion area supervision information; A de-occlusion unit configured to perform a de-occlusion operation on the scene image corresponding to the target view in the set of scene images based on the de-occluded three-dimensional model to obtain a de-occluded scene image of the target view.
8. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Method, device and equipment for removing shielding of real-time mixed reality head-mounted display
CN120997890A