Method executed by electronic equipment and electronic equipment

By introducing semantic information of spatial points into the NeRF method, determining 2D and 3D semantic features and predicting density information, the problem of low rendering quality of occlusion areas in multi-object scenes is solved, and a higher quality 3D object reconstruction is achieved.

CN120047599APending Publication Date: 2025-05-27BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311585850.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing NeRF-based methods are difficult to achieve high-quality rendering or 3D reconstruction in scenarios containing multiple objects, especially when there is occlusion between objects.

Method used

By introducing semantic information of spatial points, the 2D and 3D semantic features of each spatial point are determined, the density information is predicted, and the 3D content reconstruction model is trained to improve the accuracy of density prediction and the rendering quality in the occluded area.

Benefits of technology

The density prediction accuracy and rendering quality of occluded areas in multi-object scenes are improved, and a higher quality 3D object reconstruction is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047599A_ABST
    Figure CN120047599A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method executed by electronic equipment and the electronic equipment, and relates to the field of artificial intelligence. The method comprises the following steps: based on a multi-view image, determining 2D semantic features of each spatial point under multiple views; determining a first 3D semantic feature of each spatial point based on the 2D semantic feature of each spatial point under multiple view angles; based on the first 3D semantic feature of each spatial point, predicting density information of each spatial point; and training a 3D content reconstruction model based on the density information of each spatial point. Optionally, the method performed by the electronic device may be performed using an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision, and more specifically, to a method executed by an electronic device, the electronic device, and a storage medium. Background Art

[0002] Neural Radiance Fields (NeRF) is one of the hottest research directions in the field of computer vision. The problem it aims to solve is how to generate images from new perspectives given some captured images. NeRF can model the scene as a continuous multi-dimensional radiation field and store it implicitly in the neural network. It only needs to input several multi-perspective images with camera poses to train a neural radiance field model, and based on this model, clear images from any perspective can be rendered. In addition, 3D object reconstruction can also be performed based on NeRF. 3D object reconstruction aims to restore the geometry, texture and other models of the target object from a video or image sequence.

[0003] Although existing NeRF-based methods can render object images from new perspectives or reconstruct the 3D model of an object in a scene containing a single object, it is difficult to achieve high-quality rendering or reconstruction in a scene containing multiple objects. Summary of the invention

[0004] The present disclosure provides a method executed by an electronic device, an electronic device, and a storage medium to at least solve the above-mentioned problems in the related art.

[0005] According to a first aspect of an embodiment of the present disclosure, a method executed by an electronic device is provided, including: determining, based on multi-view images, 2D semantic features of each spatial point under multi-views; determining, based on the 2D semantic features of each spatial point under multi-views, first 3D semantic features of each spatial point; predicting density information of each spatial point based on the first 3D semantic features of each spatial point; and training a 3D content reconstruction model based on the density information of each spatial point.

[0006] Optionally, the method further includes: acquiring a display viewing angle and target object information input by a user; using the trained 3D content reconstruction model to render an image of the target object at the display viewing angle and displaying the rendered image.

[0007] Optionally, based on the 2D semantic features of each spatial point under multiple perspectives, the first 3D semantic features of each spatial point are determined, including: based on the 2D semantic features of each spatial point under multiple perspectives, determining the first semantic feature consistency of each spatial point under different perspectives and / or the second semantic feature consistency of each spatial point with the target object; based on the first semantic feature consistency and / or the second semantic feature consistency, determining the first 3D semantic feature of each spatial point.

[0008] Optionally, determining the 2D semantic features of each spatial point under multiple perspectives includes: determining an unobstructed mask of the target object under multiple perspectives; and determining the 2D semantic features of each spatial point under multiple perspectives based on the unobstructed mask of the target object under multiple perspectives.

[0009] Optionally, determining an unobstructed mask of the target object under multiple perspectives includes: obtaining a second 3D semantic feature of each spatial point based on the multi-perspective image using a neural network; and determining an unobstructed mask of the target object under multiple perspectives based on the second 3D semantic feature of each spatial point.

[0010] Optionally, based on the second 3D semantic features of each spatial point, an unobstructed mask of the target object under multiple perspectives is determined, including: determining the sampling probability of each spatial point under multiple perspectives based on the second 3D semantic features of each spatial point; sampling the spatial points based on the sampling probability of each spatial point under multiple perspectives; for each perspective, based on the second 3D semantic features of each spatial point, screening the spatial points belonging to the target object from the sampled spatial points, rendering the screened spatial points, and obtaining an unobstructed mask of the target object under the perspective.

[0011] Optionally, determining the consistency of the first semantic features of each spatial point under different perspectives includes: based on the 2D semantic features of each spatial point under multiple perspectives and the mean of the 2D semantic features, using a first attention network to determine the consistency of the first semantic features of each spatial point under different perspectives.

[0012] Optionally, determining the consistency of the second semantic features of each spatial point with the target object includes: determining the consistency of the second semantic features of each spatial point with the target object based on the 2D semantic features of each spatial point under multiple perspectives and the 2D semantic features of the spatial points of the target object under multiple perspectives.

[0013] Optionally, based on the first semantic feature consistency and / or the second semantic feature consistency, the first 3D semantic feature of each spatial point is determined, including: based on the first semantic feature consistency and the second semantic feature consistency, respectively determining a first weight and a second weight; based on the 2D semantic features of each spatial point under multiple perspectives, the first weight and the second weight, using a second attention network to determine the first 3D semantic feature of each spatial point.

[0014] Optionally, the method also includes: determining the image features of the spatial points of the target object under an unobstructed perspective and / or the viewing distances between each perspective; predicting the color information of each spatial point under multiple perspectives based on the image features of the spatial points of the target object under an unobstructed perspective and / or the viewing distances between each perspective; training a 3D content reconstruction model based on the density information of each spatial point, including: training a 3D content reconstruction model based on the density information of each spatial point and the color information of each spatial point under multiple perspectives.

[0015] Optionally, determining the image features of the target object's spatial point in an unobstructed perspective includes: determining overlapping information of the target object's spatial point in multiple perspectives based on an occlusion mask and an unobstructed mask of the target object in multiple perspectives; determining the image features of the target object's spatial point in an unobstructed perspective based on the overlapping information of the target object's spatial point in multiple perspectives and the image features of the multi-perspective images; wherein the overlapping information represents the probability that the target object's spatial point and the spatial points of other objects are at the same pixel position in the two-dimensional image.

[0016] Optionally, predicting the color information of each spatial point under multiple perspectives includes: determining the 3D color features of each spatial point based on the image features of the spatial point of the target object under an unobstructed perspective and / or the perspective distance between each perspective; determining the color information of each spatial point under multiple perspectives based on the 3D color features of each spatial point.

[0017] Optionally, determining the 3D color features of each spatial point includes: using the viewing distance between each viewing angle to weight the image features of the spatial point of the target object at an unobstructed viewing angle to obtain a first weighted image feature; determining a third weight based on the image features of the spatial point of the target object at an unobstructed viewing angle and the first weighted image feature; obtaining a fourth weight based on the viewing distance between each viewing angle; and determining the 3D color features of each spatial point using a third attention network based on the image features of the spatial point of the target object at an unobstructed viewing angle, the third weight and the fourth weight.

[0018] Optionally, the semantic features include: object category features and / or object instance features.

[0019] According to a second aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, prompt the at least one processor to execute the method described above.

[0020] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided. When the instructions are executed by a processor, the method described above is implemented.

[0021] According to an embodiment of the present disclosure, in order to solve the problem of inaccurate density prediction of spatial points, the accuracy of density prediction of spatial points is improved by introducing semantic information of spatial points. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The basic process of NeRF is shown.

[0023] Figure 2 A NeRF-based 3D reconstruction method is shown.

[0024] Figure 3 An example of a reconstructed object is shown.

[0025] Figure 4 A flowchart of an image processing method according to an embodiment of the present disclosure is shown.

[0026] Figure 5 A schematic diagram of obtaining the first semantic feature and the first density information according to an embodiment of the present disclosure is shown.

[0027] Figure 6 A schematic diagram of an instance-aware density correction method according to an embodiment of the present disclosure is shown.

[0028] Figure 7 Another schematic diagram of the instance-aware density correction method according to an embodiment of the present disclosure is shown.

[0029] Figure 8 Another schematic diagram of the instance-aware density correction method according to an embodiment of the present disclosure is shown.

[0030] Fig. 9 A schematic diagram of obtaining an unobstructed segmented image according to an embodiment of the present disclosure is shown.

[0031] Fig.10 A schematic diagram of extracting multi-view 2D semantic features of spatial points and target objects according to an embodiment of the present disclosure is shown.

[0032] Fig.11 A schematic diagram of a semantically guided density correction network module according to an embodiment of the present disclosure is shown.

[0033] Fig.12 A schematic diagram of a color correction method based on instance perception according to an embodiment of the present disclosure is shown.

[0034] Fig.13 Another schematic diagram of the example-aware color correction method according to an embodiment of the present disclosure is shown.

[0035] Fig.14A diagram showing the overlap of target space points under multiple viewing angles according to an embodiment of the present disclosure.

[0036] Fig.15 A schematic diagram of an overlap and viewing angle distance guided color correction network according to an embodiment of the present disclosure is shown.

[0037] Fig.16 A NeRF-based multi-object 3D model reconstruction method according to an embodiment of the present disclosure is shown.

[0038] Fig.17 A schematic diagram of a specific application scenario according to an embodiment of the present disclosure is shown. Fig.17 An example of reconstructing a specified object from an occluded scene is shown.

[0039] Fig.18 A schematic structural diagram of an electronic device applicable to an embodiment of the present invention is shown in FIG. DETAILED DESCRIPTION

[0040] The following description with reference to the accompanying drawings is provided to facilitate a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should be considered as exemplary only. Therefore, one of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, for the sake of clarity and conciseness, descriptions of well-known functions and structures may be omitted.

[0041] The terms and expressions used in the following specification and claims are not limited to their dictionary meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustration purposes only and not for the purpose of limiting the present disclosure as defined in the appended claims and their equivalents.

[0042] It should be understood that the singular forms "a", "an", and "the" may also include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more such surfaces. When we refer to an element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or the one element and the other element may establish a connection relationship through an intermediate element. In addition, "connected" or "coupled" as used herein may include wireless connection or wireless coupling.

[0043] The term "include" or "may include" refers to the presence of the corresponding disclosed functions, operations or components that can be used in various embodiments of the present disclosure, rather than limiting the presence of one or more additional functions, operations or features. In addition, the term "include" or "have" may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components or combinations thereof, but should not be interpreted as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components or combinations thereof.

[0044] The term "or" used in various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple, or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" may be implemented as parameter A including A1 or A2 or A3, or may be implemented as parameter A including at least two of the three items A1, A2, and A3.

[0045] Unless defined differently, all terms (including technical terms or scientific terms) used in the present disclosure have the same meanings as understood by those skilled in the art described in the present disclosure. Common terms as defined in dictionaries are interpreted as having meanings consistent with the context in the relevant technical field, and should not be interpreted ideally or overly formally unless clearly defined in the present disclosure.

[0046] At least some functions of the device or electronic device provided in the embodiments of the present disclosure can be implemented by an AI model, such as at least one module among multiple modules of the device or electronic device can be implemented by an AI model. Functions associated with AI can be performed by non-volatile memory, volatile memory and processor.

[0047] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or pure graphics processing units, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor, such as a neural processing unit (NPU).

[0048] The one or more processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are provided by training or learning.

[0049] Here, providing by learning means obtaining a predefined operating rule or an AI model with desired characteristics by applying a learning algorithm to a plurality of learning data. The learning can be performed in the device or electronic device itself in which the AI ​​according to the embodiment is executed, and / or can be implemented by a separate server / system.

[0050] The AI ​​model may include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations by calculating between the input data of the layer (such as the calculation results of the previous layer and / or the input data of the AI ​​model) and the multiple weight values ​​of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.

[0051] A learning algorithm is a method of using a plurality of learning data to train a predetermined target device (e.g., a robot) to enable, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0052] The method provided in the present disclosure may involve one or more technical fields such as speech, language, image, video or data intelligence.

[0053] Optionally, when it comes to the field of speech or language, in a method performed by an electronic device according to the present disclosure, a speech signal as an analog signal can be received via a speech input device (e.g., a microphone), and the speech portion can be converted into computer-readable text using an automatic speech recognition (ASR) model. The user's speech intention can be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or the NLU model can be an artificial intelligence model. The artificial intelligence model can be processed by an artificial intelligence dedicated processor designed in a hardware structure specified for artificial intelligence model processing. Language understanding is a technology for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.

[0054] Optionally, when it comes to the field of images or videos, in the method performed by the electronic device according to the present disclosure, output data can be obtained by using image data as input data of an artificial intelligence model. The method of the present disclosure may be related to the field of visual understanding of artificial intelligence technology, which is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.

[0055] Optionally, when it comes to the field of intelligent data processing, in the method performed by the electronic device according to the present disclosure, in the inference or prediction stage, an artificial intelligence model can be used to perform predictions by using real-time input data. The processor of the electronic device can perform preprocessing operations on the data to convert it into a form suitable for use as an input to the artificial intelligence model. Inference prediction is a technology for logical reasoning and prediction by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning or recommendation.

[0056] In the present application, an artificial intelligence model can be obtained by training. Here, "obtained by training" means obtaining a predefined operating rule or artificial intelligence model configured to perform a desired feature (or purpose) by training a basic artificial intelligence model with multiple training data through a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and the neural network calculation is performed by calculating between the calculation result of the previous layer and the multiple weight values.

[0057] The following describes several optional embodiments to illustrate the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure. It should be noted that the following embodiments can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.

[0058] Figure 1 The basic process of NeRF is shown. Figure 1 The left side shows the use of a camera to shoot a purple sphere in a scene, where the camera is currently located at a spatial point o, the shooting angle is d, and d is a unit vector in Cartesian space. Let a spatial point in the sphere be x, and the spatial point is also located on the ray of o along the direction of d. The ray can be represented as o+td, where t represents the straight-line distance from the point on the ray to point o. Figure 1 The right side of shows the scene on the left reconstructed using NeRF technology, which is represented as a set of 3D space points. Each space point has two attributes, one of which is the density (density) that is independent of the viewing angle, denoted as σ, and the other is the color that is related to the viewing angle, denoted as RGB. Density can be understood as the opacity of a space point, representing the probability that a ray is blocked when passing through the space point, and this probability is differentiable. For example, Figure 1 The surface and interior of the purple sphere in the figure are solid, so they have high density values, while the air outside the sphere has low density values. During model training, the NeRF model learns to predict the density and color of a point in space from the three-dimensional position encoding γ(x) and the viewing direction encoding γ(d), where γ() represents the encoding method, which is usually implemented by trigonometric functions, and γ(x)60 and γ(d) 24 The subscript number indicates the dimension of the encoded vector obtained after encoding. Figure 1 The numbers 256 and 128 in the blue box on the right represent the dimensions of the output vector of the corresponding layer of the neural network. After the model training is completed, the NeRF model can predict the density and color of any spatial point in the corresponding scene at any viewing angle. Therefore, accurate density and color prediction has a crucial impact on the quality of 3D object reconstruction.

[0059] NeRF is used to reconstruct all content in a video or image scene. To reconstruct a specific object in the scene, a segmented image with the target object is usually used as the input of the NeRF model to filter out non-target objects during the reconstruction process. Figure 2 A NeRF-based 3D reconstruction method is shown. Figure 2 After obtaining multi-view images with camera poses, a semantic segmentation module (such as an instance segmentation model) is first used to obtain segmented images of the target object (or specific object). These segmented images are then combined with the multi-view images as input to the NeRF model. After the NeRF model is trained, a 3D model or a new view image of the target object can be obtained. Figure 3 An example of reconstructing an object is shown. First, a segmented image of the object to be reconstructed is extracted from a multi-view image using a semantic segmentation method, and then these paired images (i.e., the segmented image of the object to be reconstructed and the corresponding multi-view image) are used as input to train a NeRF model containing only the object. After the model training is completed, the NeRF model can be used to render the object image under a new perspective, or the mesh model of the object can be extracted from the NeRF model. This stage can be called the inference stage.

[0060] Although the above method can achieve the reconstruction of the 3D model of the object in a scene containing a single object, it is difficult to achieve high-quality reconstruction in a scene containing multiple objects. Because in a scene containing multiple objects, the objects will block each other at some viewing angles, resulting in the image collected at this viewing angle being unable to capture the local information of the target object (target object). Due to the lack of information about the target object at some viewing angles, the above NeRF-based reconstruction method cannot accurately predict the density and color of the target object in the area where the information is lost, and ultimately causes the reconstructed 3D model to be partially missing. To solve the above problems, the present disclosure provides the following technical solutions.

[0061] The embodiments of the present disclosure propose an instance-aware density correction method to solve the problem of inaccurate density prediction of spatial points caused by occlusion. Considering that density prediction in object reconstruction is related to the semantics of the object, in order to improve the accuracy of density prediction, the present disclosure introduces semantic features of spatial points (such as object categories or instance information) to increase the density prediction value of spatial points belonging to the target object and reduce the density of other non-target object points, thereby improving the accuracy of density prediction of spatial points in occluded areas.

[0062] Figure 4 A flowchart of a method executed by an electronic device according to an embodiment of the present disclosure is shown. Figure 4 The method shown can be applied to object reconstruction in a scene where the target object (also referred to as the target object) is not occluded, and can also be applied to object reconstruction in a scene where the target object is occluded.

[0063] Reference Figure 4 In step S401, based on the multi-view images, the 2D semantic features of each spatial point under the multi-view are determined. The present disclosure takes into account that density prediction in object reconstruction is related to the semantics of the object, and therefore introduces the semantic features of the spatial points to predict the density information of the spatial points. Here, the semantic features may include object category features and / or object instance features.

[0064] As an example, based on the multi-view images, a neural network can be used to obtain the second 3D semantic features of each spatial point, and based on the second 3D semantic features of each spatial point, the unobstructed mask (also referred to as a mask) of the target object under multiple viewpoints can be determined. Based on the unobstructed mask of the target object under multiple viewpoints, the 2D semantic features of each spatial point under multiple viewpoints can be determined.

[0065] For example, for a certain scene, it can be photographed from multiple perspectives to obtain scene images of multiple perspectives (i.e., multi-perspective images). A NeRF model with semantic output, such as a semantic NeRF model (Semantic NeRF), can be used to obtain a rough 3D semantic feature (i.e., a second 3D semantic feature) of any spatial point in the scene.

[0066] Next, the sampling probability of each spatial point under multiple perspectives can be determined based on the second 3D semantic features of each spatial point, and spatial point sampling can be performed based on the sampling probability of each spatial point under multiple perspectives. For each perspective, based on the second 3D semantic features of each spatial point, spatial points belonging to the target object can be screened from the sampled spatial points, and the screened spatial points can be rendered to obtain an unobstructed mask of the target object under that perspective. Fig. 9 Describe in detail how to obtain the unoccluded mask.

[0067] Based on the above unobstructed mask, the semantic features of each spatial point at each viewing angle can be obtained respectively. For each spatial point, the semantic features of the spatial point at each viewing angle are fused to obtain the 2D semantic features of each spatial point at multiple viewing angles. Fig.10 Describe in detail how to obtain 2D semantic features of spatial points under multiple views.

[0068] In step S402, a first 3D semantic feature of each spatial point is determined based on the 2D semantic features of each spatial point under multiple viewing angles.

[0069] As an example, based on the 2D semantic features of each spatial point under multiple perspectives, the consistency of the first semantic features of each spatial point under different perspectives and / or the consistency of the second semantic features of each spatial point and the target object can be determined. Then, based on the consistency of the first semantic features and / or the consistency of the second semantic features, the first 3D semantic features of each spatial point are determined.

[0070] In the present disclosure, the first semantic feature consistency can be understood as the features of the spatial point under multiple perspectives are similar or consistent, and therefore, the first semantic feature consistency can also be referred to as inter-perspective consistency / similarity. The second semantic feature consistency can be understood as the features of the spatial point under multiple perspectives are similar or consistent with the features of the target object, and therefore, the second semantic feature consistency can also be referred to as target object consistency / similarity.

[0071] For the similarity of the first semantic features, the first attention network can be used to determine the consistency of the first semantic features of each spatial point under different perspectives based on the 2D semantic features of each spatial point under multiple perspectives and the mean of the 2D semantic features.

[0072] For the second semantic feature similarity, based on the 2D semantic features of each spatial point under multiple perspectives and the 2D semantic features of the spatial points of the target object under multiple perspectives, the consistency of the second semantic features of each spatial point with the target object is determined.

[0073] Next, the first weight and the second weight may be determined based on the consistency of the first semantic feature and the consistency of the second semantic feature, respectively. Based on the 2D semantic features of each spatial point under multiple perspectives, the first weight and the second weight, the second attention network is used to determine the first 3D semantic features of each spatial point. Fig.11 How to obtain the first 3D semantic feature is described in detail.

[0074] In step S403, based on the first 3D semantic features of each spatial point, the density information of each spatial point is predicted. According to an embodiment of the present disclosure, rough density prediction information can be obtained first, for example, Semantic NeRF can be used to obtain the rough 3D semantic features (i.e., the second 3D semantic features) and rough density prediction values ​​of any spatial point in a scene. Then, the rough density prediction information is corrected using the first 3D semantic features of each spatial point to determine / predict the density information of each spatial point.

[0075] The following will refer to Figure 7 , Figure 8 and Fig.11 Describe in detail how to predict the final density information.

[0076] In step S404, a 3D content reconstruction model is trained based on the density information of each spatial point.

[0077] After obtaining the 3D content reconstruction model, the user can input the display perspective and target object information, and the 3D content reconstruction model can render the image of the target object under the display perspective and display the rendered image. Alternatively, the 3D content reconstruction model can perform 3D reconstruction on the target object.

[0078] Figure 5 A schematic diagram of obtaining a second 3D semantic feature and a rough density prediction value according to an embodiment of the present disclosure is shown.

[0079] Reference Figure 5 , Figure 5 The left side of represents the reconstructed scene. First, the camera is used to collect images from multiple perspectives. Then, any pixel in these images is randomly selected. The ray representation corresponding to the pixel in space is calculated using the camera pose. A certain number of spatial points are randomly sampled along the ray. After encoding the position information x and the perspective information d of the spatial point, γ(x) and γ(d) are obtained and input into the NeFR network (such as Semantic NeRF). Finally, the NeFR network predicts the coarse density σ and 3D semantic feature s corresponding to each spatial point.

[0080] Figure 6 A schematic diagram of an instance-aware density correction method according to an embodiment of the present disclosure is shown.

[0081] Reference Figure 6 , taking the rough 3D semantic features as input, projecting them to 2D, obtaining multi-view 2D semantic features, performing feature fusion on them, and then back-translating them to 3D to obtain the fine 3D semantic features of the spatial points (i.e., the first 3D semantic features), and using the fine 3D semantic features to correct the rough density prediction values, so as to ultimately improve the accuracy of density prediction.

[0082] In order to obtain accurate semantic features of spatial points, the present disclosure designs two modules, namely, the semantic-aware sampling and rendering module and the semantic-guided density correction network module. The following will describe in detail how to use these two modules.

[0083] Figure 7 Another schematic diagram of the instance-aware density correction method according to an embodiment of the present disclosure is shown.

[0084] Reference Figure 7 ,In S1, Semantic NeRF can be used to obtain rough 3D semantic features (i.e., the second 3D semantic features).

[0085] In S2, the semantic-aware sampling and rendering module is used to perform sampling and rendering to obtain an unobstructed segmented image of the target object (i.e., an unobstructed mask).

[0086] The semantic-aware sampling and rendering module may include a semantic-aware sampling module and a semantic-aware rendering module. The semantic-aware sampling module may adjust the sampling probability of these spatial points during rendering according to the rough 3D semantic features of each spatial point. By increasing the sampling probability of the target object and reducing the sampling probability of the non-target object, an unobstructed segmented image of the target object can be generated at any viewing angle during rendering. The semantic-aware rendering module may selectively render the sampled points based on the rough 3D semantic features during the rendering process, wherein selective rendering refers to using the probability that the spatial point belongs to the target object as a weight during rendering, which can reduce the impact of the sampling points of the non-target object on the rendering result, thereby improving the quality of the unobstructed segmented image of the target object.

[0087] In S3, a feature extraction module is used to extract multi-view 2D semantic features from the unobstructed segmented image. For example, feature extraction is performed from the unobstructed segmented image and pixel alignment features are calculated to obtain multi-view 2D semantic features. Fig.10 Describe in detail how to obtain multi-view 2D semantic features.

[0088] In S4, fine 3D semantic features of the spatial points are extracted from the multi-view 2D semantic features. Specifically, the rough 3D semantic features, the position information x of the spatial points and the multi-view 2D semantic features can be input into the semantic-guided density correction network module, and the fine 3D semantic features can be extracted through the semantic-guided density correction network module to improve the accuracy of the spatial point density prediction of the occluded area.

[0089] Figure 8 Another schematic diagram of the instance-aware density correction method according to an embodiment of the present disclosure is shown. Figure 8The method shown demonstrates the role of the semantic-aware sampling and rendering module and the semantic-guided density correction network module in reconstructing the target object.

[0090] Reference Figure 8 , assuming that the target object (green sphere) is occluded by other objects (blue cube), the red curve on the ray represents the sampling probability corresponding to each spatial point during rendering. First, the semantic-aware sampling and rendering module uses rough 3D semantic features to adjust the sampling probability so that the sampling center of gravity is offset to the target object. Sampling using the sampling probability at this time can obtain an unobstructed segmented image of the target object (i.e., an unobstructed mask), which can be used to extract multi-view 2D semantic features. Then, the semantic-guided density correction network module assigns a higher density value to the target object (green sphere) based on the multi-view 2D semantic features of the spatial points, and assigns a lower density value to other objects (blue cubes). Figure 8 The red curve on the far right indicates that the sampling center of gravity is moving closer to the target object.

[0091] Fig. 9 A schematic diagram of obtaining an unobstructed segmented image according to an embodiment of the present disclosure is shown. Assume that there are two objects that need 3D reconstruction, a blue cube and a green sphere. Under certain camera viewing angles, the green sphere is obstructed by the blue cube.

[0092] Reference Fig. 9 In S2a, first, the semantic-aware sampling module can adjust the sampling probability of the spatial point during rendering according to the rough 3D semantic features of the spatial point. Specifically, Fig. 9 The purple curve in the figure represents the rough density prediction value of different spatial points on a sampling ray, while the yellow curve represents the probability that the spatial point belongs to the target object. The rough density prediction value of each spatial point is multiplied by the sampling probability of the spatial point belonging to the target object, and then normalized to obtain Fig. 9 The blue curve in the figure represents the sampling probability distribution of the spatial points on the ray during rendering. This distribution can guide the sampling center to the target object, thereby reducing the sampling probability of non-target object points. Unlike traditional NeRF, which is only based on rough density value sampling, this method can render and generate unobstructed target object segmentation images at any viewing angle.

[0093] In addition, to improve the quality of the rendered segmented image, semantic-aware rendering can be further adopted. Fig. 9In S2b, it is assumed that 6 spatial points are sampled during rendering, including 2 spatial points of non-target objects (blue points) and 4 spatial points of target objects (green points). During volume rendering, the coarse 3D semantic features will be used as the rendering weight of each spatial point to reduce the pixel contribution of the spatial points of non-target objects to the segmented image, and the quality of the unobstructed segmented image will be further improved by filtering the non-target object sampling points. Finally, the rendered unobstructed segmented image will be used to fuse the fine 3D semantic features after feature extraction.

[0094] Fig.10 A schematic diagram of extracting multi-view 2D semantic features of spatial points and target objects according to an embodiment of the present disclosure is shown. Here, 2D semantic features under multiple perspectives may be referred to as multi-view 2D semantic features or 2D multi-view semantic features, and multi-view 2D semantic features may be formed based on 2D semantic features of a spatial point or target object under multiple perspectives, and multi-view 2D semantic features may be used to calculate the first semantic feature consistency (similarity between perspectives) and the second semantic feature consistency (similarity of target objects).

[0095] After obtaining the unobstructed segmented image of the target object, a feature extraction network can be used to obtain multi-view 2D semantic features from the unobstructed segmented image. In particular, the multi-view 2D semantic features belonging to the spatial points of the target object have both inter-view similarity (i.e., the 2D semantic features of the spatial points at each view are similar) and target object similarity (the 2D semantic features of the spatial points at each view are similar to the 2D semantic features of the target object).

[0096] Reference Fig.10 , assuming there are two objects in the scene, a green target object and a blue non-target object. 1 is a spatial point belonging to the target object, and P 2 and P 3 Spatial points belonging to non-target objects. Fig.10 As shown in the left part, the target object space point P 1 The projection at any viewing angle will fall into the segmented area of ​​the unobstructed segmented image (i.e. the green area). Rather than the target object space point P 2 and P 3 Only when the target object is blocked will it be projected into the segmented area (such as P in perspective 1). 2 ), otherwise these points will fall outside the segmentation area (such as P 2 and P 3 ).

[0097] like Fig.10 As shown in the lower right part of the figure, using any feature extraction network, we can get 2D semantic features from the above unobstructed segmented image, and use the camera pose and projection relationship to get P1 , P 2 and P 3 The position of the point on the corresponding feature map. Aggregate the corresponding features under each perspective to obtain and They represent the 2D semantic features of each point under multiple perspectives. Fig.10 It can be found that only and It has consistent 2D semantic features across different viewpoints, and the features show similarities between viewpoints.

[0098] Furthermore, in order to calculate the similarity of multi-view 2D semantic features to the target object, a set of 2D points O tar Extract multi-view 2D semantic features of target object points. Fig.10 As shown in the upper right part of The center of the segmented area in each view, the features corresponding to these positions are the standard 2D semantic features of the target object. Aggregating the 2D semantic features of the target object in each view can be obtained Compare and It can be found that has the highest similarity with the target object, The similarity is second, and has the lowest similarity.

[0099] Therefore, only the target object space point P 1 The multi-view 2D semantic features of the proposed method simultaneously show the similarity between views and the similarity with the target object (i.e., target object similarity). Therefore, these two similarities can be used to extract fine 3D semantic features of fused spatial points from the multi-view 2D semantic features to distinguish target object points from non-target object points.

[0100] Fig.11 A schematic diagram of a semantically guided density correction network module according to an embodiment of the present disclosure is shown.

[0101] The structure of the semantically guided density correction network is as follows Fig.11As shown, it consists of an inter-perspective similarity calculation module (including the first attention network), a dual-weighted attention module (including the second attention network), and an MLP-based density prediction layer. Specifically, the inter-perspective similarity calculation module calculates the inter-perspective similarity of the 2D semantic features of spatial points at different perspectives through a cross-attention mechanism. Then the dual-weighted attention module combines the inter-perspective similarity of the 2D semantic features and the target object similarity, fuses the 2D semantic features at different perspectives, and extracts fine 3D semantic features from them. Guided by the two similarities, only the spatial points belonging to the target object can obtain rich and fine 3D semantic features, which can be used as the input of the density prediction layer to improve the performance of density prediction.

[0102] As an example, Fig.11 The lower left part of the figure is the view similarity calculation module, whose input is the multi-view 2D semantic feature P of any spatial point in the scene. test , according to Fig.10 The method shown here obtains P test First, the mean of the input features is calculated along the viewing dimension, which is used as the query vector Q for the attention mechanism (such as the first attention network) mean The 2D features of the same spatial point at each viewing angle will be used as the key vector. Taking viewing angle 1 as an example, the 2D features of the spatial point at viewing angle 1 will be used as the key vector K 1 , and so on. Then, the cross attention mechanism is used to calculate the similarity between the 2D features of the spatial point at each perspective and the mean features. Finally, the similarities under all perspectives are averaged to obtain the similarity between perspectives w 2 (i.e. the first weight), which corresponds to Fig.11 The purple line in the .

[0103] Fig.11 The upper left part is the dual-weight attention module, whose input is the multi-view 2D semantic features of the target object O tar , multi-view 2D semantic features P of spatial points test and the similarity between perspectives w 2 First, the multi-view 2D semantic features of the target object will be used to extract the query vector Q tar , and then extract the key vector and value vector corresponding to each perspective from the multi-perspective 2D semantic features of the spatial point. For example, the key vector and value vector of perspective 1 are respectively denoted as K 1 and V 1 Finally, an attention mechanism (such as the second attention network) is used to calculate the similarity w between the features of the spatial point at each viewpoint and the features of the target object. 1x (i.e., the second weight). Taking perspective 1 as an example, the attention mechanism is used to calculate the similarity w between the features of the spatial point in perspective 1 and the features of the target object 11, which corresponds to Fig.11 The yellow straight line in the figure, and so on. 1x With w 2 The weights are used together to guide the fusion of 2D semantic features under different perspectives, thereby extracting fine 3D semantic features. Specifically, taking perspective 1 as an example, w 11 With w 2 Both are scalars. First, we multiply them together to get a comprehensive weight, and then we multiply the weight by the value vector V of view 1. 1 Multiply them to get the weighted value vector under view 1. Finally, add the weighted value vectors of all viewpoints to get the detailed 3D semantic features of the spatial point. Since the feature fusion process considers both the similarity between viewpoints and the similarity of the target object, only the spatial points that meet both similarities can obtain rich 3D semantic features after feature fusion.

[0104] Fig.11 The right half of the figure shows the prediction process of the modified density. After obtaining the fine 3D semantic features, the fine 3D semantic features and the rough density information are sent as input to the multi-layer perceptron to obtain the modified density prediction value and the modified density features, thereby improving the performance of density prediction.

[0105] According to an embodiment of the present disclosure, an instance-aware color correction method is also proposed to solve the problem of inaccurate color prediction of spatial points caused by occlusion.

[0106] As an example, determine the image features of the target object's spatial point under an unobstructed perspective and / or the viewing distances between each perspective, and predict the color information of each spatial point under multiple perspectives based on the image features of the target object's spatial point under an unobstructed perspective and / or the viewing distances between each perspective.

[0107] After obtaining the color information of each spatial point under multiple viewing angles, a 3D content reconstruction model may be trained based on the density information of each spatial point and the color information of each spatial point under multiple viewing angles.

[0108] In the process of obtaining the image features of the target object's spatial point under an unobstructed viewing angle, the overlapping information of the target object's spatial point under multiple viewing angles can be determined based on the occlusion mask and the unobstructed mask of the target object under multiple viewing angles, wherein the overlapping information represents the probability that the target object's spatial point and the spatial points of other objects are at the same pixel position in the two-dimensional image. Then, based on the overlapping information of the target object's spatial point under multiple viewing angles and the image features of the multi-view images, the image features of the target object's spatial point under an unobstructed viewing angle are determined. Fig.14 Describe in detail the overlap of spatial points under multiple perspectives.

[0109] In the process of predicting the color information of each spatial point under multiple perspectives, the 3D color features of each spatial point can be determined based on the image features of the spatial point of the target object under an unobstructed perspective and / or the perspective distance between each perspective, and then the color information of each spatial point under multiple perspectives can be determined based on the 3D color features of each spatial point.

[0110] In the process of determining the 3D color features of each spatial point, the image features of the spatial point of the target object under an unobstructed perspective can be weighted using the perspective distance between each perspective to obtain a first weighted image feature; a third weight is determined based on the image features of the spatial point of the target object under an unobstructed perspective and the first weighted image feature; a fourth weight is obtained based on the perspective distance between each perspective; and the 3D color features of each spatial point are determined using a third attention network based on the image features of the spatial point of the target object under an unobstructed perspective, the third weight, and the fourth weight. Fig.12 , Fig.13 and Fig.15 Describe in detail how to determine the 3D color features of each spatial point and the color information of each spatial point under multiple viewing angles.

[0111] Fig.12 A schematic diagram of a color correction method based on instance perception according to an embodiment of the present disclosure is shown.

[0112] Reference Fig.12 First, a color image and a corresponding unobstructed segmented image (unobstructed mask) and an occluded segmented image (occluded mask) are obtained. Here, the color image may be an image taken from multiple perspectives (i.e., a multi-perspective 2D image). The occluded segmented image and the unobstructed segmented image may be obtained in the same manner as described above. Fig. 9 The method shown is obtained in a similar manner.

[0113] The color image and the corresponding occluded segmented image and unoccluded segmented image are used as input. The pixel aligned image features, occluded segmented map features, and unoccluded segmented map features of the spatial points at each viewing angle are calculated through feature extraction and camera pose relationship. The two types of segmented image features are used to predict the overlap of spatial points at the current viewing angle (expressed by the size of the overlap value). The overlap value is used to filter the unoccluded 2D color features (i.e., the image features of the spatial points at the unoccluded viewing angle) from the multi-view 2D image features.

[0114] The distance between the reference view and the test view (i.e., the view distance between each view) is calculated, and the unobstructed 2D color features of different view are fused with this as the weight to obtain a fine 3D color feature (i.e., the 3D color feature of the spatial point). Here, the test view can represent the current view, and the reference view can represent other view except the current view. In other words, for a certain view, the angles formed by the rays corresponding to the spatial point at the view and the rays corresponding to other view are calculated as the distance between the view and the other view.

[0115] The refined 3D color features will be used as input to the color correction network to obtain the color information of spatial points under multiple perspectives, thereby improving the accuracy of the model's color prediction in the occluded area of ​​the target object.

[0116] In the above method, the color feature extraction process is as follows: Fig.13 shown. Fig.13 Another schematic diagram of the example-aware color correction method according to an embodiment of the present disclosure is shown.

[0117] Reference Fig.13 , the multi-view 2D image features are used as the input of the whole method, and refined 3D color features are obtained by screening and fusing them through overlap values ​​and view distances, so as to ultimately improve the accuracy of color prediction.

[0118] Previous studies have shown that image features are very useful for color prediction. However, not all image features are useful at all viewing angles. In particular, due to the overlap of spatial position relationships at some viewing angles, overlap means that the spatial points of the target object and the spatial points of other objects overlap at the same pixel position in the 2D image after being projected at a specific viewing angle. An example is Fig.14 As shown, Fig.14 A diagram showing the overlap of target space points under multiple viewing angles according to an embodiment of the present disclosure.

[0119] Reference Fig.14 The camera collected images from 8 different viewing angles around a blue cube and a colored sphere (half orange and half green). Fig.14 The images at viewing angles 1, 2, and 5 are given below. There is a spatial point P on the surface of the sphere 1 , due to P 1 It is blocked by the blue cube in view 1 and by the sphere itself in view 4, 5, and 6, so the pixel information in the images at these view angles cannot correctly reflect P 1 The color of P is only contained in the images at viewing angles 2, 3, 7, and 8. 1 Correct color information. Combined Fig.12After obtaining the overlap value of the spatial point at each viewing angle, for each viewing angle, when the overlap value at the current viewing angle meets the preset condition, the color feature at the viewing angle can be selected, and when the overlap value at the current viewing angle does not meet the preset condition, the color feature at the viewing angle can be removed. In other words, the overlap value of the spatial point at each viewing angle can be used to screen the image features that can correctly reflect the color of the spatial point.

[0120] To this end, the present invention designs a color correction network guided by overlap and viewing distance, which selectively fuses accurate image features of spatial points from multi-view images using overlap values ​​and viewing distances, thereby improving the accuracy of color prediction of the NeRF model in the occluded area of ​​the target object.

[0121] Fig.15 A schematic diagram of an overlap and viewing angle distance guided color correction network according to an embodiment of the present disclosure is shown.

[0122] The color correction network guided by overlap and viewing distance has three inputs, namely, overlap value, pixel-aligned image features, and test and reference viewing angles. The overlap value is obtained by the overlap prediction network, whose input is the pixel-aligned segmented images of the target object with and without occlusion. Specifically, for any point in space, the feature vector of the corresponding pixel position can be extracted from the two segmented images through the projection relationship. These two vectors will be used as inputs of the overlap value prediction network to predict the overlap value of the point at the current viewing angle, and the overlap value is a scalar ranging from 0 to 1. The overlap value of the spatial point at different viewing angles can be obtained through the overlap prediction network. The overlap value of different viewing angles can be multiplied and combined with the pixel-aligned image features of the corresponding viewing angle to obtain the unoccluded multi-view image features (also called unoccluded multi-view color features, that is, the image features of the spatial point of the target object at the unoccluded viewing angle).

[0123] The unobstructed multi-view image features (i.e., the image features filtered based on the overlap value) will be combined with the distance weight to extract fine 3D color features. Taking view 1 as an example, the unobstructed multi-view image features are first used to calculate the key vector K of the third attention network. 1 Sum value vector V 1 . Then the first distance weight (corresponding to the fourth weight) is calculated using the distance between the test view and the reference view. In particular, the larger the distance between views, the smaller the weight, and the smaller the distance between views, the larger the weight. The first distance weight of the spatial point at each view is a scalar. Using it as a weight to weight the unobstructed features at the corresponding view can obtain a mean feature vector (i.e., the first weighted image feature), which is used to calculate the query vector Q of the third attention network. testUsing the third attention network to calculate the similarity between the unobstructed feature and the mean feature vector at each viewing angle, the second distance weight (corresponding to the third weight) can be obtained, corresponding to Fig.15 The yellow straight line in , and then the first distance weight (corresponding to the purple straight line) at the corresponding viewing angle is used to perform double weighting on the unobstructed image features to obtain the fine 3D color features of the spatial points. Specifically, the first distance weight and the second distance weight are two scalars, and the comprehensive weight can be obtained by multiplying the two. The comprehensive weight is the value vector V at the corresponding viewing angle. 1 Multiplying together can obtain a weighted value vector. For other viewing angles, calculations are performed in a similar manner. Finally, the weighted value vectors of all viewing angles are added together to obtain a fine 3D color feature.

[0124] Reference Fig.14 and Fig.15 First, the overlap prediction network predicts the overlap value of the target object's spatial point at each viewpoint based on the occluded segmentation image and the unoccluded segmentation image of the target object. For the viewpoints occluded by non-target objects, a lower weight (for example, view 1) is assigned when fusing its image features, thereby screening out the unoccluded multi-view image features. Then, the distance between the test viewpoint and each reference viewpoint is calculated, and a dual-weighted attention module (i.e., the third attention network) is used to extract fine 3D color features from the unoccluded multi-view image features. During the extraction process, since the distance between the test viewpoint and the reference viewpoint is taken into account, a lower weight is assigned to the self-occluded viewpoint (for example, viewpoints 4, 5, and 6), so that the fused 3D color features can correctly reflect the color information of the spatial points, and finally realize the correction of the color prediction of the spatial points in the occluded area.

[0125] Then, an MLP-based color prediction layer is used to obtain the final color features using the refined 3D color features, the modified density features, and the color features.

[0126] Fig.16 A NeRF-based multi-object 3D model reconstruction method according to an embodiment of the present disclosure is shown. Fig.16 The presented method can reconstruct a 3D representation of a target object from training data.

[0127] Reference Fig.16 ,In step S1, a rough 3D semantic feature is obtained.

[0128] First, we can use existing NeRF models with semantic output, such as Semantic NeRF, to obtain rough 3D semantic features and rough density prediction values ​​of any spatial point in the scene. The specific process is as follows: Figure 5 As shown, Figure 5The left side of is the reconstructed scene. First, the camera is used to collect images from multiple perspectives as training data. Then, any pixel in the training image is randomly selected, and the ray corresponding to the pixel in space is calculated using the camera pose. A certain number of spatial points are randomly sampled along the ray, and the position x and the perspective information d of the spatial point are encoded and input into the NeRF model. Finally, the NeRF model can predict the rough density σ and 3D semantic information s corresponding to each spatial point.

[0129] In step S2, an unobstructed segmented image is obtained.

[0130] Based on the rough 3D semantic features obtained in step S1, a semantic-aware sampling and rendering method is used for sampling and rendering to obtain an unobstructed segmented image of the target object at each viewing angle. These segmented images will be used to extract fine 3D semantic features after feature extraction.

[0131] The generation process of the unobstructed segmented image is as follows: Fig. 9 As shown. Assume that there are two objects that need 3D reconstruction, a blue cube and a green sphere. Under certain camera viewing angles, the green sphere is blocked by the blue cube. Fig. 9 In S2a, first, the semantic-aware sampling module can adjust the sampling probability of the point during rendering according to the rough 3D semantics of the spatial point. Specifically, Fig. 9 The purple curve in the figure represents the rough density prediction value of different spatial points on a sampling ray, while the yellow curve represents the probability that the point belongs to the target object. The rough density prediction value of each point is multiplied by the sampling probability of the point belonging to the target object, and then normalized to obtain Fig. 9 The blue curve in the figure represents the sampling probability distribution of the points on the ray when rendering. This distribution guides the sampling center to the target object, thereby reducing the sampling probability of non-target object points.

[0132] Furthermore, to improve the quality of the rendered segmented image, e.g. Fig. 9 As shown in S2b in Figure 2, it is assumed that 6 spatial points are sampled during rendering, including 2 non-target object points (blue points) and 4 target object points (green points). During volume rendering, the coarse 3D semantic features will be used as the rendering weight of each point to reduce the pixel contribution of non-target object points to the segmented image, and the quality of the unobstructed segmented image will be further improved by filtering the non-target object sampling points. Finally, the rendered unobstructed segmented image will be used to fuse the fine 3D semantic features after feature extraction.

[0133] Step S3, perform feature extraction to obtain multi-view 2D semantic features.

[0134] After obtaining the unobstructed segmented image of the target object, a feature extraction network is used to extract the multi-view 2D semantic features of the spatial points and the target object from the segmented image. In particular, the multi-view 2D semantic features of the spatial points of the target object have both inter-view similarity (the 2D semantic features of the spatial points in each view are similar) and target object similarity (the 2D semantic features of the spatial points in each view are similar to the 2D semantic features of the target object). An example is Fig.10 shown.

[0135] In step S4, a semantically guided density correction network is used to extract fine 3D semantic features and perform corrected density prediction.

[0136] After obtaining multi-view 2D semantic features, the semantic-guided density correction network will fuse 2D semantic features based on the similarity between viewpoints and the similarity with the target object to obtain refined 3D semantic features. Ultimately, the refined 3D semantic features will be used to predict the density value of the target object in space and improve the density prediction of the occluded area.

[0137] In step S5, unobstructed multi-view image features are extracted.

[0138] In order to improve the spatial point color prediction performance in the occluded area, the unoccluded image features are first extracted from the multi-view 2D image features using the overlap value. Fig.15 As shown in the figure, the input of the overlap prediction network is the occluded and unoccluded segmented images. By comparing the differences between the spatial points in the two segmented images, the network can predict whether the spatial points overlap at that perspective and output the overlap prediction value. Then the overlap prediction value of the spatial point at each perspective will be used as a weight to guide the fusion of image features from different perspectives, thereby obtaining the unoccluded multi-perspective image features.

[0139] In step S6, fine 3D color features are extracted and the color prediction is corrected.

[0140] The unobstructed multi-view image features obtained above will be used as the key vector K and value vector V of the attention network 1 , which can be used to further extract fine 3D color features from it in combination with the viewing distance.

[0141] Specifically, we first calculate the distance between the test view and the reference view of the spatial point. These distances are used as weights to calculate the mean vector of the multi-view unoccluded features, which is used as the query vector Q of the attention network. test Finally, the attention network is used to calculate the weight of each unobstructed feature and the mean feature, and the weight is weighted again in combination with the corresponding viewing distance. The selective fusion of double weights is used to extract fine 3D color features of spatial points from multiple unobstructed image features.

[0142] Fig.16 The right side of the figure shows the prediction process of the corrected color. After obtaining the refined 3D color features, they will be used together with the color features obtained in step S1 and step S4 and the corrected density features as the input of the multi-layer perceptron to predict the corrected color value, so as to improve the color prediction accuracy of the occluded area.

[0143] The corrected color values ​​and corrected density values ​​will be used to obtain a planar image using NeRF's volume rendering technology, and ultimately the aforementioned network will be optimized based on the color differences and semantic differences of image pixels.

[0144] Fig.17 A schematic diagram of a specific application scenario according to an embodiment of the present disclosure is shown. Fig.17 An example of reconstructing a specified object from an occluded scene is shown.

[0145] Reference Fig.17 , the hamburger and cup in the scene are the target objects that the user wants to reconstruct, but due to the mutual occlusion between the two in some viewing angles, the existing methods cannot obtain good reconstruction results. The method proposed in the present invention aims to solve the above problems. The specific steps are as follows: first, the user collects image data in the scene, and provides the category of the object to be reconstructed (hamburger and cup) in an interactive manner such as click and text. Note that providing the category here is optional. If the user does not provide the category of the object to be reconstructed, the proposed method will reconstruct all objects in the scene during training.

[0146] The model training process is as follows Fig.17 As shown in , all collected image data will first be used to train the Semantic NeRF of the current scene, which models the rough semantic information and rough density information of all spatial points in the current scene. Then, according to Fig.16 The operations of steps S2-S6 will optimize the network parameters in the instance-aware density correction module and the instance-aware color correction module in turn according to the category of the object to be reconstructed provided by the user. Fig.17 In the scenario described, we first set the hamburger as the target object to be reconstructed, and the cup as other objects. Then, we randomly sample pixels from the training image as training samples. For the spatial sampling points on the ray corresponding to the pixels, we follow Fig.16 The operation of step S2 - performing semantic-aware sampling and rendering to obtain an unobstructed segmented image, and then following the steps Fig.16 The operation of step S3 is to perform feature extraction to obtain multi-view 2D semantic features, according to Fig.16 The operation of step S4 - using the semantically guided density correction network, extracts fine 3D semantic features and performs corrected density prediction, according to Fig.16The operation of step S5 is to extract the unobstructed multi-view image features according to Fig.16 The operation of step S6 is to extract fine 3D color features and correct color prediction. Finally, after volume rendering, the pixel prediction value is obtained, and the difference between it and the corresponding pixel true value in the training data set will be used to guide the optimization of network parameters. Then, the cup is set as the target object, the hamburger is set as other objects and the above training steps are repeated to optimize the network again. Since the network input contains the category information of the object to be reconstructed, different objects can share the same network, that is, this method can achieve the reconstruction of multiple objects through one network.

[0147] After the model training is completed, users can view the reconstructed 3D model in two ways. An example is Fig.17 As shown in the figure, in the first method, the user specifies the target object (such as a hamburger) through interactive methods such as text and clicks, and sets the viewing angle parameters to obtain the rendered image of the 3D model of the target object at a specific viewing angle. In the second method, by combining with the existing NeRF-oriented mesh extraction method, the user can obtain the mesh model of the specified object and finally view and edit it in the 3D software.

[0148] According to the embodiments of the present disclosure, the present disclosure may also be applied to the following user scenarios:

[0149] 1. 3D content sharing: When traveling outdoors, users want to take photos of landmark buildings in scenic spots and share them, but the buildings are blocked by tourists or other objects; users want to take photos of museum exhibits and share them in 3D, but the exhibits block each other and users cannot move the exhibits. Users want to share dishes at a dinner party, but there are many dishes and they block each other.

[0150] 2. 3D face models, 3D anthropomorphic expressions, and user-defined 3D resource libraries in mobile phone albums. The present invention can improve the ability of mobile phones to obtain 3D models and reduce the cost of obtaining 3D models. Users can easily reconstruct various object models from obstructed images. It provides more diverse and user-defined 3D content on mobile phones and other devices.

[0151] This disclosure proposes a 3D object reconstruction method based on neural radiation field and corresponding application scenarios. In order to solve the problem of inaccurate spatial point density prediction caused by occlusion, this disclosure designs a density correction method based on instance perception, which introduces the semantic information of spatial points to improve the accuracy of density prediction of spatial points in occluded areas; in order to solve the problem of inaccurate color prediction of spatial points caused by occlusion, this disclosure adopts a color correction method based on instance perception, which selectively fuses 2D image features under different perspectives through overlapping conditions and viewing angle distance to improve the accuracy of color prediction of the model in the occluded area of ​​the target object.

[0152] An embodiment of the present disclosure also provides an electronic device, which includes a processor and, optionally, may also include at least one transceiver and / or at least one memory coupled to the at least one processor, wherein the at least one processor is configured to execute the steps of the method provided in any optional embodiment of the present disclosure.

[0153] Fig.18 A schematic diagram of the structure of an electronic device applicable to an embodiment of the present invention is shown in FIG. Fig.18 As shown, Fig.18 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, each of the processor 4001, the memory 4003 and the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure. Optionally, the electronic device may be a first network node, a second network node or a third network node.

[0154] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0155] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.18Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0156] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.

[0157] The memory 4003 is used to store computer programs or executable instructions for executing the embodiments of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the above method embodiments.

[0158] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program or instructions stored thereon. When the computer program or instructions are executed by at least one processor, the steps and corresponding contents of the aforementioned method embodiment can be executed or implemented.

[0159] The embodiments of the present disclosure also provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.

[0160] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described in the text.

[0161] It should be understood that, although the flowchart of the embodiment of the present disclosure indicates each operation step by arrows, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios with different execution times, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present disclosure does not limit this.

[0162] The above text and drawings are provided only as examples to help readers understand the present disclosure. They are not intended and should not be interpreted as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, based on the contents disclosed herein, it is obvious to those skilled in the art that the embodiments and examples shown can be changed without departing from the scope of the present disclosure, and other similar implementation means based on the technical ideas of the present disclosure are adopted, which also fall within the protection scope of the embodiments of the present disclosure.

Claims

1. A method performed by an electronic device, include: Based on multi-view images, determine the 2D semantic features of each spatial point under multiple viewpoints; Determine a first 3D semantic feature of each spatial point based on the 2D semantic features of each spatial point under multiple perspectives; Predicting density information of each spatial point based on the first 3D semantic feature of each spatial point; Based on the density information of each spatial point, the 3D content reconstruction model is trained.

2. The method according to claim 1, further comprising: include: Obtaining the display viewing angle and target object information input by the user; The trained 3D content reconstruction model is used to render an image of the target object at the display viewing angle and display the rendered image.

3. The method according to claim 1, in, Based on the 2D semantic features of each spatial point under multiple perspectives, a first 3D semantic feature of each spatial point is determined, including: Based on the 2D semantic features of each spatial point under multiple perspectives, determining the consistency of the first semantic features of each spatial point under different perspectives and / or the consistency of the second semantic features of each spatial point and the target object; Based on the first semantic feature consistency and / or the second semantic feature consistency, a first 3D semantic feature of each spatial point is determined.

4. The method according to claim 1, in, Determine the 2D semantic features of each spatial point under multiple perspectives, including: Determine the unobstructed mask of the target object under multiple perspectives; Based on the unobstructed mask of the target object under multiple perspectives, the 2D semantic features of each spatial point under multiple perspectives are determined.

5. The method according to claim 4, in, Determine the unobstructed mask of the target object under multiple perspectives, including: Based on multi-view images, a neural network is used to obtain the second 3D semantic features of each spatial point; Based on the second 3D semantic features of each spatial point, the unobstructed mask of the target object under multiple perspectives is determined.

6. The method according to claim 5, in, Based on the second 3D semantic features of each spatial point, determine the unobstructed mask of the target object under multiple perspectives, including: Determine the sampling probability of each spatial point under multiple viewing angles based on the second 3D semantic features of each spatial point; Perform spatial point sampling based on the sampling probability of each spatial point under multiple viewing angles; For each viewing angle, based on the second 3D semantic features of each spatial point, the spatial points belonging to the target object are screened from the sampled spatial points, and the screened spatial points are rendered to obtain an unobstructed mask of the target object at the viewing angle.

7. The method according to claim 3, in, Determine the consistency of the first semantic features of each spatial point under different viewing angles, including: Based on the 2D semantic features of each spatial point under multiple perspectives and the mean of the 2D semantic features, the first attention network is used to determine the consistency of the first semantic features of each spatial point under different perspectives.

8. The method according to claim 3, in, Determine the consistency of the second semantic feature of each spatial point with the target object, including: Based on the 2D semantic features of each spatial point under multiple perspectives and the 2D semantic features of the spatial points of the target object under multiple perspectives, the consistency of the second semantic features of each spatial point and the target object is determined.

9. The method according to claim 3, in, Determining a first 3D semantic feature of each spatial point based on the first semantic feature consistency and / or the second semantic feature consistency includes: Determine a first weight and a second weight based on the first semantic feature consistency and the second semantic feature consistency, respectively; Based on the 2D semantic features of each spatial point under multiple perspectives, the first weight and the second weight, a second attention network is used to determine the first 3D semantic features of each spatial point.

10. The method of claim 1, further comprising: include: Determine the image features of the spatial point of the target object under unobstructed viewing angles and / or the viewing angle distances between each viewing angle; Predicting the color information of each spatial point under multiple viewing angles based on the image features of the spatial point of the target object under unobstructed viewing angles and / or the viewing angle distances between each viewing angle; Based on the density information of each spatial point, the 3D content reconstruction model is trained, including: Based on the density information of each spatial point and the color information of each spatial point under multiple perspectives, a 3D content reconstruction model is trained.

11. The method according to claim 10, in, Determine the image features of the spatial points of the target object under unobstructed viewing angles, including: Based on the occlusion mask and the unocclusion mask of the target object under multiple perspectives, the overlapping information of the spatial points of the target object under multiple perspectives is determined; Based on the overlapping information of the spatial points of the target object under multiple viewing angles and the image features of the multi-view images, the image features of the spatial points of the target object under unobstructed viewing angles are determined; The overlapping information indicates the probability that the spatial point of the target object and the spatial points of other objects are at the same pixel position in the two-dimensional image.

12. The method according to claim 10, in, Predict the color information of each spatial point under multiple perspectives, including: Determine the 3D color feature of each spatial point based on the image feature of the spatial point of the target object under an unobstructed viewing angle and / or the viewing angle distance between each viewing angle; Based on the 3D color features of each spatial point, the color information of each spatial point under multiple viewing angles is determined.

13. The method according to claim 12, in, Determine the 3D color characteristics of each spatial point, including: Using the viewing angle distance between each viewing angle, weighting the image features of the spatial point of the target object under the unobstructed viewing angle, to obtain a first weighted image feature; Determining a third weight based on an image feature of a spatial point of the target object under an unobstructed viewing angle and the first weighted image feature; Based on the perspective distance between each perspective, a fourth weight is obtained; Based on the image features of the spatial points of the target object under an unobstructed viewing angle, the third weight and the fourth weight, a third attention network is used to determine the 3D color features of each spatial point.

14. The method of claim 1, in, Semantic features include: object category features and / or object instance features.

15. An electronic device, It is characterized in that include: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to perform the method of any one of claims 1 to 14.

16. A computer-readable storage medium storing instructions, It is characterized in that When the instructions are executed by a processor, the method according to any one of claims 1 to 14 is implemented.