Scene completion method and device, electronic equipment and storage medium
By obtaining semantic segmentation and feature fusion of 3D point clouds and 2D images, and utilizing a 3D GAN generator and building specification data, the problems of low efficiency and poor accuracy in traditional scene completion methods are solved, and a structurally complete and semantically accurate 3D scene model is generated.
Patent Information
- Application Number
- CN202511156718.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Traditional scene completion methods rely on manual calibration, which is inefficient and has poor completion accuracy. The reconstructed models often have problems such as structural misalignment.
By acquiring three-dimensional point clouds and multiple two-dimensional images, semantic segmentation and coordinate alignment are performed, geometric, texture, and semantic features are extracted, and fused to generate semantic point clouds. A three-dimensional GAN generator is used to complete missing areas of the scene, and building specification data is combined to ensure the rationality of the generated scene structure and semantic logic.
It can quickly and accurately complete the missing content of the scene and generate a structurally complete and semantically accurate three-dimensional scene model, overcoming the problems of low efficiency and poor accuracy in traditional methods.
Smart Images

Figure CN120726243A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a scene completion method, device, electronic device, and storage medium. Background Art
[0002] A point cloud is a collection of 3D data consisting of millions of randomly scattered spatial points. Each point in a point cloud typically contains its position coordinates and may also include attribute data such as color and normals, typically derived from sampling by laser scanners or depth cameras. The core goal of point cloud reconstruction is to recover a complete and continuous 3D surface model of an object or scene from this incomplete raw point cloud data, thereby achieving scene completion.
[0003] Traditional scene completion methods rely on manual calibration, which is inefficient and has poor completion accuracy. The reconstructed models often have problems such as structural misalignment. Summary of the Invention
[0004] The embodiments of the present application provide a scene completion method, device, electronic device and storage medium, which can quickly and accurately complete the missing content in the scene to be modeled, thereby improving the efficiency of scene completion.
[0005] The present invention provides a method for completing a scene, including: Obtain a 3D point cloud and multiple 2D images of the scene to be modeled. The 3D point cloud is composed of multiple points, and the 2D image is a picture obtained by photographing the scene to be modeled from any perspective. Perform semantic segmentation on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label for each pixel; Align the coordinates of the semantic image with each point in the 3D point cloud to obtain a semantic point cloud, which is a 3D point cloud with a semantic label associated with each point. Obtain fusion features of 3D point cloud, 2D image, and semantic point cloud; Determine the missing areas in the scene to be modeled based on the semantic point cloud; Based on the fusion features, the missing areas of the scene to be modeled are completed to obtain a complete scene.
[0006] The embodiment of the present application further provides a scene completion device, comprising: An acquisition unit is used to acquire a three-dimensional point cloud and multiple two-dimensional images of the scene to be modeled, where the three-dimensional point cloud is composed of multiple points, and the two-dimensional image is a picture obtained by shooting the scene to be modeled from any perspective; The semantic unit is used to perform semantic segmentation on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label for each pixel; An alignment unit is used to align the coordinates of the semantic image with each point in the three-dimensional point cloud to obtain a semantic point cloud, which is a three-dimensional point cloud with a semantic label associated with each point; Fusion unit, used to obtain fusion features of 3D point cloud, 2D image and semantic point cloud; A missing unit is used to determine the missing area of the scene to be modeled based on the semantic point cloud; The completion unit is used to complete the missing areas of the scene to be modeled based on the fusion features to obtain a complete scene.
[0007] In some embodiments, obtaining fusion features of a three-dimensional point cloud, a two-dimensional image, and a semantic point cloud includes: Extract geometric features of 3D point clouds; Extract texture features of two-dimensional images; Extract semantic spatial relationship features of semantic point clouds; The geometric features of the 3D point cloud, the texture features of the 2D image, and the semantic space relationship graph are fused to obtain the fused features; In some embodiments, extracting geometric features of a three-dimensional point cloud includes: Calculate the neighborhood normal vector angle of each point in the 3D point cloud to obtain the curvature distribution map of the 3D point cloud; According to the curvature distribution map, high curvature areas and low curvature areas are determined in the three-dimensional point cloud; Extract features from high curvature areas to obtain dense point features; Feature extraction is performed on low curvature areas to obtain sparse point features; The dense point features and sparse point features are fused to obtain the geometric feature vector.
[0008] In some embodiments, feature extraction is performed on high curvature areas to obtain dense point features, including: Determine the FPFH descriptor of each point in the high curvature area and obtain the first feature; The first feature is input into a feature extraction network to obtain a dense point feature, which represents the geometric features of the first feature at multiple scales.
[0009] In some embodiments, feature extraction is performed on the low curvature region to obtain sparse point features, including: Determine the FPFH descriptor of each point in the low curvature area to obtain the second feature; A three-layer MLP network is used to extract sparse point features, which represent the global contour features of the second feature.
[0010] In some embodiments, extracting semantic spatial relationship features of a semantic point cloud includes: Construct a spatial relationship graph of the semantic point cloud. The spatial relationship graph includes nodes and edges. Nodes represent objects in the scene to be modeled, and edges represent the spatial relationships between objects. A graph convolutional network is used to extract the semantic spatial relationship features of the spatial relationship graph.
[0011] In some embodiments, based on the fusion features, the missing areas of the scene to be modeled are completed to obtain a complete scene, including: Determine geometric constraints based on preset building code data; Generate a binary semantic mask of the scene missing area of the scene to be modeled; Input the fused features, geometric constraints, and binary semantic masks into the 3D GAN generator and output the completed point cloud; Obtain the texture map corresponding to the missing area of the scene in the two-dimensional image; Combine the 3D point cloud, the completed point cloud and the texture map to get the complete scene.
[0012] An embodiment of the present application also provides an electronic device, including a memory storing multiple instructions; the processor loads instructions from the memory to execute the steps of any scene completion method provided in the embodiment of the present application.
[0013] An embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the steps in any scene completion method provided in the embodiment of the present application.
[0014] The embodiments of the present application can obtain a three-dimensional point cloud and multiple two-dimensional images of the scene to be modeled, where the three-dimensional point cloud is composed of multiple points, and the two-dimensional image is a picture obtained by shooting the scene to be modeled from any perspective; perform semantic segmentation on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label marked for each pixel; perform coordinate alignment processing on the semantic image and each point in the three-dimensional point cloud to obtain a semantic point cloud, which is a three-dimensional point cloud with a semantic label associated with each point; obtain fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud; based on the semantic point cloud, determine the scene missing area in the scene to be modeled; based on the fusion features, complete the scene missing area of the scene to be modeled to obtain a complete scene.
[0015] This application uses semantic segmentation to label two-dimensional images at the pixel level and aligns them with the precise position of the point cloud, seamlessly migrating the rich semantic information in the image to the three-dimensional point cloud to construct a semantic point cloud. This allows the semantic point cloud to be used to distinguish different objects in the point cloud, overcoming the lack of semantic information in traditional point clouds. Then, the features of the three-dimensional point cloud, two-dimensional image, and semantic point cloud are fused to form a complementary joint feature expression: the geometric information of the three-dimensional point cloud ensures spatial accuracy, the texture of the two-dimensional image supplements surface details, and the semantic labels of the semantic point cloud give the scene structured understanding ability, together forming a more robust scene representation basis. Finally, the semantic point cloud can accurately locate areas with semantic incoherence or geometric breaks, such as occluded object parts. Compared with the traditional method that relies solely on geometric hole detection, the semantic consistency proposed in this solution becomes a more reliable basis for identifying missing areas. Among them, the completion process uses fused features including geometry, texture, and semantics to ensure that the generated scene structure is both geometrically reasonable and semantically logical.
[0016] This application can generate a structurally complete and semantically accurate 3D scene model under the challenge of complex occlusion, and semantically guided completion significantly reduces unreasonable completion content. As a result, the embodiments of this application can quickly and accurately complete the missing content in the scene to be modeled. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 Schematic diagram of the process of the scene completion method provided in the embodiment of the present application; Figure 2 is a structural diagram of a scene completion device provided in an embodiment of the present application; Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0020] Embodiments of the present application provide a scene completion method, device, electronic device, and storage medium.
[0021] The scene completion device can be integrated into an electronic device, such as a terminal or a server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC); the server can be a single server or a server cluster consisting of multiple servers.
[0022] In some embodiments, the scene completion device can also be integrated into multiple electronic devices. For example, the scene completion device can be integrated into multiple servers, and the scene completion method of the present application can be implemented by multiple servers.
[0023] In some embodiments, the terminal may also function as a server to implement some or all of the server's functions.
[0024] For example, the electronic device can be a terminal, which can obtain a three-dimensional point cloud and multiple two-dimensional images of the scene to be modeled, where the three-dimensional point cloud is composed of multiple points, and the two-dimensional image is a picture obtained by shooting the scene to be modeled from any perspective; semantic segmentation is performed on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label for each pixel; coordinate alignment is performed on the semantic image with each point in the three-dimensional point cloud to obtain a semantic point cloud, which is a three-dimensional point cloud with a semantic label associated with each point; fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud are obtained; based on the semantic point cloud, the scene missing area in the scene to be modeled is determined; based on the fusion features, the scene missing area of the scene to be modeled is completed to obtain a complete scene.
[0025] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0026] In this embodiment, a scene completion method is provided. Figure 1 , the specific process of the scene completion method can be as follows: 110. Obtain a three-dimensional point cloud and multiple two-dimensional images of the scene to be modeled. The three-dimensional point cloud is composed of multiple points, and the two-dimensional image is a picture obtained by shooting the scene to be modeled from any perspective.
[0027] The scene to be modeled refers to the virtual 3D scene to be completed. This virtual 3D scene may contain one or more missing areas. A virtual 3D scene is a virtual three-dimensional spatial environment completely simulated using digital information in a computer. This virtual 3D scene can be composed of one or more virtual objects. For example, a house scene can be composed of doors, windows, walls, roofs, etc.
[0028] Both virtual objects and virtual 3D scenes can contain data such as virtual models and material maps. Material maps define the visual properties and details of the model surface, such as color, texture, smoothness, and reflectivity.
[0029] A 3D point cloud is a collection of discrete points in a 3D space collected by sensors such as lidar. Each point can contain location information, such as coordinate information.
[0030] 120. Semantic segmentation is performed on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label for each pixel.
[0031] Semantic segmentation is a process that classifies each pixel in an image into a predefined object category, such as a wall, window, or roof. In some embodiments, semantic segmentation can be implemented using a neural network model. Specifically, a two-dimensional image is input into the neural network model, and the model outputs pixel-level semantic annotations as a semantic image.
[0032] Each pixel in the semantic image may be annotated with a semantic label. In some embodiments, the label value may correspond to an object category.
[0033] 130. The semantic image is aligned with each point in the three-dimensional point cloud to obtain a semantic point cloud. The semantic point cloud is a three-dimensional point cloud in which a semantic label is associated with each point.
[0034] In some embodiments, a transformation matrix can be established between the coordinate system of the two-dimensional image and the coordinate system of the three-dimensional point cloud, and the transformation matrix can be used to achieve coordinate alignment, thereby projecting the semantic image into the spatial system of the three-dimensional point cloud, and achieving one-to-one alignment of pixels in the semantic image with each point in the three-dimensional point cloud.
[0035] 140. Obtain fusion features of 3D point cloud, 2D image and semantic point cloud.
[0036] In some embodiments, obtaining fusion features of a three-dimensional point cloud, a two-dimensional image, and a semantic point cloud includes: Extract geometric features of 3D point clouds; Extract texture features of two-dimensional images; Extract semantic spatial relationship features of semantic point clouds; The geometric features of the three-dimensional point cloud, the texture features of the two-dimensional image, and the semantic space relationship graph are fused to obtain the fused features.
[0037] Among them, geometric features are feature vectors that characterize the spatial structure of objects; texture features are feature vectors that characterize the visual properties of the object surface; and semantic spatial relationship features are structured feature vectors that characterize the topological relationships between different objects.
[0038] This embodiment fuses geometric features, texture features, and a semantic space relationship graph to obtain cross-modal fusion features, thereby achieving a coupled representation of geometry, texture, and semantics.
[0039] In some embodiments, extracting geometric features of a three-dimensional point cloud includes: Calculate the neighborhood normal vector angle of each point in the 3D point cloud to obtain the curvature distribution map of the 3D point cloud; According to the curvature distribution map, high curvature areas and low curvature areas are determined in the three-dimensional point cloud; Extract features from high curvature areas to obtain dense point features; Feature extraction is performed on low curvature areas to obtain sparse point features; The dense point features and sparse point features are fused to obtain the geometric feature vector.
[0040] Dense point features are local detailed geometric features extracted from high curvature areas. In some embodiments, dense point features can be extracted through a multi-scale convolutional network, which can better represent the microstructure of high curvature areas.
[0041] Sparse point features are global contour features extracted from low curvature areas. In some embodiments, sparse point features can be extracted through a fully connected network, which can better represent the macro structure of low curvature areas.
[0042] Among them, the curvature distribution map K p Characterizes the curvature of the surface at point p, which can be obtained by the angle between the normal vectors of the neighborhood. Therefore, the curvature distribution map K p The calculation formula is:
[0043] Among them, n p is the normal vector of point p, n pi is the normal vector of the i-th neighbor point of point p, k is the total number of neighbor points of point p; a neighbor point refers to a point adjacent to point p, that is, there is an edge between point p and its i-th neighbor point.
[0044] In some embodiments, points with K greater than a preset threshold may be allocated in a high curvature region, and points with K not greater than the preset threshold may be allocated in a low curvature region.
[0045] In some embodiments, voxels in a high curvature region may be set to a first size, and voxels in a low curvature region may be set to a second size, thereby ensuring a balance between performance and accuracy.
[0046] For example, in some embodiments, the preset threshold may be set to 0.3, the first dimension is 0.05 meters (m), and the second dimension is 0.3 meters, thus obtaining:
[0047] That is, there are 8,000 points per cubic meter in the high curvature area and 125 points per cubic meter in the low curvature area.
[0048] In some embodiments, feature extraction is performed on high curvature areas to obtain dense point features, including: Determine the FPFH descriptor of each point in the high curvature area and obtain the first feature; The first feature is input into a feature extraction network to obtain a dense point feature, which represents the geometric features of the first feature at multiple scales.
[0049] In some embodiments, before determining the FPFH descriptor for each point in the high curvature region, the high curvature region may be downsampled to reduce the data size of the high curvature region and improve subsequent processing efficiency. Voxel grid downsampling is a downsampling method that, for example, can reduce the number of points in a region by dividing the point cloud into a voxel grid and averaging the points.
[0050] Among them, the FPFH descriptor (Fast Point Feature Histogram) is a feature descriptor used for 3D point cloud data processing, mainly used to describe the local geometric features of each point in the point cloud; FPFH quantifies the spatial geometric relationships between a point in the point cloud and the points in its neighborhood into a histogram, thereby generating a discriminative feature vector; among them, the spatial geometric relationships can include normal differences, distances, angles, etc.
[0051] Among them, the feature extraction network can use the Multi-Scale Grouping (MSG) module to extract features from high curvature areas. The feature extraction network includes a sampling module, a grouping module, and a feature extraction module, among which: 1. Sampling module: The sampling module can achieve uniform coverage sampling by iteratively selecting the point farthest from the existing center point through the farthest point sampling formula; the farthest point sampling formula is as follows:
[0052] Among them, S is the center point of the point cloud set, and the point cloud set is the point in the high curvature area; S k+1 Refers to the k+1th sampling center point, K is the number of sampling points; ||ps|| is the Euclidean distance between point p and center point S.
[0053] 2. Grouping module: The grouping module can aggregate neighborhood point sets within multiple radius scales, that is, collect neighborhood points within a radius r with the center point as the center of the sphere to construct a local geometric context; the formula is as follows:
[0054] Among them, r is the radius scale, G k is the neighborhood point set of the kth sampling center point.
[0055] Among them, the radius scale r can take multiple values to capture features from micro-texture to macro-contour and achieve coordinated extraction of details and global features.
[0056] 3. Feature extraction module: The feature extraction module extracts local features for each group and outputs multi-scale geometric features through shared MLP and maximum pooling; its formula is as follows: f k =MaxPool(MLP(Concat(p j ,FPFH(p j )))), p j ∈G k f k r1 = MLP(G k r1 ), f k r2 = MLP(G k r2 ) f k out =W res ·f k l-1 +Conv(Concat(f k r1 , f k r2 )) F global =max(f k 3 )∈R 256 in, f k out :For the kThe output feature vector of the center point; f k l-1 is the input feature of the previous layer; fk ( r 1), fk ( r 2) local features at different radius scales; W res is a learnable residual weight matrix used to adjust the influence of input features; Conv is a 1×1 convolutional layer used for multi-scale feature fusion; MaxPool is the maximum pooling layer; Concat is a feature fusion operation.
[0057] In some embodiments, feature extraction is performed on the low curvature region to obtain sparse point features, including: Determine the FPFH descriptor of each point in the downsampled low curvature area to obtain the second feature; A three-layer MLP network is used to extract sparse point features, which represent the global contour features of the second feature.
[0058] In some embodiments, before determining the FPFH descriptor of each point in the low curvature region, the low curvature region may be downsampled to reduce the data size of the low curvature region and improve subsequent processing efficiency.
[0059] A three-layer MLP (Multi-Layer Perceptron) network can be used to extract sparse point features. This MLP network consists of three fully connected layers. The input layer projects the second feature to 128 dimensions; the hidden layer learns nonlinear transformations using the ReLU activation function; and the output layer compresses the hidden layer output to a 64-dimensional feature vector and introduces a residual connection to preserve the original FPFH information.
[0060] In some embodiments, extracting semantic spatial relationship features of a semantic point cloud includes: Construct a spatial relationship graph G = (V, R) of the semantic point cloud. The spatial relationship graph consists of nodes V and edges R. Nodes represent objects in the scene to be modeled, and edges represent the spatial relationships between objects. A graph convolutional network is used to extract the semantic spatial relationship features of the spatial relationship graph.
[0061] In some embodiments, edges may have weights that characterize the strength of spatial relationships. Therefore, the spatial relationship graph G may also include a weight matrix Wr(l) corresponding to the spatial relationships. Therefore, the formula for extracting the semantic spatial relationship features of the spatial relationship graph using a graph convolutional network is as follows:
[0062] Among them, Hi l+1 represents the output feature of node i in the l+1 layer, σ is the activation function, r is a discrete variable, N r (i) is the set of neighbor nodes that have a spatial relationship r with node i; c i,r is the normalization coefficient, and j is the neighbor node of node i.
[0063] In some embodiments, extracting texture features of a two-dimensional image includes: Generate multiple key points of the two-dimensional image and corresponding first descriptors; Calculating confidence scores of matching pairs between the two-dimensional image and at least one two-dimensional image from another perspective based on the key points and the first descriptor; Filtering matching pairs with a confidence level greater than a preset threshold to obtain a first matching set; Identify low-texture areas in a two-dimensional image and generate a second descriptor of the low-texture area; Performing a nearest neighbor ratio test on the first descriptor and the second descriptor, screening complementary matching pairs that meet a ratio threshold, and obtaining a second matching set; The first matching set and the second matching set are fused to obtain texture features containing sparse correspondences.
[0064] Among them, texture features refer to the pixel-level correspondences and their associated descriptors established through key point matching, which are used to transfer texture information in cross-modal alignment; key points are pixel locations with significant geometric characteristics in the image, such as window corners, door frame edges, etc.; confidence score is a quantitative value of matching reliability; low-texture areas are image areas that lack significant gradients, such as solid-color walls.
[0065] In some embodiments, the matching pair confidence M can be calculated based on the graph neural network of the attention mechanism ij The formula is as follows: M ij =Sim(d i A ,d j B )+λ·Attn ij cross Among them, Sim is the cosine similarity; d i A is the first descriptor of the i-th key point in image A; d j B is the first descriptor of the jth key point in image B; λ is the attention weight coefficient; Attn ijcross is the cross-image attention weight of the graph neural network.
[0066] In some embodiments, the nearest neighbor ratio test can be implemented by the following formula:
[0067] Among them, d i is the first descriptor of the pixel to be matched; d j (1) is the second descriptor of the nearest pixel; d j (2) is the second descriptor of the next neighboring pixel.
[0068] In some embodiments, after fusing the geometric features of the 3D point cloud, the texture features of the 2D image, and the semantic space relationship graph to obtain the fused features, a cross-modal attention mechanism and multi-head self-attention can be added. Therefore: The embedding vector can be processed through three layers of Transformer decoder, each layer performing: 1. Multi-head self-attention calculation: The input features are split into 8 heads, and each head calculates the attention weight in a 64-dimensional subspace; 2. Cross-modal attention calculation: Perform attention interaction in both directions from point cloud to image and from image to point cloud, and add 3D relative position offset to point cloud features; 3. Feedforward network transformation: nonlinear mapping is performed through a fully connected layer with a hidden layer dimension of 1024.
[0069] Then: Concatenate point cloud and image features to generate a 64-channel mixed feature map; Dynamically adjust the contribution of each modality through learnable weights; Based on the classification confidence weighted voting learned from the validation set, unified fusion features are output.
[0070] The calculation formula for the relative position offset is:
[0071] Where Δp ij W is the relative coordinate offset of the point cloud, such as the physical distance between the window frame corner and the window texture; Q p With W K p is the position projection weight, Q i is the point cloud feature query vector, K j is the image feature key vector, d k is the scaling factor.
[0072] The weight formula is: α=Softmax(W•[Fp||Fi||Fs]+b); Among them, Fp is the point cloud geometric feature; Fi is the image texture feature; Fs is the semantic relationship feature; W is the learnable weight matrix; b is the bias term; α is the modal weight.
[0073] 150. Based on the semantic point cloud, determine the missing area in the scene to be modeled.
[0074] For example, a semantic segmentation module used in a convolutional neural network can be used to classify the scene to be modeled, thereby obtaining the scene missing areas in the scene to be modeled.
[0075] In some embodiments, a semantic mask of the missing region of the scene may be generated, and the texture of the corresponding region may be extracted from the image so that the texture map of the missing region of the scene may be considered during completion in subsequent steps.
[0076] 160. Based on the fusion features, the missing areas of the scene to be modeled are completed to obtain a complete scene.
[0077] In some embodiments, based on the fusion features, the missing areas of the scene to be modeled are completed to obtain a complete scene, including: Determine geometric constraints based on preset building code data; Generate a binary semantic mask of the scene missing area of the scene to be modeled; Input the fused features, geometric constraints, and binary semantic masks into the 3D GAN generator and output the completed point cloud; Obtain the texture map corresponding to the missing area of the scene in the two-dimensional image; Combine the 3D point cloud, the completed point cloud and the texture map to get the complete scene.
[0078] Among them, preset building specification data refers to a set of structured rules stored in an external building specification database, such as the height of a Gothic arch should be 1.5 times the span; this data is based on standards and constraints in the architectural field and is accessed and transmitted through the extended MCP protocol.
[0079] In some embodiments, a large language model can be used as an intelligent intermediary to parse and extract relevant information from user instructions to form a knowledge database, thereby providing these building specification data as an input source.
[0080] Geometric constraints refer to specific geometric restriction parameters derived by the system based on preset building specification data, such as component size ratios, location requirements, and so on. In some embodiments, the specification data can be encoded into an operational constraint template through a standardized MCP protocol interface, associated with semantic tags, and ultimately encoded into a 128-dimensional condition vector to ensure that the generated geometric shape strictly adheres to architectural regulations.
[0081] Among them, the binary semantic mask is a binary image representation based on semantic segmentation technology.
[0082] As can be seen from the above, the embodiments of the present application can obtain a three-dimensional point cloud and multiple two-dimensional images of the scene to be modeled, where the three-dimensional point cloud is composed of multiple points, and the two-dimensional image is a picture obtained by photographing the scene to be modeled from any perspective; perform semantic segmentation on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label for each pixel; perform coordinate alignment processing on the semantic image and each point in the three-dimensional point cloud to obtain a semantic point cloud, which is a three-dimensional point cloud with a semantic label associated with each point; obtain the fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud; based on the semantic point cloud, determine the scene missing area in the scene to be modeled; based on the fusion features, complete the scene missing area of the scene to be modeled to obtain a complete scene. Therefore, this solution can quickly and accurately complete the scene missing content in the scene to be modeled.
[0083] To better implement the above method, the present application also provides a scene completion device. The scene completion device can be integrated into an electronic device, such as a terminal or a server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc. The server can be a single server or a server cluster consisting of multiple servers.
[0084] For example, in this embodiment, the method of the embodiment of the present application will be described in detail by taking the specific integration of the scene completion device into the terminal as an example.
[0085] For example, Figure 2 As shown, the scene completion device may include an acquisition unit 301, a semantic unit 302, an alignment unit 303, a fusion unit 304, a missing unit 305, and a completion unit 306, as follows: (1) Acquisition unit 301, which is used to acquire a 3D point cloud and multiple 2D images of the scene to be modeled. The 3D point cloud is composed of multiple points, and the 2D image is a picture obtained by shooting the scene to be modeled from any perspective.
[0086] (2) Semantic unit 302.
[0087] The semantic unit 302 is used to perform semantic segmentation on the two-dimensional image to obtain a semantic image. The semantic image is a two-dimensional image in which each pixel is annotated with a semantic label.
[0088] (3) Alignment unit 303.
[0089] The alignment unit 303 is used to perform coordinate alignment processing on the semantic image and each point in the three-dimensional point cloud to obtain a semantic point cloud, which is a three-dimensional point cloud in which a semantic label is associated with each point.
[0090] (4) Fusion unit 304.
[0091] The fusion unit 304 is used to obtain fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud.
[0092] (v) Missing unit 305.
[0093] The missing unit 305 is used to determine the scene missing area in the scene to be modeled based on the semantic point cloud.
[0094] (6) Complete unit 306.
[0095] The completion unit 306 is used to complete the missing areas of the scene to be modeled based on the fusion features to obtain a complete scene.
[0096] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0097] In the embodiments of the present application, the term "module" or "unit" refers to a computer program product or a portion of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal. It can be implemented in whole or in part using software, hardware (such as processing circuits or memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the functionality of the module or unit.
[0098] As can be seen from the above, the scene completion device of this embodiment obtains a three-dimensional point cloud and multiple two-dimensional images of the scene to be modeled by an acquisition unit, where the three-dimensional point cloud is composed of multiple points, and the two-dimensional image is a picture obtained by shooting the scene to be modeled from any perspective; the semantic unit performs semantic segmentation on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label annotated for each pixel; the alignment unit performs coordinate alignment processing on the semantic image and each point in the three-dimensional point cloud to obtain a semantic point cloud, which is a three-dimensional point cloud with a semantic label associated with each point; the fusion unit obtains the fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud; the missing unit determines the scene missing area in the scene to be modeled based on the semantic point cloud; the completion unit completes the scene missing area of the scene to be modeled based on the fusion features to obtain a complete scene.
[0099] Therefore, the embodiments of the present application can improve the speed and accuracy of completing the missing content in the scene to be modeled.
[0100] The present application also provides an electronic device, which may be a terminal, a server, or the like. The terminal may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or the like; the server may be a single server or a server cluster consisting of multiple servers, or the like.
[0101] In some embodiments, the scene completion device can also be integrated into multiple electronic devices. For example, the scene completion device can be integrated into multiple servers, and the scene completion method of the present application can be implemented by multiple servers.
[0102] In this embodiment, the electronic device of this embodiment is a terminal as an example for detailed description, for example, Figure 3 , which shows a schematic structural diagram of an electronic device involved in an embodiment of the present application, specifically: The electronic device may include one or more processing core processors 410, one or more computer-readable storage media memories 420, a power supply 430, an input module 440, and a communication module 450. It will be understood by those skilled in the art that Figure 3 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently. Processor 410 is the control center of the electronic device, connecting all parts of the electronic device using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 420 and accessing data stored in memory 420, it performs various functions of the electronic device and processes data, thereby performing overall testing of the electronic device. In some embodiments, processor 410 may include one or more processing cores. In some embodiments, processor 410 may integrate an application processor and a modem processor, wherein the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 410.
[0103] Memory 420 can be used to store software programs and modules. Processor 410 executes various functional applications and data processing by running the software programs and modules stored in memory 420. Memory 420 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as sound playback or image playback); the data storage area may store data generated based on the use of the electronic device. Memory 420 may also include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 420 may also include a memory controller to provide processor 410 with access to memory 420.
[0104] The electronic device also includes a power supply 430 for supplying power to various components. In some embodiments, the power supply 430 can be logically connected to the processor 410 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 430 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0105] The electronic device may further include an input module 440 , which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0106] The electronic device may further include a communication module 450. In some embodiments, the communication module 450 may include a wireless module. The electronic device may perform short-range wireless transmission via the wireless module of the communication module 450, thereby providing the user with wireless broadband Internet access. For example, the communication module 450 may be used to help the user send and receive emails, browse web pages, and access streaming media.
[0107] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 410 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 420 according to the following instructions, and the processor 410 will run the application programs stored in the memory 420 to implement various functions as follows: Obtain a 3D point cloud and multiple 2D images of the scene to be modeled. The 3D point cloud is composed of multiple points, and the 2D image is a picture obtained by photographing the scene to be modeled from any perspective. Perform semantic segmentation on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label for each pixel; Align the coordinates of the semantic image with each point in the 3D point cloud to obtain a semantic point cloud, which is a 3D point cloud with a semantic label associated with each point. Obtain fusion features of 3D point cloud, 2D image, and semantic point cloud; Determine the missing areas in the scene to be modeled based on the semantic point cloud; Based on the fusion features, the missing areas of the scene to be modeled are completed to obtain a complete scene.
[0108] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0109] From the above, it can be seen that the embodiments of the present application can quickly and accurately complete the missing content in the scene to be modeled.
[0110] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0111] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the scene completion methods provided in the embodiments of the present application. For example, the instructions can execute the following steps: Obtain a 3D point cloud and multiple 2D images of the scene to be modeled. The 3D point cloud is composed of multiple points, and the 2D image is a picture obtained by photographing the scene to be modeled from any perspective. Perform semantic segmentation on the two-dimensional image to obtain a semantic image, which is a two-dimensional image with a semantic label for each pixel; Align the coordinates of the semantic image with each point in the 3D point cloud to obtain a semantic point cloud, which is a 3D point cloud with a semantic label associated with each point. Obtain fusion features of 3D point cloud, 2D image, and semantic point cloud; Determine the missing areas in the scene to be modeled based on the semantic point cloud; Based on the fusion features, the missing areas of the scene to be modeled are completed to obtain a complete scene.
[0112] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0113] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in various optional implementations of the scene reconstruction or scene completion aspects provided in the above-mentioned embodiments.
[0114] Since the instructions stored in the storage medium can execute the steps in any scene completion method provided in the embodiments of the present application, the beneficial effects that can be achieved by any scene completion method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0115] The above is a detailed introduction to a scene completion method, device, electronic device and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A scene completion method, characterized in that: include: Acquire a three-dimensional point cloud and multiple two-dimensional images of the scene to be modeled, wherein the three-dimensional point cloud is composed of multiple points, and the two-dimensional images are pictures obtained by photographing the scene to be modeled from any perspective; Performing semantic segmentation on the two-dimensional image to obtain a semantic image, wherein the semantic image is the two-dimensional image with a semantic label for each pixel; Performing coordinate alignment processing on the semantic image and each point in the three-dimensional point cloud to obtain a semantic point cloud, wherein the semantic point cloud is the three-dimensional point cloud in which each point is associated with a semantic label; Acquire fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud; Determining a scene missing area in the scene to be modeled based on the semantic point cloud; Based on the fusion features, the missing areas of the scene to be modeled are completed to obtain a complete scene.
2. The scene completion method according to claim 1, wherein: The obtaining of fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud includes: Extracting geometric features of the three-dimensional point cloud; extracting texture features of the two-dimensional image; Extracting semantic spatial relationship features of the semantic point cloud; The geometric features of the three-dimensional point cloud, the texture features of the two-dimensional image, and the semantic space relationship graph are fused to obtain fused features.
3. The scene completion method according to claim 2, wherein: The extracting geometric features of the three-dimensional point cloud includes: Calculating the neighborhood normal vector angle of each point in the three-dimensional point cloud to obtain a curvature distribution map of the three-dimensional point cloud; Determining high curvature areas and low curvature areas in the three-dimensional point cloud according to the curvature distribution map; Performing feature extraction on the high curvature area to obtain dense point features; Performing feature extraction on the low curvature area to obtain sparse point features; Perform feature fusion processing on the dense point features and the sparse point features to obtain a geometric feature vector.
4. The scene completion method according to claim 3, wherein: The extracting features of the high curvature area to obtain dense point features includes: Determine the FPFH descriptor of each point in the high curvature area to obtain a first feature; The first feature is input into a feature extraction network to obtain a dense point feature, where the dense point feature represents geometric features of the first feature at multiple scales.
5. The scene completion method according to claim 3, wherein: The extracting features of the low curvature area to obtain sparse point features includes: Determine the FPFH descriptor of each point in the low curvature area to obtain a second feature; A three-layer MLP network is used to extract sparse point features, and the sparse point features represent the global contour features of the second feature.
6. The scene completion method according to claim 2, wherein: The extracting of semantic spatial relationship features of the semantic point cloud includes: Constructing a spatial relationship graph of the semantic point cloud, the spatial relationship graph including nodes and edges, the nodes representing objects in the scene to be modeled, and the edges representing spatial relationships between the objects; A graph convolutional network is used to extract the semantic spatial relationship features of the spatial relationship graph.
7. The scene completion method according to claim 1, wherein: The method of completing the missing regions of the scene to be modeled based on the fusion features to obtain a complete scene includes: Determine geometric constraints based on preset building code data; Generating a binary semantic mask of a scene missing area of the scene to be modeled; Inputting the fused features, the geometric constraints, and the binary semantic mask into a 3D GAN generator, and outputting a completed point cloud; Obtaining a texture map corresponding to the missing area of the scene in the two-dimensional image; The three-dimensional point cloud, the completed point cloud, and the texture map are combined to obtain a complete scene.
8. A scene completion device, characterized in that: include: An acquisition unit, configured to acquire a three-dimensional point cloud and a plurality of two-dimensional images of a scene to be modeled, wherein the three-dimensional point cloud is composed of a plurality of points, and the two-dimensional images are pictures obtained by photographing the scene to be modeled from any perspective; A semantic unit, configured to perform semantic segmentation on the two-dimensional image to obtain a semantic image, wherein the semantic image is the two-dimensional image with a semantic label for each pixel; an alignment unit, configured to perform coordinate alignment processing on the semantic image and each point in the three-dimensional point cloud to obtain a semantic point cloud, wherein the semantic point cloud is the three-dimensional point cloud in which a semantic label is associated with each point; a fusion unit, configured to obtain fusion features of the three-dimensional point cloud, the two-dimensional image, and the semantic point cloud; A missing unit, configured to determine a scene missing area in the scene to be modeled based on the semantic point cloud; The completion unit is used to complete the missing scene area of the scene to be modeled based on the fusion features to obtain a complete scene.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the scene completion method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, which are suitable for loading by a processor to execute the steps in the scene completion method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Large-scene airborne point cloud semantic modeling method
CN110120097A
Semantic scene completion method and system, vehicle and computer readable storage medium
CN117854069A
Point cloud semantic scene completion system
CN120070902A
Real scene three-dimensional modeling method and system fusing laser point cloud and image
CN120147563A
Cited By
Three-dimensional point cloud completion method and device based on two-dimensional information guidance and readable medium
CN121353551A
Point cloud completion methods, devices, media and equipment for underground parking lot scenes
CN122415639A
Underground parking lot scene point cloud completion method and device, medium and equipment
CN122415639B