Building model reconstruction method, device, equipment and medium based on diffusion model
Through the building model reconstruction method based on the diffusion model, the feature extraction, diffusion and edge denoising modules are used to solve the problem of redundant information in the existing technology, and achieve efficient building reconstruction and improvement of edge detail accuracy.
Patent Information
- Application Number
- CN202510586852.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Existing building reconstruction methods generate unnecessary redundant information when representing regular building structures, increasing the computational and storage burden.
A building model reconstruction method based on the diffusion model is adopted. The encoding features and point-level attention weights of point cloud data are obtained through the feature extraction module. The noise vector randomly sampled from the Gaussian noise distribution by the diffusion module is used, and edge denoising is performed through the edge denoising module to reconstruct the building model.
It achieves the reconstruction of buildings with fewer facets, reduces the computational and storage burden, and improves the reconstruction accuracy of edge details by combining encoding features and point-level attention weights.
Smart Images

Figure CN120107493B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a method, device, equipment and medium for reconstructing a building model based on a diffusion model. Background Art
[0002] Building reconstruction is a core research topic at the intersection of computer vision, photogrammetry, and computer graphics. Its goal is to recover the geometric representation of 3D building structures from raw sensor data (such as point clouds and images). With the rapid advancement of urbanization, the demand for 3D building models continues to rise. Its applications span a wide range of fields, including smart city planning, autonomous driving navigation, augmented and virtual reality, disaster simulation and assessment, and cultural heritage preservation.
[0003] Existing building reconstruction methods generally use triangular or polygonal meshes to represent building surfaces. Although this representation can accurately capture details and support texture mapping, it will generate unnecessary redundant information when representing regular building structures, thereby increasing the computational and storage burden.
[0004] Therefore existing technology still needs to be improved and improved. Summary of the Invention
[0005] The technical problem to be solved by the present application is to provide a method, device, equipment and medium for reconstructing a building model based on a diffusion model in response to the deficiencies of the existing technology.
[0006] In order to solve the above technical problems, the first aspect of the present application provides a building model reconstruction method based on a diffusion model, which uses a building model reconstruction model. The building model reconstruction model includes a feature extraction module and a diffusion model, and the diffusion model includes a diffusion module and an edge denoising module. The building model reconstruction method based on the diffusion model specifically includes:
[0007] Obtain the encoding features and point-level attention weights of point cloud data through the feature extraction module;
[0008] Gaussian noise randomly sampled from the Gaussian noise distribution through the diffusion module to obtain a noise vector;
[0009] The noise vector is edge denoised according to the encoding features and the point-level attention weights by an edge denoising module to obtain a denoised wireframe, and a building model is reconstructed based on the denoised wireframe.
[0010] In the building model reconstruction method based on the diffusion model, the feature extraction module includes an encoding unit, an embedding learning unit, and a point edge attention generation unit; the step of obtaining the encoding features and point-level attention weights of the point cloud data through the feature extraction module specifically includes:
[0011] Extracting point embedding features of the point cloud data by the embedding learning unit;
[0012] Extracting encoding features of the point cloud data according to the point embedding features by the encoding unit;
[0013] The point-level attention weights of the point cloud data are generated according to the point embedding features by the point edge attention generation unit.
[0014] The building model reconstruction method based on the diffusion model, wherein edge denoising of the noise vector according to the encoding feature and the point-level attention weight specifically includes:
[0015] generating an intermediate noise vector based on the noise vector using a self-attention mechanism;
[0016] A denoised wireframe is determined according to the encoded features, the point-level attention weights and the intermediate noise vector using a cross-attention mechanism.
[0017] In the building model reconstruction method based on the diffusion model, the step of determining the denoised wireframe based on the encoding features, the point-level attention weights, and the intermediate noise vector using the cross-attention mechanism specifically includes:
[0018] Fusing the encoding features and the point-level attention weights to form a point-level attention feature map;
[0019] Constructing a query vector based on the intermediate noise vector, and constructing a value vector and a key vector based on the point-level attention feature map;
[0020] A denoised wireframe is determined based on the query vector, the value vector, and the key vector using a cross-attention mechanism.
[0021] The method for reconstructing a building model based on a diffusion model, wherein the step of reconstructing the building model based on the denoised wireframe specifically includes:
[0022] The denoised wireframe is post-processed using a non-maximum suppression method and a clustering method, and a building model is reconstructed based on the post-processed denoised wireframe.
[0023] In the building model reconstruction method based on the diffusion model, during the training process of the building model reconstruction model, the noise vector acquisition process specifically includes:
[0024] Obtaining annotation structural parameters of each annotation wireframe of the training point cloud data, wherein the annotation structural parameters include a vector representation of each edge in the annotation wireframe, and the vector representation of the edge includes a midpoint coordinate and an offset;
[0025] The annotation structured parameters of the annotation wireframe are diffused into Gaussian noise to obtain a noise vector.
[0026] The building model reconstruction method based on the diffusion model, wherein the loss function used in the training process of the building model reconstruction model includes structured parameter prediction loss and attention weight loss, wherein the generation process of the point-level attention weight label used by the attention weight loss specifically includes:
[0027] Get the vertex set and edge set of the annotated wireframe corresponding to the training point cloud data;
[0028] For each training point in the training point cloud data, searching for the nearest edge and nearest point corresponding to the training point in the vertex set and edge set, and calculating a first distance from the training point to the nearest vertex and a second distance to the nearest edge;
[0029] If the first distance is less than or equal to the second distance, use the first distance as the target distance corresponding to the training point, and use the default projection factor as the projection factor corresponding to the training point;
[0030] If the first distance is greater than the second distance, the second distance is used as the target distance corresponding to the training point, and the ratio of the distance from the projection position of the training point on the nearest edge to the edge endpoint is used as the projection factor;
[0031] Construct the point-level attention weight labels corresponding to the training point cloud data based on the target distance and projection factor corresponding to all training points.
[0032] A second aspect of the present application provides a building model reconstruction device based on a diffusion model, wherein the building model reconstruction device based on the diffusion model specifically includes:
[0033] Feature extraction module, used to obtain the encoding features and point-level attention weights of point cloud data;
[0034] a diffusion module for randomly sampling Gaussian noise from a Gaussian noise distribution to obtain a noise vector;
[0035] An edge denoising module is used to perform edge denoising on the noise vector according to the encoding feature and the point-level attention weight to obtain a denoised wireframe, and reconstruct a building model based on the denoised wireframe.
[0036] A third aspect of the present application provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in any of the above-described diffusion model-based building model reconstruction methods.
[0037] A fourth aspect of the present application provides a terminal device, comprising: a processor and a memory;
[0038] The memory stores a computer-readable program executable by the processor;
[0039] When the processor executes the computer-readable program, the processor implements the steps in any of the above-mentioned methods for reconstructing a building model based on a diffusion model.
[0040] Beneficial effects: Compared with the prior art, the present application provides a method, device, equipment and medium for reconstructing a building model based on a diffusion model, the method comprising obtaining the coding features and point-level attention weights of point cloud data through a feature extraction module; randomly sampling Gaussian noise from a Gaussian noise distribution through a diffusion module to obtain a noise vector; performing edge denoising on the noise vector according to the coding features and the point-level attention weights through an edge denoising module to obtain a denoised wireframe, and reconstructing the building model based on the denoised wireframe. The present application uses wireframe technology to represent the surface of the building, thereby achieving reconstruction of the building with a smaller number of facets, thereby reducing the burden of calculation and storage. In addition, the present application combines coding features and point-level attention weights, and utilizes a diffusion model to denoise the noise vector to obtain the structural parameters of the wireframe. This not only enables progressive refinement of the building wireframe, but also extracts geometric information from the point-level attention weights, further improving the reconstruction accuracy of edge details. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 A diagram of an application environment for the building model reconstruction method based on the diffusion model provided in an embodiment of the present application.
[0043] Figure 2 This is a flowchart of a building model reconstruction method based on a diffusion model provided in an embodiment of the present application.
[0044] Figure 3Schematic flow chart of the training process for rebuilding the model for a building.
[0045] Figure 4 This is a principle block diagram of a building model reconstruction device based on a diffusion model provided in an embodiment of the present application.
[0046] Figure 5 This is a block diagram of the principles of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] The present application provides a method, apparatus, device, and medium for reconstructing a building model based on a diffusion model. To clarify the purpose, technical solution, and effects of this application, the present application is further described below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate this application and are not intended to limit this application.
[0048] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0049] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0050] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0051] Research has shown that building reconstruction is a core research topic at the intersection of computer vision, photogrammetry, and computer graphics. Its goal is to recover the geometric representation of 3D building structures from raw sensor data (such as point clouds and images). With the rapid advancement of urbanization, the demand for 3D building models continues to rise, with applications spanning a wide range of fields, including smart city planning, autonomous driving navigation, augmented and virtual reality, disaster simulation and assessment, and cultural heritage preservation.
[0052] Existing building reconstruction methods generally use triangular or polygonal meshes to represent building surfaces. Although this representation can accurately capture details and support texture mapping, it will generate unnecessary redundant information when representing regular building structures, thereby increasing the computational and storage burden.
[0053] In order to solve the above problems, in an embodiment of the present application, the coding features and point-level attention weights of the point cloud data are obtained through a feature extraction module; Gaussian noise is randomly sampled from the Gaussian noise distribution through a diffusion module to obtain a noise vector; the edge denoising module performs edge denoising on the noise vector according to the coding features and the point-level attention weights to obtain a denoised wireframe, and the building model is reconstructed based on the denoised wireframe. The present application uses wireframe technology to represent the surface of the building, and realizes the reconstruction of the building with a smaller number of facets, thereby reducing the burden of calculation and storage. In addition, the present application combines coding features and point-level attention weights, and uses a diffusion model to denoise the noise vector to obtain the structural parameters of the wireframe. This not only enables the progressive refinement of the building wireframe, but also extracts geometric information from the point-level attention weights, further improving the reconstruction accuracy of edge details.
[0054] An application environment diagram of the building model reconstruction method based on the diffusion model provided in the embodiment of the present application can be as follows: Figure 1 As shown. Figure 1, the application scenario includes a user terminal 110 and a server 120. The user terminal 110 and the server 120 are connected via a network. The user terminal 110 can specifically be a desktop user terminal or a mobile user terminal, and the mobile user terminal can specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server 120 or a server 120 cluster composed of multiple servers 120. Among them, the user terminal 110 can send point cloud data to the server 120. The server 120 is deployed with a building model reconstruction model, and after receiving the point cloud data sent by the user terminal 110, it obtains the encoding features and point-level attention weights of the point cloud data through the feature extraction module; randomly samples Gaussian noise from the Gaussian noise distribution through the diffusion module to obtain a noise vector; performs edge denoising on the noise vector according to the encoding features and the point-level attention weights through the edge denoising module to obtain a denoised wireframe, and reconstructs the building model based on the denoised wireframe. The user terminal displays the reconstructed building model.
[0055] Of course, in other application scenarios, only the user terminal 110 may be included, and the building reconstruction process is directly executed in the user terminal 110, that is, the user terminal 110 is deployed with a building model reconstruction model. After obtaining the point cloud data, the encoding features and point-level attention weights of the point cloud data are obtained through the feature extraction module; Gaussian noise is randomly sampled from the Gaussian noise distribution through the diffusion module to obtain a noise vector; the edge denoising module performs edge denoising on the noise vector according to the encoding features and the point-level attention weights to obtain a denoised wireframe, and the building model is reconstructed based on the denoised wireframe.
[0056] Furthermore, the building model reconstruction model includes a feature extraction module and a diffusion model. The diffusion model includes a diffusion module and an edge denoising module. The feature extraction module is connected to the diffusion module and the edge denoising module, respectively. The diffusion module is connected to the edge denoising module. The feature extraction module is used to extract features from point cloud data to obtain encoding features and point-level attention weights. The diffusion module is used to obtain a noise vector. The edge denoising module is used to progressively denoise the noise vector based on the encoding features and point-level attention weights to obtain structural parameters of the wireframe to reconstruct the building model.
[0057] Apply the building model as above to rebuild the model, such as Figure 2 and Figure 3 As shown, the building model reconstruction method based on the diffusion model provided in the embodiment of the present application specifically includes:
[0058] S10. Obtain the encoding features and point-level attention weights of the point cloud data through the feature extraction module.
[0059] Specifically, the point cloud data is building point cloud data, which can be obtained by collecting data on the building through sensing equipment such as laser radar, wherein each point in the point cloud data can include three-dimensional coordinates, RGB values and reflection intensity.
[0060] Encoded features are obtained by encoding point cloud data. They include both local detail features and global structural information within the point cloud data, providing knowledge for building model reconstruction. Point-level attention weights are potential edge features identified from point cloud data. They provide geometric prior knowledge for the denoising process, improving the accuracy of edge detail reconstruction.
[0061] In one implementation, the feature extraction module includes an encoding unit, an embedding learning unit, and an edge attention generation unit; obtaining the encoding features and point-level attention weights of the point cloud data through the feature extraction module specifically includes:
[0062] Extracting point embedding features of the point cloud data by the embedding learning unit;
[0063] Extracting encoding features of the point cloud data according to the point embedding features by the encoding unit;
[0064] The edge attention generation unit generates point-level attention weights for the point cloud data according to the point embedding features.
[0065] Specifically, the embedding learning unit is used to generate point-level features , represents the feature dimension, Indicates the number of points in the point cloud data. The input of the embedding learning unit is the point cloud data, and the output is the point embedding features of the point cloud data. The embedding learning unit may include a multilayer perceptron network and an embedding network, and the multilayer perceptron network is connected to the embedding network. The multilayer perceptron is used to embed each point in the point cloud data. Mapping to a high-dimensional feature space to obtain high-dimensional point features, the embedding network is used to capture the local geometric characteristics in the high-dimensional point features to generate point embedding features. The embedding network can include several cascaded fully connected blocks, each of which includes a fully connected layer, batch normalization and ReLU activation function cascaded in sequence.
[0066] The encoding unit is used to perform global context modeling on the point embedding features to generate encoded features that contain rich local details and global structural information. The encoding unit comprises several cascaded layers of self-attention modules, each of which includes a multi-head self-attention mechanism (e.g., 4 heads) and a feedforward network. The point embedding features pass through each layer of self-attention modules in turn, where global context modeling is performed on the point embedding features to produce encoded features that contain both local details and global structural information. Of course, in practical applications, the encoding unit can also adopt other structures, such as a Transformer encoder.
[0067] The edge attention generation unit is used to predict the confidence that each point in the point cloud data belongs to the edge, so as to identify edge features from the point cloud data. The point-level attention weight is used as a geometric guide in the denoising process to provide auxiliary information for the denoising process. The point-level attention weight is generated by the edge attention generation unit. The points in the point-level attention weight correspond one-to-one to the points in the point cloud data, and each point is used to represent the confidence that the point in the corresponding point cloud data belongs to the edge. Point-level attention weight. As shown in Figure 3 As shown, the edge attention generation unit can include a multi-layer perceptron network (such as a three-layer fully connected lightweight MLP network, etc.) and an activation function layer (such as a sigmoid function, etc.), and the point embedding feature Mapped to 1D space through a multi-layer perceptron network, and then returned to normalization to the (0,1) interval through an activation function layer to obtain point-level attention weights Of course, in practical applications, the edge attention generation unit can also adopt other network structures that can predict the confidence that a point belongs to the edge.
[0068] S20, randomly sampling Gaussian noise from the Gaussian noise distribution through a diffusion module to obtain a noise vector.
[0069] Specifically, the noise vector is a Gaussian noise randomly sampled from a Gaussian noise distribution, which is a structural parameter of the wireframe containing the noise. The wireframe can be represented as an edge set , is the number of edges of the wireframe, and 6 is the vector representation of the edge, which includes the coordinates of the edge midpoint and the offset of the edge endpoint relative to the edge midpoint , accordingly, the vector representation of the edge can be ( , , , , , ). Of course, in practical applications, the vector representation of the edge may also be in other ways. For example, the vector representation of the edge may include the coordinates of one endpoint of the edge and the offset of the other endpoint relative to the endpoint.
[0070] S30. Perform edge denoising on the noise vector according to the encoding feature and the point-level attention weight through an edge denoising module to obtain a denoised wireframe, and reconstruct a building model based on the denoised wireframe.
[0071] Specifically, after obtaining the noise vector, the edge denoising module in the diffusion model performs backward denoising on the noise vector using the encoded features and the point-level attention weights as conditions to obtain a denoised wireframe. The edge denoising module can include several expansion layers, each of which is used to perform backward denoising on the noise vector and output an edge prediction result. The edge prediction result can include edge existence confidence, edge midpoint coordinates, and edge offset.
[0072] Furthermore, denoising the edge of the noise vector according to the encoding feature and the point-level attention weight specifically includes:
[0073] generating an intermediate noise vector based on the noise vector using a self-attention mechanism;
[0074] A denoised wireframe is determined according to the encoded features, the point-level attention weights and the intermediate noise vector using a cross-attention mechanism.
[0075] Specifically, the self-attention mechanism is used to calculate the dependency between edges. When the self-attention mechanism is used to generate an intermediate noise vector according to the noise vector, the noise vector can be Perform linear transformation to generate query ,key Sum ,in, , and Both represent linear transformation weight matrices, and then use the multi-head self-attention mechanism to perform edge dependencies based on queries, keys, and values to obtain the intermediate noise vector .
[0076] Further, after obtaining the intermediate noise vector After that, a cross-attention mechanism can be used to cross-learn the intermediate noise vector, the encoding feature, and the point-level attention weight to determine the denoised wireframe. Accordingly, the cross-attention mechanism is used to determine the denoised wireframe according to the encoding feature, the point-level attention weight, and the intermediate noise edge, specifically including:
[0077] Fusing the encoding features and the point-level attention weights to form a point-level attention feature map;
[0078] Constructing a query vector based on the intermediate noise edge, and constructing a value vector and a key vector based on the point-level attention feature map;
[0079] A denoised wireframe is determined based on the query vector, the value vector, and the key vector using a cross-attention mechanism.
[0080] Specifically, the point-level attention feature map is obtained by fusing the coding features and the point-level attention weights. For example, the coding features and the point-level attention weights can be element-wise multiplied to form a point-level attention feature map, so that the point-level attention feature map can make the building model reconstruction model focus on the point cloud area related to the edge, thereby improving the denoising accuracy.
[0081] Furthermore, when constructing a query vector based on the intermediate noise edge, the intermediate noise edge can be first input into a feedforward neural network, the intermediate noise edge is processed by the feedforward neural network, and a query vector is constructed based on the processed intermediate noise edge, wherein the query vector can be expressed as:
[0082] ,
[0083] in, represents a feedforward neural network, represents the weight coefficient, Indicates the offset, Represents the activation function.
[0084] Furthermore, after obtaining the point-level attention feature map, a linear transformation is performed on the point-level attention feature map to generate a value vector and a key vector, where the value vector and the key vector can be expressed as:
[0085] ,
[0086] ,
[0087] in, represents the key vector, represents a value vector, , represents the linear transformation weight matrix, represents the encoding feature, represents the point-level attention weight, Represents element-wise multiplication.
[0088] In one embodiment, after obtaining the denoised wireframe, the building model can be reconstructed based on the denoised wireframe. The reconstruction of the building model based on the denoised wireframe can adopt an existing reconstruction method based on wireframe representation, which is not specifically limited here. Furthermore, when reconstructing the building model based on the denoised wireframe, the denoised wireframe can be post-processed to ensure the topological rationality and geometric consistency of the wireframe. Based on this, the reconstruction of the building model based on the denoised wireframe specifically includes:
[0089] The denoised wireframe is post-processed using a non-maximum suppression method and a clustering method, and a building model is reconstructed based on the post-processed denoised wireframe.
[0090] Specifically, the non-maximum suppression method is used to remove redundant edges using edge confidence. Specifically, when the denoised wireframe similarity exceeds a preset threshold, only the denoised wireframes with high confidence are retained, reducing the number of similar edges. A clustering method is used to merge the endpoints of the denoised wireframes to avoid dangling edges and structural breaks. This clustering method can employ the DBSCAN clustering method, with a pre-set clustering threshold to extract high-density regions as vertex candidates.
[0091] In one embodiment, the training process for the building model reconstruction model is essentially the same as the inference process for the building model reconstruction model. The only difference between the two lies in the method for obtaining the noise vector and the optimization of the loss function formed during the training process to obtain the trained building model reconstruction model, and the use of the trained building model reconstruction model for inference. In other words, the building model reconstruction model used in this application is the trained building model reconstruction model. The following describes two differences between the training process and the inference process.
[0092] During the training process, the noise vector is formed based on the annotated wireframe of the training point cloud data. Specifically, the process of obtaining the noise vector includes:
[0093] Obtain the annotation structure parameters of each annotation wireframe of the training point cloud data;
[0094] The annotation structured parameters of the annotation wireframe are forward diffused into Gaussian noise to obtain a noise wireframe.
[0095] Specifically, each annotation wireframe of the training point cloud data is represented as an edge set The annotation structure parameters of the annotation wireframe include the vector representation of each edge in the edge set corresponding to the annotation wireframe, and the vector representation includes the midpoint coordinates and offset , the annotation structure parameters of the annotation wireframe can be packaged as dimensional vector, Indicates edge data. However, in practical applications, due to the different number of edges of different buildings, when obtaining the annotation structure parameters of the annotation wireframe, the annotation structure parameters of the annotation wireframe can be filled to a fixed length. Among them, when filling the annotation structure parameters of the annotation wireframe to a fixed length When doing so, you can use methods such as repeatedly filling the edges in the wireframe.
[0096] Furthermore, after obtaining the annotation structured parameters of the annotation wireframe, a scaling factor (e.g., 2.0) can be applied to the annotation structured parameters of the annotation wireframe to enhance the signal strength. The enhanced annotation structured parameters of the annotation wireframe are then diffused into Gaussian noise to obtain a noise vector. Of course, the annotation structured parameters of the annotation wireframe can also be directly diffused into Gaussian noise to obtain a noise vector. The diffusion can be performed using methods such as cosine variance. The noise wireframe can be expressed as:
[0097]
[0098] in, Represents the time step Noise wireframe, represents the original wireframe, represents standard Gaussian noise, represents the cumulative coefficient of variance.
[0099] During the training process, the loss function of the building model reconstruction model includes structured parameter prediction loss and attention weight loss. The structured parameter prediction loss is used to reflect the difference between the predicted edge and the true edge, and the attention weight loss is used to represent the difference between the point-level attention weight extracted based on the training point cloud data and the point-level attention weight label. Among them, the structured parameter prediction loss can include midpoint loss, component length loss, and quadrant classification loss. The midpoint loss is used to regress the terminal position, the component length loss is used to regress the component length, and the quadrant classification loss is used to regress the quadrant in which the wireframe is located. The midpoint loss, component length loss, and quadrant classification loss are respectively expressed as:
[0100] ,
[0101] ,
[0102] ,
[0103] in, represents the midpoint loss term, represents the component length loss term, represents the quadrant classification loss term, represents the forecast midpoint, represents the midpoint label, represents the length of the prediction component, Indicates the component length label, represents the prediction quadrant, Indicates the quadrant label.
[0104] Furthermore, the attention weight loss can be expressed as:
[0105] ,
[0106] in, represents the attention weight loss, represents the number of training points in the training point cloud data, represents the point-level attention weight extracted based on the training point cloud data, Represents the point-level attention weight label.
[0107] In one implementation, the process of generating the point-level attention weight label specifically includes:
[0108] Get the vertex set and edge set of the annotated wireframe corresponding to the training point cloud data;
[0109] For each training point in the training point cloud data, searching for the nearest edge and nearest point corresponding to the training point in the vertex set and edge set, and calculating a first distance from the training point to the nearest vertex and a second distance to the nearest edge;
[0110] Determine the target distance and the projection factor corresponding to the training point according to the first distance and the second distance;
[0111] Construct the point-level attention weight labels corresponding to the training point cloud data based on the target distance and projection factor corresponding to all training points.
[0112] Specifically, the vertex set V and edge set E of the wireframe corresponding to the training point cloud data P are obtained. Thus, after obtaining the vertex set V and edge set E, for each training point in the training point cloud data P, , search the distance to its vertex and the nearest edge , and then calculate the training points separately To the nearest vertex The first distance and the second distance to the nearest edge .like > (Indicates that the edge is closer than the vertex), using training points The ratio of the distance from the projected position on the nearest edge to the edge endpoint Update the default projection factor. Otherwise, if <= (meaning the edge is farther than the vertex), the default projection factor remains unchanged. However, the first distance and the second distance The minimum value in is taken as the target distance , to obtain the target distance list D and projection factor list F corresponding to the training point cloud data P. Finally, the target distance list is normalized, that is, , so that the smaller the distance point is, the larger the factor value is obtained, and the point-level attention weight label is calculated based on the normalized target distance list, projection factor list and preset balance ratio r , where the attention weight label is expressed as:
[0113] .
[0114] This application determines the point-level attention weight labels by combining the target distance list, the projection factor list and the preset balance ratio, determines the correlation between the point and the geometric structure by the nearest distance, further adjusts the weight by the projection position, emphasizes the endpoints of the edge, and uses the balance ratio to control the relative importance of the two factors, guiding subsequent processing to pay more attention to the key parts of the geometric structure, ensuring that points near the edge endpoints can obtain higher weights.
[0115] In summary, this embodiment provides a method for reconstructing a building model based on a diffusion model, the method comprising obtaining the coding features and point-level attention weights of point cloud data through a feature extraction module; randomly sampling Gaussian noise from a Gaussian noise distribution through a diffusion module to obtain a noise vector; performing edge denoising on the noise vector according to the coding features and the point-level attention weights through an edge denoising module to obtain a denoised wireframe, and reconstructing the building model based on the denoised wireframe. The present application uses wireframe technology to represent the surface of the building, thereby achieving reconstruction of the building with a smaller number of facets, thereby reducing the burden of calculation and storage. In addition, the present application combines coding features and point-level attention weights, and utilizes a diffusion model to denoise the noise vector to obtain the structural parameters of the wireframe. This not only enables progressive refinement of the building wireframe, but also extracts geometric information from the point-level attention weights, further improving the reconstruction accuracy of edge details.
[0116] Based on the above-mentioned building model reconstruction method based on the diffusion model, this embodiment provides a building model reconstruction device based on the diffusion model, such as Figure 4 As shown, the building model reconstruction device based on the diffusion model specifically includes:
[0117] Feature extraction module 100, used to obtain the encoding features and point-level attention weights of point cloud data;
[0118] Diffusion module 200 for randomly sampling Gaussian noise from a Gaussian noise distribution to obtain a noise edge;
[0119] The edge denoising module 300 is configured to perform edge denoising on the noise vector according to the encoding feature and the point-level attention weight to obtain a denoised wireframe, and reconstruct a building model based on the denoised wireframe.
[0120] Based on the above-mentioned diffusion model-based building model reconstruction method, this embodiment provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in the diffusion model-based building model reconstruction method as described in the above-mentioned embodiment.
[0121] Based on the above-mentioned building model reconstruction method based on the diffusion model, the present application also provides a terminal device, such as Figure 5 As shown, it includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via bus 24. The display screen 21 is configured to display a preset user guidance interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logic instructions in the memory 22 to execute the method described in the above embodiment.
[0122] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0123] The memory 22, as a computer-readable storage medium, can be configured to store software programs or computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes the software programs, instructions, or modules stored in the memory 22 to perform functional applications and data processing, thereby implementing the methods in the above embodiments.
[0124] The memory 22 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal device. In addition, the memory 22 may include high-speed random access memory and non-volatile memory. For example, various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, may also be transient storage media.
[0125] In addition, the specific process of loading and executing the multiple instructions in the storage medium and the processor in the terminal device has been described in detail in the above method and will not be described here one by one.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A building model reconstruction method based on a diffusion model, characterized in that: A building model reconstruction model is applied, wherein the building model reconstruction model includes a feature extraction module and a diffusion model, wherein the diffusion model includes a diffusion module and an edge denoising module; the building model reconstruction method based on the diffusion model specifically includes: The feature extraction module obtains the encoding features and point-level attention weights of the point cloud data, where the point-level attention weights are the potential edge features identified from the point cloud data; Gaussian noise randomly sampled from the Gaussian noise distribution through the diffusion module to obtain a noise vector; Performing edge denoising on the noise vector according to the encoding feature and the point-level attention weight by an edge denoising module to obtain a denoised wireframe, and reconstructing a building model based on the denoised wireframe; The feature extraction module includes an encoding unit, an embedding learning unit, and a point-edge attention generation unit; the encoding features and point-level attention weights of the point cloud data obtained by the feature extraction module specifically include: Extracting point embedding features of the point cloud data by the embedding learning unit; Extracting encoding features of the point cloud data according to the point embedding features by the encoding unit; The point-level attention weights of the point cloud data are generated according to the point embedding features by the point edge attention generation unit.
2. The building model reconstruction method based on the diffusion model according to claim 1, characterized in that: Performing edge denoising on the noise vector according to the encoding feature and the point-level attention weight specifically includes: generating an intermediate noise vector based on the noise vector using a self-attention mechanism; A denoised wireframe is determined according to the encoded features, the point-level attention weights and the intermediate noise vector using a cross-attention mechanism.
3. The building model reconstruction method based on the diffusion model according to claim 2, characterized in that: Determining the denoised wireframe using the cross-attention mechanism according to the encoding features, the point-level attention weights, and the intermediate noise vector specifically includes: Fusing the encoding features and the point-level attention weights to form a point-level attention feature map; Constructing a query vector based on the intermediate noise vector, and constructing a value vector and a key vector based on the point-level attention feature map; A denoised wireframe is determined based on the query vector, the value vector, and the key vector using a cross-attention mechanism.
4. The building model reconstruction method based on the diffusion model according to claim 1, characterized in that: The reconstructing the building model based on the denoised wireframe specifically includes: The denoised wireframe is post-processed using a non-maximum suppression method and a clustering method, and a building model is reconstructed based on the post-processed denoised wireframe.
5. The building model reconstruction method based on the diffusion model according to claim 1, characterized in that: During the training process of the building model reconstruction model, the process of obtaining the noise vector specifically includes: Obtaining annotation structural parameters of each annotation wireframe of the training point cloud data, wherein the annotation structural parameters include a vector representation of each edge in the annotation wireframe, and the vector representation of the edge includes a midpoint coordinate and an offset; The annotation structured parameters of the annotation wireframe are diffused into Gaussian noise to obtain a noise vector.
6. The building model reconstruction method based on the diffusion model according to claim 1, characterized in that: The loss function used in the training process of the building model reconstruction model includes structured parameter prediction loss and attention weight loss, wherein the generation process of the point-level attention weight label used by the attention weight loss specifically includes: Get the vertex set and edge set of the annotated wireframe corresponding to the training point cloud data; For each training point in the training point cloud data, searching for the nearest edge and nearest point corresponding to the training point in the vertex set and edge set, and calculating a first distance from the training point to the nearest vertex and a second distance to the nearest edge; If the first distance is less than or equal to the second distance, use the first distance as the target distance corresponding to the training point, and use the default projection factor as the projection factor corresponding to the training point; If the first distance is greater than the second distance, the second distance is used as the target distance corresponding to the training point, and the ratio of the distance from the projection position of the training point on the nearest edge to the edge endpoint is used as the projection factor; Construct the point-level attention weight labels corresponding to the training point cloud data based on the target distance and projection factor corresponding to all training points.
7. A building model reconstruction device based on a diffusion model, characterized in that: The building model reconstruction device based on the diffusion model specifically includes: A feature extraction module is used to obtain the encoding features and point-level attention weights of the point cloud data, where the point-level attention weights are potential edge features identified from the point cloud data; a diffusion module for randomly sampling Gaussian noise from a Gaussian noise distribution to obtain a noise vector; an edge denoising module, configured to perform edge denoising on the noise vector according to the encoding feature and the point-level attention weight to obtain a denoised wireframe, and reconstruct a building model based on the denoised wireframe; The feature extraction module includes an encoding unit, an embedding learning unit, and a point edge attention generation unit; the acquisition of encoding features and point-level attention weights of point cloud data specifically includes: Extracting point embedding features of the point cloud data by the embedding learning unit; Extracting encoding features of the point cloud data according to the point embedding features by the encoding unit; The point-level attention weights of the point cloud data are generated according to the point embedding features by the point edge attention generation unit.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the building model reconstruction method based on the diffusion model according to any one of claims 1 to 6.
9. A terminal device, characterized in that: include: processor and memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the processor implements the steps of the building model reconstruction method based on the diffusion model according to any one of claims 1 to 6.