Point cloud-to-MESH real estate building cluster 3D reconstruction method based on chessboard folding convolution
Through the point cloud to MESH method based on chessboard folding convolution, the problems of large computational complexity and insufficient sampling of 3D reconstruction methods are solved, efficient multi-scale feature extraction and refined reconstruction are achieved, and 3D reconstruction adapts to complexity and diversity.
Patent Information
- Application Number
- CN202510832952.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing 3D reconstruction methods are computationally intensive, have insufficient sampling, and are highly complex, making it difficult to effectively render community building complex models.
A point cloud to MESH method based on chessboard folding convolution is adopted. A 3D reconstruction model is constructed through multi-layer chessboard folding convolution, feature fusion and hierarchical perception weights. The pre-trained model is used to process the reconstructed data, generate a three-dimensional structure diagram and optimize the model parameters.
It realizes adaptive extraction of multi-scale features, improves reconstruction accuracy and detail expression ability, adapts to the complexity and diversity of input images, and avoids the limitations of fixed architecture on specific scenarios.
Smart Images

Figure CN120707770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of 3D modeling and reconstruction, and in particular to a 3D reconstruction method for real estate building clusters based on chessboard folded convolution and point cloud to MESH conversion. Background Art
[0002] Current technologies primarily model point clouds, but these methods are computationally expensive, and insufficient sampling can lead to model collapse. Due to the complexity of community building clusters, these methods can be difficult to render in these environments and are hardware-dependent. Summary of the Invention
[0003] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a 3D reconstruction method for real estate building clusters based on chessboard folded convolution from point cloud to MESH, so as to solve the problems of excessive computation, insufficient sampling and high complexity in existing 3D reconstruction methods.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] A 3D reconstruction method for real estate building clusters based on chessboard folded convolution from point cloud to MESH, including:
[0006] Collect image data of the target building to obtain data to be reconstructed;
[0007] The data to be reconstructed is input into a pre-trained 3D reconstruction model for processing to obtain a three-dimensional structure diagram; the training process of the 3D reconstruction model includes:
[0008] Collecting several groups of data to be reconstructed;
[0009] Build the original network model;
[0010] Using the FoldConv module of the original network model to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features;
[0011] Using the FeatureFusion module of the original network model to perform feature fusion on the multi-layer convolution features to obtain a three-dimensional feature representation;
[0012] Processing the three-dimensional feature representation using a multilayer perceptron of the original network model to obtain hierarchical perception weights;
[0013] Multiplying and accumulating the hierarchical perception weight and the three-dimensional feature representation to obtain a final feature representation;
[0014] Mapping the final feature representation to a three-dimensional space using a decoding function of the original network model to obtain an iterative structure graph;
[0015] The iterative structure graph is calculated using a loss function to obtain a total loss value, and the total loss value is used to optimize and iterate the parameters of the original network model to obtain the trained 3D reconstruction model.
[0016] Preferably, the iterative structure graph is calculated using a loss function to obtain a total loss value, and the original network model is optimized and iterated using the total loss value to obtain the trained 3D reconstruction model, including:
[0017] Define the loss function; the expression of the loss function is: E total =E+λ1E c +λ2E d ;in, E total is the total loss value; E, E c 、E d They are respectively the generation graph loss, feature extraction loss, and dynamic layer perception loss; λ1 and λ2 are respectively the first weight hyperparameter and the second weight hyperparameter; Z is the iterative structure graph; Z * is the actual structure diagram corresponding to the iteration structure diagram;
[0018] Calculating the loss value during the training of the original network model using the loss function to obtain the total loss value;
[0019] The model parameters of the original network model are optimized using the Adam optimizer according to the total loss value, and the 3D reconstructed model is obtained after iteration.
[0020] Preferably, determining the data to be reconstructed includes:
[0021] Collect 2D images of the target building;
[0022] A three-dimensional tensor that determines the chessboard layout of the model;
[0023] The two-dimensional image and the three-dimensional tensor are spliced to obtain the data to be reconstructed.
[0024] Preferably, the FoldConv module of the original network model is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features, including:
[0025] Set the normalization function;
[0026] The two-dimensional coordinates of the two-dimensional image are normalized using the normalization function; the normalization function includes: Wherein, x and y are the horizontal and vertical coordinates of the two-dimensional coordinates respectively; W and H are the width and height of the image respectively; xnorm 、y norm are the normalized horizontal and vertical coordinates respectively.
[0027] Preferably, the FoldConv module of the original network model is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features, including:
[0028] Determine the folded convolution kernel and folding parameters of the first layer;
[0029] According to the folding convolution kernel and the folding parameters of the first layer, the FoldConv module is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain the initialization feature; the expression of the initialization feature is: Z (0) =FoldConv(I,W f(1) ,α1); where Z (0) is the initialization feature; I is the data to be reconstructed; W f(1) , α1 are the folded convolution kernel and the folding parameter of the first layer respectively; FoldConv(·) represents multi-layer chessboard folded convolution.
[0030] Preferably, the FeatureFusion module of the original network model is used to perform feature fusion on the multi-layer convolution features to obtain a three-dimensional feature representation, including:
[0031] Determining the folded convolution kernel and the folding parameters of the lth layer;
[0032] Obtaining the three-dimensional feature representation of the l-1th layer;
[0033] According to the folded convolution kernel and the folding parameters of the lth layer, the FeatureFusion module is used to perform feature fusion on the three-dimensional feature representation of the l-1th layer to obtain the three-dimensional feature representation of the lth layer; the expression of the three-dimensional feature representation of the lth layer is:
[0034] Z (l) =FeatureFusion(Z (l-1) ,W f(l) ,α l );
[0035] Among them, Z (l) is the three-dimensional feature representation of the lth layer; FeatureFusion(·) represents the feature fusion operation.
[0036] Preferably, the three-dimensional feature representation is processed using a multilayer perceptron of the original network model to obtain hierarchical perceptual weights, including:
[0037] The three-dimensional feature representation of the lth layer is processed by the multilayer perceptron to obtain the perceptual features of the lth layer; the perceptual features of the lth layer are expressed as: g (l) =MLP(Z (l) ); where g (l) is the perceptual feature of the lth layer; MLP(·) represents multi-layer perceptual processing;
[0038] The activation function is used to calculate the perceptual features of the lth layer to obtain the hierarchical perceptual weight of the lth layer; the expression of the hierarchical perceptual weight of the lth layer is: (l) =σ(g (l) ); where ω (l) is the hierarchical perception weight of layer l; σ(·) is the activation function.
[0039] Preferably, a 3D reconstruction system for real estate building clusters based on chessboard folded convolution and point cloud to MESH conversion includes:
[0040] Model generation module, used to build and train 3D reconstruction models;
[0041] A data acquisition module, used to determine the data to be reconstructed;
[0042] A multi-layer convolution module is used to perform multi-layer chessboard folding convolution on the data to be reconstructed using the FoldConv module of the original network model to obtain multi-layer convolution features;
[0043] A feature fusion module, configured to perform feature fusion on the multi-layer convolution features using the FeatureFusion module of the original network model to obtain a three-dimensional feature representation;
[0044] a perception weight extraction module, configured to process the three-dimensional feature representation using a multilayer perceptron of the original network model to obtain hierarchical perception weights;
[0045] A weighted fusion module, configured to multiply and accumulate the hierarchical perception weights and the three-dimensional feature representation to obtain a final feature representation;
[0046] The spatial mapping module is used to map the final feature representation into a three-dimensional space using the decoding function of the original network model to obtain a three-dimensional structure graph.
[0047] Preferably, an electronic device comprises: at least one processor, and a memory communicatively connected to the processor; wherein the memory stores instructions that can be executed by the processor, and the instructions are executed by the processor so that the processor can execute the aforementioned 3D reconstruction method of real estate building clusters based on chessboard folding convolution and converting point cloud to MESH.
[0048] Preferably, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the aforementioned 3D reconstruction method of real estate building clusters by converting point cloud to MESH based on chessboard folding convolution.
[0049] The present invention discloses the following technical effects:
[0050] The present invention provides a 3D reconstruction method for real estate building clusters by converting point clouds into meshes based on chessboard folding convolution. By performing multi-layer chessboard folding convolution on the data to be reconstructed, the dimensional limitation of traditional convolution in 3D space is solved, and adaptive extraction of multi-scale features is achieved. By performing feature fusion on the multi-layer convolution features, the defect of poor correlation between point clouds and mesh structures in the existing technology is solved, and the reconstruction accuracy and detail expression ability are improved. Through hierarchical perception weights, the limitations of fixed architecture on specific scenes are solved, and automatic adaptation to the complexity and diversity of input images is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 A schematic diagram of the 3D reconstruction process of a real estate building cluster based on chessboard folded convolution from point cloud to MESH provided in an embodiment of the present invention;
[0053] Figure 2 A schematic diagram of the technical framework provided by an embodiment of the present invention;
[0054] Figure 3 This is a schematic diagram of the final 3D effect provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0056] The purpose of the present invention is to provide a 3D reconstruction method for real estate building clusters based on chessboard folded convolution from point cloud to MESH, so as to solve the problems of excessive computation, insufficient sampling and high complexity in existing 3D reconstruction methods.
[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0058] Figure 1 A schematic diagram of the 3D reconstruction process of real estate building clusters based on chessboard folded convolution from point cloud to MESH provided in an embodiment of the present invention. Figure 2 A schematic diagram of the technical framework provided by an embodiment of the present invention, such as Figure 1 and Figure 2 As shown, the present invention provides a 3D reconstruction method for real estate building clusters based on point cloud to MESH conversion based on chessboard folding convolution, comprising:
[0059] Step 100: Collect image data of the target building to obtain data to be reconstructed;
[0060] Step 200: Input the data to be reconstructed into a pre-trained 3D reconstruction model for processing to obtain a three-dimensional structure diagram. The training process of the 3D reconstruction model includes:
[0061] Step 201: Collect several groups of data to be reconstructed;
[0062] Step 202: constructing an original network model;
[0063] Step 203: using the FoldConv module of the original network model to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features;
[0064] Step 204: using the FeatureFusion module of the original network model to perform feature fusion on the multi-layer convolution features to obtain a three-dimensional feature representation;
[0065] Step 205: Processing the three-dimensional feature representation using the multi-layer perceptron of the original network model to obtain hierarchical perceptual weights;
[0066] Step 206: multiplying and accumulating the hierarchical perception weight and the three-dimensional feature representation to obtain a final feature representation;
[0067] Step 207: Mapping the final feature representation to a three-dimensional space using the decoding function of the original network model to obtain an iterative structure graph;
[0068] Step 208: Calculate the iterative structure graph using a loss function to obtain a total loss value, and use the total loss value to optimize and iterate the parameters of the original network model to obtain the trained 3D reconstruction model.
[0069] Furthermore, the iterative structure graph is calculated using a loss function to obtain a total loss value, and the original network model is optimized and iterated using the total loss value to obtain the trained 3D reconstruction model, including:
[0070] Define the loss function; the expression of the loss function is: E total =E+λ1E c +λ2E d ;in, E total is the total loss value; E, E c 、E d They are respectively the generation graph loss, feature extraction loss, and dynamic layer perception loss; λ1 and λ2 are respectively the first weight hyperparameter and the second weight hyperparameter; Z is the iterative structure graph; Z * is the actual structure diagram corresponding to the iteration structure diagram;
[0071] Calculating the loss value during the training of the original network model using the loss function to obtain the total loss value;
[0072] The model parameters of the original network model are optimized using the Adam optimizer according to the total loss value, and the 3D reconstructed model is obtained after iteration.
[0073] Specifically, determining the data to be reconstructed includes:
[0074] Collect 2D images of the target building;
[0075] A three-dimensional tensor that determines the chessboard layout of the model;
[0076] The two-dimensional image and the three-dimensional tensor are spliced to obtain the data to be reconstructed.
[0077] Preferably, the FoldConv module of the original network model is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features, including:
[0078] Set the normalization function;
[0079] The two-dimensional coordinates of the two-dimensional image are normalized using the normalization function; the normalization function includes: Wherein, x and y are the horizontal and vertical coordinates of the two-dimensional coordinates respectively; W and H are the width and height of the image respectively; x norm 、y norm are the normalized horizontal and vertical coordinates respectively.
[0080] Furthermore, the FoldConv module of the original network model is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features, including:
[0081] Determine the folded convolution kernel and folding parameters of the first layer;
[0082] According to the folding convolution kernel and the folding parameters of the first layer, the FoldConv module is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain the initialization feature; the expression of the initialization feature is: Z (0) =FoldConv(I,W f(1) ,α1); where Z (0) is the initialization feature; I is the data to be reconstructed; W f(1) , α1 are the folded convolution kernel and the folding parameter of the first layer respectively; FoldConv(·) represents multi-layer chessboard folded convolution.
[0083] Specifically, the FeatureFusion module of the original network model is used to perform feature fusion on the multi-layer convolution features to obtain a three-dimensional feature representation, including:
[0084] Determining the folded convolution kernel and the folding parameters of the lth layer;
[0085] Obtaining the three-dimensional feature representation of the l-1th layer;
[0086] According to the folded convolution kernel and the folding parameters of the lth layer, the FeatureFusion module is used to perform feature fusion on the three-dimensional feature representation of the l-1th layer to obtain the three-dimensional feature representation of the lth layer; the expression of the three-dimensional feature representation of the lth layer is:
[0087] Z (l) =FeatureFusion(Z (l-1) ,W f(l) ,α l );
[0088] Among them, Z (l) is the three-dimensional feature representation of the lth layer; FeatureFusion(·) represents the feature fusion operation.
[0089] Furthermore, the three-dimensional feature representation is processed using a multilayer perceptron of the original network model to obtain hierarchical perception weights, including:
[0090] The three-dimensional feature representation of the lth layer is processed by the multilayer perceptron to obtain the perceptual features of the lth layer; the perceptual features of the lth layer are expressed as: g (l) =MLP(Z(l) ); where g (l) is the perceptual feature of the lth layer; MLP(·) represents multi-layer perceptual processing;
[0091] The activation function is used to calculate the perceptual features of the lth layer to obtain the hierarchical perceptual weight of the lth layer; the expression of the hierarchical perceptual weight of the lth layer is: (l) =σ(g (l) ); where ω (l) is the hierarchical perception weight of layer l; σ(·) is the activation function.
[0092] Specifically, the overall theory of this embodiment is as follows:
[0093] 1) Input data
[0094] Assume that there are image data from different views, recorded as X = x1, x2, ... where x e Represents the image matrix of the e-th view. At the same time, the layout information of the chessboard needs to be input, denoted as C, which represents the coordinate information and label (such as color, edge, etc.) of each grid on the chessboard. The input data X and C will be used as the input of the network.
[0095] 2) Network architecture initialization
[0096] Design a two-encoder chessboard folded convolutional network, one for source encoding and one for target encoding. The network architecture is as follows:
[0097] Z=g θ (X,C)
[0098] Among them, Z represents the output feature matrix of the network, θ represents the network parameters, g θ Represents the network forward propagation function. The network structure mainly includes a source encoder and a target encoder, which encode the source image and the chessboard structure respectively.
[0099] 3) Feature extraction
[0100] First, the features of the multi-view image are extracted through the source encoder:
[0101]
[0102] Among them, Z s represents the feature matrix output by the source encoder, It is the feature extraction function of the source encoder, which mainly extracts two-dimensional coordinates and RGB features.
[0103] Next, the features of the chessboard layout are extracted through the target encoder:
[0104]
[0105] Among them, Z t represents the point cloud feature matrix output by the target encoder, It is the feature extraction function of the target encoder, where the RGB color and the two-dimensional coordinate tensor are mapped.
[0106] Furthermore, a network architecture based on chessboard folded convolution is designed and implemented. Convolution kernel selection and design: Determine how the convolution kernel slides across the chessboard layout to extract key local features. Network depth: Given the complexity of the chessboard layout, multiple layers of convolution are required to extract higher-level abstract features. Let K be the convolution kernel, which performs a convolution operation on the chessboard layout C. The goal of convolution is to extract local features by applying the convolution kernel K to different regions of the chessboard, thereby obtaining a final feature tensor Z. The convolution kernel K is a structured filter that can extract higher-order features from the chessboard data. Specifically, when processing the chessboard, the convolution kernel may focus on specific patterns or positional relationships. These patterns may not be global, but rather local structural features, such as the relative positions of chess pieces on the board. Convolution operation: In each convolution step, the convolution kernel K scans the input feature map C (chessboard layout) and calculates the convolution result. By adjusting the size and stride of the convolution kernel, the receptive field and the granularity of feature extraction can be controlled.
[0107] Preferably, the input data is represented as follows. Assume there is a three-dimensional tensor X∈R with a chessboard layout H×W×D , where H and W represent height and width respectively, and D represents depth. The depth value corresponding to each pixel position (h, w) is D h,w . In addition, the chessboard image I∈R H×W×C2 , where C2 represents the number of channels. The input data consists of a chessboard image and a depth map, namely:
[0108] D=[I,X]
[0109] Furthermore, the chessboard folded convolution is defined. The chessboard folded convolution is a special convolution operation used to map two-dimensional image information to three-dimensional space. Its basic idea is to fold the input data through multiple convolution kernels to generate a multi-layer three-dimensional feature map. Specifically, a folded convolution kernel W is defined. f ∈R K×K×D , where K is the size of the convolution kernel. The folded convolution operation can be expressed as:
[0110] Normalized 2D coordinates convert the 2D coordinates (x, y) to normalized coordinates in the range [-1, 1]:
[0111]
[0112] Determine the level of the chessboard grid. According to the level l of the chessboard grid (counting from 0), determine the chessboard grid where each point is located:
[0113] l=min{i|x norm <-α l 2 i+1 or y norm <-α l 2 i+1}
[0114] Where i is the checkerboard grid level starting from 0, α l is the folding parameter of the lth layer, ranging from 0 to 1.
[0115] Calculate the Z-axis depth based on the level l of the chessboard and the maximum number of layers L:
[0116]
[0117] Where L is the maximum checkerboard grid level (adjustable to control the folding depth). Calculate the 3D coordinates based on the normalized 2D coordinates and depth values:
[0118] X=x+x norm
[0119] Y=y+y norm
[0120] Z as described above
[0121] Keep the RGB values unchanged: RGB values are directly mapped to corresponding points in three-dimensional space:
[0122] R,G,B=Color(x,y)
[0123] The above process is uniformly represented by FoldConv, where R, G, and B represent the color values of the coordinate pixels, and Color(x,y) represents the color value corresponding to the extracted coordinates (x,y).
[0124] Furthermore, multi-layer chessboard folded convolutions are required to extract higher-level abstract features, taking into account the complexity of the chessboard layout. Furthermore, three-dimensional folding is also required. This step is responsible for converting two dimensions to three dimensions. To enhance the network's expressive power, multi-layer chessboard folded convolutions are introduced. Each layer's folded convolution kernel has different parameters, extracting features at different levels. The folded convolution operation for layer 1 is:
[0125] Z (l) =FoldConv(Z (l-1) ,W f(l) ,α l )
[0126] Preferably, dynamic hierarchical perception. In order to capture the dynamic characteristics of the chessboard layout, a dynamic hierarchical perception module is introduced after each layer of chessboard folding convolution. This module processes the features of each layer through a multi-layer perceptron (MLP) and generates hierarchical perception weights. Specifically, the hierarchical perception module of the lth layer is:
[0127] g (l) =MLP(Z (l) )
[0128] ω (l) =σ(g (l) )
[0129] Specifically, in order to optimize the network, a loss function needs to be defined. Assuming there is a target three-dimensional structure graph Z, the loss function can be defined as:
[0130]
[0131] In order to further improve network performance, dynamic layer-aware loss and optimization loss are introduced:
[0132] E total =E+λ1E c +λ2E d
[0133] Furthermore, the optimization process is performed. The Adam optimizer is used to minimize the total loss function:
[0134]
[0135] Where α is the learning rate, is the gradient of the total loss function with respect to the parameter θ, and θ' is the optimized parameter.
[0136] Specifically, initial feature extraction: Assume there is an input two-dimensional image I∈R H×W×C , where H and W represent the height and width respectively, and C represents the number of channels. The low-level features of the image are extracted through the initialized chessboard folded convolution operation:
[0137] Z (0) =FoldConv(I,W f(1) ,α1)
[0138] Feature fusion:
[0139] Through multi-layer chessboard folding convolution operations, features are gradually fused to generate a more advanced three-dimensional feature representation. The feature fusion process of the first layer is:
[0140] Z (l) =FeatureFusion(Z (l-1) ,W f(l) ,αl )
[0141] Dynamic hierarchical perception: The dynamic hierarchical perception module of the lth layer is:
[0142] g (l) =MLP(Z (l) )
[0143] ω (l) =σ(g (l) )
[0144] Combination of multiple layers of folded convolution: Through multiple iterations of the chessboard folded convolution process, the point cloud and mesh structure can be gradually reconstructed. Specifically, the weighted summation of all layer perception weights and the corresponding feature maps is performed to obtain the final feature representation:
[0145]
[0146] Where L is the number of chessboard folded convolution layers.
[0147] 3D reconstruction output:
[0148] Through the decoding function g, the final feature representation Z (L) Map to three-dimensional space and generate point cloud and Mesh structure:
[0149]
[0150] in, is the generated three-dimensional structure diagram, and g is the decoding function.
[0151] In order to optimize the network, a loss function needs to be defined. Assuming there is a target three-dimensional structure graph Z, the loss function can be defined as:
[0152]
[0153] In order to further improve network performance, dynamic layer-aware loss and optimization loss are introduced:
[0154] E total =E+λ1E c +λ2E d
[0155] Among them, E c is the feature extraction loss, E d is the dynamic layer-aware loss, and λ1 and λ2 are weight hyperparameters.
[0156] Optimization process:
[0157] Use the Adam optimizer to minimize the total loss function:
[0158]
[0159] The network architecture based on chessboard folded convolution effectively maps two-dimensional image information to three-dimensional space through multi-layer folded convolution and dynamic hierarchical perception modules, and can capture the multi-scale features of the chessboard layout. Through multiple iterations of the chessboard folded convolution process, the point cloud and Mesh structure can be gradually reconstructed to improve the accuracy and detail expression of 3D reconstruction. Finally, the 3D reconstruction task is completed, and the generated Mesh model is compared and optimized with the original 3D data to ensure the accuracy of the reconstruction results. Through multiple iterations of chessboard folded convolution and dynamic hierarchical perception modules, a complete point cloud and Mesh structure are constructed. At this point, the generated Mesh model needs to be compared and optimized with the original 3D data to further improve the accuracy and detail expression of reconstruction. Through the above steps, a high-quality 3D reconstruction result can be obtained, which further improves the performance and practicality of the model. The final 3D effect diagram is for reference Figure 3 .
[0160] The beneficial effects of the present invention are as follows:
[0161] By performing multi-layer chessboard folding convolution on the data to be reconstructed, the present invention can effectively map two-dimensional image information into three-dimensional space, realize the adaptive extraction of multi-scale features, and can simultaneously capture details and large scene structures, avoiding the dimensionality limitation of traditional convolution in 3D space; by performing feature fusion on the multi-layer convolution features, the point cloud and mesh structure can promote each other, gradually improving the reconstruction accuracy and detail expression ability; through hierarchical perception weights, it can automatically adapt to the complexity and diversity of the input image, avoiding the limitation of fixed architecture on specific scenes.
[0162] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0163] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A 3D reconstruction method for real estate building clusters based on point cloud to MESH conversion based on chessboard folding convolution, characterized by: include: Collect image data of the target building to obtain data to be reconstructed; Inputting the data to be reconstructed into a pre-trained 3D reconstruction model for processing to obtain a three-dimensional structure diagram; The training process of the 3D reconstruction model includes: Collecting several groups of data to be reconstructed; Build the original network model; Using the FoldConv module of the original network model to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features; Using the FeatureFusion module of the original network model to perform feature fusion on the multi-layer convolution features to obtain a three-dimensional feature representation; Processing the three-dimensional feature representation using a multilayer perceptron of the original network model to obtain hierarchical perception weights; Multiplying and accumulating the hierarchical perception weight and the three-dimensional feature representation to obtain a final feature representation; Mapping the final feature representation to a three-dimensional space using a decoding function of the original network model to obtain an iterative structure graph; The iterative structure graph is calculated using a loss function to obtain a total loss value, and the total loss value is used to optimize and iterate the parameters of the original network model to obtain the trained 3D reconstruction model.
2. The 3D reconstruction method of real estate building clusters based on point cloud to MESH based on chessboard folding convolution according to claim 1 is characterized in that: The iterative structure graph is calculated using a loss function to obtain a total loss value, and the original network model is optimized and iterated using the total loss value to obtain the trained 3D reconstruction model, including: Define the loss function; the expression of the loss function is: E total =E+λ1E c +λ2E d ;in, E total is the total loss value; E, E c 、E d They are respectively the generation graph loss, feature extraction loss, and dynamic layer perception loss; λ1 and λ2 are respectively the first weight hyperparameter and the second weight hyperparameter; Z is the iterative structure graph; Z * is the actual structure diagram corresponding to the iteration structure diagram; Calculating the loss value during the training of the original network model using the loss function to obtain the total loss value; The model parameters of the original network model are optimized using the Adam optimizer according to the total loss value, and the 3D reconstructed model is obtained after iteration.
3. The 3D reconstruction method of real estate building clusters based on point cloud to MESH based on chessboard folding convolution according to claim 1 is characterized in that: Determine the data to be reconstructed, including: Collect 2D images of the target building; A three-dimensional tensor that determines the chessboard layout of the model; The two-dimensional image and the three-dimensional tensor are spliced to obtain the data to be reconstructed.
4. The 3D reconstruction method of real estate building clusters based on point cloud to MESH based on chessboard folding convolution according to claim 3 is characterized in that: The FoldConv module of the original network model is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features, including: Set the normalization function; The two-dimensional coordinates of the two-dimensional image are normalized using the normalization function; the normalization function includes: Wherein, x and y are the horizontal and vertical coordinates of the two-dimensional coordinates respectively; W and H are the width and height of the image respectively; x norm 、y norm are the normalized horizontal and vertical coordinates respectively.
5. The 3D reconstruction method of real estate building clusters based on point cloud to MESH based on chessboard folding convolution according to claim 3 is characterized in that: The FoldConv module of the original network model is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain multi-layer convolution features, including: Determine the folded convolution kernel and folding parameters of the first layer; According to the folding convolution kernel and the folding parameters of the first layer, the FoldConv module is used to perform multi-layer chessboard folding convolution on the data to be reconstructed to obtain the initialization feature; the expression of the initialization feature is: Z (0) =FoldConv(I,W f(1) ,α1); where Z (0) is the initialization feature; I is the data to be reconstructed; W f(1) , α1 are the folded convolution kernel and the folding parameter of the first layer respectively; FoldConv(·) represents multi-layer chessboard folded convolution.
6. The method for 3D reconstruction of real estate building clusters based on point cloud to MESH conversion based on chessboard folding convolution according to claim 5 is characterized in that: The FeatureFusion module of the original network model is used to perform feature fusion on the multi-layer convolution features to obtain a three-dimensional feature representation, including: Determining the folded convolution kernel and the folding parameters of the lth layer; Obtaining the three-dimensional feature representation of the l-1th layer; According to the folded convolution kernel and the folding parameters of the lth layer, the FeatureFusion module is used to perform feature fusion on the three-dimensional feature representation of the l-1th layer to obtain the three-dimensional feature representation of the lth layer; the expression of the three-dimensional feature representation of the lth layer is: WITH (l) =FeatureFusion(Z (l-1) ,IN f(l) ,α l ); Among them, Z (l) is the three-dimensional feature representation of the lth layer; FeatureFusion(·) represents the feature fusion operation.
7. The method for 3D reconstruction of real estate building clusters based on point cloud to MESH conversion based on chessboard folding convolution according to claim 6 is characterized in that: The three-dimensional feature representation is processed using the multilayer perceptron of the original network model to obtain hierarchical perception weights, including: The three-dimensional feature representation of the lth layer is processed by the multilayer perceptron to obtain the perceptual features of the lth layer; the perceptual features of the lth layer are expressed as: g (l) =MLP(Z (l) ); where g (l) is the perceptual feature of the lth layer; MLP(·) represents multi-layer perceptual processing; The activation function is used to calculate the perceptual features of the lth layer to obtain the hierarchical perceptual weight of the lth layer; the expression of the hierarchical perceptual weight of the lth layer is: (l) =σ(g (l) ); where ω (l) is the hierarchical perception weight of layer l; σ(·) is the activation function.
8. A 3D reconstruction system for real estate building clusters based on chessboard folded convolution and point cloud to MESH conversion, characterized by: The method for 3D reconstruction of real estate building clusters based on point cloud to MESH conversion based on chessboard folding convolution as described in claim 1 comprises: Model generation module, used to build and train 3D reconstruction models; A data acquisition module, used to determine the data to be reconstructed; A multi-layer convolution module is used to perform multi-layer chessboard folding convolution on the data to be reconstructed using the FoldConv module of the original network model to obtain multi-layer convolution features; A feature fusion module, configured to perform feature fusion on the multi-layer convolution features using the FeatureFusion module of the original network model to obtain a three-dimensional feature representation; a perception weight extraction module, configured to process the three-dimensional feature representation using a multilayer perceptron of the original network model to obtain hierarchical perception weights; A weighted fusion module, configured to multiply and accumulate the hierarchical perception weights and the three-dimensional feature representation to obtain a final feature representation; The spatial mapping module is used to map the final feature representation into a three-dimensional space using the decoding function of the original network model to obtain a three-dimensional structure graph.
9. An electronic device, characterized in that: include: At least one processor, and a memory communicatively connected to the processor; wherein the memory stores instructions that can be executed by the processor, and the instructions are executed by the processor so that the processor can execute a 3D reconstruction method for real estate building clusters based on chessboard folded convolution and point cloud to MESH according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute a 3D reconstruction method for real estate building clusters based on chessboard folding convolution and converting point cloud to MESH according to any one of claims 1 to 7.
Citation Information
Patent Citations
Ultrasonic image super-resolution reconstruction method and device based on multi-scale feature fusion
CN116258631A
Dynamic three-dimensional reconstruction method and system for local operation scene of engineering machinery
CN117292076A