Tranform-based virtual assembly error detection method
Through the virtual assembly error detection method based on Transformer, the problem of poor mechanized registration accuracy is solved, and the assembly error detection with high accuracy and strong adaptability is achieved, and the manufacturing deviation of prefabricated components is quantified.
Patent Information
- Application Number
- CN202510343356.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In the existing virtual assembly error detection method, the problem of poor accuracy and low adaptability caused by mechanized registration cannot be effectively quantified.
Using a virtual assembly error detection method based on Transformer, the assembly error is quantified through denoising, initialization registration, enhanced feature generation, similarity matrix filtering and re-registration, combined with voxelization technology.
High-precision and highly adaptable assembly error detection is achieved, and the rotation error is 14.25% and translation error is reduced by 18.68%, achieving the geometric deviation quantization accuracy of millimeter-level.
Smart Images

Figure CN120259383A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of assembly error detection, and particularly to a virtual assembly error detection method based on Transformer. Background Art
[0002] Prefabricated structures have been widely used in the construction and infrastructure fields due to their advantages in improving construction quality, shortening construction periods, saving resources, and addressing labor shortages. They not only represent an important direction for the future development of the industry but also align with the current trends of green building and intelligent construction. With the transformation and upgrading of the construction industry, prefabricated structures play a crucial role. Prefabricated components are prefabricated in factories. First, refined designs are completed using Computer Aided Design (CAD) and Building Information Modeling (BIM), and then industrial production is carried out using molds. However, manufacturing deviations are an inherent risk in the production of prefabricated components, and severe deviations can lead to structural safety hazards, construction delays, or even catastrophic accidents. Therefore, detecting the deviations of prefabricated components is crucial for the successful implementation of prefabricated buildings and is also a basic step in subsequent assembly processes.
[0003] Before transporting prefabricated components to the construction site, problems that may occur during on-site assembly must be addressed to ensure a smooth assembly process. However, physical pre-assembly has significant limitations, including high labor costs, large space requirements, and the risk of damaging components during disassembly. To overcome the limitations of physical pre-assembly methods, Virtual Trial Assembly (VTA) has been introduced as an innovative solution. VTA uses digital technology to simulate component assembly in a computer-generated virtual environment, which helps to identify problems early and develop optimal solutions before actual construction activities. It is worth noting that VTA can significantly reduce time and costs while maintaining accuracy.
[0004] Currently, various VTA systems have been applied to the assembly inspection of building structures. Many scholars in the academic community have also conducted extensive research on the VTA method. The geometric deviation detection of prefabricated components applicable to VTA involves analyzing manufacturing deviations through point cloud data (PCD) and BIM, and the steps are as follows: First, establish a digital model of each component and use 3D laser scanning technology to obtain the point cloud data of the actual component; Second, use the registration algorithm to align the scanned point cloud with the BIM model; Finally, after the registration is completed, quantify the geometric deviation between the actual component and the BIM model. In this process, the registration algorithm is a key element, and its accuracy directly affects the accuracy of deviation detection. However, the registration algorithms used for VTA mainly rely on manually defined rules and usually execute according to predetermined algorithm steps. For example, the Iterative Closest Point (ICP) algorithm relies on specific geometric data representations, such as point cloud normal vectors, planes, and edges, as well as manually defined processes such as linear optimization and iterative update to align the point cloud. This method is very mechanical and may fall into local optima or not converge at all. In addition, it cannot adjust its registration strategy according to the scene complexity or data characteristics. Summary of the Invention
[0005] The present invention provides a virtual assembly error detection method based on Transformer to solve the defects of poor accuracy and low adaptability of assembly error detection caused by mechanical registration in the prior art, and to improve the accuracy and adaptability of assembly error detection.
[0006] The present invention provides a virtual assembly error detection method based on Transformer, including:
[0007] Use a scanner to obtain the source point cloud of each component in the component, and convert the BIM model of the component into the target point cloud of each component after segmentation;
[0008] After denoising the source point cloud of each component, perform initial registration on the source point cloud and target point cloud of each component based on the context guidance module to obtain the registered source point cloud;
[0009] After enhancing the registered source point cloud, use Transformer to generate the features of the source point cloud and target point cloud based on the self-attention mechanism;
[0010] Determine the similarity matrix between the features of the source point cloud and target point cloud, filter the source point cloud according to the similarity matrix, and perform re-registration on the target point cloud corresponding to the filtered source point cloud;
[0011] Voxelize the source point cloud and the target point cloud after re-registration, and determine the assembly error of the component according to the distance between the corresponding voxels in the source point cloud and the target points.
[0012] According to a virtual assembly error detection method based on Transformer provided by the present invention, denoise the source point cloud of each component, including:
[0013] After denoising the source point cloud of each component using a k-d tree, use a bilateral filter to denoise the source point cloud again.
[0014] According to a virtual assembly error detection method based on Transformer provided by the present invention, use the following formula to denoise the source point cloud again using a bilateral filter:
[0015]
[0016] where p i is the geometric position of the i-th point in the source point cloud, n i is the normal vector of the i-th point, α i is the bilateral filtering factor of the i-th point, ω i is the curvature of the i-th point, is the updated geometric position of the i-th point, is the smoothing weight, is the feature domain weight, p j is the distance to the k1 nearest neighboring points of p i , σ c is the influence factor of the distance from p i to each neighboring point on this point, σ s is the influence factor of the projected distance of each neighboring point on the normal vector at p i on this point.
[0017] According to a virtual assembly error detection method based on Transformer provided by the present invention, initialize the registration of the source point cloud and the target point cloud of each component through the following formula, and obtain the registered source point cloud, including:
[0018] v = h φ (C(max(f θ (X)), max(f θ (Y))))
[0019] where X is the source point cloud of each component, Y is the target point cloud of each component, f θ is the encoder module shared by the source point cloud and the target point cloud of each component, max represents max pooling, C represents the connection operation, h φIt is a transformation decoder. The first four values in v represent the quaternion rotation vector, which is used to calculate the initial transformation matrix R0, and the last three values represent the translation vector t0.
[0020] A virtual assembly error detection method based on Transformer provided by the present invention enhances the registered source point cloud through the following formula, including:
[0021]
[0022] where x is the geometric coordinate of any point in the registered source point cloud, m i is the geometric coordinate of the i-th point among the k adjacent points closest to x, is the enhanced feature of m i PPF i is the point-to-point feature of the i-th point, is the enhanced feature of x, μ θ contains multiple convolutional layers, and max is the max pooling layer.
[0023] A virtual assembly error detection method based on Transformer provided by the present invention filters the source point cloud according to the following formula based on the similarity matrix:
[0024]
[0025] where, is the feature of the point X' in the registered source point cloud X', and the feature of the registered source point cloud X' s1 and the feature of the corresponding target point cloud between the similarity matrix M is the number of points in the source point cloud, M is the number of points in the source point cloud, by deleting the non-overlapping points in X' to satisfy obtained, top-prob represents the index of obtaining a certain amount of the highest similarity score in the probability ratio of prob, X' s2 is the filtered source point cloud, and the similarity matrix of the representative overlapping points is obtained is the feature of X' s2 of.
[0026] A virtual assembly error detection method based on Transformer provided by the present invention determines the target point cloud corresponding to the filtered source point cloud for re-registration through the following formula:
[0027] J = argmax top-k H s2,i
[0028]
[0029] Among them, J represents the index of the first k points with the largest similarity scores in Y with respect to X' s2 The i-th point x in s2,i The updated geometric coordinates of the target point corresponding to the i-th point, y i is the target point cloud, and ω ij is the weight between the i-th point and y i among them.
[0030] According to a virtual assembly error detection method based on Transformer provided by the present invention, before using a scanner to obtain the source point cloud of each component in a component and converting the BIM model of the component into the target point cloud of each component after segmentation, the registration model is trained through the following loss function:
[0031]
[0032] Among them, L f is the re-registration loss, L init is the initial registration loss, L ol is the overlap score loss, X is the source point cloud sample, R and R0 are the rotation matrices for re-registration and initial registration respectively, t and t0 are the translation vectors for re-registration and initial registration respectively, and are the configured rotation matrix and translation vector annotated for the source point cloud sample respectively, S X and S Y are the overlap scores between the i-th point and the j-th point in the source point cloud sample, and the overlap scores between the i-th point and the j-th point in the target point cloud sample corresponding to the source point cloud sample respectively, N is the number of points in the source point cloud sample, M is the number of points in the target point cloud sample, and correspond to the annotated overlap scores of S X and S Y respectively.
[0033] According to a virtual assembly error detection method based on Transformer provided by the present invention, the calculation formulas of S X and S Y are:
[0034]
[0035] Among them, r is the operation of expanding the dimension and repeating the feature, is the overlap decoder, C is the connection operation, F X and F Y are the features obtained after the source point cloud sample and the corresponding target point cloud sample are processed by a shared encoder module, and are obtained after pooling through the corresponding F X and F Y respectively.
[0036] The virtual assembly error detection method based on Transformer provided by the present invention proposes a denoising algorithm for the noise problem in large-scale scanned point cloud data; develops a deep learning-driven point cloud registration network, which achieves high-precision registration while ensuring the spatial consistency of the point cloud data; and develops a geometric deviation quantification algorithm using the aligned data, which effectively quantifies the deviation value between the actual component and its corresponding theoretical model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 is a schematic flow chart of the virtual assembly error detection method based on Transformer provided by the present invention;
[0039] Figure 2 is a schematic framework diagram of the virtual assembly error detection method based on Transformer provided by the present invention;
[0040] Figure 3 is a schematic framework diagram of the context-guided module in the virtual assembly error detection method based on Transformer provided by the present invention;
[0041] Figure 4 is a schematic structural diagram of the Transformer-based feature matching module in the virtual assembly error detection method based on Transformer provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0043] The following will be combined with Figure 1Describe the Transformer-based virtual assembly error detection method of the present invention, including:
[0044] Step 101, use a scanner to obtain the source point cloud of each component in the component, and after segmenting the BIM model of the component, convert it into the target point cloud of each component;
[0045] A bracket-type or hand-held laser scanner can be used to comprehensively scan the component to obtain the source point cloud of the component.
[0046] To ensure the smooth progress of subsequent registration experiments, Revit software can be used to finely process the BIM model, systematically divide it into multiple independent sub-components. This subdivision process improves the accuracy and efficiency of registration. Subsequently, the segmented BIM model is converted into corresponding point cloud data to ensure the consistency of the model data and the measured point cloud in terms of data type.
[0047] Step 102, after denoising the source point cloud of each component, perform initial registration on the source point cloud and target point cloud of each component based on the context guidance module to obtain the registered source point cloud;
[0048] The correspondence generated by the repeated patterns in the environment produces a large number of outliers (such as points on the ground, ceiling, or flat walls of the building), which hinders the quality of pose estimation. Therefore, if these points can be detected and filtered out, the registration method can be made more robust. Point cloud denoising not only prevents subsequent steps from converging to a poor solution, but also helps to enhance the saliency of features by pre-discarding redundant and non-significant points.
[0049] The goal of registration is to align two unordered point clouds, namely the source point cloud P and the target point cloud Q. To achieve registration, a correspondence relationship between the two point clouds needs to be established, and then the adverse effects of outliers are suppressed through robust estimation. Formally, assume that the Kth pair of point clouds (or the Kth correspondence relationship) obtained through matching consists of the 3D point m i ∈P and the 3D point n j ∈Q. Then, the Kth correspondence relationship can be modeled as:
[0050] n j =Rm i +t+σ i,j
[0051] where, R∈SO(3) is the rotation matrix, is the translation vector, is the measurement noise. If the initial correspondence relationship is represented as τ, the objective function can be represented as:
[0052]
[0053] Among them, ρ represents a robust loss function designed to suppress errors caused by incorrect correspondences or outliers, and ε represents the outliers that have been filtered out before optimization. Therefore, this formula attempts to estimate the relative pose between the source point cloud and the target point cloud while being robust to outliers.
[0054] Step 103: After enhancing the registered source point cloud, use a Transformer to generate the features of the source point cloud and the target point cloud based on the self-attention mechanism.
[0055] Step 104: Determine the similarity matrix between the features of the source point cloud and the target point cloud, filter the source point cloud according to the similarity matrix, and re-register the target point cloud corresponding to the filtered source point cloud.
[0056] Step 105: Voxelize the source point cloud and the target point cloud after re-registration, and determine the assembly error of the component according to the distance between the corresponding voxels in the source point cloud and the target point cloud.
[0057] This study aims to answer three main research questions: (1) How to quickly denoise large-scale point cloud data without sacrificing accuracy; (2) How to develop a highly robust learning-based registration algorithm; (3) How to quantify geometric deviations based on aligned data. To address these issues, this study has carried out the following work: (i) Developed a denoising method applicable to large-scale point clouds; (ii) Extracted features from the denoised point clouds through machine learning algorithms and trained a point cloud registration neural network; (iii) Proposed a voxel-based high-precision deviation calculation method. The innovation of this study lies in designing a Transformer-enhanced point cloud registration method that effectively combines the global features of point clouds, and proposing a voxel-based geometric deviation quantification method to ensure high-precision quantification of manufacturing deviations.
[0058] This embodiment proposes a method for detecting component manufacturing deviations through point cloud registration, which is used to quantify the geometric dimension deviations between the actual component and the theoretical BIM model. This method mainly includes four key steps: (1) Steel component point cloud acquisition and BIM model segmentation; (2) Noise removal of large-scale scanned point cloud data; (3) Deep learning-based point cloud registration network; (4) Deviation detection algorithm using aligned data. The overall framework of this method is as Figure 2 shown.
[0059] Through experiments with actual data, the following key results were obtained: (1) The proposed denoising algorithm effectively reduces noise and smooths the point cloud while preserving key geometric features; (2) Compared with GeoTransformer, this registration algorithm shows higher accuracy, reducing the rotation error (Error(R)) by 14.25% and the translation error (Error(t)) by 18.68%. Notably, each pairwise registration operation consumes only 0.109 seconds per frame; (3) The geometric deviation detection algorithm achieves millimeter-level accuracy, with an average geometric deviation of 2.73 millimeters observed in the dataset. The innovation of this study lies in proposing a comprehensive point cloud registration network that covers an integrated process from point cloud scanning, BIM segmentation, point cloud denoising, registration to deviation quantification, demonstrating high accuracy, computational efficiency, and versatility in field measurements.
[0060] In this embodiment, a denoising algorithm is proposed for the noise problem in large-scale scanned point cloud data; a deep learning-driven point cloud registration network is developed, which achieves high-precision registration while ensuring the spatial consistency of point cloud data; using the aligned data, a geometric deviation quantification algorithm is developed, which effectively quantifies the deviation value between the actual component and its corresponding theoretical model.
[0061] Based on the above embodiment, in this embodiment, the source point cloud of each component is denoised, including:
[0062] After denoising the source point cloud of each component using the k-d tree, the source point cloud is denoised again using a bilateral filter.
[0063] Due to the large number and irregular distribution of point clouds, this embodiment proposes an outlier noise point filtering algorithm combining the k-d tree and the bilateral filter. This method first roughly denoises using the k-d tree algorithm, filtering out isolated outliers in the dataset, and then improves the bilateral filter to smooth the point cloud.
[0064] The three-dimensional raw point cloud data obtained by scanning is scattered and disordered. According to the k-d tree algorithm, the point cloud topological relationship is established, and the average distance between any point in the calculated point cloud and each point in its neighborhood follows a Gaussian distribution with a mean of μ and a standard deviation of δ. If the average distance of a certain point from all points in its neighborhood exceeds the set threshold μ + λδ, then this point is determined as a noise point and removed from the point cloud. Before using the k-d tree algorithm to search for neighborhood points, the initial value of k and the standard deviation multiple parameter λ need to be set. The selection of the initial value of k directly affects the noise point removal effect of the point cloud model. To determine the values of k and λ, this embodiment adds a set of random noises to the Bunny point cloud and conducts experiments.
[0065] Using the k-d tree algorithm can eliminate most of the outliers. However, there are still small-scale noise points mixed with the main point cloud data in the point cloud, which are difficult to eliminate by common methods based on Euclidean distance and point cloud density judgment. In this embodiment, the curvature ω of the sampling points i is used as a parameter to improve the bilateral filtering factor, so as to filter out small-scale noise points and smooth the entire point cloud.
[0066] On the basis of the above embodiment, in this embodiment, the source point cloud is denoised again by using a bilateral filter through the following formula:
[0067]
[0068] where p i is the geometric position of the i-th point in the source point cloud, n i is the normal vector of the i-th point, α i is the bilateral filtering factor of the i-th point, ω i is the curvature of the i-th point, is the updated geometric position of the i-th point, and are both Gaussian kernel functions, is the smoothing weight (spatial weight) of the distance from p i to the neighborhood points, is the feature domain weight (influence weight), p j is the k1 nearest adjacent points to p i , σ c is the influence factor of the distance from p i to each adjacent point on this point, generally taking the radius of the neighborhood, σ s is the influence factor of the projected distance of each adjacent point on the normal vector at p i on this point, generally taking the standard deviation of the near neighbor points. When σ c is determined, the smoothing distance of the point cloud in the normal direction is proportional to σ s , and <> is the inner product operation.
[0069] The surface normal is an important property of the geometric body surface. By fitting the local surface of the point cloud, the normal vector can be well estimated. The principal component analysis method can be used to estimate the normal vector of the sampling points. For the sampling point p i , perform a k-nearest neighbor search to obtain k neighborhood points N(p i ) to approximately replace the local neighborhood surface S of the point p i , that is, p i ∈S, and construct a covariance matrix C for the neighborhood points:
[0070]
[0071] The three eigenvalues of the covariance matrix are denoted as λ0, λ1, and λ2, and the eigenvector corresponding to the smallest eigenvalue is used as the normal vector n of the local surface of the point cloud at p i at the point i . The curvature at any sampled point in the point cloud can be characterized by the curvature of the local surface fitted by the point and its neighboring points, and the principal component analysis method is used to estimate the curvature ω i at the point p i which can be expressed as:
[0072]
[0073] The bilateral filtering factor depends on the local neighborhood feature information. When the local area of the point cloud model is relatively sharp, it is difficult to distinguish the normal direction of the sampled points, so the noise cannot be accurately found. To optimize the neighborhood of the sampled points and improve the accuracy of noise point search, in this embodiment, the curvature ω i of the sampled points is used as a parameter to improve the bilateral filtering factor α i .
[0074] The complete steps of denoising include: (1) initializing the value of k, constructing a k-d tree, and calculating the k nearest neighbors of the point cloud data point p i ; (2) using PCA to estimate the normal vector n i and the curvature value ω i of the point cloud; (3) calculating the smoothing weight parameter x and the feature domain weight parameter y of each sampled point p i ; (4) calculating the smoothing weight and the feature domain weight (5) calculating the bilateral filtering factor α i ; (6) calculating the geometric position of the point p i to obtain the denoised point cloud model.
[0075] Based on the above embodiment, in this embodiment, the source point cloud and the target point cloud of each component are initialized and registered through the following formula based on the context guidance module to obtain the registered source point cloud, including:
[0076]
[0077] where X is the source point cloud of each component, Y is the target point cloud of each component, and f θ is the encoder module shared by the source point cloud and the target point cloud of each component, and this module learns the high-dimensional features corresponding to X and Y max represents max pooling, C represents the concatenation operation, and h φ The transformation decoder is a simple multi-layer perceptron, and the first four values in v represent the quaternion rotation vector used to calculate the initial transformation matrix R0, and the last three values represent the translation vector The transformed source point cloud X' can be expressed as X' = X · R T + t T It is calculated that the summation of t0 is carried out row by row.
[0078] Rough initial registration is beneficial to point feature learning. In addition, the overlapping points between point cloud pairs can greatly improve the matching accuracy, and the point cloud matching accuracy is crucial for registration. Therefore, a context-guided module is proposed to initialize the registration and predict the scores of overlapping points, and its overall structure is as Figure 3 shown.
[0079] The modified PointNet++ network is used as the encoder of the context-guided module. The encoder module is used to extract global features and regress the 7D vector representing the rotation and translation transformation.
[0080] Based on the above embodiments, in this embodiment, the registered source point cloud is enhanced by the following formula, including:
[0081]
[0082] where x is the geometric coordinate of any point in the registered source point cloud, m i is the geometric coordinate of the i-th point among the k nearest neighbor points to x, is m i 's enhanced feature, PPF i is the point-to-point feature of the i-th point, is the enhanced feature of x, μ θ contains several 1-D convolutional layers, and max is the max pooling layer. The point pair feature PPF describes the relative position and pose of two directed points.
[0083] After coarse registration, the fine registration is completed through the point correspondence between the transformed source point cloud X' and the target point cloud Y. At this stage, a Transformer-based feature matching module is proposed to remove the non-representative points in the source point cloud X', and its structure is as Figure 4 shown.
[0084] Before the Transformer, the point cloud feature enhancement is carried out according to the method proposed in RPMNet. For the point x ∈ X', k nearest neighbor points are obtained through a predefined radius For m i ∈ M k 's enhanced feature is denoted as It connects the spatial coordinates and the point pair feature
[0085] The point features need to consider the global information of the point cloud. Transformer integrates point features based on the global point cloud input, showing advantages over local features. Given the enhanced features of the transformed source point cloud X' are The point features can be generated based on self-attention, as shown in the following formula:
[0086]
[0087] where · represents matrix multiplication, and i is the coefficient of self-attention. and are the learned weights, C0 is a hyperparameter used to control the computational amount. softmax represents row normalization, and norm represents column normalization. The output features for X' and Y can be expressed as and
[0088] Based on the above embodiments, in this embodiment, the source point cloud is filtered according to the similarity matrix by the following formula:
[0089]
[0090] where is the feature of point X' s1 in the registered source point cloud X', the feature of the registered source point cloud X' and the feature of the corresponding target point cloud the similarity matrix between M is the number of points in the source point cloud, X' s1 is obtained by deleting the non-overlapping points in X' to satisfy top-prob represents the index of obtaining a certain amount of the highest similarity scores in the probability ratio of prob, X' s2 is the filtered source point cloud, and the similarity matrix of the representative overlapping points is the feature of X' s2 thus transforming the part-to-part registration into part-to-whole registration.
[0091] In the feature matching removal stage, first calculate the similarity matrix However, the points in X' may not find valid corresponding points in Y, and this correspondence matrix satisfies The non-overlapping points in X' can be deleted, so as to ensure that the remaining points X' s1 satisfy Therefore, it can solve the part-to-part registration problem.
[0092] In this way, representative overlapping points can be obtained in two steps. First, according to the overlapping score S predicted by the context guidance module X select the top N points as overlapping points Secondly, perform feature matching to further eliminate the points without feature description, and obtain the final representative points X' s2 .
[0093] Based on the above embodiments, in this embodiment, the target point cloud corresponding to the filtered source point cloud is determined by the following formula for re-registration:
[0094] J = argmax top-k H s2,i
[0095]
[0096] where J represents the index in Y that has the top k maximum similarity scores for the i-th point x s2 in X', y s2,i is the updated geometric coordinate of the target point corresponding to the i-th point, Y is the target point cloud, and ω i is the weight between the i-th point and y ij i .
[0097] For each point x s2,i ∈X' s2 , the corresponding point y i is calculated. To generate the corresponding set, each point pair is defined as (x s2,i , y i ), and its weight is Then use weighted singular value decomposition to estimate the exact transformation R1 ∈ SO(3) and
[0098] Based on the above embodiments, in this embodiment, before using the scanner to obtain the source point cloud of each component in the component and converting the BIM model of the component into the target point cloud of each component, the registration model is trained through the following loss function:
[0099]
[0100] where L f is the re-registration loss, L init is the initialization registration loss, L ol is the overlapping score loss, X is the source point cloud sample, R and R0 are the rotation matrices for re-registration and initialization registration respectively, t and t0 are the translation vectors for re-registration and initialization registration respectively, and They are the configuration rotation matrix and translation vector for the ground truth of the source point cloud sample, S X and S Y are respectively the overlap scores between the i-th point and the j-th point in the source point cloud sample, and the overlap scores between the i-th point and the j-th point in the corresponding target point cloud sample of the source point cloud sample. N is the number of points in the source point cloud sample, and M is the number of points in the target point cloud sample. and correspond to the ground truth overlap scores of S X and S Y respectively.
[0101] The rotation error Error(R) and translation error Error(t) are used as evaluation indicators for registration, and their calculation formulas can be expressed as:
[0102]
[0103] where R and t represent the rotation matrix and translation vector, and tr represents the matrix trace. In addition, the mean absolute error of the Euler angles and translation vector is also used as an evaluation indicator, and its calculation formula is:
[0104]
[0105] where N represents the number of samples.
[0106] Based on the above embodiments, in this embodiment, the calculation formulas of S X and S Y are:
[0107]
[0108] where r is the operation of expanding dimensions and repeating features, is the overlap decoder, The parameters of are shared in the input point cloud. C is the connection operation, and F X and F Y are the features obtained after the source point cloud sample and the corresponding target point cloud sample are processed by the shared encoder module. and are respectively obtained by pooling through the corresponding F X and F Y respectively.
[0109] The common information of the source point cloud and the target point cloud is crucial for the prediction of overlapping points. Therefore, a simple information interaction module is designed to predict overlapping points based on global features, which can be obtained from the initialization step, thus saving computational effort. Here, the overlapping point prediction problem is formulated as a binary classification problem.
[0110] The registered point cloud pairs cannot directly calculate the geometric deviation between them. Although the two groups of point clouds are in the same coordinate system, the number of points may not be the same. Additionally, since the deep neural network used for point cloud registration is a black box model, the index between points is unknown. These two reasons make it unrealistic to calculate the distance between each pair of points through brute-force methods. Voxel-based methods are widely used in point cloud processing. By dividing the space into regular three-dimensional voxel grids, they can solve the problem of uneven point cloud density and thus unify the scale of the point cloud.
[0111] To quantify the geometric deviation between the registered point clouds, a local region definition method based on voxel grids is adopted. First, the entire point cloud needs to be voxelized, that is, the three-dimensional space is divided into cube voxel units of the same size. The size of the voxel is usually set according to the characteristics of the point cloud and the resolution requirements. The basic steps of voxelization can be simply described as follows: Given a point cloud dataset P = {p1, p2, … p n}, define the side lengths of the voxel Δx, Δy, Δz, and map each point p i = (x i , y i , z i ) in the point cloud to the corresponding voxel V k . All points belonging to the same voxel are regarded as the point set P k = {p j | p j ∈ V k}. For the point set P k in each voxel V k , a series of statistical quantities are used to describe the geometric characteristics and distribution of the voxel. Its centroid or mean coordinate is:
[0112]
[0113] where k is the number of points in the voxel V k . The standard deviation can be expressed as:
[0114]
[0115] In the source point cloud and the target point cloud, for two corresponding voxels V k and V l , the distance difference between them can be expressed as:
[0116]
[0117] Due to the different densities of points in the voxel, a certain weight ω k, which can effectively balance the contributions of different voxels to the geometric deviation. The weight value is calculated by the normalized standard deviation (it should be noted that the calculation of the weight is only for the target point cloud voxels), and can be expressed as:
[0118]
[0119] Based on the above analysis, by averaging the weighted geometric errors of all voxels, the global geometric deviation of the target point cloud can be obtained, and the calculation formula is expressed as:
[0120]
[0121] Among them, N g is the number of global voxels.
[0122] The above algorithm is used to calculate the global error of the target point cloud, that is, to evaluate the overall manufacturing quality of the component. The local geometric deviation can reveal the manufacturing accuracy of the component in a specific area. If the local deviation in a certain area is large, it indicates that there are defects in that area of the component. The local error can be quantified by the geometric distribution, centroid of the points in the local area, and the differences from the neighboring points. Statistically calculate the distance values between the current voxel and its surrounding neighboring voxels and the corresponding voxels of the target point cloud, and define its local area. If the distance value between the current voxel V k and the corresponding voxel V l is greater than the global error E global , then the neighboring voxels are clustered using the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm. For the clustered voxels, they can be regarded as the local area L(V k ). The average geometric deviation of the local area can be calculated by the following formula:
[0123]
[0124] Among them, N l is the number of voxels in the local area.
[0125] To verify the performance of the registration network, data training was carried out on the collected dataset in this paper. All experimental processes were conducted on a computer equipped with RTX 4070Ti Super and 64GB Inter-i9, with the computer system being Ubuntu20.04. Relevant algorithms and network models were developed using Python 3.8.0 and Pytorch 1.18.0. The initial learning rate was set to 0.0001 in the experiment, the number of training epochs was set to 600, the batch size was set to 8, the learning rate was changed using the cosine annealing schedule, non-iterative adaptive learning was performed using the Adam optimizer, and CUDA was used to accelerate the model training process. The neighbor radius used to calculate features in PointNet++ was set to 0.3. For the context-guided module, the number of convolutional filters during the encoding process was [64, 64, 64, 128, 512], and the number of filters in the transformation and overlapping decoders was [512, 512, 256, 7] and [512, 512, 256, 2] respectively.
[0126] In this study, to deeply explore the detection method of geometric deviation in virtual assembly, we conducted systematic experimental verification on the effectiveness of the algorithm, and carried out a detailed evaluation of the entire process of point cloud denoising preprocessing, point cloud BIM registration, and virtual assembly geometric deviation analysis based on the above dataset. By removing noise from the large-scale scanned point cloud data, we effectively reduced the random noise in the data, ensuring a high-quality basis for subsequent processing. The effectiveness of noise removal provided accurate input data for the point cloud registration step and provided more robust support for the deep learning-driven point cloud registration network. The bilateral filtering point cloud denoising method we proposed, combined with the deep learning network for point cloud registration, can effectively handle complex geometric shapes and high-noise environments. Secondly, the geometric error detection algorithm based on the registration result can accurately identify the geometric deviation in the virtual assembly process, with an accuracy reaching the millimeter level. To comprehensively evaluate the performance of the proposed method, this study conducted experimental analysis from multiple dimensions, including data processing time, registration accuracy, error accuracy, and its feasibility in practical applications. The comparison of the registration algorithm with the SOTA method further verified the advantages of the proposed method in multiple key indicators. The following conclusion analysis section will deeply explore these experimental results:
[0127] (1) The bilateral filtering denoising algorithm for point clouds proposed in this paper can effectively denoise and smooth 3D point cloud data, while well preserving the geometric features of the model. In this experiment, to test the effect of the point cloud denoising algorithm, two component point clouds were selected, and noise with a variance of 0 and a standard deviation of 0.02m was added to the point cloud models respectively. We compared the algorithm proposed in this paper with traditional filtering denoising algorithms (MRPCA, MLS) and deep learning-based point cloud denoising algorithms (PCN, Score). The experimental results are shown in Table 1. The results show that the proposed denoising algorithm significantly reduces the influence of random noise while retaining the effective point cloud information. Compared with traditional and learning-based algorithms, P2S is reduced by 13.27% and 10.97% respectively, and the CD index is reduced by 20% and 18.28% respectively. By comparing the denoising results of different algorithms, it can be seen that the algorithm we proposed effectively smooths the smoothness of the point cloud surface while denoising, reducing the surface fluctuations caused by noise. This feature can significantly improve the visual quality of the point cloud. Finally, through comparing the geometric feature parameters (such as edge sharpness, surface continuity, etc.) before and after denoising, the results show that the proposed algorithm effectively retains the main geometric features of the model during the denoising process and does not cause significant losses to the key structural details. This advantage makes the algorithm have higher application value in virtual assembly and precise geometric analysis.
[0128] Table 1
[0129]
[0130] Reviewing the existing literature, most of the deviation detection methods in precast component manufacturing rely on pre-developed software or do not deeply study the overall process of deviation detection. These methods usually only utilize partial feature information of the component, such as corner points and holes, during the registration process, while ignoring the overall information of the component. This oversight leads to a decline in registration accuracy. To solve these problems, this paper proposes a comprehensive and efficient component deviation detection method. First, a fast large-scale point cloud denoising technology is introduced to reduce noise interference. Second, a registration network containing global point cloud information is trained to ensure the precise registration between the point cloud data and the BIM model. Finally, a deviation quantification algorithm based on the aligned data is developed, which can effectively quantify the manufacturing deviation of the component.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A virtual assembly error detection method based on Transformer, characterized in that Including: Using a scanner to obtain the source point cloud of each component in the component, and after segmenting the BIM model of the component, converting it into the target point cloud of each component; After denoising the source point cloud of each component, based on the context-guided module, initial registration of the source point cloud and target point cloud of each component is performed to obtain the registered source point cloud; After enhancing the registered source point cloud, use Transformer to generate the features of the source point cloud and target point cloud based on the self-attention mechanism; Determine the similarity matrix between the features of the source point cloud and target point cloud, filter the source point cloud according to the similarity matrix, and perform re-registration on the target point cloud corresponding to the filtered source point cloud; Voxelize the re-registered source point cloud and target point cloud, and determine the assembly error of the component according to the distance between the corresponding voxels in the source point cloud and target point.
2. The virtual assembly error detection method based on Transformer according to claim 1, wherein Denoising the source point cloud of each component includes: After denoising the source point cloud of each component using a k-d tree, use a bilateral filter to perform secondary denoising on the source point cloud.
3. The virtual assembly error detection method based on Transformer according to claim 2, wherein The source point cloud is subjected to secondary denoising using a bilateral filter through the following formula: where p i is the geometric position of the i-th point in the source point cloud, n i is the normal vector of the i-th point, α i is the bilateral filtering factor of the i-th point, ω i is the curvature of the i-th point, is the updated geometric position of the i-th point, is the smoothing weight, is the feature domain weight, p j is the distance to p i of the nearest k1 adjacent points, σ c is the influence factor of the distance from p i to each adjacent point on this point, σ s is the influence factor of the projection distance of each adjacent point on the normal vector at p i on this point.
4. The virtual assembly error detection method based on Transformer according to claim 1, characterized in that Initial registration of the source point cloud and target point cloud of each component based on the context-guided module through the following formula to obtain the registered source point cloud, including: v = h φ (C(max(f θ (X)), max(f θ (Y)))) Among them, X is the source point cloud of each component, Y is the target point cloud of each component, and f θ is the encoder module shared by the source point cloud and the target point cloud of each component, max represents max pooling, C represents the concatenation operation, and h φ is the transformation decoder. The first four values in v represent the quaternion rotation vector, which is used to calculate the initial transformation matrix R0, and the last three values represent the translation vector t0.
5. The virtual assembly error detection method based on Transformer according to claim 1, characterized in that Enhancing the registered source point cloud through the following formula, including: where x is the geometric coordinate of any point in the registered source point cloud, and m i is the geometric coordinate of the i-th point among the k nearest neighboring points to x, and is the enhanced feature of m i , PPF i is the point pair feature of the i-th point, is the enhanced feature of x, μ θ contains multiple convolutional layers, and max is the max pooling layer.
6. The virtual assembly error detection method based on Transformer according to claim 1, characterized in that Filtering the source point cloud according to the similarity matrix through the following formula: Among them, is the feature of point X' in the registered source point cloud X', s1 the feature of the registered source point cloud X', and the feature of the corresponding target point cloud The similarity matrix between M is the number of points in the source point cloud, X' s1 is obtained by deleting the non-overlapping points in X' to satisfy obtained, top-prob represents the index that obtains a certain amount of the highest similarity score in the probability ratio of prob, X' s2 is the filtered source point cloud, and the similarity matrix of the representative overlapping points is obtained is the feature of X' s2 of.
7. The method for detecting virtual assembly errors based on Transformer according to claim 6, wherein Determining the re-registration of the target point cloud corresponding to the filtered source point cloud through the following formula: J = argmax top-k H s2,i Among them, J represents the index of the first k points in Y with the largest similarity scores to the i-th point x' s2 in s2 , y s2,i is the updated geometric coordinate of the target point corresponding to the i-th point, Y is the target point cloud, ω i is the weight between the i-th point and y ij i 8. The virtual assembly error detection method based on Transformer according to any one of claims 1-7, characterized in that Before using a scanner to obtain the source point cloud of each component in the component and converting the BIM model of the component into the target point cloud of each component after segmentation, the registration model is trained through the following loss function: Among them, L f is the re-registration loss, L init is the initial registration loss, L ol is the overlap score loss, X is the source point cloud sample, R and R0 are the rotation matrices of re-registration and initial registration respectively, t and t0 are the translation vectors of re-registration and initial registration respectively, and are the configured rotation matrix and translation vector of the source point cloud sample annotation respectively, S X and S Y are the overlap scores of the i-th point and the j-th point in the source point cloud sample, and the overlap scores of the i-th point and the j-th point in the corresponding target point cloud sample of the source point cloud sample respectively. N is the number of points in the source point cloud sample, and M is the number of points in the target point cloud sample. and correspond to the labeled overlap scores of S X and S Y respectively.
9. The method for detecting virtual assembly errors based on Transformer according to claim 8, characterized in that, S X and S Y The calculation formula is as follows: Among them, r is an operation to expand the dimension and repeat features, is an overlapping decoder, C is a connection operation, F X and F Y are features obtained after the source point cloud sample and the corresponding target point cloud sample are processed by a shared encoder module, and are respectively obtained after pooling through the corresponding F X and F Y
Citation Information
Patent Citations
Subway station construction quality evaluation method based on laser scanning technology
CN115018249A
Modularized integrated building component intelligent detection method and system
CN117455905A
Road element point cloud BIM reverse modeling method and system based on deep learning
CN118643561A
Method for automatically generating BIM model based on point cloud data
CN119131262A
Apparatus and method for searching for global minimum of point cloud registration error
US20220254095A1
Cited By
BIM (Building Information Modeling)-based steel structural member anti-collision detection method
CN120563516A
BIM-based steel structure component collision detection method
CN120563516B