Component automatic segmentation identification method suitable for bridge point cloud
By combining principal component analysis, large visual models, and multimodal large language models, we have achieved fast and accurate component-level segmentation and classification of bridge point clouds, solving the problem of low segmentation efficiency of bridge point clouds in existing technologies and realizing automated instance segmentation of bridge point clouds.
Patent Information
- Application Number
- CN202511882123.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-12-15
AI Technical Summary
Existing bridge point cloud segmentation methods suffer from large data volume, limited semantics, and difficulty in segmentation. Manual segmentation is time-consuming and prone to errors, while neural network training is time-consuming and data is scarce, resulting in low efficiency in bridge 3D reconstruction.
Principal component analysis was used to obtain the main orientation of the bridge point cloud and perform coordinate transformation. The point cloud was divided into three-dimensional meshes and rendered into two-dimensional video frames. Initial segmentation was performed using the visual large model SAM and SAM2. The nearest neighbor pairing algorithm for cluster centers of segmented components and error point iterative repair were combined. Finally, the components were classified using the multimodal large language model CLIP.
It achieves fast and accurate component-level segmentation and classification of bridge point clouds, improves the segmentation efficiency of bridge point clouds, solves the problems of difficulty and inefficiency in manual segmentation, and realizes automated instance segmentation of bridge point clouds.
Smart Images

Figure CN121305097A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bridge point cloud instance segmentation technology, and in particular to an automated segmentation and recognition method for bridge point clouds. Background Technology
[0002] Bridges are a crucial component of transportation infrastructure, and instance segmentation of bridge point clouds is essential for reconstructing existing bridge models. However, common methods for acquiring bridge point clouds rely on 3D laser scanners, resulting in large datasets with limited semantic meaning and difficulties in segmentation. When performing 3D bridge reconstruction, further segmentation of the point cloud data is necessary. Manually labeled segmentation methods are challenging, time-consuming, and prone to errors. Existing point cloud segmentation methods largely rely on neural networks for automatic segmentation, which suffers from lengthy training processes, scarce training data, and a gap between accuracy and modeling precision. Therefore, an automated component segmentation and recognition method suitable for bridge point clouds is urgently needed. This is highly valuable as it can significantly improve the segmentation efficiency of bridge point clouds and has substantial application potential. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by providing an automated segmentation and recognition method for bridge point clouds. This invention only requires the collection and analysis of bridge point cloud data, thereby achieving rapid and accurate component-level instance segmentation.
[0004] The objective of this invention is achieved through the following technical solution: an automated segmentation and recognition method for bridge point clouds, comprising the following steps: (1) The principal component analysis method is used to obtain the main direction of the bridge point cloud and coordinate transformation is performed to obtain the calibrated bridge point cloud. The calibrated bridge point cloud is divided into three-dimensional grids based on spatial location information, and the index information of the points in each grid is saved. Based on the visibility principle when the camera is shooting, the multi-view visible three-dimensional grid is extracted and rendered into a two-dimensional video frame to realize the dimensionality reduction of the three-dimensional point cloud. (2) The main view of the two-dimensional video frame is used as the input of the visual large model SAM to obtain the initial segmentation mask of each component. The prompt box generated by each segmentation mask is used together with the two-dimensional video frame as the input of the SAM2 model to obtain the segmentation mask of each component under each view. (3) Project the segmentation results obtained in step (2) under the diagonal view, match the segmentation results of the same components on both sides based on the nearest pairing algorithm of the cluster center of the segmentation components, calculate the score of the index point according to the segmentation mask of each view, and realize the component-level segmentation of the bridge point cloud. (4) Optimize the component-level segmentation results obtained in step (3) based on the error point iterative repair algorithm; (5) Based on the optimized component-level segmentation results obtained in step (4), pixel points are extracted using the segmentation mask coordinates of each viewpoint, and blank pixels around the perimeter are added to generate component images. The cosine similarity between each component image and prior knowledge text information is calculated based on the multimodal large language model CLIP to achieve the classification and recognition of bridge components.
[0005] Further, in step (1), the step of obtaining the principal direction of the bridge point cloud using principal component analysis and performing coordinate transformation to obtain the calibrated bridge point cloud specifically includes: The bridge point cloud is formally represented as an n x 3 3D point cloud coordinate matrix P. For this coordinate matrix P, the point cloud coordinate data is first decentered, and then the covariance matrix of the coordinate matrix is calculated. Singular value decomposition is performed on the covariance matrix to obtain three eigenvalues and their corresponding eigenvectors. The eigenvector corresponding to the largest eigenvalue is taken as the principal direction of the 3D point cloud, the eigenvector corresponding to the smallest eigenvalue is taken as the normal vector of the 3D point cloud, and the eigenvector corresponding to the other eigenvalue is taken as the secondary principal direction of the 3D point cloud. The principal direction is set as the x-axis, the secondary principal direction is set as the y-axis, and the normal vector is set as the z-axis. The 3D point cloud of the bridge is rotated based on the set x-axis, y-axis, and z-axis to perform coordinate transformation to obtain the calibrated bridge point cloud.
[0006] Further, in step (1), rendering it as a two-dimensional video frame specifically includes: The visible 3D mesh point cloud to video frame rendering algorithm is used to render multi-view visible 3D meshes into 2D video frames to generate video frame images from various viewpoints. The calculation formula is as follows: In the formula, Represents the i-th point in the point cloud. The angle between the reference viewpoint and the reference viewpoint is The corresponding 3D mesh index from the perspective of , , and Let x, y, and z represent the coordinates of the i-th point in the point cloud, respectively. , and These represent the indices of the 3D mesh on the x-axis, y-axis, and z-axis, respectively. The size of the 3D mesh; For rotation matrix, ; For point cloud coordinates, ; Represents the offset vector. , Let the coordinates be the center coordinates of the point cloud. This represents the minimum value of the point cloud in the three coordinate directions; correspondingly, it represents the minimum value of each group in a two-dimensional video frame. , minimum value The grid closest to the virtual camera point is the visible grid, and the rest are invisible grids. The average RGB value of all points in the visible grid is calculated as the pixel value of that grid.
[0007] Furthermore, step (2) specifically includes: The Visual Acuity Model (SAM) is used to globally segment the reference view (i.e., the main viewpoint of the 2D video frame) to obtain initial segmentation masks for each bridge component. The background mask is then removed based on the RGB values of the pixels in the mask, resulting in initial segmentation masks for each bridge component. Coordinate pairs are formed by selecting the top-left and bottom-right pixel coordinates of each component's initial segmentation mask, and these coordinate pairs are used to generate tooltips. The rendered 2D video frame images from each viewpoint are then rotated according to the specified angles. The video frames are divided into two groups clockwise and synthesized separately. Together with the prompt box, they are input into the SAM2 model for segmentation to obtain the final result. The component segmentation mask from the perspective of each component, among which The viewing angles are 0°, 15°, 165°, 180°, 195°, and 345°.
[0008] Furthermore, step (3) specifically includes: First, all visible points are extracted from the 0° and 180° viewpoints and projected onto the xOz plane. The centers of the visible points corresponding to different masks after projection are calculated from the 0° and 180° viewpoints. The visible points are determined based on the index information of the points contained in the pixel coordinate pairs. Then, for each center from the 0° viewpoint, the coordinates of the nearest center from the 180° viewpoint are found and paired. The center coordinate index from the 180° viewpoint is modified based on the category index of the mask from the 0° viewpoint. Next, the category scores are calculated to integrate the segmentation results from different viewpoints. For each point in the bridge point cloud corresponding to the mask, the score matrix of each point for each category is calculated. The category with the highest score is selected as the category of the point to achieve component-level segmentation of the bridge point cloud.
[0009] Furthermore, step (4) specifically includes: Based on the component-level segmentation results obtained in step (3), for the incorrectly segmented points, under the normal vector threshold constraint, according to the radius interval... Iteratively calculate the category of the most frequent point within the radius r of each point, from smallest to largest, and use that category as the category of the incorrectly segmented point; where and These represent the minimum radius at the start of the iteration and the maximum radius at the end of the iteration, respectively.
[0010] Furthermore, step (5) specifically includes: Based on the optimized component-level segmentation results obtained in step (4), Image groups generated by adding blank pixels of a specified width around the segmentation masks of each component from different viewpoints serve as the image input for the multimodal large language model CLIP. m text sentences containing semantic and geometric information about the bridge structure are used as the text input for CLIP. CLIP calculates the cosine similarity between each component image and the prior knowledge text information to obtain the corresponding probability, outputting the probability that each mask image belongs to each sentence. Finally, the predicted probabilities of each segmentation mask under all viewpoints are summed to obtain the final score probability of all masks. The text information corresponding to the highest score probability of each segmentation mask contains the component category of that mask, and the final classification result is... The calculation formula is as follows: In the formula, It is the category corresponding to component n. It is the x-th text vector. It is the mask image vector of component n in the i-th view, and k is the viewpoint. The total number of corresponding views; through this process, the text information corresponding to the segmentation mask of each component is determined, thereby realizing the classification and recognition of bridge components.
[0011] The beneficial effects of this invention are as follows: This invention only requires the collection of bridge point cloud data for analysis and processing, enabling rapid component-level segmentation of the bridge point cloud, and rapid classification and identification of components through the segmentation results. This achieves rapid and accurate component-level instance segmentation, and further classifies the component categories based on the segmentation results, ultimately realizing automated instance segmentation of bridge point clouds. This invention analyzes and processes bridge point clouds obtained by laser scanners through a series of methods such as video frame rendering, image segmentation, and cosine similarity calculation, enabling accurate and stable instance segmentation of bridges, and solving the problems of difficulty and inefficiency in manual segmentation. Attached Figure Description
[0012] Figure 1 This is a flowchart of the automated segmentation and recognition method for bridge point clouds according to the present invention; Figure 2 This is a diagram illustrating the effect of generating two-dimensional video frames from point clouds based on visible points of a multi-view virtual mesh, according to the present invention. Figure 3 This is a schematic diagram illustrating the effect of segmenting images from various viewpoints based on the visual large models SAM and SAM2 of the present invention. Figure 4 This is a schematic diagram of the two-sided view matching effect of the clustering center nearest pairing algorithm based on segmentation components of the present invention; Figure 5 This is a segmentation effect diagram of each component of the bridge point cloud according to the present invention; wherein, Figure 5 (a) in the image shows the segmentation result before optimization; Figure 5 (b) in the image shows the optimized segmentation result; Figure 6 This is a semantic classification effect diagram of the bridge components of the present invention. Detailed Implementation
[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not intended to limit this application.
[0014] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0015] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "in response to determination," or "includes." Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process or method. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0016] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.
[0017] See Figure 1 The present invention provides an automated segmentation and recognition method for bridge point clouds, which specifically includes the following steps: (1) The principal direction of the bridge point cloud is obtained by using Principal Component Analysis (PCA) and coordinate transformation is performed to obtain the calibrated bridge point cloud. The calibrated bridge point cloud is divided into three-dimensional grids based on spatial location information, and the index information of each point in the grid is saved. The multi-view visible three-dimensional grid is extracted based on the visibility principle when the camera is shooting, and it is rendered into a two-dimensional video frame to realize the dimensionality reduction of the three-dimensional point cloud.
[0018] Further, in step (1), the principal component analysis method is used to obtain the principal direction of the bridge point cloud, and coordinate transformation is performed to obtain the calibrated bridge point cloud. Specifically, the bridge point cloud is formally represented as an n-row, 3-column three-dimensional point cloud coordinate matrix P. For this coordinate matrix P, the point cloud coordinate data in it is first decentered, and then the covariance matrix of the coordinate matrix is calculated. The singular value decomposition of the covariance matrix is performed to obtain three eigenvalues and their corresponding eigenvectors. The eigenvector corresponding to the largest eigenvalue is taken as the principal direction (i.e., traffic direction) of the three-dimensional point cloud, the eigenvector corresponding to the smallest eigenvalue is taken as the normal vector of the three-dimensional point cloud, and the eigenvector corresponding to the other eigenvalue is taken as the secondary principal direction of the three-dimensional point cloud. The principal direction is set as the x-axis, the secondary principal direction is set as the y-axis, and the normal vector is set as the z-axis. The three-dimensional point cloud of the bridge is rotated based on the set x-axis, y-axis, and z-axis, and coordinate transformation is performed on it to obtain the calibrated bridge point cloud.
[0019] Further, in step (1), rendering it into a two-dimensional video frame specifically includes: using a point cloud-to-video frame rendering algorithm for visible 3D meshes to render multi-view visible 3D meshes into two-dimensional video frames to generate video frame images from each viewpoint, the calculation formula of which is: In the formula, Represents the i-th point in the point cloud. The angle between the reference viewpoint and the reference viewpoint is The corresponding 3D mesh index from the perspective of , , and Let x, y, and z represent the coordinates of the i-th point in the point cloud, respectively. , and These represent the indices of the 3D mesh on the x-axis, y-axis, and z-axis, respectively. The size of the 3D mesh; For rotation matrix, ; For point cloud coordinates, ; Represents the offset vector. , Let the coordinates be the center coordinates of the point cloud. This represents the minimum value of the point cloud in the three coordinate directions. Correspondingly, each group in a two-dimensional video frame... , minimum value The grid closest to the virtual camera point is the visible point grid, and the remaining grids are invisible point grids. The average RGB value of all points within the visible point grid is calculated as the pixel value of that grid. This ultimately achieves dimensionality reduction of the 3D point cloud, as shown in the image. Figure 2 As shown, Figure 2 This is a rendering of the effect of generating two-dimensional video frames by reducing the dimensionality of point clouds based on visible points of a multi-view virtual mesh.
[0020] It should be understood that the visible 3D mesh point cloud to video frame rendering algorithm is an existing technology that converts dynamic 3D point cloud data into continuous video frames. Its core lies in efficiently utilizing mesh visibility information to optimize the rendering process.
[0021] (2) The main viewpoint of the two-dimensional video frame is used as the input of the visual large model SAM (Segment Anything Model, abbreviated as SAM) to obtain the initial segmentation mask of each component. The prompt box generated by each segmentation mask is used together with the two-dimensional video frame as the input of the SAM2 (Segment Anything in Images and Videos, abbreviated as SAM2) model to obtain the segmentation mask of each component under each viewpoint.
[0022] It should be understood that SAM is a general-purpose image segmentation model designed to achieve fast and flexible zero-shot image segmentation tasks. It is a fundamental model in the field of computer vision, capable of segmenting any object in an image through simple user interaction without requiring training for a specific task. SAM2 is an extension of SAM in the video domain, aiming to transfer SAM's powerful zero-shot image segmentation capabilities to video, enabling frame-by-frame or cross-frame segmentation of objects in videos. This approach combines SAM's general segmentation capabilities with video temporal information.
[0023] Specifically, the visual large model (SAM) is used to analyze the baseline view. That is, global segmentation is performed on the main view of the 2D video frame (i.e., the 2D video frame with a 0° view) to obtain the initial segmentation mask for each component of the bridge. This includes a background mask that does not belong to any component of the bridge. Because the image is rendered based on point clouds, the background is a single black. Therefore, the background mask can be removed based on the RGB values of the pixels in the mask, and finally the initial segmentation mask of each component of the bridge after the background mask is removed is obtained. This represents the initial segmentation mask for the first component. This represents the initial segmentation mask for the second component. This represents the initial segmentation mask for the nth component. For each of the bridge components' initial segmentation masks, the pixel coordinates of the top-left and bottom-right corners are selected to form a coordinate pair. ,in This represents the coordinate pair of the first component. This represents the coordinate pair of the second component. This represents the coordinate pair of the nth component. and These represent the pixel coordinates of the top-left and bottom-right corners, respectively. Using the coordinates obtained above... A prompt box is generated; furthermore, to balance SAM2's segmentation performance and computation speed, the minimum number of rendered images is used for segmentation. The rendered 2D video frame images from each viewpoint are then sorted according to their rotation angles. The video frames are divided into two groups clockwise and synthesized separately. These frames, along with the prompt box, are then input into the SAM2 model for segmentation, resulting in... Segmentation mask of each component from different perspectives The viewing angles were selected as 0°, 15°, 165°, 180°, 195°, and 345°. The final segmentation effect is as follows. Figure 3 As shown.
[0024] In other embodiments, new cue points can be generated based on the segmentation mask and input into the SAM model for further fine segmentation. This is based on the pixel coordinates corresponding to... The index information of the points contained in the array can determine the visible points under each mask. and invisible points .
[0025] (3) Project the segmentation results obtained in step (2) under the diagonal view, match the segmentation results of the same components on both sides based on the nearest pairing algorithm of the cluster center of the segmentation components, calculate the score of the index point according to the segmentation mask of each view, and realize the component-level segmentation of the bridge point cloud.
[0026] In this embodiment, the segmentation results from the diagonal viewpoint are projected, and the segmentation results of identical components on both sides are matched based on the nearest neighbor pairing algorithm of the cluster centers of the segmented components. Specifically, this includes: firstly, extracting all visible points from the 0° and 180° viewpoints and projecting them onto the xOz plane, and calculating the centers of the visible points corresponding to different masks after projection from the 0° and 180° viewpoints. and ,in The center of each component mask corresponding to the visible point at a 0° viewing angle. The center of each visible point corresponding to the mask of each component is defined at a 180° viewpoint. Then, for each center at a 0° viewpoint, the coordinates of the nearest center at a 180° viewpoint are found and paired. The center coordinate indices at a 180° viewpoint are modified based on the category index of the mask at a 0° viewpoint. This allows the segmentation results from the two sets of images to be linked, with the matching effect as shown below. Figure 4 As shown. Then, the category scores are statistically analyzed to integrate the segmentation results from different perspectives, for the mask... The scoring formula for each point p in the corresponding bridge point cloud is: In the formula, Let p be the final score for class n, and k be the viewpoint. Quantity, From the perspective Score the point at that point. From the perspective The point below, The representative point is the visible point. The representative point is an invisible point. Using the above formula, the score matrix H for each category of the bridge point cloud can be calculated. Selecting the category with the highest score as the category of that point achieves component-level segmentation of the bridge point cloud.
[0027] (4) Optimize the component-level segmentation results obtained in step (3) based on the error point iterative repair algorithm, such as... Figure 5 As shown.
[0028] It should be noted that although the component-level segmentation results of the bridge point cloud obtained in step (3) have high segmentation accuracy, there will still be some incorrectly segmented points. These points can be divided into three categories: points that are not included in any mask and therefore are not classified. Points with the same scores in multiple categories A few points of incorrect segmentation at the edges of bridge components ,like Figure 5 As shown in (a), the misclassified Class I and Class II points are represented in black. and It can be easily obtained from the score matrix H, and by observation, it can be found that the vast majority of... The point cloud will be in a state detached from the main category. DBSCAN clustering can be performed on the point cloud of each category, and all smaller clusters except the largest cluster can be removed from that category to obtain... Therefore, in this embodiment, the segmentation result is optimized based on an iterative error point repair algorithm, and the optimized segmentation result is shown in the figure below. Figure 5 As shown in (b) of the diagram.
[0029] Specifically, based on the component-level segmentation results obtained in step (3), for the three types of incorrectly segmented points... , , Under the normal vector threshold constraint, based on the radius interval Iteratively calculate the category of the most frequent point within the radius r of each point, from smallest to largest, and use that category as the category of the incorrectly segmented point; where and These represent the minimum radius at the start of the iteration and the maximum radius at the end of the iteration, respectively.
[0030] It should be understood that the normal vector of the plane is obtained by fitting the plane using points within a certain neighborhood radius of each erroneous point. Normal vector similarity indicates that two points are more likely to be on the same plane, or more likely to belong to the same bridge component. The normal vector threshold constraint means that the subsequent radius interval iterative optimization process will only proceed if the normal vectors of the erroneous point are similar to those of its surrounding points.
[0031] (5) Based on the optimized component-level segmentation results obtained in step (4), pixel points are extracted using the segmentation mask coordinates of each viewpoint, and blank pixels around the perimeter are added to generate component images. The cosine similarity between each component image and prior knowledge text information is calculated based on the multimodal large language model CLIP (Contrastive Language–Image Pre-training, abbreviated as CLIP) to achieve the classification and recognition of bridge components.
[0032] Specifically, based on the optimized component-level segmentation results obtained in step (4), Segmentation mask of each component from different perspectives Image groups are generated by adding blank pixels of a specified width around the perimeter. As the image input to the multimodal large language model CLIP, m text sentences containing semantic and geometric feature information of the bridge structure are used as CLIP's text input. CLIP obtains the corresponding probability by calculating the cosine similarity between each component image and the prior knowledge text information, and outputs the probability that each mask image belongs to each sentence. Finally, the prediction probabilities of each segmentation mask M under all views are summed to obtain the final score probabilities of all masks. The text information corresponding to the highest score probability of each segmentation mask contains the component category of that mask, and the final classification result is... The specific formula is as follows: In the formula, It is the category corresponding to component n. It is the x-th text vector. It is the mask image vector of component n in the i-th view, and k is the viewpoint. The total number of corresponding views. Through this process, the text information corresponding to the segmentation mask of each component can be determined. Based on the association between the mask and the points in the point cloud, semantic recognition of the bridge component in the point cloud can be achieved.
[0033] In summary, this invention only requires the collection and analysis of bridge point cloud data to quickly achieve component-level segmentation of the bridge point cloud, and enables rapid classification and identification of components through the segmentation results, thereby achieving fast and accurate instance segmentation. This invention analyzes and processes bridge point clouds obtained by laser scanners through a series of methods such as video frame rendering, image segmentation, and cosine similarity calculation, enabling accurate and stable instance segmentation of bridges and solving the problems of difficulty and inefficiency in manual segmentation.
[0034] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automated segmentation identification of components suitable for bridge point clouds, characterized in that, The method comprises the following steps: (1) obtaining the main direction of the bridge point cloud by using the principal component analysis method, and performing coordinate transformation to obtain the calibrated bridge point cloud; The calibrated bridge point cloud is divided into three-dimensional grids based on spatial position information, and the index information of the points in each grid is saved. Based on the visibility principle when the camera is shooting, the multi-view visible three-dimensional grid is extracted, and it is rendered into a two-dimensional video frame to realize the dimension reduction processing of the three-dimensional point cloud; (2) taking the main view angle of the two-dimensional video frame as the input of the visual large model SAM to obtain the initial segmentation mask of each component, and taking the prompt box generated from each segmentation mask together with the two-dimensional video frame as the input of the SAM2 model to obtain the segmentation mask of each component under each view angle; (3) projecting the segmentation result under the diagonal view angle obtained in step (2), matching the segmentation results of the same components on both sides based on the nearest pairing algorithm of the clustering center of the segmented components, and calculating the score of the index point according to the segmentation mask under each view angle to realize the component-level segmentation of the bridge point cloud; (4) optimizing the component-level segmentation result obtained in step (3) based on the error point iterative repair algorithm; (5) based on the optimized component-level segmentation result obtained in step (4), extracting the pixel points using the segmentation mask coordinates of each view angle, adding the surrounding blank pixels to generate the component image, and calculating the cosine similarity between each component image and the prior knowledge text information based on the multi-modal large language model CLIP to realize the classification and identification of the bridge components.
2. The method for automated segmentation identification of components suitable for bridge point clouds according to claim 1, characterized in that, In step (1), the principal component analysis method is used to obtain the main direction of the bridge point cloud, and coordinate transformation is performed to obtain the calibrated bridge point cloud, which specifically comprises: The bridge point cloud is formally represented as a three-dimensional point cloud coordinate matrix P of n rows and 3 columns. For the coordinate matrix P, the point cloud coordinate data is first decentralized, then the covariance matrix of the coordinate matrix is calculated, the covariance matrix is singular value decomposed to obtain three eigenvalues and their corresponding eigenvectors, the eigenvector corresponding to the maximum eigenvalue is taken as the main direction of the three-dimensional point cloud, the eigenvector corresponding to the minimum eigenvalue is taken as the normal vector of the three-dimensional point cloud, and the eigenvector corresponding to the other eigenvalue is taken as the secondary main direction of the three-dimensional point cloud. The main direction is set as the x-axis, the secondary main direction is set as the y-axis, and the normal vector is set as the z-axis. The three-dimensional point cloud of the bridge is rotated based on the set x-axis, y-axis and z-axis, and coordinate transformation is performed to obtain the calibrated bridge point cloud.
3. The method for automated segmentation identification of components suitable for bridge point clouds of claim 1, wherein, In step (1), the two-dimensional video frame is rendered, specifically including: The point cloud of the visible three-dimensional grid is rendered into a two-dimensional video frame using a three-dimensional grid point cloud to video frame rendering algorithm to generate video frame images under each view angle, and the calculation formula is: wherein, represents the i-th point in the point cloud the angle between the reference view angle and the view angle of the i-th point in the point cloud the corresponding three-dimensional grid index under the view angle of , , and respectively represent the coordinates of the i-th point in the point cloud on the x-axis, y-axis and z-axis, , and respectively represent the indices of the three-dimensional grid on the x-axis, y-axis and z-axis; is the size of the three-dimensional grid; is the rotation matrix, ; is the point cloud coordinate, ; represents the offset vector, , is the point cloud center coordinate, is the minimum value of the point cloud in the three coordinate directions; correspondingly, the minimum value of each group , in the two-dimensional video frame The grid corresponding to the grid distance closest to the virtual camera point is the visible point grid, and the remaining grids are the invisible point grids. The RGB average value of all points in the visible point grid is calculated as the pixel value of the grid.
4. The method for automated segmentation identification of components suitable for bridge point clouds of claim 1, wherein, Step (2) specifically includes: The benchmark view, i.e., the main view of the two-dimensional video frame, is globally segmented by using a visual large model SAM to obtain initial segmentation masks of each component of the bridge, and the initial segmentation masks of each component of the bridge are obtained by removing the background mask according to the pixel RGB value in the mask; the upper left corner and the lower right corner pixel coordinates of the initial segmentation mask of each component are selected to form a coordinate pair, and a prompt box is generated by using the coordinate pair; the two-dimensional video frame images under each view angle rendered are divided into two groups clockwise at an angle of 90° The two groups are synthesized into video frames, and the video frames are input into a SAM2 model together with the prompt box for segmentation to obtain the component segmentation mask under the view angle, wherein The view angle is 0°, 15°, 165°, 180°, 195° or 345°.
5. The method for automated segmentation identification of components suitable for bridge point clouds of claim 1, wherein, Step (3) specifically includes: Firstly, all visible points under 0° and 180° view angles are extracted and projected to the xOz plane, and the centers of the visible points corresponding to different masks after projection under 0° and 180° view angles are calculated, wherein the visible points are determined according to the index information of the points contained in the pixel coordinates; then for each center under 0° view angle, the nearest center coordinate under 180° view angle is found and paired, and the center coordinate index under 180° view angle is modified based on the category index of the mask under 0° view angle; then the category score is counted to integrate the segmentation results under different view angles, for each point in the bridge point cloud corresponding to the mask, the score matrix of each point to each category is calculated, and the category with the highest score is selected as the category of the point, so as to realize the component-level segmentation of the bridge point cloud.
6. The method for automated segmentation identification of components suitable for bridge point clouds of claim 1, wherein, The step (4) specifically comprises: Based on the component-level segmentation results obtained in step (3), for the incorrectly segmented points, under the normal vector threshold constraint, according to the radius interval... Iteratively calculate the category of the most frequent point within the radius r of each point, from smallest to largest, and use that category as the category of the incorrectly segmented point; where and These represent the minimum radius at the start of the iteration and the maximum radius at the end of the iteration, respectively.
7. The method for automated segmentation identification of components suitable for bridge point clouds of claim 1, wherein, The step (5) specifically comprises: Based on the optimized component-level segmentation result obtained in step (4), the following steps are performed: An image group is generated by adding a specified width of blank pixel points to the periphery of each component segmentation mask under the perspective, which is used as the image input of the multi-modal large language model CLIP. m text statements containing semantic information and geometric feature information of the bridge structure are used as the text input of the CLIP. The CLIP calculates the cosine similarity between each component image and the prior knowledge text information to obtain the corresponding probability, and outputs the probability that each mask image belongs to each statement. Finally, the prediction probabilities of each segmentation mask under all perspectives are summed to obtain the final score probability of all masks. The text information corresponding to the highest score probability of each segmentation mask contains the component category of the mask, and the final classification result is The calculation formula is: In the formula, is the category corresponding to component n, is the xth text vector, is the mask image vector of component n in the ith view, and k is the view angle corresponding to the total number of views; through this process, the text information corresponding to each component segmentation mask is determined, and the classification recognition of the bridge component is realized.
Citation Information
Patent Citations
Bridge member identification method based on unmanned aerial vehicle point cloud reconstruction and three-dimensional synthetic data
CN121074718A
Method and system for processing point-cloud data
US20230419659A1
Single building three-dimensional reconstruction method based on point cloud semantic segmentation and structure fitting
WO2024077812A1
KR20250080971A