Multi-modal fusion bridge disease detection and three-dimensional point cloud registration method

Through a multimodal fusion bridge defect detection method, combined with an encoder and decoder framework, bridge semantic segmentation and three-dimensional point cloud registration are performed, which solves the accuracy problems of bridge defect detection and quantitative calculation, and realizes accurate assessment and rapid processing of defect information.

CN120672668AActive Publication Date: 2025-09-19FUJIAN TRANSPORTATION RES INST CO LTD +1

Patent Information

Application Number
CN202510687355.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-19
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing technologies for bridge defect detection have problems with similarity of defect types and interference factors, point cloud data quality issues, and multimodal data fusion difficulties, resulting in insufficient accuracy in defect detection and quantitative calculation, making it difficult to achieve accurate defect information assessment.

Method used

A multimodal fusion bridge defect detection method is adopted. By constructing a bridge defect dataset, the encoder is combined with spatial information mining, the adaptive pyramid context module and the lightweight decoder to perform bridge semantic segmentation, and the group matching module and mask matching module are used for 3D point cloud registration to ensure local and scene semantic consistency.

Benefits of technology

It improves the accuracy of bridge defect detection and 3D point cloud registration, realizes the precise calculation and evaluation of defect information, and adapts to the needs of rapid processing of large-scale bridge data in actual engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672668A_ABST
    Figure CN120672668A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of bridge health monitoring, and discloses a multi-modal fusion bridge disease detection and three-dimensional point cloud registration method, which specifically comprises the following steps: S1, constructing a bridge disease data set by adopting four public data sets CFD, CrackTree200, Crack500 and CrackSeg9k and autonomously acquired data samples; a bridge disease semantic segmentation result is introduced as a semantic clue, a multi-level semantic consistency registration frame is constructed, intra-class mismatching is inhibited in combination with a scene semantic consistency mask matching module, and the three-dimensional point cloud registration precision is improved; a multi-radius annular region semantic feature extraction and clustering optimization strategy is adopted, the feature integrity of sparse point cloud data is enhanced, RGB-D multi-scale features are fused based on a lightweight encoder-decoder framework, accurate segmentation of complex diseases is realized through a double-branch attention module and an adaptive pyramid context module, and the accuracy of segmentation of the complex diseases is improved. And semantic information is deeply fused with the three-dimensional point cloud to form closed-loop feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of beam health monitoring, and specifically provides a multimodal fusion bridge disease detection and three-dimensional point cloud registration method. Background Art

[0002] A bridge is a building structure that spans rivers, canyons, roads or other obstacles, providing passageways for vehicles and pedestrians. As its service life increases, the bridge gradually develops defects such as concrete cracks, surface slurry peeling, and exposed rebar cavities due to multiple factors such as concrete material aging, construction defects, and long-term alternating loads. These defects not only cause damage to the bridge structure, but also significantly reduce its bearing capacity and durability.

[0003] At present, the research on bridge defect detection based on machine vision mainly focuses on a single task, such as crack detection or crack segmentation. However, in practical applications, bridge defect detection not only needs to identify the type of defect, but also needs to perform quantitative calculation of the defect parameters, which involves the construction and analysis of three-dimensional point cloud data. At present, the difficulties of bridge image defect detection and three-dimensional point cloud data registration are mainly reflected in the following aspects: similarity of defect types and interference factors, point cloud data quality issues and multimodal data fusion problems; existing technical solutions are mainly aimed at solving the problem of defect detection, and have failed to effectively solve the problems in defect detection and subsequent quantitative calculation, such as three-dimensional point cloud registration problems, which makes the accurate detection and parameterized calculation of defects in practical applications challenging. Specifically, the shortcomings of existing technical solutions are reflected in the following aspects: separation of defect detection and quantitative calculation, Some technical solutions mainly focus on defect detection and fail to effectively integrate defect detection and quantitative calculation, resulting in difficulty in achieving accurate calculation and evaluation of defect information in practical applications; the problem of three-dimensional point cloud registration: due to the sparsity, irregularity and noise interference of bridge point cloud data, existing technical solutions have great difficulties in three-dimensional point cloud registration, which affects the accuracy of defect detection and quantitative calculation; insufficient multimodal data fusion: existing technical solutions have deficiencies in multimodal data fusion and fail to fully utilize the complementarity of image semantic information and three-dimensional point cloud data, resulting in limited accuracy of defect detection and quantitative calculation; in summary, existing technical solutions have obvious deficiencies in bridge defect detection and quantitative calculation. Therefore, it is urgently necessary to carry out research on multimodal fusion bridge defect detection and three-dimensional point cloud registration to achieve accurate calculation and evaluation of bridge defect information. Summary of the Invention

[0004] The purpose of the present invention is to provide a multimodal fusion bridge defect detection and three-dimensional point cloud registration method to solve the problems raised in the above background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a multimodal fusion bridge defect detection and 3D point cloud registration method, the specific steps of which are as follows: S1. Constructing a bridge disease dataset The bridge dataset was constructed using four public datasets: CFD, CrackTree200, Crack500, and CrackSeg9k, as well as self-collected data samples. The Crack500 dataset consists of 500 high-resolution bridge surface images with a resolution of 2272×1704 pixels, collected from bridge structures in multiple regions around the world, covering a wide range of crack morphologies and size characteristics. The CrackSeg9k dataset was constructed by integrating multiple small open source datasets, showing significant diversity in surface materials, environmental backgrounds, lighting conditions, exposure parameters, and crack morphologies. S2, divide the bridge disease dataset constructed in S1 into a training set and a test set; S3. Bridge Defect Detection S3.1. Using encoders to mine spatial information represents the input feature map, where 、 、 Represent the number of channels, height and width respectively; firstly, global average pooling is used to reduce the spatial dimension of the feature map to obtain the first vectors :

[0006] The spatial pyramid attention is introduced to transform features of different scales, where and Represent the concatenation layer and pooling operation respectively; Indicates that the kernel size is The vector size adjustment operation performed by the standard convolution layer of , the attention pyramid is represented as:

[0007] Using ReLU function and Sigmoid function As the activation layer, the final fusion feature map is expressed as:

[0008] in, and represent channel branches and space branches respectively; Represents element-wise multiplication operation; S3.2. Adaptive Pyramid Context Module A globally guided local affinity adaptive semantic module is introduced to integrate information at different scales. Contextual information with a larger receptive field helps understand shared features in adjacent local regions of different damage categories, improving bridge damage segmentation performance. To further reduce computational complexity, nearest neighbor upsampling is used. S3.3 Lightweight decoder The dual-branch design of the lightweight decoder uses asymmetric convolution , the convolution kernel sizes are 1x3 and 3x1, where Indicates the number of times used; in order to supplement the continuity of information and obtain close-range feature information, branch Using depth-wise separable convolution, where Indicates the number of times the feature map passes through asymmetric convolution; the middle branch based on The branch uses dilated convolution, which is expressed as follows:

[0009] in, For the decoder feature maps, assuming ; The decoder module integrates features from long-range and short-range:

[0010] in, express The weights of the convolutional layers, Represents the output of the middle branch; after the summation operation on these branches, the channel shuffle operation is performed Realizes feature interaction between branches; S3.4. Adaptive Pyramid Context Module To effectively integrate features from relevant image regions and their corresponding pixel semantic labels, an adaptive pyramid context module is introduced to capture multi-scale contextual details and improve segmentation performance. S4. After the bridge defect detection is completed, the bridge semantic segmentation results are used as semantic information to perform bridge 3D point cloud data registration; S4.1. Objective function definition Given two point clouds (source point cloud) and (target point cloud); input includes 、 、 、 and ,in and From and Randomly select key points from represents the label mapping function, where is a set of semantic labels; S4.2. Group matching module To solve the mismatch problem between disease categories, local semantic consistency is expressed as follows:

[0011] in, Indicates the search radius The local semantic features under the indicator function For judging points and Whether local semantic consistency is satisfied; The mismatch problem between disease categories is solved by avoiding matching key points that do not satisfy local semantic consistency. First, the matching group is defined as a combination of two disease key point sets:

[0012] in, The disease type in the source key point is The sub-keypoint set of is the set of diseased sub-keypoints in the target keypoint; S4.3 Mask Matching Constructed using the Euclidean clustering method with different radii The clustering of semantic instances is then defined as the set of cluster centers:

[0013] in, It is The center point of the cluster; Then, with the key point as the center, the bridge is divided into of annular area, of which the The radius of the annular area is , and then fill in the binary multi-ring semantic features by :

[0014] in, Query A circular area surrounding Disease label set; The mask matching module formalizes the feature matching method as a selection process, using the matching score matrix as input:

[0015] in Represents the matching score matrix based on the similarity of local descriptors; function Used to select matching pairs with high matching scores, output ,in express is a potential matching inlier; subsequently, by using a refined scene consistency mask Reweight the matching score matrix to effectively suppress intra-class mismatching; To build , first calculate the similarity matrix ,in Indicates key points and Similarity between, then, using the binary selection function The matrix The first line of each The maximum value is set to 1 and the rest are set to 0. This process generates a scene consistency mask. ,in express and Satisfy scene semantic consistency. Finally, based on the masked matching score matrix ,Match selection can be expressed as an extension of the formula; S5. Utilize the constructed three-dimensional point cloud model with complete information to perform quantitative calculation of subsequent disease parameters; S6. Use the test set data to test the bridge defect detection method and the 3D point cloud registration method.

[0016] Preferably, the matching group is defined as a combination of two disease key point sets as described in S3.2, and the matching group The key points in satisfy local semantic consistency because their local semantic features all contain On the contrary, if two key points do not exist in any matching group, they cannot be matched because they do not satisfy local semantic consistency. Therefore, it is only necessary to perform key point matching within each matching group and then merge the results to obtain a set of final correspondences without inter-class mismatches:

[0017] in, Represents a matching function.

[0018] Preferably, if the category of the binary multi-ring semantic feature in S4.3 is The disease is located in No. In a circular area, .

[0019] Preferably, when constructing the bridge defect dataset as described in S1, a multimodal data fusion strategy is adopted to align visible light images, infrared thermal imaging data and laser point cloud data in time series, and unify the data resolution through standardization processing to ensure the spatiotemporal consistency of multi-source data.

[0020] Preferably, the encoder described in S3.1 adopts a residual connection structure to fuse shallow features with deep features through skip connections to enhance the ability to express spatial information, wherein the number of residual blocks is dynamically adjusted according to the scale of the input data.

[0021] Preferably, the lightweight decoder described in S3.3 realizes multi-scale feature interaction through channel shuffling operation, divides the feature map into several subgroups along the channel dimension, and randomly permutes and reorganizes the subgroups to improve feature reuse efficiency.

[0022] Preferably, as described in S4.3, mask matching is applied to each group generated by the group matching module to generate a set of high-quality correspondences that simultaneously meet local consistency and scene semantic consistency.

[0023] Preferably, the bridge disease dataset in S2 is divided into a training set and a test set in a ratio of 7:3.

[0024] Preferably, the test method described in S6 uses intersection-over-union (IoU) and root mean square error (RMSE) as evaluation indicators, wherein IoU is used to quantify the accuracy of defect segmentation, and RMSE is used to evaluate the geometric error of point cloud registration.

[0025] The beneficial effects of the present invention are as follows: By introducing the semantic segmentation results of bridge defects as semantic clues, a multi-level semantic consistency registration framework is constructed, and the scene semantic consistency mask matching module is combined to suppress intra-class mismatching and improve the accuracy of 3D point cloud registration. A multi-radius annular area semantic feature extraction and clustering optimization strategy is adopted to enhance the feature integrity of sparse point cloud data. At the same time, based on a lightweight encoder-decoder framework, RGB-D multi-scale features are fused, and the precise segmentation of complex defects is achieved through a dual-branch attention module and an adaptive pyramid context module. The semantic information is deeply integrated with the 3D point cloud to form a closed-loop feedback. The decoder adopts a lightweight residual unit and asymmetric convolution design to reduce redundant calculations and optimize the registration process. It improves computational efficiency while ensuring millimeter-level detection accuracy, adapts to the rapid processing requirements of large-scale bridge data in actual engineering, and provides reliable support for quantitative disease assessment and health monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0028] like Figure 1 As shown, the embodiment of the present invention provides a multimodal fusion bridge defect detection and three-dimensional point cloud registration method, the specific steps are as follows: S1. Constructing a bridge disease dataset The bridge dataset was constructed using four public datasets: CFD, CrackTree200, Crack500, and CrackSeg9k, as well as self-collected data samples. The Crack500 dataset consists of 500 high-resolution bridge surface images with a resolution of 2272×1704 pixels, collected from bridge structures in multiple regions around the world, covering a wide range of crack morphologies and size characteristics. The CrackSeg9k dataset was constructed by integrating multiple small open source datasets, showing significant diversity in surface materials, environmental backgrounds, lighting conditions, exposure parameters, and crack morphologies. S2, divide the bridge disease dataset constructed in S1 into a training set and a test set; S3. Bridge Defect Detection S3.1. Using encoders to mine spatial information represents the input feature map, where 、 、 Represent the number of channels, height and width respectively; firstly, global average pooling is used to reduce the spatial dimension of the feature map to obtain the first vectors :

[0029] The spatial pyramid attention is introduced to transform features of different scales, where and Represent the concatenation layer and pooling operation respectively; Indicates that the kernel size is The vector size adjustment operation performed by the standard convolution layer of , the attention pyramid is represented as:

[0030] Using ReLU function and Sigmoid function As the activation layer, the final fusion feature map is expressed as:

[0031] in, and represent channel branches and space branches respectively; Represents element-wise multiplication operation; S3.2. Adaptive Pyramid Context Module A globally guided local affinity adaptive semantic module is introduced to integrate information at different scales. Contextual information with a larger receptive field helps understand shared features in adjacent local regions of different damage categories, improving bridge damage segmentation performance. To further reduce computational complexity, nearest neighbor upsampling is used. S3.3 Lightweight decoder The dual-branch design of the lightweight decoder uses asymmetric convolution , the convolution kernel sizes are 1x3 and 3x1, where Indicates the number of times used; in order to supplement the continuity of information and obtain close-range feature information, branch Using depth-wise separable convolution, where Indicates the number of times the feature map passes through asymmetric convolution; the middle branch based on The branch uses dilated convolution, which is expressed as follows:

[0032] in, For the decoder feature maps, assuming ; The decoder module integrates features from long-range and short-range:

[0033] in, express The weights of the convolutional layers, Represents the output of the middle branch; after the summation operation on these branches, the channel shuffle operation is performed Realizes feature interaction between branches; S3.4. Adaptive Pyramid Context Module To effectively integrate features from relevant image regions and their corresponding pixel semantic labels, an adaptive pyramid context module is introduced to capture multi-scale contextual details and improve segmentation performance. S4. After the bridge defect detection is completed, the bridge semantic segmentation results are used as semantic information to perform bridge 3D point cloud data registration; S4.1. Objective function definition Given two point clouds (source point cloud) and (target point cloud); input includes 、 、 、 and ,in and From and Randomly select key points from represents the label mapping function, where is a set of semantic labels; S4.2. Group matching module To solve the mismatch problem between disease categories, local semantic consistency is expressed as follows:

[0034] in, Indicates the search radius The local semantic features under the indicator function For judging points and Whether local semantic consistency is satisfied; The mismatch problem between disease categories is solved by avoiding matching key points that do not satisfy local semantic consistency. First, the matching group is defined as a combination of two disease key point sets:

[0035] in, The disease type in the source key point is The sub-keypoint set of is the set of diseased sub-keypoints in the target keypoint; S4.3 Mask Matching Constructed using the Euclidean clustering method with different radii The clustering of semantic instances is then defined as the set of cluster centers:

[0036] in, It is The center point of the cluster; Then, with the key point as the center, the bridge is divided into of annular area, of which the The radius of the annular area is , and then fill in the binary multi-ring semantic features by :

[0037] in, Query A circular area surrounding Disease label set; The mask matching module formalizes the feature matching method as a selection process, using the matching score matrix as input:

[0038] in Represents the matching score matrix based on the similarity of local descriptors; function Used to select matching pairs with high matching scores, output ,in express is a potential matching inlier; subsequently, by using a refined scene consistency mask Reweight the matching score matrix to effectively suppress intra-class mismatching; To build , first calculate the similarity matrix ,in Indicates key points and Similarity between, then, using the binary selection function The matrix The first line of each The maximum value is set to 1 and the rest are set to 0. This process generates a scene consistency mask. ,in express and Satisfy scene semantic consistency. Finally, based on the masked matching score matrix ,Match selection can be expressed as an extension of the formula; S5. Utilize the constructed three-dimensional point cloud model with complete information to perform quantitative calculation of subsequent disease parameters; S6. Use the test set data to test the bridge defect detection method and the 3D point cloud registration method.

[0039] By designing a lightweight encoder-decoder framework to explore the correlation and complementary clues between RGB images and depth images, the performance of RGB-D bridge disease semantic segmentation is improved. A dual-branch attention fusion module is used to learn deep feature expression and cross-modal information, enhancing the model's understanding of multimodal data. An adaptive pyramid context module is used to extract multi-level contextual information, improving the model's ability to handle complex diseases. The semantic segmentation results of bridge diseases are introduced as semantic clues. A multi-level semantic consistency registration module is used to solve the three-dimensional point cloud registration problem, optimizing the accuracy of point cloud registration and reducing dependence on geometric information. Finally, a mask matching module based on scene semantic consistency is designed to suppress intra-class mismatching problems and improve the reliability of registration.

[0040] Among them, S3.2 defines a matching group as a combination of two disease key point sets. The key points in satisfy local semantic consistency because their local semantic features all contain On the contrary, if two key points do not exist in any matching group, they cannot be matched because they do not satisfy local semantic consistency. Therefore, it is only necessary to perform key point matching within each matching group and then merge the results to obtain a set of final correspondences without inter-class mismatches:

[0041] in, Represents a matching function.

[0042] By defining a matching group of defect key points that satisfies local semantic consistency, the mismatching problem between different defect categories is effectively avoided, the accuracy and reliability of 3D point cloud registration are improved, and computational redundancy is reduced.

[0043] Among them, in the binary multi-ring semantic feature in S4.3, if the category is The disease is located in No. In a circular area, .

[0044] By extracting the distribution features of disease categories within the annular area, the semantic expression ability of complex spatially distributed diseases is enhanced, and the ability to judge the semantic consistency of the scene is improved.

[0045] Among them, when constructing the bridge disease dataset in S1, a multimodal data fusion strategy is adopted to align visible light images, infrared thermal imaging data and laser point cloud data in time series, and unify the data resolution through standardization processing to ensure the spatiotemporal consistency of multi-source data.

[0046] The adoption of multimodal data fusion strategies and ensuring spatiotemporal consistency solves the fusion deviation problem caused by resolution or time series differences in multi-source data, providing a highly consistent data foundation for subsequent disease detection and registration.

[0047] Among them, the encoder in S3.1 adopts a residual connection structure, which fuses shallow features with deep features through jump connections to enhance the ability to express spatial information. The number of residual blocks is dynamically adjusted according to the scale of the input data.

[0048] A dynamic residual connection structure is introduced into the encoder, which enhances the expression ability of spatial information through the jump fusion of shallow and deep features, and improves the model's ability to capture the detailed features of bridge diseases.

[0049] Among them, the lightweight decoder in S3.3 realizes multi-scale feature interaction through channel shuffling operation, divides the feature map into several subgroups along the channel dimension, and randomly permutes and reorganizes the subgroups to improve feature reuse efficiency.

[0050] Multi-scale feature interaction is achieved through channel shuffling operations, which enhances the efficiency of feature reuse and ensures the effective fusion of multi-level semantic information while reducing the computational complexity of the decoder.

[0051] Among them, S4.3 generates a set of high-quality correspondences by applying mask matching to each group generated by the group matching module, which satisfies both local consistency and scene semantic consistency.

[0052] Mask matching is applied to each group generated by the group matching module. Through the dual constraints of local consistency and scene semantic consistency, high-confidence correspondences are further screened out, reducing the probability of intra-class mismatching.

[0053] Among them, the bridge disease dataset in S2 is divided into training set and test set in a ratio of 7:3.

[0054] The data set is divided into a ratio of 7:3 to ensure the adequacy of training data and the representativeness of test data, which improves the generalization ability and robustness of the model in actual engineering.

[0055] Among them, the test method in S6 uses intersection-over-union (IoU) and root mean square error (RMSE) as evaluation indicators, where IoU is used to quantify the accuracy of defect segmentation, and RMSE is used to evaluate the geometric error of point cloud registration.

[0056] The intersection-over-union ratio and root mean square error are used as evaluation indicators to quantify the segmentation accuracy and registration geometric error respectively, providing a multi-dimensional performance evaluation standard and a quantifiable scientific basis for technical optimization.

[0057] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0058] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal fusion bridge defect detection and 3D point cloud registration method, characterized by: The specific steps are as follows: S1. Constructing a bridge disease dataset The bridge dataset was constructed using four public datasets: CFD, CrackTree200, Crack500, and CrackSeg9k, as well as self-collected data samples. The Crack500 dataset consists of 500 high-resolution bridge surface images with a resolution of 2272×1704 pixels, collected from bridge structures in multiple regions around the world, covering a wide range of crack morphologies and size characteristics. The CrackSeg9k dataset was constructed by integrating multiple small open source datasets, showing significant diversity in surface materials, environmental backgrounds, lighting conditions, exposure parameters, and crack morphologies. S2, divide the bridge disease dataset constructed in S1 into a training set and a test set; S3. Bridge Defect Detection S3.

1. Using encoders to mine spatial information represents the input feature map, where 、 、 Represent the number of channels, height and width respectively; firstly, global average pooling is used to reduce the spatial dimension of the feature map to obtain the first vectors : ; The spatial pyramid attention is introduced to transform features of different scales, where and Represent the concatenation layer and pooling operation respectively; Indicates that the kernel size is The vector size adjustment operation performed by the standard convolution layer of , the attention pyramid is represented as: ; Using ReLU function and Sigmoid function As the activation layer, the final fusion feature map is expressed as: ; in, and represent channel branches and space branches respectively; Represents element-wise multiplication operation; S3.

2. Adaptive Pyramid Context Module A globally guided local affinity adaptive semantic module is introduced to integrate information at different scales. Contextual information with a larger receptive field helps understand shared features in adjacent local regions of different damage categories, improving bridge damage segmentation performance. To further reduce computational complexity, nearest neighbor upsampling is used. S3.3 Lightweight decoder The dual-branch design of the lightweight decoder uses asymmetric convolution , the convolution kernel sizes are 1x3 and 3x1, where Indicates the number of times used; in order to supplement the continuity of information and obtain close-range feature information, branch Using depth-wise separable convolution, where Indicates the number of times the feature map passes through asymmetric convolution; the middle branch based on The branch uses dilated convolution, which is expressed as follows: ; in, For the decoder feature maps, assuming ; The decoder module integrates features from long-range and short-range: ; in, express The weights of the convolutional layers, Represents the output of the middle branch; after the summation operation on these branches, the channel shuffle operation is performed Realizes feature interaction between branches; S3.

4. Adaptive Pyramid Context Module To effectively integrate features from relevant image regions and their corresponding pixel semantic labels, an adaptive pyramid context module is introduced to capture multi-scale contextual details and improve segmentation performance. S4. After the bridge defect detection is completed, the bridge semantic segmentation results are used as semantic information to perform bridge 3D point cloud data registration; S4.

1. Objective function definition Given two point clouds (source point cloud) and (target point cloud); input includes 、 、 、 and ,in and From and Randomly select key points from represents the label mapping function, where is a set of semantic labels; S4.

2. Group matching module To solve the mismatch problem between disease categories, local semantic consistency is expressed as follows: ; in, Indicates the search radius The local semantic features under the indicator function For judging points and Whether local semantic consistency is satisfied; The mismatch problem between disease categories is solved by avoiding matching key points that do not satisfy local semantic consistency. First, the matching group is defined as a combination of two disease key point sets: ; in, The disease type in the source key point is The sub-keypoint set of is the set of diseased sub-keypoints in the target keypoint; S4.3 Mask Matching Constructed using the Euclidean clustering method with different radii The clustering of semantic instances is then defined as the set of cluster centers: ; in, It is The center point of a cluster; Then, with the key point as the center, the bridge is divided into of annular area, of which the The radius of the annular area is , and then fill in the binary multi-ring semantic features by : ; in, Query Surrounding the annular area Disease label set; The mask matching module formalizes the feature matching method as a selection process, using the matching score matrix as input: ; in Represents the matching score matrix based on the similarity of local descriptors; function Used to select matching pairs with high matching scores, output ,in express is a potential matching inlier; subsequently, by using a refined scene consistency mask Reweight the matching score matrix to effectively suppress intra-class mismatching; To build , first calculate the similarity matrix ,in Indicates key points and Similarity between, then, using the binary selection function The matrix The first line of each The maximum value is set to 1 and the rest are set to 0. This process generates a scene consistency mask. ,in express and Satisfy scene semantic consistency. Finally, based on the masked matching score matrix ,Match selection can be expressed as an extension of the formula; S5. Utilize the constructed three-dimensional point cloud model with complete information to perform quantitative calculation of subsequent disease parameters; S6. Use the test set data to test the bridge defect detection method and the 3D point cloud registration method.

2. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: As described in S3.2, a matching group is defined as a combination of two disease key point sets. The key points in satisfy local semantic consistency because their local semantic features all contain On the contrary, if two key points do not exist in any matching group, they cannot be matched because they do not satisfy local semantic consistency. Therefore, it is only necessary to perform key point matching within each matching group and then merge the results to obtain a set of final correspondences without inter-class mismatches: ; in, Represents a matching function.

3. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: In the binary multi-ring semantic feature described in S4.3, if the category is The disease is located in No. In a circular area, .

4. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: When constructing the bridge defect dataset as described in S1, a multimodal data fusion strategy was adopted to align visible light images, infrared thermal imaging data, and laser point cloud data in time series, and to unify the data resolution through standardization processing to ensure the spatiotemporal consistency of multi-source data.

5. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: The encoder described in S3.1 adopts a residual connection structure to fuse shallow features with deep features through skip connections to enhance the ability to express spatial information. The number of residual blocks is dynamically adjusted according to the scale of the input data.

6. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: The lightweight decoder described in S3.3 realizes multi-scale feature interaction through channel shuffling operation, divides the feature map into several subgroups along the channel dimension, and randomly permutes and reorganizes the subgroups to improve feature reuse efficiency.

7. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: As described in S4.3, mask matching is applied to each group generated by the group matching module to generate a set of high-quality correspondences that satisfy both local consistency and scene semantic consistency.

8. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: The bridge disease dataset described in S2 is divided into a training set and a test set in a ratio of 7:

3.

9. The multimodal fusion bridge defect detection and 3D point cloud registration method according to claim 1 is characterized by: The test method described in S6 uses intersection-over-union (IoU) and root mean square error (RMSE) as evaluation indicators, where IoU is used to quantify the accuracy of defect segmentation and RMSE is used to evaluate the geometric error of point cloud registration.

Citation Information

Patent Citations

  • Bridge component three-dimensional point cloud segmentation method based on multi-view data fusion

    CN117876397A

  • Bridge apparent disease detection method and system based on laser point cloud and visual image fusion

    CN119246518A

  • Bridge structure crack position identification method based on deep learning and computer vision

    CN119295660A

  • Road disease automatic identification method based on sparse self-encoding network

    CN119963927A

  • Image segmentation method and system for pavement disease based on deep learning

    US20210319561A1

Cited By

  • Existing bridge high-precision three-dimensional reconstruction method based on laser point cloud data

    CN121053328A

  • High-precision three-dimensional reconstruction method for existing bridge based on laser point cloud data

    CN121053328B

  • Seismic zone bridge member damage characteristic reproduction method based on three-dimensional mesoscopic model

    CN121280629A

  • Bridge type rapid classification system and method for unmanned aerial vehicle autonomous path planning

    CN121323653A

  • Unmanned aerial vehicle laser and vision fusion inspection method and system for bridge bottom disease detection

    CN121353265A