A method for three-dimensional segmentation of post-fire RC structure burst damage

By integrating 3D point cloud and deep learning into an innovative detection framework, the problems of low efficiency and insufficient 3D information parsing in the detection of concrete structure damage after fire are solved. It achieves efficient and accurate damage segmentation and quantitative analysis, and supports multi-scenario adaptation and real-time detection.

CN120471939BActive Publication Date: 2025-11-18QINGDAO UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510546510.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-11-18
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Existing technologies are inefficient in detecting damage to concrete structures after fires, are prone to causing secondary damage, cannot accurately obtain three-dimensional information and volume parameters, and have insufficient model generalization ability, making it difficult to meet the timeliness and accuracy requirements of large-scale post-disaster assessment.

Method used

An innovative detection framework integrating 3D point cloud and deep learning is adopted. By constructing the KP-FCNN network, multi-scale dilated convolution, cross attention mechanism and linear deformable convolution module are introduced to realize dynamic receptive field expansion and adaptive feature extraction, forming KPConvDef-DC and KP-G-FCNN network architecture, which supports seamless adaptation and real-time detection of multi-source heterogeneous devices.

Benefits of technology

It achieves efficient non-destructive testing, accurate three-dimensional damage segmentation, improves adaptability to multiple scenarios, outputs structured data to support damage quantification analysis, and meets the real-time detection needs in complex fire scene environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471939B_ABST
    Figure CN120471939B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of concrete structure damage detection, and particularly relates to a three-dimensional segmentation method for burst damage of RC structure after fire. Firstly, the types and three-dimensional characteristics of burst damage of RC structure components after fire are determined, and a corresponding three-dimensional point cloud dataset is constructed for network training and verification. Then, based on the KP-FCNN network structure, the KPConv layer is improved and optimized, the detection and segmentation accuracy is improved, the model size is reduced, and the inference time is significantly improved. Under different damage conditions, the best segmentation accuracy of 82.3% is achieved, automatic damage segmentation of burst damage of concrete structure after fire is realized, and technical support is provided for subsequent three-dimensional quantification of damage and deployment to unmanned aerial vehicle systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of concrete structure damage detection technology, specifically to a three-dimensional segmentation method for burst damage of RC structures after a fire. Background Technology

[0002] Damage detection technology for concrete structures is mainly divided into two major technical systems: non-destructive testing (NDT) and destructive testing (DT). In the field of NDT, the hammer-rebound method is a representative technique, which involves striking the surface of a component with a hammer and quantifying the degree of concrete damage based on the rebound value. However, this method has significant limitations: firstly, it requires a high degree of flatness on the tested surface, making it difficult to effectively detect areas of concrete spalling and cracking common after fires; secondly, the point-by-point striking method results in low work efficiency, making it difficult to meet the timeliness requirements of large-scale post-disaster assessments.

[0003] In the field of destructive testing, core drilling is commonly used. This technique involves drilling concrete core samples for laboratory analysis. While it can obtain accurate damage parameters, its inherent drawbacks cannot be ignored. The sampling process not only causes secondary damage to the already damaged structure, but the testing cycle also takes 2-3 weeks, making it difficult to meet the timeliness requirements of post-disaster emergency decision-making. Furthermore, the dust and vibration generated by on-site core drilling operations pose a potential threat to the personal safety of testing personnel, further limiting the widespread application of this technology.

[0004] In recent years, while deep learning-based image detection technology has made progress in the field of intelligent detection, it still faces fundamental limitations in post-fire concrete damage assessment. Existing image segmentation algorithms generally suffer from the following technical bottlenecks: 1) Sensitivity to environmental factors such as lighting conditions and shooting angles, leading to significant fluctuations in detection accuracy; 2) Lack of depth dimension data in two-dimensional images, making it impossible to accurately quantify the volumetric parameters and spatial distribution characteristics of burst damage; 3) Insufficient model generalization ability, resulting in poor adaptability to damage morphologies under different fire scenarios; 4) High computational complexity, hindering real-time detection and large-scale applications. These shortcomings prevent existing technologies from effectively acquiring three-dimensional damage information, thus restricting the development of temperature field inversion analysis and structural safety assessment.

[0005] In summary, research on the segmentation of damage to concrete structures after fire is still insufficient. Therefore, it is necessary to propose a lightweight 3D reconstruction and 3D target segmentation method for fire-corroded concrete structure damage that integrates 3D point cloud and deep learning. This method can be used to detect fire-damaged defects and provide theoretical and data support for subsequent reinforcement and repair. Summary of the Invention

[0006] This invention provides a three-dimensional segmentation method for burst damage in RC structures after a fire, addressing the following existing technical problems: First, low detection efficiency. For example, the hammer-bounce method for detecting damaged components is time-consuming and cumbersome. Second, it easily leads to secondary damage to the damaged components. For example, when using core drilling, the sampling object is usually a severely damaged component, thus the sampling process exacerbates the damage and also poses a threat to the safety of on-site inspection personnel. Finally, insufficient three-dimensional information analysis capability. Image segmentation boundaries are blurry, resulting in an error rate of 15%-20% in identifying burst damage areas; the inability to obtain the three-dimensional morphology and volume parameters of the damage leads to a lack of key input for temperature field inversion calculations; and weak model generalization ability. For example, image-based neural networks have low segmentation accuracy and cannot obtain the three-dimensional shape and volume of the damage. To address these technical pain points, this invention proposes an innovative detection framework integrating three-dimensional point clouds and deep learning, which has the following breakthrough advantages: 1. Highly efficient non-destructive testing; 2. Accurate three-dimensional damage segmentation; 3. Intelligent feature extraction adaptable to multiple scenarios.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows:

[0008] A method for three-dimensional segmentation of burst damage in RC structures after a fire includes the following steps:

[0009] Step 1: Construct a 3D point cloud database of RC structure damage after fire;

[0010] Step 2: Use KP-FCNN as the baseline network to construct a 3D point cloud segmentation network for concrete structures;

[0011] Step 3: The performance of the 3D point cloud segmentation network is systematically evaluated using average intersection-union comparison and ablation experiments are conducted.

[0012] Preferably, step 1 includes the following specific methods: The first method is to take photos of concrete burst damage using a regular high-definition camera. During the shooting, care should be taken to take photos of the same damaged component from different angles. To obtain a better 3D reconstruction point cloud effect based on 2D images, it is necessary to ensure that each adjacent pair of photos has at least 50% overlap. Then, all the photos are imported into Colmap software for 3D point cloud reconstruction. The second method utilizes the built-in LiDAR and camera of an iPhone in conjunction with 3D Scanner. The app scans damaged components and automatically generates 3D reconstructed point cloud maps; the third method uses a depth camera to photograph the fire-damaged components and directly obtains point cloud maps; the fourth method uses an R8 LiDAR to scan the fire-damaged components and obtain point cloud maps; in addition, there is a method for synthesizing data, which first models the fire-damaged components using 3D modeling software, then imports the exported .obj file into CloudCompare software for upsampling to obtain high-density 3D point cloud maps; all the datasets obtained by the above methods are imported into CloudCompare software for point cloud segmentation and annotation, after outlining the edge contour of the explosion damage, the label is set to "1", and the background label is set to "0"; after all annotations are completed, a txt file is exported, and the original labeled dataset is expanded using data augmentation techniques, and then the expanded dataset is divided into training, validation and test sets in a 7:2:1 ratio to facilitate subsequent network training.

[0013] Preferably, step 2, the improvement of the baseline network, specifically includes:

[0014] (21) Introducing multi-scale dilated convolution to realize a dynamic, multi-level receptive field expansion mechanism; this mechanism forms a multi-scale feature perception capability by combining convolution kernels with different dilation rates, thereby improving the ability to capture complex geometric structures while maintaining parameter efficiency.

[0015] (22) Introduce a cross-attention mechanism to achieve a dynamic balance between global semantic association and local feature enhancement; by dynamically establishing the semantic association of distant point cloud features, the limitations of the local receptive field are overcome, and an adaptive weight allocation strategy is used to achieve intelligent suppression of noise interference.

[0016] (23) Introduce a linear deformable convolution module. Introduce the linear deformable convolution module into the KPConv layer to realize a dynamic and adaptive kernel adjustment mechanism. Learn dynamic offset for each convolution kernel in a data-driven manner so that it can adaptively adjust its shape and size according to the local geometric features of the input point cloud.

[0017] Preferably, in step 2, the cross-attention module and the dilated convolutional layer are deeply integrated to form a "local-global" collaborative feature enhancement paradigm. That is, the dilated convolution expands the local perception range through sparse sampling, while the cross-attention further aggregates cross-regional contextual information on this basis.

[0018] The beneficial effects of the present invention on the three-dimensional segmentation method for burst damage of RC structures after fire:

[0019] 1. This invention provides a method for expanding a 3D point cloud damage dataset based on a synthetic 3D model: Given that neural network training requires a large amount of data, relying solely on real data is insufficient to meet training needs. Therefore, this invention proposes a method for synthesizing data. First, fire-damaged components are modeled using 3D modeling software (such as SketchUp Pro) to generate virtual damaged components. Then, the exported .obj file is imported into CloudCompare software for upsampling to obtain a high-density 3D point cloud map. This method can fundamentally expand the amount of original data, improving the sufficiency and accuracy of neural network training.

[0020] 2. This invention proposes a dynamic receptive field expansion mechanism based on multi-scale dilated convolution: Addressing the insufficient multi-scale defect capture caused by the fixed receptive field of KPConv, a multi-level dilated convolution collaborative optimization method is proposed. A three-level feature extraction network is constructed by combining convolution kernels with dilation rates of 1, 3, and 5, achieving dynamic perception and cross-scale fusion of local details (dilation rate 1), mesoscale (dilation rate 3), and large-scale (dilation rate 5) features. This mechanism expands the effective receptive field to more than 25 times that of traditional convolution while maintaining parameter efficiency, significantly improving the multi-scale characterization ability of concrete burst damage.

[0021] 3. This invention proposes a cross-attention mechanism and a local-global collaborative optimization paradigm: This invention embeds a cross-attention mechanism into the KPConv network architecture, breaking through the field-of-view limitations of local convolution by dynamically modeling the semantic associations of distant point cloud features. Specifically, it includes: Global semantic association modeling: explicitly establishing the spatial dependency relationship of cross-regional burst damage to capture the overall pattern of damage distribution; Noise suppression and adaptive weight allocation: dynamically selecting key features through attention weights to suppress the interference of scanning noise and sparse regions; Local-global feature collaboration: deeply integrating the expanded receptive field of dilated convolution with the global information aggregation of cross-attention to form a dual optimization mechanism that takes into account both local geometric details and global spatial layout.

[0022] 4. This invention provides a geometrically adaptive kernel adjustment technique for linearly deformable convolution: A data-driven linearly deformable convolution (LDConv) module is proposed, which dynamically learns the kernel offset to achieve adaptive adjustment of the kernel shape and coverage. It achieves the following effects: Geometric deformation-sensitive feature extraction: For irregular deformation of post-disaster concrete (such as spalling boundaries), the kernel sampling position is dynamically adjusted to enhance the representation ability of complex topological structures; Multi-dimensional attention collaborative optimization: Kernel deformation perception is achieved in the spatial dimension, complementing the channel attention mechanism and improving robustness to heterogeneous features such as aggregate distribution and reinforcement differences; Lightweight plug-and-play design: A low-parameter modular structure is adopted, seamlessly integrated into the KPConv framework, supporting efficient training and real-time processing.

[0023] 5. This invention provides a multi-module collaborative optimization 3D point cloud segmentation network architecture: Based on the aforementioned technologies, the KPConvDef-DC and KP-G-FCNN network architectures integrate multi-scale dilated convolution, cross-attention, and linearly deformable convolution modules to form a collaborative technical system of "dynamic receptive field expansion - global semantic association - geometric adaptive optimization." This architecture overcomes the limitations of traditional 3D segmentation models in complex post-disaster scenarios, significantly improving the segmentation accuracy (MIoU up to 82.31%), noise robustness, and multi-scale adaptability of burst damage, providing a highly reliable solution for post-fire concrete structure safety assessment.

[0024] 6. This invention has the following advantages in practical applications: a) Seamless adaptability to multi-source heterogeneous devices: It supports instant access from multi-modal point cloud acquisition terminals such as mobile phones, radar, and depth cameras. Through lightweight model architecture design, it achieves standardized preprocessing and real-time inference of cross-platform image data, meeting the flexible detection needs in complex fire scene environments. b) Dynamic deployment engineering adaptability: It provides multi-level deployment solutions from embedded devices to server clusters, and has developed a dedicated inference acceleration interface specifically for UAV inspection systems, realizing an automated closed-loop operation of aerial point cloud acquisition, damage identification, and result feedback. c) Data foundation for damage quantification analysis: The output point cloud segmentation results are stored in a structured data format, accurately locating key damage features such as the area, depth, and volume of concrete bursting areas, providing high-precision input data for subsequent damage level assessment, structural bearing capacity calculation, and temperature field calculation. Attached Figure Description

[0025] Figure 1 This is a flowchart of the method steps of the present invention.

[0026] Figure 2It is the KP-FCNN network architecture (Note: KP-FCNN: point cloud segmentation network; points: point cloud; features: features; classes: classes; KPConv: Kernel Point Convolution; strided KPConv: strided kernel point convolution; 1conv: 1 convolution; Neur: Ups.+Concat: upsampling + connection in neural network; Skip link: skip link).

[0027] Figure 3 It is a KPConv layer network architecture (Note: KPConv layer; Features dimension; Maxpool; ReLU: Rectified Linear Unit, a widely used activation function in deep learning; BN: (Batch Normalization) is a technique in deep learning, mainly used to accelerate the training process of neural networks and improve the stability of models).

[0028] Figure 4 It is the KPConvDef network architecture (Note: PointConv: point convolution; offset field: refers to the field used to represent the offset of the pixel position where the convolution kernel is applied in the deformable convolution operation; offsets: offset; Deformable convolution: deformable convolution; input feature map: input feature map; out feature map: output feature map).

[0029] Figure 5 It is a Cross-Attention network architecture (Note: Cross-Attention: cross-attention mechanism; Embedding size: embedding dimension; Number of tokens: number of tokens; input: input; output: output; new: new; weights: weights; Attention matrix: attention matrix; Softmax: a classification function).

[0030] Figure 6 It is the KPConvDef-DC network architecture (Note: Bias + Weight: offset + weight; Offset(x,y) and weight coefficients: bias(x,y) and weight coefficients).

[0031] Figure 7It is the LDConv network architecture (Note: Generating Algorithm; Based on the size of N; Generating initial sampled shapes; Original Coordinate; Modified Coordinate; Resample; Initial sampled shapes are adjusted by offsets; Norm; SiLU: a novel activation function).

[0032] Figure 8 It is a KPConv-G network architecture.

[0033] Figure 9 It is the KP-G-FCNN network architecture.

[0034] Figure 10 This is a diagram showing the predicted results.

[0035] Figure 11 This is a Loss value curve.

[0036] Figure 12 This is a curve of Miou values. Detailed Implementation

[0037] The following description provides a detailed explanation of the embodiments of the present invention in a step-by-step manner. This description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0038] In the description of this invention, it should be noted that the terms "upper," "lower," "left," "right," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or a specific orientational structure and operation. Therefore, they should not be construed as limiting this invention.

[0039] Example 1

[0040] A method for three-dimensional segmentation of burst damage in RC structures after a fire includes the following steps:

[0041] Step 1: Construct a 3D point cloud database of RC structure damage after fire;

[0042] Step 2: Use KP-FCNN as the baseline network to construct a 3D point cloud segmentation network for concrete structures;

[0043] Step 3: The performance of the 3D point cloud segmentation network is systematically evaluated using average intersection-union comparison and ablation experiments are conducted.

[0044] Example 2

[0045] Based on Example 1, this example discloses:

[0046] Data on concrete structure bursting damage after a fire has long been scarce in the civil engineering field due to multiple technical bottlenecks and safety risks encountered during its acquisition process. Firstly, the structural safety hazards at high-temperature disaster sites severely limit the frequency of on-site investigations, resulting in an extremely short window for raw data acquisition. Secondly, the spatiotemporal distribution of damage samples exhibits significant heterogeneity, making it difficult for existing detection methods to fully capture three-dimensional damage characteristics. Furthermore, publicly available databases contain highly fragmented data on such special working conditions, with a lack of uniformity in data formats and acquisition standards among different institutions, severely restricting the generalization ability of damage models. To address this technical predicament, this paper summarizes and analyzes… Four methods were proposed to construct a 3D point cloud database of post-fire RC structure damage that can be used to train 3D point cloud segmentation models. Specifically, the first method involves taking photos of concrete burst damage using a regular high-definition camera. During the photography process, it is important to capture the same damaged component from different angles. To obtain good 3D reconstruction point cloud results based on 2D images, it is essential to ensure that every two adjacent photos have at least 50% overlap. All the images are then imported into Colmap software for 3D point cloud reconstruction. The second method utilizes the built-in LiDAR and camera of an iPhone in conjunction with 3D... The Scanner App scans the damaged components and automatically generates a 3D reconstructed point cloud map; Method 3 uses a depth camera to photograph the fire-damaged components and then directly obtains the point cloud map; Method 4 uses an R8 LiDAR to scan the fire-damaged components and then obtains the point cloud map. In addition, since achieving good results in neural network training requires a large amount of data, training the neural network with only real data is insufficient. To fundamentally expand the amount of original data, this invention proposes a method for synthesizing data. First, the fire-damaged components are modeled using 3D modeling software (e.g., SketchUp Pro), and then the exported .obj file is imported into CloudCompare software for upsampling, thereby obtaining a high-density 3D point cloud map. The specific workflows of the five methods are as follows: Figure 1 As shown, the comprehensive application of these methods can provide richer and more accurate data support for the study of post-fire concrete structure bursting damage, and promote technological progress and application development in this field.

[0047] In this embodiment, the number of point cloud models in the constructed concrete structure fire damage point cloud dataset is 1000, including 700 real data and 300 synthetic data. All of them are imported into CloudCompare software for point cloud segmentation and annotation. After outlining the edge contour of the burst damage, the label is set to "1", and the background is labeled to "0". After all the annotations are completed, a txt file is exported. The original annotated dataset is expanded using data augmentation techniques. The expanded dataset is then divided into training, validation and test sets in a 7:2:1 ratio to facilitate subsequent network training.

[0048] Example 3

[0049] Based on Example 2, this example discloses a specific implementation method for step 2:

[0050] 2.1 Regarding the baseline network:

[0051] KP-FCNN is a fully convolutional network specifically designed for segmentation of complex scenes. Its network architecture is as follows: Figure 2 As shown, its core innovation lies in combining the advantages of Kernel Point Convolution (KPConv) and Fully Convolutional Network (FCN). In the task of 3D point cloud segmentation of concrete structure cracking after fire, the use of KP-FCNN as the baseline network has significant necessity and technical advantages, mainly reflected in the following three aspects: First, accurate modeling of complex geometric structures. Post-fire concrete structures often suffer from irregular deformations such as cracking and spalling due to high temperatures, resulting in point cloud data characterized by uneven distribution, strong noise interference, and complex local geometric features. KP-FCNN, by introducing the Kernel Point Convolution (KPConv) module, utilizes learnable kernel points to dynamically capture local geometric structures, effectively adapting to the irregular shapes of cracked concrete fragments. Secondly, it features deep fusion of multi-scale contextual information; semantic classification of fire-damaged areas (such as cracked areas, undamaged areas, and exposed rebar areas) requires combining local details with global context. KP-FCNN's fully convolutional U-shaped structure achieves multi-level feature fusion through skip connections. Finally, it offers flexibility and scalability in 3D point cloud processing; post-fire RC structure point cloud data often suffers from missing data, noise, and scale variations. KP-FCNN's fully convolutional nature allows it to directly process raw point clouds without voxelization or projection preprocessing, avoiding information loss. Furthermore, the network structure possesses excellent scalability.

[0052] KPConv (Kernel Point Convolution) is a point cloud semantic segmentation model based on convolution operations. The KPConv network uses a novel point convolution operation, defining convolution kernel points in Euclidean space and performing convolution operations with points in the point cloud. It also extends deformable convolution to point convolution operations. Its network architecture is as follows: Figure 3 As shown; however, it has the following limitations in the task of segmenting concrete structure bursts:

[0053] (1) Fixed local receptive field, insufficient multi-scale feature capture: KPConvDef (e.g.) Figure 4 The convolutional kernels of the model perform feature aggregation within a fixed radius, which cannot adaptively capture geometric structures of different scales (such as the small cracks and large damaged areas of concrete bursting); for complex, multi-scale bursting defects in concrete blocks (such as differences in crack width and depth), the receptive field of a single scale may miss key details.

[0054] (2) Weak ability to model long-distance dependencies: KPConvDef mainly relies on local neighborhood information and lacks the ability to model the global association of distant point cloud features; for example, the bursting area in a concrete block may span a large spatial range (such as through cracks), and local convolution is difficult to capture these long-distance dependencies.

[0055] (3) Insufficient robustness to noise and sparse regions: In concrete point cloud data, scanning noise, occlusion or sparse sampling regions may lead to incomplete local features; KPConvDef's fixed kernel point strategy is sensitive to noise and cannot effectively fuse multi-scale context information to enhance robustness.

[0056] 2.2 Network Improvements:

[0057] 2.2.1 Introduction of Multi-Scale Dilated Convolution:

[0058] Multi-scale dilated convolution enables a dynamic, multi-level receptive field expansion mechanism. This mechanism combines convolutional kernels with different dilation rates to achieve multi-scale feature perception, thus significantly improving the ability to capture complex geometric structures while maintaining parameter efficiency. Introducing the multi-scale dilated convolution module into KPConvDef breaks its fixed local receptive field limitation, enabling it to adapt to multi-scale features in concrete defects, ranging from micro-bursts to large-area bursts and bursts at different depths. Specific effects are as follows:

[0059] (1) Enhance multi-scale feature fusion capability:

[0060] Multi-level receptive field collaboration: A three-level feature extraction network is constructed using convolutional kernels with dilation rates of 1, 3, and 5. The convolution with a dilation rate of 1 focuses on local details, while the convolutions with dilation rates of 3 and 5 cover medium- and large-scale regions (such as large-area burst surfaces), respectively. The features of the three branches are spliced ​​through channels to achieve cross-scale information fusion and form a complete burst defect representation.

[0061] (2) Parameter efficiency optimization:

[0062] Advantages of dilated convolution: Compared with traditional stacked convolution, the dilated sampling mode of dilated convolution can expand the effective receptive field to more than 5 times the radius without increasing the number of parameters (for example, the coverage area is increased by 25 times when the dilation rate is 5), which allows the model to capture a wider range of contextual information at a lower computational cost.

[0063] (3) Overcoming geometric limitations:

[0064] Dynamically adaptable to defect morphology: By combining multiple levels of expansion rate, the contradiction that the fixed radius kernel of KPConvDef is difficult to capture at the same time as millimeter-level and centimeter-level spalling is resolved, which can effectively improve the accuracy of the model in detecting shallower bursts.

[0065] 2.2.2 Introducing a cross-attention mechanism:

[0066] Cross-attention mechanism (such as...) Figure 5 Introduced into the KPConvDef network architecture, this improved mechanism achieves a dynamic balance between global semantic association and local feature enhancement. By dynamically establishing semantic associations between distant point cloud features, it effectively overcomes the limitations of the local receptive field of traditional convolutional kernels. At the same time, it achieves intelligent suppression of noise interference through an adaptive weight allocation strategy. Specifically, the cross-attention module explicitly models the long-distance dependencies between different regions in the point cloud, enabling the model to capture the global distribution pattern of burst damage in concrete structures after a fire, such as the spatial association between cross-regional burst propagation paths and densely damaged areas.

[0067] At the feature optimization level, this invention deeply integrates the cross-attention module with the dilated convolutional layer to form a "local-global" collaborative feature enhancement paradigm. The dilated convolution expands the local perception range through sparse sampling, while the cross-attention further aggregates cross-regional contextual information on this basis. This feature optimization mechanism from a dual perspective enables the model to not only finely depict the local geometric details of the damaged area (such as burst shape and size), but also accurately understand the overall spatial layout of the damage pattern.

[0068] In summary, an improved scheme of "multi-scale dilated convolution + cross-attention mechanism" is proposed in the KPConvDef network architecture. This scheme can overcome the limitation of fixed receptive field, enhance the ability to capture multi-scale geometric structures of burst defects, and improve long-distance dependency modeling and noise robustness through global correlation modeling and dynamic weight allocation. The improved KPConvDef-DC network architecture is as follows: Figure 6 As shown.

[0069] 2.2.3 Introducing the Linear Deformable Convolution module:

[0070] This invention introduces a Linear Deformable Convolution (LDConv) module into the KPConv layer. LDConv enables a dynamic, adaptive kernel adjustment mechanism, and its network structure is as follows: Figure 7 As shown, this mechanism breaks through the fixed geometric constraints of traditional convolutional kernels, learning dynamic offsets for each kernel through a data-driven approach, enabling it to adaptively adjust its shape and size based on the local geometric features of the input point cloud. This dynamic kernel adjustment mechanism demonstrates significant advantages for the complex characteristics of segmentation tasks involving the bursting damage of concrete structures after fires.

[0071] (1) Improve feature extraction capabilities:

[0072] While traditional KPConv can effectively process 3D point clouds, its fixed kernel shape has limitations when facing irregular deformation of concrete after disasters. After introducing a linear deformable convolution module, the model can adaptively adjust the sampling position and coverage of the convolution kernel according to the local topology of the damaged area (such as the spalling boundary). This mechanism significantly enhances the context awareness capability and effectively improves the capture of cross-scale damage features.

[0073] (2) Multidimensional extended attention mechanism:

[0074] Linear deformable convolution effectively complements the traditional channel attention mechanism by dynamically adjusting the kernel in the spatial dimension. The model can not only focus on the significantly damaged area, but also adapt the feature response to the geometric heterogeneity of the concrete structure (such as aggregate distribution and reinforcement differences). This multi-dimensional optimization strategy enables the feature representation to simultaneously include spatial deformation sensitivity and semantic discriminability, and exhibits stronger robustness in complex post-disaster scenarios.

[0075] (3) Easy to integrate and optimize:

[0076] This module adopts a lightweight design and is seamlessly integrated into the KPConv framework as a plug-and-play component. This not only avoids large-scale modifications to the original network structure, but also simplifies the model optimization process, reduces the consumption of computing resources, and accelerates the training process. This efficient integration feature provides technical support for the real-time processing of large-scale point cloud data after disasters.

[0077] In summary, the introduction of linearly deformable convolutional modules enables the KPConv layers to form a collaborative mechanism of "geometric adaptation - feature optimization - efficient training," providing a new solution for the refined segmentation of post-fire concrete structure burst damage; the improved KPConv-G layer network architecture is as follows: Figure 8 As shown, the KPConv layer of the original network architecture is replaced with it to form the KP-G-FCNN network architecture, the architecture of which is as follows. Figure 9 As shown.

[0078] Example 4

[0079] Based on Example 3, this example discloses an implementation method for systematically evaluating the performance of the KP-G-FCNN network model, as follows:

[0080] 1. Training strategies and evaluation metrics:

[0081] The network was trained on a custom workstation equipped with an RTX 3090 graphics card with 24GB of VRAM, utilizing CUDA to accelerate network training. The workstation also integrated two E52699 v3 CPUs, each operating at 2.3GHz with 64 cores and 128 threads. The network architecture was developed using the Python-based PyCharm integrated development environment and the PyTorch library for deep learning applications.

[0082] The KP-G-FCNN network employs the AdamW optimizer (β1 = 0.9, β2 = 0.999, weight decay = 0.01) with a piecewise constant decay strategy. The initial learning rate is set to 0.01, decaying by a factor of 0.1 every 5 epochs. A learning rate warm-up technique is used, with the first two epochs using a factor of 0.1 to alleviate the instability of the optimizer in the initial stage. Data augmentation methods include geometric transformations (applying random rotations (±15°), scaling (0.8-1.2 times), and translations (±0.2 meters) to the point cloud to improve the model's robustness to spatial transformations) and neighborhood perturbations. (Randomly adding Gaussian noise (standard deviation 0.01 meters) in the local neighborhood to simulate measurement errors in real-world scenarios), density adjustment (enhancing the model's adaptability to different point cloud densities through random downsampling (retaining 80%-100% of points), etc.); In network training, to prevent excessive deformation leading to feature loss and to prevent overfitting, a regularization strategy of kernel point offset constraint and weight decay is adopted; the training batch is set to 16, the training cycle is 300 rounds, and the loss function adopts cross-entropy loss combined with label smoothing technology, with the smoothing parameter set to 0.1.

[0083] This invention uses the mean intersection over union (MIoU) to systematically evaluate the performance of the KP-G-FCNN network model. The mean intersection over union (MIoU) is used as the core segmentation accuracy index. By calculating the average overlap between the predicted area and the real labeled area in each category, the model's positioning accuracy for fire-damaged areas is quantitatively characterized. The calculation formula is shown in equations (1)-(4).

[0084]

[0085]

[0086]

[0087]

[0088] In equation (1), the numerator is the sum of all correctly predicted pixels, the denominator is the total number of pixels in the image, and FN (False Negative), FP (False Positive), TN (True Negative), and TP (True Positive) are the values ​​in the confusion matrix.

[0089] 2. Ablation test results:

[0090] This invention relates to an improved method for 3D point cloud segmentation models based on multi-module collaborative optimization. The effectiveness and synergistic advantages of each technical module were verified through systematic ablation experiments. Using the KP-FCNN model as a benchmark, the experiment quantitatively analyzed the impact of each module on the segmentation accuracy (MIoU) by progressively introducing the Dilated Conv module, CA attention mechanism, and LD Conv module. The experimental results are shown in Table 1. Based on the benchmark model KP-FCNN (MIoU 69.21%), the impact of the introduction and combination strategies of each module on the segmentation accuracy is as follows:

[0091] (1) Single module analysis:

[0092] The Dilated Conv module (number 2) improved MIoU to 76.45%, indicating that it effectively enhanced the ability to extract local features by expanding the receptive field;

[0093] After the introduction of the CA attention mechanism (number 3), the MIoU reached 74.67%, which verified the optimization effect of channel attention on feature selection.

[0094] The LD Conv module (number 4) performed best, with an MIoU of 78.64%, demonstrating its significant advantage in hierarchical dynamic convolution for modeling complex spatial structures.

[0095] (2) Synergistic effect of dual modules:

[0096] The combination of Dilated Conv and CA (number 5) improved MIoU to 79.41%, which is significantly higher than the single module gain, demonstrating the complementarity of feature enhancement and attention selection;

[0097] The combination of CA and LD Conv (number 7) achieved an MIoU of 81.85%, which is the optimal dual-module scheme, indicating that the synergistic optimization of attention mechanism and dynamic convolution can significantly improve the efficiency of global-local feature fusion.

[0098] The MIoU of the combination of Dilated Conv and LD Conv (serial number 6) is 78.48%, which is slightly lower than that of a single LD Conv module. It is speculated that this is because both focus on local feature expansion, resulting in a redundancy effect.

[0099] (3) Full module integration optimization:

[0100] When all modules (number 8) are integrated, the MIoU reaches 82.31%, which is 0.46% higher than the optimal dual-module combination (number 7), verifying the cumulative gain effect of multi-module collaborative optimization. Although the marginal benefit is reduced, the scheme maximizes the segmentation accuracy through comprehensive feature enhancement, attention guidance and dynamic convolution mechanism.

[0101] Table 1 Ablation Test Results

[0102] Serial Number KP-FCNN DilatedConv CA LDConv MIoU 1 √ 69.21% 2 √ √ 76.45% 3 √ √ 74.67% 4 √ √ √ 78.64% 5 √ √ √ 79.41% 6 √ √ √ 78.48% 7 √ √ √ 81.85% 8 √ √ √ √ 82.31%

[0103] The above results show that the multi-module collaborative optimization method proposed in this invention significantly improves the performance of the 3D point cloud segmentation model through the organic combination of differentiated technical paths. Among them, the synergistic effect of the CA attention mechanism and the LD Conv module is particularly prominent.

[0104] The loss function graphs of KP-G-FCNN and KP-FCNN networks are shown below. Figure 11 As shown, the MIoU value varies with the number of training epochs. Figure 12 As shown. By Figure 11 , 12 It can be known that:

[0105] (1) Regarding the loss curves, both the KP-G-FCNN and KP-FCNN models showed a rapid decrease in loss values ​​during the initial 30 time steps. Subsequently, between time steps 30 and 100, the rate of decrease in loss values ​​slowed down, eventually stabilizing after 250 training time steps until the training process was complete. Notably, no overfitting was observed during either training or validation.

[0106] (2) Regarding the MIOU value curve, in the first 50 time steps, both the KP-G-FCNN and KP-FCNN models showed a rapid increase in loss value. Subsequently, KP-G-FCNN continued to rise steadily to 0.828 and then stopped rising, while KP-FCNN rose to 0.692 and then stopped rising. Therefore, the improved model improved the MIOU value by 13.6%.

[0107] For detailed images of the model prediction results of this invention, please refer to the diagrams below. Figure 10 As shown in the figure, the model achieves accurate segmentation of both the original and synthetic data, with good results, and can accurately segment burst damage of different scales.

Claims

1. A three-dimensional segmentation method for burst damage of RC structures after a fire, characterized in that, Includes the following steps: Step 1: Construct a 3D point cloud database of RC structure damage after fire; Step 2: Use KP-FCNN as the baseline network to construct a 3D point cloud segmentation network for concrete structures; Step 3: Systematically evaluate the performance of the 3D point cloud segmentation network using the average intersection-union ratio and conduct ablation experiments; Step 2, the improvement of the baseline network, specifically includes: (21) Introduce multi-scale dilated convolution to realize a dynamic, multi-level receptive field expansion mechanism; This mechanism combines convolutional kernels with different dilation rates to form a multi-scale feature perception capability, thereby improving the ability to capture complex geometric structures while maintaining parameter efficiency. (22) Introduce a cross-attention mechanism to achieve a dynamic balance between global semantic association and local feature enhancement; by dynamically establishing the semantic association of distant point cloud features, the limitations of the local receptive field are overcome, and an adaptive weight allocation strategy is used to achieve intelligent suppression of noise interference. (23) Introduce a linear deformable convolution module. Introduce the linear deformable convolution module into the KPConv layer to realize a dynamic and adaptive kernel adjustment mechanism. Learn dynamic offset for each convolution kernel in a data-driven manner so that it can adaptively adjust its shape and size according to the local geometric features of the input point cloud.

2. A three-dimensional segmentation method for RC structure burst damage after a fire, as described in claim 1, characterized in that step 1 includes the following specific methods: The first method is to take photos of the concrete burst damage using a regular high-definition camera. During the shooting, care should be taken to take photos of the same damaged component from different angles. To obtain a better three-dimensional reconstruction point cloud effect based on two-dimensional images, it is necessary to ensure that each adjacent pair of photos has at least 50% overlap. Then, all the photos are imported into Colmap software for three-dimensional point cloud reconstruction; The second method utilizes the built-in LiDAR and camera of an iPhone in conjunction with 3D Scanner. The app scans damaged components and automatically generates 3D reconstructed point cloud maps; the third method uses a depth camera to photograph the fire-damaged components and directly obtains point cloud maps; the fourth method uses an R8 LiDAR to scan the fire-damaged components and obtain point cloud maps; in addition, there is a method for synthesizing data, which first models the fire-damaged components using 3D modeling software, then imports the exported .obj file into CloudCompare software for upsampling to obtain high-density 3D point cloud maps; all the datasets obtained by the above methods are imported into CloudCompare software for point cloud segmentation and annotation, after outlining the edge contour of the explosion damage, the label is set to "1", and the background label is set to "0"; after all annotations are completed, a txt file is exported, and the original labeled dataset is expanded using data augmentation techniques, and then the expanded dataset is divided into training, validation and test sets in a 7:2:1 ratio to facilitate subsequent network training.

3. The three-dimensional segmentation method for post-fire RC structure burst damage as described in claim 2, characterized in that, In step 2, the cross-attention module and the dilated convolutional layer are deeply integrated to form a "local-global" collaborative feature enhancement paradigm. That is, the dilated convolution expands the local perception range through sparse sampling, while the cross-attention further aggregates cross-regional contextual information on this basis.

Citation Information

Patent Citations

  • Complex special-shaped curved surface three-dimensional segmentation method and system based on robot vision

    CN111028238A

  • Image segmentation method for burn wound area

    CN119399469A