Building component fire loss real-time segmentation method based on FastSAM-P

Through the real-time fire loss segmentation method of building components based on FastSAM-P, combined with DSCConv, RFB and PPSM modules, the accuracy and speed problems in the damage detection of building components after fire are solved, and efficient damage segmentation and real-time detection are achieved.

CN120147643AActive Publication Date: 2025-06-13QINGDAO UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510294697.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art has problems such as limited detection accuracy, slow segmentation speed and poor applicability in the detection of damage to building components after fire, which limits its wide application.

Method used

The real-time fire loss segmentation method of building components based on FastSAM-P is adopted, and the real-time fire loss segmentation network architecture based on FastSAM-P is constructed, including the DSCConv module, RFB module and PPSM module to improve segmentation accuracy and inference speed.

Benefits of technology

Real-time segmentation of fire damage of building components is realized, synergistic optimization of segmentation accuracy and reasoning speed is improved, and has significant technological advancement and industrial application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147643A_ABST
    Figure CN120147643A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of fire damage detection methods, in particular to a FastSAM-P-based real-time segmentation method for fire loss of a building component, which comprises the following steps of: 1, determining types and characteristics of two damages, namely concrete burst and steel bar exposure on the surface of an RC structural component after a fire, and constructing a corresponding data set for network training and verification; 2, formulating a training strategy and an evaluation index; and step 3, constructing a building component fire loss real-time segmentation network architecture based on FastSAM-P. Compared with a traditional segmentation network, collaborative optimization of segmentation precision and reasoning speed is realized in a building fire loss identification task, and the method has remarkable technical advancement and industrial application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fire damage detection methods, and particularly to a real-time segmentation method for fire damage of building components based on FastSAM-P. Background Art

[0002] The damage detection technology of building components is roughly divided into two categories: non-destructive testing and destructive testing. In the scope of non-destructive testing, the impact echo method, as a landmark means, strikes the surface of the building component by applying a force hammer, and quantifies the damage condition of the concrete according to the rebound depth. However, the premise of this technology is that the surface of the component needs to be relatively flat, and it is ineffective for the surface of the component with concrete spalling phenomenon. In addition, the detection efficiency of the impact echo method is relatively low, time-consuming and laborious, which limits its wide application.

[0003] For the field of destructive testing, the core drilling method is a typical representative. This method samples the damaged component with a drill at the disaster site, and then transfers the sample to the laboratory, where professional technicians use precision equipment to deeply analyze the sample to accurately judge the degree of damage. Nevertheless, the core drilling method is not perfect either. Its detection cycle is relatively long, generally taking 2 to 3 weeks, and it is extremely easy to cause additional damage to the component during the sampling process, thus posing a potential threat to the safety of on-site inspectors.

[0004] Currently, although image detection technologies based on deep learning have made certain progress, they still face a series of challenges, mainly including limited detection accuracy, slow segmentation speed, and poor applicability. These factors seriously restrict their potential for wide application. Specifically, these technologies often fail to meet the high-precision requirements in practical applications, the processing speed is insufficient to meet the needs of rapid response, and they show low adaptability in the face of diverse scenarios, thus limiting their effective deployment in a wider range of fields.

[0005] Therefore, it is particularly important to develop an automatic, real-time and non-destructive fire damage detection technology for building components, in order to greatly improve the detection efficiency and meet the urgent needs in practical applications. Summary of the Invention

[0006] The present invention provides a real-time segmentation method for fire damage of building components based on FastSAM-P, which can realize automatic and real-time damage segmentation of concrete structures after fire, and provide technical support for subsequent damage quantification and deployment to unmanned aerial vehicles.

[0007] To achieve the above object, the technical solution of the present invention is as follows: A real-time segmentation method for fire damage of building components based on FastSAM-P, comprising the following steps: Step 1: Identify the types and characteristics of two types of damages, namely surface concrete spalling and steel bar exposure in RC structural components after a fire, and construct a corresponding dataset for network training and validation; Step 2: Develop training strategies and evaluation metrics; Step 3: Construct a real-time segmentation network architecture for fire damage of building components based on FastSAM-P.

[0008] Preferably, Step 1 includes the following specific steps: The dataset is taken from the post-fire site and the fire laboratory, with a total of 1500 images. These images clarify the types and characteristics of two types of damages, namely surface concrete spalling and steel bar exposure in RC structural components after a fire. During shooting, the image size is set to a square of 1:1 for taking pictures, and then the images are imported into the labelme software for annotation. The generated json files are converted into XML files. Through data augmentation, the dataset is expanded to 8 times the original quantity. The expanded dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1.

[0009] Preferably, in Step 2, a multi-stage optimization strategy is adopted to train the FastSAM-P segmentation model to improve the recognition accuracy of concrete structure spalling and steel bar exposure damages after a fire. During training, the batch size is set to 16, the total number of iterations is 100, and training is carried out for 500 rounds. And a phased learning rate regulation is implemented: initially, a linear warm-up strategy is adopted to alleviate gradient oscillation, and after the training is stable, the cosine annealing algorithm is used to gradually reduce the learning rate. To suppress the overfitting phenomenon, an L2 regularization term is introduced into the network optimization objective, and the model complexity is reduced and the generalization ability is improved by constraining the L2 norm of the weight parameters. The loss function adopts a weighted combination of BCE Loss + Dice Loss. Through this combination strategy, both the pixel-level classification accuracy, i.e., the advantage of BCE loss, and the regional structure consistency, i.e., the characteristics of Dice loss, are maintained, as shown in Equation (3): (3); In Equation (3), λ is a hyperparameter used to balance the weights between BCE loss and Dice loss.

[0010] Preferably, in step 2, a multi-dimensional quantization index system is used to systematically evaluate the performance of the FastSAM-P model as follows: 1) The mean intersection over union is used as the core segmentation accuracy index. By calculating the average overlap degree between the predicted region and the true annotation region for each category, it quantitatively characterizes the localization accuracy of the model for the fire damage region; 2) The number of parameters is used as an evaluation index for model complexity, reflecting the model's storage requirements and computational resource consumption, and providing a basis for the feasibility of edge device deployment; 3) The frame rate is used as an evaluation index for real-time performance. Under the NVIDIA RTX 3090 graphics card environment, the average inference time of a single image is measured to evaluate the application potential of the model in real-time detection scenarios at the fire scene.

[0011] Preferably, step 3 includes the following contents: (1) Select the FastSAM network as the baseline network, and replace the C2F module with the DSCConv module in the backbone part of YOLOv8; Based on the SCConv module, by adding the Deformable Convolution module, the DSCConv module is formed; (2) Replace the SPPF module with the RFB module in the backbone part of YOLOv8; (3) Replace the FPN module in the FastSAM network with the PPSM module: Based on the pyramid pooling module, by adding the attention module, the PPSM is formed, and rich context information is extracted through multi-scale pooling operations, thereby enhancing the network's recognition ability for multi-scale targets in the concrete fire damage segmentation task.

[0012] Preferably, in step 3, the PPSM performs global pooling operations on the input feature map at different scales to generate multi-scale feature representations, quickly extracts features through the attention module, and fuses them with the original feature map through upsampling, finally realizing the organic combination of global context information and local detail features; The PPSM generates context information at multiple scales by performing pooling operations on the feature map at different scales, specifically including the following 4 scales of pooling: (31) Global pooling (1×1): Perform global average pooling on the feature map to capture the global context information of the entire image; (32) Medium-scale pooling (2×2): Divide the feature map into 2×2 sub-regions and perform pooling to extract semantic information of larger regions; (33) Smaller-scale pooling (3×3 or 4×4): Divide the feature map into 3×3 or 4×4 sub-regions and perform pooling to focus on feature extraction of smaller regions; (34) Original-scale features: Retain the unpooled original feature map to retain high-resolution detail information; Then the PPSM restores the pooled features at different scales to the size of the original feature map through upsampling operations, and then performs pixel-by-pixel stitching with the original feature map.

[0013] Preferably, the specific steps of step (34) are as follows: The pooling feature maps of each scale are upsampled by bilinear interpolation to restore to the same resolution as the original feature map. Then, the upsampled multi-scale feature maps are concatenated with the original feature map in the channel dimension to form a fused feature representation. Finally, by fusing global context information and local detail features, the model can simultaneously process large wall areas and narrow boundary areas, thereby improving the segmentation accuracy.

[0014] The beneficial effects of the real-time segmentation method for building component fire damage based on FastSAM-P of the present invention: Compared with traditional segmentation networks, the present invention realizes the collaborative optimization of segmentation accuracy and inference speed in the building fire damage recognition task, and has significant technological advancement and industrial application value. Its application advantages are as follows: 1. Seamless adaptability of multi-source heterogeneous devices: It supports the instant access of multi-modal image acquisition terminals such as mobile phones, high-definition cameras, and drones. Through the lightweight model architecture design, it realizes the standardized preprocessing and real-time inference of cross-platform image data, meeting the flexible detection requirements in the complex environment of the fire scene; 2. End-to-end edge computing ability: The parameter scale of the optimized FastSAM-P network is reduced, and it can complete the real-time segmentation of 1080P images at 65.11fps on mobile devices such as drones equipped with edge computing modules, significantly reducing the dependence on cloud computing; 3. Engineering adaptability of dynamic deployment: It provides a multi-level deployment solution from embedded devices to server clusters, and especially develops a dedicated inference acceleration interface for the drone inspection system to realize the automated operation closed-loop of aerial image acquisition - damage recognition - result feedback; 4. Data foundation for damage quantification analysis: The output pixel-level segmentation results are stored in a structured data format, accurately locating key damage features such as the area of concrete spalling and the exposed form of steel bars, providing high-precision input data for subsequent damage level assessment and structural bearing capacity calculation. Description of the Drawings

[0015] Figure 1 is the Labelme image annotation interface.

[0016] Figure 2 is the FastSAM network architecture (in the figure, YOLOv8-Seg Backbone: YOLOv8-Seg backbone network; FPN: Feature Pyramid Network; Detect Branch: Detection Branch; Mask Branch: Mask Branch; P3, P4, P5: P3, P4, P5 (levels, remain unchanged); Mask Coeff.: Mask Coefficient; Detect: Detection; NMS: Non-Maximum Suppression; Crop: Cropping; Pred: Prediction; Threshold: Threshold; Image Encode: Image Encoding; Text Encoder: Text Encoding; Point-prompt: Point Prompt; Box-prompt: Box Prompt; Text-prompt: Text Prompt).

[0017] Figure 3 It is the network architecture of the DSConv module (in the figure: Spatial and Channel reconstruction Convolution: Spatial and Channel Reconstruction Convolution; Previous ConvBlock: Previous Convolution Block; NextConvBlock: Next Convolution Block; 1×1 Conv: 1×1 Convolution; Input Feature: Input Feature; Spatial-Refined Feature: Spatially Optimized Feature; Channel-Refined Feature: Channel-Optimized Feature; ResBlock +SCConv: Residual Block + SCConv; DFConv: Deformable Convolution).

[0018] Figure 4 It is the network architecture of the RFB module (in the figure: 3×3Conv Rate = 1: 3×3 Convolution Rate = 1; Concat 1×1Conv: Concatenate 1×1 Convolution; Spatial array of pRF in hV4: Spatial Array of pRF in hV4; Final spatial array of receptive field: Final Spatial Array of Receptive Field).

[0019] Figure 5 It is the comparison of the backbone network structures before and after improvement (in the figure: Conv: Convolutional Layer; SCConv: Spatial and Channel Reconstruction Convolution).

[0020] Figure 6 It is the structural diagram of the PPSM module (in the figure: POOL: Pooling Layer; UPSAMPLE: Upsampling Layer).

[0021] Figure 7It is the improved FastSAM-P network architecture (in the figure: Detect Branch: Detection Branch; MaskBranch: Mask Branch; Mask Coeff: Mask Coefficient; Detect: Detection; Crop: Cropping; Pred: Prediction; Threshold: Threshold; Image Encode: Image Encoding; Text Encoder: Text Encoding; Point-prompt: Point Prompt; Box-prompt: Box Prompt; Text-prompt: Text Prompt).

[0022] Figure 8 It is the performance comparison between FastSAM-P and FastSAM networks (where Figure a is the MIoU comparison; Figure b is the comparison of Params and FPS).

[0023] Figure 9 It is the comparison of the segmentation effects of different networks. Specific implementation manners

[0024] As described below, the embodiments of the present invention are described in detail in a step-by-step progressive manner. This description is only for the preferred embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0025] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, as well as a specific orientation structure and operation. Therefore, it should not be construed as a limitation to the present invention.

[0026] Embodiment 1: A real-time segmentation method for fire damage of building components based on FastSAM-P includes the following steps: Step 1: Identify the types and characteristics of two types of damages, namely concrete spalling and steel bar exposure on the surface of RC structural components after a fire, and construct a corresponding dataset for network training and verification; Step 2: Develop training strategies and evaluation indicators; Step 3: Construct a real-time segmentation network architecture for fire damage of building components based on FastSAM-P.

[0027] Embodiment 2: Based on Embodiment 1, this embodiment discloses that the specific steps of the said Step 1 include the following: At present, there is little research on the damage segmentation of RC structures after fire, and there is a lack of benchmark datasets for post-fire damage segmentation. Therefore, the present invention constructs a dataset for post-fire damage segmentation of building components; the dataset is taken from the scenes after fires and fire laboratories, with a total of 1500 images, which clarify the types and characteristics of two types of damage, namely, concrete spalling on the surface of RC structural components and steel bar exposure after fire; when taking pictures, the image size is set to a square of 1:1 for taking pictures, and then the images are imported into the labelme software ( Figure 1 ) for annotation, the generated json file is converted into an XML file, and through data augmentation, the dataset is expanded to 8 times the original quantity, and the expanded dataset is divided into a training set, a validation set and a test set according to the ratio of 7:2:1.

[0028] Example 3: Based on Example 1 and 2, this example discloses that in step 2, the training strategy is as follows: The training of the FastSAM-P network is carried out on a custom workstation equipped with a Precision RTX 3090 graphics card, which has 24GB of VRAM and uses CUDA to accelerate network training; in addition, this workstation integrates two E5 2699 v3 CPUs, each with a working frequency of 2.3 GHz, 64 cores and 128 threads; this network architecture is developed using the Python-based PyCharm integrated development environment and the PyTorch library for deep learning applications; a multi-stage optimization strategy is adopted to train the FastSAM-P segmentation model to improve the recognition accuracy of the bursting and steel bar exposure damage of the concrete structure after a fire; during the training process, the batch size is set to 16, the total number of iterations is 100, and it is trained for 500 rounds, and a phased learning rate adjustment is implemented: initially, a linear warm-up strategy is adopted to alleviate gradient oscillation, and after the training is stable, the cosine annealing algorithm is used to gradually reduce the learning rate; to suppress the overfitting phenomenon, an L2 regularization term is introduced into the network optimization objective, and the model complexity is reduced and the generalization ability is improved by constraining the L2 norm of the weight parameters; BCE Loss (Binary Cross Entropy Loss) refers to pixel-by-pixel classification, which measures the difference between the predicted segmentation mask and the true label. When the dataset is imbalanced, different weights can be applied to the losses of different classes to mitigate the impact of class imbalance, and its calculation method is shown in Equation (1); Dice Loss (Dice CoefficientLoss) is a loss function widely used in image segmentation tasks, especially suitable for pixel-level binary or multi-classification tasks, and its calculation formula is shown in Equation (2); the loss function of the present invention adopts a weighted combination of BCE Loss + Dice Loss. Through this combination strategy, both the pixel-level classification accuracy (the advantage of BCE loss) and the regional structure consistency (the characteristics of Dice loss) are maintained. For the segmentation task of the bursting and steel bar exposure damage of the concrete structure after a fire, this method effectively improves the model's ability to capture fine-grained features and regional consistency, and the specific form is shown in Equation 3: (1); (2); (3); In Equation (1), N is the number of samples, y i is the true label of the i-th sample, and p i is the probability that the model predicts the i-th sample as the positive class; in Equation (2), P i and G irespectively represent the values of the predicted result and the ground truth label at the \(i\)-th pixel; in Equation (3), \(\lambda\) is a hyperparameter used to balance the weights between the BCE loss and the Dice loss.

[0029] Example 4: Based on Examples 1, 2, and 3, this example discloses that in step 2, a multi-dimensional quantization index system is used to systematically evaluate the performance of the FastSAM-P model, as follows: 1) Mean Intersection over Union (MIoU) is used as the core segmentation accuracy index. By calculating the average overlap degree of the predicted region and the ground truth annotation region for each category, it quantitatively characterizes the localization accuracy of the model for the fire damage region; 2) The number of parameters is used as an evaluation index for model complexity, reflecting the storage requirements and computational resource consumption of the model, and providing a basis for the feasibility of edge device deployment; 3) Frames Per Second (FPS) is used as a real-time evaluation index. Under the NVIDIA RTX 3090 graphics card environment, the average inference time of a single image is measured to evaluate the application potential of the model in the real-time detection scenario of the fire scene. This index system takes into account the evaluation of the balance between model accuracy and efficiency: MIoU verifies the model's ability to identify fine-grained features such as concrete spalling and steel bar exposure from the perspective of pixel-level segmentation accuracy; the number of parameters index reflects the lightweight level of the model, corroborating the applicability of the algorithm on mobile terminals; the FPS index quantifies the actual deployment value of the model in disaster emergency scenarios.

[0030] Example 5: Based on Examples 1, 2, 3, and 4, this example discloses that step 3 includes: In order to accurately detect and segment the spalling and steel bar exposure damage of the concrete structure after a fire and achieve real-time detection, the present invention selects the FastSAM (Fast Segment Anything) network as the baseline network; FastSAM is an accelerated alternative proposed by the Institute of Automation, Chinese Academy of Sciences, aiming to solve the problems of SAM (Segment Anything Model) in terms of computational speed and resource occupancy; while maintaining performance comparable to SAM, FastSAM boosts the inference speed by 50 times, achieving the real-time "segment everything" function; FastSAM is based on the instance segmentation of YOLOv8-seg. During detection, it integrates the instance segmentation branch and adopts a two-stage algorithm of "full instance segmentation + instruction-based mask output"; in the first stage, a convolutional neural network (CNN) is used for full instance segmentation to generate segmentation masks for all instances in the image; in the second stage, according to the prompts provided by the user (such as points, boxes, texts, etc.), the corresponding segmentation regions are output; its network framework is as Figure 2As shown; this design not only improves computational efficiency but also maintains performance comparable to SAM; In the backbone architecture of YOLOv8, the traditional C2F module is replaced with an advanced Deformable-Spatial and Channel reconstruction Convolution (DSCConv) module. This improvement is as Figure 3 shown, aiming to significantly improve the segmentation accuracy of the model for complex damage features such as concrete spalling and steel bar exposure after a fire; Although the C2F module effectively reduces model parameters and accelerates the training process through its unique cross-stage local connection mechanism, when faced with damage detection tasks such as steel bar exposure on the concrete surface, which have significant local deformations and highly complex texture features, its convolutional kernels based on fixed geometric structures show limitations in capturing fine-grained features; In addition, traditional convolutional operations are limited by fixed sampling positions and are difficult to flexibly adjust the receptive field to adapt to the actual shape of the damage target, which not only leads to spatial redundancy in the feature map but also may cause the loss of key information; To address this challenge, the DSCConv module is introduced, which cleverly combines the advantages of deformable convolution ( Figure 2 ) and spatial channel reconstruction convolution (SCConv) to construct a feature extraction framework with highly adaptive perception capabilities; Through the learnable spatial offset parameters introduced by the deformable convolutional layer, DSCConv can dynamically adjust the receptive field distribution according to the actual shape of the damage target, thus significantly enhancing the ability to capture and represent irregular structural features, which is particularly crucial for identifying complex damages such as concrete spalling and steel bar exposure after a fire; At the level of spatial feature optimization, the spatial reconstruction unit (SRU) in the DSCConv module effectively suppresses redundant features through a feature weight separation mechanism, and at the same time uses a feature recalibration method to strengthen the spatial attention of key regions, which further improves the spatial discrimination of the feature map, enabling the model to more accurately identify the specific location and shape of concrete damage; In terms of channel dimension optimization, the channel reconstruction unit (CRU) in the DSCConv module constructs a lightweight channel attention mechanism through innovative channel splitting transformation and feature fusion strategies. This mechanism significantly reduces channel redundancy while effectively reducing the computational complexity of the model, ensuring the efficiency and real-time performance of the model when processing complex damage features; In summary, the introduction of the DSCConv module not only overcomes the limitations of traditional convolutions when dealing with complex damage features but also significantly improves the performance of the network in the tasks of post-fire concrete spalling and steel bar exposure damage segmentation through its adaptive perception ability and refined feature optimization strategy. Its network structure diagram is as shown in Figure 3 shown; Secondly, in the backbone part of YOLOv8, the SPPF module is replaced with the Receptive Field Block Modified (RFB) module. Its network architecture is as shown in Figure 4 shown; First of all, by simulating the receptive field structure of the human visual system and combining convolutional kernels of different sizes and dilated convolutions, the RFB module can extract multi-scale features at a single level; since the widths, lengths, and shapes of steel bar exposures vary greatly, there are both fine micro-steel bar exposures and larger structural steel bar exposures; the multi-scale feature extraction ability of the RFB module enables the model to capture steel bar exposure features of different sizes simultaneously, thereby improving the integrity and accuracy of segmentation; Secondly, the RFB module fuses features from different receptive fields through a multi-branch structure, enhancing the feature expression ability; in the tasks of concrete spalling and steel bar exposure segmentation, the steel bar exposure may have complex shapes (such as intersections), and the spalling may have irregular edges. The RFB module can better capture these details and improve the segmentation accuracy of the spalling and steel bar exposure shapes; In addition, there may be noise (such as stains, soot) or complex textures on the surface of post-fire concrete. The multi-branch feature fusion ability of the RFB module helps to distinguish damage from background interference and reduce false detections and missed detections; Finally, compared with the pooling operation of the SPPF module, the RFB module retains more spatial information through convolutional operations, reducing information loss; this is particularly important for the task of concrete structure damage segmentation because steel bar exposures are usually slender structures, and spatial details (such as the width and direction of steel bar exposures) are crucial for the segmentation results; the RFB module can better retain the detailed information of steel bar exposures, especially fine steel bar exposures, and avoid incomplete segmentation due to information loss; In summary, in the backbone part of YOLOv8, replacing the C2F module and the SPPF module with the SCConv module and the RFB module respectively significantly improves the segmentation performance and robustness of the model. The backbone network structures before and after improvement are as shown in Figure 5 shown; The improvements of the present invention also include: replacing FPN with PPSM. Specifically, based on the Pyramid Pooling Module (PPM), by adding the Shuffle-Attention attention module, the Pyramid Pooling Shuffle Module (PPSM) is formed. It extracts rich context information through multi-scale pooling operations, thereby enhancing the model's recognition ability for multi-scale targets in the concrete fire damage segmentation task. Specifically, PPSM performs global pooling operations of different scales on the input feature map to generate multi-scale feature representations, quickly extracts features through the Shuffle-Attention attention module, and fuses with the original feature map through upsampling, ultimately realizing the organic combination of global context information and local detail features. PPSM generates context information of multiple scales by performing pooling operations of different scales on the feature map. Specifically, it includes the following four scales of pooling: (1) Global pooling (1×1): Perform global average pooling on the feature map to capture the global context information of the entire image. This global information helps the model understand the overall distribution and structure of the wall. (2) Medium-scale pooling (2×2): Divide the feature map into 2×2 sub-regions and perform pooling to extract semantic information of larger regions. This scale is suitable for identifying large wall areas. (3) Smaller-scale pooling (3×3 or 4×4): Divide the feature map into 3×3 or 4×4 sub-regions and perform pooling, focusing on feature extraction of smaller regions. This scale helps identify narrow burst boundaries or local details. (4) Original-scale features: Retain the unpooled original feature map to retain high-resolution detail information. This original feature is crucial for accurately segmenting the wall edge and small-scale structures. (5) The PPSM restores the pooled features of different scales to the size of the original feature map through an upsampling operation, and then performs pixel-by-pixel concatenation with the original feature map. The specific steps are as follows: The pooled feature map of each scale is upsampled by bilinear interpolation to restore the same resolution as the original feature map; then the upsampled multi-scale feature maps are concatenated with the original feature map in the channel dimension to form a fused feature representation; finally, by fusing global context information and local detail features, the model can simultaneously process large wall areas and narrow boundary areas, thereby improving the segmentation accuracy; therefore, the PPSM can capture rich context information from global to local and adapt to the wall segmentation requirements of different scales; the retention of the original feature map ensures that high-resolution detail information is not lost, thereby improving the segmentation accuracy of the edges of burst and steel bar exposure damages; the introduction of global context information enhances the model's adaptability to complex scenes, enabling it to better handle challenges such as occlusion and lighting changes; in summary, the added Shuffle-Attention attention mechanism has a simple structure and high computational efficiency; the specific network structure is as follows Figure 6 as shown; In summary, based on the FastSAM network structure, by improving the backbone and optimizing the network modules, the model has improved detection and segmentation accuracy, reduced the model size, and significantly shortened the inference time. The improved FastSAM-P network architecture is as Figure 7 shown.

[0031] Example 6: Ablation experiment: The present invention relates to an improved method for an image segmentation model based on multi-module collaborative optimization. The effectiveness and collaborative advantages of each technical module are verified through systematic ablation experiments. The experiments are based on the FastSam model. By gradually introducing the DSCConv module, the RFB module, and the PPSM module, the effects of each module on the model parameter quantity, segmentation accuracy (MIoU), and processing speed (FPS) are quantitatively analyzed. The experimental results are shown in Table 1, and the specific analysis is as follows: (1) Analysis of the independent action of the module: DSCConv (serial number 2): After introduction, the MIoU increased by 10.5% (68.4% → 78.9%), but the parameter quantity increased by 0.38M (12.10M → 12.48M), and the FPS decreased by 5.11 frames / s (45.23 → 40.12); indicating that this module significantly improves the segmentation accuracy by increasing the model complexity, but sacrifices the inference speed; RFB (Serial number 3): The MIoU increased by 10.7% (from 68.4% to 79.1%), the number of parameters decreased by 1.83M (from 12.10M to 10.27M), and the FPS increased significantly by 13.55 frames / s (from 45.23 to 58.78); This shows that the RFB module not only optimizes the model lightweight (the number of parameters is reduced by 15.1%), but also significantly improves the accuracy and speed, and has efficient feature expression ability; PPSM (Serial number 4): The MIoU increased by 7.1% (from 68.4% to 75.5%), the number of parameters decreased by 1.95M (from 12.10M to 10.15M), and the FPS increased by 14.0 frames / s (from 45.23 to 59.21); This module performs optimally in terms of speed optimization, but the accuracy gain is weaker than the RFB module, and it may be more suitable for scenarios with high real-time requirements; (2) The collaborative optimization effect of modules: DSCConv + RFB (Serial number 5): The MIoU reached 81.2%, a 12.8% increase compared to the baseline, but the FPS is lower than that of using RFB alone (56.52 vs. 58.78), indicating that the complexity of DSCConv offsets part of the speed advantage; DSCConv + PPSM (Serial number 6): The MIoU increased significantly to 85.7%, and the FPS remained at 58.27 frames / s, indicating that PPSM effectively alleviates the speed loss of DSCConv, and the two cooperate to improve the balance between accuracy and efficiency; RFB + PPSM (Serial number 7): The MIoU reached 86.1%, the number of parameters is the lowest (9.31M), and the FPS is the highest (69.18 frames / s); This combination performs optimally in lightweight and speed optimization, verifying the complementarity of RFB and PPSM; Full-module combination (Serial number 8): The MIoU reached 92.0%, a 23.6% increase compared to the baseline, but the FPS dropped to 65.11 frames / s; Although the number of parameters (9.90M) is lower than the baseline, the superposition of multiple modules leads to an increase in the computational load, indicating that the limit of model performance requires a balance between accuracy and speed requirements; Table 1 Results of ablation experiments: ; The present invention solves the contradiction between model lightweight and accuracy improvement through modular design, where: the DSCConv module reduces redundant parameters through dynamic channel compression; the RFB module enhances feature discriminability through multi-level receptive fields; the PPSM module optimizes the inference speed through parallel computing; Experimental data shows that the synergistic effect of the three breaks through the performance limitations of a single module (such as the MIoU and FPS of the dual-module combination in Experiments 6 - 7 are both lower than those in Experiment 8), forming a complete technical solution with significant technical gains.

[0032] Example 7: Comparative Experiment: The accuracy performance comparison between the improved FastSAM-P network of the present invention and the original FastSAM model is as follows Figure 8 As shown, to further verify the technical advantages of the FastSAM-P model, the present invention selects four existing advanced segmentation models, namely YOLOv11, MB-SPPF-UNet, OCRNet, and SegFormer, as the comparison benchmarks, and conducts model training and testing respectively according to the same training strategy using the same dataset; the experimental results show that the improved FastSAM-P model is significantly superior to the existing technical solutions in terms of segmentation accuracy indicators. The comparison effect diagrams of the prediction results of each model are shown in detail Figure 9 as shown; from Figure 9 it can be seen that compared with other network models, the FastSAM-P has the best segmentation effect and can segment burst and exposed steel bars of different scales, which is significantly higher than the unimproved model FastSAM.

[0033] Working principle of the present invention: 1. Network architecture optimization module: Aiming at the deficiencies of the traditional FastSAM network in building damage segmentation tasks, it is proposed to use the SCConv module to replace the original C2F module in the Backbone structure, and effectively reduce the redundant feature calculation through the spatial and channel recalibration mechanism; at the same time, the RFB module is used to replace the SPPF module, and the dilated convolution is used to enhance the multi-scale feature fusion ability, enhance the recognition accuracy of the edges and shapes of exposed steel bars and bursts, and improve the feature expression ability while reducing the number of parameters; 2. Feature pyramid reconstruction unit: Innovatively improve the standard feature pyramid network (FPN) into a pyramid segmentation module (PPSM) introducing the Shuffle-Attention attention module. Through the learnable multi-scale feature aggregation mechanism, efficient cross-level feature fusion is realized, and the ability to capture subtle damage features such as concrete bursts and steel bar exposure is significantly improved; 3. Real-time inference acceleration mechanism: Through the structural optimization of the above modules, a new FastSAM-P network architecture is constructed, which can achieve a real-time processing speed of 65.11 FPS for 1080P resolution images on mobile devices while maintaining pixel-level segmentation accuracy; 4. Dedicated dataset construction method: A semantic segmentation dataset containing multi-scene fire-damaged building components is established, and a damage image library with pixel-level annotations is generated through data augmentation strategies to provide a domain-adapted benchmark dataset for model training; Compared with traditional segmentation networks, the present invention realizes the coordinated optimization of segmentation accuracy and inference speed in the building fire damage recognition task, and has significant technical advancement and industrial application value.

Claims

1. A real-time segmentation method for fire damage of building components based on FastSAM-P, characterized by comprising the following steps: Step 1: Identify the types and characteristics of two types of damage, concrete bursting and exposed steel bars on the surface of RC structural components after fire, and build corresponding data sets for network training and verification; Step 2: Develop training strategies and evaluation indicators; Step 3: Construct a real-time segmentation network architecture for building component fire damage based on FastSAM-P.

2. The method for real-time segmentation of building component fire damage based on FastSAM-P as claimed in claim 1, wherein said step 1 comprises the following specific steps: The dataset was taken from the scene after the fire and the fire laboratory, with a total of 1,500 images. These images clearly show the types and characteristics of two types of damage: concrete bursting and exposed steel bars on the surface of RC structural components after the fire. When shooting, the image size was set to a 1:1 square, and then the image was imported into the labelme software for annotation. The generated json file was converted into an XML file. Through data enhancement, the dataset was expanded to 8 times the original number, and the expanded dataset was divided into training set, validation set and test set in a ratio of 7:2:

1.

3. The method for real-time segmentation of fire damage of building components based on FastSAM-P as claimed in claim 2, characterized in that: In step 2, a multi-stage optimization strategy is used to train the FastSAM-P segmentation model to improve the recognition accuracy of concrete structure bursting and exposed reinforcement damage after fire; During the training process, the batch size was set to 16, the total number of iterations was 100, and 500 rounds of training were performed. The learning rate was controlled in stages: a linear warm-up strategy was used to alleviate gradient oscillation in the early stage, and the cosine annealing algorithm was used to gradually reduce the learning rate after the training was stable. In order to suppress overfitting, the L2 regularization term is introduced into the network optimization objective. By constraining the L2 norm of the weight parameter, the model complexity is reduced and the generalization ability is improved. The loss function adopts a weighted combination of BCE Loss + DiceLoss. Through this combination strategy, the pixel-level classification accuracy, i.e. the advantage of BCE loss, is maintained, while the regional structure consistency, i.e. the characteristic of Dice loss, is strengthened, as shown in formula (3): (3); In formula (3), λ is a hyperparameter used to balance the weight between BCE loss and Dice loss.

4. The real-time segmentation method for fire damage of building components based on FastSAM-P as described in claim 3 is characterized in that, in the step 2, a multidimensional quantitative indicator system is used to systematically evaluate the performance of the FastSAM-P model, specifically as follows: 1) The average intersection-over-union ratio is used as the core segmentation accuracy indicator, and the average overlap between the predicted area and the real labeled area in each category is calculated to quantitatively characterize the positioning accuracy of the model for the fire damage area; 2) The parameter quantity is used as a model complexity evaluation indicator to reflect the model storage requirements and computing resource consumption, and provide a basis for the feasibility of edge device deployment; 3) The frame rate is used as a real-time evaluation indicator to measure the average inference time of a single image in the NVIDIA RTX 3090 graphics card environment to evaluate the application potential of the model in real-time detection scenarios at fire scenes.

5. The method for real-time segmentation of building components with fire damage based on FastSAM-P as claimed in claim 4, characterized in that the step 3 comprises the following contents: (1) replacing the C2F module with the DSCConv module in the backbone part of YOLOv8; forming the DSCConv module by adding the Deformable Convolution module based on the SCConv module; (2) replacing the SPPF module with the RFB module in the backbone part of YOLOv8; (3) replacing the FPN module in the FastSAM network with the PPSM module: forming the PPSM by adding the attention module based on the pyramid pooling module, and extracting rich context information through the multi-scale pooling operation, thereby enhancing the network's recognition ability of multi-scale targets in the concrete fire damage segmentation task.

6. The method for real-time segmentation of fire damage of building components based on FastSAM-P as claimed in claim 5, characterized in that: In the step 3, PPSM performs global pooling operations of different scales on the input feature map to generate multi-scale feature representations, quickly extracts features through the attention module, and fuses them with the original feature map through upsampling, finally realizing the organic combination of global context information and local detail features; PPSM generates context information of multiple scales by performing pooling operations of different scales on the feature map, specifically including the following four scales of pooling: (31) Global pooling (1×1): Global average pooling is performed on the feature map to capture the global context information of the entire image; (32) Medium scale pooling (2×2): The feature map is divided into 2×2 sub-regions and pooled to extract semantic information of larger regions; (33) Smaller scale pooling (3×3 or 4×4): The feature map is divided into 3×3 or 4×4 sub-regions and pooled to focus on feature extraction of smaller regions; (34) Original scale features: The unpooled original feature map is retained to preserve high-resolution detail information; then PPSM restores the pooled features of different scales to the size of the original feature map through upsampling operations, and then splices them pixel by pixel with the original feature map.

7. The method for real-time segmentation of fire damage of building components based on FastSAM-P as claimed in claim 6, characterized in that the step (34) comprises the following specific steps: the pooled feature map of each scale is upsampled by bilinear interpolation to restore it to the same resolution as the original feature map, and then the upsampled multi-scale feature map is spliced ​​with the original feature map in the channel dimension to form a fused feature representation, and finally by fusing global context information and local detail features, the model processes large wall areas and narrow boundary areas at the same time, thereby improving the segmentation accuracy.

Citation Information

Patent Citations

  • Post-fire concrete structure damage segmentation method based on image semantic segmentation

    CN118172667A

  • Detection method for infrared thermal image damage area of coal rock

    US20250061701A1