Edge-guided multistage fire semantic segmentation network construction method

Through the edge-guided multi-level fire semantic segmentation network, ResNet-50 is used to extract RGB-T image features and combined with a multi-level cross-modal feature fusion module to solve the problem of insufficient feature fusion of multimodal data in fire recognition and achieve efficient and accurate flame segmentation.

CN120689637APending Publication Date: 2025-09-23GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510702688.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing fire recognition models have difficulty identifying small flames in low visibility and low resolution conditions. Single-modal models easily misidentify high-temperature objects. Multimodal models ignore specific features and insufficient cross-modal feature fusion, resulting in difficulty in flame recognition.

Method used

An edge-guided multi-level fire semantic segmentation network is designed. The ResNet-50 backbone network is used to extract RGB-T image features. Multi-level cross-modal feature fusion is performed through the boundary enhancement module, fusion activation module and cross-localization module. Edge supervision is introduced to improve the accuracy of flame segmentation.

Benefits of technology

The accuracy and robustness of flame segmentation are improved, and the detailed information and semantic information of multimodal data are fully utilized to achieve efficient and accurate segmentation of flames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689637A_ABST
    Figure CN120689637A_ABST
Patent Text Reader

Abstract

An edge-guided multistage fire semantic segmentation network construction method mainly comprises the following steps: respectively extracting RGB and TIR image features through a dual ResNet50 backbone network, generating five groups of cross-modal features, and enhancing the features through a shared convolutional layer; a boundary enhancement module is designed, flame features are captured by adopting multi-scale cavity convolution, and the edge detail segmentation precision is improved in combination with an edge supervision mechanism; constructing a fusion activation module, and integrating a global context and a coordinate attention mechanism to realize fine region activation; a cross positioning module is developed, and flame accurate space positioning is achieved through feature cross-correlation operation. And a joint loss function is adopted, and semantic loss Lsem constructed by cross entropy loss and Lovasz-softmax loss is combined with boundary loss Ledge to form a total loss function to improve the network segmentation performance. According to the method, through mechanisms such as multi-modal information fusion, multi-scale feature capture, edge detail attention and multi-target supervision, the flame segmentation task is excellent, and an efficient and accurate solution is provided for fire monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular relates to an edge-guided multi-level fire semantic segmentation network construction method. Background Art

[0002] In current fire identification tasks, fire semantic segmentation models can provide pixel-level fire location predictions, improving the accuracy of early fire identification. Currently, fire segmentation can be divided into single-modal segmentation networks and multimodal fusion fire segmentation networks based on data type. Single-modal segmentation networks are mainly based on RGB images and thermal infrared (TIR) ​​images. These two fire identification models have achieved good performance, but they still have certain limitations in certain scenarios. For example, in conditions of low visibility and low image resolution, some small flames are difficult to identify; single-modal models based on TIR images will identify some high-temperature objects as flames; and single-modal models have difficulty understanding contextual information when there is occlusion, making flame identification difficult.

[0003] Most RGB-T-based multimodal fusion fire recognition models focus on developing fusion strategies to learn shared representations of the two modalities, ignoring the influence of model-specific features. Furthermore, multi-level cross-modal feature fusion between the encoder and decoder typically uses the same module, ignoring the attributes of features at different levels.

[0004] To solve the above problems, the present invention discloses an edge-guided multi-level fire semantic segmentation network construction method. Summary of the Invention

[0005] In response to the problems existing in multimodal data flame segmentation, such as insufficient feature fusion, insufficient utilization of specific features and a single multi-level feature fusion strategy, this paper discloses an edge-guided multi-level fire semantic segmentation network construction method, aiming to improve the accuracy, robustness and detail representation ability of flame segmentation.

[0006] The advantage of this invention lies in the design of three modules for multi-level cross-modal feature fusion between the encoder and decoder, which fuse RGB-T data to extract high, medium, and low-level features. At the same time, edge supervision is introduced in the high-level feature module, allowing the network to pay more attention to edge details and improve the accuracy of the segmentation results. To achieve the above technical objectives, the present invention adopts the following technical solutions.

[0007] The technical solution of the present invention is: a method for constructing an edge-guided multi-level fire semantic segmentation network, the specific steps of which are as follows:

[0008] Step 1: Two parallel ResNet-50 backbone networks are used to extract features from RGB images and TIR images, generating five sets of basic cross-modal features. A shared convolutional layer is then used to enhance the RGB and TIR image features, reducing the number of model parameters while allowing the model to learn effective features from both modal images, preparing for modal fusion.

[0009] Step 2: Design a Boundary Enhanced Module (BEM) to extract valuable information from multimodal low-level features. In addition, edge supervision is used to enable the model to find regions of interest more quickly. This module acts on the first two sets of basic cross-modal features.

[0010] Step 3: Design a Fusion Activation Module (FAM), which can smooth features, suppress background interference, and further highlight the exact area of ​​the flame in the multimodal intermediate features. This module acts on the two intermediate sets of basic cross-modal features.

[0011] Step 4: Design a cross-localized module (CLM), which deeply fuses the features of RGB images and thermal infrared images to achieve accurate spatial positioning of the segmented region from multimodal high-level features. This module acts on the last set of basic cross-module features.

[0012] Step 5: Use five decoders to decode the five basic cross-modal features after the three feature modules to generate the final flame segmentation result.

[0013] In Step 1, ResNet-50 is used as the backbone network to extract features from RGB images and TIR images, generating five sets of basic cross-modal features with feature numbers of {64, 256, 512, 1024, 2048}, respectively. Subsequently, shared convolutional layers are used to enhance the RGB and TIR image features, while reducing the number of model parameters. The output feature numbers are {64, 128, 256, 256, 512}, respectively.

[0014] In the above Step 2, the boundary enhancement module is divided into feature fusion and multi-scale information extraction. For the feature fusion part, the RGB image features are fused by element addition. and TIR imaging features , and then use element multiplication to fuse the RGB image and thermal infrared image features with it twice, and finally concatenate the obtained features and output them, which is recorded as . This process can be expressed as:

[0015]

[0016] in represents element-wise multiplication, represents element-wise addition, Represents feature splicing, and i represents the position of the module in the model.

[0017] In the multi-scale information extraction part, first, the fused features After four parallel 3×3 convolution modules, the receptive fields of the convolution modules are {1, 3, 5, 7}. Then, the features of different scales are concatenated, and the features are linearly combined through a 1×1 convolution and optimized through residual connection. Finally, a 3×3 convolution is used to output the features. , an edge head module is added to the output part and edge monitoring is applied to obtain more accurate edge information and achieve edge enhancement of target information. This process can be expressed as:

[0018]

[0019] in represents element-wise multiplication, represents element-wise addition, Represents feature concatenation, conv() represents a two-dimensional 3×3 convolution, and i represents the position of the module in the model.

[0020] In the above Step 3, the fusion activation module is divided into spatial attention and global context information fusion. The first part is to use the coordinate attention module to process the comprehensive information features and generate a spatial attention map. First, two 1×1 convolution kernels are used to linearize the input RGB image and thermal infrared image features. Then, the linearized features of the RGB image and infrared image are combined using element addition operation. Finally, the features are input into the coordinate attention module to obtain , so that the network can learn the spatial relationship of features and better integrate and understand the features of each position in the two images. This process can be expressed as:

[0021]

[0022] Where CoordAtt represents the Coordinate Attention mechanism, represents element addition, conv() represents a two-dimensional 1×1 convolution, and i represents the position of the module in the model.

[0023] The second part is global context information fusion, which further fuses the RGB image and thermal infrared image features through geometric feature matching and semantic context information. First, similarly, two 1×1 convolution kernels are used to linearize the input RGB image and thermal infrared image features. Then, the linearized features of the RGB image and the infrared image are combined using element-wise multiplication to obtain , and use element addition to add features With spatial attention map Input the combination into the global context module to get , thereby highlighting the global context features and finally achieving fine-grained regional activation of features. This process can be expressed as:

[0024]

[0025] Among them, GC stands for Global Context Model. represents element addition, conv() represents a two-dimensional 1×1 convolution, and i represents the position of the module in the model.

[0026] In Step 4, the cross-localization module is divided into two parts: cross-layer association and salient feature map generation. The first part is cross-layer association. First, the RGB image features are respectively and TIR image features The dimension of is reshaped to C×HW, and the reshaped RGB image features are transposed, and its dimension becomes HW×C. Then a learnable linear weight matrix W is defined for the RGB image features c , so that subsequent operations can better capture the correlation of different modal features and enhance the robustness of the module. Finally, the correlation matrix R of the RGB image and the TIR image is calculated by element multiplication to achieve cross-layer correlation of multimodal image features. The above process can be expressed as:

[0027]

[0028] in Represents element-wise multiplication, and reshape() represents a dimension reshaping operation.

[0029] The second part is the generation of salient feature maps. In order to deeply fuse the RGB image features and thermal infrared image features, first, the correlation matrix R is normalized along the rows and columns using the softmax function to determine the salient area position of the high-level features. Then the correlation matrix R is matrix-multiplied with the reshaped RGB image features and TIR image features to obtain the feature and . This process can be expressed as:

[0030]

[0031] in Represents element multiplication, reshape() represents dimension reshaping operation, and softmax() represents normalization using softmax function.

[0032] Finally, the features 、 Compared with the original RGB image features and TIR image features Add them together to complete the feature location of the salient area, reshape the result to change the dimension to C×H×W, and use 3×3 two-dimensional convolution operation to integrate the features, and finally obtain the fused salient features. , expressed as:

[0033]

[0034] The cross-localization module effectively fuses the features of RGB and TIR images through a cross-attention mechanism. This module captures the correlation between the two modalities at the feature level and highlights important features through the attention mechanism, thereby improving the expressiveness and discriminability of features. This makes the CLM module crucial in multimodal feature fusion tasks, providing richer feature representations for subsequent analysis and recognition tasks.

[0035] The beneficial effects of the present invention are as follows: a semantic segmentation method is used to segment flames in an image to realize fire recognition. Three modules, namely a boundary enhancement module (BEM), a fusion activation module (FAM), and a cross localization module (CLM), are designed to fuse and enhance the features of multimodal data (RGB image and TIR image) at different levels, making full use of low-level detail information and high-level semantic information, solving the problem of ignoring feature attributes of different levels of multi-level cross-modal features of multimodal data in fire recognition, and introducing edge supervision for flame segmentation for the first time, thereby realizing efficient and accurate segmentation of flames. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is the network model flow chart of the present invention

[0037] Figure 2 Schematic diagram of the boundary enhancement module of the present invention

[0038] Figure 3 Schematic diagram of the fusion activation module of the present invention

[0039] Figure 4 Schematic diagram of the cross positioning module of the present invention DETAILED DESCRIPTION

[0040] The present invention will be described in detail with reference to the accompanying drawings. Examples of embodiments of the embodiments described in detail are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0041] A method for constructing an edge-guided multi-level fire semantic segmentation network, the specific steps are as follows:

[0042] Step 1: Obtain the RGB image and TIR image to be recognized; use two parallel ResNet-50 backbone networks to extract features from the RGB image and TIR image respectively, generate five basic cross-modal features, and use a shared convolutional layer to enhance the RGB and TIR image features respectively, reducing the number of model parameters and allowing the model to learn the effective features of the two modal images, making preliminary preparations for modal fusion. Figure 1 As shown, specifically:

[0043] Step 1.1: Use ResNet-50 as the backbone network to extract features from RGB images and TIR images respectively, and generate 5 sets of basic cross-modal features , the feature dimensions are {64, 256, 512, 1024, 2048} respectively;

[0044] Step 1.2: Use a shared convolutional layer to enhance the RGB and TIR image features, reducing the number of model parameters while allowing the model to learn effective features of both modal images. The dimensions of the 10 basic cross-modal features output are {64, 128, 256, 256, 512}, respectively.

[0045] Step 2: Design a Boundary Enhanced Module (BEM) to fuse multimodal low-level features , extracting valuable information from multimodal low-level features, and using edge supervision to help the model find the area of ​​interest faster, combined with Figure 2 As shown, specifically:

[0046] Step 2.1: Feature fusion: First, fuse the RGB image features by element-wise addition and TIR imaging features , and then use element multiplication to fuse the RGB image and thermal infrared image with it twice, and finally output the obtained feature splicing, which is recorded as . This process can be expressed as:

[0047]

[0048] in represents element-wise multiplication, represents element-wise addition, Represents feature splicing, and i represents the position of the module in the model.

[0049] Step 2.2: Multi-scale information extraction: First, the fused features After four parallel 3×3 convolution modules, the receptive fields of the convolution modules are {1, 3, 5, 7} respectively. Then, the features of different scales are concatenated, linearly combined using a 1×1 convolution, and optimized through residual connections. Finally, a 3×3 convolution is used to output the features. , this process can be expressed as:

[0050]

[0051] in represents element-wise multiplication, represents element-wise addition, Represents feature concatenation, conv() represents a two-dimensional 3×3 convolution, and i represents the position of the module in the model.

[0052] Step 2.3: Edge Supervision: Attach an edge head module to the output part and apply edge monitoring, using the cross entropy loss function (CrossEntropyLoss) boundary loss L edge , in order to obtain more accurate edge information and achieve edge enhancement of target information.

[0053] Step 3: Design a Fusion Activation Module (FAM) to fuse mid-level features , capturing the local structure and texture information of the flame, further activating the exact area of ​​the flame in the multimodal intermediate features of different scales, combined with Figure 3 As shown, specifically:

[0054] Step 3.1: Spatial attention map generation: First, use two 1×1 convolution kernels to linearize the input RGB image and TIR image features. Then, use element-wise addition to combine the linearized features of the RGB image and infrared image. Finally, input the features into the coordinate attention module to obtain , so that the network can learn the spatial relationship of features and better integrate and understand the features of each position in the two images. This process can be expressed as:

[0055]

[0056] Where CoordAtt represents the Coordinate Attention mechanism, represents element addition, conv() represents a two-dimensional 1×1 convolution, and i represents the position of the module in the model.

[0057] Step 3.2: Global context information fusion: First, similarly, two 1×1 convolution kernels are used to linearize the input RGB image and TIR image features. Then, the linearized features of the RGB image and the infrared image are combined using element-wise multiplication to obtain , and use element addition to add features With spatial attention map Input the combination into the global context module to get , thereby highlighting the global context features and finally achieving fine-grained regional activation of features. This process can be expressed as:

[0058]

[0059] Among them, GC stands for Global Context Model. represents element addition, conv() represents a two-dimensional 1×1 convolution, and i represents the position of the module in the model.

[0060] Step 4: Design a cross-localized module (CLM) to transform the RGB image features and thermal infrared image characteristics Perform deep fusion to achieve accurate spatial positioning of the flame area from multimodal advanced features, combined with Figure 4 As shown, specifically:

[0061] Step 4.1: Cross-layer correlation: First, the RGB image features are and TIR image features The dimension of is reshaped to C×HW, and the reshaped RGB image features are transposed, and its dimension becomes HW×C. Then a learnable linear weight matrix W is defined for the RGB image features c , so that subsequent operations can better capture the correlation of different modal features and enhance the robustness of the module. Finally, the feature correlation matrix R of the RGB image and the TIR image is calculated by element multiplication to achieve cross-layer correlation of multimodal image features. The above process can be expressed as:

[0062]

[0063] in Represents element-wise multiplication, and reshape() represents a dimension reshaping operation.

[0064] Step 4.2: Significant feature map: In order to achieve deep fusion of RGB image features and TIR image features, first, the correlation matrix R is normalized along the rows and columns using the softmax function to determine the salient area position of the high-level features. Then the correlation matrix R is matrix multiplied with the reshaped RGB image features and TIR image features to obtain the feature and . This process can be expressed as:

[0065]

[0066] in Represents element multiplication, reshape() represents dimension reshaping operation, and softmax() represents normalization using softmax function.

[0067] Finally, the features 、 Compared with the original RGB image features and TIR image features Add them together to complete the feature location of the salient area, reshape the result to change the dimension to C×H×W, and use 3×3 two-dimensional convolution operation to integrate the features, and finally obtain the fused salient features. , expressed as:

[0068]

[0069] in represents element-wise addition, reshape() represents dimension reshaping operation, and conv() represents two-dimensional 3×3 convolution.

[0070] Step 5: Use five decoders to fusion the five cross-modal features of the three feature modules Decode and generate the final flame segmentation result, such as Figure 1 As shown, specifically:

[0071] Step 5.1: Use five decoders to fusion five cross-modal features after three feature modules Decoding is performed, where each decoder is fused with the fusion features of the previous layer input using element addition to generate the flame segmentation result.

[0072] Step 5.2: During model training, the loss function L is constructed using the cross entropy loss function (CrossEntropyLoss) and the Lovasz-softmax loss function sem Supervise the model segmentation results and jointly construct the model loss function L with the boundary loss mentioned in Step 2.3 total , which can be expressed as:

[0073]

[0074] The present invention proposes a flame segmentation network with significant advantages in many aspects. It uses the ResNet50 backbone network to extract RGB and TIR image features respectively, obtains 5 groups of basic cross-modal features, and uses the shared parameter convolution layer for feature enhancement, fully mining the complementary information of the two modal data. The first two groups of features extracted by the network are defined as low-level features, the middle two groups of features are defined as intermediate features, and the last group of features is defined as high-level features. For this purpose, three multi-level cross-modal feature fusion modules, BEM, FAM, and CLM, are carefully designed to achieve effective feature fusion at different levels, thereby enhancing feature quality and information transmission efficiency: the BEM module introduces multi-scale hole convolution to capture multi-scale flame features; the FAM module combines global context and coordinate attention mechanism to capture long-distance dependencies and spatial information; the CLM module deeply mines feature associations through feature cross-correlation operations to improve feature expression capabilities. At the same time, the edge supervision mechanism is introduced to make the network pay more attention to the details of the flame edge and improve the boundary accuracy of the segmentation result. A joint loss function is used in model training, and the semantic loss L constructed by the cross entropy loss and the Lovasz-softmax loss is combined. sem With the boundary loss L edge Combined to form the total loss function L total , enabling the model to more effectively learn flame characteristics, improving segmentation accuracy and robustness. In summary, the network proposed in this paper integrates the advantages of multimodal information fusion, multi-scale feature capture, edge detail attention, and multi-target supervision, and performs well in flame segmentation tasks, providing an efficient and accurate solution for fire monitoring.

[0075] The above embodiments are intended to illustrate the present invention only and are not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions fall within the scope of the present invention, and the scope of patent protection of the present invention shall be defined by the claims.

[0076] The technical contents not described in detail in this invention are all well-known technologies.

Claims

1. A method for constructing an edge-guided multi-level fire semantic segmentation network, characterized by: Step 1: Input the RGB image and TIR image to be recognized, and use two parallel ResNet-50 backbone networks to extract features respectively, obtaining five sets of basic cross-modal features. Then, a shared convolutional layer is used to enhance the RGB and TIR image features and reduce the number of model parameters. Step 2: Design a boundary enhancement module and introduce multi-scale dilated convolution to extract valuable information from multimodal low-level features. In addition, use the cross entropy loss function L edge Perform edge supervision to help the model find the region of interest faster. This module acts on the first two sets of basic cross-modal features. Step 3: Design a fusion activation module that combines global context and coordinate attention mechanisms to capture long-range dependencies and spatial information, achieving fine-grained regional activation of flame features. This module acts on the two intermediate sets of basic cross-modal features. Step 4: Design a cross-localization module to deeply fuse the features of RGB images and thermal infrared images to achieve accurate spatial localization of the segmented region from multimodal high-level features. This module acts on the last set of basic cross-module features. Step 5: Use five decoders to fusion the five cross-modal features of the feature module Decode and generate the final flame segmentation result, where semantic segmentation uses the cross entropy loss function (CrossEntropyLoss) and Lovasz-softmax loss function to construct the loss function L sem The model is supervised and compared with the boundary loss function L edge The joint construction model training loss function is expressed as: 。 2. The edge-guided multi-level fire semantic segmentation network construction method according to claim 1 is characterized by: Step 2 is specifically as follows: Step 2.1: Feature fusion: First, fuse the RGB image features by element-wise addition and TIR imaging features , and then use element multiplication to fuse the RGB image and thermal infrared image with it twice, and finally output the obtained feature splicing, which is recorded as , this process can be expressed as: ; in represents element-wise multiplication, represents element-wise addition, Represents feature splicing, i represents the position of the module in the model; Step 2.2: Multi-scale information extraction: First, the fused features After four parallel 3×3 convolution modules, the receptive fields of the convolution modules are {1, 3, 5, 7} respectively; then, the features of different scales are concatenated, linearly combined using a 1×1 convolution, and optimized through residual connections; finally, a 3×3 convolution is used to output the features , this process can be expressed as: ; Where conv( ) represents a two-dimensional 3×3 convolution; Step 2.3: Edge Supervision: Attach an edge head module to the output part and apply edge monitoring, using the cross entropy loss function (CrossEntropyLoss) boundary loss L edge , in order to obtain more accurate edge information and achieve edge enhancement of target information.

3. The edge-guided multi-level fire semantic segmentation network construction method according to claim 1 is characterized by: Step 3 is as follows: Step 3.1: Spatial attention map generation: First, use two 1×1 convolution kernels to align the input RGB image features. and TIR image features Linearize, then use element addition operation to combine the linearized features of RGB image and infrared image; finally, input the features into the coordinate attention module to obtain , this process is expressed as: ; Where CoordAtt stands for Coordinate Attention mechanism; Step 3.2: Global context information fusion: First, the RGB image features after 1×1 two-dimensional convolution are and TIR image features Perform element-wise multiplication to obtain features Then, we use element-wise addition to add the features and features Input the combination into the global context module to get , in order to highlight the global context features and finally realize the fine area activation of the features. This process is expressed as: ; Among them, GC stands for Global Context Model.

4. The edge-guided multi-level fire semantic segmentation network construction method according to claim 1 is characterized by: Step 4 is as follows: Step 4.1: Cross-layer correlation: First, the RGB image features are and TIR image features The dimension of is reshaped to C×HW, and the reshaped RGB image features are transposed, and its dimension becomes HW×C; then a learnable linear weight matrix W is defined for the RGB image features c Finally, the feature correlation matrix R of the RGB image and the TIR image is calculated by element multiplication. The above process can be expressed as: ; Where reshape() represents the dimension reshaping operation; Step 4.2: Significant feature map: First, use the softmax function to normalize the correlation matrix R along the rows and columns respectively, and then perform matrix multiplication of the correlation matrix R with the reshaped RGB image features and TIR image features to obtain the features and , this process can be expressed as: ; Where softmax() represents the use of softmax function for normalization; Finally, the features 、 Compared with the original RGB image features and TIR image features Add them together and reshape their dimensions to C×H×W, and use 3×3 two-dimensional convolution operation to integrate features, and finally obtain the fused significant features , expressed as: ; Where conv( ) represents a two-dimensional 3×3 convolution.