Multi-scale interactive fusion smoke image segmentation method

Through the Transformer-CNN multi-scale interactive fusion method, a smoke segmentation data set is constructed and an encoder and decoding module is designed, which solves the problem of incomplete information capture in smoke segmentation, and realizes accurate extraction of smoke areas and early warning of fire.

CN120298689AActive Publication Date: 2025-07-11SOUTHWEST JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510348806.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing smoke segmentation method is difficult to capture sufficient context information and spatial details at the same time, resulting in inaccurate smoke segmentation, lack of data sets in real scenarios and effective global information-dependent methods.

Method used

Using the Transformer-CNN multi-scale interactive fusion method, the smoke segmentation data set is constructed and the Transformer encoder and CNN encoder are designed, combining multi-level feature complementary modules and multi-scale fusion decoding modules to achieve accurate extraction of smoke areas.

Benefits of technology

It improves the accuracy of smoke area information and can provide richer smoke area information, which is of great significance to early warning of fire.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298689A_ABST
    Figure CN120298689A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image segmentation, and discloses a multi-scale interactive fusion smoke image segmentation method, which comprises the steps of constructing a smoke segmentation data set; using the smoke segmentation data set to train a smoke segmentation model; the smoke segmentation model comprises a Transform encoder, a CNN (Convolutional Neural Network) encoder, a multi-level feature complementation module and a multi-scale fusion decoding module; the method comprises the following steps of: respectively inputting an original image into a Transform encoder and a CNN (Convolutional Neural Network) encoder; the Ftransformeri output by the Transformer encoder and the FCNNi output by the CNN encoder are jointly input into a multi-level feature complementation module to obtain a feature map Fi; and inputting the n feature maps Fi into a multi-scale fusion decoding module for fusion to obtain a smoke segmentation result. The method is used for extracting the smog area, can provide more accurate smog area information, and is of great significance to early warning of a fire.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and particularly relates to a smoke image segmentation method with multi-scale interactive fusion. Background Art

[0002] Visual fire detection methods need to identify and locate the specific positions of smoke and flames. Since smoke usually appears first in most fires, detecting smoke can provide earlier fire early warnings than flames. Smoke image segmentation is a per-pixel recognition task for the entire image. By separating smoke and its fine boundaries from the image, more abundant information can be provided. However, the characteristics of smoke such as translucency, non-rigidity, and blurred boundaries pose great challenges to recognition. Compared with other objects, smoke has many particularities and uncertainties as it changes over time and in different environments. In addition, the diffusibility of smoke and the angles and distances of video surveillance shooting are not fixed. Therefore, the scale sizes of smoke targets vary greatly in video images.

[0003] However, at present, the research and application of smoke segmentation for real scenarios are relatively few. There is a lack of smoke segmentation datasets for real scenarios and effective methods to solve the dependence on global information in smoke segmentation. Existing smoke segmentation methods are difficult to capture sufficient context information and spatial details simultaneously because of the contradiction between local spatial details and global semantic information. Therefore, the smoke segmentation model needs to be improved. Summary of the Invention

[0004] The purpose of the present invention is to propose a smoke image segmentation method with multi-scale interactive fusion for extracting smoke regions, which can provide more accurate smoke region information and is of great significance for early fire warning.

[0005] To achieve the above-mentioned invention purpose, the embodiments of the present invention provide the following technical solutions:

[0006] A smoke image segmentation method with multi-scale interactive fusion, comprising the following steps:

[0007] Step 1, constructing a smoke segmentation dataset;

[0008] Step 2, training a smoke segmentation model using the smoke segmentation dataset;

[0009] The smoke segmentation model includes a Transformer encoder, a CNN encoder, a multi-level feature complementary module, and a multi-scale fusion decoding module; the original image is respectively input into the Transformer encoder and the CNN encoder; the feature map F output by the Transformer encoder transformeri and the feature map F output by the CNN encoder CNNi are jointly input into the multi-level feature complementary module to obtain the feature map Fi , where \(i = 1, 2, \ldots, n\), and \(n\) is the number of levels of the multi-level feature complementary module; \(n\) feature maps \(F\) i are input into the multi-scale fusion decoding module for fusion to obtain the smoke segmentation result.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0011] The present invention discloses a smoke image segmentation method based on Transformer-CNN multi-scale interaction fusion, which realizes the segmentation of smoke in the early stage of a fire. First, two encoders, Transformer and CNN, are designed to fuse the ability of the CNN encoder to capture spatial context information and the ability of the Transformer encoder to extract long-range dependencies, improving the global modeling ability of the model and retaining richer local detailed features. In addition, the high-level semantic information of the image helps to determine the category to which a pixel belongs, while the low-level image information is beneficial for obtaining a fine segmentation result. For the complementary features generated by the dual channels, a multi-level feature complementary module (MFCM) is designed to flexibly focus on the information interaction between different levels of the two encoding paths, improving the expressive ability of the features. Finally, a multi-scale fusion decoding module (MFD) is designed to balance the global and local information by fusing image information at different scales and performing information interaction at feature layers of different scales.

[0012] The method designed by the present invention can extract the smoke area, can provide more accurate smoke area information, and is of great significance for the early warning of a fire. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 It is a schematic diagram of the network structures of the Transformer encoder, CNN encoder, and multi-level feature complementary module of the present invention;

[0015] Figure 2 It is a schematic diagram of the network structure of the Trans-Block layer in the Transformer encoder of the present invention;

[0016] Figure 3Schematic diagram of the network structure of the MFCM layer in the multi-level feature complementary module of the present invention;

[0017] Figure 4 Schematic diagram of the network structure of the multi-scale fusion decoding module of the present invention. Specific embodiments

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention to be protected, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0019] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance, or implying any such actual relationship or order between these entities or operations. In addition, the terms "connected", "connected", etc. can be directly connected between components, or indirectly connected through other components.

[0020] The present invention is realized through the following technical solutions. A multi-scale interactive fusion smoke image segmentation method includes the following steps:

[0021] Step 1, construct a smoke segmentation data set.

[0022] The specific steps of step 1 include the following steps:

[0023] Step 1-1, an ordinary optical camera faces the flame and smoke at multiple angles outdoors;

[0024] Step 1-2, continuously collect optical images in a time segment with a fixed duration;

[0025] Step 1-3, change the duration of the time segment (such as 0.1s, 0.2s,..., 1s), and then return to step 1-2 until step 1-3 is repeated 10 times;

[0026] Step 1-4, randomly scale and change the relative position of the ordinary optical camera, and then return to step 1-1 until step 1-4 is repeated 100 times;

[0027] Steps 1-5, replace the outdoor flame and smoke scene, and then return to Step 1-1 until Step 1-5 is repeated several times;

[0028] Step 1-6, to ensure the generalization of the model, add non-smoke videos (such as mountain roads, heavy fog, buildings, clouds, shaking leaves, etc.) to the dataset, and the length of each video is about 10 minutes;

[0029] Step 1-7, segment and label the smoke data in the collected optical images. The segmentation and labeling should ensure that the smoke area in each frame is accurately labeled as the foreground (smoke label), and the rest is labeled as the background (background label); for the non-smoke data, label it according to the corresponding category (such as buildings, leaves, etc.);

[0030] Step 1-8, jointly form a smoke segmentation dataset from the smoke image data and the non-smoke image data.

[0031] Step 2, use the smoke segmentation dataset to train the smoke segmentation model.

[0032] The smoke segmentation dataset is split into a training set and a test set according to a ratio of 9:1, and carries the corresponding label data (i.e., mask). Data augmentation is performed by means of scaling, cropping, mirroring, and adding noise to expand the scale of the training set and the test set. Specifically, the smoke image data and the non-smoke image data in the training set and the test set are expanded to 3 times the original by methods such as scaling outwards by 15%, randomly cropping to 224×224, mirror flipping, and adding Gaussian noise.

[0033] After preprocessing the image data, use the training set as the input of the smoke segmentation model and the corresponding label data as the output of the smoke segmentation model to train the model, so that the model can output smoke features after feature extraction and inference.

[0034] The smoke segmentation model is deployed based on the multi-scale interaction and fusion of Transformer-CNN. Please refer to Figure 1, the smoke segmentation model includes a Transformer encoder, a CNN encoder, a multi-level feature complementary module, and a multi-scale fusion decoding module. Among them, the Transformer encoder is used to extract global long-range dependency information; the CNN encoder is used to extract local context information, which helps to refine the local details of the segmentation mask; the multi-level feature complementary module is used to fuse the global long-range dependency information and the local context information, integrate the multi-dimensional data from the two encoders, and output features with rich levels and scales. This combination aims to provide accurate smoke segmentation, capturing a wider background and tiny details. For the problem that smoke will mix with air during the diffusion process, forming a gradually changing concentration field, making the target boundary present a transitional area with blurred edges, an edge-sensitive loss function is designed for the smoke segmentation model, which can accurately locate the edge area of the smoke, pay more attention to the edge features of the smoke, and effectively deal with problems such as blurred edges, texture loss, and dynamic changes in smoke segmentation.

[0035] Specifically, please refer to Figure 1 , the Transformer encoder includes an Embedding layer, a first semantic segmentation unit, a second semantic segmentation unit, a third semantic segmentation unit, and a fourth semantic segmentation unit connected in sequence. The first semantic segmentation unit includes a Trans-Block1 layer and a Trans-Block2 layer connected in sequence, where the input end of the Trans-Block1 layer is connected to the output end of the Embedding layer; the second semantic segmentation unit includes a PatchMerging1 layer, a Trans-Block3 layer, and a Trans-Block4 layer connected in sequence, where the input end of the Patch Merging1 layer is connected to the output end of the Trans-Block2 layer; the third semantic segmentation unit includes a Patch Merging2 layer, a Trans-Block5 layer, and a Trans-Block6 layer connected in sequence, where the input end of the Patch Merging2 layer is connected to the output end of the Trans-Block4 layer; the fourth semantic segmentation unit includes a Patch Merging3 layer, a Trans-Block7 layer, and a Trans-Block8 layer connected in sequence, where the input end of the Patch Merging3 layer is connected to the output end of the Trans-Block6 layer.

[0036] The four groups of stacked semantic segmentation units in the Transformer encoder can generate feature maps of different scales and perform global interaction at fine and coarse resolutions.

[0037] The structure of each Trans-Block layer in the Transformer encoder is the same. Please refer to Figure 2, any Trans-Block layer includes an Efficient Self-Attention layer, a Mix-FFN layer, and a LayerNorm layer connected in sequence.

[0038] The Efficient Self-Attention layer uses a self-attention matrix to perform feature pooling on the feature map, reducing the dimension of the output features to obtain multi-scale feature outputs. For the Efficient Self-Attention layer, by calculating the correlation between different parts of the input information, the input information is weighted by attention. This mechanism enables the model to obtain a global view of the input information at a shallow layer. Specifically, the input information X is mapped into a query matrix (Q), a key matrix (K), and a value matrix (V). The correlation between the query matrix and different key matrices is calculated, that is, the weight coefficients of different values are calculated; then the weighted average result with the value matrix is used as the attention value. The formula is as follows:

[0039]

[0040] where softmax is the activation function, d k is the scaling coefficient.

[0041] For the Mix-FFN layer, a Conv with a kernel size of 3×3 and a multi-layer perceptron (MLP) are mixed into each feed-forward network (FFN) for feature mapping to obtain the feature X output by the Trans-Block layer out . The formula is as follows:

[0042] X out = LayerNorm(GELU(Conv 3×3 (MLP(X in )))) + X in ;

[0043] where X in represents the output features of the Efficient Self-Attention layer, GELU is the activation function based on Gaussian error, MLP is the multi-layer perceptron, and LayerNorm is the normalization layer.

[0044] Please continue to refer to Figure 1 , the CNN encoder includes a convolutional block, a CNN Layer1 unit, a CNNLayer2 unit, a CNN Layer3 unit, and a CNN Layer4 unit connected in sequence. The convolutional block includes a Conv layer, a BN layer, and a ReLU activation function. The multi-level feature complementary module includes an MFCM1 layer, an MFCM2 layer, an MFCM3 layer, and an MFCM4 layer.

[0045] The original images are respectively input into the Embedding layer of the Transformer encoder and the convolutional blocks of the CNN encoder, and image processing is performed in the channels of their respective encoders. The feature map F output by the Trans-Block2 layer Transformer1 and the feature map F output by the CNNLayer1 unit CNN1 are input into the MFCM1 layer for interaction together. After integration, a feature map F1 with rich levels and scales is output; the feature map F output by the Trans-Block4 layer Transformer2 and the feature map F output by the CNN Layer2 unit CNN2 are input into the MFCM2 layer together to obtain the feature map F2; the feature map F output by the Trans-Block6 layer Transformer3 and the feature map F output by the CNN Layer3 unit CNN3 are input into the MFCM3 layer together to obtain the feature map F3; the feature map F output by the Trans-Block8 layer Transformer4 and the feature map F output by the CNN Layer4 unit CNN4 are input into the MFCM4 layer together to obtain the feature map F4.

[0046] The structures of each MFCM layer in the multi-level feature complementary module are the same. Please refer to Figure 3 , and any MFCM layer includes a Transformer integration channel, a CNN integration channel, and a feature complementary channel. In the Transformer integration channel, after the input feature map F Transformer passes through a 1×1Conv, weight coefficients W v1 , W k1 , W q1 are generated, and are mapped to a value matrix (V1), a key matrix (K1), and a query matrix (Q1) based on the weight coefficients; in the CNN integration channel, after the input feature map F CNN passes through a 1×1Conv, weight coefficients W v2 , W k2 , W q2 are generated, and are mapped to a value matrix (V2), a key matrix (K2), and a query matrix (Q2) based on the weight coefficients; then, in the Transformer integration channel, the key matrix (K1) and the query matrix (Q2) are fused through a Softmax activation function to generate an Attention Map1, and then the value matrix (V1) and the Attention Map1 are fused and passed through a 1×1Conv, thereby generating the feature map F Transformer`; In the CNN integration channel, after fusing the query matrix (Q1) and the key matrix (K2) with the Softmax activation function to generate Attention Map2, then fusing the value matrix (V2) and Attention Map2 and passing through 1×1Conv, thereby generating the feature map F CNN` ; Finally, for the feature map F Transformer` and the feature map F CNN` After performing a concatenation operation, passing through 3×3Conv, thereby generating the feature map F i (i = 1, 2, 3, 4).

[0047] As Figure 4 shown, the multi-scale fusion decoding module includes an upsampling layer, a concatenation layer, 3×3Conv, and 1×1Conv. Inputting the feature maps F1, F2, F3, and F4 into the upsampling layer for upsampling at different scales to generate feature maps F 1` 、F 2` 、F 3` 、F 4` ; Inputting the feature maps F 1` 、F 2` 、F 3` 、F 4` together into the concatenation layer for concatenation, and then inputting them into 3×3Conv and 1×1Conv in sequence, finally obtaining the smoke segmentation result output by the smoke segmentation model.

[0048] The edge-sensitive loss function designed in this scheme is extracted by the adaptive threshold Canny operator, which can accurately locate the edge region of the smoke. The edge-sensitive loss function is as follows:

[0049]

[0050] Among them, E T is the aversion edge region extracted by the adaptive threshold Canny operator, j represents the index of the j-th pixel point in E T ; η represents the balance coefficient; y j represents the true pixel value, p j represents the predicted pixel value; γ is the scaling parameter between the weighted cross-entropy loss of edge pixels and the gradient consistency loss;

[0051] is the weighted cross-entropy loss of edge pixels; L Gradient is the gradient consistency loss, and there is:

[0052]

[0053] Among them, N represents the number of edge pixel points, i represents the index of the i-th pixel point; p iDenote the predicted pixel value, y i Denote the true pixel value.

[0054] The weighted cross-entropy loss of edge pixels can enhance the accuracy of edge pixel classification, and the gradient consistency loss ensures more stable gradient information at the edges. The balance coefficient is used to flexibly adjust the weights of the two characteristics of smoke, namely semi-transparency and dynamic variability, to adapt to different smoke scenarios.

[0055] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for segmenting smoke images with multi-scale interactive fusion, characterized in that: It includes the following steps: Step 1, construct a smoke segmentation dataset; Step 2, use the smoke segmentation dataset to train the smoke segmentation model; The smoke segmentation model includes a Transformer encoder, a CNN encoder, a multi-level feature complementary module, and a multi-scale fusion decoding module; the original image is input into the Transformer encoder and the CNN encoder respectively; the feature map F output by the Transformer encoder transformeri and the feature map F output by the CNN encoder CNNi are jointly input into the multi-level feature complementary module to obtain the feature map F i , i = 1, 2,..., n, where n is the number of levels of the multi-level feature complementary module; the n feature maps F i are input into the multi-scale fusion decoding module for fusion to obtain the smoke segmentation result.

2. The multi-scale interactive fusion-based smoke image segmentation method according to claim 1, wherein: The Transformer encoder includes an Embedding layer, a first semantic segmentation unit, a second semantic segmentation unit, a third semantic segmentation unit, and a fourth semantic segmentation unit connected in sequence; The first semantic segmentation unit includes a Trans-Block1 layer and a Trans-Block2 layer connected in sequence, wherein the input end of the Trans-Block1 layer is connected to the output end of the Embedding layer; the Trans-Block2 layer outputs a feature map F Transformer1 ; The second semantic segmentation unit includes a Patch Merging1 layer, a Trans-Block3 layer, and a Trans-Block4 layer connected in sequence, where the input end of the Patch Merging1 layer is connected to the output end of the Trans-Block2 layer; the Trans-Block4 layer outputs a feature map F Transformer2 ; The third semantic segmentation unit includes a Patch Merging2 layer, a Trans-Block5 layer, and a Trans-Block6 layer connected in sequence. The input end of the Patch Merging2 layer is connected to the output end of the Trans-Block4 layer; the Trans-Block6 layer outputs a feature map F Transformer3 ; The fourth semantic segmentation unit includes a PatchMerging3 layer, a Trans-Block7 layer, and a Trans-Block8 layer connected in sequence, where the input end of the PatchMerging3 layer is connected to the output end of the Trans-Block6 layer; the Trans-Block8 layer outputs a feature map F Transformer4 .

3. A method for segmenting smoke images with multi-scale interaction and fusion according to claim 2, characterized in that: The CNN encoder includes a convolutional block, a CNNLayer1 unit, a CNNLayer2 unit, a CNNLayer3 unit, and a CNNLayer4 unit connected in sequence; The output feature map F of the CNNLayer1 unit CNN1 、The output feature map F of the CNNLayer2 unit CNN2 、The output feature map F of the CNN Layer3 unit CNN3 、The output feature map F of the CNNLayer4 unit CNN4 。 4. A multi-scale interactive fusion smoke image segmentation method according to claim 3, characterized in that: The multi-level feature complementary module includes an MFCM1 layer, an MFCM2 layer, an MFCM3 layer, and an MFCM4 layer; Input the feature map F Transformer1 and the feature map F CNN1 into the MFCM1 layer together to obtain the feature map F1; Input the feature map F Transformer2 and the feature map F CNN2 into the MFCM2 layer together to obtain the feature map F2; Input the feature map F Transformer3 and the feature map F CNN3 into the MFCM3 layer together to obtain the feature map F3; Input the feature map F Transformer4 and the feature map F CNN4 into the MFCM4 layer together to obtain the feature map F4.

5. A method for segmenting smoke images with multi-scale interactive fusion according to claim 4, characterized in that: The multi-scale fusion decoding module includes an upsampling layer, a splicing layer, a 3×3Conv, and a 1×1Conv; Input the feature maps F1, F2, F3, and F4 into the upsampling layer for upsampling at different scales to generate feature maps F with the same scale. 1` , feature map F 2` , feature map F 3` , feature map F 4` ; Input the feature maps F 1` , feature map F 2` , feature map F 3` , feature map F 4` into the splicing layer for splicing together, and then input them into 3×3Conv and 1×1Conv in sequence, and finally obtain the smoke segmentation result output by the smoke segmentation model.

Citation Information

Patent Citations

  • Multi-scale fusion smoke segmentation method based on depth separable convolution

    CN116012395A

  • Image processing method and system based on double-branch multi-scale semantic segmentation network

    CN116580241A

  • Multi-scale Transform image semantic segmentation method based on convolution local enhancement

    CN117058392A

  • Image semantic segmentation method fusing space detail context and multi-scale interaction

    CN118230323A

  • Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment

    WO2024230038A1