Risk source multi-scale identification method based on improved attention mechanism

Through the cross-stage local network and feature pyramid network extracting multi-scale features, combined with the Vision Transformer model that improves attention mechanism, the problem of difficulty in identifying traditional fire detection methods in polluted environments is solved, and efficient and accurate detection of fire risk sources is achieved.

CN120495782APending Publication Date: 2025-08-15NANTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510680328.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional fire detection methods are prone to blockage in polluted environments, making it difficult to accurately obtain air samples, and deep learning models have high dependence on large-scale data, high labeling costs, insufficient generalization capabilities, and difficult to identify small targets and irregular morphological risk sources.

Method used

Multi-scale features are extracted by cross-stage local network, feature pyramids and path aggregation network, combined with improved multi-head self-attention, channel and spatial attention mechanisms, and global semantic feature modeling of image is performed through the Vision Transformer model to generate risk source type classification results.

Benefits of technology

It improves the accuracy and robustness of identification of small targets and irregular morphological risk sources, reduces the amount of calculation, and achieves seamless conversion and efficient detection in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495782A_ABST
    Figure CN120495782A_ABST
Patent Text Reader

Abstract

The invention discloses a risk source multi-scale identification method based on an improved attention mechanism, and relates to the field of image identification, and the method comprises the specific steps: firstly obtaining a preset number of smoke and fire image data sets marked with risk source type labels; secondly, multi-scale feature information is effectively extracted and fused through a feature pyramid network, a path aggregation network and a cross-stage local network combination structure; and then, a Vision Transform model is used to carry out image block segmentation on the feature image, and through an improved multi-head self-attention mechanism, a channel attention mechanism and a space attention mechanism, accurate modeling of global semantic features of the image and efficient fusion of cross-dimension features are realized. According to the method, a small target risk source and an irregular form risk source can be accurately detected in different environments, and the accuracy and robustness of risk source detection are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a multi-scale risk source recognition method based on an improved attention mechanism. Background Art

[0002] There are many causes of fire, and the harm and losses caused by fire are becoming increasingly serious. Therefore, detecting fire in advance can help reduce harm and minimize losses.

[0003] Currently, traditional methods for early fire detection rely primarily on temperature and humidity sensors or smoke sensors. These sensors collect air samples on-site and monitor changes in ambient temperature, humidity, or smoke concentration to provide real-time warnings. However, in polluted environments, sensor sampling holes can easily become clogged, making it difficult to obtain accurate air samples, thus hindering early fire detection. Furthermore, temperature and humidity sensors or smoke sensors are typically only suitable for small or indoor spaces.

[0004] Although traditional deep learning models can improve the accuracy of fire detection, they have problems such as strong dependence on large-scale data, high labeling costs, and insufficient generalization capabilities in unseen abnormal scenarios. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a multi-scale risk source identification method based on an improved attention mechanism.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A multi-scale risk source identification method based on an improved attention mechanism is proposed. According to steps S1 and S2, a risk source identification model for identifying the risk source type corresponding to a fireworks image is constructed. According to step A, multi-scale risk source identification is implemented:

[0008] Step S1: obtaining a preset number of fireworks image datasets that have been labeled with risk source type labels, and forming a sample set with a single fireworks image and its corresponding risk source type label as a sample;

[0009] Step S2: Constructing a model to be trained that includes a feature extraction module, an enhanced feature extraction module, and an output module. Based on the sample set, the model to be identified is trained with the fireworks image in the sample as input and the risk source label corresponding to the fireworks image in the sample as output, thereby obtaining a risk source identification model for identifying the risk source type corresponding to the fireworks image.

[0010] In step A, the fireworks image to be identified is input into the target backbone network containing the trained risk source identification model to obtain the risk source type corresponding to the fireworks image to be identified.

[0011] Furthermore, the input end of the feature extraction module constitutes the input end of the risk source identification model, the output end of the feature extraction module is sequentially connected in series with the enhanced feature extraction module and the output module; the output end of the output module constitutes the output end of the risk source identification model;

[0012] The feature extraction module is used to extract and output the multi-scale fusion features of the input image; the enhanced feature extraction module is used to extract and output the semantic features of the input image based on the multi-scale fusion features; and the output module is used to generate a classification result of the image risk source type based on the semantic features of the input image.

[0013] Furthermore, the feature extraction module includes a cross-stage local network module, a feature pyramid module, and a path aggregation module;

[0014] The input end of the cross-stage local network module constitutes the input end of the feature extraction module, the output end of the cross-stage local network module is connected in series with the feature pyramid module and the path aggregation module in sequence; the output end of the path aggregation module constitutes the output end of the feature extraction module;

[0015] The cross-stage local network module receives the input image, uses the CSP1_X module and the CSP2_X module to generate basic features containing different levels, and outputs the basic features to the feature pyramid module; the feature pyramid module is a top-down feature pyramid structure with decreasing image scale. Based on the high-level strong semantic basic features at the top layer and the low-level high-resolution basic features at the bottom layer, the feature pyramid module uses a top-down feature fusion operation to fuse the high-level strong semantic basic features layer by layer into the low-level high-resolution basic features through upsampling to generate a multi-scale feature map containing features at different levels; the path aggregation module is a bottom-up path aggregation structure. For the multi-scale feature map containing features at different levels, the low-level multi-scale feature map is fused layer by layer into the high-level multi-scale feature map through convolution and downsampling operations to generate a multi-scale fused feature map containing features at different levels, that is, the multi-scale fused feature map corresponding to the input image.

[0016] Furthermore, the enhanced feature extraction module includes an embedding vector module, a position encoding module, a Transformer module, and an MLP module, and the Transformer module includes a self-attention module, a channel attention module, and a spatial attention module;

[0017] The input end of the embedding vector module constitutes the input end of the enhanced feature extraction module, the output end of the embedding vector module is connected in series with the position encoding module, the self-attention module, the channel attention module, the spatial attention module, and the MLP module; the output end of the MLP module constitutes the output end of the enhanced feature extraction module;

[0018] The embedding vector module receives the multi-scale fusion feature map, and based on the preset sub-block feature map size, uses Patch cutting to divide the multi-scale fusion feature map into sub-block feature maps, further uses linear projection to obtain the embedding vector corresponding to each sub-block feature map, and outputs the embedding vector corresponding to each sub-block feature map to the position encoding module; the position encoding module uses sine and cosine functions to generate the position code corresponding to each sub-block feature map, and outputs the embedding vector and position code corresponding to each sub-block feature map to the self-attention module; the self-attention module uses the multi-head attention mechanism to extract the global features corresponding to the multi-scale fusion feature map based on the embedding vector and position code corresponding to each sub-block feature map, and generates a multi-scale self-attention feature map. ; The channel attention module performs global average pooling and global maximum pooling on the multi-scale self-attention feature map based on the channel attention mechanism, and uses the multi-layer perceptron to generate the weight coefficients of each feature channel corresponding to the multi-scale self-attention feature map, and further uses each weight coefficient to recalibrate the multi-scale self-attention feature map to generate a weighted multi-scale self-attention feature map; The spatial attention module uses the spatial attention mechanism to perform average pooling and maximum pooling on the weighted multi-scale self-attention feature map, and uses splicing and convolution operations to generate a multi-scale comprehensive attention feature map; For the multi-scale comprehensive attention feature map, the MLP module uses nonlinear transformation and feature refinement to obtain the global semantic feature vector, that is, the global semantic feature vector corresponding to the input image.

[0019] Furthermore, the output module processes the global semantic feature vector using a fully connected layer to generate a classification result of the risk source type, that is, the risk source type corresponding to the input image.

[0020] Furthermore, after training the risk source identification model, it also includes: using TensorRT to convert the risk source identification model into ONNX format for reasoning acceleration.

[0021] The beneficial effects brought about by adopting the above technical solution are:

[0022] (1) The present invention adopts a cross-stage local network to obtain basic features at different levels, effectively reducing the redundant information and computational complexity of the feature graph, and solving the problem of difficulty in identifying "small target" risk sources in the prior art;

[0023] (2) The present invention adopts feature pyramid and path aggregation to effectively integrate high-level strong semantic basic features and low-level high-resolution basic features, thereby enhancing the robustness of the model to complex backgrounds and noise;

[0024] (3) The present invention adopts multi-head attention mechanism, channel attention mechanism, and spatial attention mechanism to effectively integrate multi-scale features, significantly improving the recognition ability of multi-scale fireworks targets, with higher classification accuracy and generalization ability, solving the problem of difficulty in identifying "irregular morphology" risk sources in the existing technology;

[0025] (4) The present invention converts the risk source identification model into ONNX format through TensorRT for inference acceleration, so that the risk source identification model can be seamlessly converted between different deep learning frameworks and has good detection effects in different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Flowchart of the present invention;

[0027] Figure 2 This is the fireworks image dataset of the present invention;

[0028] Figure 3 This is a structural diagram of the feature pyramid module and path aggregation module of the present invention;

[0029] Figure 4 This is the architecture diagram of the enhanced feature extraction module of the present invention;

[0030] Figure 5 This is a flowchart of the multi-level attention mechanism of the present invention;

[0031] Figure 6 This is a comparison chart of the inference time of the present invention in different network structures. DETAILED DESCRIPTION

[0032] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0033] refer to Figure 1 A multi-scale risk source identification method based on an improved attention mechanism is proposed. According to steps S1 and S2, a risk source identification model for identifying the risk source type corresponding to the fireworks image is constructed. According to step A, multi-scale risk source identification is implemented:

[0034] Step S1, reference Figure 2 , through collection, web crawling and other methods, a dataset of 20,000 fireworks images with risk source type labels was obtained. The dataset included scenes such as power stations, buildings, transportation, and forests, and a sample set was formed with single fireworks images and their corresponding risk source type labels as samples;

[0035] Step S2: Constructing a model to be trained that includes a feature extraction module, an enhanced feature extraction module, and an output module. Based on the sample set, the model to be identified is trained with the fireworks image in the sample as input and the risk source label corresponding to the fireworks image in the sample as output, thereby obtaining a risk source identification model for identifying the risk source type corresponding to the fireworks image.

[0036] In step A, the fireworks image to be identified is input into the target backbone network containing the trained risk source identification model to obtain the risk source type corresponding to the fireworks image to be identified.

[0037] Furthermore, the input end of the feature extraction module constitutes the input end of the risk source identification model, the output end of the feature extraction module is sequentially connected in series with the enhanced feature extraction module and the output module; the output end of the output module constitutes the output end of the risk source identification model;

[0038] The feature extraction module is used to extract and output the multi-scale fusion features of the input image; the enhanced feature extraction module is used to extract and output the semantic features of the input image based on the multi-scale fusion features; and the output module is used to generate a classification result of the image risk source type based on the semantic features of the input image.

[0039] Further, refer to Figure 3 , the feature extraction module includes a cross-stage local network module, a feature pyramid module, and a path aggregation module;

[0040] The input end of the cross-stage local network module constitutes the input end of the feature extraction module, the output end of the cross-stage local network module is connected in series with the feature pyramid module and the path aggregation module in sequence; the output end of the path aggregation module constitutes the output end of the feature extraction module;

[0041] The cross-stage local network module receives the input image, uses the CSP1_X module and the CSP2_X module to generate basic features containing different levels, and outputs the basic features to the feature pyramid module; the feature pyramid module is a top-down feature pyramid structure with decreasing image scale. Based on the high-level strong semantic basic features at the top layer and the low-level high-resolution basic features at the bottom layer, the feature pyramid module uses a top-down feature fusion operation to fuse the high-level strong semantic basic features layer by layer into the low-level high-resolution basic features through upsampling to generate a multi-scale feature map containing features at different levels; the path aggregation module is a bottom-up path aggregation structure. For the multi-scale feature map containing features at different levels, the low-level multi-scale feature map is fused layer by layer into the high-level multi-scale feature map through convolution and downsampling operations to generate a multi-scale fused feature map containing features at different levels, that is, the multi-scale fused feature map corresponding to the input image.

[0042] Further, refer to Figure 4 and Figure 5 Based on the Vision Transformer (ViT) model, the enhanced feature extraction module includes an embedding vector module, a position encoding module, a Transformer module, and an MLP module, and the Transformer module includes a self-attention module, a channel attention module, and a spatial attention module;

[0043] The input end of the embedding vector module constitutes the input end of the enhanced feature extraction module, the output end of the embedding vector module is connected in series with the position encoding module, the self-attention module, the channel attention module, the spatial attention module, and the MLP module; the output end of the MLP module constitutes the output end of the enhanced feature extraction module;

[0044] The embedding vector module receives the multi-scale fusion feature map, and based on the preset sub-block feature map size, uses Patch cutting to divide the multi-scale fusion feature map into sub-block feature maps, further uses linear projection to obtain the embedding vector corresponding to each sub-block feature map, and outputs the embedding vector corresponding to each sub-block feature map to the position encoding module; the position encoding module uses sine and cosine functions to generate the position code corresponding to each sub-block feature map, and outputs the embedding vector and position code corresponding to each sub-block feature map to the self-attention module; the self-attention module uses the multi-head attention mechanism to extract the global features corresponding to the multi-scale fusion feature map based on the embedding vector and position code corresponding to each sub-block feature map, and generates a multi-scale self-attention feature map. ; The channel attention module performs global average pooling and global maximum pooling on the multi-scale self-attention feature map based on the channel attention mechanism, and uses the multi-layer perceptron to generate the weight coefficients of each feature channel corresponding to the multi-scale self-attention feature map, and further uses each weight coefficient to recalibrate the multi-scale self-attention feature map to generate a weighted multi-scale self-attention feature map; The spatial attention module uses the spatial attention mechanism to perform average pooling and maximum pooling on the weighted multi-scale self-attention feature map, and uses splicing and convolution operations to generate a multi-scale comprehensive attention feature map; For the multi-scale comprehensive attention feature map, the MLP module uses nonlinear transformation and feature refinement to obtain the global semantic feature vector, that is, the global semantic feature vector corresponding to the input image.

[0045] Specifically, the multi-head attention mechanism, channel attention mechanism, and spatial attention mechanism are as follows:

[0046] MultiHead(Q,K,V)=Concat(head1,...,head h )W O ;

[0047] head i=Attention(QW i Q ,KW i K ,VW i V )

[0048] Among them, MultiHead(Q,K,V) is a multi-head attention mechanism, Q represents the embedding vector of each sub-block feature map generated by the embedding vector module, which is combined with the position encoding as the target feature of the image to be identified, K is used to measure the similarity between the features of the area to be identified and the features of all image areas, and V is the information required for feature fusion. i is the i-th attention head, W i Q ,W i K ,W i V is the weight of the i-th attention head, W O is the output weight matrix;

[0049] CA(X)=σ(W ca GlobalPool(X)

[0050] Among them, CA(X) is the channel attention mechanism, X is the multi-scale self-attention feature map processed by the multi-head self-attention mechanism; σ is the activation function, W ca is the weight of the channel attention mechanism, and GlobalPool(X) performs a global pooling operation on the multi-scale self-attention feature map;

[0051] SA(X)=σ(W ca ·Conv(avgpool(X)+maxpool(X)))

[0052] Where SA(X) is the spatial attention mechanism, X is the weighted multi-scale self-attention feature map after recalibration by the channel attention mechanism; σ is the activation function, avgpool(X) is the average pooling operation on the weighted multi-scale self-attention feature map, maxpool(X) is the maximum pooling operation on the weighted multi-scale self-attention feature map, and W sa is the weight of the spatial attention mechanism, and Conv is the convolution operation.

[0053] Furthermore, the output module processes the global semantic feature vector using a fully connected layer to generate a classification result of the risk source type, that is, the risk source type corresponding to the input image.

[0054] Furthermore, after training the risk source identification model, the method further includes: using TensorRT to convert the risk source identification model into ONNX format to accelerate inference and improve the performance and efficiency of the model; and using ONNX to migrate the trained risk source identification model to different deep learning frameworks and inference engines;

[0055] Specifically, refer to the following ONNX formula to build an ONNX intermediate model consisting of nodes and edges, where nodes represent operations in the neural network and edges indicate the direction of data flow. Furthermore, the dynamic dimension input of the image is canceled, and opset_version is set to 11. dummy_input is adjusted according to the optimal input of the model. Based on the data loading characteristics of the stdc network, batchsize is set to 1, and the trained PTH model category is unified with the converted category. The netron tool is used to visualize the converted network structure, with average pooling as AveragePool and activation functions as Relu and sigmoid, respectively:

[0056] ONNX(G)=(V,E)

[0057] Among them, G is the computational graph of the ONNX model, which contains all nodes and edges in the model; V represents the node set, and E represents the edge set.

[0058] Furthermore, based on different backbone networks, ONNX is used to convert the trained risk source identification model to different backbone networks. The ONNX intermediate model time consumption is analyzed, that is, the server inference time and post-processing time of different feature networks are analyzed, and the accuracy, i.e., the mAP value, is compared. In this embodiment, the backbone networks are selected as ResNet18, ResNet34, and ResNet50, and the comparative analysis results are as follows: Figure 6 shown.

[0059] from Figure 6 According to the available data, the NMS (non-maximum suppression) time remains basically consistent, which shows that the impact of the NMS stage on the inference time is relatively small, and the main differences are concentrated in the network structure and complexity. ResNet18 performs best in inference time, taking only 15.4 milliseconds to complete an inference, but the accuracy is relatively low, only 0.873. In comparison, ResNet50 has the longest inference time of 28.3 milliseconds, but the highest accuracy, reaching 0.936. ResNet34 is between ResNet18 and ResNet50, with both inference time and accuracy in the middle. The results show that the risk source identification model proposed in the present invention can be migrated to different deep learning frameworks according to the target requirements, and the risk source identification model of the present invention has good detection effects in different environments.

[0060] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural change made to the above embodiment based on the technical essence of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A multi-scale risk source identification method based on an improved attention mechanism, characterized in that: According to steps S1 and S2, a risk source identification model is constructed for identifying the risk source type corresponding to the fireworks image. According to step A, multi-scale identification of risk sources is implemented: Step S1: obtaining a preset number of fireworks image datasets that have been labeled with risk source type labels, and forming a sample set with a single fireworks image and its corresponding risk source type label as a sample; Step S2: Constructing a model to be trained that includes a feature extraction module, an enhanced feature extraction module, and an output module. Based on the sample set, the model to be identified is trained with the fireworks image in the sample as input and the risk source label corresponding to the fireworks image in the sample as output, thereby obtaining a risk source identification model for identifying the risk source type corresponding to the fireworks image. In step A, the fireworks image to be identified is input into the target backbone network containing the trained risk source identification model to obtain the risk source type corresponding to the fireworks image to be identified.

2. The multi-scale risk source identification method based on the improved attention mechanism according to claim 1 is characterized in that: The input end of the feature extraction module constitutes the input end of the risk source identification model, the output end of the feature extraction module is connected in series with the enhanced feature extraction module and the output module; the output end of the output module constitutes the output end of the risk source identification model; The feature extraction module is used to extract and output multi-scale fusion features of the input image; The enhanced feature extraction module is used to extract and output the semantic features of the input image based on the multi-scale fusion features; The output module is used to generate a classification result of the image risk source type based on the semantic features of the input image.

3. The multi-scale risk source identification method based on the improved attention mechanism according to claim 1 is characterized in that: The feature extraction module includes a cross-stage local network module, a feature pyramid module, and a path aggregation module; The input end of the cross-stage local network module constitutes the input end of the feature extraction module, the output end of the cross-stage local network module is connected in series with the feature pyramid module and the path aggregation module in sequence; the output end of the path aggregation module constitutes the output end of the feature extraction module; The cross-stage local network module receives the input image, uses the CSP1_X module and the CSP2_X module to generate basic features containing different levels, and outputs the basic features to the feature pyramid module; The feature pyramid module is a top-down feature pyramid structure with decreasing image scale. It is based on the high-level strong semantic basic features at the top layer and the low-level high-resolution basic features at the bottom layer. The feature pyramid module uses a top-down feature fusion operation to fuse the high-level strong semantic basic features layer by layer into the low-level high-resolution basic features through upsampling to generate a multi-scale feature map containing features at different levels; the path aggregation module is a bottom-up path aggregation structure. For the multi-scale feature map containing features at different levels, the low-level multi-scale feature map is fused layer by layer into the high-level multi-scale feature map through convolution and downsampling operations to generate a multi-scale fused feature map containing features at different levels, that is, the multi-scale fused feature map corresponding to the input image.

4. The multi-scale risk source identification method based on the improved attention mechanism according to claim 3 is characterized in that: The enhanced feature extraction module includes an embedding vector module, a position encoding module, a Transformer module, and an MLP module, and the Transformer module includes a self-attention module, a channel attention module, and a spatial attention module; The input end of the embedding vector module constitutes the input end of the enhanced feature extraction module, the output end of the embedding vector module is connected in series with the position encoding module, the self-attention module, the channel attention module, the spatial attention module, and the MLP module; the output end of the MLP module constitutes the output end of the enhanced feature extraction module; The embedding vector module receives the multi-scale fusion feature map and divides the multi-scale fusion feature map into sub-block feature maps based on the preset sub-block feature map size using patch cutting. The embedding vector corresponding to each sub-block feature map is further obtained by linear projection, and the embedding vector corresponding to each sub-block feature map is output to the position encoding module. The position encoding module uses the sine and cosine functions to generate the position code corresponding to each sub-block feature map, and outputs the embedding vector and position code corresponding to each sub-block feature map to the self-attention module. The self-attention module uses the multi-head attention mechanism to extract the global features corresponding to the multi-scale fusion feature map based on the embedding vector and position code corresponding to each sub-block feature map, and generates a multi-scale self-attention feature map; The channel attention module performs global average pooling and global maximum pooling on the multi-scale self-attention feature map based on the channel attention mechanism, and uses the multi-layer perceptron to generate the weight coefficients of each feature channel corresponding to the multi-scale self-attention feature map. The multi-scale self-attention feature map is further recalibrated using each weight coefficient to generate a weighted multi-scale self-attention feature map; the spatial attention module uses the spatial attention mechanism to perform average pooling and maximum pooling on the weighted multi-scale self-attention feature map, and uses splicing and convolution operations to generate a multi-scale comprehensive attention feature map; for the multi-scale comprehensive attention feature map, the MLP module uses nonlinear transformation and feature refinement to obtain the global semantic feature vector, that is, the global semantic feature vector corresponding to the input image.

5. The multi-scale risk source identification method based on the improved attention mechanism according to claim 4 is characterized in that: The output module processes the global semantic feature vector using a fully connected layer to generate a classification result of the risk source type, that is, the risk source type corresponding to the input image.

6. The multi-scale risk source identification method based on the improved attention mechanism according to claim 1 is characterized in that: After training the risk source identification model, it also includes: using TensorRT to convert the risk source identification model into ONNX format for reasoning acceleration.

Citation Information

Cited By

  • Image defect similarity search method with multi-scale feature extraction and attention enhancement

    CN122550990A