A fire zone edge accurate segmentation method and system based on multi-level attention mechanism

Through the precise fire zone edge segmentation method based on the multi-level attention mechanism, the problems of edge complexity, scale difference and temporal correlation in fire zone segmentation are solved, and higher precision and stable fire zone edge segmentation are achieved, thereby improving the reliability of fire monitoring and emergency response.

CN120355732BActive Publication Date: 2025-09-09浙江省气候中心(浙江省生态遥感中心浙江省农业气象中心)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510864136.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-09
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing fire zone segmentation methods based on radar reflectivity data have shortcomings in segmentation performance in processing target edge areas, target scale adaptability and temporal correlation between data, making it difficult to accurately process fire areas with complex edges and scale differences.

Method used

A multi-level attention mechanism is adopted to extract global and local features, combine feature fusion and credibility judgment strategies, and combine multi-scale feature fusion technology to solve the complex details and overall distribution patterns of the fire area edge and improve the segmentation accuracy and integrity.

Benefits of technology

It significantly improves the accuracy of fire zone edge segmentation and adaptability to complex scenarios, provides stable segmentation performance and anti-interference capabilities, and supports fire monitoring and emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355732B_ABST
    Figure CN120355732B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of fire zone image segmentation, and specifically to a method and system for accurately segmenting fire zone edges based on a multi-level attention mechanism. The method mainly includes a fire zone data acquisition step, a fire zone feature extraction step, a fire zone feature fusion step, and a conflict verification step. The multi-level attention mechanism is used to achieve collaborative extraction of global and local features, effectively capture the complex details and overall distribution patterns of the fire zone edge, and significantly improve the accuracy and integrity of fire zone edge segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fire zone image segmentation, and in particular to a method and system for accurately segmenting fire zone edges based on a multi-level attention mechanism. Background Art

[0002] Fire zone segmentation is a technology that monitors and analyzes target areas, accurately segmenting and predicting fire locations. By monitoring fire activity in real time, this technology helps relevant departments promptly identify fire hazards and implement emergency measures. Currently, fire zone segmentation has been widely applied in a variety of fields, including forest fire prevention and control, industrial fire supervision, urban residential fire monitoring, and wildfire monitoring, providing timely fire zone location and analysis. Therefore, fire zone segmentation technology has important research significance and can provide strong support for fire early warning, post-disaster rescue, and the optimization of fire prevention and control strategies. Fire zone segmentation is typically solved by analyzing radar reflectivity images. The heat and gases generated by a fire significantly change the temperature and density of the surrounding air, creating strong convection, which in turn creates specific signal patterns in the radar reflectivity image. Currently, researchers typically use optical flow and deep learning methods to process radar reflectivity images. Traditional optical flow methods can estimate dynamic changes in radar images, but they often suffer from low resolution and difficulty accurately processing small objects. In contrast, deep learning methods, with their powerful feature extraction and learning capabilities, can effectively process multi-scale radar echo targets, significantly improve the accuracy of recognition and segmentation, and have gradually become the focus of this research field.

[0003] However, existing fire zone segmentation methods based on radar reflectivity data still have some problems. First, the radar reflectivity map of the fire point area usually has significant edge complexity. Fire points are often concentrated in a specific area, resulting in the reflection intensity in this area being much higher than the surrounding area. The edge area of ​​the fire point has a weaker reflection intensity, showing an irregular edge complexity, which leads to poor performance in the target edge area segmentation. Second, the fire point area shows significant scale differences in the radar reflectivity map. From small-scale initial fire points to large-scale burning areas, the fire area scale varies greatly in different situations, resulting in insufficient adaptability to the target scale. Finally, the radar reflectivity map of the fire point usually shows obvious temporal changes. The spread speed of the fire, intensity changes, and airflow disturbances all have significant temporal correlations, but existing methods do not consider the temporal correlation between data.

[0004] Therefore, the present application proposes a method and system for detecting loose pipes in rail trains to solve the above-mentioned problems. Summary of the Invention

[0005] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a method and system for accurate fire zone edge segmentation based on a multi-level attention mechanism.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A fire zone edge accurate segmentation method based on a multi-level attention mechanism, including:

[0008] a fire area data acquisition step, obtaining a radar reflectivity map of the target area as an image to be inspected;

[0009] In the fire zone feature extraction step, the image to be inspected is processed and adjusted, and then classified according to feature type and introduced into a multi-level attention module to perform global feature extraction and local feature extraction respectively to obtain overall distribution features and local distribution features;

[0010] a fire zone feature fusion step, fusing the overall distribution feature and the local distribution feature, and entering a conflict verification step when there is a difference in feature extraction results of the overall distribution feature and the local distribution feature for the same pixel point in the image to be inspected;

[0011] In the conflict verification step, the overall distribution characteristics and local distribution characteristics of the pixel points with differences are used to determine whether the pixel point features are retained through the credibility judgment strategy, and the edge segmentation result is output based on the judgment results.

[0012] As a further improvement of the present invention, the fire area feature extraction step is provided with a multi-level attention module, which includes a global branch and a local branch. The image to be inspected is classified in the channel dimension according to the feature type, and is respectively imported into the global branch and the local branch to extract the overall distribution features and the local distribution features.

[0013] As a further improvement of the present invention, the global branch includes taking the image to be inspected as the input feature and performing preliminary feature transformation through a 1×1 convolution layer, and generating feature representations of Query, Key and Value respectively through three 3×3 depth convolutions, calculating the similarity between features through the matrix dot product of Query and Key, and normalizing the similarity to generate attention weights, and then weighting the attention weights and the Value features and re-integrating the input features to obtain the output F of the attention mechanism. att , the F att After adjusting the feature channel through 1×1 convolution and adding it to the original input feature, a 3×3 convolution is used to enhance the feature, and finally the overall distribution feature is extracted.

[0014] As a further improvement of the present invention, the global branch configuration includes:

[0015] ;

[0016] ;

[0017] Among them, F Att represents the feature representation of Query, Key, and Value under the attention mechanism. a is a learnable scaling parameter used to control the amplitude of the product of the Key and Query matrices. Attention(Q, K, V) represents the attention mechanism construction function, where Q, K, and V represent Query, Key, and Value, respectively. Represents the normalization function for calculating feature similarity between Query and Key, F Global Represents the overall distribution characteristics, F Gin Represents the initial features of the input convolutional layer of the image to be detected.

[0018] As a further improvement of the present invention, the local branch includes taking the image to be inspected as input features to adjust the channel dimension through a 1×1 convolution layer, and performing channel shuffling to rearrange the channel order to enhance the initial feature expression ability of the image to be inspected, and further extracting features with significant irregularities and prominent fire point areas through a 3×3 convolution to obtain local features, and further deeply mining the local features through a 3×3×3 convolution to obtain local distribution features.

[0019] As a further improvement of the present invention, the local feature extraction of the local branch includes dividing the image to be inspected into several local image blocks, generating a non-feature area mask for each local image block through 1×1 convolution and activation function, the non-feature area mask is calculated by setting the fire area feature threshold and the gradient threshold obtained by training the historical fire area edge data, performing 3×3 hole convolution on each local image block to extract spatial context features, generating a local feature map, multiplying the local feature map and the corresponding non-feature area mask element by element, deleting the non-fire area edge area with low reflectivity and no gradient change to obtain effective local features, and performing 3×3×3 depth-separable convolution on the deleted feature map to obtain local distribution features.

[0020] As a further improvement of the present invention, the credibility judgment strategy includes: when the pixel point with differences has obvious fire zone feature expression in the overall distribution feature, and has not been removed in the local distribution feature due to the ineffective feature deletion operation, and at the same time, the local area around the pixel point presents a continuous change consistent with the fire zone edge feature in the local distribution feature, and the feature distribution of the overall distribution feature in the local area is relatively smooth and lacks such continuous change, it is determined that the extraction result of the local distribution feature is more accurate, and the feature of the pixel point in the local distribution feature is retained for feature fusion; when the pixel point is determined to be a non-fire zone feature in the local distribution feature due to the ineffective feature deletion operation, but has obvious fire zone feature expression in the overall distribution feature, and in the global area of ​​the image to be inspected where the pixel point is located, the overall distribution feature presents an overall distribution trend consistent with the fire zone edge feature, and the feature distribution of the local distribution feature in the global area is relatively scattered and lacks an overall trend, it is determined that the extraction result of the overall distribution feature is more accurate, and the feature of the pixel point in the overall distribution feature is retained for feature fusion.

[0021] As a further improvement of the present invention, the credibility judgment strategy also includes, when the feature expressions of the pixel point in both the overall distribution feature and the local distribution feature are relatively vague and it is difficult to directly determine whether it is a fire zone feature, constructing a dual-modal credibility quantization formula to calculate the credibility, the dual-modal credibility quantization formula is configured as follows:

[0022] ;

[0023] Among them, C represents the feature credibility, G(p) is the global branch reflectivity value of the pixel point, μ G Represents the global fire zone characteristic mean, obtained through historical data, σ G represents the global gradient variance, indicating the mutation of the overall distribution, L(p) is the reflectivity value retained by the local branch after non-feature deletion, M(p) represents the mask value generated by the local branch, which is used to reflect whether the point is judged as a valid feature, μ L It represents the average gradient of the effective features in the local block. α and β are the balance coefficients determined by cross-validation based on the stability of the global branch and the detail sensitivity of the local branch. When C is greater than the preset confidence value, the pixel is determined to be the edge of the fire area.

[0024] As a further improvement of the present invention, it also includes a multi-scale feature fusion step. After acquiring the data, the image to be inspected is obtained by fusing spatial information at different scales to obtain a multi-scale feature map, and the different spatial level features of the fire area are captured according to the multi-scale feature map. The image to be inspected after the multi-scale feature fusion is input into the multi-level attention module for feature extraction.

[0025] A fire zone edge accurate segmentation system based on a multi-level attention mechanism, including:

[0026] The fire area data acquisition module obtains the radar reflectivity map of the target area as the image to be inspected;

[0027] The fire zone feature extraction module processes and adjusts the image to be inspected, classifies it according to feature type, and imports it into the multi-level attention module to perform global feature extraction and local feature extraction to obtain overall distribution features and local distribution features;

[0028] A fire zone feature fusion module is used to fuse the overall distribution feature and the local distribution feature. When there is a difference between the feature extraction results of the overall distribution feature and the local distribution feature for the same pixel point in the image to be inspected, a conflict verification module is entered.

[0029] The conflict verification module uses the credibility judgment strategy to determine whether the overall distribution characteristics and local distribution characteristics of the pixel points with differences should be retained, and outputs the edge segmentation result based on the judgment results.

[0030] The beneficial effects of the present invention are: collaborative extraction of global and local features is achieved through a multi-level attention mechanism, the complex details and overall distribution patterns of the fire zone edge are effectively captured, and the accuracy and completeness of the fire zone edge segmentation are significantly improved. The conflict verification strategy combined with the dual-modal credibility quantification formula can intelligently distinguish the authenticity of feature conflict areas, ensuring stable segmentation performance in scenarios with feature ambiguity or noise interference. Multi-scale feature fusion technology enhances adaptability to fire zones of different sizes, while the non-feature area masking mechanism of local branches specifically suppresses low-reflectivity background interference, further optimizing the expression of edge details. Compared with traditional methods, the present invention has made breakthrough progress in the precise positioning of fire zone edges, adaptability to complex scenes, and anti-interference capabilities, providing reliable technical support for fire monitoring and emergency response. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flow chart of the fire zone edge accurate segmentation method based on the multi-level attention mechanism of the present invention;

[0032] Figure 2 It is a system block diagram of the multi-level attention module of the present invention. DETAILED DESCRIPTION

[0033] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom," "top," "inner," and "outer" refer to directions toward or away from the geometric center of a particular component, respectively.

[0034] At present, fire zone segmentation technology has important applications in many fields such as forest fire prevention and control, industrial area fire supervision, urban residential area fire monitoring and wildfire monitoring. However, the existing fire zone segmentation method based on radar reflectivity data has shortcomings in processing the segmentation performance of target edge areas, target scale adaptability and temporal correlation between data. In order to overcome these limitations, this application proposes a fire zone edge accurate segmentation method and system based on a multi-level attention mechanism.

[0035] like Figures 1 to 2 As shown, including:

[0036] Fire area data acquisition steps: Obtain the radar reflectivity map of the target area as the image to be inspected to provide basic data for subsequent feature extraction.

[0037] Fire zone feature extraction: After data processing and adjustment, the image to be inspected is classified according to feature type and fed into a multi-level attention module for global feature extraction and local feature extraction, respectively, to obtain overall and local distribution features. Global feature extraction captures the overall distribution of the fire zone, while local feature extraction focuses on the detailed characteristics of the fire zone.

[0038] Fire zone feature fusion: The global and local distribution features are fused. If the feature extraction results for the same pixel in the image under inspection differ between the global and local distribution features, the conflict verification step is initiated. Feature fusion can comprehensively utilize global and local features to improve segmentation accuracy.

[0039] Conflict Verification: The overall and local distribution features of the discrepant pixels are evaluated using a credibility judgment strategy to determine whether the pixel features should be retained. The resulting edge segmentation results are then output based on the credibility judgment strategy. This credibility judgment strategy resolves conflicts in feature extraction results and ensures the accuracy of the segmentation results.

[0040] Through the above steps, the technical problem of accurate segmentation of the fire zone edge is solved, making the segmentation result more accurate and reliable.

[0041] Existing fire zone segmentation methods often have difficulty capturing detailed features when processing target edge areas, resulting in low segmentation accuracy. In addition, the scales of fire zones vary greatly, from small-scale initial fire points to large-scale burning areas, and existing methods perform poorly when adapting to targets of different scales. At the same time, the spread speed, intensity changes, and airflow disturbances of fires all have significant time correlations. The failure to effectively utilize the temporal correlation information between data is also a major shortcoming of existing methods. Therefore, this application proposes a method for accurate segmentation of fire zone edges based on a multi-level attention mechanism. This method uses a multi-level attention module to simultaneously extract global features and local features to improve the segmentation performance of target edge areas; uses a feature fusion step to comprehensively utilize global and local features to improve segmentation accuracy; and resolves conflicts in feature extraction results through a conflict verification step to ensure the accuracy of segmentation results.

[0042] When introducing the main content, naturally introduce and explain key terms. For example, the multi-level attention mechanism is a technique that can simultaneously focus on global and local features, feature fusion refers to the combined utilization of features from different sources, and conflict verification refers to the judgment and processing of discrepant features. Furthermore, when describing the implementation environment, you can mention using radar reflectivity images as input data, extracting and fusion features through the multi-level attention module, and ultimately outputting accurate fire zone edge segmentation results.

[0043] First, the fire zone data acquisition step is the foundation of the entire method. By obtaining a radar reflectivity map of the target area, it provides the necessary data support for subsequent feature extraction. Next, the fire zone feature extraction step is crucial. A multi-level attention module is used to extract global and local features. Global feature extraction captures the overall distribution of the fire zone, while local feature extraction focuses on the detailed characteristics of the fire zone. The fire zone feature fusion step then fuses the global and local distribution features. When the feature extraction results for the same pixel differ, the conflict verification step proceeds. Finally, the conflict verification step uses a credibility judgment strategy to determine whether the pixel feature should be retained. Combined with the judgment results, the edge segmentation result is output.

[0044] Compared with existing technologies, the method in this application has significant advantages in processing target edge regions, adapting to targets of different scales, and leveraging temporal correlations between data. Through a multi-level attention mechanism, it can simultaneously focus on global and local features, improving segmentation performance for target edge regions. Through the feature fusion step, it can comprehensively utilize global and local features to improve segmentation accuracy. Through the conflict verification step, it can resolve conflicts in feature extraction results and ensure the accuracy of segmentation results.

[0045] Specifically, the fire area data acquisition step first obtains the radar reflectivity map of the target area to provide basic data for subsequent feature extraction. In the fire area feature extraction step, after data processing and adjustment, the image to be inspected is classified according to the feature type and imported into the multi-level attention module to perform global feature extraction and local feature extraction respectively to obtain overall distribution features and local distribution features. Global feature extraction captures the overall distribution of the fire area, while local feature extraction focuses on the detailed features of the fire area. Then, in the fire area feature fusion step, the overall distribution features and local distribution features are fused. When there is a difference between the feature extraction results of the two for the same pixel point, the conflict verification step is entered. In the conflict verification step, the credibility judgment strategy is used to determine whether the pixel point features are retained, and the edge segmentation result is output in combination with the judgment result. Through the above steps, the technical problem of accurate segmentation of the fire area edge is solved, making the segmentation result more accurate and reliable.

[0046] Furthermore, if Figures 1 to 2 As shown, the present application also proposes that the fire area feature extraction step is provided with a multi-level attention module, which includes a global branch and a local branch. The image to be inspected is classified in the channel dimension according to the feature type, and the global branch and the local branch are respectively imported to extract the overall distribution features and the local distribution features.

[0047] A multi-level attention module is introduced into the fire zone feature extraction step. This module consists of a global branch and a local branch. The global branch is used to extract the overall distribution features of the image to be inspected, while the local branch is used to extract the local distribution features of the image to be inspected. By classifying the features in the channel dimension, it is ensured that the features are imported into the global and local branches respectively for feature extraction. The collaborative work of the global and local branches can effectively extract global and local features for precise segmentation of fire zone edges, solving the problem of how to effectively extract global and local features in precise segmentation of fire zone edges.

[0048] The multi-level attention module can be implemented as follows: First, the image to be inspected is classified according to its feature type in the channel dimension and fed into the global and local branches. The global branch takes the image to be inspected as input, performs preliminary feature transformation via a 1×1 convolutional layer, and then generates query, key, and value feature representations through three 3×3 depthwise convolutions. Feature similarity is calculated using the dot product of the query and key matrices, and the similarities are normalized to generate attention weights. The attention weights are weighted and summed with the value features before the input features are re-integrated to obtain the output of the attention mechanism. The global features are adjusted in feature channels via a 1×1 convolution, added to the original input features, and enhanced via a 3×3 convolution to ultimately extract the overall distribution features. The local branch takes the image to be inspected as input, adjusts the channel dimension via a 1×1 convolutional layer, and performs channel shuffling to rearrange the channel order, enhancing the initial feature representation of the image to be inspected. Features with significant irregularities and prominent hotspots are further extracted via a 3×3 convolution to obtain local features. The local features are further mined through a 3×3×3 convolution to finally obtain the local distribution features.

[0049] By introducing a multi-level attention module, this application can effectively solve the problem of global and local feature extraction in the precise segmentation of fire zone edges. Compared with existing technologies, this method can more comprehensively extract the feature information of the image to be inspected through the collaborative work of global and local branches, significantly improving the accuracy and precision of fire zone edge segmentation. Especially when dealing with complex edge areas, this method can better capture the overall distribution and local detail features of the fire zone, thereby achieving a more accurate segmentation effect.

[0050] Furthermore, if Figures 1 to 2 As shown, the present application also proposes to use the image to be inspected as the input feature and perform preliminary feature transformation through a 1×1 convolution layer, and generate feature representations of Query, Key and Value respectively through three 3×3 deep convolutions, calculate the similarity between features through the matrix dot product of Query and Key, and normalize the similarity to generate attention weights, add the attention weights and the Value features together and then reintegrate the input features to obtain the output Fatt of the attention mechanism, adjust the feature channel of Fatt through 1×1 convolution and add it to the original input features, and then perform feature enhancement through a 3×3 convolution to finally extract the overall distribution features.

[0051] The technical solution involves performing a preliminary feature transformation on the image to be inspected through a 1×1 convolutional layer to generate Query, Key, and Value feature representations. The similarity between features is calculated through the matrix dot product of the Query and Key, and the similarity is normalized to generate attention weights. The attention weights are weighted and summed with the Value features before reintegrating the input features to obtain the output Fatt of the attention mechanism. The feature channels are then adjusted through a 1×1 convolution and added to the original input features. Feature enhancement is then performed through a 3×3 convolution to ultimately extract the overall distribution features. This solution effectively extracts the overall distribution features of the image through the attention mechanism, solving the problem of accurate segmentation of the fire zone edge.

[0052] Specifically, the image to be inspected first undergoes preliminary feature transformation through a 1×1 convolution layer to reduce computational complexity and integrate preliminary channel information. Next, three 3×3 depthwise convolutions are used to generate Query, Key, and Value feature representations, which are used to construct the attention mechanism. The similarity between features is then calculated through the matrix dot product of Query and Key, and the similarity is normalized using the Softmax function to generate the attention weight. Subsequently, the attention weight is weighted and summed with the Value feature, and the input features are reintegrated to obtain the output Fatt of the attention mechanism. Next, Fatt adjusts the feature channel through a 1×1 convolution and adds it to the original input features retained by the jump connection. This not only retains the basic information of the input features, but also improves the gradient fluidity. Finally, a 3×3 convolution is used to further optimize and enhance the feature expression to obtain the overall distribution feature.

[0053] As a preferred embodiment, the similarity calculation between features can adopt different matrix dot product methods, such as inner product, outer product, or other methods suitable for feature similarity calculation. The normalization process of attention weight can use different normalization functions, such as Softmax, Sigmoid, etc., to adapt to different application scenarios. The 3×3 convolution in the feature enhancement process can further optimize the size and number of convolution kernels to improve the effect of feature extraction.

[0054] This approach effectively extracts the overall distribution features of the image through the attention mechanism, achieving precise segmentation of the fire zone edge. Compared with existing technologies, this solution not only more accurately extracts the features of the fire zone edge, but also effectively handles complex edge conditions in the image, improving the accuracy and stability of segmentation.

[0055] Furthermore, if Figures 1 to 2As shown, the present application also proposes that the global branch configuration is: FAtt represents the feature representation of Query, Key and Value under the attention mechanism, a is a learnable scaling parameter used to control the amplitude of the product of the Key and Query matrices, Attention(Q, K, V) represents the attention mechanism construction function, where Q, K, V represent Query, Key and Value respectively, representing the normalization function for calculating the feature similarity of Query and Key, FGlobal represents the overall distribution feature, and FGin represents the initial feature of the input convolution layer of the image to be tested.

[0056] The global branch uses the attention mechanism to extract the overall distribution characteristics of the fire zone edge by configuring technical features such as FAtt, a, and Attention (Q, K, V). FAtt, as the output of the attention mechanism, contains the feature representations of Query, Key, and Value. The amplitude of the product of the Key and Query matrices is controlled by the learnable scaling parameter a, thereby enhancing the flexibility of feature extraction. The Attention (Q, K, V) construction function calculates the similarity between Query and Key, performs normalization processing, generates attention weights, and then weighted sums them with the Value feature to obtain the global feature FGlobal. FGin represents the initial features of the convolutional layer of the input image to be inspected. The above technical features work together to effectively extract the overall distribution characteristics of the fire zone edge.

[0057] FAtt represents the feature representation of Query, Key, and Value under the attention mechanism. Specifically, Query, Key, and Value are obtained by performing different linear transformations on the input features. The learnable scaling parameter a is used to control the amplitude of the product of the Key and Query matrices. This allows the weights of features to be automatically adjusted during training to avoid excessive amplification or reduction of features. Attention (Q, K, V) is the core of the attention mechanism. It generates attention weights by calculating the similarity between Query and Key and normalizing them using the Softmax function. The attention weights are then weighted and summed with the Value features to obtain the final output of the attention mechanism, namely the global feature FGlobal. FGin represents the initial features of the input convolution layer of the image to be inspected. These initial features extract low-level features of the image through convolution operations, providing a basis for subsequent feature extraction.

[0058] As a preferred implementation, FAtt can be implemented using a multi-head attention mechanism. This divides the input features into multiple subspaces, independently calculates attention within each subspace, and then concatenates the results, further improving the accuracy and robustness of feature extraction. a can be implemented using learnable scalars or vectors to more flexibly control the relationships between different features. Attention (Q, K, V) can be implemented using a multi-layer perceptron (MLP) to enhance the model's expressiveness.

[0059] This application implements technical features such as FAtt, a, and Attention (Q, K, V), and utilizes an attention mechanism to extract the overall distribution characteristics of fire zone edges, achieving accurate extraction of fire zone edge features. Compared to existing technologies, this application enhances the flexibility and robustness of feature extraction and improves the accuracy of fire zone edge segmentation by introducing learnable scaling parameters and a multi-head attention mechanism. As a result, this application more effectively addresses the inaccurate fire zone edge feature extraction issue in existing technologies, providing a more precise and reliable fire zone edge segmentation method.

[0060] Furthermore, if Figures 1 to 2 As shown, the present application also proposes to use the image to be inspected as the input feature to adjust the channel dimension through a 1×1 convolution layer, and perform channel shuffling to rearrange the channel order to enhance the initial feature expression ability of the image to be inspected, and further extract the features with significant irregularities and prominent fire point areas through a 3×3 convolution to obtain local features, and further deeply mine the local features through a 3×3×3 convolution to obtain local distribution features.

[0061] The function of these technical features is to first adjust and shuffle the image channels through a multi-step convolution operation to enhance the expressive power of the initial features, then extract the features of significant irregularities and prominent fire spots through 3×3 convolution, and finally further explore the local distribution features through 3×3×3 convolution. Through the above technical solution, this application solves the problem of feature extraction of significant irregularities and prominent fire spots in the image to be inspected, ensuring that the local distribution features of the fire area can be accurately identified and extracted, and improving the performance of accurate segmentation of the fire area edge.

[0062] In the technical solution of the present application, the image to be inspected is first subjected to channel dimension adjustment through a 1×1 convolution layer. This step is mainly to reduce computational complexity and integrate preliminary channel information. Next, a channel shuffling operation is performed to promote information interaction between channels by rearranging the channel order, thereby enhancing feature expression capabilities. Then, a 3×3 convolution is used to further extract detailed features of the local area, especially for areas with significant local irregularities and prominent fire points in reflection intensity. This step can better capture their local feature patterns. Finally, a 3×3×3 convolution is used to deeply mine local features, integrate spatial dimensions and contextual information, and obtain local distribution features.

[0063] As a preferred implementation, the 1×1 convolutional layer can use standard convolution operations with a kernel size of 1 and a stride of 1. Channel shuffling can be achieved by randomly disrupting the channel order or using a preset shuffling pattern. The 3×3 convolutional layer and the 3×3×3 convolutional layer can use standard convolution operations with kernel sizes of 3 and 3×3, respectively, and a stride of 1.

[0064] This application uses multi-step convolution operations to effectively extract significant irregularities and prominent fire spot features from the image under inspection. In particular, the channel shuffling operation enhances the expressiveness of the initial features. Compared to existing technologies, this application's solution can more accurately identify and extract the local distribution characteristics of the fire zone, improving the performance of precise segmentation of the fire zone edge.

[0065] Furthermore, if Figures 1 to 2 As shown, the present application also proposes to divide the image to be inspected into several local image blocks, process each local image block one by one, and generate a non-feature area mask. The non-feature area mask is calculated by the fire area feature threshold and the gradient threshold, and is used to identify the non-feature area. Then, each local image block extracts the spatial context features through a 3×3 hole convolution to generate a local feature map. The local feature map and the non-feature area mask are multiplied element by element, and the non-fire area edge areas with low reflectivity and no gradient change are deleted to obtain effective local features. Finally, a 3×3×3 depth-separable convolution is performed on the deleted feature map to further extract local distribution features.

[0066] The local branch extracts local features by dividing the image to be inspected into several local image blocks, processing each local image block one by one to generate a non-feature region mask. The non-feature region mask is calculated using the fire region feature threshold and gradient threshold and is used to identify non-feature regions. Each local image block then undergoes a 3×3 dilated convolution to extract spatial context features and generate a local feature map. The local feature map and the non-feature region mask are element-wise multiplied to remove non-fire region edge regions with low reflectivity and no gradient change, thereby obtaining effective local features. Finally, a 3×3×3 depthwise separable convolution is performed on the deleted feature map to further extract local distribution features.

[0067] Furthermore, if Figures 1 to 2 As shown in the figure, the generation of non-feature region masks can be achieved by inputting each local image patch into a 1×1 convolutional layer, and generating a preliminary non-feature region mask through an activation function. The final non-feature region mask is then calculated using the set fire region feature threshold and the gradient threshold obtained through training with historical fire region edge data. This ensures the accuracy and effectiveness of the non-feature region mask.

[0068] In addition, the use of 3×3 dilated convolution can be achieved by inputting each local image block into a 3×3 dilated convolution layer to extract spatial context features and generate a local feature map. The use of dilated convolution can expand the receptive field without increasing computational complexity, thereby better capturing spatial context information.

[0069] Finally, the use of 3×3×3 depthwise separable convolution can be achieved by feeding the deleted feature map into a 3×3×3 depthwise separable convolution layer to further extract local distribution features. The use of depthwise separable convolution can improve the efficiency and effectiveness of feature extraction while maintaining low computational complexity.

[0070] Through the above-mentioned technical features, this application can effectively address the problem of poor performance in fire zone edge segmentation. Specifically, by generating a non-feature area mask and deleting non-fire zone edge areas with low reflectivity and no gradient change, the accuracy and effectiveness of feature extraction can be improved, thereby improving the performance of fire zone edge segmentation. Compared with existing technologies, this application has higher accuracy and robustness when processing fire zone edge segmentation.

[0071] Furthermore, if Figures 1 to 2As shown, the present application also proposes that the credibility judgment strategy includes: when the pixel point with differences has obvious fire zone feature expression in the overall distribution feature, and has not been removed in the local distribution feature due to the ineffective feature deletion operation, and the local area around the pixel point presents a continuous change in the local distribution feature consistent with the fire zone edge feature, and the feature distribution of the overall distribution feature in the local area is relatively smooth and lacks such continuous change, it is determined that the extraction result of the local distribution feature is more accurate, and the feature of the pixel point in the local distribution feature is retained for feature fusion; when the pixel point is determined to be a non-fire zone feature in the local distribution feature due to the ineffective feature deletion operation, but has obvious fire zone feature expression in the overall distribution feature, and in the global area of ​​the image to be inspected where the pixel point is located, the overall distribution feature presents an overall distribution trend consistent with the fire zone edge feature, and the feature distribution of the local distribution feature in the global area is relatively scattered and lacks an overall trend, it is determined that the extraction result of the overall distribution feature is more accurate, and the feature of the pixel point in the overall distribution feature is retained for feature fusion.

[0072] The technical features of the credibility judgment strategy include determining the representation of fire zone features in both the overall and local distribution features of a pixel, as well as the distribution of features in the local and global regions surrounding the pixel. Its purpose is to compare the extraction results of the overall and local distribution features to determine which feature extraction result is more accurate, thereby retaining the more accurate feature for feature fusion. This approach effectively addresses the issue of discrepancies in feature extraction results during fire zone edge segmentation and improves accuracy.

[0073] When implementing the credibility judgment strategy, a variety of methods can be used. Specifically, the fire zone feature expression of the pixel point in the overall distribution feature and the local distribution feature can be judged by setting a threshold. For example, when the feature value of a certain pixel point is higher than the preset fire zone feature threshold, the pixel point is considered to have obvious fire zone feature expression. Further, if Figures 1 to 2 As shown, the accuracy of feature extraction results can be determined by analyzing the distribution of features in the local and global regions surrounding a pixel. For example, by calculating the continuous change in pixel feature values ​​within a local region and the overall trend of feature distribution within the global region, we can determine whether the extraction results of local or global distribution features are more accurate. Furthermore, feature extraction results can be further optimized and verified using machine learning algorithms in conjunction with historical data.

[0074] The credibility judgment strategy proposed in this application comprehensively considers the fire zone feature expression of the pixel point in the overall distribution features and local distribution features, as well as the feature distribution of the local area and the global area around the pixel point when there are differences in the feature extraction results, thereby determining a more accurate feature extraction result for feature fusion. In this way, the problem of differences in feature extraction results in fire zone edge segmentation can be effectively solved, and the accuracy of fire zone edge segmentation can be improved. Compared with the existing technology, this application can more accurately retain the fire zone features by introducing a credibility judgment strategy, thereby improving the effect and reliability of fire zone edge segmentation.

[0075] Furthermore, if Figures 1 to 2 As shown, the present application also proposes that when the feature expressions of the pixel point in both the overall distribution features and the local distribution features are relatively vague, and it is difficult to directly determine whether it is a fire zone feature, a dual-modal credibility quantification formula is constructed to calculate the credibility. The dual-modal credibility quantification formula is configured as follows: wherein C represents the feature credibility, G(p) is the global branch reflectivity value of the pixel point, μG represents the global fire zone feature mean, which is obtained through historical data, σG represents the global gradient variance, which represents the mutation of the overall distribution, L(p) is the reflectivity value retained by the local branch after non-feature deletion, M(p) represents the mask value generated by the local branch, which is used to reflect whether the point is judged to be a valid feature, μL represents the average gradient of the valid feature in the local block, and α and β are balance coefficients determined by cross-validation based on the stability of the global branch and the detail sensitivity of the local branch. When C is greater than the preset credibility value, the pixel point is judged to be the edge of the fire zone.

[0076] This application constructs a dual-modal credibility quantification formula to calculate the credibility of the pixel to determine whether it is a fire zone feature. The dual-modal credibility quantification formula comprehensively considers parameters such as the global branch reflectivity value, the global fire zone feature mean, the global gradient variance, the reflectivity value of the local branch, the mask value generated by the local branch, the average gradient of the effective features in the local block, and is adjusted through a balance coefficient. When the calculated credibility value is greater than the preset value, the pixel is determined to be the edge of the fire zone. Through this technical solution, the dual-modal credibility quantification formula can be used to accurately determine whether the pixel is a fire zone feature when both the overall distribution characteristics and the local distribution characteristics are relatively fuzzy, thereby improving the accuracy and reliability of the fire zone edge segmentation.

[0077] Specifically, when the feature expression of a pixel is relatively vague, the various parameters in the dual-modal credibility quantification formula can be obtained in the following way: the global fire zone feature mean μG and the global gradient variance σG can be obtained through statistical analysis of historical data, and the mask value M(p) generated by the local branch is generated by the non-feature deletion operation of the local branch. The global branch reflectivity value G(p) and the local branch reflectivity value L(p) can be extracted respectively by the global branch and the local branch in the multi-level attention mechanism. The balance coefficients α and β are determined by cross-validation to ensure that the stability of the global branch and the detail sensitivity of the local branch can be effectively balanced under different circumstances. Therefore, when the calculated credibility value C is greater than the preset credibility value, the pixel is determined to be the edge of the fire zone.

[0078] Therefore, the application of the dual-modal credibility quantification formula can accurately determine whether a pixel point is a fire zone feature even when the feature expression is ambiguous, avoiding misjudgment caused by feature ambiguity and improving the accuracy and reliability of fire zone edge segmentation. Compared with the existing technology, this application effectively solves the problem of fire zone edge segmentation in the case of ambiguous feature expression by comprehensively considering global and local features and introducing a credibility quantification formula for precise judgment, thereby improving the segmentation effect.

[0079] Furthermore, if Figures 1 to 2 As shown, the present application also proposes a multi-scale feature fusion step. After acquiring the data, the image to be inspected is obtained by fusing spatial information at different scales to obtain a multi-scale feature map, and the different spatial level features of the fire area are captured according to the multi-scale feature map, and the image to be inspected after the multi-scale feature fusion is input into the multi-level attention module for feature extraction.

[0080] The multiscale feature fusion step fuses spatial information at different scales to obtain a multiscale feature map. This step captures the different spatial hierarchical features of the fire zone. The image to be inspected, resulting from the multiscale feature fusion, is then fed into the multi-level attention module for feature extraction, ensuring accurate identification and segmentation of the fire zone edges at different scales. This multiscale feature fusion step solves the problem of capturing features at different scales in fire zone edge segmentation. By fusing spatial information at different scales, the model can more comprehensively capture the various characteristics of the fire zone, thereby improving the accuracy and precision of edge segmentation.

[0081] Specifically, the multi-scale feature fusion step can be implemented in the following ways: one way is to use a multi-scale convolutional layer to extract feature maps of different scales, and then fuse these feature maps. Another way is to use a pyramid pooling operation to pool feature maps of different scales and then fuse them. In addition, the expressive power of feature maps of different scales can be enhanced through a multi-scale attention mechanism. As a preferred embodiment, the multi-scale feature map can be input into a multi-level attention module for further processing to ensure that the edge features of the fire area can be accurately captured at different scales.

[0082] Therefore, through the multi-scale feature fusion step, this application can achieve higher accuracy and precision in the fire zone edge segmentation task. Compared with existing technologies, the method of this application can more comprehensively capture the various characteristics of the fire zone, especially the significantly enhanced feature expression ability at different scales, thereby improving the edge segmentation effect. This technical solution not only solves the problem of capturing features at different scales in fire zone edge segmentation, but also improves the performance and adaptability of the overall model.

[0083] Furthermore, if Figures 1 to 2 As shown, the present application also proposes a fire zone edge precision segmentation system based on a multi-level attention mechanism, including: a fire zone data acquisition module, which obtains a radar reflectivity map of the target area as an image to be inspected; a fire zone feature extraction module, which processes and adjusts the image to be inspected, classifies it according to the feature type, and imports it into a multi-level attention module to perform global feature extraction and local feature extraction respectively to obtain overall distribution features and local distribution features; a fire zone feature fusion module, which fuses the overall distribution features and the local distribution features, and enters a conflict verification module when there is a difference in the feature extraction results of the overall distribution features and the local distribution features for the same pixel point in the image to be inspected; the conflict verification module uses a credibility judgment strategy to judge whether the pixel point features are retained based on the overall distribution features and the local distribution features of the pixel point with differences, and outputs the edge segmentation result based on the judgment result.

[0084] The fire zone data acquisition module obtains the radar reflectivity map of the target area as the image to be inspected, solving the problem of data source. The fire zone feature extraction module performs classification after data processing and adjustment, and imports the multi-level attention module to perform global feature extraction and local feature extraction respectively, solving the problem of feature extraction accuracy. The fire zone feature fusion module fuses the overall and local features, and enters the conflict verification module when there are differences, solving the problem of feature fusion accuracy. The conflict verification module determines whether the feature should be retained through a credibility judgment strategy, and outputs the edge segmentation result, solving the problem of the accuracy of the final edge segmentation. This technical solution ensures the accurate segmentation of the fire zone edge through a multi-level attention mechanism and feature fusion strategy, combined with a conflict verification module.

[0085] The fire area data acquisition module can obtain the radar reflectivity map of the target area through the radar equipment and convert it into the image to be inspected. The fire area feature extraction module includes a data processing unit and a multi-level attention module. The data processing unit is used to pre-process and classify the features of the image to be inspected. The multi-level attention module includes a global branch and a local branch. The global branch is used to extract the overall distribution features, and the local branch is used to extract the local distribution features. The fire area feature fusion module is used to fuse the overall distribution features and the local distribution features, and when there is a difference in the feature extraction results, the features of the difference pixels are passed to the conflict verification module. The conflict verification module makes a decision on whether to retain or discard the pixel features through a credibility judgment strategy, and outputs the final edge segmentation result.

[0086] The system in this application uses a multi-level attention mechanism and feature fusion strategy, combined with a conflict verification module, to ensure accurate segmentation of fire zone edges. Compared with existing technologies, the system in this application significantly improves the accuracy of feature extraction and feature fusion, can more effectively handle the complexity of fire zone edges, and provide higher-precision segmentation results.

[0087] The above shows and describes the basic features, principles, and advantages of the present invention. It should be noted that the present invention is not limited to the above embodiments, which are only some embodiments. Without departing from the spirit and scope of the present invention, various improvements and supplements made are considered to be within the scope of protection of the present invention.

Claims

1. A fire zone edge accurate segmentation method based on a multi-level attention mechanism, characterized by: include: a fire area data acquisition step, obtaining a radar reflectivity map of the target area as an image to be inspected; In the fire zone feature extraction step, after data processing and adjustment of the image to be inspected, the image is classified according to feature type and introduced into a multi-level attention module for global feature extraction and local feature extraction to obtain overall distribution features and local distribution features; the multi-level attention module includes a global branch and a local branch; a fire zone feature fusion step, fusing the overall distribution feature and the local distribution feature, and entering a conflict verification step when there is a difference in feature extraction results of the overall distribution feature and the local distribution feature for the same pixel point in the image to be inspected; In the conflict verification step, the overall distribution characteristics and local distribution characteristics of the pixel points with differences are used to determine whether the pixel point features are retained through the credibility judgment strategy, and the edge segmentation result is output based on the judgment results; The global branch includes taking the image to be inspected as the input feature and performing preliminary feature transformation through a 1×1 convolution layer, and generating feature representations of Query, Key, and Value respectively through three 3×3 depth convolutions, calculating the similarity between features through the matrix dot product of Query and Key, and normalizing the similarity to generate attention weights, and then re-integrating the input features after weighted summation of the attention weights and the Value features to obtain the output F of the attention mechanism. att , the F att After adjusting the feature channel through 1×1 convolution and adding it to the original input feature, a 3×3 convolution is performed to enhance the feature, and finally the overall distribution feature is extracted; The local branch includes: using the image to be inspected as an input feature to adjust the channel dimension through a 1×1 convolution layer, and performing channel shuffling to rearrange the channel order to enhance the initial feature expression capability of the image to be inspected; further extracting features with significant irregularities and prominent fire point areas through a 3×3 convolution to obtain local features; and further performing deep mining on the local features through a 3×3×3 convolution to obtain local distribution features; The credibility judgment strategy includes: when the pixel point with differences has obvious fire zone feature expression in the overall distribution feature and has not been removed in the local distribution feature due to the ineffective feature deletion operation, and at the same time, the local area around the pixel point shows a continuous change in the local distribution feature that is consistent with the fire zone edge feature, and the overall distribution feature in the local area has a relatively smooth feature distribution and lacks such continuous change, the extraction result of the local distribution feature is determined to be more accurate, and the feature of the pixel point in the local distribution feature is retained for feature fusion; When the pixel point is judged as a non-fire area feature in the local distribution feature due to the ineffective feature deletion operation, but has obvious fire area feature expression in the overall distribution feature, and in the global area of ​​the image to be inspected where the pixel point is located, the overall distribution feature shows an overall distribution trend consistent with the edge feature of the fire area, and the feature distribution of the local distribution feature in the global area is relatively scattered and lacks an overall trend, it is judged that the extraction result of the overall distribution feature is more accurate, and the feature of the pixel point in the overall distribution feature is retained for feature fusion.

2. The method for accurately segmenting fire zone edges based on a multi-level attention mechanism according to claim 1 is characterized in that: The fire zone feature extraction step is provided with a multi-level attention module, which classifies the image to be inspected in the channel dimension according to the feature type, and imports the global branch and the local branch to extract the overall distribution features and the local distribution features respectively.

3. The method for accurately segmenting fire zone edges based on a multi-level attention mechanism according to claim 1 is characterized in that: The global branch configuration includes: ; ; Among them, F Att represents the feature representation of Query, Key, and Value under the attention mechanism. a is a learnable scaling parameter used to control the amplitude of the product of the Key and Query matrices. Attention(Q, K, V) represents the attention mechanism construction function, where Q, K, and V represent Query, Key, and Value, respectively. Represents the normalization function for calculating feature similarity between Query and Key, F Global Represents the overall distribution characteristics, F Gin Represents the initial features of the input convolutional layer of the image to be detected.

4. The method for accurately segmenting fire zone edges based on a multi-level attention mechanism according to claim 1 is characterized in that: The local feature extraction of the local branch includes dividing the image to be inspected into several local image blocks, generating a non-feature area mask for each local image block through 1×1 convolution and activation function, wherein the non-feature area mask is calculated by setting the fire area feature threshold and the gradient threshold obtained by training the historical fire area edge data, performing 3×3 hole convolution on each local image block to extract spatial context features, generating a local feature map, multiplying the local feature map and the corresponding non-feature area mask element by element, deleting the non-fire area edge area with low reflectivity and no gradient change to obtain effective local features, and performing 3×3×3 depth-separable convolution on the deleted feature map to obtain local distribution features.

5. The method for accurately segmenting fire zone edges based on a multi-level attention mechanism according to claim 1 is characterized in that: The credibility judgment strategy also includes constructing a dual-modal credibility quantization formula to calculate credibility when the feature expressions of the pixel point in both the overall distribution feature and the local distribution feature are relatively vague, making it difficult to directly determine whether it is a fire zone feature. The dual-modal credibility quantization formula is configured as follows: ; Among them, C represents the feature credibility, G(p) is the global branch reflectivity value of the pixel point, μ G Represents the global fire zone characteristic mean, obtained through historical data, σ G represents the global gradient variance, indicating the mutation of the overall distribution, L(p) is the reflectivity value retained by the local branch after non-feature deletion, M(p) represents the mask value generated by the local branch, which is used to reflect whether the point is judged as a valid feature, μ L It represents the average gradient of the effective features in the local block. α and β are the balance coefficients determined by cross-validation based on the stability of the global branch and the detail sensitivity of the local branch. When C is greater than the preset confidence value, the pixel is determined to be the edge of the fire area.

6. The method for accurately segmenting fire zone edges based on a multi-level attention mechanism according to claim 1 is characterized in that: It also includes a multi-scale feature fusion step. After acquiring the data, the image to be inspected is obtained by fusing spatial information at different scales to obtain a multi-scale feature map, and the different spatial level features of the fire area are captured according to the multi-scale feature map. The image to be inspected after the multi-scale feature fusion is input into the multi-level attention module for feature extraction.

7. A fire zone edge accurate segmentation system based on a multi-level attention mechanism, characterized by: include: The fire area data acquisition module obtains the radar reflectivity map of the target area as the image to be inspected; The fire zone feature extraction module processes and adjusts the image to be inspected, classifies it according to feature type, and imports it into a multi-level attention module to extract global features and local features respectively to obtain overall distribution features and local distribution features; the multi-level attention module includes a global branch and a local branch; A fire zone feature fusion module is used to fuse the overall distribution feature and the local distribution feature. When there is a difference between the feature extraction results of the overall distribution feature and the local distribution feature for the same pixel point in the image to be inspected, a conflict verification module is entered. The conflict verification module uses the credibility judgment strategy to determine whether the overall distribution characteristics and local distribution characteristics of the pixel points with differences are retained, and outputs the edge segmentation result based on the judgment result; the global branch includes taking the image to be inspected as the input feature through a 1×1 convolution layer for preliminary feature transformation, and generating the feature representations of Query, Key and Value respectively through three 3×3 deep convolutions, calculating the similarity between the features through the matrix dot product of Query and Key, and normalizing the similarity to generate the attention weight, and then re-integrating the input features after weighted summation of the attention weight and the Value feature to obtain the output F of the attention mechanism att , the F att After adjusting the feature channel through 1×1 convolution and adding it to the original input feature, a 3×3 convolution is performed to enhance the feature, and finally the overall distribution feature is extracted; The local branch includes: using the image to be inspected as an input feature to adjust the channel dimension through a 1×1 convolution layer, and performing channel shuffling to rearrange the channel order to enhance the initial feature expression capability of the image to be inspected; further extracting features with significant irregularities and prominent fire point areas through a 3×3 convolution to obtain local features; and further performing deep mining on the local features through a 3×3×3 convolution to obtain local distribution features; The credibility judgment strategy includes: when the pixel point with differences has obvious fire zone feature expression in the overall distribution feature and has not been removed in the local distribution feature due to the ineffective feature deletion operation, and at the same time, the local area around the pixel point shows a continuous change in the local distribution feature that is consistent with the fire zone edge feature, and the overall distribution feature in the local area has a relatively smooth feature distribution and lacks such continuous change, the extraction result of the local distribution feature is determined to be more accurate, and the feature of the pixel point in the local distribution feature is retained for feature fusion; When the pixel point is judged as a non-fire area feature in the local distribution feature due to the ineffective feature deletion operation, but has obvious fire area feature expression in the overall distribution feature, and in the global area of ​​the image to be inspected where the pixel point is located, the overall distribution feature shows an overall distribution trend consistent with the edge feature of the fire area, and the feature distribution of the local distribution feature in the global area is relatively scattered and lacks an overall trend, it is judged that the extraction result of the overall distribution feature is more accurate, and the feature of the pixel point in the overall distribution feature is retained for feature fusion.

Citation Information

Patent Citations

  • High-precision flame positioning method and system based on bilingual sense attention mechanism

    CN113393521A

  • Automobile part image super-resolution method based on dual-channel residual attention

    CN118570065A