Coal-fired power plant boiler burner flame combustion area segmentation method and device and medium
By introducing attention mechanism and multi-level fusion detection head in the YOLOv8 model, combined with the RepNCSPELAN4 module, the problem of insufficient deep feature processing in flame segmentation is solved, and higher flame segmentation accuracy and model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510274062.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
AI Technical Summary
The existing YOLOv8 model lacks deep feature processing in flame segmentation tasks, resulting in insufficient accuracy when identifying flames with complex morphology and irregular boundaries.
By introducing attention mechanism (CBAM) and multi-level fusion detection head (MAH) into the YOLOv8 model, combined with the RepNCSPELAN4 module, the model's processing ability of deep features is enhanced, and the effective fusion of deep and shallow features is achieved.
It improves the pixel-level segmentation accuracy of the flame, can more accurately identify complex flame patterns and boundaries, and enhances the model's feature extraction and generalization capabilities.
Smart Images

Figure CN120107595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and in particular to a flame combustion area segmentation method for a coal-fired power station boiler burner based on YOLOv8-MAH. Background Art
[0002] In the context of promoting green and low-carbon development, coal-fired power plants need to improve their peak load-shaving capabilities to support the consumption of new energy and ensure stable operation of the power grid. The flame state of a single burner is crucial to the overall combustion efficiency and safety of the furnace. Unstable burner flame combustion can lead to local overheating and uneven temperature distribution, which will affect the overall combustion effect of the furnace and increase pollutant emissions. Therefore, burner flame detection is a key means to ensure safe, environmentally friendly and efficient operation of power plants.
[0003] In recent years, with the rapid development of computer vision, semantic segmentation algorithms based on deep learning have been widely used, and a series of networks have been used in flame segmentation. However, the actual flame detection and segmentation have high requirements for the real-time performance of the model, but the deep learning model is large in size and has high computational complexity, especially the introduction of the enhancement module, which often leads to performance degradation. The lightweight model provides an efficient solution for flame segmentation under complex working conditions by balancing accuracy and efficiency. The YOLO series of models achieves efficient inference speed and high detection accuracy through a single-stage detection architecture. Its core is to simplify the target detection task into a direct regression problem, avoiding the candidate region generation process in the multi-stage detector, thereby greatly improving the detection speed, which is very suitable for real-time tasks.
[0004] However, at present, this series of models mainly detect flames by predicting the position and category of the bounding box, and pay less attention to the precise segmentation of flames. In particular, when performing segmentation tasks, YOLOv8 only performs segmentation based on shallow feature maps, lacking detailed processing of deep features of flames. Although such shallow feature maps can quickly identify the approximate area of the flame, they are still insufficient in identifying flames with complex shapes and irregular boundaries. In practical applications, the shapes and boundaries of flames are complex and changeable, and it is difficult to meet the needs of precise segmentation by relying solely on shallow features. This has a serious impact on the subsequent extraction of flame features. Summary of the invention
[0005] Technical problem: In view of the shortcomings of the above-mentioned existing methods, the present invention provides a method, equipment and medium for segmenting the flame combustion area of a coal-fired power plant boiler burner that meets the image segmentation accuracy, which can be used for real-time detection and segmentation of the flame in the furnace under multiple load conditions.
[0006] Technical solution:
[0007] The present invention first provides a method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model, comprising the following steps:
[0008] Step 1: Obtain flame combustion images of power plant boilers under multiple working conditions through the built in-furnace image acquisition system;
[0009] Step 2: Label each burning area of a single-frame flame image to form a data set;
[0010] Step 3: Build a YOLOv8-MAH flame detection model; the YOLOv8-MAH flame detection model includes a backbone network, a neck network, and a head network; an attention mechanism module is added to the backbone network to extract channel features and spatial features in the input feature map;
[0011] Step 4: Use the data set formed in step 2 to train the YOLOv8-MAH flame detection model constructed in step 3;
[0012] Step 5: Use the YOLOv8-MAH flame detection model trained in step 4 to realize flame detection and segmentation under unprecedented working conditions.
[0013] The steps for training the YOLOv8 model in step 4 are:
[0014] Step 4.1: Add the CBAM attention mechanism to the backbone network of YOLOV8. CBAM enhances the model's ability to focus on key features through two independent attention mechanisms (channel attention and spatial attention):
[0015] The channel attention module adaptively adjusts the weight of each channel according to the importance of different feature channels, specifically:
[0016] Mc(F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))
[0017] The spatial attention module assigns weights to the spatial positions of the feature maps, specifically:
[0018] Ms(F')=σ(f 7×7 ([AvgPool(F');MaxPool(F')]))
[0019] Among them: Mc(F) is the channel attention module, Ms(F') is the spatial attention module; F is the input feature map; F' is the feature map output by the channel attention module; MLP(﹒) is a multi-layer perceptron used to generate the weight of each channel; AvgPool(﹒) is the global average pooling; MaxPool(﹒) is the global maximum pooling; σ(﹒) is the Sigmoid activation function; f 7×7 Represents a 7x7 convolution operation, which is used to generate weights of the spatial dimension.
[0020] Step 4.2: Replace the C2f module in the neck network with the RepNCSPELAN4 module, where:
[0021] RepConvN extracts detailed information through multiple convolution branches (1x1, 3x3 convolution and residual branches);
[0022] The CSP module can effectively combine shallow features with deep features through cross-stage feature fusion. This cross-stage feature reuse can help the model better handle multi-scale features in flame detection.
[0023] The ELAN4 module can effectively integrate convolutional features at different levels through feature aggregation design.
[0024] Step 4.3: Design of MAH detection head based on interaction between deep and shallow features:
[0025] The MAH detection head performs convolution and average pooling on the P4 (40×40) and P5 (20×20) layers;
[0026] The compressed P5 and P4 features are used as convolution kernels to traverse the shallow feature map and fuse and enhance the features pixel by pixel. This allows it to retain spatial detail information while also being able to perceive deeper semantic information.
[0027] Through the above pixel-level convolution operation, P5-P3 mask and P4-P3 mask are generated respectively. These two masks represent the convolution interaction results of P3 from the P5 and P4 feature maps;
[0028] The learnable parameters λ1 and λ2 perform weighted fusion of the P5-P3 mask and the P4-P3 mask, respectively. These weights are learned through the network and are used to dynamically adjust the influence of each mask;
[0029] The fused feature map is further combined with the P3 feature map.
[0030] The present invention also provides an electronic device, comprising:
[0031] one or more processors;
[0032] A memory for storing one or more programs;
[0033] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-mentioned coal-fired power plant boiler burner flame combustion area segmentation method based on the YOLOv8-MAH model.
[0034] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model.
[0035] Advantages of the invention:
[0036] 1. The multi-level fusion detection head and attention mechanism are coupled into the YOLOv8 network to achieve effective fusion of deep and shallow features and improve the pixel-level segmentation accuracy of the flame. Adding the attention mechanism to the backbone layer network of YOLOv8 can improve the feature extraction ability of the model, focus more on the key features of the flame in the early stage of training, "pay attention" to the morphology and texture features of the flame, rather than limited to the brightness features of the pixels, and further distinguish the interference between the flame and the bright background and fly ash in the boiler. The multi-level fusion detection head performs convolution and average pooling on the P4 feature map (40*40) and the P5 feature map (20*20), uses the compressed P5 and P4 features as the convolution kernel, traverses the P3 feature map of the shallow features, and fuses and enhances the features pixel by pixel. It can retain spatial detail information while also being able to perceive deeper semantic information. Through the above pixel-level convolution operation, the P5-P3 mask and P4-P3 mask are generated respectively. These two masks represent the convolution interaction results of the P5 and P4 feature maps on the P3. The learnable parameters λ1 and λ2 perform weighted fusion on the P5-P3 mask and the P4-P3 mask respectively. These weights are learned through the network and are used to dynamically adjust the influence of each mask. Finally, the fused feature map is further combined with the P3 feature map. After the interaction of deep and shallow features, the final feature map will have both the spatial information of shallow features and the semantic information of deep features.
[0037] 2. By introducing the RepNCSPELAN4 module into the neck network, reparameterization and cross-stage feature reuse are achieved, avoiding parameter redundancy in the network. The RepNCSPELAN4 module combines technologies such as RepConv, NCSP and ELAN4, and improves the feature extraction capability of the target detection network through reparameterization, multi-level feature fusion and cross-stage feature reuse. RepConv can capture different scales, shapes and edge features in flame detection through multi-branch convolution design. The shape and brightness of the flame vary, and multiple convolution branches (1x1, 3x3 convolution and residual branches) can better extract detail information. The NCSP module can effectively combine shallow features with deep features through cross-stage feature fusion. This cross-stage feature reuse can help the model better handle multi-scale features in flame detection. The ELAN4 module can effectively integrate convolution features at different levels through feature aggregation design. The constructed image acquisition system was applied under various load conditions of the unit to verify the segmentation accuracy, inference speed and generalization ability of the model. It not only solves the problem of excessive background brightness and the influence of fly ash on flame detection in the furnace, but also realizes the segmentation of the flame area and coal powder area at the burner outlet, and solves the interference of adjacent burner flames, avoiding the occurrence of false detection. The flame features extracted from the segmented combustion area can further reflect the flame combustion situation in the furnace. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of the network structure of the present invention.
[0039] Figure 2 This is a schematic diagram of the RepNCSPELAN4 module of the present invention.
[0040] Figure 3 It is a schematic diagram of the RepNCSP module of the present invention.
[0041] Figure 4 It is a schematic diagram of the RepNBottelneck module of the present invention.
[0042] Figure 5 Schematic diagram of the multi-level segmentation head MAH.
[0043] Figure 6 Schematic diagram of the data set structure of this embodiment.
[0044] Figure 7 Schematic diagram of YOLOv8-MAH model performance evaluation.
[0045] Figure 8 This is a comparison chart of YOLOv8-MAH flame segmentation effects under multiple load conditions. DETAILED DESCRIPTION
[0046] The present invention is described in detail below in conjunction with the accompanying drawings and attached tables:
[0047] Example 1
[0048] This embodiment provides a method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on YOLOv8-MAH, including the following steps:
[0049] Step 1: Flame image acquisition: obtain flame combustion images of power plant boilers under multiple working conditions through the built in-furnace image acquisition system;
[0050] Step 2: Label each burning area of a single-frame flame image to form a data set. The data set structure is as follows: Figure 3 As shown in the figure, to avoid randomness in the combustion process, images are taken every 5 minutes, and the sampling time is 1 minute. Finally, 3,000 images are collected under each combustion condition, totaling 9,000 images. 1,500 images are randomly selected under each condition, totaling 4,500 images, and the flame area is divided and labeled by manual annotation. Among them, 60% are randomly selected as training sets, 20% as validation sets, and 20% as test sets.
[0051] Step 3: Construct a flame detection model. The model design structure is as follows Figure 1 As shown;
[0052] The model consists of three parts:
[0053] The first part is the backbone network (Backbone) with the attention mechanism (CBAM) added. After adjusting the input image to (640×640×3), the Conv convolution module extracts the features of the input image; the CBAM extracts the spatial and channel features of the high-resolution feature map; the C2f module transfers the features in stages to reduce redundant feature transfers; the feature map is then input into the SPPF (Spatial Pyramid Pooling Fast) module, which captures multi-scale features through pooling windows of different sizes, enhancing the model's ability to detect objects of different sizes. Finally, the (20×20×512) feature map is input into the CBAM to achieve the feature extraction capability of the low-resolution feature map.
[0054] The second part is the Neck network. In the Neck layer, the low-resolution feature map is converted to the same size as the high-resolution feature map by Upsampling. Then, multiple feature maps of different scales are merged by concatenation. In the Neck network, the original c2f module is replaced with the RepNCSPELAN4 module, which can enhance the representation ability of the feature map and enable the model to capture the global information and long-range dependencies in the image.
[0055] RepNCSPELAN4 module (with Figure 2 As shown in the figure, a 1×1 convolution layer (Conv) is first input, and then the feature map is split (split); the segmented feature map is input into the RepNCSP module and then into the convolution module. After this step, the output is divided into two parts, one of which is input into the Concat module, and the other part continues to input a RepNCSP module and the convolution module, and then inputs the Concat module. Finally, all feature maps are spliced.
[0056] The RepNCSP module in the above structure (attached Figure 3 As shown in the figure, through cross-stage feature fusion, shallow features can be effectively combined with deep features. Specifically, the input feature map is convolved through two different paths, one of which continues to input N RepNBottleneck modules after convolution. Then the two intermediate features are spliced.
[0057] In the above structure, RepNBottleneck consists of RepConvN and a convolution module (see Appendix Figure 4 As shown in the figure, RepConvN extracts detail information through multiple convolution branches (1x1, 3x3 convolution and residual branch). The first convolution Conv1 uses a 3×3 convolution kernel with a stride of 1, and the second convolution Conv2 uses a 1×1 convolution with a stride of 1 to compress the input channels. The output of the convolution is added to the input and nonlinearly activated by the Silu activation function.
[0058] The third part is the head network, which is the last module in the target detection model and is mainly responsible for generating the final detection result based on the feature map provided by the Neck layer. The head network in the present invention is a multi-level fusion detection head (Multi-Augment Head). By performing convolution and average pooling on the P4 (40×40) and P5 (20×20) feature maps, the compressed P5 and P4 features are obtained, and used as convolution kernels for pixel-by-pixel feature fusion and enhancement with the shallow feature map. This operation retains spatial detail information and perceives deep semantic information. Through pixel-level convolution operations, P5-P3 mask and P4-P3 mask are generated respectively, indicating the convolution interaction results of P5 and P4 on P3. The learnable parameters λ1 and λ2 perform weighted fusion on the two masks and dynamically adjust their influence. Finally, the fused feature map is combined with the P3 feature map to form a feature map containing shallow spatial information and deep semantic information.
[0059] The above feature maps are subjected to bounding box regression, category prediction, and confidence prediction. An output with target location, category, confidence, and segmentation information is generated. This output provides an accurate bounding box for each flame target and also generates a detailed segmentation area of the flame, achieving the dual functions of target detection and segmentation.
[0060] Step 4: Use the data to train, test and verify the image-based flame detection model.
[0061] Step 5: Use the trained YOLOv8-MAH model to achieve flame detection and segmentation under unprecedented conditions.
[0062] The steps to train the YOLOv8 model in step 4 are:
[0063] Step 4.1: The size of the annotated flame image input into the backbone network is 640×640×3. After passing through two convolution modules and one C2f module, it is compressed into a feature map of 160×160×128 and further input into the CBAM module;
[0064] Step 4.2: CBAM uses two independent attention mechanisms (channel attention and spatial attention) to enhance the model's ability to focus on key features. This module ensures that the size of the input and output feature maps remains unchanged and only performs key feature focus learning;
[0065] The channel attention module adaptively adjusts the weight of each channel according to the importance of different feature channels, specifically:
[0066] Mc(F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))
[0067] The spatial attention module assigns weights to the spatial positions of the feature maps, specifically:
[0068] Ms(F')=σ(f 7×7 ([AvgPool(F');MaxPool(F')]))
[0069] Where: Mc(F) is the channel attention module, Ms(F”) is the spatial attention module; F is the input feature map; F′ is the feature map output by the channel attention module; MLP(﹒) is a multi-layer perceptron used to generate the weight of each channel; AvgPool(﹒) is the global average pooling; MaxPool(﹒) is the global maximum pooling; σ is the Sigmoid activation function; f 7×7 Represents a 7x7 convolution operation, which is used to generate weights of the spatial dimension.
[0070] Step 4.3: After the backbone network ends through the SPPF module, the feature map size is 20×20×512, which is input into the CBAM module for the second feature focusing, and then into the neck network;
[0071] Step 4.4: Replace the C2f module in the neck network with the RepNCSPELAN4 module, where:
[0072] RepConvN extracts detailed information through multiple convolution branches (1x1, 3x3 convolution and residual branch). The first convolution Conv1 uses a 3×3 convolution kernel with a stride of 1, and the second convolution Conv2 uses a 1×1 convolution with a stride of 1 to compress the input channels. The output of the convolution is added to the input, and nonlinear activation is performed through the Silu activation function;
[0073] The CSP module can effectively combine shallow features with deep features through cross-stage feature fusion. Specifically, the input feature map is convolved through two different paths to generate two intermediate features. The two intermediate features are then concatenated through a 1×1 convolution. Finally, the concatenated features are processed through multiple bottleneck layers.
[0074] Output=cv3(concat(m(cv2(x)),cv2(x)))
[0075] Among them, m represents the combination of multiple bottleneck modules, concat represents concatenation, and cv3 is the last convolutional layer.
[0076] Step 4.5: Design of MAH detection head based on interaction of deep and shallow features, as shown in the attached figure. Figure 2 As shown:
[0077] The MAH detection head performs convolution and average pooling on the P4 (40*40) and P5 (20*20) layers;
[0078] The compressed P5 and P4 features are used as convolution kernels to traverse the shallow feature map and fuse and enhance the features pixel by pixel. This allows it to retain spatial detail information while also being able to perceive deeper semantic information.
[0079] Through the above pixel-level convolution operation, P5-P3mask and P4-P3mask are generated respectively. These two masks represent the convolution interaction results of P3 from the P5 and P4 feature maps;
[0080] The learnable parameters λ1 and λ2 perform weighted fusion of the P5-P3 mask and the P4-P3 mask, respectively. These weights are learned through the network and are used to dynamically adjust the influence of each mask;
[0081] The fused feature map is further combined with the P3 feature map.
[0082] In order to further verify the technical effect of the present invention, the following experiment is designed to further illustrate the
[0083] Experimental environment: The program running environment is Python 3.12, the processor is Intel i9-10940X CPU, the RAM is 128GB, the GPU is NVIDIA GeForce RTX3090, and the CUDA version is 11.6.
[0084] The training set data comes from the flame images at the burner outlet of the boiler at 430MW, 440MW and 600MW loads.
[0085] To avoid randomness during the combustion process, images were taken every 5 minutes, with a sampling time of 1 minute. Finally, 3,000 images were collected under each combustion condition, totaling 9,000 images. 1,500 images were randomly selected under each condition, totaling 4,500 images, and the flame area was divided and labeled by manual annotation. 60% of them were randomly selected as training sets, 20% as validation sets, and 20% as test sets.
[0086] The application scenario of the model generalization capability is the flame image at the burner outlet under boiler loads of 270MW, 330MW, and 580MW.
[0087] The model performance is evaluated by using the recall rate Recall as shown in formula (4), the precision rate Precision as shown in formula (5), and the mean intersection over union (mIoU).
[0088] Recall rate: This indicator evaluates the proportion of all positive samples that the model can detect
[0089]
[0090] Precision: This metric evaluates the proportion of true positive samples in the prediction model.
[0091]
[0092] Intersection-over-Union ratio: measures the degree of overlap between the predicted area and the true area
[0093]
[0094] Average intersection-over-union ratio: This indicator evaluates the average detection effect of the model on all categories
[0095]
[0096] Among them, TP represents the number of positive samples correctly detected by the model, FN represents the number of positive samples that the model failed to detect, FP represents the number of samples incorrectly predicted as positive, and IoU i is the IoU of the i-th category.
[0097] The total number of iterations during model training is 300 epochs. The training results are attached. Figure 4 shown.
[0098] In order to verify the impact of CBAM attention mechanism, RepNCSPELAN4, and MAH multi-level segmentation head on the performance of YOLOv8 network, two sets of ablation experiments were designed.
[0099] Ablation experiment 1: The performance comparison of different levels of MAH segmentation head fusion is shown in Table 1. The results show that the mIoU of the P3 shallow feature is 73.9, and the inference time is 4.12ms. After fusing the two deep features P4 and P5, the mIoU increases. The P4 P5-P3 segmentation head obtained by fusing the two deep features with the P3 shallow feature increases the mIoU to 79.6, and the inference time only increases by 0.2ms.
[0100] Table 1. Experimental results of multi-level fusion ablation
[0101]
[0102] Ablation experiment 2: The performance comparison of different modules and networks added to the YOLOv8 model is shown in Table 2. The results show that with the addition of the attention mechanism CBAM and the replacement of the C2f module of the Neck part by RepNCSPELAN4, the MIOU of the model has increased significantly. After the optimization of the above modules, the mIoU of the P4 P5-P3 feature fusion layer increased by 1.6 year-on-year. Finally, after the introduction of learnable parameters λ1 and λ2, the MIOU of the YOLOV8-MAH model increased from 73.8 to 82.6, while the inference time only increased by 0.31ms.
[0103] Table 2 Ablation experiment results of different modules
[0104]
[0105]
[0106] SOTA experimental comparison: The flame segmentation performance of the YOLOv8-MAH model and other neural networks is shown in Table 3. PSPNet performed best in the mIoU index, reaching 84.7, followed by UperNet and SegFormer, 83.81 and 82.96 respectively. In contrast, the mIoU of YOLOv8 is only 73.8, the lowest among all models. After the introduction of multi-scale feature fusion and module optimization, the mIoU of the YOLOv8-MAH model increased to 82.6. In terms of inference speed, the inference time of PSPNet and UperNet is 16.19ms and 15.75ms respectively, and the inference speed is significantly slower. The inference speed of SegFormer is 8.09ms, which is faster than PSPNet and UperNet, but still nearly twice as slow as YOLOv8-MAH. Therefore, YOLOv8-MAH has a clear advantage in the comprehensive performance of mIoU and inference speed.
[0107] Table 3 Comparison of segmentation performance of different models
[0108]
[0109] Comparison of segmentation results before and after model improvement
[0110] As attached Figure 8As shown in the figure: under the 580MW working condition, the flame burns violently, and the furnace is filled with a large amount of dust particles. On the one hand, the combustion background in the furnace is too bright, and on the other hand, the image collected by the imaging system is blurred, and the effective segmentation of the flame effective combustion area cannot be achieved. YOLOv8 relies on the pixel information of the image and lacks semantic learning of the entire image, which leads to the excessive division of the flame texture details in the segmentation model, resulting in the loss of the flame combustion area. YOLOV8-MAH combines deep flame image features to segment the flame edge texture more accurately, and achieves the segmentation of the flame effective combustion area under high load.
[0111] Under the 330MW condition, the flame and pulverized coal airflow are intertwined. When the flame area or the pulverized coal area is small, YOLOv8 cannot accurately segment it. YOLOv8-MAH can not only accurately identify the flame combustion area and pulverized coal area of different sizes, but also avoid the overlap of the two areas during the division process.
[0112] Under the 270MW condition, the ignition zone at the burner outlet moves forward, resulting in an increase in the flame black dragon area. YOLOv8-MAH more accurately segments the coal powder area and the flame at the front end of the coal powder area.
[0113] Example 2
[0114] This embodiment provides an electronic device, including:
[0115] one or more processors;
[0116] A memory for storing one or more programs;
[0117] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the coal-fired power plant boiler burner flame combustion area segmentation method based on the YOLOv8-MAH model provided in Example 1.
[0118] Example 3
[0119] This embodiment provides a storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model provided in Embodiment 1 are implemented.
Claims
1. A method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model, characterized in that: The steps include: Step 1: Obtain flame combustion images of power plant boilers under multiple working conditions through the built in-furnace image acquisition system; Step 2: Label each burning area of a single-frame flame image to form a data set; Step 3: Build a YOLOv8-MAH flame detection model; the YOLOv8-MAH flame detection model includes a backbone network, a neck network, and a head network; an attention mechanism module is added to the backbone network to extract channel features and spatial features in the input feature map; Step 4: Use the data set formed in step 2 to train the YOLOv8-MAH flame detection model constructed in step 3; Step 5: Use the YOLOv8-MAH flame detection model trained in step 4 to realize flame detection and segmentation under unprecedented working conditions.
2. The method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model according to claim 1 is characterized in that: The backbone network includes Conv convolution module, CBAM module, C2f module and SPPF module; the input image is extracted through the Conv convolution module to form a flame image feature map; the spatial features and channel features of the flame image feature map are extracted through the CBAM module; the spatial features and channel features extracted by the CBAM module are then transmitted in stages through the C2f module; the transmitted features are then input into the SPPF module, and the SPPF module captures multi-scale features through pooling windows of different sizes to form a low-resolution feature map, thereby enhancing the model's detection capability for targets of different sizes; finally, the low-resolution feature map obtained by the SPPF module is input into the CBAM module to extract the spatial features and channel features of the low-resolution feature map, thereby realizing the feature extraction capability of the low-resolution feature map.
3. The method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model according to claim 2 is characterized in that: The CBAM module includes a channel attention module and a spatial attention module. The channel attention module adaptively adjusts the weight of each channel according to the importance of different feature channels. Specifically: Mc(F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) The spatial attention module assigns weights to the spatial positions of the feature maps, specifically: Ms(F')=σ(f 7×7 ([AvgPool(F');MaxPool(F')])) Among them: Mc(F) is the channel attention module, Ms(F') is the spatial attention module; F is the input feature map; F' is the feature map output by the channel attention module; MLP(﹒) is a multi-layer perceptron used to generate the weight of each channel; AvgPool(﹒) is the global average pooling; MaxPool(﹒) is the global maximum pooling; σ is the Sigmoid activation function; f 7×7 Represents a 7x7 convolution operation, which is used to generate weights of the spatial dimension.
4. The method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model according to claim 2 is characterized in that: In the neck network, the low-resolution feature map output by the backbone network is converted to the same size as the high-resolution feature map through upsampling; then, multiple feature maps of different scales are merged through a splicing operation; and the representation capability of the feature map is enhanced through the RepNCSPELAN4 module, so that the model can capture the global information and long-range dependencies in the image, and output P3 feature maps, P4 feature maps, and P5 feature maps of different resolutions.
5. The method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model according to claim 4 is characterized in that: The C2f module in the neck network is changed to the RepNCSPELAN4 module, where the RepNCSPELAN4 module first inputs a 1×1 convolution layer and then performs feature map segmentation; the segmented feature map is input into the RepNCSP module and then into the convolution module. After this step, the output is divided into two parts, one part of which is input into the Concat module, and the other part continues to input a RepNCSP module and a convolution module, and then into the Concat module; finally, all feature maps are spliced; In the above structure, the RepNCSP module can effectively combine shallow features with deep features through cross-stage feature fusion. Specifically, the input feature map is convolved through two different paths, one of which continues to input N RepNBottleneck modules after convolution, and then the two intermediate features are spliced; In the above structure, the RepNBottleneck module consists of RepConvN and a convolution module, where RepConvN extracts detail information through multiple convolution branches; the first layer of convolution Conv1 uses a 3×3 convolution kernel with a stride of 1; the second layer of convolution Conv2 uses a 1×1 convolution with a stride of 1; the input is channel compressed; the output of the convolution is added to the input, and nonlinear activation is performed through the Silu activation function.
6. The method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model according to claim 5, characterized in that: The head network is a multi-level fusion detection head; the multi-level fusion detection head obtains compressed P4 features and P5 features by convolution and average pooling on the P4 feature map and the P5 feature map, and uses the compressed P4 features and P5 features as convolution kernels to perform pixel-by-pixel feature fusion and enhancement with the P3 feature map; through convolution operations, P5-P3 mask and P4-P3 mask are generated respectively, and P5-P3 mask and P4-P3 mask represent the convolution interaction results of P5 features and P4 features on P3 features respectively; P5-P3 mask and P4-P3 mask are adjusted by using learnable parameters λ1 and λ2. The fused feature map is weighted fused with the P3 feature map; the fused feature map is combined with the P3 feature map to form a feature map containing shallow spatial information and deep semantic information; finally, bounding box regression, category prediction, and confidence prediction are performed on the feature map containing shallow spatial information and deep semantic information to generate an output with target location, category, confidence, and segmentation information. This output provides an accurate bounding box for each flame target and also generates a detailed segmentation area of the flame, realizing the dual functions of target detection and segmentation.
7. The method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model according to claim 6, characterized in that: The training method of the YOLOv8-MAH flame detection model is: Step 4.1: The annotated flame image is input into the backbone network, and after passing through two convolution modules and a C2f module, it is compressed into a feature map and further input into the CBAM module; Step 4.2: The CBAM module enhances the model's ability to focus on key features through channel attention and spatial attention. The CBAM module ensures that the size of the input and output feature maps remains unchanged and only performs key feature focus learning; Step 4.3: After the backbone network end undergoes the SPPF module, it is input into the CBAM module for a second feature focusing to obtain a low-resolution feature map, which is then input into the neck network; Step 4.4: After two upsampling and concatenation operations, input into the RepNCSPELAN4 module to obtain P3 (feature map; The P3 feature map is further processed through convolution and concatenation operations, and then input into the RepNCSPELAN4 module to obtain the P4 feature map; The P4 feature map is concatenated with the low-resolution feature map output by the backbone network after convolution, and then input into the RepNCSPELAN4 module to obtain the P5 feature map; Among them, in the RepNCSPELAN4 module in the neck network: RepConvN extracts detail information through multiple convolution branches; the first convolution Conv1 uses a 3×3 convolution kernel with a stride of 1, and the second convolution Conv2 uses a 1×1 convolution with a stride of 1 to compress the input channels; the convolution output is added to the input, and nonlinear activation is performed through the Silu activation function; The CSP module performs cross-stage feature fusion, convolves the input feature map through two different paths, and generates two intermediate features. The two intermediate features are then concatenated through a 1×1 convolution. Finally, the concatenated features are processed through multiple bottleneck modules: Output=cv3(concat(m(cv2(x)),cv2(x))) Among them, m represents the combination of multiple bottleneck modules, concat represents concatenation, and cv3 is the last convolutional layer; Step 4.5: The MAH detection head performs convolution and average pooling on the P4 feature map and the P5 feature map; The processed P5 and P4 features are used as convolution kernels, the P3 feature map is traversed, and the features are fused and enhanced pixel by pixel to generate P5-P3mask and P4-P3mask respectively; The learnable parameters λ1 and λ2 perform weighted fusion on the P5-P3 mask and the P4-P3 mask, respectively. The parameters λ1 and λ2 are learned through the network and are used to dynamically adjust the influence of each mask. The fused feature map is further combined with the P3 feature map.
8. The method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model according to claim 7, characterized in that: The size of the annotated flame image input into the backbone network is 640×640×3; after passing through two convolution modules and a C2f module, it is compressed into a feature map of 160×160×128; after passing through the SPPF module at the end of the backbone network, the feature map size is 20×20×512.
9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model as described in any one of claims 1-8.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for segmenting the flame combustion area of a coal-fired power plant boiler burner based on the YOLOv8-MAH model as described in any one of claims 1 to 8 are implemented.