Electrical fire identification method based on DB-YOLOv8 and BTFF network
The image quality is enhanced through CycleGAN and SCI algorithms, combined with dual-branch and feature fusion network, and improved the YOLOv8 network, solving the image quality problem in electrical fire recognition and achieving efficient and accurate electrical fire detection.
Patent Information
- Application Number
- CN202510492473.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In the prior art, the high-resolution electrical fire images have low brightness or poor quality, which makes it difficult to accurately extract fire characteristics, increase background complexity, and high probability of misjudgment, making it difficult for existing methods to effectively identify electrical fires.
Image enhancement is used for CycleGAN network, dual branch backbone network is designed and CBAM module is integrated, and dual FPN and binary tree feature fusion network are combined to improve feature extraction and fusion capabilities, use SCI algorithm to adjust brightness, and improve YOLOv8 network for training.
It improves the accuracy and efficiency of electrical fire detection, reduces misjudgment, improves the safety of electrical equipment, and meets the development needs of smart grids.
Smart Images

Figure CN120356005A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision, and in particular relates to an electrical fire recognition method based on DB-YOLOv8 and BTFF networks. Background Art
[0002] Electricity plays an important role in modern society. While serving mankind as an important energy source for daily life and production, it has also been affected by various factors and has continuously caused major electrical fire accidents, which has attracted much public attention. Electrical fires have become the main type of fire that affects social fire safety in my country. Machine vision is a branch of artificial intelligence and an important part of artificial intelligence. Most artificial intelligences require machine vision to a greater or lesser extent. Simply put, machine vision uses electronic video and recording equipment to capture image information to achieve a visual effect similar to that of human and animal eyes.
[0003] In the prior art, an electrical safety hazard identification system based on machine vision with application number CN202011171771.9 includes several independent image acquisition modules, and the image acquisition module includes an image acquisition unit and a lighting system, and the image acquisition module is connected to a signal transmission module in a two-way signal. The present invention sets a power supply system, an image acquisition system, a control system and a signal display system; the identification method and device based on machine vision for automatic tracking and positioning of fire source points with application number CN202010430901.X clusters the fire source point images and extracts the fire source point features, and uses a deep learning method to construct a fire source point identifier, which can quickly identify the fire source point, reduce the time for subsequent identification of the fire source point, and improve the fire-fighting efficiency of the fire-fighting robot.
[0004] In the actual image recognition process, the collected image data is generally high-resolution images, and it is inevitable that there will be pictures with poor image quality or low brightness. In low-brightness or poor-quality images, the brightness, color and texture features of the flame may be masked, making it difficult for the detection algorithm to accurately extract fire-related features. For low-brightness images, simple contrast enhancement or brightness adjustment may not be able to restore the detailed information in the image, and may even introduce new distortion. Moreover, the fire scene is often accompanied by smoke or strong light. These factors will further reduce the image quality and interfere with the recognition of fire features. Electrical fire scenes may contain complex backgrounds. The difficulty of distinguishing between the background and the flame in low-quality images increases, and misjudgment is prone to occur.
[0005] Therefore, an electrical fire identification method based on DB-YOLOv8 and BTFF network is provided to solve the technical defects mentioned in the background technology. Summary of the invention
[0006] In view of the above-mentioned technical problems, the present invention provides an electrical fire recognition method based on the DB-YOLOv8 and BTFF networks.
[0007] In the first aspect, the present invention provides an electrical fire recognition method based on the DB-YOLOv8 and BTFF networks, including the following steps: Step 1: Obtain electrical fire pictures, preprocess the electrical fire pictures to obtain preprocessed images, and annotate the preprocessed images to obtain a training dataset; Step 2: Build a dual-branch backbone network based on the YOLOv8 backbone network, add a connection layer between the two branches of the dual-branch backbone network, and fuse the CBAM module into the SPPF module; Step 3: Respectively lead out three feature layers from each branch in Step 2 for FPN feature fusion, a total of six feature layers are led out, and the six feature layers are subjected to binary tree-type feature fusion to obtain an improved YOLOv8 network model; Step 4: Use the training dataset to train the improved YOLOv8 network model to obtain an electrical fire recognition model, input the electrical fire pictures to be recognized into the electrical fire recognition model, and obtain the electrical fire recognition result.
[0008] Specifically, in Step 1, the preprocessing includes: using the CycleGAN network to augment the electrical fire pictures, and when it is judged that the brightness or quality of the augmented pictures is less than the corresponding threshold, if so, using the SCI algorithm to enhance the augmented pictures, and after traversing all the augmented images, obtain the preprocessed images.
[0009] Specifically, the CycleGAN network uses two generators and two discriminators.
[0010] Specifically, in Step 2, based on the network structure of YOLOv8l, its backbone network is transformed into a dual-branch structure, and the dual-branch structure is respectively an S-branch backbone network and an M-branch backbone network. The S-branch backbone network is used to extract low-level scale features from high-resolution images, and the M-branch backbone network is used to extract high-level scale features from low-resolution images after downsampling.
[0011] Specifically, the first layer of the S-branch backbone network uses a 3×3 convolutional kernel with a stride of 2, and each of the remaining layers uses a 3×3 convolutional kernel in combination with a C2f_1_n module. The second and third layers of the M-branch backbone network are ordinary 3×3 convolutional kernels, and the structures of the remaining layers are the same as those of the S-branch backbone network.
[0012] Specifically, the CBAM module is composed of a channel attention module and a spatial attention module. The steps of integrating the CBAM module into the SPPF module are as follows: Step 21: Input the feature map. After passing through the channel attention module, by learning channel-level attention, when the network processes the input feature map, it automatically adjusts the weights of each channel. The network focuses on the feature channels useful for fire detection, ignores redundant or irrelevant channels, generates channel weights, and applies them to the input features. Step 22: The weighted feature map is then input into the spatial attention module for attention adjustment in the spatial dimension, enabling the network to focus on the key fire areas in the image and suppressing background noise. The spatial attention module helps the model more accurately locate the target, generates spatial weights, and applies them to the spatial dimension of the feature map. Step 23: The output of the channel attention module and the spatial attention module is a feature map that enhances important channels and spatial regions.
[0013] Specifically, in step 3, after the FPN feature fusion, a dual-FPN feature fusion network is formed. The six feature layers are subjected to binary tree-type feature fusion to form a binary tree-type feature fusion network. First, the feature layers with similar channels are fused pairwise, then the three obtained feature layers are subjected to binary tree-type feature fusion again, and finally, the two obtained fused feature layers are subjected to channel concatenation.
[0014] Specifically, in step 4, the improved YOLOv8 network model uses detection accuracy, recall rate, mean average precision, and average precision as evaluation metrics.
[0015] In a second aspect, the electrical fire recognition method provided by the present invention based on the DB-YOLOv8 and BTFF networks includes: an image data preprocessing module, an image feature extraction module, an image feature fusion module, and a fusion layer inspection and recognition module; The image data preprocessing module obtains electrical fire pictures, preprocesses the electrical fire pictures to obtain preprocessed images, and labels the preprocessed images to obtain a training dataset; The image feature extraction module constructs a dual-branch backbone network based on the YOLOv8 backbone network, adds a connection layer between the two branches of the dual-branch backbone network, and integrates the CBAM module into the SPPF module; The image feature fusion module respectively extracts three feature layers from each branch in step 2 for FPN feature fusion, a total of six feature layers are extracted, and the six feature layers are subjected to binary tree-type feature fusion to obtain an improved YOLOv8 network model; The fusion layer inspection and recognition module uses the training data set to train the improved YOLOv8 network model to obtain an electrical fire recognition model, and inputs the electrical fire picture to be recognized into the electrical fire recognition model to obtain the electrical fire recognition result.
[0016] Compared with the prior art, the beneficial effects of the technical solution of this application are at least as follows: 1. The present invention discloses an electrical fire recognition method based on DB-YOLOv8 and BTFF networks. In the data preprocessing stage, the CycleGAN generative adversarial network is used to supplement part of the data set to solve the problem of insufficient data set. The SCI algorithm is used to enhance some of the images with low brightness after supplementation. Moreover, a dual-branch backbone network designed based on the YOLOv8 backbone network is innovatively proposed, and the CBAM module is fused into the SPPF module; a dual-FPN feature fusion network and a binary tree-type feature fusion network are proposed to complete efficient feature extraction and feature fusion, reducing the loss of picture details. The processed data set is input into the network model for training and the test recognition of electrical fire images, and the recognition effect is ideal.
[0017] 2. The present invention can improve the detection effect of electrical fires, effectively reduce casualties, and improve the overall safety of electrical equipment. The present invention highly combines the concept of the digital and intelligent transformation of the power grid, conforms to the development trends of informatization, digitization, and digital and intelligentization in the power industry, and meets the fundamental needs of today's intelligent power grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0019] Figure 1 It is a flowchart of the electrical fire recognition method based on DB-YOLOv8 and BTFF networks of the present invention; Figure 2 It is a block diagram of the electrical fire recognition method based on DB-YOLOv8 and BTFF networks of the present invention; Figure 3 It is a schematic diagram of the CycleGAN network structure of the electrical fire recognition method based on DB-YOLOv8 and BTFF networks of the present invention; Figure 4 It is a structure diagram of the generator of the electrical fire recognition method based on DB-YOLOv8 and BTFF networks of the present invention; Figure 5 This is the discriminator structure diagram of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 6 This is the schematic diagram of the SCI network structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 7 This is the schematic diagram of the dual-branch backbone network structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 8 This is the schematic diagram of the C2f_1_n module structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 9 This is the schematic diagram of the Concat connection layer structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 10 This is the schematic diagram of the CBAM module structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 11 This is the schematic diagram of the GSPPF structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 12 This is the schematic diagram of the FPN structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 13 This is the schematic diagram of the binary tree type feature fusion network structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 14 This is the schematic diagram of the improved YOLOv8 network structure of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention; Figure 15 This is the relationship curve diagram of the model performance and confidence of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks of the present invention. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates on the present invention in conjunction with the accompanying drawings and embodiments. Obviously, the specific embodiments described herein are only used to explain the present invention, which are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0021] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0022] Figure 1 The figure shows a flowchart of an embodiment of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks provided by the present invention. The flowchart specifically includes the following steps: Step 1: Obtain electrical fire pictures, preprocess the electrical fire pictures to obtain preprocessed images, and annotate the preprocessed images to obtain a training dataset. Among them, the preprocessing includes: using the CycleGAN network to augment the electrical fire pictures. When it is determined that the brightness or quality of the augmented pictures is less than the corresponding threshold, if so, use the SCI algorithm to enhance the augmented pictures. After traversing all the augmented images, obtain the preprocessed images.
[0023] The CycleGAN network is a generative adversarial network (GAN) architecture for unsupervised image-to-image translation. It uses two generators and two discriminators, enabling the model to learn the mapping from one domain to another without paired training data. Specifically: CycleGAN is a symmetric structure. The upper part is to map the pictures in the X (real) source domain through G (generator) to generate pictures in the Y (fake) source domain. Dy will determine whether Y (fake) is a real picture or a generated picture. Using the idea of cycle consistency, the generated picture Y (fake) is remapped through F to generate a picture in the X (cycle) source domain. To prevent the situation where the generated picture X (cycle) loses the content of the picture in the X (real) source domain, the cycle consistency loss function is calculated between the generated picture X (cycle) and the picture in the X (real) source domain, making these two pictures as similar as possible, that is , and the lower part of the structure is to first map the Y (real) source domain through F to generate pictures in the X (fake) source domain, and then map them through G (generator) to generate Y (cycle), ultimately achieving , and its structure diagram is as Figure 3 shown.
[0024] The network structure of the generator consists of three parts: an encoder, a feature transformation module, and a decoder. The encoder includes three convolutional layers for processing the input image. First, a 7×7 convolutional kernel with a stride of 1 is used; then, two 3×3 convolutional kernels with a stride of 2 are used. All three convolutional layers use the ReLU activation function. Through the processing of these convolutional layers, the encoder realizes the reduction of the image size and increases the number of channels of the feature map. The feature transformation module is located at the core of the generator and consists of nine residual modules. The role of these residual modules is to extract the detailed features in the image. Each module contains convolutional operations. The structure of the decoder is similar to that of the encoder and is mainly composed of two transposed convolutional layers and one convolutional layer. The transposed convolution first uses interpolation (zero padding) to enlarge the feature map size, and then uses a convolutional kernel to perform a convolution operation on the feature map after interpolation, finally obtaining a large feature map. The transposed convolutional layer uses a 3×3 convolutional kernel with a stride of 2 and also uses the ReLU activation function. The role of these two transposed convolutional layers is to upsample the feature map output by the feature transformation module to restore the image size and at the same time reduce the number of channels. The size of the convolutional kernel of the last convolutional layer is 7×7, and the activation function is tanh. Such a design enables the generator to effectively learn the mapping from the input domain to the target domain, convert the input image into an image in the target domain through a convolutional neural network, and maintain the main features and details of the image. The generator gradually learns through training. The feature transformation module between the encoder and the decoder plays a role of connection and transformation, generating more real images similar to the distribution of the target domain. The structure of the generator is as shown in Figure 4 shown.
[0025] In CycleGAN, the discriminator adopts the PatchGAN model instead of the traditional design that outputs a single probability value. The traditional discriminator usually outputs a probability value indicating that the input image is real, as a probability in a binary classification task. What makes PatchGAN different is that it divides the input image into multiple local regions and outputs an independent probability value for each region. These probability values form a matrix, usually 16×16, where each matrix element represents the probability that the corresponding region in the image is real. The design concept of PatchGAN is to enhance the discriminator's perception ability of local details in the image through this local region evaluation. The receptive field size corresponding to each matrix point is set to 70×70 by parameters, which means that the discriminator can focus on different local information of the image and thus make a more accurate authenticity judgment. Finally, CycleGAN will take the average of all the values in this probability matrix to obtain the final evaluation result of the entire image. In terms of the specific structure, the discriminator of CycleGAN usually consists of five convolutional layers. The size of the convolutional kernel in each layer is 4×4, the stride is 2, and the activation function is LeakyReLU. By adopting the PatchGAN model, the discriminator of CycleGAN can output a local region authenticity evaluation in the form of a matrix, making the model more efficient in capturing local features of the image. The goal of the discriminator is to judge whether the input image is a real image or a generated image through classification, and optimize the model through training so that it can identify the fake images generated by the generator and give a higher score to real images. The discriminator and the generator are in an adversarial training process. The goal of the generator is to try to deceive the discriminator so that the generated images look real enough. Finally, through continuous training, the generator can generate more and more real images. The discriminator structure is as Figure 5 shown.
[0026] The SCI self-calibrated illumination network consists of an illumination estimation and a self-calibrated module. The self-calibrated module estimates and adjusts the illumination conditions of the image through a deep learning model, automatically adjusts the under-exposed or over-exposed parts in the image, improves the exposure stability and greatly reduces the computational burden. As the input for the next-stage illumination estimation, there is a relationship between the underexposed image y and the clear image z. The most important part is the illumination component x. The SCI network structure is as Figure 6 shown. The calculation formula of the illumination estimation module is:
[0027] where F is the main function of the illumination estimation module; is the residual of the t-th stage, which can improve the exposure stability and greatly reduce the computational burden. Its function is to learn a part of the illumination component at each stage in the form of a cascaded network, and finally learn the entire illumination component; is the illumination component of the t-th stage, the initial illumination component is the original input low-luminance image; is the parameter introduced by the illumination estimation module mapping to learn the illumination component and is independent of the number of stages, that is, weight sharing is maintained at each stage.
[0028] The calculation formula of the self-calibration module is:
[0029] where G is the main function of the self-calibration module; is the target image output at the t-th stage; represents element-wise division; is the self-calibration mapping at the t-th stage; is the introduced parameterized operator with learnable parameters ; is the input used for the next stage after calibration, and the number of stages 1 ≤ t ≤ T. The self-calibration module connects the inputs of each stage except the first stage with the original low-illumination input, that is, the input of the first stage, so as to explore the convergence behavior between each stage and introduce a self-calibration mapping , to represent the difference between the input of each level and the input of the first level. The self-calibration module can ensure that the outputs of different stages can converge to the same state during training. Finally, the conversion at the t-th stage (1 ≤ t ≤ T) can be expressed as:
[0030] The loss function of SCI consists of a fidelity loss and a smoothing loss. Unsupervised learning is used to expand the capabilities of the network. The fidelity loss is to ensure the pixel-level consistency between the estimated illumination component and the input of each stage. Its formula is:
[0031] where T is the total number of stages.
[0032] The smoothing loss can make the overall brightness transition of the picture tend to be smooth and avoid over-bright or over-dark phenomena in a certain area. Its formula is:
[0033] Among them, N is the total number of pixels, i represents the i-th pixel, and N(i) represents the adjacent pixels of i in the 5×5 sliding window. represents the weight; and are the i-th and j-th illumination components respectively.
[0034] The formula for the weight is:
[0035] Among them, c is the image channel in the YUV color space; = 0.1 is the standard deviation of the Gaussian kernel.
[0036] The formula for the total loss function is:
[0037] Among them, and are the balance coefficients.
[0038] Step 2: Based on the YOLOv8 backbone network, construct a dual-branch backbone network, that is, the DB-YOLOv8 network. Add a connection layer between the two branches of the dual-branch backbone network, and fuse the CBAM module into the SPPF module. Among them, based on the network structure of YOLOv8l, transform its backbone network into a dual-branch structure. While reducing the number of parameters, make it focus on small targets and not affect the detection of large targets. The dual-branch structure is respectively the S-branch backbone network and the M-branch backbone network. The S-branch backbone network is used to extract low-level scale features from high-resolution images, and the M-branch backbone network is used to extract high-level scale features from the low-resolution images after downsampling. At the same time, a connection layer is adopted between the two branches to perform feature fusion. The structure diagram is as Figure 7 shown.
[0039] When the S-branch backbone network is used, the high-resolution image is directly input into the S-branch for feature extraction operations without downsampling, focusing on small target features. The first layer of the S-branch backbone network uses a 3×3 convolutional kernel with a stride of 2 to halve the size of the output feature map and increase the number of channels to 24. The remaining structure of the S-branch backbone network is a 3×3 convolutional kernel paired with a C2f_1_n module for feature extraction, optimizing the network structure and improving the performance of the model. The C2f_1_n module improves the computational efficiency of the model, reduces the computational cost, enhances the feature representation ability, and at the same time maintains good accuracy. Through feature separation, the amount of computation and memory usage can be significantly reduced. The structure of the C2f_1_n module is as Figure 8As shown in the figure, the input feature map is first subjected to a 1×1 convolution operation and then Split into two parts; the separated feature maps are subjected to a convolution operation with a 3×3 convolution kernel for two layers, and then the input features and the output features after convolution are concatenated through a skip connection to extract richer features; the feature maps of the two separated paths after processing are concatenated, and then a 1×1 convolution operation is performed to output the final feature map.
[0040] The second and third layers of the M-branch backbone network are ordinary 3×3 convolution kernels. The structures of the remaining layers are the same as those of the S-branch backbone network. The number of C2f_1_n modules is modified to be twice that of the S-branch, and the output number of channels is {128, 256, 512, 512}. Moreover, the M-branch backbone network adds a Concat connection layer to perform feature fusion with the S-branch backbone network, and the structure diagram of the Concat connection layer is as Figure 9 shown. The outputs of the second, third, and fourth C2f_1_n modules of the S-branch are concatenated with the inputs of the first, second, and third C2f_1_n modules of the M-branch. Feature fusion is performed through a 1×1 convolution layer, and then the number of channels is adjusted through a 1×1 convolution layer. The adjusted number of channels is {128, 256, 512} as the input of the first three C2f_1_n modules of the M-branch.
[0041] The SPPF module is a spatial pyramid pooling acceleration structure that obtains information at different scales through max-pooling operations. The CBAM module is introduced inside the SPPF module to complete the transformation of the SPPF module structure, and the feature expression ability is improved through feature fusion at different scales while reducing the computational cost.
[0042] The CBAM module consists of a channel attention module and a spatial attention module, and the structure is as Figure 10 shown. The steps of integrating the CBAM module into the SPPF module are as follows: S1: First, the input feature map passes through the channel attention module. By learning the channel-level attention, when the network processes the input feature map, it automatically adjusts the weight of each channel. The network focuses on the feature channels useful for fire detection, ignores redundant or irrelevant channels, and the network extracts more meaningful features from the feature map, enhancing the model's detection ability for complex scenes and small-target fires. The calculation formula is:
[0043] where represents the sigmoid activation function, is the channel attention weight, is the input feature map, and the channel weights can be generated and applied to the input feature; S2: Then, the weighted feature map is input into the spatial attention module to adjust the attention in the spatial dimension, enabling the network to focus on the key fire areas in the image and suppressing background noise. The spatial attention module helps the model to more accurately locate the target, and its calculation formula is:
[0044] where represents the sigmoid activation function, is the spatial attention weight, represents the channel concatenation operation, generating the spatial weight and acting on the spatial dimension of the feature map; S3: The output of the channel attention module and the spatial attention module is a feature map that enhances important channels and spatial regions.
[0045] While the feature map output by the CBAM module is input into the SPPF module, in order to further reduce the detail loss caused by spatial pyramid pooling, a grouping operation is used. The GSPPF structure is as Figure 11 shown. One group is the S group, which only performs feature fusion operations through a 1×1 convolutional layer to retain the features of small targets to the greatest extent. The other group is the M group, which halves the number of channels through a 1×1 convolutional layer and then performs spatial pyramid pooling operations. The SPPF introduces spatial pyramid pooling, which can perform pooling operations at different scales, converge features of different sizes, and effectively extract multi-scale features. The resolution of small target open flames in the image is relatively low and is usually difficult to be effectively captured by the convolutional neural network. The SPPF enhances the feature extraction at different scales, enabling better feature expression of small target open flames. The feature map output after the SPPF passes through two 1×1 convolutional layers, first performs feature fusion and then halves the number of channels, and then performs channel concatenation with the feature map from the S group, and outputs the final fused feature map through a 1×1 convolutional layer. While increasing a small amount of computational cost, it reduces the loss of small target features to enhance feature expression and improve detection accuracy.
[0046] Step 3: Each branch in Step 2 is respectively led out three feature layers for FPN feature fusion, a total of six feature layers are led out, and the six feature layers are subjected to binary tree-shaped feature fusion to obtain the improved YOLOv8 network model. Among them, after the FPN feature fusion, a dual FPN feature fusion network is formed, and after the six feature layers are subjected to binary tree-shaped feature fusion, a binary tree-shaped feature fusion network, that is, a BTFF network, is formed. First, the feature layers with similar channels are fused in pairs, then the three obtained feature layers are subjected to binary tree-shaped feature fusion again, and finally the two obtained fused feature layers are subjected to channel concatenation.
[0047] Specifically, after the FPN feature fusion, a dual-FPN feature fusion network is formed. The multi-scale feature layers output by the second, third, and fourth C2f_1_n modules in the S branch are set as , and the obtained multi-scale feature layers are processed by FPN. The processed feature layers are , and the structure diagram of FPN is as Figure 12 shown. The processing idea is to fuse the high-level (low-resolution) and low-level (high-resolution) feature maps to enhance the network's perception ability of targets at different scales. FPN upsamples the low-resolution high-level feature maps and fuses them into the corresponding high-resolution low-level feature maps. Through multi-scale feature fusion, the semantic information of the low-level feature maps is enhanced, enabling the network to capture the features of small targets from more detailed information and better detect targets at different scales. FPN outputs feature maps at multiple scales for target detection of different sizes.
[0048] The multi-scale feature layers output by the second and third C2f_1_n modules in the M branch and the feature layer output by the GSPPF module are set as , and the obtained multi-scale feature layers are processed by FPN. The processed feature layers are . After double-FPN processing, six feature layers are obtained , and these six feature layers will undergo the final binary tree-type feature fusion to further enhance the detection effect of targets at different scales.
[0049] In step 3, after the six feature layers undergo binary tree-type feature fusion to form a binary tree-type feature fusion network, the six feature layers obtained after the FPN feature fusion of the dual-branch backbone network are input into the binary tree-type feature fusion network, and its structure is as Figure 13 shown.
[0050] Among them, the three feature layers from the S branch will be halved in size through a 3×3 convolutional kernel and then paired up , and after passing through the Concat layer, is obtained. The feature layer is led out for detection; then after being paired up and adjusted in size through a 3×3 convolutional kernel and then undergoing Concat splicing, is obtained. The feature layer is led out for detection; finally undergoes Concat splicing to obtain . The feature layer is led out for detection. This structure continuously fuses low-level features into high-level features, continuously strengthens the low-level features in high-level features, and pays more attention to the low-level features in multi-scale features.
[0051] Step 4: Use the training dataset to train the improved YOLOv8 network model to obtain an electrical fire recognition model. Input the electrical fire picture to be recognized into the electrical fire recognition model to obtain the electrical fire recognition result.
[0052] In Step 4, the deployment environment of the improved YOLOv8 network model is Python 3.8, the CPU is the Intel(R) Xeon(R) Gold 5418Y processor, the GPU selects the 24GB graphics card of NVIDIA GeForce RTX 4090, the operating system is Ubuntu 20.04, the deep learning framework based on PyTorch 1.11.0 is selected, the CUDA version is 11.3, the number of training rounds is set to 200, the batch size is set to 64, the learning rate is set to 0.01, the weight decay coefficient is 0.0005, and the momentum is 0.937. The network structure of the improved YOLOv8 in the present invention is as Figure 14 shown. Conduct experimental verification on the obtained YOLOv8 network model, and use detection accuracy, recall rate, mean average precision, and average precision as evaluation indicators. Precision refers to the proportion of actual positive samples among all positive samples predicted by the model, and it measures the prediction accuracy of the model for positive samples; Recall refers to the proportion of all actual positive samples that are correctly predicted, and it measures the coverage of the model for positive samples; AP is obtained by calculating the area under the precision-recall curve (PR curve). AP comprehensively reflects the performance of precision and recall at different thresholds and reflects the detection effect of the model at all prediction thresholds. The larger the AP value, the higher the detection performance of the model; mAP is the average of AP and is used to measure the overall performance of a multi-class object detection model, which is the result of averaging the AP values of all classes. Among them, YOLOv5, YOLOv7, and YOLOv8 are selected as the comparison groups for experiments, and the comparison experiment results are shown in the following table:
[0053] The F1 score can also reflect the performance of the object detection model. It is the harmonic mean of precision (Precision) and recall (Recall), and is used to balance the performance of precision and recall. The relationship curve between the F1 score and the confidence level Figure 15 is shown.
[0054] In the data preprocessing stage of the present invention, the CycleGAN generative adversarial network is adopted to supplement part of the dataset, solve the problem of insufficient dataset, and enhance some images with low brightness after supplementation by the SCI algorithm. An innovative dual-branch backbone network designed based on the YOLOv8 backbone network is proposed, and the CBAM module is integrated into the SPPF module; a dual-FPN feature fusion network and a binary tree-shaped feature fusion network are proposed to complete efficient feature extraction and feature fusion, reducing the loss of details in the pictures.
[0055] Figure 2 The following is a schematic structural diagram of an embodiment of the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks provided by the present invention, including: an image data preprocessing module, an image feature extraction module, an image feature fusion module, and a fusion layer inspection and recognition module; The image data preprocessing module supplements part of the dataset by using the CycleGAN generative adversarial network in the data preprocessing stage and enhances some images with low brightness after supplementation by the SCI algorithm; The image feature extraction module is based on the YOLOv8 backbone network and redesigned into a dual-branch backbone network. A connection layer is added in the dual-branch backbone network to enhance the feature fusion and information exchange between the two branches, and the CBAM module is integrated into the SPPF module; The image feature fusion module respectively leads out three feature layers from each branch in step 2 for FPN feature fusion. After the two-branch feature fusion, a total of six feature layers will be led out. The six feature layers are subjected to binary tree-shaped feature fusion. First, the feature layers with similar channels are fused in pairs, then the three obtained feature layers are subjected to binary tree-shaped feature fusion again, and finally the two obtained fusion feature layers are subjected to channel splicing; The fusion layer inspection and recognition module trains and validates the network model with the preprocessed dataset, and tests and recognizes electrical fire images with the obtained weights.
[0056] According to another aspect of the embodiments of the present invention, a storage medium is provided. The storage medium stores program instructions, wherein when the program instructions run, they control the device where the storage medium is located to execute the electrical fire recognition method based on the DB-YOLOv8 and BTFF networks.
[0057] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0058] The above-mentioned embodiments only represent the preferred implementation modes of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. An electrical fire recognition method based on the DB-YOLOv8 and BTFF networks, characterized in that It includes the following steps: Step 1: Obtain electrical fire pictures, preprocess the electrical fire pictures to obtain preprocessed images, and annotate the preprocessed images to obtain a training dataset; Step 2: Based on the YOLOv8 backbone network, construct a dual-branch backbone network, add a connection layer between the two branches of the dual-branch backbone network, and fuse the CBAM module into the SPPF module; Step 3: Respectively lead out three feature layers from each branch in Step 2 for FPN feature fusion, a total of six feature layers are led out, and the six feature layers are subjected to binary tree-type feature fusion to obtain an improved YOLOv8 network model; Step 4: Use the training dataset to train the improved YOLOv8 network model to obtain an electrical fire recognition model, input the electrical fire picture to be recognized into the electrical fire recognition model, and obtain the electrical fire recognition result.
2. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 1, wherein In Step 1, the preprocessing includes: using the CycleGAN network to augment the electrical fire pictures, judging whether the brightness or quality of the augmented pictures is less than the corresponding threshold, if so, using the SCI algorithm to enhance the augmented pictures, and after traversing all the augmented images, obtain the preprocessed images.
3. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 2, characterized in that, The CycleGAN network uses two generators and two discriminators.
4. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 1, wherein In Step 2, based on the network structure of YOLOv8l, transform its backbone network into a dual-branch structure. The dual-branch structure is respectively an S-branch backbone network and an M-branch backbone network. The S-branch backbone network is used to extract low-level scale features from high-resolution images, and the M-branch backbone network is used to extract high-level scale features from low-resolution images after downsampling.
5. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 4, characterized in that, The first layer of the S-branch backbone network uses a 3×3 convolutional kernel with a stride of 2, and each of the remaining layers uses a 3×3 convolutional kernel paired with a C2f_1_n module. The second and third layers of the M-branch backbone network are ordinary 3×3 convolutional kernels, and the structures of the remaining layers are the same as those of the S-branch backbone network.
6. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 1, characterized in that, The CBAM module consists of a channel attention module and a spatial attention module. The steps of fusing the CBAM module into the SPPF module are: Step 21: Input the feature map, pass through the channel attention module, and by learning channel-level attention, enable the network to automatically adjust the weight of each channel when processing the input feature map. The network focuses on the feature channels useful for fire detection, ignores redundant or irrelevant channels, generates channel weights and acts on the input features; Step 22: The weighted feature map is then passed into the spatial attention module to adjust the attention in the spatial dimension, enabling the network to focus on the key fire areas in the image, suppressing background noise. The spatial attention module helps the model to more accurately locate the target, generates spatial weights, and acts on the spatial dimension of the feature map; Step 23: What is output after passing through the channel attention module and the spatial attention module is a feature map that enhances important channels and spatial regions.
7. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 1, wherein, In step 3, after the FPN feature fusion, a dual-FPN feature fusion network is formed. The six feature layers are subjected to binary tree-type feature fusion to form a binary tree-type feature fusion network. First, the feature layers with similar channels are fused pairwise, then the three obtained feature layers are subjected to binary tree-type feature fusion again, and finally, the channel splicing is performed on the two obtained fused feature layers.
8. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 1, wherein In step 4, the improved YOLOv8 network model uses detection accuracy, recall rate, mean average precision, and average precision as evaluation metrics.
9. The electrical fire recognition method based on the DB-YOLOv8 and BTFF networks according to claim 1, wherein, It includes: An image data preprocessing module, an image feature extraction module, an image feature fusion module, and a fusion layer inspection and recognition module; The image data preprocessing module obtains electrical fire pictures, preprocesses the electrical fire pictures to obtain preprocessed images, and labels the preprocessed images to obtain a training dataset; The image feature extraction module constructs a dual-branch backbone network based on the YOLOv8 backbone network, adds a connection layer between the two branches of the dual-branch backbone network, and fuses the CBAM module into the SPPF module; The image feature fusion module respectively extracts three feature layers from each branch in step 2 for FPN feature fusion, a total of six feature layers are extracted, and the six feature layers are subjected to binary tree-type feature fusion to obtain an improved YOLOv8 network model; The fusion layer inspection and recognition module uses the training dataset to train the improved YOLOv8 network model to obtain an electrical fire recognition model, inputs the electrical fire pictures to be recognized into the electrical fire recognition model, and obtains the electrical fire recognition result.
Citation Information
Patent Citations
Identification method and device for automatically tracking and positioning fire source point based on machine vision
CN111738082A
Electrical potential safety hazard identification system based on machine vision
CN112183475A
Fire detection method based on PD-YOLO
CN117333753A
Power field operation specification detection method based on improved YOLOv8 algorithm
CN118736307A
Data collecting method and apparatus, and storage medium and system
WO2020078385A1