Transformer substation fire and smoke real-time identification method based on improved RT-DETR

By improving the RT-DETR model, combining ContextGuidedBlock, Lion optimizer and Alpha-IoU loss function, the accuracy and real-time problems of substation fire and smoke recognition are solved, efficient fire monitoring in complex environments is achieved, and the reliability of substation safety monitoring is improved.

CN120339827APending Publication Date: 2025-07-18CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510384873.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing substation fire and smoke identification technologies are difficult to accurately and in real time to identify fire sources or smoke signals in complex environments, and it is difficult to distinguish between fire signals and environmental interference, resulting in false alarms and missed alarms, affecting safety monitoring efficiency.

Method used

Using the improved RT-DETR model, by constructing high-quality fire datasets, using the LabelImg tool to annotate images, and introducing the ContextGuidedBlock module, Lion optimizer and Alpha-IoU loss function, the model's recognition ability under different lighting conditions is optimized.

Benefits of technology

Accurate identification of substation smoke and fires under different lighting conditions improves the model's adaptability in complex environments, reduces missed judgments and false alarms, ensures all-weather fire monitoring effect, and improves the reliability and real-timeness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339827A_ABST
    Figure CN120339827A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer substation fire and smoke real-time identification method based on improved RT-DETR. The method comprises the following steps: collecting and sorting flame and smoke images shot in a transformer substation; marking the image data set by using Labelimg through visual inspection and data analysis; an original RTDETR model is improved, a BaseBlock module of an original backbone network is optimized, a ContextGuideBlock module is introduced and replaced with a BaseBlock GC module, an optimizer is changed from Adam to Lion, Alpha-IoU is adopted as a new loss function, and a new model with optimized performance is constructed; and inputting the fire and smoke images of the transformer substation into the trained improved model to generate a detection result, and accurately positioning the positions of flames and smoke and improving the recognition precision. The method provided by the invention can detect fire and smoke abnormities in the transformer substation more accurately in real time, realizes efficient positioning of fire and smoke sources, has important practical application value, and provides powerful guarantee for safe operation of a power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to target detection technology. Specifically, it relates to a method for quickly detecting flames and smoke during a substation fire. Background Art

[0002] As an important part of the power system, the substation has a complex operating environment, a wide variety of equipment, and the fire risk has always been an important hidden danger threatening safe operation. Once a fire occurs, if it cannot be detected and handled in time, it may cause serious damage to equipment, power interruption, and even threaten the safety of the surrounding environment and personnel. Therefore, how to achieve rapid and accurate identification of substation fires and smoke has become an urgent problem to be solved.

[0003] Existing fire monitoring systems mainly rely on traditional sensor technologies, such as temperature sensors and smoke sensors. Although these technologies are widely used, they are easily interfered with in complex environments and have problems of false alarms and missed alarms. In addition, their response speed to the initial stage of a fire is slow, making it difficult to meet the requirements of real-time early warning.

[0004] With the development of image recognition technology, vision-based fire and smoke detection methods have gradually attracted attention. However, in the substation environment, the characteristics of flames and smoke are often interfered with by complex backgrounds and lighting, such as day and night. These factors make traditional image detection methods insufficient in terms of accuracy and real-time performance, and unable to meet the requirements of high reliability and high efficiency for substation fire monitoring.

[0005] Therefore, there is an urgent need for a detection method that can adapt to the special environment of the substation and balance accuracy and real-time performance in fire and smoke recognition to improve the overall level of substation fire prevention and control.

[0006] To solve the above technical defects, the patent with the application number CN2022114122136 proposes a substation fire recognition method and system, which includes the following steps: S1: Obtain the key frame images in the video data and perform preprocessing of the images; S2: Use the YOLOv3 algorithm to locate the preprocessed images, and obtain the image contour of the fire through edge detection; S3: Extract the fractal dimension of the contour image and compare it with a preset threshold. When it is greater than the preset threshold, it is determined that a fire has occurred, otherwise, continue to process the video data. Although it can achieve the recognition of substation fires, the accuracy of the image recognition benchmark model used is low, it is difficult to accurately and real-time recognize, and it cannot effectively predict possible fire situations in the substation, which may lead to missed detections.

[0007] Therefore, the applicant proposes a real-time recognition method for substation fires and smoke based on improved RT-DETR. Summary of the Invention

[0008] The present invention aims to solve the limitations of existing substation fire and smoke recognition technologies; specifically, when existing technologies detect fires and smoke in complex and variable environments, there are technical problems such as difficulty in accurately and real-time identifying the fire source or smoke signal, and it is also difficult to precisely distinguish the specific differences between fire signals and other environmental interferences. These technical bottlenecks often lead to false alarms and missed alarms, thus having an adverse impact on the safety monitoring efficiency of substations.

[0009] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0010] A real-time substation fire and smoke recognition method based on improved RT-DETR, comprising the following steps:

[0011] Step 1: Collect and organize a number of fire images to construct a fire dataset;

[0012] Step 2: Label the target dataset, process the labels of the collected images one by one, construct a substation image dataset containing fire and smoke scenes, and divide this dataset into a training set, a validation set, and a test set;

[0013] Step 3: Obtain an improved RT-DETR model and use this model to recognize the collected substation fire and smoke images;

[0014] Step 4: Verify the model, including testing its robustness and detection accuracy under different lighting conditions during the day and at night; ensure that even in low-light environments at night, the model can still effectively recognize fires and smoke, guaranteeing all-weather detection capabilities.

[0015] Step 1 specifically includes the following steps:

[0016] S1.1: Extensively collect data on typical substation fire cases;

[0017] S1.2: Carefully screen and organize relevant fire images;

[0018] S1.3: Preprocess the dataset to reduce the inaccuracy and contingency of some dataset images.

[0019] 3. According to the method described in claim 1, wherein step 2 specifically includes the following steps:

[0020] S2.1: Perform annotation in the Anaconda environment. Use the labelimg image annotation tool to annotate the collected substation flame and smoke images. The recognition objects are divided into two categories: smoke and flame. Among them, smoke is annotated as "0", and flame is annotated as "1"; obtain the Pascal VOC format data label, and then convert it into a txt file suitable for training through the data format conversion code written in Python. The file includes most of the data, including the x coordinate, y coordinate, width, height and other annotation information of the center point of the bounding box;

[0021] S2.2: After completing the annotation work, divide the dataset to obtain the training set, validation set, and test set of the substation fire dataset.

[0022] In step 3, the improved RT-DETR model is specifically:

[0023] Initialize the input image and input it to the first ConvN 3*3 module. The output of the first ConvN 3*3 module is connected to the input of the first ConvN 3*3 module. The output of the first ConvN 3*3 module is connected to the input of the second ConvN 3*3 module. The output of the second ConvN 3*3 module is connected to the input of the third ConvN 3*3 module. The output of the third ConvN 3*3 module is connected to the input of the Maxpool2d module. The output of the Maxpool2d module is connected to the input of the first Basic_Block_CG module. The output of the first Basic_Block_CG module is connected to the input of the second Basic_Block_CG module. The output of the second Basic_Block_CG module is connected to the input of the third Basic_Block_CG module and the input of the first Conv 1*1 module. The output of the third Basic_Block_CG module is connected to the input of the fourth Basic_Block_CG module and the input of the second Conv 1*1 module. The output of the fourth Basic_Block_CG module is connected to the input of the first ConvN 1*1 module. The output of the first ConvN 1*1 module is connected to the input of the AIFI module;

[0024] The output of the AIFI module is connected to the inputs of the second Conv 1*1 module and the third Conv 1*1 module. The output of the third Conv 1*1 module is connected to the input of the first upsampling module. The outputs of the first upsampling module and the second Conv 1*1 module are connected to the input of the first splicing module. The output of the first splicing module is connected to the input of the first RepC3 module. The output of the first RepC3 module is connected to the input of the fourth Conv 1*1 module. The output of the fourth Conv 1*1 module is connected to the inputs of the second splicing module and the second upsampling module. The outputs of the second upsampling module and the first Conv 1*1 module are connected to the input of the third splicing module. The output of the third splicing module is connected to the input of the second RepC3 module. The output of the second RepC3 module is connected to the input of the second ConvN 3*3 module. The output of the second ConvN 3*3 module is connected to the input of the second splicing module. The output of the second splicing module is connected to the input of the third RepC3 module. The output of the third RepC3 module is connected to the input of the third ConvN 3*3 module. The output of the third ConvN 3*3 module is connected to the input of the fourth splicing module. The output of the fourth splicing module is connected to the input of the fourth RepC3 module;

[0025] The outputs of the second RepC3 module, the third RepC3 module, and the fourth RepC3 module are connected to the input of the minimum query mechanism module. The output of the minimum query mechanism module is connected to the input of the decoder module equipped with an auxiliary prediction head. The decoder module equipped with an auxiliary prediction head outputs the total features.

[0026] The structure of the Basic_Block_CG module is specifically as follows:

[0027] The initial input features of the Basic_Block_CG module are input to the ConvN3*3 module. The output of the ConvN3*3 module is connected to the input of the CG Block module. The output of the CG Block module is connected to the input of the splicing module. The output of the splicing module is connected to the input of the ReLU module. The output of the ReLU module is the refined features.

[0028] The CG Block module is specifically as follows:

[0029] The initialization input features of the CG Block module are input to the input of the 1*1Conv module and the input of the output module. The output of the 1*1Conv module is connected to the input of the local feature extractor module and the input of the surrounding context extractor module. The output of the local feature extractor module is connected to the input of the 3*3Conv module. The output of the surrounding context extractor module is connected to the input of the 3*3DConv module. The outputs of the 3*3Conv module and the 3*3DConv module are connected to the input of the concatenation module. The output of the concatenation module is connected to the input of the BN / PReLU module. The output of the BN / PReLU module is connected to the input of the GAP module and the input of the output module. The output of the GAP module is connected to the input of the first FC module. The output of the first FC module is connected to the input of the second FC module. The output of the second FC module is connected to the input of the output module.

[0030] The Alpha-IoU loss function is used as the bounding box loss function for the improved RT-DETR model.

[0031] In step 4, tests are conducted through two typical scenarios of day and night to ensure that the model can operate stably under different lighting conditions; at night, verify whether the model can still effectively identify fires or smoke under low-light or dark conditions.

[0032] Specifically, it includes verifying the accuracy and fire detection ability of the model, as well as the real-time object detection and response ability.

[0033] In step 3, it includes the evaluation of the model performance, and Precision, Recall, mean Average Precision (mAP), and Frames Per Second (FPS) are used as the main metrics.

[0034] Compared with the prior art, the present invention has the following technical effects:

[0035] 1) The present invention can accurately identify substation smoke and fires under different lighting conditions, especially in two extreme lighting scenarios of day and night, thereby significantly improving the adaptability of the model in complex environments. In practical applications, the model can automatically adjust according to the changes in scene light to ensure effective fire monitoring regardless of how the lighting conditions change;

[0036] 2) In an environment with sufficient daylight, the model can clearly identify the smoke and fire characteristics in the substation, including the initial signs of the fire and the diffusion pattern of the smoke, thereby achieving efficient and accurate fire early warning. This technical effect significantly improves the accuracy of daytime detection and reduces the risk of missed judgment;

[0037] 3) In low-light environments, especially at night or under other conditions with weak light, the model can still maintain strong anti-interference ability and accurately judge the characteristics of smoke and fire. By strengthening the ability to recognize features under low-light conditions, the model effectively reduces the recognition errors caused by insufficient light and improves the fire monitoring effect at night or in dim environments;

[0038] 4) Even when only smoke or fire exists in the scene, the model can independently identify them according to their respective characteristics, ensuring the correct distinction between fire and non-fire situations. This ability enables the model to make accurate judgments even when the information is incomplete, thereby improving the reliability of the overall system and avoiding false alarms or missed alarms;

[0039] In summary, the present invention improves the stability and accuracy of the system in various complex environments. Whether in the daytime, at night, or in environments with different light intensities, the model can efficiently identify fires, ensuring its excellent performance in practical applications;

[0040] The present invention proposes a real-time recognition method for substation fires and smoke based on improved RT-DETR, which combines the Lion optimizer and Alpha-IoU to further improve the performance of the model. The improved model can efficiently identify smoke and fires in substations, especially under two extreme lighting conditions of daytime and night, showing remarkable detection accuracy and robustness. In the well-lit daytime, the model can clearly identify the characteristics of smoke and fire, while in low-light environments, the model introduces the ContextGuidedBlock module to better process complex scenes and long-range dependence information, ensuring the accurate judgment of fire signs even at night with insufficient light. By improving RT-DETR, the model can effectively enhance the attention to key features and improve the recognition ability in complex environments. The model not only has a high fire recognition accuracy but also can monitor the safety status of substations in real time, providing strong technical support for the safe operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The following further illustrates the present invention in conjunction with the drawings and embodiments:

[0042] Figure 1 is the schematic diagram of the network structure of the present invention;

[0043] Figure 2 is the schematic diagram of the channel comparison between the BasicBlock_CG module of the present invention and the original position module;

[0044] Figure 3 is Figure 2 the schematic diagram of the CGBlock structure in

[0045] Figure 4Schematic diagram for detection when there is only smoke in the present invention;

[0046] Figure 5 Schematic diagram for detection when there is only a flame in the present invention;

[0047] Figure 6 Fire situation prediction diagram in the daytime in the present invention;

[0048] Figure 7 Fire situation prediction diagram at night in the present invention. Detailed implementation manners

[0049] A real-time recognition method for substation fire and smoke based on improved RT-DETR includes the following steps:

[0050] Step S1: Through comprehensively summarizing typical substation fire cases common on the network, collecting and organizing relevant fire images, aiming to construct a high-quality fire dataset. This dataset should not only cover a large number of fire cases, but also contain various different fire scenarios to ensure its sufficient diversity to adapt to the analysis needs in different environments. The images in the dataset have high resolution and excellent clarity, ensuring rich details and no distortion, facilitating subsequent image processing and analysis.

[0051] To ensure the representativeness and integrity of the dataset, it is necessary to collect from multiple dimensions, including covering three different situations of only fire, only smoke, and both fire and smoke occurring simultaneously, and distinguishing and judging the fire occurrence situation according to different lighting conditions during the day and at night. By ensuring that the dataset covers sufficient actual scenarios and fire situations unique to substations, the learning ability of the subsequent model can be effectively improved, thereby providing more accurate and reliable data support for applications such as fire warning systems and automatic detection systems;

[0052] Step S2: Use the LabelImg tool to finely annotate the target dataset, and process the labels of the collected images one by one. Divide the dataset according to a reasonable ratio, and finally construct a substation image dataset containing fire and smoke scenarios. This dataset will be further divided into a training set, a validation set, and a test set to ensure the comprehensiveness and accuracy of subsequent model training;

[0053] Step S3: Obtain the improved RT-DETR model, and use this model to identify the collected substation fire and smoke images. When evaluating the model performance, use precision, recall, mean average precision (mAP), and frames per second (FPS) as the main indicators.

[0054] Step S4: Validate the model in complex typical scenarios, especially test its robustness and detection accuracy under different lighting conditions of day and night. Ensure that the model can still effectively identify fires and smoke even in low-light environments at night, guaranteeing all-weather detection capabilities.

[0055] RT-DETR (Real-Time Detection Transformer) is a variant of the end-to-end object detector based on DETR. It inherits the characteristics of DETR in using the Transformer architecture and self-attention mechanism to capture long-range dependencies in images. RT-DETR effectively reduces the consumption of computing resources by innovatively decomposing the multi-scale feature interaction process into two steps: intra-scale interaction and cross-scale fusion, and omits the complex post-processing steps such as candidate region generation and non-maximum suppression (NMS) in traditional convolutional neural networks. Compared with mainstream methods such as Faster R-CNN and YOLO, RT-DETR has achieved significant improvements in both detection speed and accuracy, making it more suitable for real-time application scenarios.

[0056] In the series of RT-DETR object detection models, the official provides 12 versions in total, including RT-DETR-R18, RT-DETR-R34, RT-DETR-R50, RT-DETR-R101 with ResNet as the backbone, and RT-DETR-L, RT-DETR-X with HGNet as the backbone. For the actual needs of substation fire detection, the present invention selects RT-DETR-R18 with a relatively balanced network depth and detection accuracy as the basic model. The network structure of this model mainly consists of four core parts: the backbone network, the efficient hybrid encoder, the uncertainty-minimal query selection mechanism, and the decoder with an auxiliary prediction head. The specific structure is as Figure 1 shown.

[0057] The structure of this model is as follows: The initialized input image is fed into the first ConvN 3*3, S2 module. The output of the first ConvN 3*3, S2 module is connected to the input of the first ConvN 3*3 module. The output of the first ConvN 3*3 module is connected to the input of the second ConvN 3*3 module. The output of the second ConvN 3*3 module is connected to the input of the third ConvN 3*3 module. The output of the third ConvN 3*3 module is connected to the input of the Maxpool2d module. The output of the Maxpool2d module is connected to the input of the first Basic_Block_CG module. The output of the first Basic_Block_CG module is connected to the input of the second Basic_Block_CG module. The output of the second Basic_Block_CG module is connected to the input of the third Basic_Block_CG module and the input of the first Conv 1*1 module. The output of the third Basic_Block_CG module is connected to the input of the fourth Basic_Block_CG module and the input of the second Conv 1*1 module. The output of the fourth Basic_Block_CG module is connected to the input of the first ConvN 1*1 module. The output of the first ConvN 1*1 module is connected to the input of the AIFI module. The output of the AIFI module is connected to the input of the second Conv 1*1 module and the input of the third Conv 1*1 module. The output of the third Conv 1*1 module is connected to the input of the first upsampling module. The output of the first upsampling module and the output of the second Conv 1*1 module are connected to the input of the first concatenation module. The output of the first concatenation module is connected to the input of the first RepC3 module. The output of the first RepC3 module is connected to the input of the fourth Conv 1*1 module. The output of the fourth Conv 1*1 module is connected to the input of the second concatenation module and the input of the second upsampling module. The output of the second upsampling module and the output of the first Conv 1*1 module are connected to the input of the third concatenation module. The output of the third concatenation module is connected to the input of the second RepC3 module. The output of the second RepC3 module is connected to the input of the second ConvN 3*3, S2 module. The output of the second ConvN 3*3, S2 module is connected to the input of the second concatenation module. The output of the second concatenation module is connected to the input of the third RepC3 module. The output of the third RepC3 module is connected to the input of the third ConvN 3*3, S2 module. The output of the third ConvN 3*3, S2 module is connected to the input of the fourth concatenation module. The output of the fourth concatenation module is connected to the input of the fourth RepC3 module. The outputs of the second RepC3 module, the third RepC3 module, and the fourth RepC3 module are connected to the input of the minimum query mechanism module. The output of the minimum query mechanism module is connected to the input of the decoder module with an auxiliary prediction head. The decoder module with an auxiliary prediction head outputs the total features;

[0058] Specifically, step S1 includes:

[0059] S1.1: Widely collect data on typical substation fire cases

[0060] Widely collect typical substation fire cases that have occurred at home and abroad in recent years through network platforms. These cases should include detailed fire accident reports, on-site photos, video materials, and investigation reports of relevant departments, etc. During the collection process, ensure the authenticity and reliability of the data, and avoid introducing incorrect or misleading information. At the same time, pay attention to the diversity of cases, including substation fires of different scales and types, to ensure the comprehensiveness and representativeness of the data set.

[0061] S1.2: Carefully screen and organize relevant fire images

[0062] After collecting a large amount of case data, carefully screen and organize this data. Especially for the photos and video materials at the fire scene, the present invention needs to select those images with high resolution and good clarity as the basis of the data set. These images should be able to clearly show the situation at the fire scene, including the shape of the flames, the spread of the smoke, the degree of damage to the equipment, etc. In addition, ensure that the images have sufficient diversity, covering fire scenes at different angles, different time periods, and different weather conditions to improve the accuracy of subsequent processing and analysis. During the organization process, uniformly name and classify the images for subsequent management and use.

[0063] S1.3: Preprocess the data set to reduce the inaccuracy and randomness of some data set images. The methods of data preprocessing include rotation, folding, and grayscaling, and finally 3,611 images are obtained.

[0064] The image preprocessing method is as follows:

[0065] 1) Through the rotation operation, the present invention can simulate the presentation of images at different angles, which helps the model learn more robust feature representations. The rotation operation not only increases the diversity of the data but also alleviates the problem of image inaccuracy caused by different shooting angles to a certain extent. In actual operation, the present invention can rotate the image at a certain angle (such as 90 degrees, 180 degrees, or 270 degrees) to obtain multiple versions of image data.

[0066] 2) The folding operation is also an effective data augmentation method. It generates new image samples by folding the image horizontally or vertically. This operation can simulate the deformation of the image from different perspectives and further improve the generalization ability of the model. It should be noted that during the folding process, the coherence and integrity of the image content should be maintained to avoid introducing additional noise or distortion.

[0067] 3) Grayscale processing is also a commonly used method in data preprocessing. By converting color images into grayscale images, the present invention can remove the interference of color information on model training, enabling the model to focus more on the learning of key features such as shape and texture. Grayscale processing not only simplifies image information but also helps reduce the computational load and improve training efficiency.

[0068] After the above preprocessing operations, the present invention obtains a more abundant, diverse, and accurate dataset. According to statistics, the number of finally obtained images reaches 3611. This data scale not only meets the requirements of subsequent analysis or model training but also provides strong support for further improving the accuracy and robustness of the algorithm.

[0069] To further expand the dataset and enhance the generalization ability of the model, the present invention can consider adopting more preprocessing methods and data augmentation techniques. For example, operations such as image scaling, cropping, flipping can be performed, or perturbation factors such as noise and blur can be introduced to simulate complex situations in the real scene. Through these means, the present invention can further expand the dataset scale to twice or more of the original, thus providing a more solid data foundation for subsequent algorithm development and model training.

[0070] Specifically, step S2 includes:

[0071] S2.1: Perform annotation in the Anaconda environment. Use the labelimg image annotation tool to annotate the collected substation flame and smoke images. The recognition objects are divided into two categories: smoke and flame, where smoke is labeled as "0" and flame is labeled as "1"; obtain PascalVOC format data labels (XML labels), and then convert them into txt files suitable for training through data format conversion code written in python. The files include most of the data, including annotation information such as the x coordinate, y coordinate, width, and height of the center point of the bounding box.

[0072] S2.2: After completing the annotation work, divide the dataset to obtain the training set, validation set, and test set of the substation fire dataset, obtaining 2611 training set images, 600 validation set images, and 400 test set images;

[0073] Specifically, step S3 includes:

[0074] S3.1: Improve the RT-DETR model to enhance its accuracy and real-time performance to adapt to complex environmental changes.

[0075] (1) Optimize the Basic_Block module of the original backbone network, and introduce the Context Guided Block module to replace it with the Basic_Block_CG module.

[0076] In the backbone network of RT-DETR-R18, the BasicBlock acts as a basic building block. Its core goal is to efficiently extract and transmit information from the input feature map, providing strong support for the subsequent processing of complex tasks. To achieve this goal, two consecutive convolutional processing units are carefully designed inside the BasicBlock.

[0077] Each convolutional processing unit integrates a convolutional layer (Conv), batch normalization (BN, i.e., BatchNorm), and the ReLu activation function. This combination can significantly improve the model's feature extraction ability and training efficiency. Specifically, the convolutional layer is responsible for capturing local features in the input feature map; batch normalization accelerates model convergence and improves training stability by normalizing the features; while the ReLu activation function introduces non-linearity, enhancing the model's expressive ability.

[0078] In addition, the BasicBlock also adopts a residual connection design. The innovation lies in directly adding the input of the module to the output of the second convolutional processing unit, forming an "information direct path". This design enables the input information to cross multiple levels without loss and directly participate in the construction of subsequent features, thus effectively alleviating the problem of gradient disappearance and further improving the training stability and performance of the network.

[0079] In summary, through its carefully designed convolutional processing units and residual connection design inside, the BasicBlock achieves the goal of efficiently extracting and transmitting information from the input feature map, laying a solid foundation for the excellent performance of the RT-DETR-R18 model in subsequent complex tasks.

[0080] Although the BasicBlock effectively extracts and transmits local features in the deep learning model, its utilization of global context information is relatively limited. To make up for this deficiency, the present invention innovatively introduces the Context Guided Block (CGblock), replacing the second convolutional normalization layer in the BasicBlock to form a new Basic_Block_CG module. This improvement not only improves the computational efficiency but also enables the model to better capture long-range dependencies, enhancing its ability to use surrounding information to infer occluded or noise-disturbed parts, thus ensuring the stable substation fire detection performance of the model in various complex scenarios.

[0081] Specifically, the running process (feature extraction process) of Basic_Block_CG includes the following steps:

[0082] The context-guided module simulates the dependence of the human visual system on context information, fusing local features, surrounding context, and global context to improve the accuracy of feature extraction. This module uses a standard 3×3 convolutional layer to extract local features and dilated / convolutional layers to expand the receptive field for capturing rich surrounding context information. Both use channel convolution to reduce computational costs. The local and surrounding context features are integrated through connection layers, batch normalization, and parametric ReLU, balancing local fineness and global context to enhance object detection accuracy. Meanwhile, the global feature extractor refines the global context using global average pooling and a multi-layer perceptron, adjusting the joint features channel by channel as a weighted vector to strengthen important features and suppress irrelevant features. Additionally, global residual learning is implemented inside the module, further facilitating the learning of complex features through the direct connection of the input to the global feature extractor.

[0083] The benefits of replacing the original backbone network module with Basic_Block_CG are mainly reflected in the following aspects:

[0084] 1) Improve object detection accuracy

[0085] The CG block can significantly improve object detection accuracy by capturing local features, surrounding context, and global context information and fusing this information. This mechanism helps the model better understand complex scenes and the relationships between objects, thus performing better in detection tasks.

[0086] 2) Enhance the model's utilization of context information

[0087] The CG block can fully capture and utilize local features, surrounding context, and global context information in the image through local feature extractors, surrounding environment extractors, joint feature extractors, and global environment extractors. The fusion of this information helps the model better understand complex scenes and the relationships between objects, thus showing higher accuracy in object detection tasks. Especially when facing challenges such as occlusion, small-sized objects, or complex backgrounds, the introduction of the CG block can significantly improve the detection performance of the model, reducing false detections and missed detections.

[0088] 3) Enhance the model's generalization ability

[0089] Since the CG block can combine global context information, the model can show stronger robustness and generalization ability when dealing with various complex scenes and interference factors. This is of great significance for improving the detection accuracy of the model under different lighting conditions, painting interference, etc.

[0090] 4) Optimize feature representation

[0091] The CG block enhances the feature representation by combining local features and global context information, which helps the model better focus on important regions and suppress irrelevant backgrounds. The optimized feature representation can improve the performance of the model in computer vision tasks such as object detection and image segmentation.

[0092] In summary, introducing the CG block to replace the second convolutional normalization layer in the BasicBlock can significantly improve the object detection accuracy of the RT-DETR model, reduce the number of parameters and computational resource consumption, enhance the generalization ability of the model, and optimize the feature representation.

[0093] The Basic_Block_CG module is as Figure 2 shown, and the specific structure is as follows:

[0094] The output of the initialized input feature module is connected to the input of the ConvN3*3 module, the output of the ConvN3*3 module is connected to the input of the CG Block module, the output of the CG Block module is connected to the input of the concatenation module, the output of the concatenation module is connected to the input of the ReLU module, and the output of the ReLU module is the refined feature.

[0095] The specific structure of the CG Block module is as follows:

[0096] The output of the initialized input feature module is connected to the input and output of the 1*1Conv module, the output of the 1*1Conv module is connected to the input of the local feature extractor module and the input of the surrounding context extractor module, the output of the local feature extractor module is connected to the input of the 3*3Conv module, the output of the surrounding context extractor module is connected to the input of the 3*3DConv module, the outputs of the 3*3Conv module and the 3*3DConv module are connected to the input of the concatenation module, the output of the concatenation module is connected to the input of the BN / PReLU module, the output of the BN / PReLU module is connected to the input of the GAP module and the input of the output module, the output of the GAP module is connected to the input of the first FC module, the output of the first FC module is connected to the input of the second FC module, and the output of the second FC module is connected to the input of the output module.

[0097] S3.2: Set the Adam optimizer to the Lion optimizer. Compare the accuracy with the SGD optimizer and set up a reference experiment.

[0098] The Lion optimizer is a gradient-based optimization algorithm designed specifically for deep learning, especially suitable for complex image recognition tasks such as substation fire and smoke detection. In the substation environment, the timely and accurate detection of fire and smoke is crucial for preventing catastrophic consequences. The Lion optimizer optimizes the performance of the deep learning model in this task by dynamically adjusting the learning rate, introducing momentum acceleration, and deeply analyzing the gradient distribution of model parameters. It automatically adjusts the learning rate according to the gradient of each parameter, ensuring that the model can converge quickly and accurately, and can effectively identify fire and smoke features even in the complex substation background. At the same time, the momentum acceleration mechanism enhances the stability of parameter updates, avoids oscillations or vibrations during training, and improves the robustness of detection. The Lion optimizer of RT-DETR demonstrates high efficiency and superior performance in the training of deep learning models through its unique working principle and advantages.

[0099] In the context of substation fire and smoke detection, the Lion optimizer of RT-DETR has the following advantages:

[0100] (1) Adaptive learning rate: In the complex scenarios of substation fire and smoke detection, the adaptive learning rate feature of the Lion optimizer can ensure that the model quickly adapts to different image features and lighting conditions during training, improving the accuracy and efficiency of detection.

[0101] (2) Momentum acceleration: By introducing the concept of momentum, the Lion optimizer can accelerate gradient updates in the model training of substation fire and smoke detection, making the model converge more stably and reducing false detections caused by image noise or complex backgrounds.

[0102] (3) Balanced parameter distribution: In the substation environment, the characteristics of fire and smoke may vary due to equipment types, environmental factors, etc. The Lion optimizer enhances the model's generalization ability to different features and improves the accuracy and reliability of detection by achieving balanced parameter distribution.

[0103] In summary, in the context of substation fire and smoke detection, the Lion optimizer of RT-DETR provides an efficient and accurate deep learning model optimization solution for this critical task with its unique working principle and significant advantages.

[0104] S3.3: Use the Alpha-IoU loss function as the bounding box loss function of the model proposed in the present invention.

[0105] Use the Alpha-IoU bounding box loss function as the new bounding box loss function of RT-DETR.

[0106] The main idea of Alpha-IoU is: The principle of Alpha-IoU is to generalize the traditional IoU loss to the Power-IoU series and introduce an adjustable power parameter α to control the focus of the loss, in order to obtain better bbox regression accuracy.

[0107] Alpha-IoU is a new loss function designed to improve the performance of object detectors. Based on the traditional IoU (Intersection over Union) loss, Alpha-IoU is extended by introducing a power IoU term and a regularization term. This adjustable power parameter α allows the loss function to flexibly focus on different levels of IoU values, thus achieving fine control over the bbox (bounding box) regression accuracy. By adjusting the α parameter, the Alpha-IoU loss can achieve better detection performance in most cases, especially showing stronger robustness in small datasets and noisy scenarios. In addition, the use of the Alpha-IoU loss does not introduce additional parameters or increase the training / inference time, so it can be easily used to improve the effect of detectors.

[0108] The definition of Alpha-IoU is as follows:

[0109] Alpha-IoU (α-IoU) is an improvement to the traditional IoU (Intersection over Union) evaluation metric, which is used in object detection to evaluate the overlap between the predicted bounding box and the ground truth bounding box.

[0110] The traditional IoU metric evaluates the overlap by calculating the ratio of the intersection to the union of the predicted bounding box and the ground truth bounding box. Generally, if the IoU value is greater than a certain threshold (usually 0.5), the detection is considered correct. However, this method may not be very adaptable to objects of different sizes or shapes.

[0111] Alpha-IoU makes the IoU metric more flexible by introducing a hyperparameter α. This parameter controls the calculation method of IoU and can adjust the requirements for overlap according to different object characteristics. For example, in some cases, for slender or highly asymmetric objects, the standard IoU may not be applicable, while Alpha-IoU can handle this situation better and improve the accuracy of model evaluation.

[0112] The calculation formula for the standard IoU metric is:

[0113]

[0114] When the IoU value is greater than a certain set threshold (usually 0.5), the detection result is considered correct.

[0115] Alpha-IoU introduces the α parameter by adjusting the calculation formula of IoU. Specifically, Alpha-IoU not only focuses on the overlapping area between the predicted bounding box and the ground truth bounding box, but also takes into account the influence of the shape and size of the bounding box on the IoU calculation. In this way, by flexibly adjusting α, the calculation method of IoU can be changed, making it more adaptable to different types of objects (such as long and thin shapes, circles, etc.). When the α value is large, Alpha-IoU is more tolerant of the shape of the object and may allow some smaller overlapping areas as correct predictions; while when the α value is small, it will require a more strict overlap and reduce the tolerance for shape differences.

[0116] The calculation formula of Alpha-IoU can be expressed as:

[0117]

[0118] Here, α controls the contribution degree of the intersection and union in the calculation. By adjusting the value of α, IoU can be made more sensitive or tolerant to changes in the shape and size of the object.

[0119] The advantages of Alpha-IoU are mainly reflected in the following aspects:

[0120] (1) Traditional IoU may have a large error when facing irregular or deformed objects, especially long and thin, dense or slender objects. Alpha-IoU can flexibly adjust the overlapping calculation method by introducing the α parameter to adapt to objects of different shapes and sizes.

[0121] (2) The sensitivity of IoU to changes in the shape of the object can be controlled by adjusting the α value, which helps to improve the detection accuracy.

[0122] (3) When detecting small objects, the standard IoU may sometimes result in a low IoU value due to the relatively large error between the small object and the predicted bounding box. However, Alpha-IoU can more tolerantly evaluate the overlapping situation of small objects in some cases by adjusting the α value, thereby improving the detection accuracy of small objects.

[0123] Alpha-IoU makes the IoU metric more flexible by introducing the α parameter, which can adjust the evaluation criteria according to different object shapes, sizes, densities, etc., thereby improving the accuracy and robustness of object detection. It has more obvious advantages than traditional IoU in dealing with special scenarios and tasks.

[0124] S4: In a complex environment, the model needs to be able to handle different lighting conditions. The lighting differences between day and night can have a significant impact on object detection. Therefore, it is necessary to test through two typical scenarios of day and night to ensure that the model can operate stably under different lighting conditions. Especially at night, fires are usually difficult to detect due to insufficient lighting. Therefore, it is necessary to verify whether the model can still effectively identify fires or smoke under low-light or dark conditions.

[0125] Specifically, step S4 includes:

[0126] S4.1: Verify the accuracy and fire detection ability of the model

[0127] Ensure that the model can accurately detect fires, whether it is day or night. In particular, it is necessary to consider characteristics such as the color, shape, and brightness of the flame. The model should be able to identify small-scale fires or signs in the initial stage of a fire. Even if the fire is small or just starting to spread, the model should be able to promptly and accurately capture the fire source and give an alarm. And when a fire occurs, it is usually accompanied by smoke or flames. It is very important to verify whether the model can independently identify both of them. For example, before a substation fire spreads, there may only be the appearance of smoke in the initial stage. If the model can identify the smoke in advance and give an alarm at this time, it can effectively prevent the spread of the fire. Therefore, the model not only needs to accurately identify the flame but also be able to detect the situation with only smoke.

[0128] S4.2: Real-time object detection and response ability

[0129] In critical infrastructure such as substations, the spread of fires can be very rapid. To effectively prevent disasters, the model needs to have a fast response ability, be able to detect the fire source or smoke immediately before the fire spreads and give an early warning. This requires the model to not only have a high accuracy rate but also be able to respond within a short time to ensure that it can be processed in time before the fire spreads.

[0130] In addition to detecting the fire source or smoke, the model also needs to have the ability to track the target, that is, be able to continuously track the location and size of the fire source after the fire occurs and adjust the response measures according to the changes. This is crucial for the control and extinguishment of early fires.

[0131] When verifying this model, the real-time performance of the model is judged by verifying the detection speed of the model.

[0132] Through the ablation experiment shown in Table 1, it can be seen that when adding the ContextGuidedBlock module, Lion optimizer, and Alpha-IoU loss function at the same time, compared with the baseline model and other improvement methods, the model of the present invention is much higher than other detection models in terms of accuracy and speed, and can meet the requirements of real-time classification detection.

[0133] The following is the analysis table of the comparison results with the original model

[0134] Table 1 Comparison test results table

[0135]

[0136] The following is the analysis table of the comparison results of the loss function

[0137] Table 2 Comparison results table of the loss function

[0138]

[0139] From the experimental results in the above table, it can be seen that Alpha-IoU has the best performance and the highest accuracy. Compared with the CIoU loss function of the baseline model, Alpha-IoU has improved by 1.6% in accuracy, reaching 89.8%, the recall rate has increased by 23.5%, reaching 87%, the mAP value has been greatly improved, reaching 91.3%, and mAP0.5-0.95 has increased by 18%, reaching 60.2%.

[0140] The following is the analysis table of the comparison results of various different types of optimizers

[0141] Table 3 Comparison results table of the selection of different optimizers

[0142]

[0143] From the experimental results in the above table, it can be obtained that the Lion optimizer has a positive effect on the model. Compared with the Adam optimizer and the SGD optimizer, the Lion optimizer has improved in accuracy, recall rate, mAP0.5, and mAP0.5-0.95, indicating that this model is most suitable for the Lion optimizer.

[0144] Table 4 Comparison of the model detection accuracy under different α values

[0145]

[0146] From the experimental results in the above table, it can be obtained that when α = 3, the Alpha-IoU loss function performs the best. When α = 3, compared with the baseline model, the detection accuracy has increased to 89.8%, the recall rate has increased by 1.4%, reaching 87%, mAP0.5 has reached 91.3%, and mAP0.5-0.95 has reached 60.2%.

[0147] The platform used for the tests of this invention is Pycharm 11.0.15, the model framework is Pytorch 1.13.1, the CPU is 15vCPU Intel(R) Xeon(R) Platinum 8358P CPU @ 2.60GHz, and the GPU is RTX3090, 24GB.

[0148] When setting the hyperparameters, the number of training epochs is set to 200, and the batch size is set to 16.

[0149] In summary, based on the improved RT-DETR model, this invention achieves efficient and accurate identification of substation fires and smoke. Compared with traditional detection methods, this method demonstrates higher detection accuracy, real-time performance, and robustness in complex environments, can adapt to different lighting conditions such as day and night, and improves the reliability of fire warnings. By introducing the ContextGuidedBlock module, Lion optimizer, and Alpha-IoU (α = 3) loss function, this model enhances the ability to focus on key features, reduces the inaccuracy of model detection, and optimizes the regression accuracy of the target bounding box. In addition, this invention can effectively distinguish flames from smoke, accurately detect even in the initial stage of a fire, and improves the intelligent level of substation fire monitoring. This method not only improves the efficiency of fire monitoring but also provides important technical support for the safe operation of the power system, with significant engineering application value and social and economic benefits.

Claims

1. A real-time recognition method for substation fires and smoke based on improved RT-DETR, characterized in that, It includes the following steps: Step 1: Collect and organize several fire images to construct a fire dataset; Step 2: Annotate the target dataset, process the collected images one by one with labels, construct a substation image dataset containing fire and smoke scenes, and divide this dataset into a training set, a validation set, and a test set; Step 3: Obtain an improved RT-DETR model and use this model to identify the collected substation fire and smoke images; Step 4: Validate the model, including testing its robustness and detection accuracy under different lighting conditions during the day and at night; ensure that the model can still effectively identify fires and smoke even in low-light environments at night, guaranteeing all-weather detection capabilities.

2. The method according to claim 1, wherein Step 1 specifically includes the following steps: S1.1: Extensively collect data on typical substation fire cases; S1.2: Carefully screen and organize relevant fire images; S1.3: Preprocess the dataset to reduce the inaccuracy and randomness of some dataset images.

3. The method according to claim 1, characterized in that, Step 2 specifically includes the following steps: S2.1: Conduct annotation in the Anaconda environment, use the labelimg image annotation tool to annotate the collected substation flame and smoke images, and the recognition objects are divided into two categories: smoke and flame, where smoke is labeled as "0" and flame is labeled as "1"; obtain PascalVOC format data labels, and then convert them into txt files suitable for training through data format conversion code written in python. The files include most of the data, including annotation information such as the x coordinate, y coordinate, width, and height of the center point of the bounding box; S2.2: After completing the annotation work, divide the dataset to obtain the training set, validation set, and test set of the substation fire dataset.

4. The method according to claim 1, wherein In Step 3, the improved RT-DETR model is specifically: The initialized input image is fed into the first ConvN 3*3 module. The output of the first ConvN 3*3 module is connected to the input of the first ConvN 3*3 module. The output of the first ConvN 3*3 module is connected to the input of the second ConvN 3*3 module. The output of the second ConvN 3*3 module is connected to the input of the third ConvN 3*3 module. The output of the third ConvN 3*3 module is connected to the input of the Maxpool2d module. The output of the Maxpool2d module is connected to the input of the first Basic_Block_CG module. The output of the first Basic_Block_CG module is connected to the input of the second Basic_Block_CG module. The output of the second Basic_Block_CG module is connected to the input of the third Basic_Block_CG module and the input of the first Conv 1*1 module. The output of the third Basic_Block_CG module is connected to the input of the fourth Basic_Block_CG module and the input of the second Conv 1*1 module. The output of the fourth Basic_Block_CG module is connected to the input of the first ConvN 1*1 module. The output of the first ConvN 1*1 module is connected to the input of the AIFI module; The output of the AIFI module is connected to the input of the second Conv 1*1 module and the input of the third Conv 1*1 module. The output of the third Conv1*1 module is connected to the input of the first upsampling module. The output of the first upsampling module and the output of the second Conv 1*1 module are connected to the input of the first concatenation module. The output of the first concatenation module is connected to the input of the first RepC3 module. The output of the first RepC3 module is connected to the input of the fourth Conv 1*1 module. The output of the fourth Conv 1*1 module is connected to the input of the second concatenation module and the input of the second upsampling module. The output of the second upsampling module and the output of the first Conv 1*1 module are connected to the input of the third concatenation module. The output of the third concatenation module is connected to the input of the second RepC3 module. The output of the second RepC3 module is connected to the input of the second ConvN 3*3 module. The output of the second ConvN 3*3 module is connected to the input of the second concatenation module. The output of the second concatenation module is connected to the input of the third RepC3 module. The output of the third RepC3 module is connected to the input of the third ConvN 3*3 module. The output of the third ConvN 3*3 module is connected to the input of the fourth concatenation module. The output of the fourth concatenation module is connected to the input of the fourth RepC3 module; The outputs of the second RepC3 module, the third RepC3 module, and the fourth RepC3 module are connected to the input of the minimum query mechanism module. The output of the minimum query mechanism module is connected to the input of the decoder module equipped with an auxiliary prediction head. The decoder module equipped with an auxiliary prediction head outputs the total features.

5. The method according to claim 4, wherein The structure of the Basic_Block_CG module is specifically as follows: The initialized input features of the Basic_Block_CG module are input to the ConvN3*3 module. The output of the ConvN3*3 module is connected to the input of the CG Block module. The output of the CG Block module is connected to the input of the concatenation module. The output of the concatenation module is connected to the input of the ReLU module. The output of the ReLU module is the refined features.

6. The method according to claim 4, characterized in that, Specifically, the CG Block module is as follows: The initialized input features of the CG Block module are input to the input of the 1*1Conv module and the input of the output module. The output of the 1*1Conv module is connected to the input of the local feature extractor module and the input of the surrounding context extractor module. The output of the local feature extractor module is connected to the input of the 3*3Conv module. The output of the surrounding context extractor module is connected to the input of the 3*3DConv module. The output of the 3*3Conv module and the output of the 3*3DConv module are connected to the input of the concatenation module. The output of the concatenation module is connected to the input of the BN / PReLU module. The output of the BN / PReLU module is connected to the input of the GAP module and the input of the output module. The output of the GAP module is connected to the input of the first FC module. The output of the first FC module is connected to the input of the second FC module. The output of the second FC module is connected to the input of the output module.

7. The method according to claim 1 or 4 or 5 or 6, characterized in that, The Alpha-IoU loss function is used as the bounding box loss function of the improved RT-DETR model.

8. The method according to claim 1, wherein In step 4, tests are conducted through two typical scenarios, day and night, to ensure that the model can operate stably under different lighting conditions. At night, it is verified whether the model can effectively identify fires or smoke under low-light or dark conditions.

9. The method according to claim 8, wherein Specifically, it includes verifying the accuracy and fire detection ability of the model, as well as the real-time object detection and response ability.

10. The method according to claim 1, wherein In step 3, it includes the evaluation of the model performance, using precision, recall, mean average precision (mAP), and frames per second (FPS) as the main metrics.