Chimney detection method based on AI technology

Through the YOLO-RSOD object detection model, combined with the explicit visual center and global attention mechanism, the problem of low chimney detection accuracy in remote sensing images is solved, and high-precision and fast-responsive chimney detection is achieved.

CN120298710APending Publication Date: 2025-07-11INST OF ATMOSPHERIC PHYSICS CHINESE ACADEMY SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410543694.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing remote sensing image detection technology has the problem of low detection accuracy in chimney detection tasks, especially in remote sensing images with complex backgrounds, small targets and many similar objects, and the real-time performance is insufficient.

Method used

The YOLO-RSOD object detection model is adopted, and the combination of backbone network, neck network, feature pyramid network and head network is combined with an explicit visual center and global attention mechanism to perform feature extraction and fusion, and the decoupled object detection head is used to perform object detection to improve detection accuracy and speed.

Benefits of technology

It improves the detection accuracy and speed of remote sensing chimney images, can better handle complex backgrounds and small targets, reduces false detection and missed detection, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298710A_ABST
    Figure CN120298710A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides an AI technology-based chimney detection method, which comprises the steps of S1, collecting a remote sensing chimney image data set, dividing the remote sensing chimney image data set into a training set and a verification set, and performing data enhancement; s2, inputting the data set into a backbone network for feature extraction, and transmitting the data set to a neck network for feature information extraction; s3, transmitting feature information obtained by the neck network to a feature pyramid, performing up-down sampling to perform feature fusion, and enhancing features through an explicit visual center and a global attention mechanism to obtain a corresponding enhanced feature map; and S4, respectively inputting the enhanced feature maps into the head network to obtain a detection result of the remote sensing chimney image. A pyramid network and a decoupling detection structure are adopted to solve the common problems of complex and changeable backgrounds, large target scale change, inconsistent imaging quality, interference of a large number of similar objects and the like in remote sensing image detection, and the recognition capability of tiny targets is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a chimney detection method based on AI technology. Background Art

[0002] Industrial chimney emissions are one of the main sources of urban air pollution, and there is often an inverse relationship between the urban environmental quality and the number of chimneys. Therefore, chimney location detection is crucial for urban environmental monitoring and governance. Although target detection technology has made continuous progress in the past few years and achieved good detection results in some simple scenarios, in the chimney detection task, there are still problems such as complex remote sensing image backgrounds, small targets, and a large number of similar objects, which lead to a decrease in detection accuracy.

[0003] In the prior art, the most representative deep learning algorithms in the field of target detection include RCNN, Fast RCNN, Faster RCNN, and YOLO, etc. Among these algorithms, RCNN and its derivative algorithms belong to two-stage convolutional neural networks, which find possible target positions in the image through region proposal technology and then use the features extracted from the feature layer for target classification. The advantage of this type of detector is high accuracy, but the real-time performance is low, making it difficult to meet the requirements of rapid detection. On the other hand, YOLO is an end-to-end convolutional neural network based on regression problems, and its real-time performance has been significantly improved, but its detection accuracy is not as good as that of two-step detectors such as Faster RCNN. Although these target detection models perform well in their respective application fields, in the detection of remote sensing chimney images, due to problems such as complex backgrounds, small targets, and image quality effects, the detection accuracy is generally not ideal.

[0004] The difficulties in remote sensing image detection are mainly reflected in the following aspects: background complexity and variability, multi-scale nature of targets, image resolution and quality issues. These problems may become key factors restricting remote sensing image detection, usually resulting in the detection model being prone to false detection or missed detection when processing remote sensing images, and the existing remote sensing image detection technologies cannot ensure both detection accuracy and recognition accuracy at the same time. For example, the complex and variable background and the diversity of target sizes make it difficult for traditional detection methods to accurately locate, while image resolution and quality issues may reduce the detection ability of the model. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to propose a chimney detection method based on AI technology, which can improve the detection effect of remote sensing chimney images, increase the MAP value while maintaining the detection speed, so as to solve the problem of poor detection effect of remote sensing chimney images proposed in the above background art.

[0006] To achieve the above purpose, the present invention is realized through the following technologies:

[0007] The present invention provides a chimney detection method based on AI technology, which is characterized by including:

[0008] Step S1, collect a remote sensing chimney image dataset, divide it into a training set and a validation set, and perform data augmentation;

[0009] Step S2, input the dataset into the backbone network for feature extraction, and transmit it to the neck network to extract feature information, where the feature information includes spatial information and channel information;

[0010] Step S3, transmit the feature information obtained by the neck network to the feature pyramid, perform upsampling and downsampling for feature fusion, and strengthen the features through an explicit visual center and a global attention mechanism to obtain a corresponding enhanced feature map;

[0011] Step S4, input the enhanced feature map into the head network respectively, use four decoupled object detection heads included in the head network to perform object detection to obtain corresponding predicted feature maps, and output the final preselected boxes through non-maximum suppression to obtain the detection result of the remote sensing chimney image.

[0012] Further, in step S1, collect a remote sensing chimney image dataset, divide it into a training set and a validation set, and perform data augmentation. Specifically, obtain the remote sensing images of the chimney through satellite images to form a dataset, crop each image and label the position and size of the chimney to generate corresponding picture labels; divide the collected and labeled dataset into a training set and a validation set according to a ratio of 8:2.

[0013] Further, in step S2, input the dataset into the backbone network for feature extraction. Specifically, the backbone network receives the input of the cropped and labeled remote sensing image and extracts features; perform convolution and feature enhancement on the remote sensing image through a multi-layer convolution structure and an enhancement module, and the enhancement module is a combined module of ELAN and MP; extract feature information from the corresponding levels of the backbone network to generate feature maps of different sizes.

[0014] Further, in step S2, transmit it to the neck network to extract feature information, where the feature information includes spatial information and channel information. Specifically, input the feature maps of different sizes into the neck network; the neck network performs convolution operations on the input feature maps to extract the spatial information and channel information therein.

[0015] Further, in step S3, the feature information obtained by the neck network is transmitted to the feature pyramid, and upsampling and downsampling are performed for feature fusion. The features are enhanced through an explicit visual center and a global attention mechanism to obtain corresponding enhanced feature maps. Specifically, feature maps of different sizes generated by the neck network are input into the feature pyramid network; the feature pyramid network performs upsampling on each size of feature map. Among them, the explicit visual center and the Stem Block perform feature enhancement and smoothing on the top-level feature map; the explicit visual center includes a lightweight MLP module and an LVC module; during the downsampling process, for the transmitted feature maps, the global attention mechanism is used to focus on key regions, and the MP module is used for feature fusion.

[0016] Further, the Stem Block performs feature enhancement and smoothing on the top-level feature map. Specifically, the top-level feature map undergoes a 7×7 convolutional layer operation, and the output after convolution passes through a batch normalization layer and then through an activation function layer to enhance the non-linear processing ability.

[0017] Further, the feature pyramid network performs upsampling on each size of feature map. Specifically, the lightweight MLP module performs group normalization and depth convolution processing on the output feature Xsb of the Stem Block to enhance the feature representation and uses a residual connection; the LVC module encodes the feature Xsb using a convolution combination of 1×1, 3×3, and 1×1, and enhances the feature through a CBR block to obtain the corresponding relationship between the corresponding pixel points and position information; the output feature maps of the MLP module and the LVC module are concatenated along the channel dimension to obtain the final output of the explicit visual center.

[0018] Further, during the downsampling process, for the transmitted feature maps, the global attention mechanism is used to focus on key regions, and the MP module is used for feature fusion. Specifically, the channel attention mechanism processes the incoming feature map F1 and concatenates it with the original feature map to form an intermediate state F2; the intermediate state F2 is processed through the spatial attention mechanism, and the enhanced spatial information is concatenated to obtain the final feature output F3, enhancing the ability to recognize local spatial details.

[0019] Further, in step S4, the enhanced feature maps are respectively input into the head network, and four decoupled object detection heads included in the head network are used for object detection to obtain corresponding predicted feature maps. Specifically, four decoupled object detection heads are set and applied in the head network, and the decoupled object detection heads correspond to different sizes of enhanced feature maps; the four decoupled object detection heads perform multi-size prediction on the enhanced feature maps at different levels, generate predicted feature maps and output prediction information including the offsets of the center horizontal and vertical coordinates, width, height, bounding box confidence, and class confidence.

[0020] Further, in step S4, the final preselected bounding boxes are output through non-maximum suppression to obtain the detection result of the remote sensing chimney image. Specifically, the bounding boxes with confidence levels lower than the threshold are filtered, and the selection of the bounding boxes is optimized by calculating the intersection over union (IoU) and adjusting the confidence levels of the bounding boxes. The optimized bounding boxes are sorted by confidence level and traversed, the confidence levels of the overlapping bounding boxes are reduced, the coordinates of the bounding boxes are adjusted back to the original image size, and the final detection result is output.

[0021] Compared with the prior art, the present invention has at least one of the following technical effects:

[0022] The present invention adopts a pyramid network and a decoupled detection structure to solve the problems commonly encountered in remote sensing image detection, such as complex and variable backgrounds, large variations in target scales, inconsistent imaging qualities, and interference from a large number of similar objects. This method can better extract and fuse multi-scale features, improve the recognition ability of small targets, and maintain high accuracy and fast response during the image processing and target positioning processes. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments:

[0024] Figure 1 is the flowchart of the steps of the chimney detection method based on AI technology of the present invention;

[0025] Figure 2 is the detailed flowchart of the steps of the chimney detection method based on AI technology of the present invention;

[0026] Figure 3 is the schematic diagram of the small target detection layer of the chimney detection method based on AI technology of the present invention

[0027] Figure 4 is the flowchart of the EVC display visual center of the chimney detection method based on AI technology of the present invention;

[0028] Figure 5 is the structural diagram of the GAM global attention mechanism of the chimney detection method based on AI technology of the present invention;

[0029] Figure 6 is the structural diagram of the decoupled detection head of the chimney detection method based on AI technology of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will describe in detail a chimney detection method based on YOLO-RSOD provided by the present invention with reference to the drawings in the embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and provides a detailed implementation method and specific operation process.

[0031] Example 1

[0032] Object detection in remote sensing images has important applications in fields such as environmental monitoring and resource exploration and analysis. However, current general object detection algorithms face some challenges when processing industrial chimney remote sensing images, such as complex and variable backgrounds, large variations in target sizes, uneven imaging quality, and false detections and missed detections caused by a large number of similar objects. These problems result in generally low detection accuracy. For example, ① remote sensing images usually contain a large number of natural and man-made elements, and the background is extremely complex. Industrial areas, buildings, trees, etc. may all affect the visibility of the chimney, making it difficult to distinguish the target from the background. ② The relative size of the chimney in the remote sensing image is small, which means that in high-resolution images, the chimney may be only a few pixels in size, making it difficult for traditional object detection algorithms to accurately identify and locate. ③ There are often a large number of structures similar to chimneys in industrial areas, such as exhaust pipes, columns, etc. These similar objects are likely to cause misjudgments by the detection algorithm and reduce the detection accuracy. ④ The target scales in remote sensing images vary greatly. The chimney may be mixed with other targets and even partially occluded. This multi-scale feature increases the difficulty of detection, especially for algorithms that cannot flexibly adjust the receptive field. ⑥ The quality and resolution of remote sensing images may not be consistent. This will affect the clarity of the target, making it difficult for the detection algorithm to extract effective features and further reducing the detection accuracy.

[0033] To address these problems, the present invention proposes an efficient and low-complexity anchor-free remote sensing image detection framework based on YOLOv7 - YOLO-RSOD. Through a series of improvements such as introducing an additional remote sensing image detection head, small object data prediction, decoupled detection head structure, explicit visual center of the centralized feature pyramid network, and global attention mechanism, this framework has made significant progress in dealing with complex backgrounds, detecting small targets, and overcoming interference from similar objects, effectively overcoming the challenges in the prior art, improving the accuracy and efficiency of object detection in remote sensing images, and further data augmentation and optimized model structure. The specific implementation process is as follows:

[0034] As Figure 1 、 Figure 2 、 Figure 3 and Figure 6 shown, the present invention provides a chimney detection method based on AI technology, using the YOLO-RSOD object detection model, which includes a backbone network, a neck network, a feature pyramid network, and a head network, and includes the following steps:

[0035] Step S1, manually collect the remote sensing chimney image dataset, divide the collected dataset into a training set and a validation set, and perform data augmentation. Specifically, the operation steps are as follows: Search for industrial chimneys through satellite images and crop the image size; repeat the operation to collect a sufficient amount of dataset; use the Label Img tool for annotation to generate corresponding picture labels; divide the dataset and the corresponding picture labels into a training set and a validation set according to the ratio of 8:2.

[0036] Step S2, after the dataset is augmented, input it into the backbone network of YOLO-RSOD for feature extraction. The backbone network is an eleven-layer network structure composed of a CBS convolutional network, an ELAN gradient network, and an MP downsampling convolutional network. Extract the feature information of the fifth, seventh, ninth, and eleventh layers of the backbone network to obtain four different-sized feature maps. Specifically, the operation steps are as follows:

[0037] (1) The backbone network receives the input of the cropped and annotated remote sensing image and performs feature extraction.

[0038] (2) The backbone network adopted by YOLO-RSOD performs convolution, pooling, and normalization operations on the input image. That is, first, through four CBS modules, perform convolution and feature enhancement on the remote sensing image to enhance the feature expression ability; sequentially pass the enhanced remote sensing image through three groups of ELAN and MP combined modules and one ELAN gradient module to reduce the loss of feature information of the remote sensing image, increase the receptive field, and enable the network to learn more feature information.

[0039] (3) In step S202, extract the feature information from the fifth, seventh, ninth, and eleventh layers of the backbone network respectively to obtain four different-sized feature maps.

[0040] (4) Input the four different-sized feature maps in step S203 into the neck network. The neck network performs convolution operations on the input feature maps to enhance their expression ability for the image, and extracts the spatial information and channel information therein.

[0041] Step S3, transmit the feature information obtained by the neck network, that is, the spatial information and channel information of the feature map, to the feature pyramid, perform upsampling and downsampling for feature fusion, that is, perform feature fusion in two ways: top-down and bottom-up, respectively obtain four different-sized enhanced feature maps, introduce an explicit visual center (EVC) in the upsampling part of the feature pyramid network, and at the same time fuse the GAM global attention mechanism in the ELAN-H gradient network in the downsampling part of the feature pyramid network. The specific operations are as follows:

[0042] (1) Input the feature information of the four different-sized feature maps generated by the neck network into the feature pyramid network.

[0043] (2) The Feature Pyramid Network upsamples the feature maps of each size. Among them, the display visual center enhances the features of the top-layer (the eleventh layer of the backbone network) feature map, and at the same time, the Stem Block smooths the top-layer feature map through 7x7 convolution, batch normalization, and activation function; among them, the display visual center mainly consists of two blocks connected in parallel, namely the lightweight MLP module and the LVC module; the lightweight MLP module captures the global information of the top-layer feature map, and the LVC module aggregates the local region feature information;

[0044] The explicit visual center (EVC) structure of the core block in CFP is as Figure 4 shown. Between the top-layer feature Xin and the EVC, there is a Stem Block for feature smoothing. The Stem Block consists of a 7×7 convolution with an output channel size of 256, followed by a batch normalization layer and an activation function layer. The above process can be represented by Xsb (full name) and the formula. Xsb = σ(BN(Conv7×7(Xin))). As Figure 4 、 6 shown, the feature enhancement of the top-layer (the eleventh layer of the backbone network) feature map by the display visual center is specifically

[0045] (1) The lightweight MLP mainly consists of two residual modules: a depth convolution-based module and a channel MLP-based block, where the input of the MLP-based module is the output of the depth convolution-based module. Specifically, for the depth convolution-based module, the feature Xsb output from the StemBlock module is first processed by group normalization, then fed to the depth convolution layer, and then a residual connection is made.

[0046] (2) For the channel MLP-based module, the features output from the depth convolution-based module Xml p1 are first fed to group normalization, and then the channel MLP is implemented on these features. Then, channel scaling, DropPath, and Xmlp1 residual connection are performed in sequence.

[0047] (3) In the LVC module, the features of Xsb are first encoded by a combination of a group of convolutional layers (which consists of 1×1 convolution, 3×3 convolution, and 1×1 convolution). Then, the encoded features are processed by the CBR block, which consists of a 3×3 convolution with a BN layer and a ReLU activation function. Then, the encoded features are input into the CodeBlock, and the Code Block is used to obtain the correspondence between pixel points and position information.

[0048] (4) After obtaining the output of the Code Block, further feed Xlvc1 into the fully connected layer and the 1×1 convolutional layer to predict the features highlighting the key classes. Then multiply Xsb and the scale factor coefficient δ(·) channel by channel. Finally, perform channel-wise addition between its output and Xsb;

[0049] (5) Concatenate the resulting feature maps of the mlp module and the lvc module along the channel dimension as the final output Xout of the EVC, enabling the feature pyramid network to introduce centralized information during the upsampling process, thereby achieving efficient feature fusion.

[0050] Step S33, during the downsampling process for the transmitted feature map, focus on the key region through the global attention mechanism and use the MP module for feature fusion, as Figure 5 shown, the specific operations are as follows:

[0051] (1) Channel attention processing: For the transmitted feature map F1, use the channel attention mechanism Mc of the global attention mechanism for processing, and concatenate the processed result with the feature map F1 to obtain the intermediate state F2;

[0052] (2) Spatial attention processing: Input the intermediate state F2 into the spatial attention mechanism of the global attention mechanism, and then concatenate the result with the intermediate state F2 to obtain the final output F3.

[0053] Step S4, input the enhanced feature maps into the head network respectively, use the four decoupled object detection heads included in the head network to perform object detection to obtain the corresponding predicted feature maps, and output the final preselected boxes through non-maximum suppression to obtain the detection results of the remote sensing chimney images. The specific operations are as follows:

[0054] In the head network, set four decoupled object detection heads so that each decoupled object detection head detects the enhanced feature maps of different sizes; input the enhanced feature maps into the corresponding decoupled object detection heads for object detection; the four decoupled object detection heads will perform multi-size prediction on the enhanced feature maps of different levels to obtain the corresponding predicted feature maps; the predicted feature maps output the prediction information of the bounding boxes, and the prediction information includes the offsets of the center horizontal and vertical coordinates, width, height, the confidence of the bounding box, and the confidence of the class. Filter the bounding boxes in the predicted feature maps whose bounding box confidence is lower than the set threshold; use the Soft-NMS algorithm for bounding box management, and the formula is: Among them, Si' is the adjusted bounding box confidence, IoU(bi,M) is the intersection over union of the current bounding box bi and the bounding box M with the highest score, Si represents the initial value of the current bounding box confidence, and D is the set of target boxes; σ is a penalty factor with a value in (0,1); the adjusted bounding boxes are sorted in descending order of bounding box confidence, and starting from the bounding box with the highest score, the overlap degree between the current bounding box and the traversed bounding boxes is calculated, and the bounding box confidence of the current bounding box is reduced according to the overlap degree and a preset reduction rate; until all bounding boxes are traversed; the position information of the bounding boxes is restored to the original image size, and the detection results are output.

[0055] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.

Claims

1. A chimney detection method based on AI technology, characterized in that, Including: Step S1: Collect a remote sensing chimney image dataset, divide it into a training set and a validation set, and perform data augmentation. Step S2: Input the dataset into a backbone network for feature extraction, and transfer it to a neck network to extract feature information, where the feature information includes spatial information and channel information. Step S3: Transfer the feature information obtained by the neck network to a feature pyramid, perform upsampling and downsampling for feature fusion, and strengthen the features through an explicit visual center and a global attention mechanism to obtain corresponding enhanced feature maps. Step S4: Input the enhanced feature maps into a head network respectively, use four decoupled object detection heads included in the head network to perform object detection to obtain corresponding predicted feature maps, and output the final preselected boxes through non-maximum suppression to obtain the detection results of the remote sensing chimney images.

2. The chimney detection method according to claim 1, characterized in that, In step S1, collect a remote sensing chimney image dataset, divide it into a training set and a validation set, and perform data augmentation. Specifically, Obtain remote sensing images of chimneys through satellite imagery to form the dataset, crop each image and label the position and size of the chimney to generate corresponding image labels; divide the collected and labeled dataset into the training set and the validation set according to a ratio of 8:

2.

3. The chimney detection method according to claim 2, characterized in that, In step S2, input the dataset into a backbone network for feature extraction. Specifically, The backbone network receives the input of the cropped and labeled remote sensing images and extracts features. Perform convolution and feature enhancement on the remote sensing images through a multi-layer convolution structure and an enhancement module, where the enhancement module is a combined module of ELAN and MP, and extract the feature information from corresponding levels of the backbone network to generate feature maps of different sizes.

4. The chimney detection method according to claim 3, characterized in that In step S2, transfer it to a neck network to extract feature information, where the feature information includes spatial information and channel information. Specifically, Input the feature maps of different sizes into the neck network; the neck network performs a convolution operation on the input feature maps to extract the spatial information and the channel information therein.

5. The chimney detection method according to claim 4, characterized in that, In step S3, transfer the feature information obtained by the neck network to a feature pyramid, perform upsampling and downsampling for feature fusion, and strengthen the features through an explicit visual center and a global attention mechanism to obtain corresponding enhanced feature maps. Specifically, Input the feature maps of different sizes generated by the neck network into the feature pyramid network. The feature pyramid network performs upsampling on each size of the feature maps. Among them, the explicit visual center and the Stem Block perform feature strengthening and smoothing on the top-level feature maps; the explicit visual center includes a lightweight MLP module and an LVC module. During the downsampling process, for the transferred feature maps, focus on the key regions through the global attention mechanism and use the MP module for feature fusion.

6. The chimney detection method according to claim 5, characterized in that, The Stem Block performs feature strengthening and smoothing on the top-level feature maps. Specifically, The top-level feature maps go through a 7×7 convolutional layer operation, the output after convolution goes through a batch normalization layer, and then goes through an activation function layer to enhance the non-linear processing ability.

7. The chimney detection method according to claim 6, characterized in that The feature pyramid network upsamples the feature maps of each size. Specifically, The lightweight MLP module performs group normalization and depth convolution processing on the output feature Xsb of the Stem Block to enhance the feature representation and through residual connection; The LVC module encodes the feature Xsb using a convolution combination of 1×1, 3×3, and 1×1, and enhances the feature through the CBR block to obtain the correspondence between the corresponding pixel points and the position information; The output feature maps of the MLP module and the LVC module are aggregated and concatenated along the channel dimension to obtain the final output of the display visual center.

8. The chimney detection method according to claim 6, characterized in that, During the downsampling process, for the transmitted feature maps, the global attention mechanism is used to focus on the key regions, and the MP module is used for feature fusion. Specifically, The channel attention mechanism processes the incoming feature map F1 and concatenates it with the original feature map to form an intermediate state F2; The intermediate state F2 is processed through the spatial attention mechanism, and the enhanced spatial information is concatenated to obtain the final feature output F3, enhancing the ability to recognize local spatial details.

9. The chimney detection method according to claim 8, wherein, In step S4, the enhanced feature maps are respectively input into the head network, and the target detection is performed using four decoupled object detection heads included in the head network to obtain the corresponding predicted feature maps. Specifically, Four decoupled object detection heads are set and applied in the head network, and the decoupled object detection heads correspond to the enhanced feature maps of different sizes; the four decoupled object detection heads perform multi-size prediction on the enhanced feature maps of different levels, generate the predicted feature maps and output the prediction information including the offset of the center horizontal and vertical coordinates, width, height, the confidence of the bounding box, and the confidence of the category.

10. The chimney detection method according to claim 9, characterized in that, In step S4, the final preselected boxes are output through non-maximum suppression to obtain the detection result of the remote sensing chimney image. Specifically, The bounding boxes with confidence lower than the threshold are filtered, and the selection of the bounding boxes is optimized by calculating the intersection over union IoU and adjusting the confidence of the bounding boxes; the optimized bounding boxes are sorted by confidence and traversed, the confidence of the overlapping bounding boxes is reduced, and the coordinates of the bounding boxes are adjusted back to the original image size and the final detection result is output.