CNN (Convolutional Neural Network) smoke and fire detection method and system based on double attention mechanisms

By applying the CNN pyrotechnic detection method based on the dual attention mechanism in the tower computer room, the problems of slow response and high false alarm rate of traditional fire monitoring methods are solved, and high accuracy and timely firework detection are achieved, ensuring effective identification and tracking of small targets.

CN120088740AActive Publication Date: 2025-06-03CHINA TOWER CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510586888.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-03
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional fire monitoring methods rely on smoke detectors or temperature sensors, are slow to respond and are susceptible to environmental noise and interference factors, resulting in false alarms and reliability problems.

Method used

The CNN pyrotechnic detection method based on the dual attention mechanism is adopted. By acquiring and preprocessing the pyrotechnic historical images, selecting characteristic data, and establishing a pyrotechnic detection model with the dual attention mechanism, combining environmental perception data to perform dual authentication to issue an alarm.

Benefits of technology

It improves the accuracy and timeliness of firework detection, reduces the false alarm rate, enhances the ability to capture the characteristics of small targets and the detailed characteristics, and ensures effective identification and tracking of dangerous scenarios caused by small targets such as cigarette butts in the tower computer room.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088740A_ABST
    Figure CN120088740A_ABST
Patent Text Reader

Abstract

The invention discloses a CNN smoke and fire detection method and system based on a double attention mechanism, and the method comprises the steps: obtaining a smoke and fire historical image, carrying out the preprocessing of the smoke and fire historical image, carrying out the feature extraction of the processed smoke and fire historical image, and obtaining the feature data of the smoke and fire historical image, and establishing a smoke and fire detection model based on a dual attention mechanism according to the feature data, inputting the to-be-detected image into the smoke and fire detection model to perform smoke and fire identification and output an identification result, performing dual authentication based on the identification result and the environment perception data, and if the authentication is passed, sending out alarm information. The method can effectively detect dangerous scenes caused by small targets such as cigarette ends and mars in an iron tower machine room, and through training and learning of long-distance and small target features, the model can effectively detect features of a smoke and fire target in a camera real-time video, so that identification and tracking of the small target object are achieved, and the identification efficiency of the iron tower machine room is improved. And the small target detection accuracy of the CNN is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a CNN fireworks detection method and system based on a dual attention mechanism. Background Art

[0002] In the context of the rapid development of informatization and intelligence today, as an important part of the communication network, the iron tower machine room bears key network equipment and data centers. The equipment in the machine room is intensive, and most of them are high-value assets. Therefore, ensuring its safe and stable operation is crucial. The occurrence of a fire accident will not only cause serious damage to the equipment, but may also lead to service interruption, thereby affecting the normal operation of the entire communication network, resulting in huge economic losses and social impacts. However, traditional fire monitoring methods mostly rely on smoke detectors or temperature sensors, which often react slowly in the early stage of a fire and cannot identify potential dangers in time. In addition, noise and other interference factors in the environment may cause false alarms, thus affecting the reliability and effectiveness of the equipment. Summary of the Invention

[0003] To solve the above problems, the present invention provides a CNN fireworks detection method and system based on a dual attention mechanism to solve the problems that traditional fire monitoring methods mostly rely on smoke detectors or temperature sensors, react slowly in the early stage of a fire, resulting in the inability to identify potential dangers in time, and noise and other interference factors in the environment may cause false alarms, thereby affecting the reliability and effectiveness of the equipment.

[0004] A CNN fireworks detection method based on a dual attention mechanism includes: Obtaining historical fireworks images and preprocessing the historical fireworks images; Extracting features from the processed historical fireworks images to obtain feature data of the historical fireworks images; Establishing a fireworks detection model based on a dual attention mechanism according to the feature data; Inputting the image to be detected into the fireworks detection model for fireworks recognition and outputting the recognition result; Performing dual authentication based on the recognition result and environmental perception data, and if the authentication is passed, sending an alarm message.

[0005] According to a specific embodiment of the present invention, obtaining historical fireworks images and preprocessing the historical fireworks images includes: Obtaining historical video stream data of fireworks and performing frame division processing and data annotation on the historical video stream to obtain historical fireworks images; Performing normalization processing on the historical fireworks images.

[0006] According to a specific embodiment of the present invention, feature extraction is performed on the processed historical firework images, and the feature data of the historical firework images obtained includes: The historical firework images are divided into multiple channels based on the image resolution to obtain multi-channel image data; Initial feature extraction is respectively performed on the multi-channel image data based on the convolutional neural network algorithm to obtain multi-channel feature data of the historical firework images; Deep feature extraction is respectively performed on the multi-channel feature data based on the dual attention mechanism to obtain deep feature data of the historical firework images.

[0007] According to a specific embodiment of the present invention, deep feature extraction is respectively performed on the multi-channel feature data based on the dual attention mechanism, and the deep feature data of the historical firework images obtained includes: Channel attention features and spatial attention features are respectively extracted from the multi-channel feature data based on the dual attention mechanism, where the channel attention features include channel saliency information, and the spatial attention features include spatial position information; Feature enhancement processing and feature splicing are performed based on the channel attention features and spatial attention features to obtain deep feature data of the historical firework images.

[0008] According to a specific embodiment of the present invention, a firework detection model based on the dual attention mechanism is established according to the feature data, including: A channel attention model and a spatial attention model are respectively established according to the feature data. The channel attention model and the spatial attention model adopt a parallel structure, where the channel attention model is used to extract the saliency information of the feature channels, and the spatial attention model is used to extract the position information of the target in space.

[0009] According to a specific embodiment of the present invention, the method further includes obtaining environmental perception data, and the environmental perception data includes CO concentration, temperature, and humidity data.

[0010] According to a specific embodiment of the present invention, dual authentication is performed based on the recognition result and the environmental perception data. If the authentication is passed, an alarm message is issued, including: Based on the recognition result, it is judged whether fireworks are detected. If detected, it is judged whether the environmental perception data reaches a preset threshold. If it reaches, an alarm message is issued.

[0011] According to a specific embodiment of the present invention, the recognition result includes the firework category, target position, and confidence score.

[0012] A CNN firework detection system based on the dual attention mechanism includes: A data acquisition and processing module, which is used to obtain historical firework images and preprocess the historical firework images; A feature extraction module, which is used to extract features from the processed historical fireworks images to obtain the feature data of the historical fireworks images; A model establishment module, which is used to establish a fireworks detection model based on a dual attention mechanism according to the feature data; An identification module, which inputs the image to be detected into the fireworks detection model for fireworks identification and outputs the identification result; An alarm module, which is used to perform dual authentication based on the identification result and the environmental perception data. If the authentication is passed, an alarm message is sent.

[0013] According to a specific embodiment of the present invention, the data acquisition and processing module further includes: A data acquisition module, which is used to acquire the historical video stream data of fireworks, perform frame division processing and data annotation on the historical video stream to obtain the historical fireworks images; A data processing module, which is used to perform normalization processing on the historical fireworks images.

[0014] According to a specific embodiment of the present invention, the feature extraction module further includes: A channel division module, which is used to perform multi-channel division on the historical fireworks images based on the image resolution to obtain multi-channel image data; An initial feature extraction module, which is used to perform initial feature extraction on the multi-channel image data respectively based on the convolutional neural network algorithm to obtain the multi-channel feature data of the historical fireworks images; A deep feature extraction module, which is used to perform deep feature extraction on the multi-channel feature data respectively based on the dual attention mechanism to obtain the deep feature data of the historical fireworks images.

[0015] According to a specific embodiment of the present invention, the deep feature extraction module further includes: A dual attention feature extraction module, which is used to extract channel attention features and spatial attention features from the multi-channel feature data respectively based on the dual attention mechanism, where the channel attention features include channel saliency information and the spatial attention features include spatial position information; A feature enhancement and splicing module, which is used to perform feature enhancement processing and feature splicing based on the channel attention features and the spatial attention features to obtain the deep feature data of the historical fireworks images.

[0016] According to a specific embodiment of the present invention, it further includes: An environmental perception data acquisition module, which is used to acquire CO concentration, temperature and humidity data.

[0017] According to a specific embodiment of the present invention, the alarm module further includes: A first judgment unit, which is used to judge whether fireworks are detected based on the identification result; A second judgment unit, configured to judge whether a preset threshold is reached based on the environmental perception data; An alarm unit, configured to send an alarm message according to the judgment results of the first judgment unit and the second judgment unit.

[0018] An electronic device, comprising: a processor and a memory, where a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned CNN smoke detection method based on a dual attention mechanism.

[0019] A computer-readable storage medium, in which a computer program is stored, and the computer program is loaded and executed by the processor to implement the above-mentioned CNN smoke detection method based on a dual attention mechanism.

[0020] Compared with the prior art, a CNN smoke detection method and system based on a dual attention mechanism provided by the present invention have the following advantages: 1. A CNN model with a small number of parameters and a small target detection model based on a dual attention mechanism designed by the present invention can effectively detect dangerous scenarios caused by small targets such as cigarette butts and sparks in the iron tower machine room. Through the training and learning of the features of long distances and small targets, the model can effectively detect the features of smoke and fire targets in the real-time video of the camera, so as to achieve the recognition and tracking of small target objects, improve the accuracy of CNN for small target detection, enhance the feature representation of small targets and improve the model's ability to capture detailed features. When detecting dangerous situations caused by small targets such as cigarette butts in the iron tower machine room, it can ensure that dangerous accidents are cut off from the source.

[0021] 2. The present invention simultaneously uses CO, temperature, and humidity sensors to comprehensively assist in monitoring the safety situation in the machine room, combines the sensor detection end and the visual algorithm end to jointly discover smoke and fire dangerous scenarios. Once an anomaly is detected, it immediately verifies with the data collected by the sensor end. If the data exceeds the preset threshold, an alarm is immediately given. This mechanism not only improves the response speed of smoke and fire detection but also reduces the possibility of false alarms.

[0022] 3. The present invention also provides optional remote operation and maintenance and management functions, which can detect the occurrence of dangerous abnormal accidents such as a fire in the machine room in real time, thereby providing effective guarantee for the safe operation of the iron tower machine room. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a flowchart of a CNN firework detection method based on a dual attention mechanism provided by an embodiment of the present invention.

[0025] Figure 2 It is a flowchart of a method for obtaining historical firework images and performing preprocessing provided by an embodiment of the present invention.

[0026] Figure 3 It is a flowchart of a method for feature extraction from historical firework images provided by an embodiment of the present invention.

[0027] Figure 4 It is a flowchart of a method for performing deep feature extraction on multi-channel feature data based on a dual attention mechanism provided by an embodiment of the present invention.

[0028] Figure 5 It is an overall framework diagram of a CNN firework detection model based on a dual attention mechanism provided by an embodiment of the present invention.

[0029] Figure 6 It is a schematic diagram of an enhancement operation for feature extraction by a dual attention model provided by an embodiment of the present invention.

[0030] Figure 7 It is a schematic diagram of the dual attention feature fusion process provided by an embodiment of the present invention.

[0031] Figure 8 It is a structural diagram of a CNN firework detection system based on a dual attention mechanism provided by an embodiment of the present invention.

[0032] Figure 9 It is a structural diagram of a data acquisition and processing module provided by an embodiment of the present invention.

[0033] Figure 10 It is a structural diagram of a feature extraction module provided by an embodiment of the present invention.

[0034] Figure 11 It is a structural diagram of a deep feature extraction module provided by an embodiment of the present invention.

[0035] Figure 12 It is a structural diagram of an alarm module provided by an embodiment of the present invention.

[0036] Figure 13 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention.

[0037] Reference numerals: 00 - Environmental perception data acquisition module; 01 - Data acquisition and processing module; 02 - Feature extraction module; 03 - Model establishment module; 04 - Recognition module; 05 - Alarm module; 011 - Data acquisition module; 012 - Data processing module; 021 - Channel division module; 022 - Initial feature extraction module; 023 - Deep feature extraction module; 0231 - Dual attention feature extraction module; 0232 - Feature enhancement and splicing module; 051 - First judgment unit; 052 - Second judgment unit; 053 - Alarm unit. Detailed implementation manner

[0038] In order to enable those skilled in the art to more clearly understand the concepts and ideas of the present invention, the present invention will be described in detail below in conjunction with specific embodiments. It should be understood that the embodiments given herein are only a part of all possible embodiments of the present invention. After reading the specification of this application, those skilled in the art are capable of making improvements, modifications, or substitutions to part or all of the following embodiments, and these improvements, modifications, or substitutions are also included within the scope of protection required by the present invention.

[0039] In this article, terms such as "advance notice", "entering the station" and other similar words do not imply any order, quantity, or importance, but are only used to distinguish different elements. In this article, terms such as "a", "one" and other similar words do not mean that there is only one thing, but mean that the relevant description only refers to one of the things, and the thing may have one or more. In this article, terms such as "include", "comprise" and other similar words are intended to represent a logical relationship, and should not be regarded as representing a spatial structure relationship. For example, "A includes B" is intended to mean that logically B belongs to A, rather than meaning that B is located inside A in space. Additionally, the meanings of terms such as "include", "comprise" and other similar words should be regarded as open-ended, rather than closed. For example, "A includes B" is intended to mean that B belongs to A, but B does not necessarily constitute all of A, and A may also include other elements such as C, D, E, etc.

[0040] In this article, terms such as "embodiment", "this embodiment", "one embodiment", "a single embodiment" do not mean that the relevant description only applies to a specific embodiment, but mean that these descriptions may also apply to one or more other embodiments. Those skilled in the art should understand that in this article, any description made for a certain embodiment can be substituted, combined, or otherwise combined with the relevant descriptions in one or more other embodiments, and the new embodiments generated by substitution, combination, or other means are easily conceivable by those skilled in the art and fall within the scope of protection of the present invention.

[0041] Example 1 Additional aspects and advantages of the embodiments of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the embodiments of the present invention. In combination with Figures 1-7 , an embodiment of the present invention provides a CNN fireworks detection method based on a dual attention mechanism, including: S1: Obtain historical fireworks images and preprocess the historical fireworks images.

[0042] S2: Extract features from the processed historical fireworks images to obtain feature data of the historical fireworks images.

[0043] S3: Establish a fireworks detection model based on a dual attention mechanism according to the feature data.

[0044] S4: Input the image to be detected into the fireworks detection model for fireworks recognition and output the recognition result.

[0045] S5: Perform dual authentication based on the recognition result and environmental perception data. If the authentication passes, an alarm message is issued.

[0046] Compared with traditional fire monitoring methods that mostly rely on smoke detectors or temperature sensors, these devices often react slowly in the initial stage of a fire and cannot identify potential dangers in time. In addition, noise and other interference factors in the environment may cause false alarms, thus affecting the reliability and effectiveness of the devices. The present invention uses a small target detection algorithm based on a lightweight dual attention mechanism to perform real-time inference and detection of fireworks at the camera end. By monitoring the computer room environment in real time, small targets such as smoke and flames can be accurately identified. At the same time, combined with sensor data, the environmental changes in the computer room can be analyzed more comprehensively, providing strong support for fire early warning. When the visual algorithm detects that the fireworks match the threshold of the sensor for ignition, an alarm is issued, enabling it to quickly issue an alarm in the initial stage of a fire. This can not only improve the accuracy and timeliness of fireworks detection, but also significantly reduce the false alarm rate and enhance the safety management level of the computer room, providing guarantee for the stable operation of communication infrastructure.

[0047] Specifically, step S1 of obtaining historical fireworks images and preprocessing the historical fireworks images includes: S11: Obtain historical video stream data of fireworks, perform frame division processing and data annotation on the historical video stream to obtain historical fireworks images.

[0048] S12: Perform normalization processing on the historical fireworks images.

[0049] In a specific embodiment of the present invention, a camera is used to obtain historical video stream data of fireworks and perform frame division on the historical video stream data to obtain historical images of fireworks. After obtaining the historical images of fireworks, data annotation is performed on the images, and the images with data annotation are normalized to facilitate subsequent training of the fireworks detection model.

[0050] Specifically, step S2 performs feature extraction on the processed historical images of fireworks, and the feature data of the historical images of fireworks obtained includes: S21: Perform multi-channel division on the historical images of fireworks based on the image resolution to obtain multi-channel image data.

[0051] S22: Perform initial feature extraction on the multi-channel image data respectively based on the convolutional neural network algorithm to obtain multi-channel feature data of the historical images of fireworks.

[0052] S23: Perform deep feature extraction on the multi-channel feature data respectively based on the dual attention mechanism to obtain deep feature data of the historical images of fireworks.

[0053] Further, step S23 performs deep feature extraction on the multi-channel feature data respectively based on the dual attention mechanism, and the deep feature data of the historical images of fireworks obtained includes: S231: Extract channel attention features and spatial attention features from the multi-channel feature data respectively based on the dual attention mechanism, where the channel attention features include channel saliency information, and the spatial attention features include spatial position information.

[0054] S232: Perform feature enhancement processing and feature splicing based on the channel attention features and spatial attention features to obtain deep feature data of the historical images of fireworks.

[0055] In a specific embodiment of the present invention, by performing multi-channel division on the historical images of fireworks according to different image resolution dimensions, historical images of fireworks with multiple channels are obtained, where each channel represents a feature dimension in the image, and different channels contain different image feature information. By analyzing and extracting features from the images of different channels separately, the image can be analyzed from multiple dimensions to more accurately identify these features. After performing multi-channel division on the image, the present invention separately performs feature extraction on the image data of each channel, such as Figure 7As shown in the figure, the present invention first performs initial feature extraction on the historical firework images of each channel through a convolutional neural network algorithm to obtain the firework feature data corresponding to each channel. The lightweight convolutional neural network is particularly suitable for real-time monitoring in resource-constrained environments due to its relatively simple model structure and high computational efficiency. By applying the lightweight CNN to firework detection, efficient image processing and analysis can be achieved, and potential dangers can be detected in a timely manner. Then, a dual attention mechanism is used to perform deep feature extraction on the firework feature data of each channel to capture the spatial feature information and channel feature information in the image. The dual attention mechanism includes two branches, namely the channel attention branch and the spatial attention branch. The channel attention branch is used to capture the correlation between the feature channels in the image, and the spatial attention branch is used to capture the correlation of features in the spatial position. Through the dual attention mechanism, the spatial feature information and channel feature information of the firework feature data of each channel can be extracted, and then the extracted spatial feature information and channel feature information are subjected to feature enhancement and feature splicing to obtain the deep feature data of the historical firework image. This deep feature data can enhance the feature representation of small targets and improve the model's ability to capture detailed features, thereby more accurately detecting the spatial position information and significant feature information of small targets in the iron tower machine room. Through the lightweight CNN and the dual attention mechanism, in the task of detecting small targets in the iron tower machine room, the spatial and channel information of the features in the image can be effectively captured, and the dual attention convolution with a lightweight design has a low computational cost and is suitable for application in resource-limited environments.

[0056] Specifically, step S3 of establishing a firework detection model based on a dual attention mechanism according to the feature data includes: A channel attention model and a spatial attention model are respectively established according to the feature data. The channel attention model and the spatial attention model adopt a parallel structure, where the channel attention model is used to extract the saliency information of the feature channels, and the spatial attention model is used to extract the position information of the target in space.

[0057] In a specific embodiment of the present invention, the firework detection model based on a dual attention mechanism includes an input module, a feature extraction module, and an output module, as Figure 5As shown, the input module of the model is used to input the collected historical images of fireworks and environmental perception data. The environmental perception data includes CO concentration, temperature, and humidity data collected by sensors. Based on the detection requirements of small targets of fireworks in the iron tower machine room, the present invention embodiment designs a multi-channel input module to independently process different channel data through a lightweight convolutional neural network. Each channel represents a feature dimension in the image, and the same convolution and attention operations are shared on the same channel. After the processed multi-channel features are spliced through feature concatenation, they are used as the input of the dual attention mechanism, so as to more accurately detect the spatial position information and significant feature information of small targets. The output module of the model is used to output the recognition results of the image to be detected, including the fireworks category, target position, and confidence score. On the other hand, the output module also outputs alarm information. The feature extraction module uses a convolutional layer and a dual attention mechanism. The convolutional layer adopts the traditional CNN convolutional neural network architecture. The dual attention module divides the data output by the convolutional layer into two parallel branches, namely the channel attention branch and the spatial attention branch. The channel attention branch focuses on the correlation between feature channels in the image. Each feature channel is regarded as a unit, and a channel correlation model is established through a self-attention network. The spatial attention branch focuses on the correlation of features in spatial positions. Similarly, a self-attention mechanism is used to model spatial information. This parallel attention representation realizes feature modeling from different perspectives of the input image. Since the same convolution and attention operations are shared on the same channel, for most targets in the image, similar feature spaces are shared from the channel and spatial perspectives, that is, they have consistent representations. This consistency can be maintained stably in channels and spaces. For noise or abnormal targets, their representation patterns are random and diverse, neither sharing strong correlations in channels and spaces with normal targets nor having consistency with other noise points. The fireworks detection model based on the dual attention mechanism established by the present invention can more accurately identify and distinguish normal targets from abnormal information, making the detection effect more accurate, especially suitable for small target detection in complex environments.

[0058] In a specific embodiment of the present invention, a loss function based on feature saliency is also designed during the feature modeling process. By calculating the maximum mean difference between the target feature and the background feature, the similarity between the two features is measured, thereby improving the detection accuracy of the model for small targets.

[0059] In a specific embodiment of the present invention, the establishment process of the dual attention model is as Figure 6 shown: The dual attention model consists of two parts, namely, the channel attention module and the spatial attention module, which adopt a parallel structure. Specifically, the channel attention module calculates the average value of each channel through global average pooling operation and passes the result to a fully connected layer to calculate the attention weight of each channel. The calculation process is as follows: (1) (2) where F is the feature map, and represent channel-based attention and space-based attention respectively, represents element-wise multiplication, and represent the output feature maps after channel attention and spatial attention respectively.

[0060] The channel attention module first performs max pooling and average pooling on the input feature map in the spatial dimension respectively, then passes through a multi-layer perceptron with shared weights respectively, then adds the outputs of the two element-wise, and then passes through the sigmoid activation function to obtain the channel attention weight. Multiplying this weight with the input feature map element-wise realizes the attention mechanism on the channel. Then the output of the channel attention can be expressed as: (3) where MLP represents the multi-layer perceptron, AvgPool represents global average pooling, MaxPool represents max pooling, is the output image after channel attention.

[0061] Performing max pooling and average pooling in the spatial dimension respectively actually means only retaining the channel dimension and setting other dimensions to 1 to obtain a one-dimensional vector, and then passing through a multi-layer perceptron with shared weights. The multi-layer perceptron consists of two convolutional operations. The ReLU activation function is used after the first convolution, and the sigmoid activation function is used after the second convolution to obtain the output weight. Then each channel is multiplied by its corresponding weight. The above process can also be expressed as: (4) where, and represent the first convolutional operation and the second convolutional operation respectively, and represent average pooling and max pooling respectively, represents the sigmoid activation function.

[0062] In the spatial attention module, first, perform channel-based max pooling and average pooling operations on the input feature map. Concatenate the two resulting feature maps based on channels, then perform a convolution operation and a sigmoid activation function on them to obtain an attention weight with 1 channel and the same feature map size as the input. Then, perform element-wise multiplication with the input feature map, thus completing the spatial attention calculation. This process can be expressed as: (5) Among them, represents a convolution operation with a convolution kernel size of 7 × 7, represents global average pooling, and MaxPool represents max pooling, M s (F) is the output image based on spatial attention.

[0063] Finally, concatenate the output feature maps that have passed through channel attention and spatial attention, so that they first pass through the channel attention mechanism and then through the spatial attention mechanism, respectively extracting the information of the image in the channel dimension and the spatial dimension, in the order of channel first and then space. The specific formula is as follows: (6) Among them, represents the feature output with attention enhancement in the channel dimension first and then the feature output with attention intensity in the spatial dimension.

[0064] For the concatenated channel-spatial attention module, the features are serially processed through the channel attention and spatial attention modules in sequence, layer by layer deeply mining the significant features in the image information. The channel attention module first performs weighted distribution on the input features to enhance the model's sensitivity to key channels; then further expands the spatial correlation of the features in the spatial attention module to ensure that the final representation has strong global and local perception capabilities.

[0065] Different from the former's concatenated attention mechanism with channel first and then space, the present invention connects the two in parallel, thus greatly reducing the computational amount of the model and being more suitable for the detection effect of small targets such as fireworks and cigarette butts in real-time images. Then, use the lightweight Switch activation function to normalize it to more conveniently and efficiently capture the weight information in the channel dimension and the spatial position dimension. Combine the input feature map with the channel weights. Concatenate and calculate the features extracted by the two attentions, and process the channels through a 1×1 convolution operation. The specific calculation formula is as follows: (7) Among them, represents the parallel operation method.

[0066] For the parallel channel attention module and spatial attention module, the channel attention module starts from a global perspective and emphasizes the importance of features in specific channels based on the weighted relationships of different channel features. The spatial attention module, on the other hand, starts from local spatial information and focuses on the correlations between different positions in the image. Based on the joint representation of the parallel channel attention module and spatial attention module as well as the serial channel-spatial attention module, more comprehensive dual-dimensional attention processing of multi-level features in visual tasks can be performed, thereby enhancing the model's expressive ability and generalization ability. The channel attention module and spatial attention module in the parallel attention structure extract features independently from the channel dimension and spatial dimension respectively, ensuring that the model can fully capture the detailed features from different dimensions. In addition, the method of sharing weights further improves the model's recognition and abstraction ability for multi-perspective features.

[0067] Since the output dimensions of the channel attention module and the spatial attention module in the parallel structure are different, in order to ensure the stability of subsequent calculations and make up for this difference, the present invention unifies the output dimensions of the parallel module by adding an upsampling operation. The upsampling processes of these two ensure that channel and spatial information can interact with each other at the same scale, thereby forming a richer joint feature representation.

[0068] Specifically, in step S4, the image to be detected is input into the fireworks detection model for fireworks recognition and the recognition result is output. The recognition result includes the fireworks category, target location, and confidence score. The output recognition result can be used as the basis for judging fireworks. When fireworks are recognized, the fireworks category, target location, and confidence score are uploaded to the cloud management module for remote operation and maintenance and real-time monitoring.

[0069] Specifically, step S5 performs dual authentication based on the recognition result and environmental perception data. If the authentication is passed, an alarm message is issued, including: Based on the recognition result, it is judged whether fireworks are detected. If fireworks are detected, it is then judged whether the environmental perception data reaches a preset threshold. If it reaches, an alarm message is issued.

[0070] In a specific embodiment of the present invention, the recognition result output by the model is combined with the environmental perception data detected by the sensor. When the model detects fireworks and the environmental perception data detected by the sensor reaches the preset threshold, the alarm is immediately triggered to issue an alarm message, and at the same time, the data slice is uploaded to the platform for storage for subsequent management and maintenance by personnel. By combining the recognition result with the environmental perception data detected by the sensor, once the model detects an anomaly, it immediately verifies with the data collected by the sensor end. If the data exceeds the preset threshold, an alarm is immediately issued. This mechanism not only improves the response speed of fireworks detection but also reduces the possibility of false alarms.

[0071] Embodiment 2 Based on the above method, an embodiment of the present invention further provides a CNN fireworks detection system based on a dual attention mechanism, as Figures 8-12 shown, including: An environmental perception data acquisition module 00, configured to acquire CO concentration, temperature, and humidity data.

[0072] A data acquisition and processing module 01, configured to acquire historical fireworks images and preprocess the historical fireworks images.

[0073] A feature extraction module 02, configured to extract features from the processed historical fireworks images to obtain feature data of the historical fireworks images.

[0074] A model establishment module 03, configured to establish a fireworks detection model based on a dual attention mechanism according to the feature data.

[0075] An identification module 04, configured to input the image to be detected into the fireworks detection model for fireworks identification and output the identification result.

[0076] An alarm module 05, configured to perform dual authentication based on the identification result and the environmental perception data. If the authentication is passed, an alarm message is sent.

[0077] A CNN fireworks detection system based on a dual attention mechanism provided by the present invention uses a small target detection algorithm based on a lightweight dual attention mechanism to perform real-time inference and detection of fireworks at the camera end. By real-time monitoring of the computer room environment, small targets such as smoke and flame can be accurately identified. At the same time, combined with sensor data, the environmental changes in the computer room can be analyzed more comprehensively, providing strong support for fire warning. When the visual algorithm detects that the fireworks match the threshold of the sensor for starting a fire, an alarm is issued, enabling it to quickly issue an alarm at the initial stage of a fire. It can not only improve the accuracy and timeliness of fireworks detection, but also significantly reduce the false alarm rate, improve the safety management level of the computer room, and provide guarantee for the stable operation of communication infrastructure.

[0078] Specifically, the data acquisition and processing module 01 further includes: A data acquisition module 011, configured to acquire historical video stream data of fireworks, perform frame division processing and data annotation on the historical video stream to obtain historical fireworks images, and is also configured to acquire environmental perception data, including CO concentration, temperature, and humidity data.

[0079] A data processing module 012, configured to perform normalization processing on the historical fireworks images.

[0080] In a specific embodiment of the present invention, first, the data acquisition module 011 is used to obtain the historical video stream data of fireworks and perform frame division on the historical video stream data to obtain historical fireworks images. Then, the data processing module 012 is used to perform data annotation on the obtained historical fireworks images and perform normalization processing on the images with data annotation, so as to facilitate the subsequent training of the fireworks detection model.

[0081] Specifically, the feature extraction module 02 further includes: The channel division module 021 is used to perform multi-channel division on the historical fireworks images based on the image resolution to obtain multi-channel image data; The initial feature extraction module 022 is used to perform initial feature extraction on the multi-channel image data respectively based on the convolutional neural network algorithm to obtain multi-channel feature data of the historical fireworks images; The deep feature extraction module 023 is used to perform deep feature extraction on the multi-channel feature data respectively based on the dual attention mechanism to obtain deep feature data of the historical fireworks images.

[0082] Furthermore, the deep feature extraction module 023 further includes: The dual attention feature extraction module 0231 is used to extract channel attention features and spatial attention features from the multi-channel feature data respectively based on the dual attention mechanism, where the channel attention features include channel saliency information and the spatial attention features include spatial position information; The feature enhancement and splicing module 0232 is used to perform feature enhancement processing and feature splicing based on the channel attention features and spatial attention features to obtain deep feature data of the historical fireworks images.

[0083] In a specific embodiment of the present invention, first, the channel division module 021 divides the historical fireworks images into multiple channels according to different image resolution dimensions to obtain historical fireworks images of multiple channels. Each channel represents a feature dimension in the image, and different channels contain different image feature information. By analyzing and extracting features from the images of different channels separately, the image can be analyzed from multiple dimensions to more accurately identify these features. After performing multi-channel division on the image, the present invention uses the initial feature extraction module 022 to separately extract features from the image data of each channel, such as Figure 7As shown, the present invention first performs initial feature extraction on the historical firework images of each channel through a convolutional neural network algorithm to obtain the firework feature data corresponding to each channel. The lightweight convolutional neural network is particularly suitable for real-time monitoring in resource-constrained environments due to its relatively simple model structure and high computational efficiency. By applying the lightweight CNN to firework detection, efficient image processing and analysis can be achieved, and potential dangers can be detected in a timely manner. Then, the deep feature extraction module 023 is used to perform deep feature extraction on the firework feature data of each channel to capture the spatial feature information and channel feature information in the image. The dual attention mechanism includes two branches, namely the channel attention branch and the spatial attention branch. The channel attention branch is used to capture the correlation between the feature channels in the image, and the spatial attention branch is used to capture the correlation of the features in the spatial position. Through the dual attention feature extraction module 0231, the spatial feature information and channel feature information of the firework feature data of each channel can be extracted. Then, the feature enhancement and splicing module 0232 performs feature enhancement and feature splicing on the extracted spatial feature information and channel feature information to obtain the deep feature data of the historical firework image. This deep feature data can enhance the feature representation of small targets and improve the model's ability to capture detailed features, thereby more accurately detecting the spatial position information and significant feature information of small targets in the iron tower machine room. Through the lightweight CNN and the dual attention mechanism, the present invention can effectively capture the spatial and channel information of the features in the image in the task of detecting small targets in the iron tower machine room, and the dual attention convolution with a lightweight design has a low computational cost and is suitable for application in resource-limited environments.

[0084] Specifically, the present invention uses the model establishment module 03 to establish a firework detection model based on the dual attention mechanism.

[0085] In a specific embodiment of the present invention, the firework detection model based on the dual attention mechanism includes an input module, a feature extraction module, and an output module, as Figure 5As shown in the figure, the input module of the model is used to input the collected historical images of fireworks and environmental perception data. The environmental perception data includes CO concentration, temperature, and humidity data collected by sensors. Based on the detection requirements of small targets of fireworks in the iron tower machine room, the embodiment of the present invention designs a multi-channel input module to independently process different channel data through a lightweight convolutional neural network. Each channel represents a feature dimension in the image, and the same convolutional and attention operations are shared on the same channel. After the processed multi-channel features are spliced, they are used as the input of the dual attention mechanism, so as to more accurately detect the spatial position information and significant feature information of small targets. The output module of the model is used to output the recognition results of the image to be detected, including the type of fireworks, the target position, and the confidence score. On the other hand, the output module also outputs alarm information. The feature extraction module uses a convolutional layer and a dual attention mechanism. The convolutional layer adopts the traditional CNN convolutional neural network architecture. The dual attention module is divided into a channel attention model and a spatial attention model. The channel attention model and the spatial attention model adopt a parallel structure. The channel attention model is used to extract the significant information of the feature channels, and the spatial attention model is used to extract the position information of the target in space. The channel attention model focuses on the correlation between feature channels in the image. Each feature channel is regarded as a unit, and a channel correlation model is established through a self-attention network. The spatial attention model focuses on the correlation of features in spatial positions and also uses a self-attention mechanism to model spatial information. This parallel attention representation realizes feature modeling from different perspectives of the input image. Since the same convolutional and attention operations are shared on the same channel, for most targets in the image, similar feature spaces are shared from the channel and spatial perspectives, that is, they have consistent representations. This consistency can be maintained stably in channels and spaces. For noise or abnormal targets, their representation patterns are random and diverse, neither sharing strong correlations in channels and spaces with normal targets nor having consistency with other noise points. The fireworks detection model based on the dual attention mechanism established by the present invention can more accurately identify and distinguish normal targets from abnormal information, making the detection effect more accurate, especially suitable for small target detection in complex environments.

[0086] Specifically, the alarm module 05 further includes: The first judgment unit 051 is used to judge whether fireworks are detected based on the recognition result.

[0087] The second judgment unit 052 is used to judge whether a preset threshold is reached based on the environmental perception data.

[0088] The alarm unit 053 is used to send alarm information according to the judgment results of the first judgment unit and the second judgment unit.

[0089] In a specific embodiment of the present invention, the recognition result output by the model is combined with the environmental perception data detected by the sensor through the alarm module 05. When the first judgment unit 051 determines that smoke and fire are detected and the second judgment unit 052 determines that the detected environmental perception data reaches the preset threshold, the alarm unit 053 is immediately triggered to send an alarm message, and at the same time, the data slice is uploaded to the platform for storage, so as to facilitate subsequent personnel management and maintenance. By combining the recognition result with the environmental perception data detected by the sensor, once the model detects an anomaly, it immediately verifies with the data collected by the sensor end. If the data exceeds the preset threshold, an alarm is immediately triggered. This mechanism not only improves the response speed of smoke and fire detection but also reduces the possibility of false alarms.

[0090] Embodiment 3 To achieve the detection of small smoke and fire targets in the iron tower machine room, an embodiment of the present invention also provides a CNN smoke and fire detection product based on a dual attention mechanism to implement the above-mentioned CNN smoke and fire detection method based on a dual attention mechanism. The product includes: A sensor acquisition module, a camera vision processing module, an alarm intelligent connection module, and a cloud management module. The sensor acquisition module is used to collect environmental data such as CO concentration, temperature, and humidity in real time, providing necessary basic information for smoke and fire detection. The camera vision processing module is equipped with a small target detection algorithm based on a lightweight dual attention mechanism, providing inference and analysis capabilities for the visual information of the machine room under camera monitoring, and providing real-time protection for the normal operation of the machine room. The neural network algorithm used therein introduces a multi-scale feature fusion module with a unique lightweight dual attention module to provide small target feature information such as cigarette butts and sparks, and at the same time, it can well capture the long-range dependence of global features to obtain the feature information of small targets at a long distance. The small target detection algorithm integrated in the camera vision processing module can perform efficient real-time inference and detection at the camera end, ensuring smooth processing of the image and timely identifying smoke and fire and performing dual authentication with the environmental perception data collected by the sensor. The alarm intelligent connection module is used to combine the camera detection result with the sensor data. Once the camera detects smoke and fire and the relevant data collected by the sensor reaches the preset threshold, an alarm is immediately issued. This not only improves the response speed of smoke and fire detection but also reduces the possibility of false alarms. The cloud management module provides a convenient interface for remote operation and maintenance personnel and managers, supporting real-time monitoring and the issuance of alarm information. Through the cloud management module, users can achieve comprehensive monitoring and maintenance of the iron tower machine room, improving the overall safety management level.

[0091] In a specific embodiment of the present invention, data is first collected through CO, temperature, and humidity sensors to obtain indoor environmental perception data of the iron tower machine room. At the same time, model inference is completed at the camera end and suspicious small targets are monitored in real time. Finally, the monitored suspicious targets and sensor data are double-authenticated, and then the model is uploaded to the cloud for fusion and then sent to the camera end for update. The camera algorithm detection module analyzes the uploaded video frames based on the CNN smoke detection method based on the dual attention mechanism. Once an anomaly is detected, it immediately verifies with the data collected at the sensor end. Once the threshold is exceeded, an alarm response is immediately triggered, and the data slices are uploaded to the platform end for storage to facilitate subsequent management and maintenance work by personnel. The camera algorithm detection module is also based on the object detection optimization algorithm model and uploads the model. While ensuring the abnormality of the monitoring device, it provides a continuously updated inference model for early warning of fire danger scenarios such as detecting smoke and fire.

[0092] Embodiment 4 As Figure 13 shown, an embodiment of the present invention further provides an electronic device, including: a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned CNN smoke detection method based on the dual attention mechanism. The device in the present invention can be a server, a PC, a PAD, a mobile phone, etc.

[0093] Furthermore, an embodiment of the present invention further provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by the processor to implement the above-mentioned CNN smoke detection method based on the dual attention mechanism.

[0094] In summary, the CNN smoke detection method and system based on the dual attention mechanism described in the present invention have the following advantages: 1. A CNN model with a small number of parameters and a small target detection model based on the dual attention mechanism designed in the present invention can effectively detect dangerous scenarios caused by small targets such as cigarette butts and sparks in the iron tower machine room. Through the training and learning of the characteristics of long distances and small targets, the model can effectively detect the characteristics of smoke and fire targets in the real-time video of the camera, so as to achieve the recognition and tracking of small target objects, improve the accuracy of CNN in detecting small targets, enhance the feature representation of small targets and improve the model's ability to capture detailed features. When detecting dangerous situations caused by small targets such as cigarette butts in the iron tower machine room, it can ensure that dangerous accidents are cut off from the source.

[0095] 2. The present invention simultaneously utilizes CO, temperature, and humidity sensors to comprehensively assist in monitoring the safety situation in the computer room. It combines the sensor detection end and the visual algorithm end to jointly detect dangerous scenarios of smoke and fire. Once an abnormality is detected, it immediately verifies with the data collected by the sensor end. If the data exceeds the preset threshold, it immediately gives an alarm. This mechanism not only improves the response speed of smoke and fire detection but also reduces the possibility of false alarms.

[0096] 3. The present invention also provides optional remote operation and maintenance and management functions, which can detect the occurrence of dangerous abnormal accidents such as fires in the computer room in real time, thereby providing effective guarantee for the safe operation of the iron tower computer room.

[0097] The concept, principle, and idea of the present invention have been described in detail above in combination with specific implementation manners (including embodiments and examples). Those skilled in the art should understand that the implementation manners of the present invention are not limited to the several forms given above. After reading the application documents of the present invention, those skilled in the art can make any possible improvements, substitutions, and equivalent forms to the steps, methods, systems, and components in the above implementation manners. These improvements, substitutions, and equivalent forms should be regarded as falling within the scope of the present invention, and the protection scope of the present invention is only subject to the claims.

Claims

1. A CNN fireworks detection method based on dual attention mechanism, characterized in that: include: Acquire fireworks historical images and preprocess the fireworks historical images; Performing feature extraction on the processed fireworks historical image to obtain feature data of the fireworks historical image; Establishing a fireworks detection model based on a dual attention mechanism according to the feature data; Inputting the image to be detected into the fireworks detection model to perform fireworks recognition and output the recognition result; Double authentication is performed based on the recognition result and the environmental perception data, and if the authentication is successful, an alarm message is issued.

2. The CNN fireworks detection method based on dual attention mechanism according to claim 1, characterized in that: The acquiring of fireworks historical images and preprocessing of the fireworks historical images comprises: Acquire historical video stream data of fireworks and perform frame processing and data annotation on the historical video stream to obtain historical images of fireworks; The fireworks historical image is normalized.

3. The CNN fireworks detection method based on dual attention mechanism according to claim 1, characterized in that: The feature extraction of the processed fireworks historical image to obtain feature data of the fireworks historical image includes: Performing multi-channel division on the fireworks historical image based on image resolution to obtain multi-channel image data; Based on the convolutional neural network algorithm, initial feature extraction is performed on the multi-channel image data to obtain multi-channel feature data of the fireworks historical image; Based on the dual attention mechanism, deep feature extraction is performed on the multi-channel feature data respectively to obtain the deep feature data of the fireworks historical image.

4. The CNN fireworks detection method based on dual attention mechanism according to claim 3 is characterized in that: The deep feature extraction of the multi-channel feature data based on the dual attention mechanism is performed to obtain the deep feature data of the fireworks history image, including: Based on the dual attention mechanism, channel attention features and spatial attention features are respectively extracted from the multi-channel feature data, wherein the channel attention features include channel saliency information, and the spatial attention features include spatial position information; Feature enhancement processing and feature splicing are performed based on channel attention features and spatial attention features to obtain deep feature data of the fireworks historical image.

5. The CNN fireworks detection method based on dual attention mechanism according to claim 4 is characterized in that: The establishing of a fireworks detection model based on a dual attention mechanism according to the feature data includes: A channel attention model and a spatial attention model are respectively established according to the feature data, and the channel attention model and the spatial attention model adopt a parallel structure, wherein the channel attention model is used to extract the saliency information of the feature channel, and the spatial attention model is used to extract the position information of the target in space.

6. The CNN fireworks detection method based on dual attention mechanism according to claim 1, characterized in that: The method also includes acquiring environmental sensing data, wherein the environmental sensing data includes CO concentration, temperature and humidity data.

7. The CNN fireworks detection method based on dual attention mechanism according to claim 6, characterized in that: The dual authentication based on the recognition result and the environmental perception data, if the authentication is passed, the alarm information is issued including: Based on the recognition result, it is determined whether fireworks are detected. If detected, it is determined whether the environmental perception data reaches a preset threshold. If so, an alarm message is issued.

8. The CNN fireworks detection method based on dual attention mechanism according to claim 1, characterized in that: The recognition result includes the fireworks category, target location and confidence score.

9. A CNN fireworks detection system based on dual attention mechanism, characterized in that: include: A data acquisition and processing module, used for acquiring fireworks historical images and preprocessing the fireworks historical images; A feature extraction module, used to extract features from the processed fireworks historical image to obtain feature data of the fireworks historical image; A model building module, used to build a fireworks detection model based on a dual attention mechanism according to the feature data; A recognition module, inputting the image to be detected into the fireworks detection model to perform fireworks recognition and outputting the recognition result; The alarm module is used to perform dual authentication based on the recognition result and the environmental perception data, and if the authentication is successful, an alarm message is issued.

10. The CNN fireworks detection system based on dual attention mechanism according to claim 9, characterized in that: The data acquisition and processing module also includes: A data acquisition module is used to obtain historical video stream data of fireworks and perform frame processing and data annotation on the historical video stream to obtain historical images of fireworks; The data processing module is used to perform normalization processing on the fireworks historical image.

11. The CNN fireworks detection system based on dual attention mechanism according to claim 9, characterized in that: The feature extraction module also includes: A channel division module, used for performing multi-channel division on the fireworks historical image based on image resolution to obtain multi-channel image data; An initial feature extraction module, used to perform initial feature extraction on the multi-channel image data based on a convolutional neural network algorithm to obtain multi-channel feature data of the fireworks historical image; The deep feature extraction module is used to perform deep feature extraction on multi-channel feature data based on a dual attention mechanism to obtain deep feature data of the fireworks historical image.

12. The CNN fireworks detection system based on dual attention mechanism according to claim 11, characterized in that: The depth feature extraction module also includes: A dual attention feature extraction module, used to extract channel attention features and spatial attention features from multi-channel feature data based on a dual attention mechanism, wherein the channel attention features include channel saliency information, and the spatial attention features include spatial position information; The feature enhancement and splicing module is used to perform feature enhancement processing and feature splicing based on channel attention features and spatial attention features to obtain deep feature data of the fireworks historical image.

13. The CNN fireworks detection system based on dual attention mechanism according to claim 9, characterized in that: Also includes: Environmental perception data acquisition module, used to obtain CO concentration, temperature and humidity data.

14. The CNN fireworks detection system based on dual attention mechanism according to claim 13, characterized in that: The alarm module also includes: a first determination unit, configured to determine whether fireworks are detected based on the recognition result; A second judgment unit, used to judge whether a preset threshold is reached based on the environmental perception data; The alarm unit is used to send alarm information according to the judgment results of the first judgment unit and the second judgment unit.

15. An electronic device, characterized in that: include: A processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the CNN fireworks detection method based on the dual attention mechanism as described in any one of claims 1 to 8.

16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the CNN fireworks detection method based on the dual attention mechanism as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Smoke and fire detection early warning method and system based on YOLOV5 network

    CN114677629A

  • Smoking behavior detection method based on SDVGNet network

    CN118587762A

  • Early fire smoke detection method and system based on improved YOLOv9 algorithm

    CN118799805A

  • Smoke detection system and method

    CN119274038A

  • Subway fire detection method based on YOLOv8

    CN119942764A