Smoke and fire identification detection method and system combining infrared and visible light cameras
By combining the pyrotechnic identification and detection method of infrared and visible light cameras, and using convolutional network to process video stream data, the problems of slow response and high false alarm rate in the tower computer room are solved, and effective detection and identification of small target fireworks are achieved.
Patent Information
- Application Number
- CN202510592582.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional fire monitoring methods respond slowly in the early stages of the fire, and environmental noise and interference factors may lead to false alarms, making it difficult to effectively identify small target fireworks in the tower computer room.
The pyrotechnic identification and detection method combined with infrared and visible light cameras is adopted to obtain video stream data, pre-process, use convolutional network to process infrared and visible light feature maps, perform feature refinement and enhancement processing, and finally object detection and alarm response are carried out.
Effective detection and identification of small fireworks targets generated by equipment obstruction or blind spots in the tower computer room is realized, and the accuracy of small target detection of convolutional neural networks in complex environments is improved.
Smart Images

Figure CN120107869A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fire point detection, and in particular relates to a method, system, device and storage medium for identifying and detecting fireworks combined with infrared and visible light cameras. Background Art
[0002] In today's context of rapid development of informatization and intelligence, the tower room, as the core part of the communication network, carries key network equipment and data centers. The equipment in the room is dense and high-value, so it is very important to ensure its safe and stable operation. Fire accidents will not only seriously damage the equipment, but may also cause service interruptions, affect the normal operation of the entire communication network, and bring huge economic losses and social impacts.
[0003] Traditional fire monitoring methods mainly rely on smoke detectors or temperature sensors, but these devices often react slowly in the early stages of a fire and have difficulty identifying potential dangers in a timely manner. In addition, environmental noise and other interference factors may cause false alarms, affecting the reliability and effectiveness of the equipment.
[0004] In recent years, target detection technology based on deep learning has gradually become an effective solution to this problem. Convolutional neural networks are particularly suitable for real-time monitoring in resource-constrained environments due to their simple model structure and high computational efficiency. However, traditional methods have many challenges in detecting small targets such as fireworks in the blind spots of the field of view caused by equipment occlusion in the tower room, such as the difficulty in detecting small targets and the easy loss of feature information. In addition, it is often difficult to fully retain the spatial details of the image by relying solely on convolutional neural networks for multimodal fusion of infrared and visible light images, which is also an important problem facing current technology. Summary of the invention
[0005] In view of the problems mentioned in the above background technology, the present invention provides a method, system, device and storage medium for fire and smoke identification and detection combined with infrared and visible light cameras. The present invention provides a method for fire and smoke identification and detection combined with infrared and visible light cameras, the method comprising: Obtain video stream data, pre-process the video stream data, and obtain an infrared original feature map And the original feature map of visible light ; Using two convolutional networks, the infrared original feature map And the visible light original characteristic map to process; The processing results are subjected to feature refinement and block enhancement to obtain the detection results.
[0006] Furthermore, the infrared original feature map is respectively processed by using two convolutional networks. And the visible light original characteristic map Processing includes: The infrared original feature map And the visible light original characteristic map , respectively input into the corresponding convolution layer for extraction, and obtain the infrared feature map after convolution And the visible light feature map after convolution ; The infrared characteristic map And the visible light characteristic diagram Processing is performed to obtain the infrared global information map and visible light global information map .
[0007] Furthermore, the infrared original feature map And the visible light original characteristic map , respectively input into the corresponding convolutional layer for extraction, specifically: The infrared original feature map Enter the corresponding In the convolution layer, the infrared feature map after convolution is obtained ; The visible light original feature map Enter the corresponding In the convolution layer, the convolutional visible light feature map , the expression is:
[0008]
[0009] in, is the original infrared feature map, is the original feature map of visible light, is the infrared feature map after convolution. It is the visible light feature map after convolution.
[0010] Furthermore, the infrared characteristic map And the visible light characteristic diagram Processing is performed to obtain the infrared global information map and visible light global information feature map , including: The infrared characteristic map And the visible light characteristic diagram Perform fusion processing to obtain the global information feature map S; The global information feature map G is processed by convolution to obtain the infrared global information map and the visible light global information map , the expression is:
[0011]
[0012] in, is the infrared feature map after convolution. is the visible light feature map after convolution, represents element-by-element addition, This is the infrared global information map. is the global information map of visible light, represents the maximum pooling layer, and FC represents the fully connected layer.
[0013] Furthermore, the infrared characteristic map , the visible light characteristic graph 、The infrared global information map and the visible light global information map Perform differential comparison and calculate the information balance constraint value , the calculation formula is:
[0014] Among them, sigmoid represents the activation function applied to the binary classification task, which is used to represent the probability of detecting small target categories such as fireworks or cigarette butts.
[0015] Furthermore, when the information balance constraint value When the value is greater than or equal to the preset value, the original infrared feature map And the visible light original characteristic map Perform enhancement processing respectively to obtain infrared feature maps And the original feature map of visible light ; When the information balance constraint value When the infrared original characteristic image is less than the preset value, , Visible light original feature map Directly as infrared signature , Visible light original feature map , specifically expressed as:
[0016]
[0017]
[0018] in, The original feature map of visible light And the infrared global information map after convolution The visible light feature map obtained by performing element-by-element multiplication operation, For the original infrared feature map And the global information map of visible light after convolution The infrared feature map obtained by performing element-by-element multiplication operation, Represents element-wise multiplication.
[0019] Furthermore, the infrared characteristic image And the original feature map of visible light , perform feature refinement and block enhancement processing, including: The infrared feature map is respectively processed by using multi-scale dilated convolution and multiple enhancement modules. And the visible light original characteristic map Perform calculations; By performing a transposed convolution operation on the operation results and fusing them using a splicing operation, a feature-refined and enhanced infrared feature map is obtained. and visible light characteristics ; The infrared characteristic map , the visible light original characteristic map Respectively with the infrared feature map after feature refinement and enhancement and visible light characteristics The infrared image and the visible light image are fused.
[0020] Performing target detection on the infrared image and the visible light image; Based on the dual authentication, it is determined whether the target detection result is greater than a preset threshold. If it is greater than the preset threshold, an alarm response is performed.
[0021] The present invention also provides a fireworks recognition and detection system combining infrared and visible light cameras, the system comprising: The acquisition module is used to acquire video stream data, pre-process the video stream data, and obtain the infrared original feature map. And the original feature map of visible light ; A processing module is used to use two convolutional networks to respectively process the infrared original feature map And the visible light original characteristic map to process; The refinement and enhancement module is used to perform feature refinement and enhancement block processing on the processing result to obtain the detection result.
[0022] Furthermore, the processing module includes: Extracting feature unit, used for extracting the original infrared feature map And the visible light original characteristic map , respectively input into the corresponding convolution layer for extraction, and obtain the infrared feature map after convolution And the convolutional visible light feature map ; A processing unit, configured to convert the infrared characteristic image And the visible light characteristic diagram Processing is performed to obtain the infrared global information map and visible light global information map .
[0023] Furthermore, the system further comprises: An enhancement module is used to convert the infrared characteristic image , the visible light characteristic graph 、The infrared global information map and the visible light global information map Perform differential comparison and calculate the information balance constraint value , the calculation formula is:
[0024] Among them, Sigmoid represents the activation function applied to the binary classification task, which is used to represent the probability of detecting small target categories such as fireworks or cigarette butts.
[0025] Furthermore, the enhancement module includes: A judging unit is used to judge when the information balance constraint value When the value is greater than or equal to the preset value, the original infrared feature map And the visible light original characteristic map Perform enhancement processing respectively to obtain infrared feature maps And the original feature map of visible light ; When the information balance constraint value When the infrared original characteristic image is less than the preset value, , Visible light original feature map Directly as infrared signature , Visible light original feature map , specifically expressed as:
[0026]
[0027]
[0028] in, The original feature map of visible light And the infrared global information map after convolution The visible light feature map obtained by performing element-by-element multiplication operation, For the original infrared feature map And the global information map of visible light after convolution The infrared feature map obtained by performing element-by-element multiplication operation, Represents element-wise multiplication.
[0029] Furthermore, the refinement and enhancement module includes: A convolution unit is used to use multi-scale dilated convolution and multiple enhancement modules to respectively perform the infrared feature map And the visible light original characteristic map Perform calculations; The splicing unit is used to perform a transposed convolution operation on the operation results and fuse them using the splicing operation to obtain the infrared feature map after feature refinement and enhancement. and visible light characteristics ; A fusion unit is used to combine the infrared characteristic image , the visible light original characteristic map Respectively with the infrared feature map after feature refinement and enhancement and visible light characteristics Fusion is performed to obtain infrared images and visible light images; A detection unit, used for performing target detection on the infrared image and the visible light image; The alarm unit is used to determine whether the target detection result is greater than a preset threshold based on the dual authentication, and if so, to issue an alarm response.
[0030] The present invention also provides a device, including a processor, which is coupled to a memory; the processor is used to read and execute the computer program stored in the memory to implement the aforementioned fireworks recognition and detection method combining infrared and visible light cameras.
[0031] The present invention also provides a computer-readable storage medium storing a program or instruction. When the program or instruction is executed on a computer, the computer executes the aforementioned method for identifying and detecting fireworks by combining infrared and visible light cameras.
[0032] Compared with the prior art, the present invention has the following advantages: The present invention provides a method for identifying and detecting fireworks by combining infrared and visible light cameras. It uses two modal data obtained by two spectral imaging, and detects and identifies small targets such as fireworks through convolutional neural networks, and jointly discovers fireworks danger scenes, thereby providing effective protection for the safe operation of the computer room. At the same time, through the training and learning of long-distance and small target features, the model can effectively detect the features of fireworks targets in the real-time video of the camera, thereby achieving the recognition and tracking of small target objects, and improving the accuracy of the convolutional neural network in detecting small targets in occluded scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0034] Figure 1 A flowchart of a method for identifying and detecting fireworks by combining infrared and visible light cameras according to the present invention; Figure 2 Schematic diagram of a model framework of a convolutional neural network fireworks detection method based on a cross-modal complementary mechanism according to an embodiment of the present invention; Figure 3 Schematic diagram of the principle of the cross-modal complementary mechanism of an embodiment of the present invention; Figure 4 A schematic diagram of a feature refinement and enhancement block according to an embodiment of the present invention; Figure 5 It is a structural schematic diagram of a fireworks recognition and detection system combining infrared and visible light cameras of the present invention; Figure 6 It is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a series of steps or methods are not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes or methods.
[0037] The purpose of the present invention is to provide a method and system for identifying and detecting fireworks in combination with infrared and visible light cameras, which can effectively detect dangerous situations caused by small targets such as cigarette butts and sparks in the tower room that are blocked by equipment or in the blind area of vision, and provide optional remote operation and management functions. As a complete and reliable solution, it provides effective protection for the safety of the tower room, can detect the occurrence of dangerous abnormal accidents such as fire in the room in real time, provide intelligent protection for the normal operation of the tower, and use cameras with two imaging spectrums to conduct all-round auxiliary monitoring of the safety situation in the room.
[0038] In one embodiment of the present invention, a method for firework recognition and detection combining infrared and visible light cameras is provided, which completes model reasoning and monitors suspicious small targets in real time on the camera side in two modes, and uses the obtained sensor data for dual authentication and real-time alarm. After that, the model is uploaded to the cloud for fusion and then sent to the camera side for update. Figure 1 As shown, the method comprises the following steps: S1. Obtain video stream data, pre-process the video stream data, and obtain an infrared original feature map And the original feature map of visible light .
[0039] In this embodiment, the target area is monitored in real time by the camera end in infrared and visible light modes, and the video stream data in the two modes are obtained. The video stream data in the two modes are framed respectively to obtain the infrared original image and the visible light original image; the infrared original image and the visible light original image are annotated and normalized for small fireworks targets to obtain the original feature images in the two modes, namely, the infrared original feature image And the original feature map of visible light .
[0040] S2, using two convolutional networks, respectively, the infrared original feature map And the visible light original characteristic map to be processed.
[0041] In this embodiment, the obtained original feature map is processed by the camera algorithm processing module, and the camera algorithm processing module is equipped with a small target detection algorithm of a dual-modal camera imaging complementary mechanism, which is used to identify possible fireworks accidents under occluded objects, and the visual information of the computer room under camera monitoring provides reasoning and analysis capabilities, providing real-time protection for the normal operation of the computer room. The algorithm model design used introduces a unique cross-modal complementary mechanism, which can not only provide small target feature information such as cigarette butts and Mars in blind spots and equipment occlusions, but also capture the long-distance dependency of global features, so as to accurately identify small targets at a long distance.
[0042] In this embodiment, the cross-modal complementary mechanism is an important part of the model, which mainly divides the input data into two parallel branches to process the images imaged by infrared and visible light cameras respectively. The neural network branch that processes visible light focuses on the correlation between feature channels in the image under the visible light spectrum. Each feature channel is regarded as a unit, and the image information is analyzed by modeling the correlation between channels; the neural network branch that processes infrared focuses on the correlation of features in spatial position under the infrared spectrum, and also uses the self-attention mechanism to model spatial information. This parallel network representation method realizes the modeling of different perspective features of the input image.
[0043] For most targets in the image, they share similar feature spaces in channel and spatial perspectives, and this consistent representation can maintain stability in channel and space. For noise or abnormal targets, their representation patterns are random and diverse, neither sharing strong correlations in channels and space with normal targets, nor lacking consistency with other noise points. In this way, the model can more accurately identify and distinguish normal targets from abnormal information, making the detection effect more accurate, especially suitable for small target detection in complex environments. Combined with the application of infrared and visible light modalities, this cross-modal complementary mechanism significantly improves the system's ability to detect targets under a variety of spectral conditions, ensuring reliability and accuracy in various environments. In this embodiment, for feature learning specific to cross-modality, we set a message passing mechanism-based module in the encoder, using the two modal information to mutually optimize the extracted rational features, find the difference information between the features they extract, and further dynamically complement and interact with the differences.
[0044] In this embodiment, two convolutional networks are used to respectively process the infrared original feature map. And the visible light original characteristic map The processing includes the following steps: S21, infrared original feature map And the original feature map of visible light , respectively input into the corresponding 3 × 3 convolutional layers in the parallel connection, and extract two mapping dimensions about local information, namely the infrared feature map after convolution And the convolutional visible light feature map ; The infrared characteristic map And the visible light characteristic diagram The feature information at different positions is aggregated together to obtain a feature map S containing global information, such as Figure 3 As shown, the specific method is: Based on the original infrared feature map And the original feature map of visible light , respectively select the extracted features of the Kth layer (K = 1, 2, 3, 4), and divide K into regions, and the dimension of each region is and feed it into The maximum pooling layer is used to obtain a C×1×1 information vector. The C output neurons are used to generate a global information vector through the fully connected layer, and the global information vector is replicated H × W times to reshape the global information feature map S. Among them, C, H and W are the number of channels, height and width of these features.
[0045] Finally, convolution operation and 1×1 convolution are used to transform the global information feature map. The channels and image size are adjusted, and the infrared feature map is separated. and visible light characteristics Two feature images with the same dimensional information, namely infrared global information map and visible light global information map , using feature connection and Perform element-by-element multiplication and we get:
[0046]
[0047]
[0048]
[0049] in, is the original infrared feature map, is the original feature map of visible light, is the infrared feature map after convolution. is the visible light feature map after convolution, represents element-by-element addition, represents element-wise multiplication, This is the infrared global information map. is the global information map of visible light, represents the maximum pooling layer, and FC represents the fully connected layer.
[0050] S23, the infrared characteristic map , the visible light characteristic diagram 、The infrared global information map and the visible light global information map Perform differential comparison and calculate the information balance constraint value .
[0051] In this embodiment, an information balance constraint is designed in this module , which is similar to many dynamic complementation mechanisms based on gated devices, the specific method is: Based on the subtraction operation between pixels, the infrared global information map is Visible light characteristics , Visible light global information map With infrared signature The pixel difference weight after subtraction is processed by sigmoid to a value between 0 and 1, and then the ratio of the change with the original feature image is taken as the value of Ө.
[0052]
[0053] Among them, sigmoid represents the activation function applied to the binary classification task, which is used to represent the probability of detecting small target categories such as fireworks or cigarette butts.
[0054] In this embodiment, when the information balance constraint value When the difference is greater than or equal to the preset value, it is considered that the difference is large, and the CMIDI module can be used to perform complementary enhancement to their respective downsampling layers; And the original feature map of visible light Perform enhancement processing respectively to obtain infrared feature maps And the original feature map of visible light .
[0055] When the information balance constraint value When it is less than the preset value, it means that the information extracted by the parallel downsampling is not very different, and no feature enhancement operation is required to avoid consuming extra complex calculation and reasoning time due to minor or invalid enhancement. , Visible light original feature map Directly as infrared signature , Visible light original feature map The preset value may be 0.15. The gate control principle is calculated as follows:
[0056]
[0057]
[0058] in, The original feature map of visible light And the infrared global information map after convolution The visible light feature map obtained by performing element-by-element multiplication operation, For the original infrared feature map And the global information map of visible light after convolution The infrared feature map obtained by performing element-by-element multiplication operation, Represents element-wise multiplication.
[0059] S3. Perform feature refinement and block enhancement processing on the processing result to obtain the detection result.
[0060] In this embodiment, the feature refinement and enhancement block uses horizontal strip convolution and vertical strip convolution to convolve the feature map in the horizontal and vertical directions, respectively, to capture feature information of different shapes and directions, and to improve the richness and accuracy of feature representation.
[0061] In this embodiment, in order to capture more comprehensive feature information and better adapt to edges and textures in different directions, depth-separable convolution is added after horizontal and vertical convolutions, and channel attention is introduced. Among them, the horizontal branch can be composed of 1×K horizontal strip convolution, K×1 depth convolution and 1×1 point-by-point convolution; the vertical branch can be composed of 1×K vertical strip convolution, 1×K depth convolution and 1×1 point-by-point convolution. Through the overlap of horizontal convolution and vertical convolution, the ability of BLCM to extract features in multiple directional dimensions is improved.
[0062] In this embodiment, Figure 4 The schematic diagram of the feature refinement and enhancement block structure is as follows. First, the input end of the feature refinement and enhancement block is connected to the output characteristics of the multimodal feature fusion module, that is, the infrared feature map And the original feature map of visible light Then, five dilated convolutions with different dilation rates and five enhancement modules are used to obtain road features at multiple scales. At the same time, the enhancement modules of different branches are weighted in a shared way to reduce the number of network parameters. The outputs of the five branches are passed through different transposed convolution layers to obtain feature maps of the same dimension. Through the splicing operation, a large amount of contextual feature information in these feature maps is fused into a larger receptive field, thereby recovering the lost detail information. This process can be expressed as:
[0063]
[0064]
[0065]
[0066]
[0067]
[0068] Among them, F is the feature map that needs to be processed by feature refinement and enhancement block, that is, the infrared feature map obtained in the above steps , Visible light original feature map ; represents the i-th (i=1,2,3,4,5) layer, Represents dilated convolutions with different dilation rates; Represents the feature map after feature refinement and enhancement, that is, the infrared feature map after feature refinement and enhancement is obtained respectively and visible light characteristics ; BN represents batch normalization layer, Switch represents lightweight activation function, and Concat represents feature concatenation operation.
[0069] In this embodiment, a lightweight activation function Switch is used to eliminate the hidden danger of gradient explosion in forward propagation, the BN layer performs batch normalization, and finally a 1×1 convolution is used to change the number of channels as the output of the feature refinement and enhancement block.
[0070] In this embodiment, in order to reduce the complexity of training and enhance the plug-and-play nature of the module, a residual strategy is introduced, that is, the feature map after feature refinement and enhancement is processed based on the residual strategy. After processing, we can get the infrared detection result graph and the visible light detection result graph. The whole process can be described as:
[0071]
[0072] Among them, F is the feature map that needs to be processed by feature refinement and enhancement block, that is, the infrared feature map obtained in the above steps , Visible light original feature map ; Conv represents the feature map after feature refinement and enhancement, that is, the infrared feature map and visible light feature map after feature refinement and enhancement are obtained respectively; represents a single-layer convolution operation, is the feature map obtained by one convolution, Conv represents a 3×3 convolution operation, represents the element-by-element addition operation, The detection result graph represents the final output.
[0073] In this embodiment, the infrared detection result image and the visible light detection result image are detected by the camera algorithm detection module; If the algorithm under the infrared camera or visible light camera detects an abnormality when detecting the corresponding detection result graph, it will immediately verify the detection result of the algorithm under the visible light camera or infrared camera. Once it exceeds the preset threshold, the real-time alarm intelligent module will immediately respond to the alarm; and upload the data to the cloud for management, which is convenient for subsequent personnel to manage and maintain. While ensuring the abnormality of the monitoring equipment, it provides a continuously updated reasoning model for detecting fire hazard scenes such as fireworks.
[0074] Alarm and respond to fireworks accidents based on the recognition results, and upload the recognition results to the cloud platform.
[0075] The above contents are only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
[0076] In this embodiment, the alarm intelligent connection module is responsible for combining the detection results of the cameras of the two imaging modes. Once fireworks are detected, an alarm can be issued. This not only improves the response speed of fireworks detection, but also reduces the possibility of false alarms, and notifies the computer room management personnel and triggers the alarm as quickly as possible.
[0077] In this embodiment, the cloud management module provides a convenient interface for remote operation and maintenance and management personnel, supporting real-time monitoring and the issuance of alarm information. Through cloud management, users can achieve comprehensive monitoring and maintenance of the tower room, improving the overall safety management level.
[0078] Figure 2 Schematic diagram of the model framework of the convolutional neural network fireworks detection method based on the cross-modal complementary mechanism according to an embodiment of the present invention. Figure 2 The model includes an input module, a feature extraction module and an output module, specifically including: The input module mainly includes collecting the target video stream data through the camera, and performing frame processing, image processing, normalization and data labeling; The feature complementary enhancement module mainly includes a complementary mechanism and a lightweight feature refinement and enhancement block, which is used to perform convolution operation feature extraction on the processing results of the input module to obtain infrared images and visible light images, namely infrared feature maps. , Visible light original feature map ; Through the infrared characteristic map , Visible light original feature map Perform feature refinement and block enhancement processing to obtain infrared detection result images and visible light detection result images, and perform detection based on the camera algorithm detection module; use the alarm intelligent connection module to combine the camera detection results of the two modal imaging modes to perform dual authentication and identify dangers; The output module mainly includes the output of the detection results of the camera algorithm detection module, including classification output, positioning output and confidence score; and intelligent alarm for the dual authentication results of the feature complementary enhancement module.
[0079] Aiming at the problem of detecting small fireworks targets in the blind area of vision blocked by equipment in the tower machine room, the present invention proposes a fireworks recognition and detection method and system combining infrared and visible light cameras. The core idea is to enhance the features of small targets under occlusion conditions and improve the model's ability to capture detail features. First, the two input modal images are extracted using a convolutional neural network. The lightweight feature refinement and enhancement blocks have a low computational cost and are suitable for use in resource-limited environments. Since it is difficult to fully retain the spatial details of the image by relying solely on convolutional neural networks for multimodal fusion under infrared and visible light vision, the receptive field extracted by the model is expanded to improve the ability to capture edge feature information of small targets, thereby obtaining effective information on smaller targets.
[0080] In feature modeling, a loss function based on feature saliency is designed. The similarity between the target feature and the background feature is measured by calculating the maximum mean difference between the two features, thereby improving the model's detection accuracy for small targets. Based on the detection requirements of small fireworks targets in the tower machine room, the model input module designs multimodal input features, performs independent convolutional neural network processing on different modal data, and shares the same convolution and attention operations. The processed multi-channel features are spliced through features and used as the input of the dual attention mechanism, so as to more accurately detect the spatial position information and salient feature information of occluded small targets.
[0081] In the task of detecting small targets in the tower room, the input multimodal feature map effectively captures the spatial and channel information of the features in the image through convolutional neural networks and cross-modal complementary mechanisms. The feature splicing result is used as the output result of the two modalities, and the detection of small targets such as fireworks is more effectively completed through feature refinement and enhancement blocks.
[0082] The present invention also provides a fireworks recognition and detection system combining infrared and visible light cameras, the system comprising: The acquisition module 501 is used to acquire video stream data and pre-process the video stream data to obtain an infrared original feature map. And the original feature map of visible light ; The processing module 502 is used to use two convolutional networks to respectively process the infrared original feature map And the visible light original characteristic map to process; The refinement and enhancement module 503 is used to perform feature refinement and enhancement block processing on the processing result to obtain the detection result.
[0083] In this embodiment, the processing module includes: Extracting feature unit, used for extracting the original infrared feature map And the visible light original characteristic map , respectively input into the corresponding convolution layer for extraction, and obtain the infrared feature map after convolution And the visible light feature map after convolution ; A processing unit, configured to convert the infrared characteristic image And the visible light characteristic diagram Processing is performed to obtain the infrared global information map and visible light global information map .
[0084] In this embodiment, the system further includes: An enhancement module is used to convert the infrared characteristic image , the visible light characteristic diagram 、The infrared global information map and the visible light global information map Perform differential comparison and calculate the information balance constraint value , the calculation formula is:
[0085] Among them, Sigmoid represents the activation function applied to the binary classification task, which is used to represent the probability of detecting small target categories such as fireworks or cigarette butts.
[0086] In this embodiment, the enhancement module includes: A judging unit is used to judge when the information balance constraint value When the value is greater than or equal to the preset value, the original infrared feature map And the visible light original characteristic map Perform enhancement processing respectively to obtain infrared feature maps And the original feature map of visible light ; When the information balance constraint value When the infrared original characteristic image is less than the preset value, , Visible light original feature map Directly as infrared signature , Visible light original feature map , specifically expressed as:
[0087]
[0088]
[0089] in, The original feature map of visible light And the infrared global information map after convolution The visible light feature map obtained by performing element-by-element multiplication operation, For the original infrared feature map And the global information map of visible light after convolution The infrared feature map obtained by performing element-by-element multiplication operation, Represents element-wise multiplication.
[0090] In this embodiment, the refinement and enhancement module includes: A convolution unit is used to use multi-scale dilated convolution and multiple enhancement modules to respectively perform the infrared feature map And the visible light original characteristic map Perform calculations; The splicing unit is used to perform a transposed convolution operation on the operation results and fuse them using the splicing operation to obtain the infrared feature map after feature refinement and enhancement. and visible light characteristics ; A fusion unit is used to combine the infrared characteristic image , the visible light original characteristic map Respectively with the infrared feature map after feature refinement and enhancement and visible light characteristics Fusion is performed to obtain infrared images and visible light images; A detection unit, used for performing target detection on the infrared image and the visible light image; The alarm unit is used to determine whether the target detection result is greater than a preset threshold based on the dual authentication, and if so, to issue an alarm response.
[0091] like Figure 6 As shown, an embodiment of the present invention further provides a device, including: a processor 601, the processor 601 is coupled to a memory 602, and the processor 601 is used to read and execute a computer program stored in the memory 602 to implement a fireworks recognition and detection method combining infrared and visible light cameras as described in the above method embodiment.
[0092] An embodiment of the present invention also provides a computer-readable storage medium, which stores a program or instruction. When the above program or instruction is run on a computer, the computer executes a fireworks recognition and detection method combining infrared and visible light cameras as described in the above method embodiment.
[0093] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying and detecting fireworks by combining infrared and visible light cameras, characterized in that: The method comprises: Obtain video stream data, pre-process the video stream data, and obtain an infrared original feature map And the original feature map of visible light ; Using two convolutional networks, the infrared original feature map And the visible light original characteristic map to process; The processing results are subjected to feature refinement and block enhancement to obtain the detection results.
2. The method according to claim 1, characterized in that The two convolutional networks are used to respectively process the infrared original feature map. And the visible light original characteristic map Processing includes: The infrared original feature map And the visible light original characteristic map , respectively input into the corresponding convolution layer for extraction, and obtain the infrared feature map after convolution And the convolutional visible light feature map ; The infrared characteristic map And the visible light characteristic diagram Processing is performed to obtain the infrared global information map and visible light global information map .
3. The method according to claim 2, characterized in that The infrared original characteristic map And the visible light original characteristic map , respectively input into the corresponding convolutional layer for extraction, specifically: The infrared original feature map Enter the corresponding In the convolution layer, the infrared feature map after convolution is obtained ; The visible light original feature map Enter the corresponding In the convolution layer, the convolutional visible light feature map , the expression is: in, is the original infrared feature map, is the original feature map of visible light, is the infrared feature map after convolution. It is the visible light feature map after convolution.
4. The method according to claim 2, characterized in that: The infrared characteristic map And the visible light characteristic diagram Processing is performed to obtain the infrared global information map and visible light global information feature map , specifically including: The infrared characteristic map And the visible light characteristic diagram Perform fusion processing to obtain the global information feature map S; The global information feature map G is processed by convolution to obtain the infrared global information map and the visible light global information map , the expression is: in, is the infrared feature map after convolution. is the visible light feature map after convolution, represents element-by-element addition, This is the infrared global information map. is the global information map of visible light, represents the maximum pooling layer, and FC represents the fully connected layer.
5. The method according to claim 2, characterized in that: The infrared characteristic map , the visible light characteristic graph 、The infrared global information map and the visible light global information map Perform differential comparison and calculate the information balance constraint value , the calculation formula is: Among them, sigmoid represents the activation function applied to the binary classification task, which is used to represent the probability of detecting small target categories such as fireworks or cigarette butts.
6. The method according to claim 5, characterized in that When the information balance constraint value When the value is greater than or equal to the preset value, the original infrared feature map And the visible light original characteristic map Perform enhancement processing respectively to obtain infrared feature maps And the original feature map of visible light ; When the information balance constraint value When the infrared original characteristic image is less than the preset value, , Visible light original feature map Directly as infrared signature , Visible light original feature map , specifically expressed as: in, The original feature map of visible light And the infrared global information map after convolution The visible light feature map obtained by performing element-by-element multiplication operation, For the original infrared feature map And the global information map of visible light after convolution The infrared feature map obtained by performing element-by-element multiplication operation, Represents element-wise multiplication.
7. The method according to claim 6, characterized in that The infrared characteristic map And the original feature map of visible light , perform feature refinement and enhancement block processing, including: The infrared feature map is respectively processed by using multi-scale dilated convolution and multiple enhancement modules. And the visible light original characteristic map Perform calculations; By performing a transposed convolution operation on the calculation results and fusing them using a splicing operation, we can obtain a refined and enhanced infrared feature map. and visible light characteristics ; The infrared characteristic map , the visible light original characteristic map Respectively with the infrared feature map after feature refinement and enhancement and visible light characteristics Fusion is performed to obtain infrared images and visible light images; Performing target detection on the infrared image and the visible light image; Based on the dual authentication, it is determined whether the target detection result is greater than a preset threshold. If it is greater than the preset threshold, an alarm response is performed.
8. A fireworks recognition and detection system combining infrared and visible light cameras, characterized in that: The system comprises: The acquisition module is used to acquire video stream data, pre-process the video stream data, and obtain the infrared original feature map. And the original feature map of visible light ; A processing module is used to use two convolutional networks to respectively process the infrared original feature map And the visible light original characteristic map to process; The refinement and enhancement module is used to perform feature refinement and enhancement block processing on the processing result to obtain the detection result.
9. The system according to claim 8, characterized in that The processing module comprises: Extracting feature unit, used for extracting the original infrared feature map And the visible light original characteristic map , respectively input into the corresponding convolution layer for extraction, and obtain the infrared feature map after convolution And the convolutional visible light feature map ; A processing unit, configured to convert the infrared characteristic image And the visible light characteristic diagram Processing is performed to obtain the infrared global information map and visible light global information map .
10. The system according to claim 9, characterized in that The system further comprises: An enhancement module is used to transform the infrared characteristic image , the visible light characteristic diagram 、The infrared global information map and the visible light global information map Perform differential comparison and calculate the information balance constraint value , the calculation formula is: Among them, Sigmoid represents the activation function applied to the binary classification task, which is used to represent the probability of detecting small target categories such as fireworks or cigarette butts.
11. The system according to claim 10, characterized in that The enhancement module comprises: A judging unit is used to judge when the information balance constraint value When the value is greater than or equal to the preset value, the original infrared feature map And the visible light original characteristic map Perform enhancement processing respectively to obtain infrared feature maps And the original feature map of visible light ; When the information balance constraint value When the infrared original characteristic image is less than the preset value, , Visible light original feature map Directly as infrared signature , Visible light original feature map , specifically expressed as: in, The original feature map of visible light And the infrared global information map after convolution The visible light feature map obtained by performing element-by-element multiplication operation, For the original infrared feature map And the global information map of visible light after convolution The infrared feature map obtained by performing element-by-element multiplication operation, Represents element-wise multiplication.
12. The system according to claim 10, characterized in that The refinement and enhancement module includes: A convolution unit is used to use multi-scale dilated convolution and multiple enhancement modules to respectively perform the infrared feature map And the visible light original characteristic map Perform calculations; The splicing unit is used to perform a transposed convolution operation on the operation results and fuse them using the splicing operation to obtain the infrared feature map after feature refinement and enhancement. and visible light characteristics ; A fusion unit is used to combine the infrared characteristic image , the visible light original characteristic map Respectively with the infrared feature map after feature refinement and enhancement and visible light characteristics Fusion is performed to obtain infrared images and visible light images; A detection unit, used for performing target detection on the infrared image and the visible light image; The alarm unit is used to determine whether the target detection result is greater than a preset threshold based on the dual authentication, and if so, to issue an alarm response.
13. An electronic device, characterized in that: comprising a processor coupled to a memory; The processor is used to read and execute the computer program stored in the memory to implement a fireworks recognition and detection method combining infrared and visible light cameras as described in any one of claims 1 to 7.
14. A computer storage medium, characterized in that: A program or instruction is stored, and when the program or instruction is run on a computer, the computer is caused to execute a fireworks recognition and detection method combining infrared and visible light cameras as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Weak and small target detection method and device based on infrared and visible light feature fusion
CN116958782A
Long-distance video fire detection method based on dual-light fusion
CN118799800A
Target detection method and system based on infrared and visible light image fusion
CN119649175A