Fire target detection method and device

Through the improved three-dimensional convolutional neural network and YOLOv8 target detection network, the fire target detection is fused with multimodal video data, which solves the problems of false alarms, missed alarms and insufficient anti-interference capabilities in the existing technology, and achieves high-accuracy fire detection.

CN119964060APending Publication Date: 2025-05-09SHENYANG FIRE RES INST OF MEM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510443460.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing image fire detection methods have problems with false alarms and missed reports, and have poor anti-interference capabilities. There is a lack of detection methods that use deep learning methods to extract, fusion and utilization of multimodal information.

Method used

The improved three-dimensional convolutional neural network is adopted to fuse visual, near-infrared, thermal imaging video data, and fire target detection is performed through the YOLOv8 target detection network, and multi-scale fire image features are extracted using the backbone network layer, the neck network layer and the small target detection layer.

Benefits of technology

It realizes accurate identification of fire images, improves the accuracy of fire detection, reduces false alarms and missed alarm rates, and enhances anti-interference ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964060A_ABST
    Figure CN119964060A_ABST
Patent Text Reader

Abstract

A fire target detection method comprises the following steps: acquiring fire video data of multiple modes, grouping the fire video data of multiple modes, and processing the grouped fire video data into a video frame sequence; performing feature extraction on the processed multi-mode fire video data by using a three-dimensional convolutional neural network, and fusing the features into fire fusion features; and inputting the fire fusion feature into a YOLOv8 target detection network of an improved neck network for fire target detection, wherein the YOLOv8 target detection network of the improved neck network comprises a backbone network layer, a neck network layer and a small target detection layer. According to the method, the fire image multi-modal features including visual, near-infrared, thermal imaging and dynamic features can be extracted, the YOLOv8 target network structure is improved, and the accuracy of fire detection can be effectively improved through a target detection mode of fire multi-modal feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fire detection and alarm technology in the field of deep learning, and in particular to a method and device for fire target detection based on visible, near-infrared and thermal imaging videos of three-dimensional convolutional neural networks. Background Art

[0002] In recent years, various types of fires have occurred frequently, causing great harm to social and economic development and human life safety. At present, fire detection technologies include traditional fire detection methods such as gas alarms and smoke alarms, as well as image-based fire detection methods. With the development of video surveillance methods, the use of visual sensors to assist fire detection has attracted much attention. Compared with traditional fire detection methods such as point-type smoke and temperature sensing, the advantages of image fire detection include rapid response and the ability to obtain real-time images or videos of fire scenes, which is conducive to emergency decision-making and is an ideal solution for fire detection in large indoor spaces and outdoor environments.

[0003] However, there are still some problems with the existing image-based fire detection methods: (1) The existing image-based fire detection technology uses a single visible light image information to extract fire features. The features are relatively simple and there are problems of false alarms and missed alarms, which seriously hinders the promotion and application of image-based fire detection technology. (2) Visible light images are easily interfered by light. Infrared images themselves lack key information such as color and texture of flames. The single-mode video flame detection algorithm has poor anti-interference ability. Among the dual-channel flame detection methods, those that use infrared technology all use traditional methods to identify flame features, and some use binocular vision technology to locate flames. There is a lack of detection methods that use deep learning methods to extract, fuse and utilize multi-modal information such as visible light, near infrared, and thermal imaging. (3) Deep learning methods can extract the depth information of images for analysis, but single-frame images can only reflect the spatial characteristics of fireworks and lack the dynamic information of fireworks in time. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method and device for fire target detection, which effectively utilizes an improved three-dimensional convolutional neural network, fuses multi-spectral images and dynamic features of fire, and realizes accurate identification of fire images based on a multi-spectral composite method.

[0005] In order to solve the above technical problems, the present invention provides a method for detecting a fire target, comprising: Collecting fire video data of multiple modes, grouping the fire video data of multiple modes and processing them into video frame sequences; The 3D convolutional neural network is used to extract features from the processed multi-modal fire video data and fuse them into fire fusion features; The fire fusion feature is input into the YOLOv8 target detection network of the improved neck network for fire target detection, wherein the YOLOv8 target detection network of the improved neck network comprises a backbone network layer, a neck network layer and a small target detection layer.

[0006] Furthermore, the step of inputting the fire fusion feature into the YOLOv8 target detection network of the improved neck network for fire target detection includes: The fire fusion features are input into the backbone network layer, and multi-scale fire image features are output after passing through four stages of CBS module, C2f module and SPPF module in the backbone network layer; Input the multi-scale fire image features into the neck network layer, and output multiple fire feature maps after the neck network layer passes through convolution, upsampling, splicing, GSConv module and V-GSCSP module, wherein the GSConv module is used to perform a convolution operation on the input features, and after a first spatial convolution, a depth-separable convolution, and then a second spatial convolution, the features output by the first spatial convolution and the second spatial convolution are spliced ​​and input into the third spatial convolution, and the features are output after a Shuffle operation, and the V-GSCSP module is used to perform a convolution operation on the input features to output two branches, the first branch passes through two GSConv modules, and the second branch is spliced ​​with the output of the first branch and then passes through a convolution to output the features; Inputting the plurality of fire characteristic images into a small target detection layer, and outputting a plurality of fire detection result images of different sizes through a plurality of detection modules in the small target detection layer; The coordinate information of the fire boundary box is outputted according to the multiple fire detection result images of different sizes.

[0007] Furthermore, the multi-scale fire image features are output after the four-stage CBS module, C2f module and SPPF module in the backbone network layer, including: The first stage: after the fire fusion feature is input into a CBS module, it enters the fire image feature extraction operation, including: after passing through a CBS module, it enters the C2f module, in the C2f module, it passes through the first convolution layer and is divided into two branches, the first branch is passed to the output, and the second branch passes through three bottleneck modules, after the two branches are spliced, it passes through the second convolution layer and then outputs the fire image feature; The second stage: performing the fire image feature extraction operation on the fire image feature outputted in the first stage, and outputting the first fire image feature P1; The third stage: performing the fire image feature extraction operation on the fire image feature outputted in the second stage, and outputting the second fire image feature P2; The fourth stage: performing the fire image feature extraction operation on the fire image feature outputted in the third stage, and then outputting the third fire image feature P3 after passing through the spatial pyramid pooling fast module.

[0008] Furthermore, the multi-scale fire image features are input into the neck network layer, and after the neck network layer undergoes convolution, upsampling, concatenation, GSConv module and V-GSCSP module, multiple fire feature maps are output, including: After the third fire image feature P3 is upsampled, it is concatenated with the second fire image feature P2 and then input into the V-GSCSP module to obtain the first feature map Q1; After the first feature map Q1 is upsampled, it is concatenated with the first fire image feature P1 and input into the V-GSCSP module to obtain the second feature map Q2; The second feature map Q2 is input into the GSConv module and then concatenated with the first feature map Q1, and then input into the V-GSCSP module to obtain the first fire feature map J1; The first fire feature map J1 is input into the GSConv module, concatenated with the third fire image feature P3, and then input into the V-GSCSP module to obtain the second fire feature map J2.

[0009] Furthermore, the inputting of the plurality of fire characteristic images into the small target detection layer, and outputting a plurality of fire detection result images of different sizes through a plurality of detection modules in the small target detection layer, comprises: The second feature map Q2, the first fire feature map J1 and the second fire feature map J2 are respectively input, and respectively pass through the corresponding detection modules in the small target detection layer to output a small-size fire detection result map, a medium-size fire detection result map and a large-size fire detection result map.

[0010] Furthermore, the three-dimensional convolutional neural network is used to extract features from the processed multi-modal fire video data and fuse them into fire fusion features, including: The processed multi-modal fire video data are respectively passed through a three-dimensional convolutional neural network to output corresponding features, and after a splicing operation is performed on the output corresponding features, they are input into the three-dimensional convolutional neural network to obtain the fire fusion features.

[0011] A fire target detection device, comprising: A video acquisition unit, used for acquiring fire video data of multiple modes; A processing unit, used for grouping the multimodal fire video data and processing them into video frame sequences; used for extracting features from the processed multimodal fire video data using a three-dimensional convolutional neural network, and fusing them into fire fusion features; The detection unit is used to input the fire fusion feature into the YOLOv8 target detection network of the improved neck network to perform fire target detection. The YOLOv8 target detection network of the improved neck network includes a backbone network layer, a neck network layer and a small target detection layer.

[0012] Furthermore, the detection unit inputs the fire fusion feature into the YOLOv8 target detection network of the improved neck network for fire target detection, including: inputting the fire fusion feature into the backbone network layer, and outputting multi-scale fire image features after passing through four stages of CBS module, C2f module and SPPF module in the backbone network layer; inputting the multi-scale fire image features into the neck network layer, and outputting multiple fire feature maps after passing through convolution, upsampling, splicing, GSConv module and V-GSCSP module in the neck network layer, wherein the GSConv module is used to perform a convolution operation on the input features, and after a first spatial convolution, a depth-separable convolution, and then a second spatial convolution, the features output by the first spatial convolution and the second spatial convolution are spliced ​​and input into a third spatial convolution, and after Shuffle, the features are output. After the operation, feature output is performed, the V-GSCSP module is used to convolve the input features and output two branches, the first branch passes through two GSConv modules, and the second branch is spliced ​​with the output of the first branch and then passes through a convolution for feature output; the multiple fire feature maps are input into the small target detection layer, and multiple fire detection result maps of different sizes are output in the small target detection layer through multiple detection modules; the coordinate information of the fire boundary box is output according to the multiple fire detection result maps of different sizes.

[0013] Furthermore, the detection unit outputs multi-scale fire image features after passing through four stages of CBS modules, C2f modules and SPPF modules in the backbone network layer, including: the first stage: after the fire fusion feature is input into a CBS module, it enters the fire image feature extraction operation including: after passing through a CBS module, it enters the C2f module, in the C2f module, it is divided into two branches after passing through the first convolution layer, the first branch is transmitted to the output, and the second branch passes through three bottleneck modules, the two branches are spliced, and then pass through the second convolution layer and then output the fire image feature; the second stage: the fire image feature extraction operation is performed on the fire image feature output in the first stage, and the first fire image feature P1 is output; the third stage: the fire image feature extraction operation is performed on the fire image feature output in the second stage, and the second fire image feature P2 is output; the fourth stage: the fire image feature extraction operation is performed on the fire image feature output in the third stage, and then the third fire image feature P3 is output after passing through the spatial pyramid pooling fast module.

[0014] Further, the detection unit inputs the multi-scale fire image features into the neck network layer, and outputs multiple fire feature maps after the neck network layer passes through convolution, upsampling, splicing, GSConv module and V-GSCSP module, including: after the third fire image feature P3 is upsampled, it is spliced ​​with the second fire image feature P2 and then input into the V-GSCSP module to obtain the first feature map Q1; after the first feature map Q1 is upsampled, it is spliced ​​with the first fire image feature P1 and then input into the V-GSCSP module to obtain the second feature map Q2; after the second feature map Q2 is input into the GSConv module, it is spliced ​​with the first feature map Q1, and then input into the V-GSCSP module to obtain the first fire feature map J1; after the first fire feature map J1 is input into the GSConv module, it is spliced ​​with the third fire image feature P3, and then input into the V-GSCSP module to obtain the second fire feature map J2.

[0015] In summary, this method effectively utilizes the improved three-dimensional convolutional neural network, integrates multi-spectral images and dynamic features of fire, realizes accurate identification of fire images based on multi-spectral composite methods, and develops a real-time fire detection device with visible, near-infrared and thermal imaging modes, which provides methodological support for ultimately solving the technical bottleneck problems of false alarms of flame detection and low reliability of smoke detection in engineering applications of image fire detection technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flow chart of a method for detecting a fire target according to an embodiment of the present invention; Figure 2 is a schematic diagram of multimodal feature fusion according to an embodiment of the present invention; Figure 3 A schematic diagram of YOLOv8 of an improved neck network according to an embodiment of the present invention; Figure 4 Schematic diagram of the GSConv module of this embodiment; Figure 5 Schematic diagram of the V-GSCSP module of this embodiment; Figure 6 The figure is a schematic diagram of a fire target detection device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical scheme and advantages of the present invention clearer, the embodiments of the present invention will be described in detail in conjunction with the accompanying drawings hereinafter. It should be noted that, in the absence of conflict, the embodiments in this application and the features in the embodiments can be arbitrarily combined with each other. In order to better understand the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, other embodiments obtained by ordinary technicians in this field without making creative work are all within the scope of protection of the present invention.

[0018] Figure 1 FIG. 1 is a flow chart of a method for detecting a fire target according to an embodiment of the present invention. Figure 1 As shown, a method for detecting a fire target in this embodiment includes the following steps: Step S10: acquiring fire video data of multiple modes, grouping the acquired fire video data of multiple modes, and processing them into video frame sequences to form a fire sample data set; Fire video data is collected simultaneously through the visible camera module, near-infrared camera module, and thermal imaging camera module, and the fire video data is marked.

[0019] Among them, the camera module requires a minimum resolution of 1080×1080.

[0020] The occurrence of fireworks can be regarded as a behavior. The video sequence contains rich spatiotemporal action behavior information. Spatiotemporal dynamic features can better describe the characteristics of motion foreground than static features, thus providing more feature information for recognition and detection.

[0021] The acquired fire video data of multiple modes are preprocessed, and continuous multiple frames of visible, near infrared, and thermal imaging videos are acquired through the visible camera module, the near infrared camera module, and the thermal imaging camera module. The videos are divided into three groups, and 12 frames of pictures are intercepted from each group, with a resolution of 1080 × 1080. In other embodiments, 3 frames per second or 1 frame per second can be intercepted, etc., which can be intercepted according to actual needs.

[0022] Step S20: extracting features from the processed multi-modal fire video data using a three-dimensional convolutional neural network, and fusing the features into fire fusion features; In this embodiment, the three groups of multi-frame image data of visible, near-infrared and thermal imaging intercepted in step S10 are input into multiple three-dimensional convolutional neural networks, and the three modal features after the output of the three-dimensional convolutional neural network are then input into the three-dimensional convolutional neural network for feature extraction.

[0023] like Figure 2As shown in the figure, the three groups of images 12×1080×1080×3 obtained above are each subjected to 3DConv (3D-Convolutional Neural Networks, three-dimensional convolutional neural network), and the output features are visual features , near infrared characteristics , thermal imaging features .

[0024] The visual features obtained , near infrared characteristics , Thermal imaging features To perform splicing,

[0025] Get fire image features It is 3×640×640×128.

[0026] Fire image features Input into 3DConv to obtain the fire fusion feature F: 640×640×3.

[0027] Step S30: inputting the fire fusion features into the YOLOv8 target detection network of the improved neck network to perform fire target detection; In this embodiment, Figure 3 As shown in the figure, the YOLOv8 target detection network with improved neck network includes: backbone network layer (backbone), neck network layer (Neck), small target detection layer (Head), and outputs fire detection results. Specifically, it includes the following steps: Step S301: The fire fusion feature F: 640×640×3 obtained in step S20 is input into the backbone network layer, and passes through four stages of CBS module, C2f module and SPPF (Spatial Pyramid Pooling Faster) module.

[0028] Step S301-1: In the first stage, after the fire fusion feature F passes through a CBS module, the fire image feature extraction operation includes: passing through 1 CBS module, the CBS module of this embodiment includes Conv convolution, BN (BatchNormalization, batch normalization layer) and SiLU activation function, and the output image feature is 32×160×160. The image feature is input to the C2f module, and the C2f module passes the 32×160×160 fire image feature through the first convolution layer Conv1, and is divided into two branches. The first branch is passed to the output, and the second branch passes through 3 Bottleneck modules. The two branches are spliced ​​and passed through the second Conv2 to obtain the output. The output fire image feature is 32×160×160.

[0029] Step S301 - 2 : In the second stage, the operation of the first stage in step S301 - 1 is repeated to obtain the fire image feature P1 of 64×80×80.

[0030] Step S301 - 3 : In the third stage, step S301 - 2 is repeated to obtain the fire image feature P2 of 128×40×40.

[0031] Step S301-4: In the fourth stage, step S301-3 is repeated, and after passing through the SPPF module, the fire image feature P3 is output as 256×20×20.

[0032] Step S302: The multi-scale fire features P1: 64×80×80, P2: 128×40×40, and P3: 256×20×20 obtained in step S301 are input into the neck network layer at the same time. The neck network layer includes convolution, upsampling, splicing, V-GSCSP module and GSConv (lightweight convolution) module.

[0033] Step S3021: Figure 4 As shown in the figure, the GSConv module specifically includes spatial convolution (SC), depthwise separable convolution (DSConv), concatenation (Concat) and shuffle operations. The input features are convolved, and after a spatial convolution SC1, a depthwise separable convolution, and then a spatial convolution SC2, the features output by SC1 and SC2 are concat-concatenated and input into the spatial convolution SC3, and the features are output after the shuffle operation.

[0034] Step S3022: Figure 5As shown in Figure 1, the V-GSCSP module specifically includes convolution, splicing and GSConv modules. The input features are convolved to output two branches. The first branch passes through two GSConvs, and the second branch is spliced ​​with the output of the first branch and output after a convolution.

[0035] Step S303: The fire feature P3 is upsampled, concatenated with P2, and then sent to the V-GSCSP module described in step A3202 to obtain a feature map Q1 of size 128×40×40.

[0036] Step S304: The feature map Q1 obtained in step S303 is upsampled, concatenated with P1, and sent to the V-GSCSP module described in step A3202 to obtain a feature map Q2 of size 192×80×80.

[0037] Step S305: The fire feature P3, the fire feature graphs Q1 and Q2 obtained in steps 303 and 304 are continuously input to the next layer of the Neck module. Q2 is input to the GSConv module described in step S3021 and concatenated with Q1, and then input to the V-GSCSP module described in step S3022 to output a feature graph J1 of size 192×40×40.

[0038] Step S306: Input the feature map J1 obtained in step S305 into the GSConv module described in step S3021, concatenate it with P3 and input it into the V-GSCSP module described in step S3022, and output a feature map J2 of size 64×80×80.

[0039] Step S307: Send the fire characteristic images Q2, J1 and J2 obtained in steps S301-S306 to the Head module respectively to obtain fire detection results of small size 20×20, medium size 40×40 and large size 80×80.

[0040] Step S308: Integrate the fire detection results of the Head module in step S307 and output the coordinates of the fire boundary box that is finally detected.

[0041] In this embodiment, the input end integrates three types of video acquisition modules: visible, near-infrared, and thermal imaging. The data processing end is an embedded development board. After the model is trained, the weight parameters are saved and deployed on the embedded development board. The output end is a screen, and the alarm signal is displayed on the screen.

[0042] The embodiment of the present invention also provides a corresponding fire target detection device, such as Figure 6 As shown, the device of this embodiment includes: A video acquisition unit, used for acquiring fire video data of multiple modes; A processing unit, used for grouping the multimodal fire video data and processing them into video frame sequences; used for extracting features from the processed multimodal fire video data using a three-dimensional convolutional neural network, and fusing them into fire fusion features; The detection unit is used to input the fire fusion feature into the YOLOv8 target detection network of the improved neck network to perform fire target detection. The YOLOv8 target detection network of the improved neck network includes a backbone network layer, a neck network layer and a small target detection layer.

[0043] Preferably, the detection unit inputs the fire fusion feature into the YOLOv8 target detection network of the improved neck network for fire target detection, including: inputting the fire fusion feature into the backbone network layer, and outputting multi-scale fire image features after passing through four stages of CBS module, C2f module and SPPF module in the backbone network layer; inputting the multi-scale fire image features into the neck network layer, and outputting multiple fire feature maps after passing through convolution, upsampling, splicing, GSConv module and V-GSCSP module in the neck network layer, wherein the GSConv module is used to perform a convolution operation on the input features, and after the first spatial convolution, the depthwise separable convolution, and then the second spatial convolution, the features output by the first spatial convolution and the second spatial convolution are spliced ​​and input into the third spatial convolution, and after Shuffle, the features are output. After the operation, feature output is performed, the V-GSCSP module is used to convolve the input features and output two branches, the first branch passes through two GSConv modules, and the second branch is spliced ​​with the output of the first branch and then passes through a convolution for feature output; the multiple fire feature maps are input into the small target detection layer, and multiple fire detection result maps of different sizes are output in the small target detection layer through multiple detection modules; the coordinate information of the fire boundary box is output according to the multiple fire detection result maps of different sizes.

[0044] Preferably, the detection unit outputs multi-scale fire image features after passing through four stages of CBS module, C2f module and SPPF module in the backbone network layer, including: the first stage: after the fire fusion feature is input into a CBS module, it enters the fire image feature extraction operation including: after passing through a CBS module, it enters the C2f module, in the C2f module, it is divided into two branches after passing through the first convolution layer, the first branch is transmitted to the output, and the second branch passes through three bottleneck modules, the two branches are spliced, and then pass through the second convolution layer and then output the fire image feature; the second stage: the fire image feature extraction operation is performed on the fire image feature output in the first stage, and the first fire image feature P1 is output; the third stage: the fire image feature extraction operation is performed on the fire image feature output in the second stage, and the second fire image feature P2 is output; the fourth stage: the fire image feature extraction operation is performed on the fire image feature output in the third stage, and then the third fire image feature P3 is output after passing through the spatial pyramid pooling fast module.

[0045] Preferably, the detection unit inputs the multi-scale fire image features into the neck network layer, and outputs multiple fire feature maps after the neck network layer passes through convolution, upsampling, splicing, GSConv module and V-GSCSP module, including: after the third fire image feature P3 is upsampled, the splicing operation is performed with the second fire image feature P2, and then the input is into the V-GSCSP module to obtain the first feature map Q1; after the first feature map Q1 is upsampled, the splicing operation is performed with the first fire image feature P1, and then the input is into the V-GSCSP module to obtain the second feature map Q2; after the second feature map Q2 is input into the GSConv module, the splicing operation is performed with the first feature map Q1, and then the input is into the V-GSCSP module to obtain the first fire feature map J1; after the first fire feature map J1 is input into the GSConv module, the splicing operation is performed with the third fire image feature P3, and then the input is into the V-GSCSP module to obtain the second fire feature map J2.

[0046] In summary, the embodiment of the present invention discloses a method and device for fire target detection based on visual, near-infrared, and thermal imaging videos of three-dimensional convolutional neural networks, which can extract multimodal features of fire images, including visual, near-infrared, thermal imaging, and dynamic features, and improve the YOLOv8 target network structure. The target detection method through the fusion of fire multimodal features can effectively improve the accuracy of fire detection. This method has high application value in the field of fire detection in large spaces, warehouses, terminals, nine small places, and petrochemical parks, and greatly solves the problems of false alarms and missed alarms in existing image fire detection.

[0047] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any particular form of combination of hardware and software.

[0048] The above are only preferred embodiments of the present invention. Of course, the present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for detecting a fire target, comprising: Collecting fire video data of multiple modes, grouping the fire video data of multiple modes and processing them into video frame sequences; The 3D convolutional neural network is used to extract features from the processed multi-modal fire video data and fuse them into fire fusion features; The fire fusion feature is input into the YOLOv8 target detection network of the improved neck network for fire target detection, wherein the YOLOv8 target detection network of the improved neck network comprises a backbone network layer, a neck network layer and a small target detection layer.

2. The method according to claim 1, characterized in that: The step of inputting the fire fusion feature into the YOLOv8 target detection network of the improved neck network for fire target detection includes: The fire fusion features are input into the backbone network layer, and multi-scale fire image features are output after passing through four stages of CBS module, C2f module and SPPF module in the backbone network layer; Input the multi-scale fire image features into the neck network layer, and output multiple fire feature maps after the neck network layer passes through convolution, upsampling, splicing, GSConv module and V-GSCSP module, wherein the GSConv module is used to perform a convolution operation on the input features, and after a first spatial convolution, a depth-separable convolution, and then a second spatial convolution, the features output by the first spatial convolution and the second spatial convolution are spliced ​​and input into the third spatial convolution, and the features are output after a Shuffle operation, and the V-GSCSP module is used to perform a convolution operation on the input features to output two branches, the first branch passes through two GSConv modules, and the second branch is spliced ​​with the output of the first branch and then passes through a convolution to output the features; Inputting the plurality of fire characteristic images into a small target detection layer, and outputting a plurality of fire detection result images of different sizes through a plurality of detection modules in the small target detection layer; The coordinate information of the fire boundary box is outputted according to the multiple fire detection result images of different sizes.

3. The method according to claim 2, characterized in that: The backbone network layer outputs multi-scale fire image features after passing through four stages of CBS module, C2f module and SPPF module, including: The first stage: after the fire fusion feature is input into a CBS module, the fire image feature extraction operation includes: after passing through a CBS module, entering the C2f module, in the C2f module, it is divided into two branches after passing through the first convolution layer, the first branch is passed to the output, and the second branch passes through three bottleneck modules, after the two branches are spliced, it passes through the second convolution layer and then outputs the fire image feature; The second stage: performing the fire image feature extraction operation on the fire image feature outputted in the first stage, and outputting the first fire image feature P1; The third stage: performing the fire image feature extraction operation on the fire image feature outputted in the second stage, and outputting the second fire image feature P2; The fourth stage: performing the fire image feature extraction operation on the fire image feature outputted in the third stage, and then outputting the third fire image feature P3 after passing through the spatial pyramid pooling fast module.

4. The method according to claim 3, characterized in that: The multi-scale fire image features are input into the neck network layer, and after the neck network layer undergoes convolution, upsampling, splicing, GSConv module and V-GSCSP module, multiple fire feature maps are output, including: After the third fire image feature P3 is upsampled, it is concatenated with the second fire image feature P2 and then input into the V-GSCSP module to obtain the first feature map Q1; After the first feature map Q1 is upsampled, it is concatenated with the first fire image feature P1 and input into the V-GSCSP module to obtain the second feature map Q2; The second feature map Q2 is input into the GSConv module and then concatenated with the first feature map Q1, and then input into the V-GSCSP module to obtain the first fire feature map J1; The first fire feature map J1 is input into the GSConv module, concatenated with the third fire image feature P3, and then input into the V-GSCSP module to obtain the second fire feature map J2.

5. The method according to claim 4, characterized in that: The method of inputting the plurality of fire characteristic images into a small target detection layer, and outputting a plurality of fire detection result images of different sizes through a plurality of detection modules in the small target detection layer, comprises: The second feature map Q2, the first fire feature map J1 and the second fire feature map J2 are respectively input, and respectively pass through the corresponding detection modules in the small target detection layer to output a small-size fire detection result map, a medium-size fire detection result map and a large-size fire detection result map.

6. The method according to any one of claims 1 to 5, characterized in that: The method of using a three-dimensional convolutional neural network to extract features from the processed multi-modal fire video data and fuse them into fire fusion features includes: The processed multi-modal fire video data are respectively passed through a three-dimensional convolutional neural network to output corresponding features, and after a splicing operation is performed on the output corresponding features, they are input into the three-dimensional convolutional neural network to obtain the fire fusion features.

7. A fire target detection device, characterized in that: include: A video acquisition unit, used for acquiring fire video data of multiple modes; A processing unit, used for grouping the multimodal fire video data and processing them into a video frame sequence; It is used to extract features from processed multi-modal fire video data using a three-dimensional convolutional neural network and fuse them into fire fusion features; The detection unit is used to input the fire fusion feature into the YOLOv8 target detection network of the improved neck network to perform fire target detection. The YOLOv8 target detection network of the improved neck network includes a backbone network layer, a neck network layer and a small target detection layer.

8. The device according to claim 7, characterized in that: The detection unit inputs the fire fusion feature into the YOLOv8 target detection network of the improved neck network for fire target detection, including: inputting the fire fusion feature into the backbone network layer, and outputting multi-scale fire image features after passing through four stages of CBS module, C2f module and SPPF module in the backbone network layer; inputting the multi-scale fire image features into the neck network layer, and outputting multiple fire feature maps after passing through convolution, upsampling, splicing, GSConv module and V-GSCSP module in the neck network layer, wherein the GSConv module is used to perform convolution operation on the input features, and after passing through the first spatial convolution, depth-separable convolution, and then the second spatial convolution, the features output by the first spatial convolution and the second spatial convolution are spliced ​​and input into the third spatial convolution, and after Shuffle, the features are output. After the operation, feature output is performed, the V-GSCSP module is used to convolve the input features and output two branches, the first branch passes through two GSConv modules, and the second branch is spliced ​​with the output of the first branch and then passes through a convolution for feature output; the multiple fire feature maps are input into the small target detection layer, and multiple fire detection result maps of different sizes are output in the small target detection layer through multiple detection modules; the coordinate information of the fire boundary box is output according to the multiple fire detection result maps of different sizes.

9. The device according to claim 8, characterized in that: The detection unit outputs multi-scale fire image features after passing through four stages of CBS modules, C2f modules and SPPF modules in the backbone network layer, including: the first stage: after the fire fusion feature is input into a CBS module, it enters the fire image feature extraction operation including: after passing through a CBS module, it enters the C2f module, in the C2f module, it passes through the first convolution layer and is divided into two branches, the first branch is transmitted to the output, and the second branch passes through three bottleneck modules, the two branches are spliced, and then pass through the second convolution layer and then output the fire image feature; the second stage: the fire image feature extraction operation is performed on the fire image feature output in the first stage, and the first fire image feature P1 is output; the third stage: the fire image feature extraction operation is performed on the fire image feature output in the second stage, and the second fire image feature P2 is output; the fourth stage: the fire image feature extraction operation is performed on the fire image feature output in the third stage, and then the third fire image feature P3 is output after passing through the spatial pyramid pooling fast module.

10. The device according to claim 9, characterized in that: The detection unit inputs the multi-scale fire image features into the neck network layer, and outputs multiple fire feature maps after the neck network layer passes through convolution, upsampling, splicing, GSConv module and V-GSCSP module, including: after the third fire image feature P3 is upsampled, the splicing operation is performed with the second fire image feature P2, and then the map is input into the V-GSCSP module to obtain the first feature map Q1; after the first feature map Q1 is upsampled, the splicing operation is performed with the first fire image feature P1, and then the map is input into the V-GSCSP module to obtain the second feature map Q2; after the second feature map Q2 is input into the GSConv module, the splicing operation is performed with the first feature map Q1, and then the map is input into the V-GSCSP module to obtain the first fire feature map J1; after the first fire feature map J1 is input into the GSConv module, the splicing operation is performed with the third fire image feature P3, and then the map is input into the V-GSCSP module to obtain the second fire feature map J2.

Citation Information

Cited By

  • Steel wire rope defect detection method, steel wire rope defect detection device, medium and equipment

    CN121033003A

  • Fire smoke image detection method based on improved YOLOv8 network

    CN121095735A