Flame detection method based on depth model, flame detector and storage medium

By building a flame detector based on the depth model, using IR sensors to collect data and build a flame depth network model, the false alarm problem of traditional flame detectors in high temperature environments is solved, and accurate flame recognition and alarm are achieved.

CN120541783APending Publication Date: 2025-08-26河南驰诚电气股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510679507.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing flame detectors are prone to false alarms in places with high temperature and heat radiation, and traditional neural network models are complex and require a large hardware resource, so they cannot accurately identify flame tag types and time ranges, resulting in large identification errors.

Method used

The flame detection method based on the depth model is adopted to build a flame depth network model through IR sensor data acquisition, including a cascading backbone network, a feature extraction fusion network and a decoupling regression network, and combined with the loss function to optimize the model parameters, output flame confidence, label type and location information.

Benefits of technology

It improves the accuracy and anti-interference ability of flame detection, can accurately identify the number, range and duration of flame targets, and provides accurate alarm and emergency response reference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541783A_ABST
    Figure CN120541783A_ABST
Patent Text Reader

Abstract

The invention discloses a flame detection method based on a depth model, a flame detector and a storage medium, and relates to the field of flame detection, and the method comprises the steps: collecting IR data of different flames and interference source signals, and carrying out the data augmentation; driving a flame depth network model based on the flame label type and the target position data; flame confidence, label type confidence and target position loss are respectively calculated according to the model predicted value and label and position marking values in the training sample set, and total loss of the model is constructed based on various losses; model parameters are updated based on total loss function value back propagation, flame confidence, type confidence and a position regression value are output according to sampled multi-channel IR data after iteration model training is completed, and a final fusion result is output based on a fusion strategy. According to the model design scheme, a multi-scale feature extraction and fusion strategy is adopted, information such as types and positions of flames is recognized in cooperation with multi-level feature extraction and a multi-feature graph, and accurate classification and early warning of flame detection are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of flame detection, and in particular to a flame detection method, a flame detector, and a storage medium based on a depth model. Background Art

[0002] Flame detectors based on pyroelectric infrared sensors are already widely used. A common detection method involves detecting infrared radiation signatures, converting received infrared radiation into current or voltage signals. The signals are then analyzed for flicker frequency, energy, continuity, and rate of change, as well as correlations between signals from different channels. These characteristics are then used to determine whether the signals meet flame characteristics. While this method can detect flames to a certain extent, it can suffer from significant errors, particularly in locations with high temperatures and thermal radiation, where external interference can easily cause false alarms.

[0003] In recent years, some approaches have combined Fourier transforms to extract frequency domain features, then used BP neural network models for classification and prediction to determine whether a signal is a flame, thereby achieving fire detection. Other approaches have combined LSTM (Long Short-Term Memory) networks and GRU (Gated Recurrent Unit) networks to improve flame prediction accuracy and interference resistance. While these traditional approaches have achieved some success, their relatively complex network structures require more hardware computing resources. The use of continuous alarm threshold strategies also suffers from poor fitting. Low thresholds weaken false alarm immunity, while high thresholds degrade real-time alarm performance. Importantly, these prediction network models cannot effectively identify flame label types and time ranges, hindering accurate classification and early warning. With the advancement of artificial intelligence, models built on deep learning frameworks can further improve the classification and prediction accuracy of flame detectors and achieve more refined predictions. Summary of the Invention

[0004] This application provides a flame detection method, flame detection lamp and storage medium based on a deep model, which further improves the prediction accuracy and realizes refined prediction compared with the traditional neural network model.

[0005] In one aspect, the present application provides a flame detection method based on a depth model, the method comprising: The IR data of different flame and interference source signals are collected through various IR sensors, and the flame IR data is used as the target and the interference source IR data is used as the background for fusion to expand the training set and test set; According to the flame label type and target position data contained in the IR data in the training set, a flame deep network model based on target classification and position regression is constructed and driven; Based on the model prediction value, the label mark value and the position mark value in the training sample set, the flame confidence loss, flame label type confidence loss and target position loss are calculated respectively, and the total model loss is constructed based on various losses; The model parameters are updated based on the back propagation of the total loss function value of the model. After the iterative model training is completed, the flame confidence, flame label type confidence and position regression value are output according to the sampled multi-channel IR data, and the final fusion result is output based on the fusion strategy.

[0006] Specifically, the Flame deep network model includes a cascaded backbone network, a first feature extraction and fusion network, a second feature extraction and fusion network, and a decoupled regression network; The backbone network includes a cascaded spatial pyramid pooling layer SPPF module and a focus module. The input IR feature data is captured by the SPPF module under different field of view sizes, and then the Focus module performs fast slicing and downsampling. After multiple downsampling and feature extraction, it outputs a multi-scale feature map. The first feature extraction and fusion network includes a horizontal connection network and a first vertical fusion network. The horizontal connection network part inputs the multi-scale feature map, and the first vertical fusion network cascades the output of the backbone network to successively convert the high-level feature map into the low-level feature map, and matches and fuses the multi-scale feature map with the low-level feature map to output a multi-scale first fused feature map with flame type information. The second feature extraction and fusion network includes a horizontal connection network and a second vertical fusion network. The horizontal connection network part inputs the first fused feature map of the corresponding scale. The second vertical fusion network cascades the output of the first vertical fusion network, converts the low-level feature map into a high-level feature map one by one, and matches and fuses the first fused feature map of the corresponding scale with the high-level feature map to output a multi-scale second fused feature map with position information. The decoupling regression network sets up several decoupling heads at different scales. Each decoupling head inputs the second fused feature map of the corresponding scale through a horizontal connection network, predicts the type information and target position information contained therein through classification and regression, and outputs the flame label type, flame number and flame position.

[0007] Specifically, three groups of CBL modules and C2 feature extraction modules are alternately cascaded after the Focus module of the backbone network; the CBL modules downsample layer by layer, and the C2 feature extraction modules extract features layer by layer. Each C2 feature extraction module outputs a corresponding scale feature map and sends it to the first vertical fusion network through the horizontal connection network.

[0008] The first feature extraction and fusion network includes three groups of alternately cascaded CBL modules, upsampling modules, channel splicing units, and C2 feature extraction modules; the channel splicing unit performs channel splicing on the high-level feature map output by upsampling and the corresponding scale feature map, and the CBL module outputs the first fused feature map at the corresponding scale; The second feature extraction fusion network includes three groups of alternately cascaded CBL modules, channel stitching units and C2 feature extraction modules; the channel stitching unit performs channel stitching on the first fused feature map and the feature map output by the CBL module, and derives a second fused feature map of corresponding scale from each CBL module input and C2 feature extraction module output through a horizontal connection network.

[0009] Specifically, the decoupling head includes three parallel decoupling networks, namely a flame target decoupling network, a flame tag type decoupling network and a position decoupling network; The flame target decoupling network and the flame label type decoupling network respectively contain a cascaded CBL module, a one-dimensional convolution module, and a Sigmoid activation layer, while the position decoupling network contains a cascaded CBL module and a one-dimensional convolution module. The number of grids of the input feature maps of the three parallel decoupling networks is , confidence decoupling network output Flame label confidence, flame type decoupling network output The flame type confidence, position decoupling network output regression Flame target location information; The three parallel decoupling networks are fused through the channel splicing unit to output the flame confidence, flame label type, and flame position information of all grids at different scales.

[0010] Specifically, the SPPF module inputs the original IR feature data and captures multi-scale feature information through three sequentially cascaded pooling structures of different sizes. The output of each pooling structure is fused with the original IR feature data through a channel splicing unit, and the output does not change the feature size. The Focus module is connected to the output of the SPPF module, performs feature slicing through 64 parallel Slice slicing units, and performs fast downsampling through channel splicing.

[0011] Specifically, the CBL module includes a cascaded one-dimensional convolution module Conv1d, a normalization layer and an activation layer; The C2 feature extraction module includes a cascaded CBL module and N bottleneck modules. A residual connection is set between the CBL module and the bottleneck module. The residual connection part is the CBL module. The output of the bottleneck module and the residual connection output are fused and output through the channel splicing unit, and the output does not change the feature size. The bottleneck module consists of two cascaded CBL modules and a directly connected residual structure is set between the cascaded CBL modules, and the output is superimposed with the cascaded part; the two cascaded CBL modules halve the channels and restore the channels in turn, and the superimposed output does not change the feature size.

[0012] Specifically, the loss calculation process is as follows:

[0013]

[0014]

[0015]

[0016] Among them is the total number of labeled positive and negative samples, is the number of categories, represents the flame confidence loss, represents the confidence loss of flame type, represents the target position loss, Represents the total loss function of the model; is the target number, is the flame target mark value, is the flame confidence prediction value, is the flame tag type tag value, is the flame label type confidence prediction value, is the mark value of the flame target center position, is the predicted value of the flame target center position, is the flame target width mark value, is the flame target width prediction value; 、 and is an adjustable weight coefficient.

[0017] Specifically, outputting the fused flame type and flame position coordinates according to the sampled IR data includes: Obtain the flame tag and determine the flame target center position and flame width in the IR time series based on the flame target position; Identify multiple flame targets with the same flame category and overlapping target positions in the IR time series, and fuse them to update the flame category, number, and target position based on the target position; When at least two flames are of different types and have overlapping parts, they are directly output without fusion; When at least two flames are of the same type and have overlapping parts, the number of flames, target center position and width are fused. The fused and updated flame position is expressed as follows:

[0018] Among them and Indicates the updated flame center coordinates and flame width; and Respectively represent the starting coordinates of the near and far flame targets in the time series, and Indicates the width of near and far flame targets.

[0019] On the other hand, the present application provides a flame detector, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the deep model-based flame detection method described in any of the above aspects.

[0020] On the other hand, the present application provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set. The at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the flame detection method based on the deep model described in any of the above aspects.

[0021] The beneficial effects brought about by this application are: using a backbone network, a first feature extraction and fusion network with an FPN structure, and a second feature extraction and fusion network with a PAN structure to extract deep and shallow feature data of the feature map in layers and stages, fusing deep and shallow features in a multi-scale feature fusion manner, training flame targets and types with sample sets, and then regressing and predicting flame size and position information through path aggregation. Finally, through multiple decoupling heads, the label classification problem and the position regression problem are combined and output, which improves the accuracy of flame detection while providing a more intelligent output. The model output results use a fusion strategy to identify the number and range of flame targets, update the number of flames and the size of flames, and combine with the sampling input mechanism to accurately identify the scale, location and duration of environmental flames, providing an effective reference for subsequent accurate alarms and emergency disposal. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a hardware structure diagram of the flame detector provided in an embodiment of the present application; Figure 2 This is a flow chart of a flame detection method based on a mechanism model provided in an embodiment of the present application; Figure 3 This is a schematic diagram of grid sampling of flame detectors; Figure 4 It is an example diagram for setting positive and negative samples; Figure 5 This is a schematic diagram of the flame deep network model structure based on flame recognition; Figure 6 It is a schematic diagram of the structure of the CBL 1d convolution block; Figure 7 It is a detailed network structure diagram of the decoupling head; Figure 8 This is a schematic diagram of the SPPF1d structural design of this application; Figure 9 This is a schematic diagram of the structure of the Focus module; Figure 10 Schematic diagram of the structure of C2 1d convolution block is shown; Figure 11 It is a structural diagram of the bottleneck module; Figure 12 The IR characteristic spectrum of the four-channel flame target that may be labeled output is listed; Figure 13 This is a signal diagram of the Hanning window; Figure 14 This is a schematic diagram of the IR signal after adding the interference source; Figure 15 It is a schematic diagram before and after data augmentation; Figure 16 This is a schematic diagram of the SPPF1d+Focus1d structure built by Pytorch; Figure 17 This is a schematic diagram of the decoupled head output structure built by Pytorch; Figure 18 Schematic diagram of the precision and recall rate shown in the test in the embodiment; Figure 19 This is an example graph of IoU calculated according to the evaluation prediction scheme. DETAILED DESCRIPTION

[0023] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0024] In this document, "plurality" refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0025] Flame detectors are equipped with several IR sensors depending on their size and purpose, such as Figure 1The flame detector hardware structure shown in the figure uses a four-channel IR sensor in its IR hardware circuit. This four-channel IR sensor uses different filters to capture IR signals of different bands and wavelengths. A properly functioning flame detector collects IR analog current signals in real time. After passing them through an operational amplifier circuit, a filter circuit, and an A / D converter, it generates discrete signals of infrared radiation at different wavelengths within the detector's field of view. These signals are then analyzed and processed using the flame detection algorithm provided in this application.

[0026] Figure 2 This is a flow chart of a flame detection method based on a mechanism model provided in an embodiment of the present application, comprising the following steps: S1. Collect IR data of different flame and interference source signals through various IR sensors, and fuse the flame IR data as the target and the interference source IR data as the background to expand the training set and test set. During the data acquisition and monitoring phase, the flame detector must first collect raw IR data. This application utilizes a refined grid sampling method within the flame detector's effective detection volume. The flame detector can adjust its pitch, rotation, and detection distance as needed to capture infrared data within different fields of view. Figure 3 This is a diagram of a flame detector's grid sampling. The red dots represent the grid sampling locations. Combined with the flame detector's varying pitch angles, it can detect flames and interference sources from all directions within the effective detection space, creating a relatively complete sample set. This comprehensive sample set enables comprehensive IR signal analysis and improves model generalization performance, significantly enhancing the flame detector's alarm accuracy and false alarm immunity.

[0027] During the acquisition process, different interference source signals can be introduced according to actual needs to collect IR data of various interference sources. The interference source can be various lights, such as introducing incandescent lamps or flame test lamps, etc., and integrating them into the IR data to increase sample diversity.

[0028] In one possible implementation, this application uses a combination of positive and negative samples for data augmentation. Specifically, the grid size and the flame targets included in the sample feature map are marked, and the size and position of the flame targets within the sample feature map are identified. Next, a target grid in the sample feature map that contacts the flame target is defined. The overlap width between the target grid and the flame target is determined, and the sample type is determined based on the target ratio of the flame overlap width to the flame width.

[0029] Figure 4 This is an example diagram for setting positive and negative samples. Since IR data needs to extract feature maps, this application uses The sample feature map is represented by the grid size of , which includes grids 1 to 4. For samples containing flame targets, the flame target will overlap with at least one grid. Assuming grid 3 is the target grid, the flame target will overlap with grid 3 by a width w1. The width of the flame itself is W. This can be calculated using the following formula:

[0030] The value represents the target ratio.

[0031] The embodiment of the present application stipulates that when When , it is marked as a positive sample, otherwise ( ) is defined as a negative sample. The ratio of positive and negative samples during training is After marking one positive sample, four negative samples will be marked near the positive sample from closest to farthest away. Positive and negative samples can effectively increase the sample size and maintain diversity, which greatly helps with recognition accuracy.

[0032] In the specific fusion process, a signal of a fixed length can be randomly intercepted from the interference source sample as the background, and then one or more signals within a certain length range can be randomly intercepted from the flame sample as the target, and then fused to the background signal at a random position.

[0033] S2. Based on the flame label type and target position data contained in the IR data in the training set, a flame deep network model based on target classification and position regression is constructed and driven; The flame label and interference source label are set according to actual needs, such as setting match flame, alcohol flame, n-heptane flame, etc.; and the interference source can be set to incandescent lamp, LED, mercury lamp and halogen lamp, etc. The target position here refers to the position of the flame target in the feature map, corresponding to Figure 4 The flame deep network model constructed in this solution is designed to output multi-parameter results. For IR data containing a flame target, after model testing, it outputs parameters such as flame type, flame location, flame size (width), and flame duration (determined by width and IR duration). This AI deep model can improve flame detector performance and meet the needs of refined prediction of flame type and duration.

[0034] S3. Calculate the flame confidence loss, flame type confidence loss, and target position loss based on the model prediction value, as well as the label mark value and position mark value in the training sample set, and construct the total model loss based on various losses. The flame confidence loss is used to evaluate the accuracy of identifying flame targets from sample feature maps or actual test process feature maps, while the flame label type confidence loss is used to evaluate the accuracy of identifying the flame type. The target position loss is used to determine the accuracy of the flame's position in the IR time series. These three dimensions are the parameters that actually need to be output, so the (weighted) sum of the three losses is the final total model loss. Considering that the flame target and flame type confidence predictions are both 0-1 problems, binary cross entropy (non-cross entropy) is used. The constructed loss function is calculated as follows:

[0035]

[0036]

[0037] in is the total number of positive and negative samples marked, and negative samples do not calculate position loss, is the number of categories, is the target number, is the flame target mark value, is the flame confidence prediction value, is the flame tag type tag value, is the flame label type confidence prediction value, is the mark value of the flame target center position, is the predicted value of the flame target center position, is the flame target width mark value, is the flame target width prediction value; According to the above three loss functions, combined with the functional requirements in the actual project, the total loss function of the model is obtained , which is expressed as follows:

[0038] 、 and is an adjustable weight coefficient.

[0039] S4. During training, the model parameters are updated based on the backpropagation of the total loss function value of the model. After the iterative model training is completed, the flame confidence, flame label type confidence and position regression value are output according to the sampled multi-channel IR data, and the final fusion result is output based on the fusion strategy.

[0040] The iterative process performs multiple rounds of iterative updates based on the previously produced sample sets, training sets, test sets and other data, and continuously adjusts the model parameters until the various result parameters of the training output meet expectations.

[0041] It's important to note that the output feature map may identify more than one flame target, particularly when multiple flames overlap. In this case, target fusion techniques are required to fuse the output. For example, if there are two overlapping flames (regardless of flame type), this is likely due to segmented recognition caused by temperature and interference, while the image is actually a single, complete flame. Alternatively, wind could cause a flame to jump, misidentified as two separate flames. These situations also require flame fusion output. Fusion can be performed based on the overlapping areas, and the number, size, location, and range of the flame targets need to be redefined based on the fusion rules.

[0042] In the embodiments of this application, Figure 5 This is a schematic diagram of the flame recognition deep network model architecture. The model consists of a cascaded backbone network, a primary feature extraction and fusion network, a secondary feature extraction and fusion network, and a decoupled regression network. The backbone network employs a feature pyramid design principle, generating IR feature pyramids of varying resolutions based on the input. This allows high-resolution feature maps to focus on detailed features, while low-resolution feature maps focus on global features.

[0043] The first and second feature extraction fusion networks correspond to the two measurement dimensions of flame target and position, respectively, and together form the Neck network. The first feature extraction fusion network extracts categorical features, while the second feature extraction network extracts positional features. Categorical features are used to identify flame targets and corresponding flame labels. Positional features are then extracted for position regression prediction, determining information such as flame location, size, and range. Finally, the data determined by these dimensions is integrated and decoupled through a decoupled regression network (i.e., the Head network) for output.

[0044] The following is a detailed introduction to the composition structure of each sub-network: 1. The Backbone network includes a cascaded spatial pyramid pooling layer SPPF module and a focus module. The input IR feature data (this embodiment takes 4-channel IR data as an example, i.e. 4 1024 feature size) uses the SPPF module to capture multi-scale feature information in different fields of view. Because 4-channel IR data is less complex than images, we directly use the Focus1d module (all modules in the figures use a 1D structure, representing one-dimensional IR spectrum data) to quickly reduce the IR signal width while maintaining the IR signal characteristics by increasing the dimension.

[0045] The Backbone network's Focus module is followed by three alternate cascades of CBL modules and C2 (specifically, a 1D structure, or C21d) feature extraction modules. The CBL (specifically, a 1D structure, or CBL 1d) convolutional blocks perform layer-by-layer downsampling, while the C2 feature extraction modules extract features layer by layer, outputting a corresponding scale feature map at each C2 feature extraction module.

[0046] SPPF module output 16 The 1024-size feature map is then quickly sliced ​​and downsampled through the Focus module to output a multi-scale feature map. Figure 5 Each C2 feature extraction module output constructs 1024 8-size feature map, 2048 4-size feature map, 4096 2 size feature map, and 8192 1-size feature map and fed into the first vertical fusion network through the horizontal connection network.

[0047] exist Figure 5 In the overall network structure, the structure of CBL modules is the same, but the specific parameters are different. The specific parameters change according to the level. Figure 6 As shown in the figure, the CBL module specifically includes a cascaded one-dimensional convolution module Conv1d, a normalization layer BN1d and an activation layer Leaky reru.

[0048] Conv1d: In the one-dimensional convolution kernel, k is the convolution kernel width (k=1), s is the step size, p is the number of edge padding, and c is the number of output channels. The convolution module of 3 is used to extract IR features, 1 1 Convolution module is used to change the feature dimension, and downsampling is achieved with step size s. These parameters also correspond to different parameters of CBL 1d. For specific parameters, see Figure 5 , which will not be described in detail here.

[0049] BN1d: The benefit of adding BN layer is that it can improve the convergence speed and prediction accuracy of the model. The BN layer introduces learnable parameters during training. and , learning is performed based on the mean and variance of each training batch of data. During forward inference, the data is translated and scaled at the BN location to achieve standardization of the input data; the training input and output The calculation formula is expressed as follows:

[0050] Leaky reru: Compared with the sigmoid activation function, Leaky reru has the advantages of being more suitable for deep networks, high computational efficiency, and avoiding gradient disappearance; it is expressed as follows:

[0051] 2. The Neck-FPN network for first feature extraction and fusion consists of a horizontal connection network and a first vertical fusion network. The horizontal connection network inputs the multi-scale feature map and forms a branching structure connecting the backbone network and the first vertical fusion network. The first vertical fusion network itself cascades the output of the last C2 convolutional block of the backbone network, successively converting high-level feature maps into low-level feature maps. In the horizontal connection network, the multi-scale feature map is matched and fused with the low-level feature map, outputting a multi-scale first fused feature map containing flame type information.

[0052] Similar to the above, the first vertical fusion network consists of three sets of alternating cascade settings: CBL modules, upsampling modules, channel concatenation units, and C2 feature extraction modules. The channel concatenation unit matches and fuses the high-level feature maps output by the upsampling with the corresponding scale feature maps output by the backbone network. The CBL module outputs the first fused feature map at the corresponding scale, that is, 4096 1 size feature map, 2048 2-size feature map, 1024 4 size feature maps, and 8192 1Size feature diagram.

[0053] The purpose of this network structure is to fuse the deep IR global features and shallow IR local features extracted by the network to improve the accuracy of flame target classification, flame label type classification and flame target position regression.

[0054] 3. The structure of the second feature extraction and fusion Neck-PAN network is similar to that of the category feature extraction Neck-FPN network, consisting of a horizontal connection network and a second vertical fusion network. The horizontal connection network inputs the first fused feature map of the corresponding scale. The second vertical fusion network cascades the output of the last C2 feature extraction module of the first vertical fusion network, successively converting low-level feature maps into high-level feature maps. The first fused feature map of the corresponding scale is matched and fused with the high-level feature map, outputting a multi-scale second fused feature map with target location information.

[0055] The second feature extraction fusion network is oriented in the opposite direction to the first. This structure primarily aggregates paths, further strengthening the fusion of IR features between different layers of the network based on Neck-FPN. The second vertical fusion network consists of three alternating cascaded CBL modules, channel concatenation units (Concat), and C2 feature extraction modules. Concat concatenates the first fused feature map with the feature map output by the CBL module. At each CBL module input and C2 feature extraction module output, a second fused feature map of the corresponding scale is generated through a horizontal connection network.

[0056] Since four feature size outputs are used, the final output of the first vertical fusion network is not only used as the starting input of the second vertical fusion network, but also directly used as the output of the second fusion feature map of the first layer, that is, 2048 8-size feature map, and the rest are output 2048 in each C2 feature extraction module according to the network depth. 4-size feature map, 4096 2 size feature map, and 8192 1Size feature diagram.

[0057] 4. The decoupling regression network sets up several decoupling heads according to the scale. According to the structure matching the previous stage, there are four decoupling heads. Each decoupling head inputs the second fusion feature map of the corresponding scale through the horizontal connection network, extracts the label information and target position information contained therein, and outputs the flame label type, number of flames and flame position.

[0058] This decoupling head uses an anchor-free design. The advantage of the decoupling head is that it uses multiple models for flame classification and regression, improving the accuracy of each prediction. Because this model needs to identify flames, flame label types, and positions, the decoupling head is designed to include three parallel decoupling networks: a flame confidence decoupling network, a flame label type decoupling network, and a position decoupling network.

[0059] The flame confidence decoupling network and the flame label type decoupling network each consist of a cascaded CBL module, a one-dimensional convolution kernel, and a sigmoid activation layer. The position decoupling network also consists of a cascaded CBL module and a one-dimensional convolution kernel. These three parallel decoupling networks are fused through a channel concatenation unit to output flame confidence (i.e., number of flames), flame label type, and flame position information at the corresponding scale.

[0060] Figure 7 This is a detailed diagram of the network structure of the decoupling head. Assuming that the feature map grid width is , the number of flame types is C, then the AI ​​model decoupling head outputs Flame confidence (when the flame detector does not distinguish between flame type and position alarm, it only uses this confidence output, and the alarm function can be realized in combination with the confidence threshold). Decoupling head output The confidence level of different flame label types, and the output The target position of each flame target (the offset of each flame target center relative to the starting position of the grid and the target width).

[0061] In this structure, flame classification uses the Sigmoid activation function instead of the Softmax activation function. The Sigmoid activation function is expressed as follows:

[0062] The purpose of this design is to meet the needs of this project. That is, when the IR signal contains multiple flames, the same grid can simultaneously predict the target confidence and target box position of different flame signals. In principle, it is different from conventional image target detection, in which a grid (anchor-free, no anchor box design) or anchor box (anchor-base, with anchor box design) is only responsible for detecting one target.

[0063] It should be noted that the parameters of the CBL 1d module and Conv 1d module of the three branches are set according to the output size and are not exactly the same. Assume that the input grid width of the three parallel decoupling networks is The feature map of the confidence decoupling network outputs Flame label confidence, flame type decoupling network output The flame type confidence, position decoupling network output regression Flame target position characteristics.

[0064] In the above embodiment, the SPPF1d module (1d structure size) is a technology in deep learning. Its main function is to capture multi-scale feature information through pooling windows of different sizes, that is, to perceive fields of view of different sizes. In this application, the SPPF1d structure design is as follows Figure 8 The SPPF module inputs the original IR feature data and captures multi-scale feature information through three sequentially cascaded and different-sized pooling MaxPool structures (specifically 1d structures). The output of each pooling structure is fused with the original IR feature data through the channel splicing unit, and the output does not change the feature size (resolution W). However, three pooling structures of different sizes are used, and the number of output channels after splicing is 4 times the number of input channels. For example, if the original IR signal is 1 4 1024 (1 sample, 4 channels, 1024 resolution), becomes 1 after SPPF1d 16 1024 (1 sample, 16 channels, 1024 resolution), providing more feature information under different fields of view for the subsequent extraction of IR features of the C21d structure.

[0065] Figure 9 This is a structural diagram of the Focus module. The Focus module is connected to the output of the SPPF module, slices features through 64 parallel Slice units, and splices features through the channel splicing unit to increase the feature size and downsample the output. Figure 9 The middle module is sliced ​​and then spliced ​​at intervals of 64. For example, the output of SPPF1d is 1 16 1024 (1 sample, 16 channels, 1024 resolution), after slicing it becomes 1 1024 16 (1 sample, 1024 channels, 16-bit resolution). After processing with Focus1d, no data information is lost while achieving fast downsampling.

[0066] The C2 (specifically 1D structure) feature extraction module is used multiple times in the above model construction. Figure 10 The schematic diagram of the structure of the C21d module is shown, which specifically includes the cascaded CBL module and N bottleneck modules BottleNeck 1d n, a residual connection is set between the CBL module and the bottleneck module. The residual connection part is the CBL module. The bottleneck module output and the residual connection output are fused and output through Concat, and the output does not change the feature size. Figure 11 This is a structural diagram of the bottleneck module. The bottleneck module contains two cascaded CBL modules, and a directly connected residual structure ShortCut is set between the cascaded CBL modules. ShortCut and the cascaded part are superimposed on the output. The two cascaded CBL modules halve and restore the channels in turn, and the superimposed output does not change the feature size. That is, assuming the original input is 1 C The W feature map output is 1 after the first CBL 1d 0.5 C W feature map, after another CBL 1d output 1 C W feature map, after superimposing ShortCut, the output size is still 1 C W feature map.

[0067] In summary, the BottleNeck1d structure makes the input and output resolution and number of channels the same. The first convolution reduces the number of channels by half, and the second convolution restores it to the original number of channels, reducing the number of network parameters. The ShortCut residual direct connection structure can retain the original input features, which has the effect of alleviating the gradient vanishing problem, accelerating training convergence and improving prediction accuracy.

[0068] In the forward reasoning output stage of the model, considering special cases such as fusion strategy, the process of outputting the fused flame type and flame position coordinates can be further summarized as follows: A. Obtain the flame tag and determine the flame target center position and flame width in the IR time sequence based on the flame target position; Figure 12 The IR timing diagram of the four-channel flame target with labeled output under a possible situation is listed, where the red target box 1 and the blue target box 2 represent the two flame targets of the predicted output (assuming that the predicted target box 1 and target box 2 both meet the flame confidence threshold condition, such as 0.9). and The parameters representing the two flame targets are x, the center coordinate of the target frame (flame), and w, the flame width (i.e., flame size). For a time series graph of a fixed size, the coverage of the target frame represents the flame position and width (the specific fire location in the 3D world can be determined by reverse reasoning), and the center of the target frame is the center of the flame target.

[0069] B. Identify multiple flame targets with the same flame category and overlapping target positions in the IR time series, and fuse and update the flame label type, number and target position according to the target position; Under normal circumstances, when only one target box can be predicted and output, there is no fusion strategy, but when there are two target boxes and the two target boxes overlap (i.e. Figure 12 In the case of the display, the output needs to be fused.

[0070] C. When at least two flames are of different types and have overlapping parts, they are directly output without fusion; D. When at least two flames are of the same type and have overlapping parts, the number of flames, target center position and width are fused. The fused and updated flame position is expressed as follows:

[0071] Among them and Indicates the updated flame center coordinates and flame width; and Respectively represent the starting coordinates of the near and far flame targets in the IR time series, and Indicates the width of near and far flame targets.

[0072] Of course, in some embodiments, flame type labels can be disregarded. In other words, as long as overlapping flames are detected, both flames are output according to the aforementioned fusion strategy, thus sufficing to output only one flame type. This output is based on the fact that, since a flame has been detected, an alarm should be activated immediately, regardless of the type of flame. Furthermore, flame IR signal target detection assumes that targets of the same type do not overlap (this is different from the situation in which similarly labeled targets such as vehicles and people in a visual image may partially overlap).

[0073] In summary, the beneficial technical effects brought about by the technical solution of this application are as follows: A backbone network, a first feature extraction and fusion network with a Neck-FPN structure, and a second feature extraction and fusion network with a Neck-PAN structure are used to extract deep and shallow feature data from the feature map in a hierarchical and phased manner. The deep and shallow features are fused using a multi-scale feature fusion method. The flame targets and types are trained using a sample set, and the flame size and position information is regressed and predicted through path aggregation. Finally, multiple decoupling heads are used to combine the label classification problem with the position regression problem to achieve a fused output, improving flame detection accuracy while providing a more intelligent output. The model output uses a fusion strategy to identify the number and range of flame targets, update the number and size of flames, and combine with the sampling input mechanism to accurately identify the scale, location, and duration of ambient flames, providing an effective reference for subsequent accurate alarms and emergency response. Flame samples and interference source samples are augmented to enrich the sample set, effectively improving the generalization performance of the flame detector AI model.

[0074] The following is an example of the actual application of this solution: This solution uses the method of augmenting existing flame and interference source samples to create training and test sets. Figure 14 This is a schematic diagram of the IR signal after adding an interference source. The incandescent lamp interference source signal is augmented, and a fixed-length signal is randomly intercepted from it as the background. Then, one or more segments of the signal within a certain length range are randomly intercepted from the n-heptane flame signal as the target, and randomly superimposed on the background signal for the final fusion output.

[0075] In order to make the fusion signal smooth, the intercepted target signal is firstly subjected to the reference signal removal operation, and then the Hanning window is added. Figure 13 This is a schematic diagram of the Hanning window signal, and its generation formula is as follows:

[0076] Based on the above formula, set the target width , get the intercepted signal Taking the fusion process of a certain channel of IR data as an example, we get Figure 15 The diagrams shown before and after data augmentation show that the left image shows a single-channel fusion example, and the right image shows a four-channel signal augmentation example. This approach augments the Flame AI model's training and test sets, enriching the sample sets and improving the AI ​​model's generalization performance.

[0077] In one embodiment, Pytorch is used to build the flame AI model. Figure 16 The schematic diagram of the SPPF1d+Focus1d structure shown in the figure is as follows: Figure 17 The following is a schematic diagram of the decoupling head output structure. The test indicators used are defined as follows: Let TP (True Positive) be the correct positive prediction, FP (False Positive) be the wrong positive prediction, TN (True Positive) be the correct negative prediction, and FN (False Positive) be the wrong negative prediction, then the accuracy is is defined as:

[0078] Recall is defined as:

[0079] This example uses three typical flame IR signal samples of n-heptane, ethanol, and hydrogen for simulation testing. According to the augmentation method in Section 4, 100,000 samples are prepared as a training set and 10,000 samples are prepared as a test set (this test does not cover multiple flame target signals). After 100 rounds of training, the test is conducted on the test set to obtain Figure 18 Schematic diagram of precision and recall performance shown in, Figure 18 The left side shows the test accuracy, the middle shows the recall, and the right side shows the PR curve. mAP@0.5 = 0.994. @0.5 refers to the intersection over union (IoU) of the target prediction box and the ground-truth box. mAP is defined as follows:

[0080] Among them represents the number of categories, Indicates the The integral of class precision and recall.

[0081]

[0082] When the intersection over union (IoU) is greater than 0.5, the predicted target is considered to be of TP type.

[0083] Set the predicted box and the true box to be and , and is the center position and width of the prediction box, and The center position and width of the marker value. Figure 19 This is an example graph of IoU calculated according to the above evaluation and prediction scheme. It can be seen from the figure that the error between the predicted target box and the true target box is extremely small, meeting the requirements of high-precision alarm prediction.

[0084] The embodiment of the present application also discloses a computer-readable storage medium. Specifically, the computer-readable storage medium is used to store a computer program, and when the computer program is executed by the processor, the method in the above-mentioned method implementation is implemented. Those skilled in the art will understand that all or part of the processes in the above-mentioned method implementation of the present application can be completed by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the implementation of the above-mentioned methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.

[0085] This specific embodiment is merely an explanation of the present invention and is not intended to limit the present invention. After reading this specification, those skilled in the art may make non-creative modifications to this embodiment as needed. However, as long as such modifications are within the scope of the claims of the present invention, they are protected by patent law.

Claims

1. A flame detection method based on a deep model, characterized in that: The method comprises: The IR data of different flame and interference source signals are collected through various IR sensors, and the flame IR data is used as the target and the interference source IR data is used as the background for fusion to expand the training set and test set; According to the flame label type and target position data contained in the IR data in the training set, a flame deep network model based on target classification and position regression is constructed and driven; Based on the model prediction value, the label mark value and the position mark value in the training sample set, the flame confidence loss, flame label type confidence loss and target position loss are calculated respectively, and the total model loss is constructed based on various losses; The model parameters are updated based on the back propagation of the total loss function value of the model. After the iterative model training is completed, the flame confidence, flame label type confidence and position regression value are output according to the sampled multi-channel IR data, and the final fusion result is output based on the fusion strategy.

2. The method according to claim 1, characterized in that The Flame deep network model includes a cascaded backbone network, a first feature extraction and fusion network, a second feature extraction and fusion network, and a decoupled regression network; The backbone network includes a cascaded spatial pyramid pooling layer SPPF module and a focus module. The input IR feature data is captured by the SPPF module under different field of view sizes, and then the Focus module performs fast slicing and downsampling. After multiple downsampling and feature extraction, it outputs a multi-scale feature map. The first feature extraction and fusion network includes a horizontal connection network and a first vertical fusion network. The horizontal connection network part inputs the multi-scale feature map, and the first vertical fusion network cascades the output of the backbone network to successively convert the high-level feature map into the low-level feature map, and matches and fuses the multi-scale feature map with the low-level feature map to output a multi-scale first fused feature map with flame type information. The second feature extraction and fusion network includes a horizontal connection network and a second vertical fusion network. The horizontal connection network part inputs the first fused feature map of the corresponding scale. The second vertical fusion network cascades the output of the first vertical fusion network, converts the low-level feature map into a high-level feature map one by one, and matches and fuses the first fused feature map of the corresponding scale with the high-level feature map to output a multi-scale second fused feature map with position information. The decoupling regression network sets up several decoupling heads at different scales. Each decoupling head inputs the second fused feature map of the corresponding scale through a horizontal connection network, predicts the type information and target position information contained therein through classification and regression, and outputs the flame label type, flame number and flame position.

3. The method according to claim 2, characterized in that After the Focus module of the backbone network, three groups of CBL modules and C2 feature extraction modules are alternately cascaded; the CBL modules downsample layer by layer, and the C2 feature extraction modules extract features layer by layer. Each C2 feature extraction module outputs a corresponding scale feature map, which is then fed into the first vertical fusion network through the horizontal connection network; The first feature extraction and fusion network includes three groups of alternately cascaded CBL modules, upsampling modules, channel splicing units, and C2 feature extraction modules; the channel splicing unit performs channel splicing on the high-level feature map output by upsampling and the corresponding scale feature map, and the CBL module outputs the first fused feature map at the corresponding scale; The second feature extraction fusion network includes three groups of alternately cascaded CBL modules, channel stitching units and C2 feature extraction modules; the channel stitching unit performs channel stitching on the first fused feature map and the feature map output by the CBL module, and derives a second fused feature map of corresponding scale from each CBL module input and C2 feature extraction module output through a horizontal connection network.

4. The method according to claim 2, characterized in that The decoupling head includes three parallel decoupling networks, namely a flame target decoupling network, a flame tag type decoupling network and a position decoupling network; The flame target decoupling network and the flame label type decoupling network respectively contain a cascaded CBL module, a one-dimensional convolution module, and a Sigmoid activation layer, while the position decoupling network contains a cascaded CBL module and a one-dimensional convolution module. The number of grids of the input feature maps of the three parallel decoupling networks is , confidence decoupling network output Flame label confidence, flame type decoupling network output The flame type confidence, position decoupling network output regression Flame target location information; The three parallel decoupling networks are fused through the channel splicing unit to output the flame confidence, flame label type, and flame position information of all grids at different scales.

5. The method according to claim 3 or 4, characterized in that The SPPF module inputs the original IR feature data and captures multi-scale feature information through three sequentially cascaded pooling structures of different sizes. The output of each pooling structure is fused with the original IR feature data through a channel splicing unit, and the output does not change the feature size. The Focus module is connected to the output of the SPPF module, performs feature slicing through 64 parallel Slice slicing units, and performs fast downsampling through channel splicing.

6. The method according to any one of claims 2 to 4, characterized in that: The CBL module includes a cascaded one-dimensional convolution module Conv1d, a normalization layer and an activation layer; The C2 feature extraction module includes a cascaded CBL module and N bottleneck modules. A residual connection is set between the CBL module and the bottleneck module. The residual connection part is the CBL module. The output of the bottleneck module and the residual connection output are fused and output through the channel splicing unit, and the output does not change the feature size. The bottleneck module consists of two cascaded CBL modules and a directly connected residual structure is set between the cascaded CBL modules, and the output is superimposed with the cascaded part; the two cascaded CBL modules halve the channels and restore the channels in turn, and the superimposed output does not change the feature size.

7. The method according to claim 3, characterized in that The loss calculation process is as follows: ; ; ; ; Among them is the total number of labeled positive and negative samples, is the number of categories, represents the flame confidence loss, represents the confidence loss of flame type, represents the target position loss, Represents the total loss function of the model; is the target number, is the flame target mark value, is the flame confidence prediction value, is the flame tag type tag value, is the flame label type confidence prediction value, is the mark value of the flame target center position, is the predicted value of the flame target center position, is the flame target width mark value, is the flame target width prediction value; 、 and is an adjustable weight coefficient.

8. The method according to claim 1 or 7, characterized in that The outputting of the fused flame type and flame position coordinates according to the sampled IR data includes: Obtain the flame tag and determine the flame target center position and flame width in the IR time series based on the flame target position; Identify multiple flame targets with the same flame category and overlapping target positions in the IR time series, and fuse them to update the flame category, number, and target position based on the target position; When at least two flames are of different types and have overlapping parts, they are directly output without fusion; When at least two flames are of the same type and have overlapping parts, the number of flames, target center position and width are fused. The fused and updated flame position is expressed as follows: ; Among them and Indicates the updated flame center coordinates and flame width; and Respectively represent the starting coordinates of the near and far flame targets in the time series, and Indicates the width of near and far flame targets.

9. A flame detector, characterized in that: The flame detector includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the deep model-based flame detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the flame detection method based on a deep model as described in any one of claims 1 to 7.