Smoke detection method and device based on context feature interaction, equipment and medium

CN117830691BActive Publication Date: 2026-09-29SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311653523.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2026-09-29
Estimated Expiration
2043-12-05

AI Technical Summary

Benefits of technology

[0023]本发明实施例提供了一种基于上下文特征交互的烟雾检测方法、装置、设备及介质。其中,方法包括:收集样本烟雾图像集,并对所述样本烟雾图像集中的每一图像进行预处理,得到训练图像数据;将所述训练图像数据输入至检测网络模型中,通过引入重参数化大核卷积的特征提取模块对所述训练图像数据进行特征提取,得到多层初始特征,其中,所述初始特征包括第一层初始特征、第二层初始特征和第三层初始特征;基于所述初始特征构建特征金字塔,并基于所述特征金字塔进行特征融合,得到初始融合特征;通过上下文特征交互增强模块基于所述初始融合特征进行上下文特征交互增强融合,得到目标融合特征;基于所述目标融合特征进行目标检测,得到多个目标的类别信息和位置信息,并基于所述类别信息和位置信息进行损失计算,生成模型损失值;基于所述模型损失值调整所述检测网络模型的模型参数,并采用反向传播的方式对所述检测网络模型进行训练,得到目标检测网络模型;获取待检测烟雾图像,并基于所述目标检测网络模型对所述待检测烟雾图像进行检测,生成目标检测结果。本发明实施例通过将上下文特征进行交互融合增强,有利于提高烟雾检测的准确性,同时在检测到烟雾发生时便能够及时发现烟雾,有利于提高烟雾检测的效率,并且无需部署多个烟雾传感器,有利于降低烟雾检测的成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117830691B_ABST
    Figure CN117830691B_ABST
Patent Text Reader

Abstract

The application relates to a smoke detection method and device based on context feature interaction, equipment and medium, wherein the method comprises the following steps: preprocessing a collected sample smoke image set to obtain training image data; performing feature extraction on the training image data through a feature extraction module of a reparameterization large kernel convolution to obtain multi-layer initial features; performing preliminary fusion on the initial features to obtain initial fusion features; performing context feature interaction enhancement fusion on the initial fusion features through a context feature interaction enhancement module to obtain target fusion features; performing target detection and loss calculation based on the target fusion features to generate a model loss value; training a detection network model based on the model loss value to obtain a target detection network model; and detecting a to-be-detected smoke image based on the target detection network model to generate a target detection result. The application improves the accuracy and efficiency of smoke detection and reduces the cost of smoke detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a smoke detection method, apparatus, device, and medium based on contextual feature interaction. Background Technology

[0002] Early detection of potential fire hazards is a major concern. Smoke typically rises and is visible in the initial stages before a fire spreads. Smoke provides an early indication of fire risk. However, smoke has natural properties such as being deformable and a non-rigid substance, which can interfere with smoke detection. Furthermore, no single feature can perfectly characterize smoke. Additionally, objects similar to smoke or that can interfere with detection, such as clouds, water vapor, fog, shadows, sky, and glass reflections, can also contribute to smoke detection.

[0003] Currently, most smoke detection systems are based on sensors. However, because smoke takes time to propagate, smoke sensors cannot detect it promptly and accurately. Furthermore, smoke can spread in multiple directions, resulting in low efficiency when detecting large areas. Additionally, smoke sensors are expensive, leading to high costs for smoke detection. Summary of the Invention

[0004] The purpose of this application is to propose a smoke detection method, apparatus, device, and medium based on context feature interaction, so as to improve the accuracy and efficiency of smoke detection and reduce the cost of smoke detection.

[0005] To address the aforementioned technical problems, embodiments of this application provide a smoke detection method based on contextual feature interaction, comprising:

[0006] Collect a set of sample smoke images and preprocess each image in the set to obtain training image data;

[0007] The training image data is input into the detection network model, and the feature extraction module of reparameterized large kernel convolution is introduced to extract features from the training image data to obtain multi-layer initial features, wherein the initial features include a first layer initial features, a second layer initial features and a third layer initial features;

[0008] A feature pyramid is constructed based on the initial features, and feature fusion is performed based on the feature pyramid to obtain the initial fused features;

[0009] The target fused feature is obtained by performing context feature interaction enhancement fusion based on the initial fused feature through the context feature interaction enhancement module.

[0010] Target detection is performed based on the target fusion features to obtain the category information and location information of multiple targets, and loss is calculated based on the category information and location information to generate model loss values;

[0011] The model parameters of the detection network model are adjusted based on the model loss value, and the detection network model is trained using backpropagation to obtain the target detection network model.

[0012] A smoke image to be detected is acquired, and the smoke image is detected based on the target detection network model to generate a target detection result.

[0013] To address the aforementioned technical problems, embodiments of this application provide a smoke detection device based on contextual feature interaction, comprising:

[0014] A smoke image collection unit is used to collect a sample smoke image set and preprocess each image in the sample smoke image set to obtain training image data;

[0015] An initial feature extraction unit is used to input the training image data into the detection network model, and to extract features from the training image data by introducing a feature extraction module with reparameterized large kernel convolution to obtain multi-layer initial features, wherein the initial features include a first layer initial features, a second layer initial features, and a third layer initial features;

[0016] An initial feature fusion unit is used to construct a feature pyramid based on the initial features and perform feature fusion based on the feature pyramid to obtain initial fused features;

[0017] The target feature fusion unit is used to perform context feature interaction enhancement fusion based on the initial fusion features through the context feature interaction enhancement module to obtain the target fusion features;

[0018] The model loss calculation unit is used to perform target detection based on the target fusion features, obtain the category information and location information of multiple targets, and perform loss calculation based on the category information and location information to generate the model loss value;

[0019] The model training unit is used to adjust the model parameters of the detection network model based on the model loss value, and to train the detection network model using backpropagation to obtain the target detection network model.

[0020] The detection result generation unit is used to acquire the smoke image to be detected, and to detect the smoke image to be detected based on the target detection network model to generate the target detection result.

[0021] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is to provide a computer device, including one or more processors; and a memory for storing one or more programs, such that the one or more processors implement the smoke detection method based on context feature interaction as described above.

[0022] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is: a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the smoke detection method based on context feature interaction as described above.

[0023] This invention provides a smoke detection method, apparatus, device, and medium based on context feature interaction. The method includes: collecting a set of sample smoke images and preprocessing each image in the set to obtain training image data; inputting the training image data into a detection network model, and extracting features from the training image data by introducing a feature extraction module with reparameterized large kernel convolution to obtain multi-layer initial features, wherein the initial features include a first layer initial feature, a second layer initial feature, and a third layer initial feature; constructing a feature pyramid based on the initial features, and performing feature fusion based on the feature pyramid to obtain initial fused features; performing context feature interaction enhancement fusion based on the initial fused features through a context feature interaction enhancement module to obtain target fused features; performing target detection based on the target fused features to obtain category information and location information of multiple targets, and calculating loss based on the category information and location information to generate model loss values; adjusting the model parameters of the detection network model based on the model loss values, and training the detection network model using backpropagation to obtain a target detection network model; acquiring smoke images to be detected, and detecting the smoke images to be detected based on the target detection network model to generate target detection results. The embodiments of the present invention enhance the accuracy of smoke detection by interactively fusing contextual features. At the same time, it can detect smoke in a timely manner when it occurs, which improves the efficiency of smoke detection. Furthermore, it eliminates the need to deploy multiple smoke sensors, thereby reducing the cost of smoke detection. Attached Figure Description

[0024] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1This is a flowchart illustrating the implementation of the smoke detection method based on context feature interaction provided in this application embodiment;

[0026] Figure 2 This is a schematic diagram of the overall structure of the detection network model provided in the embodiments of this application;

[0027] Figure 3 This is a flowchart illustrating the implementation of a sub-process in the smoke detection method based on context feature interaction provided in this application embodiment;

[0028] Figure 4 This is a flowchart illustrating the implementation of a sub-process in the smoke detection method based on context feature interaction provided in this application embodiment;

[0029] Figure 5 This is a schematic diagram of the introduced reparameterized large kernel convolution model provided in an embodiment of this application;

[0030] Figure 6 This is a flowchart illustrating the implementation of a sub-process in the smoke detection method based on context feature interaction provided in this application embodiment;

[0031] Figure 7 This is a flowchart illustrating the implementation of a sub-process in the smoke detection method based on context feature interaction provided in this application embodiment;

[0032] Figure 8 This is a flowchart illustrating the implementation of a sub-process in the smoke detection method based on context feature interaction provided in this application embodiment;

[0033] Figure 9 This is a flowchart of a sub-process in the smoke detection method based on context feature interaction provided in the embodiments of this application;

[0034] Figure 10 This is a schematic diagram of the intersection of bounding boxes provided in an embodiment of this application;

[0035] Figure 11 This is a schematic diagram of a smoke detection device based on context feature interaction provided in an embodiment of this application;

[0036] Figure 12 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0038] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0039] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0040] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0041] It should be noted that the smoke detection method based on context feature interaction provided in this application is generally executed by a server, and correspondingly, the smoke detection device based on context feature interaction is generally configured in the server.

[0042] Please see Figure 1 and Figure 2 , Figure 1 This paper illustrates a specific implementation of a smoke detection method based on contextual feature interaction. Figure 2 This is a schematic diagram of the overall structure of the detection network model provided in the embodiments of this application.

[0043] It should be noted that if substantially the same result is obtained, the method of this invention is not based on... Figure 1 Limited to the order of the processes shown, this method includes the following steps:

[0044] S1: Collect a set of sample smoke images and preprocess each image in the set to obtain training image data.

[0045] Specifically, this application uses Ultralytics' YOLOv5 model as the base model and improves the model by proposing a Contaxual Interaction Enhancement Module (CIEM), introducing heavily parameterized large kernel convolution, and improving the regression loss function IoU Loss, so as to achieve a real-time and accurate smoke detection method that meets the requirements of industrial application.

[0046] In this embodiment, smoke image data needs to be collected to train the detection network model. Therefore, in this embodiment, a sample smoke image set is collected, and each image in the sample smoke image set is preprocessed to obtain training image data. The preprocessing involves performing Mosaic data augmentation on the smoke images, which involves stitching the smoke images together by randomly scaling, cropping, and arranging them to generate the corresponding training image data.

[0047] In one specific embodiment, due to the lack of publicly available datasets in the field of smoke detection research, only a limited number of low-resolution videos are publicly available. Furthermore, datasets containing challenging environmental conditions such as fog and haze are scarce. Therefore, this application proposes a new smoke dataset, aiming to inspire more researchers to discover innovative solutions for early smoke detection. To ensure the dataset's rationality and effectiveness, this application collected and labeled smoke images from different scenarios, integrating them into a dataset containing approximately 24,776 smoke images. This dataset consists of only a single category of smoke, encompassing various scenes from simple to complex, single or multiple objects, and indoor and outdoor smoke simulation images. This diverse inclusion ensures the dataset's universality and applicability to models. Furthermore, this application meticulously categorizes the dataset, forming a well-classified and comprehensive resource, including MS-COCO, VOC, and YOLO formats.

[0048] S2: Input the training image data into the detection network model, and extract features from the training image data by introducing a feature extraction module with reparameterized large kernel convolution to obtain multi-layer initial features.

[0049] The initial features include a first layer of initial features, a second layer of initial features, and a third layer of initial features.

[0050] like Figure 2 As shown, the improved YOLOv5-s model is used as the detection network model, which mainly includes a Backbone (referred to as the feature extraction module in this application) for feature extraction, a Neck module, a context feature interaction enhancement module CIEM, and a decoupled detection head.

[0051] In this embodiment, the feature extraction module includes modules C3, C4, and C5. This application improves upon the C3 module of YOLOv5-s by reparameterizing a portion of the C3 module into a large kernel convolution module, resulting in a larger receptive field and capturing finer details and contextual information of smoke objects, thereby improving smoke target detection performance. This application uses the features output after passing through module C3 as the initial features of the first layer, the features output after passing through module C4 as the initial features of the second layer, and the features output after passing through module C5 and spatial pyramid pooling as the initial features of the third layer.

[0052] Please see Figure 3 , Figure 3 A specific implementation of step S2 is shown below:

[0053] S21: Input the training image data into the detection network model.

[0054] S22: The C3 module, which incorporates reparameterized large kernel convolution, extracts features from the training image to generate the initial features of the first layer.

[0055] Please see Figure 4 and Figure 5 , Figure 4 One specific implementation of step S22 is shown. Figure 5 This is a schematic diagram of the introduced reparameterized large kernel convolution model provided in an embodiment of this application, detailed below:

[0056] S221: In the C3 module, which introduces reparameterized large kernel convolution, the training image is convolved by two different convolutional blocks to obtain two convolutional features.

[0057] S222: Perform batch standardization and superposition processing on the two convolutional features to generate basic features.

[0058] S223: Perform residual connections on the different basic features, and reparameterize the basic features after residual connection to obtain the initial features of the first layer.

[0059] Specifically, in the C3 module, which introduces reparameterized large kernel convolution, the training image is convolved with two different convolutional blocks (3×3 and 27×27 blocks respectively) to obtain two convolutional features. These two features are then batch normalized (BN operation), and finally superimposed to generate basic features. Residual connections are then made between these basic features, and the resulting basic features are reparameterized to obtain the initial features of the first layer. This embodiment, through the combination of residual connections, helps to provide a more comprehensive understanding of the input data for the detection network model, enabling it to capture complex details and global information. Compared to the original C3 module, the large kernel convolution module reduces the number of parameters, thereby reducing model complexity and computational burden.

[0060] S23: Extract features from the first layer initial features based on the C4 module to generate the second layer initial features.

[0061] S24: The second layer initial features are processed by feature extraction and spatial pyramid pooling based on the C5 module to generate the third layer initial features.

[0062] Specifically, after the C3 module, which introduces heavily parameterized large kernel convolution, outputs the first layer of initial features, the C4 module performs feature extraction to generate the second layer of initial features. The second layer of initial features is then processed by the C5 module for feature extraction and spatial pyramid pooling to generate the third layer of initial features.

[0063] S3: Construct a feature pyramid based on the initial features, and perform feature fusion based on the feature pyramid to obtain initial fused features.

[0064] like Figure 2 As shown, this application constructs a feature pyramid based on initial features using the Neck module, and performs feature fusion based on the feature pyramid to obtain initial fused features. In this embodiment, in order to combine with the subsequent context feature interaction enhancement processing in the context feature interaction enhancement module, the following steps are taken: Figure 4 In the Neck module, the feature output after passing through the P3 module is used as the first fused feature, the feature output after passing through the N3 module is used as the second fused feature, the feature output after passing through the N4 module is used as the third fused feature, and the feature output after passing through the N5 module is used as the fourth fused feature.

[0065] S4: The context feature interaction enhancement module performs context feature interaction enhancement fusion based on the initial fusion features to obtain the target fusion features.

[0066] Specifically, this embodiment adds a Contaxtual Interaction Enhancement Module (CIEM) to the Neck and Head modules of the original YOLOv5-s model. Since classification and regression tasks have different preferences for feature context, existing methods typically use a decoupled Head to learn different context features for each character. However, regression tasks are fine-grained and require more low-level boundary-aware features to accurately regress bounding boxes, while classification tasks are coarse-grained and require more high-level semantic context information. Low-level features have more boundary-aware details but lack contextual semantic information, while high-level features are the opposite. Therefore, to address the complex properties of smoke, such as its non-fixed shape and constantly changing appearance, the model is further decoupled. The use of depthwise separable convolution helps the model learn more complex feature representations. In this embodiment, by upsampling, convolution, and depthwise convolution operations on the initial fused features, the feature map fuses contexts at different scales to obtain a fused multi-scale feature context, playing a role similar to global modeling in a Vision Transformer.

[0067] Please see Figure 6 , Figure 6 A specific implementation of step S4 is shown below:

[0068] S41: In the context feature interaction enhancement module, the first fused feature is convolved to obtain the first target fused feature.

[0069] S42: Stack the first layer initial features with the first target fusion features to obtain the second target fusion features.

[0070] like Figure 2 As shown, in the context feature interaction enhancement module, the first fused feature is convolved to obtain the first target fused feature; the first layer initial feature and the first target fused feature are stacked to obtain the second target fused feature.

[0071] S43: Based on the first fusion feature, the third fusion feature and the fourth fusion feature, feature fusion is performed to generate a third target fusion feature.

[0072] Please see Figure 7 , Figure 7 A specific implementation of step S43 is shown below:

[0073] S431: The third fusion feature is upsampled and convolutionally processed, and then added to the first fusion feature to obtain the first added feature.

[0074] S432: Upsample and perform depth convolution on the fourth fusion feature to obtain the first depth feature.

[0075] S433: The first additive feature is processed by convolution through three layers of convolutional blocks and then added to the first depth feature to obtain the third target fusion feature.

[0076] Specifically, the third fusion feature is upsampled and convolved, then added to the first fusion feature to obtain the first added feature. Next, the fourth fusion feature is upsampled and depthwise convolved to obtain the first depth feature. Finally, the first added feature is convolved through three layers of convolutional blocks and added to the first depth feature to obtain the third target fusion feature. Figure 2 In this code, "Conv" indicates convolution processing, "Upsample" indicates upsampling processing, "DWConv" indicates depthwise convolution processing, "C" indicates stacking processing, and "+" indicates addition processing.

[0077] S44: After performing convolution processing on the third fusion feature, it is stacked with the second fusion feature to generate the fourth target fusion feature.

[0078] S45: Based on the third fusion feature and the fourth fusion feature, perform feature fusion to generate the fifth target fusion feature.

[0079] Please see Figure 8 , Figure 8 A specific implementation of step S45 is shown below:

[0080] S451: The fourth fusion feature is upsampled and depthwise convolved, and then added to the third fusion feature to obtain the second added feature.

[0081] S452: The second additive feature is processed by deep convolution through three layers of deep convolution blocks and then added to the fourth fusion feature to obtain the fifth target fusion feature.

[0082] Specifically, it is necessary to fuse the fourth fusion feature and the third fusion feature. Therefore, in this embodiment, the fourth fusion feature is upsampled and deep convolutioned and then added to the third fusion feature to obtain the second added feature. Then, the second added feature is deep convolutioned through three layers of deep convolution blocks and then added to the fourth fusion feature to obtain the fifth target fusion feature.

[0083] S46: After performing convolution processing on the fourth fusion feature, it is stacked with the third fusion feature to generate the sixth target fusion feature.

[0084] The target fusion features include a first target fusion feature, a second target fusion feature, a third target fusion feature, a fourth target fusion feature, a fifth target fusion feature, and a sixth target fusion feature. Therefore, in this embodiment, after passing through the context feature interaction enhancement module, the first target fusion feature, the second target fusion feature, the third target fusion feature, the fourth target fusion feature, the fifth target fusion feature, and the sixth target fusion feature are output.

[0085] S5: Target detection is performed based on the target fusion features to obtain the category information and location information of multiple targets, and loss calculation is performed based on the category information and location information to generate model loss values.

[0086] Specifically, the output of object detection is to obtain the category information and location information of a certain object. In the embodiments of this application, object detection is performed on the fused features of each object by decoupling the detection head to obtain the category information and location information of multiple objects, and loss is calculated based on the category information and location information to generate the model loss value.

[0087] Please see Figure 9 and Figure 10 , Figure 9 One specific implementation of step S5 is shown. Figure 10 This is a schematic diagram of the intersection of bounding boxes provided in an embodiment of this application, detailed below:

[0088] S51: The decoupled detection head is used to perform target detection on the target fusion features to obtain the category information and location information of multiple targets.

[0089] like Figure 2 As shown, in the decoupling detection head, the first target fusion feature, the second target fusion feature, the third target fusion feature, the fourth target fusion feature, the fifth target fusion feature, and the sixth target fusion feature are all processed by two convolutional blocks to obtain the decoupling information corresponding to each feature. Specifically, the first target fusion feature is decoupled as Reg0 Small, the second target fusion feature is decoupled as Cls0 Small, the third target fusion feature is decoupled as Reg1 Medium, the fourth target fusion feature is decoupled as Cls1 Medium, the fifth target fusion feature is decoupled as Reg2 Large, and the sixth target fusion feature is decoupled as Cls2 Large.

[0090] S52: The binary cross-entropy loss function is used as the classification loss function and the regression loss function, and the preset loss function is used as the confidence loss function. Based on the category information and location information, the loss is calculated to obtain the classification loss value, the regression loss value and the confidence loss value.

[0091] S53: Add the classification loss value, the regression loss value, and the confidence loss value together to obtain the model loss value.

[0092] Specifically, the loss function of the detection network model in this embodiment includes: classification loss Cls_loss, regression loss Box_Loss, and confidence loss Objectness_loss. The model loss value = classification loss value + regression loss value + confidence loss value. A binary cross-entropy loss function is used as both the classification and regression loss functions, and a preset loss function is used as the confidence loss function. Loss is calculated based on category and location information to obtain the classification loss value, regression loss value, and confidence loss value. Existing confidence loss uses the IoU loss function for calculation. However, since IoU Loss is invariant to bounding box scales, it can better train the detection network model. However, when the predicted bounding box and the ground truth bounding box do not overlap, there is a gradient vanishing problem, which leads to a decrease in convergence speed and detection accuracy. Therefore, there are related variants of IoU Loss: Generalized IoU (GIoU), Distance-IoU (DIoU), and Complete IoU (CIoU). GIoU incorporates a penalty term into the IoU loss to mitigate the gradient vanishing problem. DIoU and CIoU, under the penalty condition, consider the center distance and aspect ratio between the predicted and ground truth boxes. Similarly, the preset loss function used in this embodiment is:

[0093]

[0094] Where B1 is the ground truth bounding box, B2 is the predicted bounding box, and C is the minimum bounding box of the ground truth bounding box and the predicted bounding box. α is the confidence loss value, a single Power parameter.

[0095] S6: Adjust the model parameters of the detection network model based on the model loss value, and train the detection network model using backpropagation to obtain the target detection network model.

[0096] Specifically, the model parameters of the detection network model are adjusted based on the model loss value, and the training image data is re-output into the detection network model in batches for training using backpropagation until a preset number of iterations or a preset value of model loss value is reached, thereby obtaining the target detection network model.

[0097] S7: Acquire the smoke image to be detected, and perform detection on the smoke image based on the target detection network model to generate target detection results.

[0098] Specifically, since the above steps have yielded the target detection network model, in practical applications, the smoke image to be detected is acquired, and the smoke image to be detected is detected based on the target detection network model to generate the target detection result.

[0099] In this embodiment, a sample smoke image set is collected, and each image in the sample smoke image set is preprocessed to obtain training image data. The training image data is input into a detection network model, and a feature extraction module with reparameterized large kernel convolution is introduced to extract features from the training image data to obtain multi-layer initial features, wherein the initial features include a first layer initial feature, a second layer initial feature, and a third layer initial feature. A feature pyramid is constructed based on the initial features, and feature fusion is performed based on the feature pyramid to obtain initial fused features. A context feature interaction enhancement module is used to perform context feature interaction enhancement fusion based on the initial fused features to obtain target fused features. Target detection is performed based on the target fused features to obtain the category information and location information of multiple targets, and loss is calculated based on the category information and location information to generate model loss values. The model parameters of the detection network model are adjusted based on the model loss values, and the detection network model is trained using backpropagation to obtain a target detection network model. A smoke image to be detected is acquired, and the smoke image to be detected is detected based on the target detection network model to generate target detection results. The embodiments of the present invention enhance the accuracy of smoke detection by interactively fusing contextual features. At the same time, it can detect smoke in a timely manner when it occurs, which improves the efficiency of smoke detection. Furthermore, it eliminates the need to deploy multiple smoke sensors, thereby reducing the cost of smoke detection.

[0100] Please refer to Figure 11 As a response to the above Figure 1 The implementation of the method shown in this application provides an embodiment of a smoke detection device based on context feature interaction. This device embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0101] like Figure 11 As shown, the smoke detection device based on contextual feature interaction in this embodiment includes: a smoke image collection unit 81, an initial feature extraction unit 82, an initial feature fusion unit 83, a target feature fusion unit 84, a model loss calculation unit 85, a model training unit 86, and a detection result generation unit 87, wherein:

[0102] The smoke image collection unit 81 is used to collect a sample smoke image set and preprocess each image in the sample smoke image set to obtain training image data.

[0103] The initial feature extraction unit 82 is used to input the training image data into the detection network model, and to extract features from the training image data by introducing a feature extraction module with reparameterized large kernel convolution to obtain multi-layer initial features, wherein the initial features include a first layer initial features, a second layer initial features and a third layer initial features;

[0104] The initial feature fusion unit 83 is used to construct a feature pyramid based on the initial features and perform feature fusion based on the feature pyramid to obtain initial fused features;

[0105] The target feature fusion unit 84 is used to perform context feature interaction enhancement fusion based on the initial fusion features through the context feature interaction enhancement module to obtain the target fusion features;

[0106] The model loss calculation unit 85 is used to perform target detection based on the target fusion features, obtain the category information and location information of multiple targets, and perform loss calculation based on the category information and location information to generate the model loss value;

[0107] The model training unit 86 is used to adjust the model parameters of the detection network model based on the model loss value, and to train the detection network model using backpropagation to obtain the target detection network model.

[0108] The detection result generation unit 87 is used to acquire the smoke image to be detected, and to detect the smoke image to be detected based on the target detection network model to generate the target detection result.

[0109] Furthermore, the feature extraction module includes modules C3, C4, and C5, and the initial feature extraction unit 82 includes:

[0110] A data input unit is used to input the training image data into the detection network model;

[0111] The first feature extraction unit is used to extract features from the training image by the C3 module which introduces reparameterized large kernel convolution, and generate the first layer of initial features.

[0112] The second feature extraction unit is used to extract features from the first layer initial features based on the C4 module to generate the second layer initial features.

[0113] The third feature extraction unit is used to perform feature extraction and spatial pyramid pooling processing on the second layer initial features based on the C5 module to generate the third layer initial features.

[0114] Furthermore, the first feature extraction unit includes:

[0115] The first convolution processing unit is used to perform convolution processing on the training image through two different convolution blocks in the C3 module, which introduces reparameterized large kernel convolution, to obtain two convolution features.

[0116] The basic feature generation unit is used to perform batch standardization and superposition processing on the two convolutional features to generate basic features;

[0117] The reparameterization processing unit is used to perform residual connection on the different basic features and reparameterize the basic features after residual connection to obtain the initial features of the first layer.

[0118] Furthermore, the initial fusion feature includes a first fusion feature, a second fusion feature, a third fusion feature, and a fourth fusion feature, and the target feature fusion unit 84 includes:

[0119] The first fusion feature processing unit is used to perform convolution processing on the first fusion feature in the context feature interaction enhancement module to obtain the first target fusion feature;

[0120] The second fusion feature processing unit is used to stack the first layer initial features with the first target fusion features to obtain the second target fusion features;

[0121] The third fusion feature processing unit is used to perform feature fusion based on the first fusion feature, the third fusion feature and the fourth fusion feature to generate a third target fusion feature;

[0122] The fourth fusion feature processing unit is used to perform convolution processing on the third fusion feature and then stack it with the second fusion feature to generate a fourth target fusion feature;

[0123] The fifth fusion feature processing unit is used to perform feature fusion based on the third fusion feature and the fourth fusion feature to generate a fifth target fusion feature;

[0124] The sixth fusion feature processing unit is used to perform convolution processing on the fourth fusion feature and then stack it with the third fusion feature to generate a sixth target fusion feature, wherein the target fusion feature includes the first target fusion feature, the second target fusion feature, the third target fusion feature, the fourth target fusion feature, the fifth target fusion feature and the sixth target fusion feature.

[0125] Furthermore, the third fusion feature processing unit includes:

[0126] The first addition processing unit is used to upsample and convolve the third fusion feature and then add it to the first fusion feature to obtain the first addition feature;

[0127] The first deep feature generation unit is used to upsample and perform deep convolution processing on the fourth fused feature to obtain the first deep feature;

[0128] The third target fusion feature generation unit is used to perform convolution processing on the first additive feature through three layers of convolutional blocks and then add it to the first depth feature to obtain the third target fusion feature.

[0129] Furthermore, the fifth fusion feature processing unit includes:

[0130] The second addition processing unit is used to upsample and perform depth convolution on the fourth fusion feature, and then add it to the third fusion feature to obtain the second addition feature;

[0131] The fifth target fusion feature generation unit is used to perform depth convolution processing on the second additive feature through three layers of depth convolution blocks, and then add it to the fourth fusion feature to obtain the fifth target fusion feature.

[0132] Furthermore, the model loss calculation unit 85 includes:

[0133] The target detection unit is used to perform target detection on the target fusion features using a decoupled detection head to obtain the category information and location information of multiple targets;

[0134] The loss calculation unit is used to use the binary cross-entropy loss function as the classification loss function and the regression loss function, and the preset loss function as the confidence loss function, and to perform loss calculation based on the category information and location information to obtain the classification loss value, the regression loss value and the confidence loss value.

[0135] The model loss value generation unit is used to add the classification loss value, the regression loss value and the confidence loss value to obtain the model loss value;

[0136] Furthermore, the preset loss function is:

[0137]

[0138] Where B1 is the ground truth bounding box, B2 is the predicted bounding box, and C is the minimum bounding box of the ground truth bounding box and the predicted bounding box. α is the confidence loss value, a single Power parameter.

[0139] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 12 , Figure 12 This is a basic structural block diagram of the computer device in this embodiment.

[0140] Computer device 9 includes a memory 91, a processor 92, and a network interface 93 that are interconnected via a system bus. It should be noted that only a computer device 9 with three components—memory 91, processor 92, and network interface 93—is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0141] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through methods such as keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0142] The memory 91 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 91 may be an internal storage unit of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 91 may also be an external storage device of the computer device 9, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 9. Of course, the memory 91 may also include both internal storage units and external storage devices of the computer device 9. In this embodiment, the memory 91 is typically used to store the operating system and various application software installed on the computer device 9, such as the program code of a smoke detection method based on context feature interaction. In addition, the memory 91 can also be used to temporarily store various types of data that have been output or will be output.

[0143] In some embodiments, processor 92 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 92 is typically used to control the overall operation of computer device 9. In this embodiment, processor 92 is used to run program code stored in memory 91 or process data, for example, to run the program code of the aforementioned smoke detection method based on context feature interaction, to implement various embodiments of the smoke detection method based on context feature interaction.

[0144] The network interface 93 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 9 and other electronic devices.

[0145] This application also provides another embodiment, namely, providing a computer-readable storage medium storing a computer program that can be executed by at least one processor to cause the at least one processor to perform the steps of the smoke detection method based on context feature interaction as described above.

[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0147] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A smoke detection method based on contextual feature interaction, characterized in that, include: Collect a set of sample smoke images and preprocess each image in the set to obtain training image data; The training image data is input into the detection network model, and the feature extraction module of reparameterized large kernel convolution is introduced to extract features from the training image data to obtain multi-layer initial features, wherein the initial features include a first layer initial features, a second layer initial features and a third layer initial features; A feature pyramid is constructed based on the initial features, and feature fusion is performed based on the feature pyramid to obtain the initial fused features; The target fused feature is obtained by performing context feature interaction enhancement fusion based on the initial fused feature through the context feature interaction enhancement module. Target detection is performed based on the target fusion features to obtain the category information and location information of multiple targets, and loss is calculated based on the category information and location information to generate model loss values; The model parameters of the detection network model are adjusted based on the model loss value, and the detection network model is trained using backpropagation to obtain the target detection network model. Acquire a smoke image to be detected, and perform detection on the smoke image based on the target detection network model to generate a target detection result; The initial fusion features include a first fusion feature, a second fusion feature, a third fusion feature, and a fourth fusion feature. The context feature interaction enhancement module performs context feature interaction enhancement fusion based on the initial fusion features to obtain the target fusion features, including: In the context feature interaction enhancement module, the first fused feature is convolved to obtain the first target fused feature; The first layer initial features are stacked with the first target fused features to obtain the second target fused features; Based on the first fusion feature, the third fusion feature and the fourth fusion feature, feature fusion is performed to generate a third target fusion feature; The third fusion feature is convolved and then stacked with the second fusion feature to generate a fourth target fusion feature; Based on the third fusion feature and the fourth fusion feature, feature fusion is performed to generate a fifth target fusion feature; The fourth fusion feature is convolved and then stacked with the third fusion feature to generate a sixth target fusion feature. The target fusion feature includes the first target fusion feature, the second target fusion feature, the third target fusion feature, the fourth target fusion feature, the fifth target fusion feature, and the sixth target fusion feature.

2. The smoke detection method based on contextual feature interaction according to claim 1, characterized in that, The feature extraction module includes modules C3, C4, and C5. The training image data is input into the detection network model, and features are extracted from the training image data by introducing a reparameterized large-kernel convolution feature extraction module to obtain multi-layer initial features, including: The training image data is input into the detection network model; The C3 module, which incorporates reparameterized large kernel convolution, extracts features from the training image to generate the initial features of the first layer. The first layer initial features are extracted based on the C4 module to generate the second layer initial features; The second layer initial features are processed by feature extraction and spatial pyramid pooling based on the C5 module to generate the third layer initial features.

3. The smoke detection method based on contextual feature interaction according to claim 2, characterized in that, The C3 module, which incorporates reparameterized large kernel convolution, extracts features from the training image to generate the first layer of initial features, including: In the C3 module, which introduces reparameterized large kernel convolution, the training image is convolved by two different convolutional blocks to obtain two convolutional features. Batch standardization and superposition processing are performed on the two convolutional features to generate basic features; The different basic features are residually connected, and the basic features after residual connection are reparameterized to obtain the initial features of the first layer.

4. The smoke detection method based on contextual feature interaction according to claim 1, characterized in that, The feature fusion is performed based on the first fusion feature, the third fusion feature, and the fourth fusion feature. Generate third target fusion features, including: The third fusion feature is upsampled and convolutionally processed, and then added to the first fusion feature to obtain the first additive feature; The fourth fused feature is upsampled and depthwise convolutional to obtain the first depth feature; The first additive feature is processed by convolution through three layers of convolutional blocks and then added to the first depth feature to obtain the third target fusion feature.

5. The smoke detection method based on contextual feature interaction according to claim 1, characterized in that, The feature fusion is performed based on the third fusion feature and the fourth fusion feature. Generate the fifth objective fusion feature, including: The fourth fusion feature is upsampled and depthwise convolved, and then added to the third fusion feature to obtain the second added feature; The second additive feature is processed by deep convolution through three layers of deep convolution blocks and then added to the fourth fusion feature to obtain the fifth target fusion feature.

6. The smoke detection method based on contextual feature interaction according to any one of claims 1 to 5, characterized in that, The process of target detection based on the target fusion features, obtaining category and location information of multiple targets, and calculating loss based on the category and location information to generate model loss values ​​includes: A decoupled detection head is used to perform target detection on the target fusion features to obtain the category information and location information of multiple targets; The binary cross-entropy loss function is used as the classification loss function and the regression loss function, and a preset loss function is used as the confidence loss function. Based on the category information and location information, the loss is calculated to obtain the classification loss value, the regression loss value and the confidence loss value. The classification loss value, the regression loss value, and the confidence loss value are added together to obtain the model loss value. The preset loss function is: ; in, For the true frame, For the prediction box, The smallest bounding box is the sum of the ground truth bounding box and the predicted bounding box. The confidence loss value is... For a single Power parameter.

7. A smoke detection device based on contextual feature interaction, characterized in that, include: smoke An image collection unit is used to collect a set of sample smoke images and preprocess each image in the set of sample smoke images to obtain training image data; An initial feature extraction unit is used to input the training image data into the detection network model, and to extract features from the training image data by introducing a feature extraction module with reparameterized large kernel convolution to obtain multi-layer initial features, wherein the initial features include a first layer initial features, a second layer initial features, and a third layer initial features; An initial feature fusion unit is used to construct a feature pyramid based on the initial features and perform feature fusion based on the feature pyramid to obtain initial fused features; The target feature fusion unit is used to perform context feature interaction enhancement fusion based on the initial fusion features through the context feature interaction enhancement module to obtain the target fusion features; The model loss calculation unit is used to perform target detection based on the target fusion features, obtain the category information and location information of multiple targets, and perform loss calculation based on the category information and location information to generate the model loss value; The model training unit is used to adjust the model parameters of the detection network model based on the model loss value, and to train the detection network model using backpropagation to obtain the target detection network model. The detection result generation unit is used to acquire the smoke image to be detected, and to detect the smoke image to be detected based on the target detection network model to generate the target detection result; The initial fusion feature includes a first fusion feature, a second fusion feature, a third fusion feature, and a fourth fusion feature, and the target feature fusion unit includes: The first fusion feature processing unit is used to perform convolution processing on the first fusion feature in the context feature interaction enhancement module to obtain the first target fusion feature; The second fusion feature processing unit is used to stack the first layer initial features with the first target fusion features to obtain the second target fusion features; The third fusion feature processing unit is used to perform feature fusion based on the first fusion feature, the third fusion feature and the fourth fusion feature to generate a third target fusion feature; The fourth fusion feature processing unit is used to perform convolution processing on the third fusion feature and then stack it with the second fusion feature to generate a fourth target fusion feature; The fifth fusion feature processing unit is used to perform feature fusion based on the third fusion feature and the fourth fusion feature to generate a fifth target fusion feature; The sixth fusion feature processing unit is used to perform convolution processing on the fourth fusion feature and then stack it with the third fusion feature to generate a sixth target fusion feature, wherein the target fusion feature includes the first target fusion feature, the second target fusion feature, the third target fusion feature, the fourth target fusion feature, the fifth target fusion feature and the sixth target fusion feature.

8. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the smoke detection method based on context feature interaction as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the smoke detection method based on context feature interaction as described in any one of claims 1 to 6.