Flame detection method based on visual large model

Through feature extraction, bidirectional feature interaction and adaptive point representation of large visual models, combined with multi-task loss functions, the problem of insufficient fire source detection accuracy of traditional fire detection systems in complex backgrounds is solved, and high-precision and robust flame detection is achieved.

CN120765901APending Publication Date: 2025-10-10CHINA YANGTZE POWER
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510793634.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Traditional fire detection systems lack accuracy in detecting small-scale fires in complex backgrounds. Existing deep learning models find it difficult to effectively capture detailed information about fire sources and are easily affected by environmental interference.

Method used

A flame detection method based on a large visual model is adopted to improve the accuracy and robustness of fire source positioning by combining a feature extractor with a context aggregation module, a two-way feature interaction module, a dynamic adaptive point representation module and a multi-task loss function.

Benefits of technology

The accuracy and robustness of fire source detection are improved in complex backgrounds, making it suitable for real-time monitoring and fire warning in scenarios such as hydropower plants, reducing false detections and missed detections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765901A_ABST
    Figure CN120765901A_ABST
Patent Text Reader

Abstract

The invention discloses a flame detection method based on a visual large model. The method comprises the following steps: extracting a multi-scale context semantic enhancement feature map through a feature extractor in combination with a context aggregation module; the fire source positioning accuracy is improved by fusing multi-scale information through cross-layer feature two-way convolution transmission by using a two-way feature interaction module; sensing geometric features of a fire source through a dynamic adaptive point representation module, and screening samples to suppress interference in combination with density clustering and spatial constraint; a multi-task loss function containing classification, regression (generalized IoU and center distance IoU loss), multi-scale supervision and spatial constraint is constructed, and a model is trained in combination with an AdamW optimizer and a dynamic attenuation learning rate strategy. According to the method, small-scale fire source features can be effectively captured, complex background false detection is reduced, the detection precision on multiple data sets is superior to that of a traditional method, and the method is suitable for real-time monitoring and fire early warning of scenes such as hydraulic power plants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of flame detection based on a large visual model, and in particular relates to a flame detection method based on a large visual model. Background Art

[0002] Traditional fire detection systems primarily rely on smoke or flame detectors. While these systems can detect fires to a certain extent, their accuracy is significantly challenged when dealing with small-scale fires in complex backgrounds. Early fire detection methods were primarily based on simple image processing and machine learning techniques. For example, they utilized color information, shape changes, and optical flow analysis to assess the spread and direction of fire sources. While these methods can improve the efficiency and accuracy of fire detection in certain scenarios, they typically require professional involvement in feature design and screening. This human intervention is not only time-consuming but also affects the ultimate effectiveness of detection. Traditional methods rely on single pixels as sensing units, resulting in limited detection accuracy and segmentation effects.

[0003] In recent years, the rapid development of deep learning technology has provided new approaches for accurate fire location and detection. Object detection methods based on convolutional neural networks (CNNs) increase network depth to capture detailed features of fire sources, avoiding the errors caused by relying solely on expert experience. These methods improve the representation of fire source features by increasing the network's receptive field and focusing on prominent areas of the fire source. However, the inherent receptive field limitations of convolutional networks make it difficult to fully capture detailed information about the fire source, especially when detecting small-scale changes in the fire source. Compared to traditional deep learning models, the Transformer architecture, with its powerful feature extraction capabilities, can more accurately locate the fire source area. The Transformer model utilizes a self-attention mechanism to fully utilize global information during feature extraction, improving fire detection accuracy. In fire scenes with complex image semantics, the Transformer can better model the differences between the characteristics of the fire source area and surrounding tissue structures. However, when fire images have complex semantic content and are subject to external environmental interference such as lighting and weather, traditional methods struggle to capture long-range features of the fire source and can also lose detailed semantic information during feature information transfer. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a flame detection method based on a large visual model, which can effectively capture the characteristics of small-scale fire sources, reduce false detections in complex backgrounds, and has better detection accuracy than traditional methods on multiple data sets. It is suitable for real-time monitoring and fire warning in scenarios such as hydropower plants.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: A flame detection method based on a large visual model, S1, extracts an enhanced feature map containing multi-scale contextual semantics from an input image through a feature extractor combined with a context aggregation module; S2. Using the bidirectional feature interaction module, through bidirectional convolution transfer of cross-layer features, multi-scale context information is integrated to improve the accuracy of fire source positioning; S3, through the dynamic adaptive point representation module to perceive the geometric characteristics of the fire source, combined with density clustering and spatial constraints, to screen effective samples and suppress interference points; S4. Construct a multi-task loss function to supervise model training and combine it with dynamic optimization strategy to improve detection accuracy and generalization ability.

[0006] Preferably, the sub-steps of S1 are: S1.1. The input image is convolved with the feature extractor to generate a multi-level feature map containing basic visual information. S1.2. A context aggregation module with a masking strategy is embedded in the patch merging stage of each feature extractor layer. This module uses a masking mechanism to filter out redundant background information, focusing on the salient features of the flame. The module then divides the feature map into multiple subspaces, each of which independently learns different feature dimensions, enhancing the contextual semantic representation of the fire source area. S1.3. The enhanced feature map is further divided into four independent subspaces, corresponding to different scale characteristics of the fire source, to achieve refined feature capture of small-scale fire sources.

[0007] Preferably, the sub-steps of S2 are: S2.1. Upsample or downsample feature maps at different levels to the same resolution, providing a unified data basis for cross-layer interaction. S2.2. Perform bidirectional convolution transfer through the bidirectional feature interaction module: Forward transfer: The low-level feature map is transferred upward layer by layer through convolution, integrating the detailed information into the high-level features and enhancing the detailed expression of the high-level features; Backward transfer: The high-level feature map is transferred downward layer by layer through convolution, integrating semantic information into the low-level features and enhancing the semantic discrimination ability of the low-level features; 2.3. By alternately aggregating forward and reverse information flows, we establish interactive relationships between features at different levels, achieve the fusion of multi-scale contextual information, improve feature semantic consistency, and reduce false detections in complex backgrounds.

[0008] Preferably, the sub-steps of S3 are: 3.1. In the detection phase, a dynamic adaptive point representation module is introduced to assign a direction offset parameter and a weight coefficient to each sampling point. Deformable convolution is used to adaptively capture the geometric features of the fire source area, improving positioning accuracy in complex backgrounds. 3.2. Use density clustering algorithm to perform cluster analysis on adaptive points and divide them into three types of samples: Positive samples: high-density clusters corresponding to real flame areas, used by the model to learn effective detection features; Negative samples: low-density background points, corresponding to non-flame areas, used by the model to distinguish the background; Discrete samples: isolated outliers, corresponding to noise or interference areas, require further constraint processing; 3.3. Apply a weighted loss function with spatial constraints to discrete samples: By calculating the spatial distance between discrete points and positive samples, a penalty weight is imposed on points that deviate far from the cluster, suppressing their interference with the detection results and enhancing the robustness of the model in complex scenarios.

[0009] Preferably, the sub-steps of S4 are: 4.1. Construct a joint loss function consisting of four parts: Classification loss L cls :Use cross entropy loss to supervise the classification accuracy of the model for flame and non-flame areas and output category probability distribution; Regression loss L reg : Combine the generalized IoU loss and the center distance IoU loss to optimize the position matching between the predicted bounding box and the real box, and output accurate positioning coordinates; Multi-scale supervision loss L ms : Introducing auxiliary loss functions at each stage of feature extraction to prevent gradient vanishing and ensure the stability of feature learning at each layer; Spatial constraint loss L sc : Penalize the spatial position of discrete sample points to enhance the model's ability to focus on the flame area; 4.2. Use the AdamW optimizer with an initial learning rate of 0.00025 and weight decay to prevent overfitting. After every 6 rounds of training, the learning rate is dynamically decayed at a fixed rate to prevent the model from falling into the local optimum in complex feature space; Through a maximum of 36 rounds of iterative training, the parameters of components such as the feature extractor, bidirectional interaction module, and adaptive point representation module are gradually optimized, and finally a flame detection model with robust detection capabilities is generated.

[0010] Preferably, the joint loss function formula is: ; : Regression loss weight coefficient, multi-scale supervision loss weight coefficient, spatial constraint loss weight coefficient.

[0011] Preferably, the classification loss formula: ; N : The total number of samples, including positive samples, negative samples and discrete samples; : true label of the sample; : The model’s predicted value of flame probability for sample i, generated by the classification branch output by the dynamic adaptive point representation module.

[0012] Preferably, the regression loss formula: ; Generalized IoU loss : ; B : The real flame bounding box is obtained from data annotation; : The flame bounding box predicted by the model, output by the regression branch of the dynamic adaptive point representation module; C :Contains B and The minimum enclosing rectangle of ;. Center distance IoU loss : ; b : The center point of the ground truth bounding box; : Predict the center point of the bounding box, calculated by the regression branch of the model; d: The diagonal length of the minimum circumscribed rectangle C.

[0013] Preferably, the multi-scale supervision loss formula: ; : Corresponding to Stages1~Stages4; : Auxiliary loss of the k-th layer feature, where: Stages 1 is low-level features: focusing on edges and colors, Use regression loss to optimize detail features; Stages 2-3 are mid-level features: focusing on texture and dynamics. Classification loss is used to distinguish between flames and interference; Stage 4 is a high-level feature: focusing on global semantics, Use GIoU+DIoU loss to optimize positioning; : hierarchical weights.

[0014] Preferably, the spatial constraint loss Formula: ; M: number of discrete sample points; : discrete point coordinates; G: positive sample geometric center; : penalty weight.

[0015] The present application can achieve the following beneficial effects: The innovation of the present application in the field of fire detection includes the organic combination of multiple modules and methods. First, it adopts a context aggregation module combined with a mask strategy to enhance the attention ability of the feature extraction stage to the fire source area. Specifically, this module divides the feature map into multiple subspaces for independent learning, thereby suppressing irrelevant noise while extracting significant features, and improving the model's ability to capture the context semantics of the fire source area. Second, the present application innovatively introduces a bidirectional feature interaction module, which aggregates and interacts between different levels of features through bidirectional information flow, establishing a cooperative relationship between global and local features. This bidirectional pyramid convolution strategy not only captures multi-scale fire source information, but also enhances the model's adaptability to complex fire source areas. In addition, the model contains a dynamic adaptive point representation and a discrete point space constraint strategy, which uses adaptive direction conversion and density clustering algorithm to filter out the optimal detection points, and uses a spatial constraint loss function to punish the discrete points. Through this strategy, the present application can accurately locate the fire source and filter out the interference points in the complex background. Finally, the present application designs a multi-task loss function, including classification, regression, multi-scale supervision and spatial constraint. The loss function performs multi-scale supervision at each stage of feature extraction to ensure the stability and accuracy of the model, and uses spatial constraint loss to constrain the discrete points, thereby strengthening the overall performance of fire source detection.

[0016] This invention offers significant advantages in accuracy, adaptability, and computational efficiency. First, experiments on various fire datasets demonstrate that the model's detection accuracy and robustness surpass those of traditional methods. Through multi-scale feature aggregation, the invention effectively captures small-scale fire source information, achieving a deep fusion of global and local features, enabling the model to demonstrate higher accuracy in detecting small fire sources. Second, the invention exhibits robust adaptability to complex backgrounds. Dynamically adaptive point representation and spatial constraint strategies enable the model to adapt to complex background interference in diverse fire scenarios, significantly reducing false and missed detections caused by environmental factors such as lighting and color. Furthermore, despite incorporating multiple enhancement modules, its detection efficiency remains excellent, meeting high real-time requirements and enabling its application in real-time monitoring of fire early warning systems. By introducing multi-task supervision into the loss function, the model's multi-scenario generalization capabilities are significantly enhanced, making it applicable not only to forest fire monitoring but also to other scenarios requiring high-precision fire source detection. Overall, the invention strikes a good balance between accuracy, efficiency, and robustness, providing an advanced and efficient solution for fire early warning and emergency response, demonstrating strong application potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below with reference to the accompanying drawings and examples: Figure 1 This is the overall architecture diagram of the visual large-scale model flame detection method of the present invention; Figure 2 Flowchart of the present invention; Figure 3 This is a comparison chart of the detection accuracy of multiple data sets of the present invention. DETAILED DESCRIPTION

[0018] The preferred solution is Figures 1 to 3 As shown in the figure, a flame detection method based on a large visual model is proposed. First, the monitoring cameras and sensors of the hydropower plant are used to collect flame images and other environmental image data, and data enhancement is performed. Subsequently, based on the multi-scale feature extraction of SWin Transformer and Mask-CAM, and the CBPCM bidirectional convolution module, the model can effectively extract flame features and enhance the flame detection capability at multiple scales. Through the dynamic adaptive point learning module, the model can flexibly adjust the geometric shape of the flame area, screen key points, reduce environmental interference, and then perform sample screening and spatial constraints to ensure detection accuracy. After the model is deployed in the hydropower plant, the monitoring camera images are analyzed in real time. When a flame is detected, the alarm system is immediately triggered, and the fire-fighting equipment can be linked to realize intelligent flame warning and emergency response. Regular data updates and model optimization can ensure its high efficiency in various environments, making it an important part of fire protection in hydropower plants. Specifically, it includes the following steps: S1, extracting enhanced feature maps containing multi-scale contextual semantics from the input image through the feature extractor combined with the context aggregation module; S1.1. The input image is convolved with a feature extractor (e.g., Swin Transformer) to generate a multi-level feature map containing basic visual information (e.g., edges, colors). S1.2. A context aggregation module with a masking strategy is embedded in the patch merging stage of each feature extractor layer. This module uses a masking mechanism to filter out redundant background information, focusing on salient flame features (such as dynamic texture). The module also partitions the feature map into multiple subspaces, each of which independently learns different feature dimensions (such as color and shape subspaces), enhancing the contextual semantic representation of the fire source area. S1.3. The enhanced feature map is further divided into four independent subspaces, corresponding to different scale characteristics of the fire source (such as global flame shape and local flame flicker details), to achieve refined feature capture of small-scale fire sources.

[0019] First, a feature extractor extracts multi-level features from the input image. To enhance detail capture, the feature module embeds a context aggregation module within each layer's patch merging. This module uses a masking strategy to focus on salient features while reducing redundant information transfer. By partitioning the features into multiple subspaces and independently learning different features in each subspace, this module effectively enhances the model's ability to represent context in the fire source area.

[0020] The extracted feature map is further divided into four subspaces, each focusing on different scale features of the fire source area. This can more effectively capture the global and local information of the fire source, making the model more accurate in detecting small-scale fire sources.

[0021] Building on traditional pyramid convolution, this paper employs a bidirectional feature interaction module that establishes interactive relationships between features at different levels through a bidirectional information flow mechanism. First, features at each layer are scaled up or down to the same resolution, and then transferred through forward and backward convolutions to enhance the interactivity between features. This bidirectional information flow not only captures multi-scale contextual information but also reduces potential misjudgments in complex backgrounds.

[0022] S2. Using the bidirectional feature interaction module, through bidirectional convolution transfer of cross-layer features, multi-scale context information is integrated to improve the accuracy of fire source positioning; S2.1. Upsample or downsample feature maps at different levels (e.g., low-level detail features, high-level semantic features) to the same resolution, providing a unified data foundation for cross-layer interaction. S2.2. Perform bidirectional convolutional transfer via the bidirectional feature interaction module (CBPCM): Forward transfer: The low-level feature map is transferred upward layer by layer through convolution, incorporating detailed information (such as flame edges) into high-level features, enhancing the detailed expression of high-level features; Backward transfer: The high-level feature map is transferred downward layer by layer through convolution, incorporating semantic information (such as the semantics of the "flame" category) into the low-level features, thereby enhancing the semantic discrimination ability of the low-level features; 2.3. By alternately aggregating forward and backward information flows, we establish interactive relationships between features at different levels, achieve the fusion of multi-scale contextual information (such as the association between the global flame area and the local flame dynamics), improve feature semantic consistency, and reduce false detections in complex backgrounds.

[0023] By utilizing bidirectional feature interaction to maintain forward and reverse information flow in cross-layer aggregation, the model can fully fuse features at different levels when capturing the fire source area, thereby enhancing the consistency of contextual semantics and improving the accuracy of fire source positioning.

[0024] S3, through the dynamic adaptive point representation module to perceive the geometric characteristics of the fire source, combined with density clustering and spatial constraints, to screen effective samples and suppress interference points; 3.1. A dynamic adaptive point representation module is introduced during the detection phase. Each sampling point is assigned a direction offset parameter (describing the changing trend of the flame shape) and a weight coefficient (indicating the importance of the feature). Deformable convolution (DCNv2+) is used to adaptively capture the geometric features of the fire source area (such as irregular flame edges), improving positioning accuracy in complex backgrounds. 3.2. Use density clustering algorithm (such as DBSCAN) to cluster the adaptive points and divide them into three types of samples: Positive samples: high-density clusters corresponding to real flame areas, used by the model to learn effective detection features; Negative samples: low-density background points, corresponding to non-flame areas, used by the model to distinguish the background; Discrete samples: isolated outliers, corresponding to noise or interference areas, require further constraint processing; 3.3. Apply a weighted loss function with spatial constraints to discrete samples: By calculating the spatial distance between discrete points and positive samples, a penalty weight is imposed on points that deviate far from the cluster, suppressing their interference with the detection results and enhancing the robustness of the model in complex scenarios.

[0025] To better adapt to changes in fire source detection scenarios, the present invention incorporates a dynamic adaptive point representation module during the detection phase. This module assigns a directional offset and weight coefficient to each sampling point, enabling the model to accurately perceive the geometric characteristics of the fire source area. Furthermore, the adaptive direction conversion function helps the model more accurately locate the fire source, significantly enhancing detection effectiveness, especially in complex backgrounds.

[0026] Next, using a density clustering algorithm, the module divides the adaptive points into positive, negative, and discrete samples. Discrete points are penalized using a spatially constrained loss function, aiming to reduce their impact on overall detection accuracy and further improve the model's detection capabilities in complex scenarios.

[0027] S4. Construct a multi-task loss function to supervise model training and combine it with dynamic optimization strategy to improve detection accuracy and generalization ability.

[0028] 4.1. Construct a joint loss function consisting of four parts: Classification loss (L cls ): Using cross entropy loss, the supervised model classifies the flame and non-flame areas accurately and outputs the category probability distribution; Regression loss (L reg ): Combine the generalized IoU loss (GIoU Loss) and the center distance IoU loss (DIoU Loss) to optimize the position matching between the predicted bounding box and the real box, and output accurate positioning coordinates; Multi-scale supervision loss (L ms ): Introducing auxiliary loss functions at each stage of feature extraction (such as different layers of Swin Transformer) to prevent gradient disappearance and ensure the stability of feature learning at each layer; Spatial constraint loss (L sc ): Penalize the spatial position of discrete sample points to enhance the model's ability to focus on the flame area; 4.2. Use the AdamW optimizer with an initial learning rate of 0.00025 and weight decay to prevent overfitting. After every 6 rounds of training, the learning rate is dynamically decayed at a fixed rate (e.g., multiplied by 0.1) to prevent the model from falling into a local optimum in a complex feature space. Through a maximum of 36 rounds of iterative training, the parameters of components such as the feature extractor, bidirectional interaction module, and adaptive point representation module are gradually optimized, and finally a flame detection model with robust detection capabilities is generated.

[0029] Joint loss function formula: ; : Regression loss weight coefficient, multi-scale supervision loss weight coefficient, spatial constraint loss weight coefficient.

[0030] Classification loss formula:

[0031] Parameters introduced: N : corresponds to the “total number of samples” in the invention, including positive samples, negative samples, and discrete samples (the document does not specify the sample type ratio, which needs to be dynamically determined based on the density clustering results); : The true label of the sample, which is the flame area (1) or the non-flame area (0) in the invented scene; : The model’s predicted value of flame probability for sample i, generated by the classification branch output by the dynamic adaptive point representation module.

[0032] Regression loss formula: ; 1. Generalized IoU Loss

[0033] ; B: The actual flame bounding box, obtained from data annotation (e.g., the flame area annotation in the hydropower plant monitoring video); : The flame bounding box predicted by the model, output by the regression branch of the dynamic adaptive point representation module; C: contains B and The minimum bounding rectangle corresponds to the "global area under complex background" in the invention; 2. Center distance IoU loss : ; b: The center point of the real bounding box, corresponding to the "geometric center of the fire source area" in the invention; : Predict the center point of the bounding box, calculated by the regression branch of the model; d: The diagonal length of the minimum circumscribed rectangle C, reflecting the "spatial scope of the complex background" in the invention.

[0034] Multi-scale supervision loss formula: ; :Corresponding to the invention Figure 1Stages1~Stages4 in (i.e., the 4 levels of Swin Transformer); : The auxiliary loss of the k-th layer feature can be set as follows according to the invention of “introducing multi-scale supervision at each stage of feature extraction”: Stages1 (low-level features): focus on edges and colors, Use regression loss to optimize detail features; Stages 2-3 (middle-level features): focus on texture and dynamics, Use classification loss to distinguish between flames and interference; Stage 4 (high-level features): Focus on global semantics, Use GIoU+DIoU loss to optimize positioning; : Level weight, according to the requirement of “preventing gradient disappearance” in the invention, the low-level feature weight can be set to (like ).

[0035] Spatial Constraint Loss formula: ; M: the number of discrete sample points, obtained by the "density clustering algorithm" in the invention (such as isolated points outside the density threshold in the DBSCAN algorithm); : coordinates of discrete points, corresponding to the “discrete points that interfere with overall performance” in the invention; G: The geometric center of the positive sample, calculated as , where P is the number of positive sample points (high-density flame area points determined by density clustering); : Penalty weight, according to the "space constraint strategy" in the invention, can be set as (The lower the density of discrete points, the higher the weight).

[0036] This paper introduces a multi-task loss function to supervise the model's performance in fire source detection. First, a classification loss ensures that the model correctly identifies the fire source area. Second, a regression loss uses a generalized IoU loss and a center distance IoU loss to optimize the position of the bounding box, resulting in more accurate fire source localization.

[0037] To prevent the network from falling into local optima due to vanishing or exploding gradients during feature extraction, the model incorporates multi-scale supervised losses at each stage of feature extraction. Furthermore, a spatially constrained loss penalizes discrete sample points, further improving the model's accuracy in capturing the fire source area. This not only enhances the model's focus on the fire source area but also ensures overall detection stability and generalization capabilities.

[0038] During model training, the AdamW optimizer was used, with an initial learning rate of 0.00025. After every six rounds of training, the learning rate was dynamically decayed to ensure the model did not fall into local optima in complex feature spaces. Through a maximum of 36 rounds of iterative training, the model effectively optimized feature extraction, point representation, and spatial constraint strategies, ultimately achieving robust fire source detection.

[0039] This paper combines context aggregation, bidirectional feature interaction, adaptive direction point allocation, and multi-task loss optimization to build a highly accurate and robust fire source detection system. Experimental results on various datasets demonstrate its excellent detection performance in complex fire source environments, making it an effective tool for fire early warning systems.

[0040] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions set forth in the claims, including equivalent alternatives to the technical features of the technical solutions set forth in the claims. In other words, equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A flame detection method based on a large visual model, characterized by: The following steps are involved: S1, extracts enhanced feature maps containing multi-scale contextual semantics from the input image through the feature extractor combined with the context aggregation module; S2. Using the bidirectional feature interaction module, through bidirectional convolution transfer of cross-layer features, multi-scale context information is integrated to improve the accuracy of fire source positioning; S3, through the dynamic adaptive point representation module to perceive the geometric characteristics of the fire source, combined with density clustering and spatial constraints, to screen effective samples and suppress interference points; S4. Construct a multi-task loss function to supervise model training and combine it with dynamic optimization strategy to improve detection accuracy and generalization ability.

2. The flame detection method based on a visual large model according to claim 1, characterized in that: The sub-steps of S1 are: S1.

1. The input image is convolved with the feature extractor to generate a multi-level feature map containing basic visual information. S1.

2. A context aggregation module with a masking strategy is embedded in the patch merging stage of each feature extractor layer. This module uses a masking mechanism to filter out redundant background information, focusing on the salient features of the flame. The module then divides the feature map into multiple subspaces, each of which independently learns different feature dimensions, enhancing the contextual semantic representation of the fire source area. S1.

3. The enhanced feature map is further divided into four independent subspaces, corresponding to different scale characteristics of the fire source, to achieve refined feature capture of small-scale fire sources.

3. The flame detection method based on a visual large model according to claim 1, characterized in that: The sub-steps of S2 are: S2.

1. Upsample or downsample feature maps at different levels to the same resolution, providing a unified data basis for cross-layer interaction. S2.

2. Perform bidirectional convolution transfer through the bidirectional feature interaction module: Forward transfer: The low-level feature map is transferred upward layer by layer through convolution, integrating the detailed information into the high-level features and enhancing the detailed expression of the high-level features; Backward transfer: The high-level feature map is transferred downward layer by layer through convolution, integrating semantic information into the low-level features and enhancing the semantic discrimination ability of the low-level features; 2.

3. By alternately aggregating forward and reverse information flows, we establish interactive relationships between features at different levels and achieve the fusion of multi-scale contextual information. Improve feature semantic consistency and reduce false detection in complex backgrounds.

4. The flame detection method based on a visual large model according to claim 1, characterized in that: The sub-steps of S3 are: 3.

1. In the detection phase, a dynamic adaptive point representation module is introduced to assign a direction offset parameter and a weight coefficient to each sampling point. Deformable convolution is used to adaptively capture the geometric features of the fire source area, improving positioning accuracy in complex backgrounds. 3.

2. Use density clustering algorithm to perform cluster analysis on adaptive points and divide them into three types of samples: Positive samples: high-density clusters corresponding to real flame areas, used by the model to learn effective detection features; Negative samples: low-density background points, corresponding to non-flame areas, used by the model to distinguish the background; Discrete samples: isolated outliers, corresponding to noise or interference areas, require further constraint processing; 3.

3. Apply a weighted loss function with spatial constraints to discrete samples: By calculating the spatial distance between discrete points and positive samples, a penalty weight is imposed on points that deviate far from the cluster, suppressing their interference with the detection results and enhancing the robustness of the model in complex scenarios.

5. The flame detection method based on a visual large model according to claim 1, characterized in that: The sub-steps of S4 are: 4.

1. Construct a joint loss function consisting of four parts: Classification loss L cls :Use cross entropy loss to supervise the classification accuracy of the model for flame and non-flame areas and output category probability distribution; Regression loss L reg : Combine the generalized IoU loss and the center distance IoU loss to optimize the position matching between the predicted bounding box and the real box, and output accurate positioning coordinates; Multi-scale supervision loss L ms : Introducing auxiliary loss functions at each stage of feature extraction to prevent gradient vanishing and ensure the stability of feature learning at each layer; Spatial constraint loss L sc : Penalize the spatial position of discrete sample points to enhance the model's ability to focus on the flame area; 4.

2. Use the AdamW optimizer with an initial learning rate of 0.00025 and weight decay to prevent overfitting. After every 6 rounds of training, the learning rate is dynamically decayed at a fixed rate to prevent the model from falling into the local optimum in complex feature spaces; Through a maximum of 36 rounds of iterative training, the parameters of components such as the feature extractor, bidirectional interaction module, and adaptive point representation module are gradually optimized, and finally a flame detection model with robust detection capabilities is generated.

6. The flame detection method based on a visual large model according to claim 5, characterized in that: Joint loss function formula: ; : Regression loss weight coefficient, multi-scale supervision loss weight coefficient, spatial constraint loss weight coefficient.

7. The flame detection method based on a visual large model according to claim 1, characterized in that: Classification loss formula: ; N : The total number of samples, including positive samples, negative samples and discrete samples; : true label of the sample; : The model's predicted value of flame probability for sample i, generated by the classification branch output by the dynamic adaptive point representation module.

8. The flame detection method based on a visual large model according to claim 1, characterized in that: Regression Loss formula: ; Generalized IoU loss : ; B : The real flame bounding box is obtained from data annotation; : The flame bounding box predicted by the model, output by the regression branch of the dynamic adaptive point representation module; C :Contains B and The minimum enclosing rectangle of ;. Center distance IoU loss : ; b : The center point of the ground truth bounding box; : Predict the center point of the bounding box, calculated by the model regression branch; d: The diagonal length of the minimum circumscribed rectangle C.

9. The flame detection method based on a visual large model according to claim 1, characterized in that: Multi-scale supervision loss formula: ; : Corresponding to Stages1~Stages4; : Auxiliary loss of the k-th layer feature, where: Stages 1 is low-level features: focusing on edges and colors, Use regression loss to optimize detail features; Stages 2-3 are mid-level features: focusing on texture and dynamics. Classification loss is used to distinguish between flames and interference; Stage 4 is a high-level feature: focusing on global semantics, Use GIoU+DIoU loss to optimize positioning; : Level weight.

10. The flame detection method based on a visual large model according to claim 1, characterized in that: Spatial Constraint Loss formula: ; M: number of discrete sample points; : discrete point coordinates; G: geometric center of positive sample; : Penalty weight.

Citation Information

Cited By

  • Smoke and fire detection method and system based on deep reinforcement learning and feature fusion

    CN121659243A