Large industry model-based wildfire detection method, device, equipment and medium

By combining industry big models of convolutional neural networks and Transformer transformers, the problems of low efficiency and insufficient global feature capture in wildfire detection are solved, and efficient and accurate wildfire detection is achieved, which is suitable for real-time applications in resource-constrained environments.

CN120564047APending Publication Date: 2025-08-29SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510704389.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Existing wildfire detection methods have problems such as low efficiency, limited coverage, high cost and unstable performance in harsh weather conditions, and deep learning models have shortcomings in capturing global features and handling long-distance dependencies.

Method used

The industry large model based on hybrid neural networks is adopted, combined with convolutional neural networks and Transformer transformers, through slice sampling, bidirectional attention mechanism and feature fusion, the capture ability of image features is improved, the conversion of image features from spatial dimension to channel dimension is realized, and the calculation complexity is reduced through model pruning and quantization technology.

Benefits of technology

It improves the accuracy and operating efficiency of wildfire detection, is suitable for real-time detection in resource-constrained environments, reduces computing costs and storage requirements, and enhances the understanding of global features and the ability to capture local features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564047A_ABST
    Figure CN120564047A_ABST
Patent Text Reader

Abstract

The invention discloses a wildfire detection method and device based on an industry large model, equipment and a medium, and relates to the technical field of artificial intelligence. Comprising the steps of performing slice sampling operation on target wildfire image data through a backbone network of an industry large model, performing superposition processing on a plurality of obtained image slices to obtain converted image slices, and then performing feature extraction on the converted image slices to obtain extracted image slice features; carrying out feature fusion on the extracted image slice features based on a bidirectional attention mechanism through a neck network of an industry large model so as to obtain corresponding image feature data; and generating a corresponding detection result based on the image feature data through a detection head of the industry large model, and positioning a wildfire target in the target wildfire image data based on the detection result. Therefore, the wildfire detection precision and the operation efficiency can be improved by combining the advantages of the converter of the convolutional neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a wildfire detection method, device, equipment and medium based on an industry large model. Background Art

[0002] As global warming continues, wildfires have become a serious safety hazard, posing a significant threat to human life, property, and the natural environment. Traditional wildfire detection methods rely primarily on manual observation or advanced sensing technologies, such as infrared sensors and thermal imagers. However, these methods have significant limitations: manual observation is inefficient and prone to missed detections, while sensor coverage is limited, making it difficult to monitor large areas in real time. Furthermore, sensors are expensive to install and maintain, and their performance is unstable in adverse weather conditions. Therefore, the development of an efficient and reliable wildfire detection technology is urgent.

[0003] In recent years, deep learning technology has demonstrated tremendous potential for wildfire detection. By training on large-scale fire event datasets, deep learning models can significantly improve the accuracy and robustness of fire source detection. Currently, deep learning methods for wildfire detection fall into two main categories: those based on convolutional neural networks (CNNs) and those based on transformers. CNNs extract local features of images through convolution operations, possessing powerful local feature capture capabilities and are widely used in real-time detection tasks. However, CNNs have limitations in capturing global features and establishing long-range dependencies. Meanwhile, transformer-based models, while excelling at capturing global features and handling long-range dependencies through self-attention mechanisms, require high computational resources and are slow to train and infer, hindering their widespread application in real-time projects.

[0004] As can be seen from the above, how to improve the accuracy of wildfire detection and enhance operational efficiency is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the present invention aims to provide a wildfire detection method, apparatus, device, and medium based on an industry-wide model, which can improve wildfire detection accuracy and enhance operational efficiency by combining the advantages of convolutional neural network transformers. The specific solution is as follows:

[0006] In a first aspect, the present application provides a wildfire detection method based on an industry-wide model, wherein the industry-wide model is an industry-wide model based on a hybrid neural network, and the hybrid neural network is a network constructed based on a convolutional neural network and a Transformer transformer; wherein the method comprises:

[0007] Through the backbone network of the industry large model, a slice sampling operation is performed on the target wildfire image data to obtain a plurality of image slices corresponding to the target wildfire image data, and the plurality of image slices are superimposed to achieve the conversion of image features from the spatial dimension to the channel dimension to obtain converted image slices, and then feature extraction is performed on the converted image slices to obtain extracted image slice features; wherein the Transformer converter is located in the backbone network;

[0008] Through the neck network of the industry large model, the extracted image slice features are fused based on the bidirectional attention mechanism to obtain corresponding image feature data;

[0009] The detection head of the industry large model generates corresponding detection results based on the image feature data, and locates the wildfire target in the target wildfire image data based on the detection results.

[0010] Optionally, the backbone network consists of a slicing operation module, a depth-wise separable convolution and a bottleneck layer; the depth-wise separable convolution consists of a depth-wise convolution layer and a point-wise convolution; the bottleneck layer includes the Transformer transformer and the bottleneck residual convolution; the neck network consists of a C2f convolution layer, a bidirectional attention mechanism and a preset splicing module; the detection head is a decoupled head structure.

[0011] Optionally, performing a slice sampling operation on the target wildfire image data through the backbone network of the industry large model to obtain a plurality of image slices corresponding to the target wildfire image data, and performing a superposition process on the plurality of image slices to achieve conversion of image features from a spatial dimension to a channel dimension to obtain converted image slices, including:

[0012] Utilizing a slicing operation module in a backbone network of the industry large model to perform a slicing sampling operation on the target wildfire image data to obtain a plurality of image slices corresponding to the target wildfire image data;

[0013] The plurality of image slices are superimposed using the depthwise separable convolution in the backbone network of the industry large model to achieve conversion of image features from the spatial dimension to the channel dimension, so as to obtain converted image slices.

[0014] Optionally, the performing feature extraction on the converted image slice includes:

[0015] The Transformer transformer in the bottleneck layer of the backbone network of the industry model and the bottleneck residual convolution are used to capture the global context information of the converted image slices through layer normalization and sliding window multi-head self-attention mechanism to complete feature extraction of the converted image slices.

[0016] Optionally, the performing feature fusion on the extracted image slice features based on a bidirectional attention mechanism to obtain corresponding image feature data includes:

[0017] The C2f convolution layer and the preset splicing module in the neck network of the large industry model are used to strengthen the feature representation from the channel and spatial dimensions through the bidirectional attention mechanism to perform feature fusion to obtain the corresponding image feature data.

[0018] Optionally, after performing feature extraction on the converted image slice, the method further includes:

[0019] The extracted image slice features of different sizes are processed using a preset fast spatial pyramid pooling layer to generate extracted image slice features with a fixed dimension output, so that the extracted image slice features with a fixed dimension output are fused based on a bidirectional attention mechanism to obtain corresponding image feature data.

[0020] Optionally, the wildfire detection method based on the industry big model further includes:

[0021] Performing structured and unstructured pruning operations on the industry big model to remove redundant parameters and structures of the industry big model;

[0022] The model parameters of the industry large model after pruning are converted from 32-bit floating point numbers to 16-bit floating point numbers through FP16 quantization to obtain the target industry large model.

[0023] In a second aspect, the present application provides a wildfire detection device based on an industry large model, wherein the industry large model is an industry large model based on a hybrid neural network, and the hybrid neural network is a network constructed based on a convolutional neural network and a Transformer transformer; wherein the device includes:

[0024] An image slice feature extraction module is configured to perform a slice sampling operation on the target wildfire image data through the backbone network of the industry large model to obtain a plurality of image slices corresponding to the target wildfire image data, and perform a superposition process on the plurality of image slices to achieve conversion of image features from the spatial dimension to the channel dimension to obtain converted image slices, and then perform feature extraction on the converted image slices to obtain extracted image slice features; wherein the Transformer converter is located in the backbone network;

[0025] An image data feature extraction module is used to perform feature fusion on the extracted image slice features based on a bidirectional attention mechanism through the neck network of the industry large model to obtain corresponding image feature data;

[0026] The wildfire target positioning module is used to generate corresponding detection results based on the image feature data through the detection head of the industry large model, and to locate the wildfire target in the target wildfire image data based on the detection results.

[0027] In a third aspect, the present application provides an electronic device, comprising:

[0028] Memory, used to store computer programs;

[0029] A processor is used to execute the computer program to implement the aforementioned wildfire detection method based on the industry big model.

[0030] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned wildfire detection method based on the industry big model.

[0031] The present application provides a wildfire detection method based on an industry large model. First, a slice sampling operation is performed on the target wildfire image data through the backbone network of the industry large model to obtain a number of image slices corresponding to the target wildfire image data, and the several image slices are superimposed to realize the conversion of image features from spatial dimensions to channel dimensions to obtain converted image slices, and then feature extraction is performed on the converted image slices to obtain extracted image slice features; then, through the neck network of the industry large model, feature fusion is performed on the extracted image slice features based on a bidirectional attention mechanism to obtain corresponding image feature data; finally, through the detection head of the industry large model, a corresponding detection result is generated based on the image feature data, and the wildfire target in the target wildfire image data is located based on the detection result.

[0032] As can be seen from the above, this application constructs the industry-wide model using a hybrid network based on a convolutional neural network and a transformer. The industry-wide model then performs a slice sampling operation on target wildfire image data, overlaying the resulting image slices to transform image features from the spatial dimension to the channel dimension, generating transformed image slices. Feature extraction is then performed on the transformed image slices to obtain extracted image slice features. Specifically, the convolutional neural network extracts local image features, while the transformer captures global image features. Using the neck network of the industry-wide model, the extracted image slice features are fused using a bidirectional attention mechanism to obtain corresponding image feature data. Specifically, the local features extracted by the convolutional neural network and the global features captured by the transformer are fused to obtain image feature information. This approach not only retains the powerful local feature capture capabilities of CNNs, but also enhances understanding of global features through the transformer, thereby improving the model's overall detection performance. The bidirectional attention mechanism further enhances the model's accuracy, strengthening feature representation in both the channel and spatial dimensions. The channel attention mechanism evaluates the importance of each channel, generates weights, and applies them to the original feature map to highlight important features. The spatial attention mechanism focuses on specific locations in the feature map, similarly calculating weights to enhance the information representation of these areas. By combining the advantages of the transformer model of convolutional neural networks, this approach can improve wildfire detection accuracy and enhance operational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0034] Figure 1 A flow chart of a wildfire detection method based on an industry large model disclosed in this application;

[0035] Figure 2 This is a diagram showing the overall architecture of a large hybrid neural network model disclosed in this application;

[0036] Figure 3 A schematic diagram of a wildfire detection device based on an industry large model disclosed in this application;

[0037] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] In recent years, deep learning technology has demonstrated tremendous potential for wildfire detection. By training on large-scale fire event datasets, deep learning models can significantly improve the accuracy and robustness of fire source detection. Currently, deep learning methods for wildfire detection fall into two main categories: those based on convolutional neural networks (CNNs) and those based on transformers. CNNs extract local features of images through convolution operations, possessing powerful local feature capture capabilities and are widely used in real-time detection tasks. However, CNNs lack the ability to capture global features and establish long-range dependencies. On the other hand, transformer-based models excel at capturing global features and handling long-range dependencies through self-attention mechanisms, but their high computational resource requirements and slow training and inference speeds hinder widespread application in real-time projects. Therefore, this application proposes a wildfire detection solution based on industry-wide large-scale models that combines the advantages of transformers with convolutional neural networks to improve wildfire detection accuracy and operational efficiency.

[0040] See also Figure 1 As shown, the embodiment of the present application discloses a wildfire detection method based on an industry large model, wherein the industry large model is an industry large model based on a hybrid neural network, and the hybrid neural network is a network constructed based on a convolutional neural network and a Transformer transformer; wherein the method includes:

[0041] Step S11: Perform a slice sampling operation on the target wildfire image data through the backbone network of the industry large model to obtain a number of image slices corresponding to the target wildfire image data, and perform superposition processing on the several image slices to realize the conversion of image features from the spatial dimension to the channel dimension to obtain converted image slices, and then perform feature extraction on the converted image slices to obtain extracted image slice features.

[0042] In this embodiment, the industry large model is an industry large model based on a hybrid neural network, including a backbone network, a neck network and a detection head; the backbone network is composed of a slicing operation module, a depth-wise separable convolution and a bottleneck layer; the depth-wise separable convolution is composed of a depth-wise convolution layer and a point-by-point convolution; the bottleneck layer includes a transformer and a bottleneck residual convolution; the neck network is composed of a C2f convolution layer, a bidirectional attention mechanism and a preset splicing module; the detection head is a decoupling head structure.

[0043] It is worth mentioning that the embodiment of the present application significantly reduces the complexity of the model through model pruning and quantization technology, improves its operating efficiency on edge devices, and enables it to be efficiently deployed in resource-constrained environments. Specifically, structured and unstructured pruning operations are performed on the industry large model to remove redundant parameters and structures of the industry large model; the model parameters of the industry large model after the pruning operation are converted from 32-bit floating point numbers to 16-bit floating point numbers through FP16 quantization to obtain the target industry large model, which can further reduce storage requirements and computing overhead. In other words, the pruned and quantized model significantly improves the operating speed and energy efficiency while maintaining a high detection accuracy, which is very suitable for real-time wildfire detection applications in resource-constrained environments.

[0044] In this embodiment, the backbone network uses a slice sampling module to perform up and down sampling operations. The SliceSamp module consists of a slice operation and a depthwise separable convolution. The slice operation is responsible for slicing image features and superimposing the feature slices to achieve the conversion of image features from the spatial dimension to the channel dimension. The depthwise separable convolution consists of a depthwise convolution layer with a convolution kernel of 3×3 and a pointwise convolution with a convolution kernel of 1×1. Specifically, the backbone network of the industry model performs a slice sampling operation on the target wildfire image data to obtain a number of image slices corresponding to the target wildfire image data, and superimposes the image slices to achieve the conversion of image features from the spatial dimension to the channel dimension to obtain the converted image slices. This may include: using the slice operation module in the backbone network of the industry model to perform a slice sampling operation on the target wildfire image data to obtain a number of image slices corresponding to the target wildfire image data; and using the depthwise separable convolution in the backbone network of the industry model to superimpose the image slices to achieve the conversion of image features from the spatial dimension to the channel dimension to obtain the converted image slices. Furthermore, the transformer in the bottleneck layer of the backbone network of the large industry model and the bottleneck residual convolution are used to capture the global context information of the converted image slices through layer normalization and sliding window multi-head self-attention mechanism to complete the feature extraction of the converted image slices. First, the global context information is captured through layer normalization and sliding window multi-head self-attention mechanism, and then the global perception capability is further enhanced through a normalization layer and a multi-layer perceptron. That is, the front end uses the CNN structure to extract local features of the image, while the back end uses the transformer to capture global features. Not only does it retain the powerful local feature capture capability of CNN, but it also enhances the understanding of global features through the transformer, thereby improving the overall detection performance of the model.

[0045] Furthermore, embodiments of the present application utilize a preset fast spatial pyramid pooling layer to process extracted image slice features of different sizes to generate fixed-dimensional output extracted image slice features. These features are then fused based on a bidirectional attention mechanism to obtain corresponding image feature data. In other words, the integration of the fast spatial pyramid pooling layer (SPPF) enables processing images of different sizes to generate a fixed-dimensional output.

[0046] Step S12: Through the neck network of the industry large model, feature fusion is performed on the extracted image slice features based on the bidirectional attention mechanism to obtain corresponding image feature data.

[0047] In this embodiment, the neck network uses a C2f convolutional layer, a bidirectional attention (CS Attention) mechanism, and a splicing module to optimize the feature fusion process. Specifically, the bidirectional attention-based feature fusion of the extracted image slice features to obtain corresponding image feature data can include: utilizing the C2f convolutional layer and a preset splicing module in the neck network of the industry-wide model to enhance feature representation in both the channel and spatial dimensions through a bidirectional attention mechanism to fuse the features and obtain the corresponding image feature data. In other words, the CS Attention mechanism is introduced to enhance feature representation in both the channel and spatial dimensions. The channel attention mechanism evaluates the importance of each channel, generates weights, and applies them to the original feature map to highlight important features. The position attention mechanism focuses on specific locations in the feature map, similarly calculating weights to enhance the information representation of these areas. This dual mechanism effectively improves the model's accuracy in identifying fire characteristics.

[0048] Step S13: Generate corresponding detection results based on the image feature data through the detection head of the industry large model, and locate the wildfire target in the target wildfire image data based on the detection results.

[0049] In this embodiment, the detection head generates corresponding detection results based on the image feature data and locates the wildfire target in the target wildfire image data based on the detection results. The detection head adopts a decoupled head structure for anchor-free target positioning, ensuring accurate and flexible positioning.

[0050] As can be seen from the above, the embodiment of the present application utilizes the CNN structure to extract local features of the image, while the back-end uses a transformer to capture global features. Not only does it retain the powerful local feature capture capability of CNN, but it also enhances the understanding of global features through the transformer, thereby improving the overall detection performance of the model. The dual attention mechanism further improves the accuracy of the model and can strengthen feature representation from both channel and spatial dimensions. The channel attention mechanism evaluates the importance of each channel, generates weights and applies them to the original feature map to highlight important features. The position attention mechanism focuses on specific positions in the feature map, and also strengthens the information expression of these areas by calculating weights. Fast and accurate wildfire detection is achieved through an innovative hybrid neural network architecture, efficient feature extraction methods and optimization strategies.

[0051] See also Figure 2 As shown, the embodiment of the present application discloses a specific wildfire detection method based on an industry large model, including:

[0052] The embodiment of the present application adopts a hierarchical design and integrates multiple attention mechanisms. Its overall architecture consists of three main components: a backbone network, a neck network, and a detection head. The backbone network uses a slice sampling (SliceSamp) module for up and down sampling operations. The SliceSamp module consists of a slicing operation and a depthwise separable convolution. The slicing operation is responsible for slicing image features and superimposing the feature slices to achieve the conversion of image features from spatial dimensions to channel dimensions. The depthwise separable convolution consists of a depthwise convolution layer with a convolution kernel of 3×3 and a pointwise convolution with a convolution kernel of 1×1. The SwinBottle module combines the Swin Transformer and the bottleneck residual convolution for feature extraction. The Swin Transformer module first captures global context information through layer normalization and a sliding window multi-head self-attention mechanism, and then further enhances the global perception capability through a normalization layer and a multi-layer perceptron; at the same time, a fast spatial pyramid pooling layer (SPPF) is integrated to process images of different sizes and generate a fixed-dimensional output. The neck network uses YOLOv8's C2f convolutional layer and bidirectional attention (CS Attention) mechanism, as well as the ConcatBifpn module to optimize feature fusion. Finally, the detection head adopts a decoupled head structure for anchor-free object localization, ensuring accuracy and flexibility.

[0053] It is worth mentioning that in the wildfire detection scenario, the lightweight SliceSamp module reduces the computational cost while retaining the details of smoke and flames, which is particularly suitable for the real-time monitoring needs of edge devices such as drones and cameras. In addition, the SwinBottle module integrates global and local perception capabilities, enhances the capture of irregularly shaped smoke diffusion and flame edge features, effectively suppresses background interference, and reduces false alarm rates. The embodiment of the present application also uses the SPPF and ConcatBifpn modules to achieve effective cross-scale feature fusion and dynamic fire tracking, improve the single-frame inference speed and flame positioning accuracy under thick smoke cover. Compared with existing technologies such as YOLO, the embodiment of the present application not only performs better in small target detection and dynamic smoke processing, but also has higher multi-sensor compatibility and edge deployment efficiency.

[0054] Furthermore, this embodiment of the application applies test-time augmentation (TTA) technology to improve the stability and robustness of prediction results. Visual analysis of detection results in different application scenarios further validates the superior performance and adaptability of the large-scale model architecture in real-world deployments, demonstrating significant advantages in real-time wildfire detection in resource-constrained environments.

[0055] As can be seen from the above, the embodiments of this application combine the advantages of CNN and transformers to achieve multi-scale feature extraction, significantly improving the accuracy of wildfire detection. Model pruning and FP16 quantization techniques significantly reduce the model's parameter count and runtime, enabling efficient operation on edge devices and meeting the needs of real-time detection. Test-time augmentation (TTA) technology is applied to improve the stability and robustness of prediction results, demonstrating excellent performance and high adaptability in real-world deployments.

[0056] See also Figure 3 As shown, the embodiment of the present application discloses a wildfire detection device based on an industry large model, wherein the industry large model is an industry large model based on a hybrid neural network, and the hybrid neural network is a network constructed based on a convolutional neural network and a Transformer transformer; wherein the device includes:

[0057] An image slice feature extraction module 11 is configured to perform a slice sampling operation on the target wildfire image data through the backbone network of the industry large model to obtain a plurality of image slices corresponding to the target wildfire image data, and perform a superposition process on the plurality of image slices to achieve conversion of image features from the spatial dimension to the channel dimension to obtain converted image slices, and then perform feature extraction on the converted image slices to obtain extracted image slice features; wherein the Transformer converter is located in the backbone network;

[0058] An image data feature extraction module 12 is configured to perform feature fusion on the extracted image slice features based on a bidirectional attention mechanism through the neck network of the industry large model to obtain corresponding image feature data;

[0059] The wildfire target positioning module 13 is used to generate corresponding detection results based on the image feature data through the detection head of the industry large model, and to locate the wildfire target in the target wildfire image data based on the detection results.

[0060] Among them, the backbone network consists of a slicing operation module, a depth-wise separable convolution and a bottleneck layer; the depth-wise separable convolution consists of a depth-wise convolution layer and a point-by-point convolution; the bottleneck layer includes the Transformer transformer and the bottleneck residual convolution; the neck network consists of a C2f convolution layer, a bidirectional attention mechanism and a preset splicing module; the detection head is a decoupled head structure.

[0061] As can be seen from the above, the embodiment of the present application constructs the industry-wide model using a hybrid network based on a convolutional neural network and a transformer. The industry-wide model is then used to perform a slice sampling operation on target wildfire image data. The resulting image slices are then superimposed to convert image features from the spatial dimension to the channel dimension, resulting in transformed image slices. Feature extraction is then performed on the transformed image slices to obtain extracted image slice features. Specifically, the convolutional neural network extracts local image features, and the transformer captures global image features. The extracted image slice features are then fused using a bidirectional attention mechanism within the neck network of the industry-wide model to obtain corresponding image feature data. Specifically, the local features extracted by the convolutional neural network and the global features captured by the transformer are fused to obtain image feature information. This not only retains the powerful local feature capture capabilities of CNNs, but also enhances understanding of global features through the transformer, thereby improving the model's overall detection performance. The bidirectional attention mechanism further enhances the model's accuracy, strengthening feature representation in both the channel and spatial dimensions. The channel attention mechanism evaluates the importance of each channel, generates weights, and applies them to the original feature map to highlight important features. The spatial attention mechanism focuses on specific locations in the feature map, similarly calculating weights to enhance the information representation of these areas. By combining the advantages of the transformer model of convolutional neural networks, this approach can improve wildfire detection accuracy and enhance operational efficiency.

[0062] In some specific implementations, the image slice feature extraction module 11 may specifically include:

[0063] An image slice generation unit, configured to perform a slice sampling operation on the target wildfire image data using a slice operation module in the backbone network of the industry large model, so as to obtain a plurality of image slices corresponding to the target wildfire image data;

[0064] An image slice processing unit, configured to perform a stacking process on the plurality of image slices using a depthwise separable convolution in the backbone network of the industry large model to convert image features from a spatial dimension to a channel dimension, thereby obtaining converted image slices;

[0065] The image slice feature extraction unit is used to utilize the transformer in the bottleneck layer in the backbone network of the industry large model and the bottleneck residual convolution to capture the global context information of the converted image slice through layer normalization and sliding window multi-head self-attention mechanism to complete the feature extraction of the converted image slice.

[0066] In some specific implementations, the image data feature extraction module 12 may specifically include:

[0067] The image data feature extraction unit is used to utilize the C2f convolution layer and the preset splicing module in the neck network of the industry large model to enhance the feature representation from the two dimensions of channel and space through the bidirectional attention mechanism, so as to perform feature fusion and obtain the corresponding image feature data.

[0068] In some specific embodiments, the wildfire detection device based on the industry big model may further include:

[0069] a feature dimension fixing unit, configured to process extracted image slice features of different sizes using a preset fast spatial pyramid pooling layer to generate extracted image slice features outputted with a fixed dimension, so as to perform feature fusion on the extracted image slice features outputted with a fixed dimension based on a bidirectional attention mechanism to obtain corresponding image feature data;

[0070] An industry large model pruning unit, configured to perform structured and unstructured pruning operations on the industry large model to remove redundant parameters and structures of the industry large model;

[0071] The industry large model quantization unit is used to convert the model parameters of the industry large model after the pruning operation from 32-bit floating point numbers to 16-bit floating point numbers through FP16 quantization to obtain the target industry large model.

[0072] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the wildfire detection method based on the industry large model disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0073] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0074] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0075] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, NetWare, Unix, Linux, etc. In addition to including computer programs capable of implementing the industry model-based wildfire detection method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs capable of performing other specific tasks.

[0076] Furthermore, this application discloses a computer-readable storage medium for storing a computer program. When executed by a processor, the computer program implements the aforementioned wildfire detection method based on a large industry model. The specific steps of this method can be found in the corresponding content disclosed in the aforementioned embodiments and will not be further described here.

[0077] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0078] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0079] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0080] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0081] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A wildfire detection method based on an industry large model, characterized in that: The industry large model is an industry large model based on a hybrid neural network, and the hybrid neural network is a network constructed based on a convolutional neural network and a Transformer transformer; wherein the method includes: Through the backbone network of the industry large model, a slice sampling operation is performed on the target wildfire image data to obtain a plurality of image slices corresponding to the target wildfire image data, and the plurality of image slices are superimposed to achieve the conversion of image features from the spatial dimension to the channel dimension to obtain converted image slices, and then feature extraction is performed on the converted image slices to obtain extracted image slice features; wherein the Transformer converter is located in the backbone network; Through the neck network of the industry large model, the extracted image slice features are fused based on the bidirectional attention mechanism to obtain corresponding image feature data; The detection head of the industry large model generates corresponding detection results based on the image feature data, and locates the wildfire target in the target wildfire image data based on the detection results.

2. The wildfire detection method based on industry large model according to claim 1 is characterized in that: The backbone network consists of a slicing operation module, a depth-wise separable convolution and a bottleneck layer; the depth-wise separable convolution consists of a depth-wise convolution layer and a point-wise convolution; the bottleneck layer includes the Transformer transformer and the bottleneck residual convolution; the neck network consists of a C2f convolution layer, a bidirectional attention mechanism and a preset splicing module; the detection head is a decoupled head structure.

3. The wildfire detection method based on industry large model according to claim 2 is characterized in that: The backbone network of the industry large model performs a slice sampling operation on the target wildfire image data to obtain a plurality of image slices corresponding to the target wildfire image data, and performs a superposition process on the plurality of image slices to achieve conversion of image features from a spatial dimension to a channel dimension to obtain converted image slices, including: Utilizing a slicing operation module in a backbone network of the industry large model to perform a slicing sampling operation on the target wildfire image data to obtain a plurality of image slices corresponding to the target wildfire image data; The plurality of image slices are superimposed using the depthwise separable convolution in the backbone network of the industry large model to achieve conversion of image features from the spatial dimension to the channel dimension, so as to obtain converted image slices.

4. The wildfire detection method based on industry large model according to claim 2 is characterized in that: The performing feature extraction on the converted image slices comprises: The Transformer transformer in the bottleneck layer of the backbone network of the industry model and the bottleneck residual convolution are used to capture the global context information of the converted image slices through layer normalization and sliding window multi-head self-attention mechanism to complete feature extraction of the converted image slices.

5. The wildfire detection method based on industry large model according to claim 2 is characterized in that: The step of performing feature fusion on the extracted image slice features based on a bidirectional attention mechanism to obtain corresponding image feature data includes: The C2f convolution layer and the preset splicing module in the neck network of the large industry model are used to strengthen the feature representation from the channel and spatial dimensions through the bidirectional attention mechanism to perform feature fusion to obtain the corresponding image feature data.

6. The wildfire detection method based on industry large model according to claim 1 is characterized in that: After performing feature extraction on the converted image slice, the method further includes: The extracted image slice features of different sizes are processed using a preset fast spatial pyramid pooling layer to generate extracted image slice features with a fixed dimension output, so that the extracted image slice features with a fixed dimension output are fused based on a bidirectional attention mechanism to obtain corresponding image feature data.

7. The wildfire detection method based on industry large model according to any one of claims 1 to 6, characterized in that: Also includes: Performing structured and unstructured pruning operations on the industry big model to remove redundant parameters and structures of the industry big model; The model parameters of the industry large model after pruning are converted from 32-bit floating point numbers to 16-bit floating point numbers through FP16 quantization to obtain the target industry large model.

8. A wildfire detection device based on an industry large model, characterized in that: The industry large model is an industry large model based on a hybrid neural network, and the hybrid neural network is a network constructed based on a convolutional neural network and a Transformer transformer; wherein the device includes: An image slice feature extraction module is configured to perform a slice sampling operation on the target wildfire image data through the backbone network of the industry large model to obtain a plurality of image slices corresponding to the target wildfire image data, and perform a superposition process on the plurality of image slices to achieve conversion of image features from the spatial dimension to the channel dimension to obtain converted image slices, and then perform feature extraction on the converted image slices to obtain extracted image slice features; wherein the Transformer converter is located in the backbone network; An image data feature extraction module is used to perform feature fusion on the extracted image slice features based on a bidirectional attention mechanism through the neck network of the industry large model to obtain corresponding image feature data; The wildfire target positioning module is used to generate corresponding detection results based on the image feature data through the detection head of the industry large model, and to locate the wildfire target in the target wildfire image data based on the detection results.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the wildfire detection method based on the industry large model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, it implements the wildfire detection method based on the industry large model as described in any one of claims 1 to 7.