Lightweight extreme fire identification method based on context-aware large kernel transformation

By improving the YOLOv8 network model, the context-aware large-core transformation module and lightweight decoupling detection head are introduced, which solves the problem of difficult to identify and distinguish different types of extreme fires in the prior art, and achieves fast, efficient and accurate extreme fire recognition, reduces the demand for computing resources, and maximizes the protection of human life and property safety.

CN120032219APending Publication Date: 2025-05-23STATE GRID FUJIAN ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411655213.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and distinguish different types of extreme fires, resulting in limited effectiveness of fire warning and response strategies.

Method used

The lightweight extreme fire recognition method based on context-aware large core transformation is adopted. By improving the YOLOv8 network model, the context-aware large core transformation module and lightweight decoupling detection head are introduced to optimize feature extraction and fusion, and the computing resource requirements are reduced.

Benefits of technology

It realizes fast, efficient and accurate identification of extreme fire types and locations, and can efficiently identify extreme fire types in real fire fields, reduce losses, and provide valuable time to make reasonable handling decisions, and maximize the protection of human life and property safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032219A_ABST
    Figure CN120032219A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight extreme fire identification method based on context-aware large kernel transformation. The method comprises the following steps: acquiring an extreme fire image; performing extreme fire type and extreme fire position labeling on the extreme fire image, generating an extreme fire data set, and dividing the extreme fire data set into a training set, a verification set and a test set; a decoupling detection head of the YOLOv8 network model is optimized, and a lightweight decoupling detection head L-Deect is formed; a bottleneck part in a C2f module of a YOLOv8 network model is improved into a context sensing module and a large-kernel transformation module, and a context sensing large-kernel transformation module CALKT is formed; a YOLOv8 network model is improved, CALKT and L-Deect are introduced to replace an original C2f module and an original Deect module of the model respectively, and an improved YOLOv8 network model is formed; performing training, parameter adjustment and evaluation on the improved YOLOv8 network model through the training set, the verification set and the test set; and S4, identifying the extreme fire by using the improved YOLOv8 network model. According to the method, the type and the position of the extreme fire can be quickly, efficiently and accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a lightweight extreme fire recognition method based on context-aware large kernel transformation. Background Art

[0002] Extreme fire is generally regarded as a natural disaster with significant impact, which will cause great damage to human society and forest ecosystems. Under different burning conditions, the intensity of extreme fire burning varies, and different phenomena will be presented, such as crown fire, fire line merging, flying fire, explosive fire, fire whirlwind, super fire, jumping fire, fire storm, etc. The behavioral characteristics of extreme fire include unpredictable changes in fire intensity, uncertainty in the spread speed and direction of advance, flying fire phenomenon, and strong winds and plumes accompanying fire. These changing behaviors pose a major threat to fire rescue personnel and may cause fire fighting work to fall short. In view of the significant threat of extreme fire to human life and property, it has become extremely urgent to develop technologies that can intelligently, accurately and in real time identify these extreme fires. It is worth noting that when facing special tasks such as extreme fires, the more complex the target detection algorithm model is, the more effective it is. Extreme fire scenes often face severe resource limitations, including equipment movement range, computing power and storage space, which requires that the model for extreme fire recognition must take into account the hardware limitations of embedded and mobile devices.

[0003] Current research on extreme fires mainly focuses on the identification and detection of conventional fires or larger-scale fires, but these studies can usually only determine whether a fire has occurred without further exploring extreme fires. There are very few patents on distinguishing between extreme fire categories. This limitation weakens the effectiveness of fire warning and response strategies, as different types of extreme fires may require different response measures. Summary of the invention

[0004] The object of the present invention is to provide a lightweight extreme fire identification method based on context-aware large kernel transformation, which is conducive to quickly, efficiently and accurately identifying the type and location of extreme fire.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is: a lightweight extreme fire identification method based on context-aware large kernel transformation, comprising the following steps:

[0006] Step S1: Acquire extreme fire images;

[0007] Step S2: annotating extreme fire types and extreme fire locations on extreme fire images to generate an extreme fire dataset, which is divided into a training set, a validation set, and a test set;

[0008] Step S3: using a lightweight structure to replace two redundant standard convolutional layers in the YOLOv8 network model, adjusting and optimizing the decoupled detection head structure of YOLOv8, and forming a lightweight decoupled detection head L-Detect; improving the bottleneck part in the C2f module of the YOLOv8 network model into two main parts: a context-aware module and a large-core transformation module, and forming a context-aware large-core transformation module CALKT; improving the YOLOv8 network model, introducing the context-aware large-core transformation module CALKT and the lightweight decoupled detection head L-Detect to replace the original C2f module and Detect module of the model, respectively, and forming an improved YOLOv8 network model; training, parameter adjustment and evaluation of the improved YOLOv8 network model are performed through training sets, validation sets and test sets;

[0009] Step S4: Use the trained improved YOLOv8 network model to identify extreme fires.

[0010] Furthermore, in step S1, extreme fire images are obtained from multiple sources, including indoor and outdoor experimental scenes of extreme fires, surveillance video data of wild fires, and network image search engines.

[0011] Furthermore, in step S2, Labelimg software is used to label extreme fire images with extreme fire categories and typical extreme fire locations to generate an extreme fire dataset.

[0012] Further, in step S2, the extreme fire types include crown fire, fire line merging, flying fire and fire whirlwind.

[0013] Furthermore, in step S2, data enhancement is performed on the acquired extreme fire image, including: horizontal and vertical flipping, affine transformation, random rotation, contrast enhancement, sharpening, noise addition and blurring; and then the extreme fire type and extreme fire location are labeled on the extreme fire image obtained by data enhancement.

[0014] Furthermore, in step S2, the extreme fire dataset is divided into a dataset according to the ratio of training set: validation set: test set = 8:1:1.

[0015] Furthermore, in step S3, the decoupled detection head structure of YOLOv8 is adjusted and optimized, specifically: the original standard convolution is replaced by the depthwise separable convolution with a 5×5 convolution kernel, so that the input feature map is reduced in dimension and lightweight through two branches.

[0016] Furthermore, in step S3, the bottleneck part in the C2f module of the YOLOv8 network model is improved into two main parts: one is a context-aware module, which first reduces the dimension of the input feature map through a 1×1 standard convolution, and then transmits the feature map in parallel to a 3×3 standard convolution and a 3×3 dilated convolution, and the two outputs are fused to obtain rich local and global details without adding additional parameters; the second is a large kernel transformation module, which adopts a large separable kernel module design, and is first composed of cascaded vertical and horizontal depthwise separable convolutions, respectively, and then cascaded vertical and horizontal large-size dilated depthwise separable convolutions to capture the geometric shape characteristics of extreme fire without adding too many parameters.

[0017] Furthermore, in step S3, the YOLOv8 network model is improved, specifically including: improving the original C2f module of the model to a context-aware large kernel transformation module, and adopting a more effective feature extraction and fusion method to achieve a wider effective receptive field and context information perception capability; improving the original Detect module of the model to a lightweight decoupled detection head L-Detect to reduce the parameter redundancy of the detection head, thereby reducing the computational requirements of the model in the inference stage.

[0018] Further, in step S4, the extreme fire image to be identified is input into the trained improved YOLOv8 network model to output the category and location of the extreme fire.

[0019] Compared with the prior art, the present invention has the following beneficial effects: the present invention provides a lightweight extreme fire identification method based on context-aware large kernel transformation, which combines context-awareness and large kernel transformation technology, focuses on capturing the detailed features of extreme fires, and takes into account the environmental information around extreme fires in a global scope, thereby achieving effective perception of global information, accurate identification of extreme fire features, and fast and efficient computational processing, and can efficiently identify the types of extreme fires in real fire scenes, which can not only significantly reduce the losses caused by extreme fires, but also provide valuable time for fire rescue personnel to make more reasonable and effective handling decisions, thereby maximizing the protection of human life and property safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of a method implementation of an embodiment of the present invention;

[0021] Figure 2 2 is a schematic diagram of the structure of a lightweight decoupling detection head L-Detect in an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of the structure of a context-aware large-kernel transformation module in an embodiment of the present invention;

[0023] Figure 4 It is a schematic diagram of the structure of the improved YOLOv8 network model in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0025] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0027] like Figure 1 As shown, this embodiment provides a lightweight extreme fire recognition method based on context-aware large kernel transformation, which generally includes: obtaining extreme fire images, constructing extreme fire data sets, designing a lightweight decoupled detection head L-Detect, constructing a context-aware large kernel transformation module CALTK, improving the YOLOv8 network model, model training, and extreme fire recognition. The implementation steps of this method are further described below.

[0028] Step S1: Acquire extreme fire images from multiple sources, including: indoor and outdoor experimental scenes of extreme fires, surveillance video data of wild fires, and network image search engines.

[0029] Step S2: Perform data enhancement on the acquired extreme fire images, and then annotate the extreme fire types and extreme fire locations on the extreme fire images obtained through data enhancement to generate an extreme fire dataset, which is divided into a training set, a validation set, and a test set.

[0030] Among them, data enhancement is performed on extreme fire images, including horizontal and vertical flipping, affine transformation, random rotation, contrast enhancement, sharpening, noise addition, and blurring. In this embodiment, in order to solve the problem of data imbalance and too few data samples, 2,330 images of crown fire, 1,422 images of fire line merging, 815 images of flying fire, 1,247 images of fire whirlwind, and a total of 5,814 typical extreme fire images were obtained through data enhancement.

[0031] Then, Labelimg software was used to annotate the extreme fire type and location of each extreme fire image to generate an extreme fire dataset. The extreme fire types include crown fire, fire line merging, flying fire, and fire whirl. The generated extreme fire dataset is an image dataset of extreme fire.

[0032] Finally, the extreme fire dataset was divided into training set: validation set: test set = 8:1:1 ratio.

[0033] Step S3: First, a lightweight structure is used to replace the two redundant standard convolutional layers in the YOLOv8 network model, and the decoupled detection head structure of YOLOv8 is adjusted and optimized to form a lightweight decoupled detection head L-Detect.

[0034] like Figure 2 As shown in the figure, this method uses a depth-separable convolution with a 5×5 convolution kernel to replace the original standard convolution, so that the input feature map is reduced in dimension and lightweight through two branches.

[0035] Secondly, the bottleneck part (Bottleneck) in the C2f module of the YOLOv8 network model is improved into two main parts: a context-aware module and a large kernel transformation module, forming a context-aware large kernel transformation module CALKT.

[0036] Among them, the context-aware module first reduces the dimension of the input feature map through a 1×1 standard convolution, and then transmits the feature map in parallel to a 3×3 standard convolution and a 3×3 dilated convolution. The two outputs are fused, and the model can obtain rich local and global details without adding additional parameters. The large kernel transformation module adopts the Large Separable Kernel (LSK) module design. The module is first composed of cascaded vertical (longitudinal) and horizontal (transverse) depthwise separable convolutions, and then cascaded vertical and horizontal large-size dilated depthwise separable convolutions. The introduction of the large kernel transformation module is conducive to the model capturing the geometric shape characteristics of extreme fire without adding too many parameters.

[0037] like Figure 3 As shown in the figure, given an input feature map X with a height of h, a width of w, and a number of channels of c, the following output will be obtained:

[0038]

[0039]

[0040] Where X, Y represent feature maps, W represents convolution kernel, k, d represent large kernel size and dilation rate respectively, [_,_] represents feature channel connection, * represents convolution operation, Represents the Hadamard product. Input feature map X 1 After passing through the context-aware module, we get the feature map X 4 , and then through the large-size separable kernel transformation to obtain X 7 Finally, the output feature map Y containing rich contextual information is obtained by fusing the original features with the transformed features.

[0041] The bottleneck structure has two stages of information flow. In the first stage, the input feature map is sent to the context-aware module. This module first performs dimensionality reduction through a 1×1 standard convolution. Subsequently, the feature map is transmitted to a 3×3 standard convolution and a 3×3 dilated convolution in parallel. The standard convolution part is a local information extraction submodule that focuses on capturing the local detail features of extreme fire images, while the dilated convolution is a global information extraction submodule that effectively captures the global environmental information in extreme fire images through its wider effective receptive field. By fusing the outputs of these two submodules, the model can simultaneously obtain rich local and global details without adding additional parameters. In the second stage, the features processed by the context-aware module are guided to the large kernel transformation module. This method adopts the Large Separable Kernel (LSK) module design, which first consists of cascaded vertical (longitudinal) and horizontal (transverse) depthwise separable convolutions, followed by cascaded vertical and horizontal large-size dilated depthwise separable convolutions. The introduction of the large kernel transformation module is conducive to the model capturing the geometric shape features of extreme fire without adding too many parameters. In addition, two residual connections are introduced in the module, one in the context-aware part and the other in the large kernel transformation part. The residual connection in the context-aware part helps the model capture additional context information while ensuring the original information flow of the input features. The residual connection in the large kernel transformation part helps fuse the features that have undergone complex transformations with the original features, maintaining the richness of the features while avoiding the loss of key information during the learning process.

[0042] On this basis, the YOLOv8 network model is improved, and the context-aware large kernel transformation module CALKT and the lightweight decoupled detection head L-Detect are introduced to replace the original C2f module and Detect module of the model respectively, forming an improved YOLOv8 network model. Specifically, the original C2f module of the model is improved to a context-aware large kernel transformation module, and a more effective feature extraction and fusion method is adopted to achieve a wider effective receptive field and context information perception capability; the original Detect module of the model is improved to a lightweight decoupled detection head L-Detect to reduce the parameter redundancy of the detection head, thereby reducing the computational requirements of the model in the inference stage.

[0043] like Figure 4As shown in the figure, the improved YOLOv8 network model structure is based on the inheritance of the benchmark model architecture, and the context-aware large kernel transformation module (CALKT) and lightweight decoupled detection head (L-Detect) designed by this method are used to replace the C2f module and the Detect module. The overall model consists of three main parts: the backbone network (Backbone), the neck (Neck) and the head (Head).

[0044] In Backbone, the extreme fire image of 640×640 pixels is first used as input, and then sent to the designed CALKT module after multiple downsampling processes. Then, after three convolution operations and two CALKT module processes, the image is converted into an 80×80 pixel feature map. At this time, the feature map not only incorporates the detailed information of the extreme flames, but also contains rich environmental context information. Subsequent downsampling and CALKT module processing further reduce the size of the feature map to 40×40 pixels. The CALKT module plays a role in enhancing the capture of the shape and texture information of the extreme fire. Finally, after the last downsampling, the feature map size is reduced to 20×20 pixels. The SPPF module further enhances the model's ability to express features without changing the scale.

[0045] In the Neck part, the 20×20 pixel abstract feature map is fused with the 40×40 and 80×80 pixel feature maps in the Backbone through bottom-up upsampling. Then it is fused into a larger scale feature map through top-down downsampling. The embedded CALKT module provides the model with rich bidirectional position information through its cascaded large-scale depthwise separable convolution, significantly reducing the computational cost of the model, while helping the model capture rich detail features and extensive contextual information, thereby comprehensively improving the model's extreme fire recognition ability.

[0046] Finally, in the Head part, the model recognizes and classifies feature maps of three different scales. The improved lightweight decoupled detection head L-Detect replaces the original Detect detection head module, thereby achieving high efficiency in the model reasoning process and improving detection speed.

[0047] In this embodiment, the model uses classification loss and regression loss as its main loss functions. Among them, the classification loss function uses the variable focus loss function VFL, which can more accurately determine whether a certain type of extreme fire belongs to this category. The formula of the VFL loss function is as follows:

[0048]

[0049] Where q is the target IoU score, p is the IoU-Aware Classification Score (IACS) predicted by the model, and α and γ are weight factors, which are set to 0.75 and 2.0 respectively.

[0050] The regression loss function uses CIoU loss and distribution focus loss DFL. CIoU loss is used to enhance the performance of the model in bounding box regression. The CIoU formula is as follows:

[0051]

[0052] Where d represents the distance between the centers of the predicted box and the target box, and c represents the distance between the diagonals of their minimum bounding rectangles. gt Represent the aspect ratios of the prediction box and the target box respectively.

[0053] DFL further enhances the model’s prediction accuracy for the target location by finely adjusting the probability distribution near the predicted value, making the prediction box more closely surround the target object. The DFL formula is as follows:

[0054] DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(yy i )log(S i+1 ))

[0055] Among them, y represents the label, S i and S i+1 Represent the predicted value y i and i+1 The probability of being at

[0056] The total loss of the model is obtained by the weighted sum of the classification loss and regression loss. The calculation formula of the total loss of the model is as follows:

[0057] Loss = λ 1 L VFL +(λ 2 L CIoU +λ 3 L DFL )

[0058] Among them, λ 1 ,λ 2 and λ 3 represents weight factors, which are set to 0.5, 7.5, and 1.5 respectively.

[0059] Then, the improved YOLOv8 network model is trained and parameter adjusted through the training set and validation set.

[0060] The extreme fire test set images or videos with the extreme fire categories labeled are input into the trained improved YOLOv8 network model. In this embodiment, the initial batch size during model training is set to 32, the initial learning rate is set to 0.01, the optimizer is SGD, the weight decay is set to 5e-4, and the momentum is set to 0.937. Finally, the category, location, and detection score of the extreme fire are output. In this embodiment, the detection accuracy of extreme fire reaches 87.4%, and the location is also correctly selected.

[0061] Step S4: Using the trained improved YOLOv8 network model to identify extreme fires. The extreme fire image to be identified is input into the trained improved YOLOv8 network model, and the category and location of the extreme fire are output.

[0062] The present invention provides a lightweight extreme fire recognition method based on context-aware large kernel transformation. The method obtains images of extreme fire datasets through indoor and outdoor experimental scenes of extreme fires, surveillance video materials of wild fires, and network image search engines. The extreme fire dataset is rich in samples. The method establishes an extreme fire dataset, which includes four types: crown fire, fire line merging, flying fire, and fire whirlwind. The data enhancement technology is used to increase the diversity of data samples and reduce the overfitting tendency of the model, which can be used by other researchers for further research on extreme fires. The method improves the original C2f module into a context-aware large kernel transformation module, and improves the original Detect module into a lightweight decoupled detection head L-Detect, and then improves the target detection architecture of the YOLOv8 network model, so that the model takes into account the environmental information around extreme fires in the global scope, while reducing the loss of computing resources and maintaining a high level of recognition accuracy.

[0063] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0064] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0065] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0067] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A lightweight extreme fire identification method based on context-aware large kernel transformation, characterized in that: The following steps are involved: Step S1: Acquire extreme fire images; Step S2: annotating extreme fire types and extreme fire locations on extreme fire images to generate an extreme fire dataset, which is divided into a training set, a validation set, and a test set; Step S3: using a lightweight structure to replace two redundant standard convolutional layers in the YOLOv8 network model, adjusting and optimizing the decoupled detection head structure of YOLOv8, and forming a lightweight decoupled detection head L-Detect; improving the bottleneck part in the C2f module of the YOLOv8 network model into two main parts: a context-aware module and a large-core transformation module, and forming a context-aware large-core transformation module CALKT; improving the YOLOv8 network model, introducing the context-aware large-core transformation module CALKT and the lightweight decoupled detection head L-Detect to replace the original C2f module and Detect module of the model, respectively, and forming an improved YOLOv8 network model; training, parameter adjustment and evaluation of the improved YOLOv8 network model are performed through training sets, validation sets and test sets; Step S4: Use the trained improved YOLOv8 network model to identify extreme fires.

2. A lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1, characterized in that: In step S1, extreme fire images are obtained from multiple sources, including indoor and outdoor experimental scenes of extreme fires, surveillance video data of wild fires, and network image search engines.

3. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S2, Labelimg software is used to label extreme fire images with extreme fire categories and typical extreme fire locations to generate an extreme fire dataset.

4. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S2, the extreme fire types include crown fire, fire line merging, flying fire and fire whirl.

5. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S2, data enhancement is performed on the acquired extreme fire image, including: horizontal and vertical flipping, affine transformation, random rotation, contrast enhancement, sharpening, noise addition, and blurring; then, the extreme fire type and extreme fire location are labeled on the extreme fire image obtained by data enhancement.

6. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S2, the extreme fire data set is divided into a data set according to the ratio of training set: validation set: test set = 8:1:

1.

7. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S3, the decoupled detection head structure of YOLOv8 is adjusted and optimized, specifically: the original standard convolution is replaced by a depth-separable convolution with a 5×5 convolution kernel, so that the input feature map is reduced in dimension and lightweight through two branches.

8. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S3, the bottleneck part in the C2f module of the YOLOv8 network model is improved into two main parts: one is a context-aware module, which first reduces the dimension of the input feature map through a 1×1 standard convolution, and then transmits the feature map in parallel to a 3×3 standard convolution and a 3×3 dilated convolution, and the two outputs are fused to obtain rich local and global details without adding additional parameters; the second is a large kernel transformation module, which adopts a large separable kernel module design, and is first composed of cascaded vertical and horizontal depthwise separable convolutions, respectively, and then cascaded vertical and horizontal large-size dilated depthwise separable convolutions to capture the geometric shape characteristics of extreme fire without adding too many parameters.

9. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S3, the YOLOv8 network model is improved, specifically including: improving the original C2f module of the model to a context-aware large kernel transformation module, and adopting a more effective feature extraction and fusion method to achieve a wider effective receptive field and context information perception capability; improving the original Detect module of the model to a lightweight decoupled detection head L-Detect to reduce the parameter redundancy of the detection head, thereby reducing the computational requirements of the model in the inference stage.

10. The lightweight extreme fire identification method based on context-aware large kernel transformation according to claim 1 is characterized in that: In step S4, the extreme fire image to be identified is input into the trained improved YOLOv8 network model, and the category and location of the extreme fire are output.