Wildfire detection method and system based on three-dimensional convolutional network and self-attention mechanism
By introducing a three-dimensional attention mechanism and consistent attention regularization in the three-dimensional convolutional network, extracting spatiotemporal features and enhancing the network distinction ability, the existing wildfire detection algorithm has solved the problems of low recognition accuracy and high false alarm rate, and achieved higher recognition accuracy and lower false alarm rate.
Patent Information
- Application Number
- CN202510021362.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The existing wildfire detection algorithm has low recognition accuracy and high false alarm rate, especially in cloud/fog/smoke and other environments, and the false alarm rate exceeds 90%.
Wildfire detection method based on three-dimensional convolutional networks and self-attention mechanisms is adopted to extract spatiotemporal features by introducing three-dimensional attention mechanisms, and consistent attention regularization is adopted in feature extraction networks to enhance the distinction ability and robustness of the network.
It improves the recognition accuracy of wildfire detection, reduces the false alarm rate, and can more accurately identify wildfire areas, reducing the impact of clouds/fog/smoke and other environments on the detection effect.
Smart Images

Figure CN119418253B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical fields of image recognition, deep learning, etc., and specifically relates to a wildfire detection method and system based on a three-dimensional convolutional network and a self-attention mechanism. Background Art
[0002] Under the influence of abnormal climate factors such as global warming, high-risk fire weather with high temperature, low humidity, continuous drought and high wind speed is becoming more and more frequent, and the number of wildfire outbreaks each year is on the rise. Some areas with high forest coverage and forest vegetation density have repeatedly experienced wildfires that endanger the safety of the power grid. Therefore, the identification and early warning of wildfires are particularly important. At present, AI intelligent recognition algorithms are often used to automatically identify wildfires, which can detect wildfires in time, prevent accidents from expanding and affecting the safety of transmission lines, and free operation and maintenance personnel from the heavy work of human eye identification, greatly improving work efficiency.
[0003] However, some conventional algorithm models have low recognition accuracy and high false alarm rates, which poses a great challenge to wildfire prevention work. The main reasons are: clouds / fog / smoke are difficult to accurately identify, the three are unpredictable, have low characteristics, and have strong similarities, which are sometimes difficult for even the human eye to distinguish. The transmission line image / video monitoring distance is long and the scale is large. In addition, the geographical and meteorological environment in some areas is complex and changeable, which increases the background interference of image information and causes a high false alarm rate. The current false alarm rate exceeds 90%. The complex and changeable field meteorological and terrain conditions also pose a huge challenge to wildfire image recognition algorithms, making the algorithm model's generalization performance weak and the recognition effect poor. Summary of the invention
[0004] In order to solve the problems of cloud / smoke / fog etc. which are difficult to accurately identify and have high false alarm rate in current wildfire detection, this application proposes a wildfire detection method and system based on a three-dimensional convolutional network and a self-attention mechanism. This application adds a three-dimensional attention mechanism on the basis of a three-dimensional convolutional network, which can more effectively extract spatiotemporal features and enhance the network's ability to distinguish. At the same time, it also utilizes consistent attention regularization to constrain the obtained spatiotemporal features at the feature level, thereby obtaining features with strong self-distinguishing commonalities and improving the accuracy of network recognition.
[0005] This application is implemented through the following technical solutions:
[0006] A forest fire detection method based on a three-dimensional convolutional network and a self-attention mechanism, the forest fire detection method comprising:
[0007] Construct training and test sets based on historical wildfire video datasets;
[0008] Preprocessing the video data in the training set and the test set, converting each video data into an array containing multiple frames of images, and performing normalization processing;
[0009] Constructing a wildfire recognition model, the wildfire recognition model comprising a three-dimensional convolutional network, a feature extraction network based on a self-attention mechanism and a consistent attention regularization mechanism, and a classification network connected in sequence;
[0010] Using the preprocessed training set to train the wildfire recognition model;
[0011] Using the preprocessed test set to test the trained wildfire recognition model to obtain a wildfire detection model;
[0012] Wildfire video data is collected in real time, and the wildfire detection model is used to perform wildfire detection.
[0013] In some embodiments, the three-dimensional convolutional network includes a three-dimensional convolutional layer, a three-dimensional pooling layer, and a Dropout layer;
[0014] The three-dimensional convolutional layer inputs multiple frames of images and initially extracts a spatiotemporal feature map;
[0015] The three-dimensional pooling layer performs dimensionality reduction processing on the feature map output by the three-dimensional convolutional layer;
[0016] The Dropout layer processes the feature map output by the three-dimensional pooling layer and outputs it to the feature extraction network.
[0017] In some embodiments, the feature extraction network includes a pooling layer and a self-differentiating feature extraction structure, and the self-differentiating feature extraction structure includes three convolutional networks, corresponding to the feature extraction of bottom-level features, middle-level features, and high-level features, respectively. A self-attention mechanism is used on the three convolutional networks to focus attention on discriminative image areas, and a consistent attention regularization mechanism is used to constrain the attention of different layers, thereby extracting wildfire target features and inputting them into a classification network for classification and recognition.
[0018] In some embodiments, the calculation formula of the three-dimensional convolutional layer is:
[0019]
[0020] in, Indicates i In the convolutional layer j The feature map is at position The value at , tanh() is the hyperbolic tangent function, and Represent the height and width of the convolution kernel respectively, represents the size of the three-dimensional kernel in the time dimension, is the deviation of the feature map, m The feature map set is i -1 layer is connected to the index of the current feature map, The previous layer is connected to the m The feature map has a spatial dimension The weight value at Indicates that the input feature map of the previous layer is at position The value at .
[0021] In some embodiments, the self-differentiating feature extraction structure adopts the ResNet50 network architecture.
[0022] In some embodiments, the self-attention mechanism includes three dilated convolutions with different dilated rates and three corresponding ordinary convolutions. The feature map is divided into three branches, which are respectively processed by the dilated convolutions with different dilated rates and then respectively processed by the corresponding ordinary convolutions. The processed three branch feature maps are then summed to obtain a heat map output.
[0023] In some embodiments, the consistent attention regularization is defined as:
[0024]
[0025] in, is the loss function of the regularization term, indicating that it depends on the input matrix and learnable parameters in neural networks , Indicates k +1 layer feature matrix, Indicates k The target feature matrix of the layer, that is, the network hopes to k The ideal output learned by the layer, Indicates k The feature matrix of the layer, is the square of the Frobenius norm, which is used to measure the matrix and The degree of difference between is the norm of the matrix, which is used to ensure the consistency between matrix features. K and K +1 indicates the number of heat maps, and are two constant weights.
[0026] In some embodiments, the loss function used by the wildfire recognition model is the sum of classification loss and consistent attention regularization.
[0027] In some embodiments, the real-time collection of wildfire video data and the use of the wildfire detection model to perform wildfire detection specifically include:
[0028] Preprocessing the wildfire video data collected in real time, converting the wildfire video data into an array containing multiple frames of images, and performing normalization processing;
[0029] The normalized data containing multiple frames of images is input into the wildfire detection model to obtain a recognition result.
[0030] On the other hand, the present application proposes a forest fire detection system based on a three-dimensional convolutional network and a self-attention mechanism, the forest fire detection system comprising:
[0031] A data set construction module, wherein the data set construction module constructs a training set and a test set based on a historical wildfire video data set;
[0032] A data preprocessing module, which preprocesses the video data in the training set and the test set, converts each video data into an array containing multiple frames of images, and performs normalization processing;
[0033] A model building module, wherein the model building module is used to build a wildfire recognition model, wherein the wildfire recognition model includes a three-dimensional convolutional network, a feature extraction network based on a self-attention mechanism and a consistent attention regularization mechanism, and a classification network connected in sequence;
[0034] A model training module, wherein the model training module uses the preprocessed training set to train the wildfire recognition model; and uses the preprocessed test set to test the trained wildfire recognition model to obtain a wildfire detection model;
[0035] And, a real-time detection module, the real-time detection module is used to collect wildfire video data in real time and use the wildfire detection model to perform wildfire detection.
[0036] The present application proposes a wildfire detection method and system based on a three-dimensional convolutional network and a self-attention mechanism. A three-dimensional attention mechanism is introduced into the three-dimensional convolutional network, which not only obtains attention at the spatial level, but also obtains attention at the temporal level, which is more helpful to mine spatiotemporal features that are effective for wildfire identification. At the same time, the attention mechanism is adopted at the bottom, middle and high feature layers, which helps to capture the relationship between the global and the local, thereby improving the recognition performance and recognition effect of the network.
[0037] The present application proposes a method and system for detecting wildfires based on a three-dimensional convolutional network and a self-attention mechanism, which can extract features with strong self-distinguishing properties. These features are extremely robust and can remain stable and accurate under different environments and conditions. Such features are particularly important in wildfire detection, because wildfire scenes are often complex and changeable, and factors such as lighting, smoke, and background will affect the detection effect. In addition, such highly distinguishable feature extraction can directly improve the accuracy of wildfire identification. Through refined feature analysis, wildfire areas can be identified more accurately, reducing the occurrence of false alarms and missed alarms.
[0038] The present application proposes a wildfire detection method and system based on a three-dimensional convolutional network and a self-attention mechanism, which can successfully identify early tiny smoke and fire signals and achieve early warning, which not only provides a valuable time window for wildfire prevention and control, but also can effectively reduce the probability of large-scale wildfires. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings described herein are used to provide a further understanding of the embodiments of the present application, constitute a part of the present application, and do not constitute a limitation on the embodiments of the present application. In the drawings:
[0040] Figure 1 A flow chart of a method according to an embodiment of the present application;
[0041] Figure 2 (a) shows some fire video images in the wildfire video dataset;
[0042] Figure 2 (b) shows some non-fire video images in the wildfire video dataset;
[0043] Figure 3 A schematic diagram of the wildfire identification model architecture of an embodiment of the present application;
[0044] Figure 4 This is a schematic diagram of the self-attention mechanism architecture of an embodiment of the present application;
[0045] Figure 5 This is a system principle block diagram of an embodiment of the present application. DETAILED DESCRIPTION
[0046] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with examples and drawings. The illustrative implementation scheme of the present application and its description are only used to explain the present application and are not intended to limit the present application.
[0047] Embodiment: In view of the problems of low recognition accuracy and high false alarm rate in current wildfire detection technology, this embodiment proposes a wildfire detection method based on a three-dimensional convolutional network and a self-attention mechanism. The method proposed in this embodiment uses a three-dimensional convolutional network to preliminarily extract the spatiotemporal features contained in the input video. The attention mechanism enables the detection network to pay more attention to the characteristics of the detected target. The use of the consistent attention regularization method allows the network to learn the similarities of attention information at each layer, and learn high-quality representation information from high-level feature maps to help the network focus on identifiable areas at lower levels, thereby extracting features with strong self-distinguishing commonalities, and then inputting the extracted features into the classifier network to obtain the final wildfire detection results.
[0048] Specific as Figure 1 As shown, the detection method proposed in this embodiment includes:
[0049] Step 100, construct a training set and a test set based on the historical wildfire video dataset.
[0050] In step 100, a historical wildfire video data set is processed to obtain a number of video data without fire and a number of video data with fire, and then divided into a training set and a test set according to a preset ratio (for example, 8:2 or 7:3, etc.).
[0051] Specifically, step 100 includes the following sub-steps:
[0052] Step 101, the historical wildfire video dataset (including both fire and no fire) is processed, including: deleting blank or too short videos, cutting the videos into segments with smaller duration differences (i.e., ensuring that all videos in the dataset have the same duration), and finally obtaining 2475 video data without fire and 1881 video data with fire, as shown in Figure 2 (a) for some video images with fire and Figure 2 (b) for some video images without fire.
[0053] Step 102, divide the processed video data set into a training set and a test set, wherein the video data volume in the training set accounts for 80% of the total video data volume, and the remaining video data is divided into a test set, which is used for subsequent network model training and testing respectively.
[0054] Step 200 , preprocessing the video data in the training set and the test set, converting each video data into an array containing multiple frames of images, and performing normalization processing.
[0055] The step 200 specifically includes the following sub-steps:
[0056] Step 201, read the corresponding video file according to the input file name, file path, size of each frame image and frame length L, and return an array containing L frames of images;
[0057] Step 202, using shuffle (random order) and file names to batch read the video files obtained in step 201, randomly shuffle the order of the file indexes according to the number of data files, and complete the reading of the array data according to the files corresponding to the indexes.
[0058] Step 203, normalize the read array data.
[0059] Step 300, construct a wildfire recognition model, which includes a three-dimensional convolutional network, a feature extraction network based on a self-attention mechanism and a consistent attention regularization mechanism, and a classification network that are connected in sequence.
[0060] like Figure 3 The specific architecture of the wildfire recognition model is shown in the figure. The three-dimensional convolutional network mainly includes three-dimensional convolutional layers, three-dimensional pooling layers and Dropout layers; the feature extraction network mainly includes pooling layers and self-distinguishing feature extraction structures, which mainly include three convolutional networks, corresponding to the feature extraction of bottom-level features, middle-level features and high-level features, and the three convolutional networks all use the self-attention mechanism to focus attention on the discriminative image areas, and use the consistent attention regularization mechanism to constrain the attention of different layers, so as to mine strong distinguishable wildfire target features that contain rich semantic information and have consistent commonalities; the classification network mainly includes global average pooling layers, fully connected layers and Softmax layers.
[0061] Among them, multiple frames of images are input into the three-dimensional convolution layer, and the spatiotemporal features are preliminarily extracted through the three-dimensional convolution layer; among them, the calculation formula of the three-dimensional convolution layer is as follows:
[0062]
[0063] in, Indicates i In the convolutional layer j The feature map is at position The value at , tanh() is the hyperbolic tangent function, and Represent the height and width of the convolution kernel respectively, represents the size of the three-dimensional kernel in the time dimension, is the deviation of the feature map, m The feature map set is i -1 layer is connected to the index of the current feature map, The previous layer is connected to the m The feature map has a spatial dimension The weight value at Indicates that the input feature map of the previous layer is at position In this embodiment, the three-dimensional convolution layer is used to extract spatial and temporal features by applying a convolution operation to a three-dimensional kernel, and the three-dimensional convolution layer has a total of 32 filters of a size of 3×3×15.
[0064] Then, the feature map output by the three-dimensional convolution layer is input to the three-dimensional pooling layer for dimensionality reduction. After the convolution process, in order to prevent overfitting, the three-dimensional pooling layer is used to compress the data. In this embodiment, the three-dimensional pooling layer can use but is not limited to 3×3×3 pooling kernel pooling, and by selecting the best feature representation in a small spatiotemporal window, the output dimension of the three-dimensional convolution layer is reduced while retaining important features and reducing the difficulty of model training. This embodiment uses maximum pooling.
[0065] In order to reduce the overfitting of the model relative to the training samples, this embodiment also adds a Dropout layer in the three-dimensional convolutional network.
[0066] The spatiotemporal feature map initially extracted by the 3D convolutional network is input into the feature extraction network, and the pooling layer is used to reduce its dimensionality. Then, the feature map output by the pooling layer is input into the self-differentiating feature extraction structure. The self-differentiating feature extraction structure can adopt the ResNet50 network architecture, and the self-attention mechanism is used in the bottom, middle and high convolutional layers to focus on the discriminative image areas. Specifically, the self-attention mechanism includes three Atrous convolutions with different atrous rates are used to obtain different local to global perception perspectives, followed by three The convolution improves the network’s expressiveness. Finally, the results of the three convolution branches are summed to obtain the heat map output. The schematic diagram is shown in Figure 4 In addition, the consistent attention regularization mechanism is used to constrain the attention of different layers, so as to mine the strong distinguishable wildfire target features that contain rich semantic information and have consistent commonalities. The definition of consistent attention regularization is as follows:
[0067]
[0068] in, is the loss function of the regularization term, indicating that it depends on the input matrix and learnable parameters in neural networks , Indicates k +1 layer feature matrix, Indicates k The target feature matrix of the layer, that is, the network hopes to k The ideal output learned by the layer, Indicates k The feature matrix of the layer, is the square of the Frobenius norm, which is used to measure the matrix and The degree of difference between is the norm of the matrix, which is used to ensure the consistency between matrix features. K and K +1 indicates the number of heat maps, and are two constant weights. In addition, and The same size is obtained by The maximum pooling operation is performed, where the hyperparameter stride in the maximum pooling is 2.
[0069] The consistent attention regularization mechanism consists of two parts, namely the consistency term (i.e., the first half of the right side of the above definition) and the sparsity term (i.e., the second half of the right side of the above definition). Among them: the purpose of the consistency term is to maintain the similarity of the heat maps learned from the low-level, medium-level, and high-level feature maps, so that the high-quality representation information learned from the high-level feature map can be used to help the network focus on the identifiable areas of the lower layers; the sparsity term tends to perform feature selection, which is conducive to eliminating and filtering some features that cause misidentification.
[0070] The feature map extracted by the feature extraction network is input into the global average pooling layer for dimensionality reduction; then the feature map after dimensionality reduction is input into the fully connected layer, and then the final classification result (i.e., fire or no fire) is obtained through the Softmax layer.
[0071] The loss function used in the wildfire recognition model consists of two parts. The first part is the classification loss, and its calculation formula is:
[0072]
[0073] in, is the loss function of classification loss, which depends on the input data X and the parameters W of the model. N Indicates the number of samples processed in the same batch, exp() represents the exponential function, represents the learned classifier List, express The transpose of represents the feature vector learned by our attention network in the feedforward, Representation Tags The corresponding columns, express The transpose of .
[0074] The second part is the consistent attention regularization, the total loss L is the sum of the two:
[0075] .
[0076] Step 400: Use the preprocessed training set and test set to train and test the wildfire recognition model to obtain a wildfire detection model. Use the training set and test set to train the constructed wildfire recognition model until the expected training requirements are met: for example, the number of training iterations is reached, or the error meets the requirements, etc.; then, use the test set to test the accuracy of the trained wildfire recognition model. If the accuracy meets the requirements, the current wildfire recognition model is used as the wildfire detection model and loaded into the corresponding terminal to realize real-time wildfire detection. Otherwise, return to retrain until the corresponding accuracy requirements are met.
[0077] Step 500, collect wildfire video data in real time, and use the wildfire detection model to perform wildfire detection.
[0078] This step 500 uses the wildfire detection model to perform real-time wildfire detection, and the specific process includes:
[0079] Step 501, collect wildfire video data in real time.
[0080] Step 502, pre-process the wildfire video data, convert the video data into an array containing multiple frames of images, and perform normalization processing. It should be noted that the processing process of step 502 is the same as that described in step 200 above, and will not be repeated here.
[0081] Step 503, input the normalized array containing multiple frames of images into the wildfire detection model to obtain a recognition result.
[0082] The wildfire detection method proposed in this embodiment utilizes a three-dimensional convolutional network to extract the spatiotemporal features contained in wildfire videos, and adopts an attention mechanism and a consistent attention regularization mechanism, so that the detection network can pay more attention to the characteristics of the detected target, extract features with strong self-distinguishing commonalities, and use features with strong self-distinguishing commonalities for classification and identification, thereby reducing the interference of similar targets such as clouds / fog / smoke on wildfire detection, improving accuracy, and reducing false alarms.
[0083] Based on the same technical concept as above, this embodiment also proposes a forest fire detection system based on a three-dimensional convolutional network and a self-attention mechanism. Figure 5 As shown, the detection system specifically includes:
[0084] A dataset construction module, which constructs a training set and a test set based on a historical wildfire video dataset.
[0085] The data preprocessing module preprocesses the video data in the training set and the test set, converts each video data into an array containing multiple frames of images, and performs normalization.
[0086] A model building module is used to build a wildfire recognition model, which includes a three-dimensional convolutional network connected in sequence, a feature extraction network based on a self-attention mechanism and a consistent attention regularization mechanism, and a classification network.
[0087] The model training module uses the training set and the test set to train and test the wildfire recognition model to obtain the wildfire detection model.
[0088] And, a real-time detection module, which collects wildfire video data in real time and uses a wildfire detection model to perform wildfire detection.
[0089] It should be noted that the specific implementation process of each functional module of the system is as described in the above method, and will not be elaborated here.
[0090] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0091] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0092] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0094] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A forest fire detection method based on a three-dimensional convolutional network and a self-attention mechanism, characterized in that: The mountain fire detection method comprises: Construct training and test sets based on historical wildfire video datasets; Preprocessing the video data in the training set and the test set, converting each video data into an array containing multiple frames of images, and performing normalization processing; Constructing a wildfire recognition model, the wildfire recognition model comprising a three-dimensional convolutional network, a feature extraction network based on a self-attention mechanism and a consistent attention regularization mechanism, and a classification network connected in sequence; Using the preprocessed training set to train the wildfire recognition model; Using the preprocessed test set to test the trained wildfire recognition model to obtain a wildfire detection model; Collecting wildfire video data in real time, and using the wildfire detection model to detect wildfires; the three-dimensional convolutional network includes a three-dimensional convolutional layer, a three-dimensional pooling layer and a Dropout layer; The three-dimensional convolutional layer inputs multiple frames of images and initially extracts a spatiotemporal feature map; The three-dimensional pooling layer performs dimensionality reduction processing on the feature map output by the three-dimensional convolutional layer; The Dropout layer processes the feature map output by the three-dimensional pooling layer and outputs it to the feature extraction network; the feature extraction network includes a pooling layer and a self-differentiating feature extraction structure, and the self-differentiating feature extraction structure includes three convolutional networks, which correspond to the feature extraction of bottom-level features, middle-level features and high-level features, respectively. The self-attention mechanism is adopted on the three convolutional networks to focus attention on the discriminative image areas, and the consistent attention regularization mechanism is adopted to constrain the attention of different layers, so as to extract the wildfire target features and input them into the classification network for classification and recognition; the calculation formula of the three-dimensional convolutional layer is: in, represents the value of the jth feature map at position (x, y, z) in the i-th convolutional layer, tanh() is the hyperbolic tangent function, P i and Q i Represent the height and width of the convolution kernel, R i represents the size of the three-dimensional kernel in the time dimension, b ij is the deviation of the feature map, m is the index of the feature map set connected to the current feature map at the i-1 layer, is the weight value of the previous layer connected to the mth feature map in the spatial dimension (p, q, r), Represents the value of the input feature map of the previous layer at position (x+p, y+p, z+r).
2. A forest fire detection method based on a three-dimensional convolutional network and a self-attention mechanism according to claim 1, characterized in that: The self-distinguishing feature extraction structure adopts the ResNet50 network architecture.
3. The method for detecting wildfires based on a three-dimensional convolutional network and a self-attention mechanism according to claim 1, characterized in that: The self-attention mechanism includes three dilated convolutions with different dilated rates and three corresponding ordinary convolutions. The feature map is divided into three branches, which are processed by the dilated convolutions with different dilated rates respectively and then processed by the corresponding ordinary convolutions respectively. After that, the processed three branch feature maps are summed to obtain the heat map output.
4. The method for detecting wildfires based on a three-dimensional convolutional network and a self-attention mechanism according to claim 1, characterized in that: The consistent attention regularization is defined as: in, is the loss function of the regularization term, indicating that it depends on the input matrix H and the learnable parameters Θ, H in the neural network k+1 represents the feature matrix of the k+1th layer, represents the target feature matrix of the kth layer, that is, the ideal output that the network hopes to learn at the kth layer, H k represents the feature matrix of the kth layer, is the square of the Frobenius norm, which is used to measure the matrix H k+1 and The degree of difference between them, || ||1 is the norm of the matrix, which is used to ensure the consistency between the matrix features, K and K+1 represent the number of heat maps, β and are two constant weights.
5. The method for detecting wildfires based on a three-dimensional convolutional network and a self-attention mechanism according to claim 1, characterized in that: The loss function used by the wildfire recognition model is the sum of classification loss and consistent attention regularization.
6. A forest fire detection method based on a three-dimensional convolutional network and a self-attention mechanism according to any one of claims 1 to 5, characterized in that: The real-time collection of wildfire video data and the use of the wildfire detection model for wildfire detection specifically include: Preprocessing the wildfire video data collected in real time, converting the wildfire video data into an array containing multiple frames of images, and performing normalization processing; The normalized data containing multiple frames of images is input into the wildfire detection model to obtain a recognition result.
7. A forest fire detection system based on a three-dimensional convolutional network and a self-attention mechanism, characterized in that: The mountain fire detection system comprises: A data set construction module, wherein the data set construction module constructs a training set and a test set based on a historical wildfire video data set; A data preprocessing module, which preprocesses the video data in the training set and the test set, converts each video data into an array containing multiple frames of images, and performs normalization processing; A model building module, wherein the model building module is used to build a wildfire recognition model, wherein the wildfire recognition model includes a three-dimensional convolutional network, a feature extraction network based on a self-attention mechanism and a consistent attention regularization mechanism, and a classification network connected in sequence; A model training module, wherein the model training module uses the preprocessed training set to train the wildfire recognition model; and uses the preprocessed test set to test the trained wildfire recognition model to obtain a wildfire detection model; and, a real-time detection module, the real-time detection module being used to collect wildfire video data in real time and perform wildfire detection using the wildfire detection model; The three-dimensional convolutional network includes a three-dimensional convolutional layer, a three-dimensional pooling layer and a Dropout layer; The three-dimensional convolutional layer inputs multiple frames of images and initially extracts a spatiotemporal feature map; The three-dimensional pooling layer performs dimensionality reduction processing on the feature map output by the three-dimensional convolutional layer; The Dropout layer processes the feature map output by the three-dimensional pooling layer and outputs it to the feature extraction network; the feature extraction network includes a pooling layer and a self-differentiating feature extraction structure, and the self-differentiating feature extraction structure includes three convolutional networks, which correspond to the feature extraction of bottom-level features, middle-level features and high-level features, respectively. The self-attention mechanism is adopted on the three convolutional networks to focus attention on the discriminative image areas, and the consistent attention regularization mechanism is adopted to constrain the attention of different layers, so as to extract the wildfire target features and input them into the classification network for classification and recognition; the calculation formula of the three-dimensional convolutional layer is: in, represents the value of the jth feature map at position (x, y, z) in the i-th convolutional layer, tanh() is the hyperbolic tangent function, P i and Q i Represent the height and width of the convolution kernel, R i represents the size of the three-dimensional kernel in the time dimension, b ij is the deviation of the feature map, m is the index of the feature map set connected to the current feature map at the i-1 layer, is the weight value of the previous layer connected to the mth feature map in the spatial dimension (p, q, r), Represents the value of the input feature map of the previous layer at position (x+p, y+p, z+r).
Citation Information
Patent Citations
Text classification model based on convolutional neural network and self-attention
CN110619045A
Mountain fire detection method based on three-dimensional convolutional neural network
CN111898440A