A Method for Detecting Abnormal Events Underground Based on Multispectral Image Fusion

Through multispectral image fusion technology, combined with improved loss function and twin network, the problem of low detection accuracy of traditional video surveillance systems in the underground environment is solved, and the accuracy and speed of abnormal event detection are achieved.

CN119863669BActive Publication Date: 2025-06-20ANHUI ZHENXIN INTERNET TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510354466.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-20
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Traditional video surveillance systems are difficult to effectively identify abnormal events in complex and changing environments underground, and the detection accuracy is reduced due to factors such as noise and light changes.

Method used

Multi-spectral imaging equipment is used to collect multiple sets of multi-spectral downhole images, and by constructing a multi-spectral image detection model, fusing the features of infrared images and visible images, and feature extraction and fusion are performed using improved weighted MSE loss function and twin network.

Benefits of technology

It improves the accuracy and reliability of underground abnormal event detection, enhances feature capture capabilities, significantly improves the accuracy of abnormal event detection, and improves the detection speed through image enhancement and model pruning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863669B_ABST
    Figure CN119863669B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting underground abnormal events based on multi-spectral image fusion. First, a multi-spectral imaging device is used to collect multiple groups of multi-spectral underground images and perform image preprocessing. Each group of multi-spectral underground images includes infrared images and visible light images. Then, a multi-spectral image detection model is constructed, which includes a feature extraction module and a feature fusion module. Finally, the multi-spectral image detection model is trained, and the loss function is set as a weighted MSE loss function. The trained multi-spectral image detection model is used to detect underground abnormal events in each group of multi-spectral underground images, and the detection results of underground abnormal events are obtained. By fusing the underground image data in different spectral bands, the present invention improves the accuracy and reliability of underground abnormal event detection, can focus on the areas containing key abnormal information, thereby enhancing its feature capture ability, and significantly improving the precision of abnormal event detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and specifically to a method for detecting underground abnormal events based on multi-spectral image fusion. Background Art

[0002] In mine production, underground safety is the primary consideration. However, due to the complex underground environment and poor lighting conditions, there are various potential safety hazards, such as fires, gas leaks, etc. Once these abnormal events occur, they may have a serious impact on mine production and personnel safety.

[0003] Traditional video surveillance systems mainly rely on image data in a single spectral band and are difficult to adapt to the complex and changeable underground environmental conditions. In the case of insufficient lighting or occlusion, traditional video surveillance methods often have difficulty effectively identifying abnormal events. And traditional image processing methods are easily affected by factors such as noise and lighting changes when processing underground images, resulting in a decrease in detection accuracy. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for detecting underground abnormal events based on multi-spectral image fusion. By fusing underground image data in different spectral bands, the accuracy and reliability of underground abnormal event detection are improved, and it can focus on the area containing key abnormal information, thereby enhancing its feature capture ability and significantly improving the precision of abnormal event detection.

[0005] The technical solution of the present invention is as follows:[[]]

[0006] A method for detecting underground abnormal events based on multi-spectral image fusion specifically includes the following steps:[[]]

[0007] (1) Collect multiple groups of multi-spectral underground images using a multi-spectral imaging device and perform image preprocessing to construct an underground image dataset. Each group of multi-spectral underground images includes an infrared image and a visible light image.

[0008] (2) Construct a multi-spectral image detection model. The multi-spectral image detection model includes a feature extraction module and a feature fusion module. The feature extraction module extracts features from the infrared image and the visible light image respectively, and the feature fusion module fuses the extracted infrared image features and visible light image features.

[0009] (3) Use the underground image dataset to train the multi-spectral image detection model. The loss function for training is a weighted MSE loss function , and the calculation formula of the weighted MSE loss function is shown in the following formula (1):

[0010] (1);

[0011] In formula (1), 、 、 represent the height, width, and number of channels of the multispectral downhole image respectively; 、 、 represent the indices of each pixel in the height, width, and number of channels of the multispectral downhole image respectively, and are used to traverse the positions and channels of each pixel in the spectral downhole image; represents the pixel value of the true value of the multispectral downhole image at the position, which is the target for model learning; represents the pixel value of the prediction value of the multispectral image detection model at the position; is the weight decay regularization parameter; is the weight at each pixel position;

[0012] (4) Use the trained multispectral image detection model to detect downhole abnormal events for each group of multispectral downhole images, and obtain the detection results of downhole abnormal events.

[0013] The image preprocessing includes image denoising and image enhancement. For the smooth area in the multispectral downhole image, an image denoising algorithm based on Gaussian filtering is used for image denoising; for the area containing edge and detail information, an edge-preserving filtering algorithm is used for image denoising; the image enhancement is to use the contrast-limited adaptive histogram equalization method to enhance the local contrast of the image.

[0014] The feature extraction module is a Siamese network. The Siamese network includes a visible light image feature extraction sub-network and an infrared image feature extraction sub-network with the same network architecture. The input of the visible light image feature extraction sub-network is the image obtained by splicing the denoised visible light image and its enhanced image in the channel dimension, and the input of the infrared image feature extraction sub-network is the image obtained by splicing the denoised infrared image and its enhanced image in the channel dimension;

[0015] Both the visible light image feature extraction sub-network and the infrared image feature extraction sub-network include 3×3 convolution, depth convolution, two ReLU functions, an attention mechanism, a residual module, and a fully connected layer. The processing process is specifically shown in the following formula (2):

[0016] (2);

[0017] In formula (2), represents 3×3 convolution, represents depth convolution, represents the ReLU function, Represents the attention mechanism, Represents the residual module, Represents the fully connected layer, Is the input of the visible light image feature extraction sub-network or the infrared image feature extraction sub-network, Is the feature map output after depth convolution, Is the output of the visible light image feature extraction sub-network or the infrared image feature extraction sub-network.

[0018] The described attention mechanism includes a channel attention mechanism and a spatial attention mechanism. The channel attention mechanism first performs global average pooling and global max pooling operations on the input feature map respectively, then fuses and concatenates the vectors after global average pooling and global max pooling together and compresses the dimension through a fully connected layer, and then uses the Sigmoid function for normalization processing. The output of the channel attention mechanism is used as the input of the spatial attention mechanism. The spatial attention mechanism first performs global average pooling and global max pooling operations on the input respectively, then fuses and concatenates the two spatial maps output after global average pooling and global max pooling processing in the channel dimension, and then uses 1×1 convolution and the Sigmoid function to compress and normalize the feature map, and outputs the attention weight map.

[0019] The described feature fusion module first performs feature concatenation on the feature maps output by the visible light image feature extraction sub-network and the infrared image feature extraction sub-network with the same network architecture, then unifies the concatenated features into the same feature space through the first 3×3 convolution, then respectively performs feature extraction on the output of the first 3×3 convolution with convolutions of different receptive field sizes, and finally fuses the features extracted by the convolutions of different receptive field sizes through the second 3×3 convolution.

[0020] The weight of each pixel position described Is calculated by the following formula (3):

[0021] (3);

[0022] In formula (3), Represents the pixel value within the window; Represents the mean value of the pixels within the window; Represents the local variance calculated within the window; Represents the number of pixels within the window; Represents the coordinate points within the window; Represents the gradient magnitude calculated using the Sobel operator; And Are the gradients in the x and y directions obtained by convolving with the Sobel operator respectively; is a hyperparameter that balances the contributions of local variance and gradient magnitude; is the index of the multi-spectral downhole image in terms of height, width, and number of channels, used to traverse the entire multi-spectral downhole image to determine the global maximum of local variance and gradient magnitude, providing a global reference for the calculation of the current pixel weight; represents taking the maximum value within the value range; represents taking the maximum value within the small window centered at the position, which is the calculated local variance; represents the gradient magnitude calculated using the Sobel operator for the image area centered at the

[0023] During the training process of the multi-spectral image detection model described above, first train the multi-spectral image detection model once, then perform multiple random pruning on the Siamese network, and select the multi-spectral image detection model composed of the Siamese network with the best effect after pruning for retraining.

[0024] Advantages of the present invention:

[0025] (1) By fusing image data in two different spectral bands of visible light and infrared, the present invention can obtain more scene information, improve the accuracy and reliability of abnormal event detection. The multi-spectral image fusion technology can overcome the limitations of single-spectral image data, improve the signal-to-noise ratio and contrast of the image. In the case of insufficient downhole lighting or occlusion, the fused image can provide more feature information for the model.

[0026] (2) The present invention uses an improved weighted MSE loss function to improve the detection performance. The improved weighted MSE loss function consists of two parts: adaptive weighting based on data features and weight decay regularization. In the detection of abnormal events in multi-spectral downhole images, the adaptive weighting part determines the weight by calculating the local variance and gradient magnitude, and assigns a higher weight to regions with rich texture or significant gradient changes, which can guide the model to focus on regions containing key abnormal information, thereby enhancing its feature capture ability and significantly improving the accuracy of abnormal event detection; while the weight decay regularization aims to prevent overfitting, which penalizes the sum of squares of the weights to constrain the growth of the weights and ensure the stability and generalization performance of the model; the two are organically combined to achieve the strengthening of local key information and the effective management of global weights; at the same time, its parameters have high flexibility and can be adjusted according to the downhole scene and abnormal types, making the loss function have good versatility and scalability, providing a strong guarantee for the effective detection of downhole abnormal events.

[0027] (3). The image enhancement of the present invention uses the Contrast Limited Adaptive Histogram Equalization (CLAHE) method to enhance the local contrast of the image. Compared with the traditional global histogram equalization, CLAHE can adaptively enhance according to the brightness information of the local area, effectively avoiding the problem of noise amplification caused by over-enhancement while enhancing details, and is particularly suitable for enhancing the details of low-contrast images. Finally, image enhancement is performed by the method of local stitching.

[0028] (4). During the training of the multi-spectral image detection model of the present invention, random pruning is performed on the model, so that the lightweight multi-spectral image detection model has a faster inference and detection speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flowchart of the present invention.

[0030] Figure 2 is an architecture diagram of the multi-spectral image detection model of the present invention.

[0031] Figure 3 is an architecture diagram of the siamese network of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] See Figures 1 - 3 , a method for detecting underground abnormal events based on multi-spectral image fusion, specifically including the following steps:

[0034] (1). Use a multi-spectral imaging device to collect multiple groups of multi-spectral underground images and perform image preprocessing to construct an underground image dataset. Each group of multi-spectral underground images includes an infrared image and a visible light image;

[0035] Image preprocessing includes image denoising and image enhancement. For the smooth regions in the multi-spectral downhole images, a denoising algorithm based on Gaussian filtering is used for image denoising, which can effectively remove noise and maintain the uniformity of the smooth regions; for the regions containing edge and detail information, an edge-preserving filtering algorithm is used for image denoising to maximize the retention of edge and texture features while removing noise and avoid the loss of important details; image enhancement is to use the Contrast Limited Adaptive Histogram Equalization (CLAHE) method to enhance the local contrast of the image. CLAHE can adaptively enhance according to the brightness information of the local region, effectively avoiding the problem of noise amplification caused by over-enhancement while enhancing details, especially suitable for enhancing the details of low-contrast images. Finally, image enhancement is performed by the method of local stitching; the enhanced image of the present invention specifically cuts the infrared image and the matched visible light image into four equal parts uniformly and splices them crosswise;

[0036] (2) Construct a multi-spectral image detection model. The multi-spectral image detection model includes a feature extraction module and a feature fusion module. The feature extraction module extracts features from the infrared image and the visible light image respectively, and the feature fusion module fuses the extracted infrared image features and visible light image features;

[0037] S21. The feature extraction module is a Siamese network. The Siamese network includes a visible light image feature extraction sub-network and an infrared image feature extraction sub-network with the same network architecture. The input of the visible light image feature extraction sub-network is the image obtained by splicing the denoised visible light image and its enhanced image in the channel dimension, and the input of the infrared image feature extraction sub-network is the image obtained by splicing the denoised infrared image and its enhanced image in the channel dimension;

[0038] Both the visible light image feature extraction sub-network and the infrared image feature extraction sub-network include 3×3 convolution, depth convolution, two ReLU functions, an attention mechanism, a residual module, and a fully connected layer. The specific processing process is shown in the following formula (2):

[0039] (2);

[0040] In formula (2), represents 3×3 convolution, represents depth convolution, represents the ReLU function, represents the attention mechanism, represents the residual module, represents the fully connected layer, is the input of the visible light image feature extraction sub-network or the infrared image feature extraction sub-network, is the feature map output after depth convolution, It is the output of the visible light image feature extraction sub-network or the infrared image feature extraction sub-network;

[0041] That is, the input feature map first extracts the local features of the image, such as edge, texture and other information, through a 3×3 convolution; then further extracts the channel-specific feature information through depth convolution, which can enhance the expression ability of the network without increasing the computational complexity while reducing the computational amount. Then, the ReLU function is applied to the output feature map of the depth convolution to enhance the non-linear expression ability of the network, and then it enters the attention mechanism;

[0042] The attention mechanism includes a channel attention mechanism and a spatial attention mechanism. The channel attention mechanism first performs global average pooling (AvgPool) and global max pooling (MaxPool) operations on the input feature map respectively, and then fuses and concatenates the vectors after global average pooling and global max pooling and compresses the dimension through a fully connected layer. After that, the Sigmoid function is used for normalization to highlight the key channel information and improve the effectiveness and pertinence of feature representation; the output of the channel attention mechanism is used as the input of the spatial attention mechanism. The spatial attention mechanism first performs global average pooling and global max pooling operations on the input respectively, and then fuses and concatenates the two spatial maps output after global average pooling and global max pooling in the channel dimension. Subsequently, a 1×1 convolution and the Sigmoid function are used to compress and normalize the feature map, and an attention weight map is output;

[0043] Then, the attention weight map is connected to the feature map processed by depth convolution (DConv) in a residual manner to obtain a new feature map, which ensures that information can be directly transmitted from the shallower layer to the deeper layer of the network, maintains the effective flow of information, and avoids performance degradation caused by increasing depth. Then, the new feature map is passed into the fully connected layer to fuse the extracted features and integrate different feature information. Finally, the ReLU function is applied again to the output of the fully connected layer for non-linear activation to obtain the final output;

[0044] S22. The feature fusion module first performs feature concatenation on the feature maps output by the visible light image feature extraction sub-network and the infrared image feature extraction sub-network with the same network architecture, and then unifies the concatenated features into the same feature space through the first 3×3 convolution. Then, convolutions with different receptive field sizes are used to extract features from the output of the first 3×3 convolution respectively. Finally, the features extracted by the three convolutions with receptive field sizes of 3×3, 5×5, and 7×7 are fused through the second 3×3 convolution;

[0045] (3) Train and test the multi-spectral image detection model using the downhole image dataset. During the training process, first train the multi-spectral image detection model once, then randomly prune the Siamese network ten times, and select the multi-spectral image detection model composed of the Siamese network with the best effect after pruning for retraining;

[0046] The loss function for training is the weighted MSE loss function , the weighted MSE loss function The calculation formula is shown in the following formula (1):

[0047] (1);

[0048] In formula (1), , , respectively represent the height, width, and number of channels of the multi-spectral downhole image; , , respectively represent the indices of each pixel in the multi-spectral downhole image in terms of height, width, and number of channels, and are used to traverse the positions and channels of each pixel in the spectral downhole image; represents the pixel value of the true value of the multi-spectral downhole image at the position, which is the target for the model to learn; represents the pixel value of the prediction value of the multi-spectral image detection model at the position; is the weight decay regularization parameter; is the weight at each pixel position, calculated by the following formula (3):

[0049] (3);

[0050] In formula (3), represents the pixel value within the window; represents the mean value of the pixels within the window; represents the local variance calculated within the window; represents the number of pixels within the window; represents the coordinate points within the window; represents the gradient magnitude calculated using the Sobel operator; and are the gradients in the x and y directions obtained by convolving with the Sobel operator respectively; is a hyperparameter that balances the contributions of local variance and gradient magnitude; is the index in terms of height, width, and number of channels of the multi - spectral down - hole image, used to traverse the entire multi - spectral down - hole image to determine the global maximum of the local variance and gradient magnitude, providing a global reference for the calculation of the current pixel weight, and realizing the comparison and normalization of local features and global features; denotes taking the maximum value within the value range; denotes the local variance calculated within a small window centered at the position of ; represents the gradient magnitude calculated for the image region centered at the position of using the Sobel operator;

[0051] (4) Use the trained multi - spectral image detection model to detect down - hole abnormal events for each group of multi - spectral down - hole images, and obtain the down - hole abnormal event detection results.

[0052] Performance analysis:

[0053] Table 1 below shows the infrared and visible light datasets for pedestrian detection collected down - hole. The dataset contains 1000 infrared images and 1000 visible light images. 80% of the data is used as the training set and 20% as the test set in the experiment. The experimental results in Table 1 show that the model proposed in this embodiment (Ours) leads other existing models (YOLOv10n and RT - DETR) with an accuracy of 0.954, indicating that it performs best in correctly identifying positive samples.

[0054] Table 1

[0055]

[0056] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting underground abnormal events based on multispectral image fusion, characterized in that: The specific steps include: (1) Using a multispectral imaging device to collect multiple groups of multispectral downhole images and perform image preprocessing to construct a downhole image data set, where each group of multispectral downhole images includes an infrared image and a visible light image; (2) Construct a multispectral image detection model. The multispectral image detection model includes a feature extraction module and a feature fusion module. The feature extraction module extracts features from infrared images and visible light images respectively, and the feature fusion module fuses the extracted infrared image features and visible light image features. (3) The multispectral image detection model is trained using the downhole image dataset, and the training loss function is the weighted MSE loss function; The weighted MSE loss function The calculation formula is shown in the following formula (1): (1); In formula (1), , , They represent the height, width and number of channels of multispectral downhole images respectively; , , Respectively represent the index of each pixel in the multispectral downhole image in terms of height, width and number of channels, and are used to traverse the position and channel of each pixel in the spectral downhole image; Indicates the true value of the multispectral downhole image The pixel value of the position is the target of model learning; Indicates the predicted value of the multispectral image detection model in The pixel value of the position; is the weight decay regularization parameter; is the weight of each pixel position, which is calculated by the following formula (3): (3); In formula (3), Represents the pixel value within the window; Represents the mean value of pixels in the window; Represents the local variance calculated within the window; Represents the number of pixels in the window; Represents the coordinate point within the window; represents the gradient magnitude calculated using the Sobel operator; and They are the gradients in the x and y directions obtained by convolution of the Sobel operator; is a hyperparameter that balances the contribution of local variance and gradient amplitude; It is the index of the multispectral downhole image in height, width and number of channels, which is used to traverse the whole image of the multispectral downhole image to determine the global maximum of the local variance and gradient amplitude, and provide a global reference for the current pixel weight calculation; Indicated in Take the maximum value within the value range; Indicates that The local variance calculated within the window centered at the position; Represents the use of the Sobel operator to The gradient magnitude calculated from the image region centered at position; (4) Use the trained multispectral image detection model to detect underground abnormal events in each group of multispectral underground images to obtain the underground abnormal event detection results.

2. The method for detecting underground abnormal events based on multispectral image fusion according to claim 1, characterized in that: The image preprocessing includes image denoising and image enhancement. For the smooth area in the multispectral downhole image, a denoising algorithm based on Gaussian filtering is used for image denoising; for the area containing edge and detail information, an edge-preserving filtering algorithm is used for image denoising; the image enhancement adopts a contrast-limited adaptive histogram equalization method to enhance the local contrast of the image.

3. The method for detecting underground abnormal events based on multispectral image fusion according to claim 2, characterized in that: The feature extraction module is a twin network, which includes a visible light image feature extraction subnetwork and an infrared image feature extraction subnetwork with the same network architecture. The input of the visible light image feature extraction subnetwork is an image obtained by splicing the denoised visible light image and its enhanced image in the channel dimension, and the input of the infrared image feature extraction subnetwork is an image obtained by splicing the denoised infrared image and its enhanced image in the channel dimension. The visible light image feature extraction subnetwork and the infrared image feature extraction subnetwork both include 3×3 convolution, depth convolution, two ReLU functions, attention mechanism, residual module and fully connected layer. The processing process is specifically shown in the following formula (2): (2); In formula (2), represents a 3×3 convolution, represents the depthwise convolution, represents the ReLU function, represents the attention mechanism, represents the residual module, represents the fully connected layer, is the input of the visible light image feature extraction subnetwork or the infrared image feature extraction subnetwork, is the feature map output after deep convolution, It is the output of the visible light image feature extraction subnetwork or the infrared image feature extraction subnetwork.

4. The method for detecting underground abnormal events based on multispectral image fusion according to claim 3, characterized in that: The attention mechanism includes a channel attention mechanism and a spatial attention mechanism. The channel attention mechanism first performs global average pooling and global maximum pooling operations on the input feature map, and then fuses and splices the vectors after global average pooling and global maximum pooling together and compresses the dimension through the fully connected layer, and then uses the Sigmoid function to normalize. The output of the channel attention mechanism is used as the input of the spatial attention mechanism. The spatial attention mechanism first performs global average pooling and global maximum pooling operations on the input, and then fuses and splices the two spatial maps output after global average pooling and global maximum pooling in the channel dimension, and then uses 1×1 convolution and Sigmoid function to compress and normalize the feature map, and outputs the attention weight map.

5. The method for detecting underground abnormal events based on multispectral image fusion according to claim 3, characterized in that: The feature fusion module first performs feature splicing on the feature maps output by the visible light image feature extraction subnetwork and the infrared image feature extraction subnetwork with the same network architecture, and then unifies the spliced ​​features into the same feature space through the first 3×3 convolution, and then respectively uses convolutions with different receptive field sizes to extract features from the output of the first 3×3 convolution, and finally uses the second 3×3 convolution to perform feature fusion on the features extracted by the convolutions with different receptive field sizes.

6. The method for detecting underground abnormal events based on multispectral image fusion according to claim 3, characterized in that: During the training process of the multispectral image detection model, the multispectral image detection model is first trained once, and then the twin network is randomly pruned multiple times, and the multispectral image detection model composed of the twin network with the best pruning effect is selected for re-training.

Citation Information

Patent Citations

  • Flame detection method and device based on visible light image and near-infrared image

    CN118691932A