Coal mine image segmentation algorithm based on Swinin-UMama
Through the coal mine image segmentation algorithm based on Swin-UMamba, combined with dual-branch dark light enhancement and multi-loss function optimization, the accuracy and robustness of image segmentation in low-light environment of coal mines is solved, and efficient image segmentation effect is achieved.
Patent Information
- Application Number
- CN202510344634.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-04
AI Technical Summary
The existing image segmentation method is poor in low light and complex backgrounds of coal mines, making it difficult to achieve high-precision and robust image segmentation.
The coal mine image segmentation algorithm based on Swin-UMamba is adopted, combining dual-branch dark light enhancement, adaptive multi-scale feature extraction of Swin-UMamba architecture and three-channel parallel processing, and through multi-loss function optimization, the accuracy and boundary clarity of image segmentation are improved.
In low-light environments, the accuracy of image segmentation and boundary clarity are significantly improved, the adaptability and robustness of the model in complex environments are enhanced, and technical support is provided for coal mine safety monitoring and management.
Smart Images

Figure CN120259331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and particularly to a coal mine image segmentation algorithm based on Swin-UMamba. Background Art
[0002] Image segmentation is a fundamental problem in computer vision. The main purpose is to divide different regions in an image so that pixels within each region belong to the same category. Common image segmentation methods include threshold-based segmentation methods, region growing methods, edge detection methods, and deep learning-based segmentation methods. Among them, deep learning-based image segmentation methods have made significant progress in recent years, especially convolutional neural networks (CNNs) and their variant architectures. Image segmentation has a wide range of applications in multiple fields.
[0003] Due to the particularity of the coal mine environment, the coal mine image segmentation problem faces many challenges, such as low light, complex background noise, dust interference, etc. Traditional image processing methods usually cannot achieve good segmentation effects under low light or complex backgrounds.
[0004] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a coal mine image segmentation algorithm based on Swin-UMamba to solve the technical problems proposed in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A coal mine image segmentation algorithm based on Swin-UMamba, at least including the following steps:
[0007] S1: Build a dual-branch low-light enhancement deep learning architecture to improve the image quality and segmentation performance under low light conditions. The dual-branch low-light enhancement deep learning architecture includes a lower branch and an upper branch;
[0008] S2: During the model training process, comprehensively use a variety of different loss functions for optimization and compensation;
[0009] S3: Build a Swin-UMamba architecture based on VSS blocks. After completing low-light enhancement, use the Swin-UMamba architecture for high-level feature extraction. The Swin-UMamba architecture has the characteristics of Adaptive Multi-Scale Feature Extraction, that is, by dynamically adjusting the receptive field and feature levels, it can effectively capture detail information of different scales in the image and adapt to the diversity and complexity of target objects in the coal mine environment;
[0010] S4: Build a three-channel parallel processing module to further enhance feature expression and segmentation effect;
[0011] S5: Combine the deep learning architecture of dual-branch low-light enhancement, the Swin-UMamba architecture, and the three-channel parallel processing module, and optimize it using the multi-loss function compensation given in S2 to obtain a coal mine image segmentation model, and use the coal mine image segmentation model to process coal mine images.
[0012] Preferably, the lower branch uses the classic U-Net architecture for feature extraction. Through its symmetric encoder and decoder structure, U-Net can effectively capture the multi-scale features of the image and retain rich spatial information, which is suitable for fine boundary recognition in segmentation tasks. The lower branch includes an encoder and a decoder, and there is a skip connection between the encoder and the decoder. The encoder is downsampling, and the decoder is upsampling.
[0013] Preferably, the upper branch introduces Large Kernel Convolution to expand the receptive field of the convolution operation, thereby enhancing the ability to capture global information and detailed features of the image. Large Kernel Convolution helps to identify more subtle features in the complex coal mine environment and improve the accuracy of segmentation. The upper branch uses a branch interaction mechanism and a multi-layer perceptron;
[0014] The representation of the branch interaction mechanism is shown in the following formula:
[0015]
[0016] x2 = Conv(x1),
[0017] x3 = Concat(DDConv73(x2),
[0018] DDConv53(x2),
[0019] DDConv13(x2))
[0020] Among them, PWConv represents pointwise convolution; Conv represents a convolution kernel size of 5; DDConv73 represents a 7x7 depthwise dilated convolution with a dilation rate of 3; DDConv53 represents a 5×5 depthwise dilated convolution with a dilation rate of 3; DWDConv13 represents a 3×3 depthwise dilated convolution with a dilation rate of 3; Concat represents concatenating features in the channel dimension;
[0021] The three parallel depthwise dilated convolutions with different kernel sizes can extract multi-scale features. Among them, the large dilation convolution (7x7) and the medium dilation convolution (5×5) have long-range modeling and a larger receptive field, just like self-attention in Transformer;
[0022] The multi-layer perceptron is a feed-forward neural network used to achieve information interaction between different layers. The multi-layer perceptron is MLP for short.
[0023] In the architecture of the upper and lower branches, the MLP realizes feature sharing and fusion by connecting the output layers of the upper and lower branches, thereby improving the expressiveness and robustness of the model. The MLP combines several fully connected layers with activation functions (such as ReLU, Sigmoid, etc.) to perform non-linear transformation on the input features, and then extracts and integrates the features of the upper and lower branches.
[0024] Preferably, the multiple different loss functions in S2 include but are not limited to cross-entropy loss, Dice loss, focal loss, margin loss, brightness loss, structural loss, color loss, total variation loss, perceptual loss, adversarial loss, and a comprehensive loss function that combines the above ten losses.
[0025] The application of the multiple different loss functions aims to optimize the model performance from multiple dimensions to ensure that the segmentation results achieve the best effects in terms of accuracy, boundary clarity, robustness, etc.
[0026] The multiple different loss functions coordinate and optimize the performance of the model on different features through different weight combinations, making up for the limitations that may exist in a single loss function, thereby achieving more refined and reliable image segmentation.
[0027] Finally, expressed according to the comprehensive loss function, refer to the following formula:
[0028] L total = λ1L CE + λ2L Dice + λ3L Focal + λ4L Edge + λ5L Brightness + λ6L Structure + λ7L Color + λ8L TV + λ9L Perceptual + λ 10 L adv
[0029] Among them, λ1, λ2, …, λ 10 respectively represent the weight coefficients of the ten loss functions, used to adjust the contribution of different loss functions to the total loss.
[0030] According to the requirements of the task, these weights are adjusted to balance the influence of different loss functions.
[0031] Preferably, the operations of the VSS block include, but are not limited to, linear transformation, layer normalization, SS2D, depthwise convolution, and output merging;
[0032] Assume the input feature is X ∈ R H×W×C , where H is the height of the image, W is the width of the image, C is the number of channels of the image, and the processing process of the VSS block can be divided into the following steps:
[0033] S3.1: Linear transformation (Linear) performs feature mapping through a weight matrix:
[0034] X lin = WX + b
[0035] where W is the weight matrix and b is the bias term;
[0036] S3.2: Layer normalization (Layer Norm) is used to accelerate training. By performing normalization operations on each layer, the formula is as follows:
[0037]
[0038] where μ is the mean, σ is the standard deviation, and γ and β are learnable scaling factors and offsets;
[0039] S3.3: SS2D (Selective Scan Space State Sequential Model) enhances feature information through scanning operations in four directions. The specific steps are as follows:
[0040] Expand operation (expand), according to the scanning direction v ∈ {1, 2, 3, 4}, expand the input feature z:
[0041] z v = expand(z, v)
[0042] The SSM operation, namely S6, performs SSM processing on the expanded feature z v , where S6 is the core Scan Space State Sequential Model (SSM), which processes the expanded features in each direction;
[0043] Merge operation (merge), merge the features in four directions:
[0044]
[0045] The merge operation integrates the features in four directions into a complete 2D feature map;
[0046] S3.4: Depthwise Convolution (DW Conv) performs convolution operations on each channel, and the calculation formula is:
[0047] F dw = DepthwiseConv(F)
[0048] S3.5: Combine the outputs. The final output of the VSS block is obtained by weighted combination of features from multiple branches. The formula is:
[0049]
[0050] where represents the weighted combination operation, X dw represents the features after depthwise convolution, X att represents the features obtained through SS2D and other operations.
[0051] Preferably, the three channels in S4 include a wavelet module extraction channel, a spatial pyramid channel, and a channel attention channel;
[0052] The wavelet module extraction channel (Wavelet Module Extraction Path) uses wavelet transform to perform multi-frequency decomposition on the image, extracts detailed features of different frequency bands, and enhances the texture information and edge details of the image;
[0053] The spatial pyramid channel (Spatial Pyramid Path) applies spatial pyramid pooling (Spatial Pyramid Pooling, SPP). Through a multi-level spatial pyramid structure, it performs multi-scale spatial aggregation on the feature map to capture spatial context information at different scales;
[0054] The channel attention channel (Channel Attention Path) introduces a channel attention mechanism. By adaptively adjusting the weights of each channel, it strengthens the information expression of important feature channels, suppresses irrelevant or redundant features, and improves the effectiveness of feature representation.
[0055] Preferably, in the wavelet module extraction channel, wavelet transform is a commonly used tool in image processing, which can decompose the image into sub-bands of different frequencies to capture the detailed information in the image;
[0056] Assume the input image X ∈ R H×W×C , where H is the height of the image, W is the width of the image, and C is the number of channels of the image. After wavelet transform processing, coefficients W k of different frequency bands are obtained, where k represents the frequency band index;
[0057] The wavelet transform formula is expressed as:
[0058]
[0059] where, represents the wavelet transform operation; f k represents different frequency bands;
[0060] Through wavelet transform, the image is decomposed into multiple frequency bands, containing different detailed information.
[0061] Preferably, in the spatial pyramid channel, spatial pyramid pooling (SPP) performs multi-scale spatial aggregation on the feature map through multi-level pooling operations, so as to capture spatial context information at different scales;
[0062] Assume the input feature map is where H is the height of the image, W is the width of the image, and C is the number of channels of the image. Through spatial pyramid pooling, pooling results of different scales are generated, and the specific pooling scales are {1×1, 2×2, 4×4};
[0063] The operation of spatial pyramid pooling is expressed as:
[0064] F SPP = SPP(F) = [Maxpool(F, 1×1), Maxpool(F, 2×2), Maxpool(F, 4×4)]
[0065] where, Maxpool(F, k×k) represents the maximum pooling operation on the feature map with a size of k×k to obtain spatial information at different scales.
[0066] Preferably, the channel attention mechanism enhances the features of important channels by adaptively adjusting the weights of each channel, while suppressing irrelevant or redundant features;
[0067] Assume the feature map is where H is the height of the image, W is the width of the image, and C is the number of channels of the image. Through the channel attention mechanism, the channel weight
[0068] A i = σ(W a ·F i + b a )
[0069] where, F i is the feature of the i-th channel, W a and b aare the learned weights and biases, σ is the activation function (such as Sigmoid), and A i is the channel attention coefficient.
[0070] The final weighted feature representation is:
[0071] F att = A ⊙ F
[0072] where ⊙ represents element-wise multiplication per channel, A is the channel attention weight, F is the input feature map, and the finally obtained F att is the feature map adjusted by the attention mechanism.
[0073] Compared with the prior art, the beneficial effects of the present invention are:
[0074] By combining dual-branch low-light enhancement, adaptive multi-scale feature extraction of the Swin-UMamba architecture, and the design of three-channel parallel processing (wavelet module, spatial pyramid, channel attention), and supplemented by the comprehensive optimization of multiple loss functions, the present invention obtains a coal mine image segmentation model. The proposed coal mine image segmentation model shows excellent performance in the image segmentation task under low-light environments in coal mines. This method not only significantly improves the segmentation accuracy and boundary clarity but also enhances the adaptability and robustness of the model in complex environments, providing strong technical support for coal mine safety monitoring and management. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0076] Figure 1 is the overall structural schematic diagram of the segmentation model of the present invention;
[0077] Figure 2 is the schematic diagram of the deep learning architecture of dual-branch low-light enhancement of the present invention;
[0078] Figure 3 is the schematic diagram of upsampling of the present invention;
[0079] Figure 4 is the schematic diagram of the feature extraction part of the Swin-UMamba architecture of the present invention;
[0080] Figure 5 is the schematic diagram of the VSS block of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0082] A coal mine image segmentation algorithm based on Swin-UMamba includes at least the following steps:
[0083] S1: Build a deep learning architecture with dual-branch low-light enhancement (refer to Figure 1 and Figure 2 ) to improve the image quality and segmentation performance under low-light conditions. The deep learning architecture with dual-branch low-light enhancement includes a lower branch and an upper branch;
[0084] The lower branch uses a classic U-Net architecture for feature extraction. Through its symmetric encoder and decoder structures, U-Net can effectively capture the multi-scale features of images and retain rich spatial information, which is suitable for fine boundary recognition in segmentation tasks. The lower branch includes an encoder and a decoder, and there is a skip connection between the encoder and the decoder. The encoder is downsampling, and the decoder is upsampling;
[0085] The theoretical formula of the U-Net architecture assumes that the input image is X ∈ R H×W×C , where H is the height of the image, W is the width of the image, and C is the number of channels of the image;
[0086] The processing process of the U-Net architecture includes at least the following steps:
[0087] S1.1.1: In the encoder part, the input image undergoes downsampling through several convolutional layers and pooling layers. The convolutional layer extracts the local features of the image through a convolutional kernel. The calculation of each convolutional layer can be expressed as:
[0088] Y k =Conv(X, W k ) + b k
[0089] where X is the input feature, W k is the convolutional kernel of the k-th layer, b k is the bias, and Y k is the output of this layer; then downsampling is performed through the pooling layer. Taking max pooling as an example:
[0090] Y pool =MaxPool(Y k )
[0091] S1.1.2: In the decoder part, the upsampling operation is used to restore the size of the feature map to the original image size, and the segmentation result is refined through the convolution operation. The upsampling layer can be represented by the transposed convolution (Deconvolution) as follows:
[0092] Y up = DeConv(Y k , W up ) + b up
[0093] where Y k is the output of a certain layer in the encoder, W up is the transposed convolution kernel, b up is the bias term, and Y up is the upsampled feature;
[0094] S1.1.3: In U-Net, the skip connection plays a role in connecting the information between the encoder and the decoder. Assuming that the output of a certain layer in the encoder is Y cnc , and a certain layer in the decoder is Y dcc , then the skip connection directly merges Y cnc with Y dcc for example, through the connection:
[0095] Y d ′ cc = concat(Y dcc , Y cnc )
[0096] S1.14: Final output. The U-Net architecture uses a convolutional layer to generate the final segmentation result:
[0097] Y output = Conv(Y d ′ cc , W final ) + b final
[0098] where Y d ′ cc is the feature after merging the decoder part and the skip connection, W final is the last convolutional kernel, b final is the bias, and the final output Y output is the segmentation result.
[0099] The upper branch introduces Large Kernel Convolution to expand the receptive field of the convolution operation, thereby enhancing the ability to capture global image information and detailed features. Large Kernel Convolution helps identify more subtle features in complex coal mine environments and improve the accuracy of segmentation. The upper branch adopts a branch interaction mechanism and a multi-layer perceptron;
[0100] The representation of the branch interaction mechanism is shown in the following formula:
[0101]
[0102] x2 = Conv(x1),
[0103] x3 = Concat(DDConv73(x2),
[0104] DDConv53(x2),
[0105] DDConv13(x2))
[0106] Among them, PWConv represents pointwise convolution; Conv represents a convolution kernel size of 5; DDConv73 represents a 7x7 depthwise dilated convolution with a dilation rate of 3; DDConv53 represents a 5x5 depthwise dilated convolution with a dilation rate of 3; DWDConv13 represents a 3x3 depthwise dilated convolution with a dilation rate of 3; Concat represents concatenating features in the channel dimension;
[0107] Three parallel depthwise dilated convolutions with different kernel sizes can extract multi-scale features. Among them, the large dilation convolution (7x7) and the medium dilation convolution (5x5) have long-range modeling and a larger receptive field, just like self-attention in Transformer;
[0108] A multi-layer perceptron is a feedforward neural network used to achieve information interaction between different levels. The multi-layer perceptron is MLP for short;
[0109] In the architecture of the upper and lower branches, MLP realizes feature sharing and fusion by connecting the output layers of the upper and lower branches, thereby enhancing the expressiveness and robustness of the model. MLP combines several fully connected layers and activation functions (such as ReLU, Sigmoid, etc.) to non-linearly transform the input features, and then extracts and integrates the features of the upper and lower branches. The basic working principle of the MLP module can be expressed by the following formula:
[0110] 1. Feature representation of the upper and lower branches
[0111] Assume that the output features of the upper and lower branches are h upper and h lower, with dimensions d1 and d2 respectively. The input to the MLP module is the concatenation or weighted combination of these two feature vectors:
[0112] h input = concat(h upper , h lower ) or h input = W · h upper + h upper + W · h lower + h lower )
[0113] where concat represents the concatenation operation, and W and b are the weighted coefficients and bias terms.
[0114] 2. Feature transfer through the hidden layer
[0115] The MLP performs a non - linear transformation on the input features through multiple hidden layers. Suppose there are l hidden layers, and the output of each layer is calculated from the output of the previous layer:
[0116] h l = σ(W l · h l-1 + b l )
[0117] where σ is the activation function, such as ReLU, and W l and b l are the weights and bias of the l - th layer respectively.
[0118] 3. Output layer
[0119] The final output of the MLP is calculated through a linear layer (usually the Softmax or Sigmoid function):
[0120] y = softmax(W out · h L + b out )
[0121] where W out and b out are the weights and bias of the output layer, and softmax is the activation function for multi - class classification.
[0122] 4. Loss function
[0123] Usually, the training of the MLP is optimized by minimizing the loss function, such as the cross - entropy loss function:
[0124] Loss):
[0125]
[0126] Among them, y i is the true label, and
[0127] S2: During the model training process, a variety of different loss functions are comprehensively used for optimization and compensation;
[0128] The various different loss functions in S2 include, but are not limited to, cross-entropy loss, Dice loss, focal loss, edge loss, brightness loss, structure loss, color loss, total variation loss, perceptual loss, adversarial loss, and a comprehensive loss function that combines the above ten losses;
[0129] The application of various different loss functions aims to optimize the model performance from multiple dimensions, ensuring that the segmentation results achieve the best effects in terms of accuracy, boundary clarity, robustness, etc.;
[0130] The various different loss functions coordinate and optimize the model's performance on different features through different weight combinations, making up for the possible limitations of a single loss function, so as to achieve more refined and reliable image segmentation;
[0131] Cross-entropy loss
[0132]
[0133] Cross-entropy loss is used in pixel classification tasks to evaluate the difference between the probability distribution predicted by the model and the target label.
[0134] Dice loss
[0135]
[0136] This loss is used for segmentation tasks, especially when dealing with class imbalance problems, to calculate the similarity between the prediction and the true segmentation.
[0137] Focal loss
[0138] L Focal = -α(1 - p t ) γ log(p t )
[0139] Focal loss emphasizes the training of difficult samples by reducing the attention to easily classified samples, and is often used to deal with class imbalance problems.
[0140] Edge loss
[0141]
[0142] The edge loss is used to ensure that the edge information of the enhanced image and the reference image is consistent, helping the network to focus on the edge details in the image.
[0143] Brightness Loss
[0144]
[0145] The brightness loss ensures that the brightness order of the enhanced image and the reference image is consistent, especially for the enhancement of low-light images.
[0146] Structure Loss
[0147]
[0148] The structure loss ensures the structural consistency of the enhanced image and the reference image by comparing the gradient information.
[0149] Color Loss
[0150]
[0151] The color loss is used to optimize the hue and saturation consistency between the enhanced image and the reference image, especially in image enhancement.
[0152] Total Variation Loss
[0153]
[0154] The total variation loss is used to smooth the image, reduce noise and preserve structural details, and is commonly used in image denoising and image enhancement tasks.
[0155] Perceptual Loss
[0156]
[0157] The perceptual loss measures the perceptual difference between the enhanced image and the reference image by extracting the deep features of the image, and is commonly used in tasks such as image super-resolution and style transfer.
[0158] Finally, it is expressed according to the comprehensive loss function, referring to the following formula:
[0159] L total = λ1L CE + λ2L Dice + λ3L Focal + λ4L Edge + λ5L Brightness + λ6L Structure + λ7LColor
[0160] +λ8L TV +λ9L Perceptual +λ 10 L adv
[0161] Among them, λ1, λ2, …, λ 10 respectively represent the weight coefficients of ten loss functions, which are used to adjust the contributions of different loss functions in the total loss;
[0162] According to the requirements of the task, these weights are adjusted to balance the impacts of different loss functions.
[0163] S3: Build a Swin-UMamba architecture based on the VSS block (refer to Figure 4 ), after completing low-light enhancement, use the Swin-UMamba architecture for advanced feature extraction. The Swin-UMamba architecture has the characteristics of Adaptive Multi-Scale Feature Extraction, that is, by dynamically adjusting the receptive field and feature levels, it can effectively capture the detailed information of different scales in the image and adapt to the diversity and complexity of target objects in the coal mine environment;
[0164] The operations of the VSS block (refer to Figure 5 ) include but are not limited to linear transformation, layer normalization, SS2D, Depthwise Convolution, and merging outputs;
[0165] Assume the input feature is X ∈ R H×W×C , where H is the height of the image, W is the width of the image, and C is the number of channels of the image. The processing process of the VSS block is divided into the following steps:
[0166] S3.1: Linear transformation (Linear) performs feature mapping through a weight matrix:
[0167] X lin = WX + b
[0168] where W is the weight matrix and b is the bias term;
[0169] S3.2: Layer normalization (Layer Norm) is used to accelerate training. By performing normalization operations on each layer, the formula is as follows:
[0170]
[0171] where μ is the mean, σ is the standard deviation, and γ and β are learnable scaling factors and offsets;
[0172] S3.3: SS2D (Selective Scan Space State Sequential Model) enhances the feature information through scanning operations in four directions. The specific steps are as follows:
[0173] Expand operation (expand), according to the scanning direction v ∈ {1, 2, 3, 4}, expand the input feature z:
[0174] z v = expand(z, v)
[0175] The SSM operation, i.e., S6, performs SSM processing on the expanded feature z v , where S6 is the core Scan Space State Sequential Model (SSM), which processes the expanded features in each direction;
[0176] Merge operation (merge), merge the features in four directions:
[0177]
[0178] The merge operation integrates the features in four directions into a complete 2D feature map;
[0179] S3.4: Depthwise Convolution (DW Conv) performs convolution operations on each channel, and the calculation formula is:
[0180] F dw = DepthwiseConv(F)
[0181] S3.5: Merge the outputs. The final output of the VSS block is obtained by weighted merging of the features from multiple branches. The formula is:
[0182]
[0183] Among them, represents the weighted merging operation, X dw represents the feature after depthwise convolution, X att represents the feature obtained through SS2D and other operations.
[0184] S4: Build a three-channel parallel processing module to further enhance the feature expression and segmentation effect;
[0185] The three channels in S4 include the wavelet module extraction channel, the spatial pyramid channel, and the channel attention channel;
[0186] The Wavelet Module Extraction Path uses wavelet transform to perform multi-frequency decomposition on images, extract detailed features in different frequency bands, and enhance the texture information and edge details of images;
[0187] The Spatial Pyramid Path applies Spatial Pyramid Pooling (SPP). Through a multi-level spatial pyramid structure, it performs multi-scale spatial aggregation on the feature map to capture spatial context information at different scales;
[0188] The Channel Attention Path introduces a channel attention mechanism. By adaptively adjusting the weights of each channel, it strengthens the information expression of important feature channels, suppresses irrelevant or redundant features, and improves the effectiveness of feature representation.
[0189] In the Wavelet Module Extraction Path, wavelet transform is a commonly used tool in image processing. It can decompose an image into sub-bands of different frequencies to capture the detailed information in the image;
[0190] Assume the input image \(X\in\mathbb{R}^{H\times W\times C}\), H×W×C where \(H\) is the height of the image, \(W\) is the width of the image, and \(C\) is the number of channels of the image. After wavelet transform processing, coefficients \(W_k\) of different frequency bands are obtained, k where \(k\) represents the frequency band index;
[0191] The wavelet transform formula is expressed as:
[0192]
[0193] where, represents the wavelet transform operation; \(f_k\) k represents different frequency bands;
[0194] Through wavelet transform, the image is decomposed into multiple frequency bands, which contain different detailed information.
[0195] In the Spatial Pyramid Path, Spatial Pyramid Pooling (SPP) performs multi-scale spatial aggregation on the feature map through multi-level pooling operations to capture spatial context information at different scales;
[0196] Assume the input feature map is where \(H\) is the height of the image, \(W\) is the width of the image, and \(C\) is the number of channels of the image. Through spatial pyramid pooling, pooling results of different scales are generated, and the specific pooling scales are \(\{1\times1, 2\times2, 4\times4\}\);
[0197] The operation of spatial pyramid pooling is expressed as:
[0198] F SPP = SPP(F) = [Maxpool(F, 1×1), Maxpool(F, 2×2), Maxpool(F, 4×4),]
[0199] where Maxpool(F, k×k) represents the max-pooling operation on the feature map with a size of k×k to obtain spatial information at different scales.
[0200] The channel attention mechanism enhances the features of important channels by adaptively adjusting the weights of each channel, while suppressing irrelevant or redundant features;
[0201] Assume the feature map is where H is the height of the image, W is the width of the image, and C is the number of channels of the image. The channel weights can be obtained through the channel attention mechanism
[0202] A i = σ(W a ·F i + b a )
[0203] where F i is the feature of the i-th channel, W a and b a are the learned weights and biases, σ is the activation function (such as Sigmoid), and A i is the channel attention coefficient.
[0204] The final weighted feature is expressed as:
[0205] F att = A ⊙ F
[0206] where ⊙ represents element-wise multiplication per channel, A is the attention weight of the channel, F is the input feature map, and the finally obtained F att is the feature map adjusted by the attention mechanism.
[0207] S5: Combine the deep learning architecture for dual-branch low-light enhancement, the Swin-UMamba architecture, and the three-channel parallel processing module, and optimize it using the multi-loss function compensation given in S2 to obtain a coal mine image segmentation model, and use the coal mine image segmentation model for coal mine image processing;
[0208] Integrating the above feature extraction and enhancement modules, the coal mine images are finally accurately segmented by the advanced segmentation network. This entire process not only improves the visualization quality of the images but also achieves efficient and robust image segmentation effects in complex low-light environments through the collaborative optimization of multiple levels and multiple modules.
[0209] In summary, the present invention also has the following advantages:
[0210] Heterogeneous model fusion: The combination of Swin Transformer and Mamba takes into account both local details and global dependencies, which is superior to a single architecture.
[0211] Scene customization design: Dual-branch enhancement, dynamic multi-scale extraction, and three-channel parallel processing are all optimized for challenges such as low light, dust, and target diversity in coal mines.
[0212] Multi-task joint optimization: By balancing segmentation accuracy and enhancement effects through a complex loss function, it surpasses the traditional "enhancement first then segmentation" pipeline mode.
[0213] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be construed as limiting the claimed rights.
Claims
1. A coal mine image segmentation algorithm based on Swin-UMamba, characterized in that: At least include the following steps: S1: Build a deep learning architecture for dual-branch low-light enhancement to improve image quality and segmentation performance under low-light conditions. The deep learning architecture for dual-branch low-light enhancement includes a lower branch and an upper branch; S2: During model training, comprehensively use a variety of different loss functions for optimization and compensation; S3: Build a Swin-UMamba architecture based on VSS blocks. After low-light enhancement is completed, use the Swin-UMamba architecture for advanced feature extraction. The Swin-UMamba architecture has the characteristics of adaptive multi-scale feature extraction, that is, by dynamically adjusting the receptive field and feature levels, it can effectively capture detailed information of different scales in the image and adapt to the diversity and complexity of target objects in the coal mine environment; S4: Build a three-channel parallel processing module to further enhance feature expression and segmentation effects; S5: Combine the deep learning architecture for dual-branch low-light enhancement, the Swin-UMamba architecture, and the three-channel parallel processing module, and use the multi-loss function compensation given in S2 for optimization to obtain a coal mine image segmentation model, and use the coal mine image segmentation model for coal mine image processing.
2. The coal mine image segmentation algorithm based on Swin-UMamba according to claim 1, wherein: The lower branch uses the classic U-Net architecture for feature extraction. U-Net can effectively capture multi-scale features of the image and retain rich spatial information through its symmetric encoder and decoder structures, and is suitable for fine boundary recognition in segmentation tasks. The lower branch includes an encoder and a decoder, and there is a skip connection between the encoder and the decoder. The encoder is downsampling, and the decoder is upsampling.
3. The coal mine image segmentation algorithm based on Swin-UMamba according to claim 2, characterized in that: The upper branch introduces large kernel convolution to expand the receptive field of the convolution operation, thereby enhancing the ability to capture global information and detailed features of the image. Large kernel convolution helps to identify more subtle features in the complex coal mine environment and improve the accuracy of segmentation. The upper branch uses a branch interaction mechanism and a multi-layer perceptron; The representation of the branch interaction mechanism is shown in the following formula: x2 = Conv(x1), x3 = Concat(DDConv73(x2), DDConv53(x2), DDConv13(x2)) where PWConv represents pointwise convolution; Conv represents a convolution kernel size of 5; DDConv73 represents a 7x7 depthwise dilated convolution with a dilation rate of 3; DDConv53 represents a 5×5 depthwise dilated convolution with a dilation rate of 3; DWDConv13 represents a 3×3 depthwise dilated convolution with a dilation rate of 3; Concat represents concatenating features in the channel dimension; The three parallel depthwise dilated convolutions with different kernel sizes can extract multi-scale features, where the large dilation convolution and the medium dilation convolution (have long-range modeling and a larger receptive field; The multi-layer perceptron is a feedforward neural network used to achieve information interaction between different levels. The multi-layer perceptron is MLP; In the architecture of the upper and lower branches, the MLP realizes the sharing and fusion of features by connecting the output layers of the upper and lower branches, thereby enhancing the expressiveness and robustness of the model. The MLP combines several fully connected layers and activation functions to perform non-linear transformation on the input features, and then extracts and integrates the features of the upper and lower branches.
4. The coal mine image segmentation algorithm based on Swin-UMamba according to claim 3, characterized in that: The various different loss functions in S2 include but are not limited to cross-entropy loss, Dice loss, focal loss, edge loss, brightness loss, structural loss, color loss, total variation loss, perceptual loss, adversarial loss, and a comprehensive loss function that combines the above ten losses; The application of the various different loss functions aims to optimize the model performance from multiple dimensions to ensure that the segmentation results achieve the best effects in terms of accuracy, boundary clarity, robustness, etc.; The various different loss functions coordinate and optimize the performance of the model on different features through different weight combinations, making up for the limitations that may exist in a single loss function, thereby achieving more refined and reliable image segmentation; Finally, expressed according to the comprehensive loss function, refer to the following formula: L total = λ1L CE + λ2L Dice + λ3L Focal + λ4L Edge + λ5L Brightness + λ6L Structure + λ7L Color + λ8L TV + λ9L Perceptual + λ 10 L adv Among them, λ1, λ2, …, λ 10 respectively represent the weight coefficients of ten loss functions, which are used to adjust the contributions of different loss functions to the total loss; According to the requirements of the task, adjust these weights to balance the influence of different loss functions.
5. The coal mine image segmentation algorithm based on Swin-UMamba according to claim 4, wherein: The operations of the VSS block include but are not limited to linear transformation, layer normalization, SS2D, depthwise convolution, and merging outputs; Suppose the input feature is \(X\in\mathbb{R}\) H×W×C , where \(H\) is the height of the image, \(W\) is the width of the image, and \(C\) is the number of channels of the image. The processing procedure of the VSS block is divided into the following steps: S3.1: Linear transformation performs feature mapping through a weight matrix: X lin = WX + b where W is the weight matrix and b is the bias term; S3.2: Layer normalization is used to accelerate training by normalizing each layer. The formula is as follows: where μ is the mean, σ is the standard deviation, and γ and β are learnable scaling factors and offsets; S3.3: SS2D enhances feature information through scanning operations in four directions. The specific steps are as follows: Expansion operation: Expand the input feature z according to the scanning direction v ∈ {1, 2, 3, 4}; z v = expand(z, v) The SSM operation, namely S6, is performed on the expanded feature z v , and SSM processing is carried out: Here, S6 is the core scanning space state sequence model, which processes the expanded features in each direction; Merging operation: Merge the features in the four directions; The merging operation integrates the features in the four directions into a complete 2D feature map; S3.4: Depthwise convolution performs convolution operations on each channel. The calculation formula is: F dw = DepthwiseConv(F) S3.5: Merging outputs: The final output of the VSS block is obtained by weighted merging of the features from multiple branches. The formula is: Among them, represents a weighted merging operation, and X dw represents the feature after depth convolution, and X att represents the feature obtained through SS2D and other operations.
6. The coal mine image segmentation algorithm based on Swin-UMamba according to claim 5, characterized in that: The three channels in S4 include the wavelet module extraction channel, the spatial pyramid channel, and the channel attention channel; The wavelet module extraction channel uses wavelet transform to perform multi-frequency decomposition on the image, extracts the detailed features of different frequency bands, and enhances the texture information and edge details of the image; The spatial pyramid channel applies spatial pyramid pooling. Through a multi-level spatial pyramid structure, it performs multi-scale spatial aggregation on the feature map to capture the spatial context information at different scales; The channel attention channel introduces a channel attention mechanism. By adaptively adjusting the weights of each channel, it strengthens the information expression of important feature channels, suppresses irrelevant or redundant features, and improves the effectiveness of feature representation.
7. The coal mine image segmentation algorithm based on Swin-UMamba according to claim 6, characterized in that: In the wavelet module extraction channel, wavelet transform is a commonly used tool in image processing that can decompose an image into sub-bands of different frequencies to capture the detailed information in the image; Assume the input image \(X\in\mathbb{R}\) H×W×C , where \(H\) is the height of the image, \(W\) is the width of the image, and \(C\) is the number of channels of the image. After wavelet transform processing, coefficients \(W\) of different frequency bands are obtained k , where \(k\) represents the frequency band index; The wavelet transform formula is expressed as: Among them, represents a wavelet transform operation; f k represents different frequency bands; Through wavelet transform, the image is decomposed into multiple frequency bands, containing different details of information.
8. A coal mine image segmentation algorithm based on Swin-UMamba according to claim 6, characterized in that: In the spatial pyramid channel, spatial pyramid pooling performs multi-scale spatial aggregation on the feature map through multi-level pooling operations, so as to capture spatial context information at different scales; Assume that the input feature map is where H is the height of the image, W is the width of the image, and C is the number of channels of the image. Through spatial pyramid pooling, pooling results of different scales are generated, and the specific pooling scales are {1×1, 2×2, 4×4}; The operation of spatial pyramid pooling is expressed as: F SPP = SPP(F) = [Maxpool(F, 1×1), Maxpool(F, 2×2), Maxpool(F, 4×4),] Among them, Maxpool(F,k×k) represents the maximum pooling operation on the feature map with a size of k×k, obtaining spatial information at different scales.
9. The coal mine image segmentation algorithm based on Swin-UMamba according to claim 6, wherein: The channel attention mechanism enhances the features of important channels by adaptively adjusting the weights of each channel, while suppressing irrelevant or redundant features; Suppose the feature map is where H is the height of the image, W is the width of the image, and C is the number of channels of the image. The channel weights are obtained through the channel attention mechanism A i = σ(W a ·F i + b a ) Among them, F i is the feature of the first i-th channel, W a and b a are the learned weights and biases, σ is the activation function, and A i is the channel attention coefficient; The final weighted feature is expressed as: F att = A ⊙ F Among them, ⊙ represents element-wise multiplication for each channel, A is the attention weight of the channel, F is the input feature map, and the finally obtained F att is the feature map adjusted by the attention mechanism.
Citation Information
Cited By
Coal mine image segmentation model and method based on VMama and multi-expert hybrid network and construction method thereof
CN121639706A
Coal mine image segmentation model, method and construction method based on VMamba and multi-expert hybrid network
CN121639706B