A weak interlayer semantic segmentation method and device based on deep learning and a medium
Patent Information
- Application Number
- CN202510895143.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
然而,传统图像处理方法依赖人工设计的特征(如边缘检测、纹理分析),难以应对软弱夹层低对比度、细长结构及复杂纹理的特性;而基于深度学习的模型虽能提取多尺度特征,但固定空洞率参数的ASPP模块对夹层曲率变化的适应性不足,且小目标分割精度低,加之缺乏地质学先验知识的引导,导致模型在数据稀缺场景下泛化能力有限
通过动态特征感知架构的自适应空洞卷积机制,根据输入图像局部特征动态调整卷积核感受野,有效增强对细长弯曲结构的边缘捕捉能力;结合跨尺度特征交互模块的浅层细节与深层语义融合策略,显著提升低对比度区域的夹层识别精度,减少漏分割现象。
Smart Images

Figure CN120807923B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of semantic segmentation, and in particular to a weak mezzanine semantic segmentation method, device, and medium based on deep learning. Background Technology
[0002] In geological exploration and engineering safety monitoring, semantic segmentation of weak interlayers (such as rock fractures and fault zones) is a crucial step in assessing geological stability and risk. However, traditional image processing methods rely on manually designed features (such as edge detection and texture analysis), which are insufficient to handle the characteristics of weak interlayers, such as low contrast, slender structures, and complex textures. While deep learning-based models can extract multi-scale features, the ASPP module with a fixed void ratio parameter is not adaptable to changes in interlayer curvature, and the segmentation accuracy for small targets is low. In addition, the lack of prior geological knowledge leads to limited generalization ability of the models in data-scarce scenarios.
[0003] Current technologies also face a contradiction between the scarcity of labeled data and deployment efficiency: labeling weak interlayers requires the participation of professionals, which is costly and difficult to obtain on a large scale; at the same time, existing models mostly use large-scale networks (such as ResNet-101), making it difficult to balance real-time performance and computational resource requirements on embedded devices. Although some studies have attempted to introduce dynamic feature perception, attention mechanisms, and adaptive loss functions, their technical solutions still suffer from problems such as parameter rigidity, insufficient integration of geological priors, and inadequate lightweight optimization.
[0004] Therefore, how to improve the segmentation accuracy and generalization ability of low-contrast weak sandwich layers has become an urgent technical problem to be solved. Summary of the Invention
[0005] This application provides a deep learning-based method, device, and medium for weak sandwich semantic segmentation, which aims to solve the following technical problem: improving the segmentation accuracy and generalization ability of low-contrast weak sandwich layers.
[0006] In a first aspect, embodiments of this application provide a deep learning-based semantic segmentation method for weak interlayers, characterized in that the method includes: acquiring weak interlayer image data, wherein the weak interlayer image data includes geological exploration images, engineering inspection images, and synthetic data; preprocessing the weak interlayer image data to generate normalized image data and enhanced image data; constructing a dynamic feature perception architecture to process the normalized image data and enhanced image data to generate a multi-scale fused feature map; the dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module; and processing the data through a geological prior-guided attention network. The multi-scale fusion feature map is processed to generate an attention-weighted feature map; the geological prior-guided attention network fuses channel attention weights, spatial attention weights, and geological structural parameter constraint factors; the attention-weighted feature map is processed based on a preset encoder-decoder structure to generate an initial segmentation prediction map; the difference between the initial segmentation prediction map and the true label is calculated through an adaptive loss function module to generate optimized gradient data; the model parameters are iteratively updated based on the optimized gradient data to generate a final semantic segmentation model; and the final semantic segmentation model is used to perform segmentation operations on the target weak interlayer image to generate an interlayer segmentation result map.
[0007] In one implementation of this application, the weak interlayer image data is preprocessed to generate normalized image data and enhanced image data. Specifically, this includes: extracting multi-level feature maps from the input image through an encoder to generate a first-level feature map, a second-level feature map, a third-level feature map, and a fourth-level feature map; processing the fourth-level feature map using an adaptive dilated convolution mechanism to generate dynamic dilation rate parameters and multi-branch output features; fusing the first-level feature map, the second-level feature map, the third-level feature map, and the multi-branch output features through a cross-scale feature interaction module to generate a multi-scale fused feature map; and introducing residual connections in the decoder to process the fourth-level feature map and the multi-scale fused feature map to generate a high-resolution feature map.
[0008] In one implementation of this application, an adaptive dilated convolution mechanism is used to process the fourth-level feature map to generate dynamic dilation rate parameters and multi-branch output features. Specifically, this includes: performing a global average pooling operation on the input feature map to generate local feature vectors; inputting the local feature vectors into a learnable parameter matrix and generating dynamic dilation rate parameters through Softmax normalization; adjusting the receptive field of the convolution kernel according to the dynamic dilation rate parameters, and performing multi-branch dilated convolution operations on the input feature map to generate multi-branch output features.
[0009] In one implementation of this application, a dynamic feature-aware architecture is constructed to process the normalized image data and enhanced image data to generate a multi-scale fusion feature map. Specifically, this includes: extracting channel features from the multi-scale fusion feature map and generating channel attention weights through a fully connected layer; extracting spatial features from the channel-weighted feature map and generating spatial attention weights through a convolutional layer; generating a layered geometric parameter map through a pre-trained geological feature extractor and converting it into a geological constraint factor; and fusing the spatial attention weights and the geological constraint factor to generate an attention-weighted feature map.
[0010] In one implementation of this application, a geometries map of interlayer geometry is generated by a pre-trained geological feature extractor and converted into a geological constraint factor. Specifically, this includes: extracting interlayer curvature parameters and extension direction parameters from the input image to generate a geological structural parameter tensor; performing a convolution operation on the geological structural parameter tensor to generate an initial constraint factor; and normalizing the initial constraint factor using the Sigmoid function to generate a geological constraint factor.
[0011] In one implementation of this application, the difference between the initial segmentation prediction map and the real label is calculated by an adaptive loss function module to generate optimized gradient data. Specifically, this includes: calculating the Dice loss value and Focal loss value between the initial segmentation prediction map and the real label; dynamically allocating the Dice loss weight and Focal loss weight according to the mezzanine density parameter to generate a weighted joint loss value; extracting the edge mask of the real label using an edge detection algorithm and calculating the edge-aware loss value; and fusing the weighted joint loss value and the edge-aware loss value to generate optimized gradient data.
[0012] In one implementation of this application, Dice loss weights and Focal loss weights are dynamically allocated based on the interlayer density parameter to generate a weighted joint loss value. Specifically, this includes: counting the number of interlayer pixels per unit area to generate an interlayer density parameter; inputting the interlayer density parameter into a Sigmoid function to generate dynamic Dice loss weights; and calculating dynamic Focal loss weights based on the dynamic Dice loss weights.
[0013] In one implementation of this application, the edge mask of the real label is extracted by an edge detection algorithm, and the edge-aware loss value is calculated. Specifically, this includes: performing Canny edge detection on the real label to generate a binary edge mask; extracting the pixel prediction values of the edge region in the initial segmentation prediction map to generate an edge prediction subset; and calculating the cross-entropy loss between the edge prediction subset and the real edge mask to generate the edge-aware loss value.
[0014] Secondly, embodiments of this application also provide a deep learning-based weak interlayer semantic segmentation device, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: acquire weak interlayer image data, the weak interlayer image data including geological exploration images, engineering inspection images, and synthetic data; preprocess the weak interlayer image data to generate normalized image data and enhanced image data; construct a dynamic feature perception architecture to process the normalized image data and enhanced image data, generating a multi-scale fusion feature map; The dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module. A geologically prior-guided attention network processes the multi-scale fused feature map to generate an attention-weighted feature map. This geologically prior-guided attention network fuses channel attention weights, spatial attention weights, and geological structural parameter constraint factors. Based on a pre-defined encoder-decoder structure, the attention-weighted feature map is processed to generate an initial segmentation prediction map. An adaptive loss function module calculates the difference between the initial segmentation prediction map and the true label to generate optimized gradient data. Based on the optimized gradient data, the model parameters are iteratively updated to generate a final semantic segmentation model. Based on the final semantic segmentation model, a segmentation operation is performed on the target weak interlayer image to generate an interlayer segmentation result map.
[0015] Thirdly, embodiments of this application also provide a non-volatile computer storage medium for weak interlayer semantic segmentation based on deep learning, storing computer-executable instructions, characterized in that the computer-executable instructions are configured to: acquire weak interlayer image data, the weak interlayer image data including geological exploration images, engineering inspection images, and synthetic data; preprocess the weak interlayer image data to generate normalized image data and enhanced image data; construct a dynamic feature perception architecture to process the normalized image data and enhanced image data to generate a multi-scale fusion feature map; the dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module; The multi-scale fused feature map is processed by a geologically prior-guided attention network to generate an attention-weighted feature map. The geologically prior-guided attention network fuses channel attention weights, spatial attention weights, and geological structural parameter constraint factors. The attention-weighted feature map is processed based on a preset encoder-decoder structure to generate an initial segmentation prediction map. The difference between the initial segmentation prediction map and the true label is calculated through an adaptive loss function module to generate optimized gradient data. The model parameters are iteratively updated based on the optimized gradient data to generate a final semantic segmentation model. Based on the final semantic segmentation model, a segmentation operation is performed on the target weak interlayer image to generate an interlayer segmentation result map.
[0016] The embodiments of this application provide a weak mezzanine semantic segmentation method, device, and medium based on deep learning, which at least includes the following technical effects: By using the adaptive dilated convolution mechanism of the dynamic feature perception architecture, the receptive field of the convolution kernel is dynamically adjusted according to the local features of the input image, which effectively enhances the ability to capture the edges of slender and curved structures. Combined with the shallow detail and deep semantic fusion strategy of the cross-scale feature interaction module, the accuracy of interlayer recognition in low-contrast areas is significantly improved, and the phenomenon of missed segmentation is reduced.
[0017] The geological prior-guided attention network integrates channel attention weights, spatial attention weights, and geological structural parameter constraint factors. By guiding the feature weighting process through interlayer curvature and extension direction parameters, it enhances the model's semantic understanding of complex geological structures and maintains stable segmentation robustness even in data-scarce scenarios.
[0018] The adaptive loss function module combines weighted Dice-FocalLoss and edge-aware loss, dynamically allocating loss weights based on the interlayer density parameter to optimize the segmentation effect of small targets. At the same time, an edge detection constraint mechanism is introduced to enhance the segmentation accuracy of the interlayer boundary and significantly reduce edge blurring error.
[0019] The lightweight deployment solution compresses the model size through knowledge distillation technology and combines parameter quantization and hardware acceleration strategies to achieve high frame rate inference on embedded devices. While ensuring segmentation accuracy, it effectively balances computational resource consumption and real-time requirements, meeting the engineering application requirements of geological exploration field monitoring. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart of a weak mezzanine semantic segmentation method based on deep learning is provided for embodiments of this application; Figure 2 A schematic diagram of a dynamic feature-aware architecture in a weak mezzanine semantic segmentation method based on deep learning, provided for an embodiment of this application; Figure 3 This is a schematic diagram of the internal structure of a weak mezzanine semantic segmentation device based on deep learning, provided as an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] This application provides a deep learning-based method, device, and medium for weak sandwich semantic segmentation, which aims to solve the following technical problem: improving the segmentation accuracy and generalization ability of low-contrast weak sandwich layers.
[0023] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0024] Figure 1 This document provides a flowchart for a weak mezzanine semantic segmentation process based on deep learning, as illustrated in an embodiment of this application. Figure 1 As shown in the figure, the weak mezzanine semantic segmentation method based on deep learning provided in this application embodiment specifically includes the following steps: Step 1: Obtain weak interlayer image data, which includes geological exploration images, engineering inspection images, and composite data.
[0025] Image data of weak interlayers refers to raw images containing geologically weak zones (such as rock fissures and faults), including geological exploration images (geological structure images taken by satellite or aerial photography), engineering inspection images (engineering structure images collected on-site), and synthetic data (simulated images generated by simulation software). This data is used to train and validate semantic segmentation models to improve the ability to identify low-contrast interlayers. Geological exploration image acquisition: Download geological images of a mining area from a satellite platform (such as Sentinel-2), covering areas with faults and rock strata. These images have a resolution of at least 10 meters per pixel to ensure the capture of slender interlayer structures.
[0026] Engineering inspection image acquisition: High-resolution images (resolution ≥ 1920×1080) were collected at a tunnel construction site B using drone equipment, with a focus on capturing images of the slope or rock surface to cover cracks and muddy interlayers.
[0027] Synthetic data generation: Simulated images are created using geological simulation software (such as Petrel) to simulate weak interlayers with different curvatures and extension directions. Simulation parameters include interlayer density (0.1~0.5 pixels / unit area) and texture complexity to supplement insufficient real data. All image data is stored uniformly in RGB format, ensuring the absence of any personally sensitive information.
[0028] Step 2: Preprocess the weak interlayer image data to generate normalized image data and enhanced image data.
[0029] Preprocessing includes normalization (standardizing the range of pixel values) and enhancement (improving data diversity through transformations), generating normalized image data (pixel values uniformly set to the range [0, 1]) and enhanced image data (images after applying geometric or color transformations). These operations aim to improve the model's generalization ability to low-contrast scenes.
[0030] Normalization: A linear mapping is performed on the original image, scaling pixel values from the original range to the [0, 1] interval. The specific formula is: Pixel value = (Original value - Minimum value) / (Maximum value - Minimum value). This eliminates the effects of lighting differences and ensures input consistency.
[0031] Data augmentation: Applying random transformations to generate enhanced image data, including: Geometric enhancement: Random rotation (angle range ±30°) and scaling (0.8~1.2x) to simulate different shooting angles.
[0032] Color Enhancement: Adjust brightness (offset ±0.2) and contrast (ratio ±0.3) to enhance the visibility of low-contrast interlayers.
[0033] The enhanced data expands the dataset size and reduces the risk of overfitting.
[0034] In a specific example, image preprocessing includes: 1) Image normalization The pixel values of the input image are normalized to the range [0, 1], using the following formula: Where I is the initial image, This is the normalized image.
[0035] 2) Data Augmentation Geometric transformations: random rotation (±30°), translation (±5%), scaling (0.8~1.2 times); Color transformation: Adjust brightness (±0.2), contrast (±0.3), saturation (±0.3); Synthetic data generation: Geological structure simulation software is used to generate interlayer distribution maps, which are then fused with real images to enhance data diversity.
[0036] 3) Data partitioning The dataset was divided into training, validation, and test sets in a 7:2:1 ratio to ensure the model's generalization ability.
[0037] Step 2.1: Extract multi-level feature maps from the input image using the encoder to generate the first-level feature map, the second-level feature map, the third-level feature map, and the fourth-level feature map.
[0038] The encoder is a convolutional neural network (such as ResNet-101) used to extract multi-level feature maps from the input image; the higher the level, the richer the semantic information. The first-level feature map ( Preserve detailed edges, second-level feature map ( ) Fusing local textures, third-level feature maps ( Extract structural features, fourth-level feature map ( It contains high-level semantics.
[0039] Input the normalized image into the pre-trained ResNet-101 encoder.
[0040] For example, output feature maps at four levels: : Resolution 56×56, number of channels 256, captures high-frequency details such as the edges of sandwich structures.
[0041] : Resolution 28×28, number of channels 512, integrates local texture information.
[0042] : Resolution 14×14, number of channels 1024, extracting sandwich structure features.
[0043] : Resolution 7×7, number of channels 2048, encoding high-level semantics such as the overall distribution of the mezzanine.
[0044] Through hierarchical progression, the model takes into account both the details of slender structures and global semantics.
[0045] Step 2.2: The fourth-level feature map is processed using an adaptive dilated convolution mechanism to generate dynamic dilation rate parameters and multi-branch output features.
[0046] The Adaptive Dilated Convolution (AAC) mechanism dynamically adjusts the dilation rate (kernel spacing) based on local image features to optimize the receptive field size. The dynamic dilation rate parameter is generated by a learnable matrix, and the multi-branch output features are a set of feature maps after convolution with different dilation rates, enhancing adaptability to changes in the curvature of the interlayer.
[0047] Step 2.2.1: Perform global average pooling on the input feature map to generate local feature vectors.
[0048] Perform global average pooling on the input feature map: The feature map (7×7×2048) is subjected to Global Average Pooling (GAP) to compress the spatial dimensions and generate a local feature vector (1×1×2048). This captures the global context of the feature map and avoids ignoring thin mezzanines.
[0049] Step 2.2.2: Input the local feature vector into the learnable parameter matrix, and generate a dynamic hole rate parameter through Softmax normalization.
[0050] Input the local feature vectors into the learnable parameter matrix: Input the local feature vectors into the learnable parameter matrix (size 2048×K, where K is the number of branches), normalize it using the Softmax function, and output the dynamic hole rate parameter ri (e.g., ri∈[1,6]). Softmax ensures that the sum of the parameters is 1, making the hole rate adapt to the input features.
[0051] Step 2.2.3: Adjust the receptive field of the convolution kernel according to the dynamic dilatation rate parameter, perform multi-branch dilated convolution operation on the input feature map, and generate multi-branch output features.
[0052] Adjusting the receptive field of the convolution kernel based on the dynamic porosity parameter: The feature map is processed using a multi-branch dilated convolution operation, with each branch using an independent dilation rate ri (e.g., branch 1: r=2, branch 2: r=4) to generate multi-branch output features (size 7×7×512). For example, in regions with high interlayer curvature, ri is automatically increased to expand the receptive field and improve the ability to capture slender structures.
[0053] Step 2.3: The first-level feature map, the second-level feature map, the third-level feature map and the multi-branch output features are fused through the cross-scale feature interaction module to generate a multi-scale fused feature map.
[0054] The Cross-Scale Feature Interaction (CSFI) module fuses feature maps from different levels (including multi-branch output features) to generate a multi-scale fused feature map. This module integrates shallow details and deep semantics through upsampling / downsampling and concatenation operations, solving the problem of low-level feature loss.
[0055] Upsampling high-level features: Feature maps upsampled to 28×28 resolution. The multi-branch output feature map is upsampled to 28×28 resolution, and bilinear interpolation is used to maintain smoothness.
[0056] Downsampling low-level features: The feature map is downsampled to 28×28 resolution, and average pooling is used to reduce noise.
[0057] Stitching all feature maps: Stitching along the channel dimension (After downsampling) , (After upsampling) and multi-branch output features (after upsampling).
[0058] Perform 1×1 convolutional fusion: Apply convolutional kernels to fuse and splice features, generating a multi-scale fused feature map (size 28×28×1024). This preserves the details of the mezzanine edges while enhancing semantic consistency.
[0059] Step 2.4: Introduce residual connections in the decoder to process the fourth-level feature map and the multi-scale fused feature map to generate a high-resolution feature map.
[0060] Residual connections directly transmit high-level features in the decoder, avoiding information loss. High-resolution feature maps are outputs generated by upsampling and fusing features, with a resolution similar to the input image (e.g., 56×56).
[0061] Will The feature map is upsampled to 28×28 resolution.
[0062] Perform residual summation: add the upsampled values... The feature is added to the multi-scale fused feature map, and the formula is expressed as: Output Feature = ↑+ Fusing feature maps.
[0063] Upsampling via decoder: Use transposed convolution to upsample the feature map to the original resolution (e.g., 224×224) to generate a high-resolution feature map, ensuring clear segmentation boundaries of the mezzanine.
[0064] Step 3: Construct a dynamic feature-aware architecture to process the normalized and enhanced image data, generating a multi-scale fused feature map. The dynamic feature-aware architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module.
[0065] The Dynamic Feature Aware Architecture (DFA) is a neural network structure for enhancing image feature extraction, comprising an Adaptive Dilated Convolution (AAC) mechanism and a Cross-Scale Feature Interaction (CSFI) module. It processes normalized image data (pixel values in the range [0, 1]) and enhanced image data (images after geometric / color transformations), generating multi-scale fused feature maps (feature maps that fuse details and semantics at different levels). This architecture improves the ability to identify weak interlayers (such as low-contrast cracks) by dynamically adjusting the receptive field and cross-scale fusion.
[0066] Applications of Adaptive Dilated Convolution (AAC): Input normalized image data or enhanced image data (resolution 224×224).
[0067] The AAC module dynamically adjusts the dilation rate (kernel spacing) of convolutional kernels. For example, on the input feature map, the dilation rate is automatically increased or decreased based on the local interlayer curvature: when the curvature is large, the dilation rate is increased (e.g., r=6) to capture slender structures, and when the curvature is small, it is decreased (e.g., r=2) to preserve details.
[0068] Output multi-scale feature branches, each branch corresponding to a feature map of a different receptive field (size 7×7×512), to ensure that the model adapts to changes in the morphology of the sandwich layer.
[0069] Cross-Scale Feature Interaction Module (CSFI) Fusion: The multi-scale feature branches of the AAC output are combined with the low-level features of the encoder (such as...). Integration.
[0070] The resolution is unified to 28×28 by upsampling (high-level features) and downsampling (low-level features), and then all features are stitched together.
[0071] A multi-scale fused feature map (size 28×28×1024) is generated by applying 1×1 convolutional fusion to stitch together the features. This feature map integrates shallow edge details (such as crack texture) and deep semantic information (such as interlayer distribution), solving the problem of omission in low-contrast regions by traditional methods.
[0072] In a specific example, such as Figure 2 As shown: The dynamic feature-aware architecture employs a combination of Adaptive Atrous Convolution (AAC) and Cross-Scale Feature Interaction (CSFI). The AAC module calculates the hole rate parameter using local features. The formula is ,in For learnable parameters, GAP is the global average pooling operation, and F is the input feature map. The CSFI module fuses shallow detail features with deep semantic features through upsampling / downsampling and concatenation operations.
[0073] AAC: Local feature extraction: on the input feature map Local features are extracted using global average pooling (GAP). .
[0074] Void rate generation: Input learnable parameters (K is the number of output channels), the hole rate parameter is generated by Softmax normalization. This enables the model to dynamically adjust its receptive field based on local features of the input image, enhancing its sensitivity to weak interlayer edges.
[0075] Dynamic convolution operation: Apply a dilation rate of 1 / 2 to each branch of the input feature map F. The convolutional kernel extracts multi-scale features, improving the ability to capture slender structures (such as fault lines) and outputs features. The formula is .
[0076] Parameter optimization: Automatic adjustment via backpropagation This allows the model to adaptively adjust the void ratio during training to adapt to changes in the curvature of the interlayer.
[0077] The AAC module calculates the void ratio parameter through local features, avoiding the insufficient adaptability of a fixed void ratio to changes in interlayer curvature. For example, in areas with large interlayer curvature, the void ratio automatically increases, enhancing the ability to detect slender structures.
[0078] CSFI: Feature pyramid construction: 4 levels in the ResNet-101 encoder Extract the output feature maps respectively .
[0079] Multi-scale fusion strategy: combining high-level features ( and Upsampling to Resolution, low-level features downsampling to Resolution, and After splicing, the data is fused using a 1×1 convolution formula. By fusing multi-scale features, edge information of slender structures is preserved, while semantic understanding of complex geological structures is enhanced.
[0080] Residual connections: Introducing residual blocks into the decoder to incorporate high-level features from the encoder. Upsampling is directly passed to the decoder to avoid information loss and improve the model's ability to model long-distance dependencies of weak sandwich structures. (Formula) The final output is then generated through pointwise convolution and upsampling.
[0081] Step 3.1: Extract the channel features of the multi-scale fusion feature map and generate channel attention weights through a fully connected layer.
[0082] Channel features are global statistics of the feature map (such as the average activation value of each channel), while channel attention weights are weight vectors (weight values for each channel) generated by the fully connected layer, used to emphasize key channels (such as feature channels related to the mezzanine).
[0083] Channel feature extraction: Input a multi-scale fused feature map (size 28×28×1024).
[0084] Perform a Global Average Pooling (GAP) operation: Calculate the average value for each channel across the spatial dimension (28×28), generating a channel feature vector (size 1×1×1024). This compresses redundant information while preserving the global context associated with the weak mezzanine.
[0085] Channel attention weight generation: The channel feature vectors are input into a fully connected layer (FC), which contains a learnable parameter matrix (size 1024×1024).
[0086] The output is normalized using the Sigmoid function to generate channel attention weight vectors (size 1×1×1024). The weight values are in the range [0, 1], with higher weights indicating that the channel is more critical for mezzanine recognition (such as texture or edge channels). For example, in areas with dense cracks, the weights automatically enhance the relevant channels.
[0087] This process enhances the model's sensitivity to key features and improves its robustness to low-contrast sandwiches.
[0088] Step 3.2: Extract spatial features from the channel-weighted feature map and generate spatial attention weights through a convolutional layer.
[0089] The channel-weighted feature map is a feature map (size 28×28×1024) after applying channel attention weights. Spatial features are the spatial distribution information of the feature map (such as location-related activation values). Spatial attention weights are weight matrices (weight values for each pixel) generated by convolutional layers, used to emphasize key spatial regions (such as mezzanine edges).
[0090] Channel-weighted feature map generation: The channel attention weight vector is multiplied channel by channel with the multi-scale fused feature map to generate a channel-weighted feature map (size 28×28×1024). This highlights important channels and weakens noisy channels (such as background regions).
[0091] Spatial feature extraction: Spatial pooling is performed on the channel-weighted feature map: max pooling is used to extract the maximum activation value at each location, and average pooling is used to extract the average activation value, generating two spatial feature matrices (size 28×28). This captures the spatial distribution pattern of the mezzanine (such as the crack extension direction).
[0092] Spatial attention weight generation: The spatial feature matrix is concatenated from the outputs of max pooling and average pooling.
[0093] Input the convolutional layer (kernel size 7×7) to generate the initial spatial weight matrix.
[0094] The spatial attention weight matrix (28×28) is output after normalization using the Sigmoid function. The weight values are in the range [0, 1], with high-weight regions corresponding to high-probability locations within the interlayer (such as the center line of a crack). For example, in the edge region of the interlayer, the weight value is close to 1, enhancing the segmentation accuracy.
[0095] Step 3.3: Generate interlayer geometric parameter maps using a pre-trained geological feature extractor and convert them into geological constraint factors.
[0096] The geological feature extractor is a pre-trained graph neural network (GNN), the interlayer geometry parameter map is a feature map containing the geometric attributes of the interlayer (such as curvature and extension direction), and the geological constraint factor is a weight factor generated by normalization, which is used to incorporate geological prior knowledge (such as interlayer morphology) into the attention mechanism.
[0097] Step 3.3.1: Extract the interlayer curvature parameters and extension direction parameters from the input image to generate the geological structure parameter tensor.
[0098] The interlayer curvature parameter describes the degree of bending of the interlayer (the larger the value, the more obvious the bending), the extension direction parameter describes the strike of the interlayer (such as the angle value), and the geological structure parameter tensor is a feature map that integrates these parameters (size 28×28×D, where D is the parameter dimension).
[0099] Input the original image or a preprocessed image (resolution 224×224).
[0100] Using a pre-trained geological feature extractor (such as a GNN-based model): extract interlayer curvature parameters (calculate local curvature values, ranging from 0 to 1) and extension direction parameters (calculate interlayer angles, ranging from 0° to 360°).
[0101] Output a geological structural parameter tensor (size 28×28×2), where the first channel is the curvature parameter and the second channel is the extension direction parameter. For example, in a curved fracture region, the curvature value is greater than 0.5.
[0102] Step 3.3.2: Perform a convolution operation on the geological structural parameter tensor to generate initial constraint factors.
[0103] The initial constraint factor is an intermediate weight matrix generated through convolution operations, used for the initial fusion of geological parameters.
[0104] Input the geological structure parameter tensor (size 28×28×2).
[0105] A convolutional layer (3×3 kernel size) is applied to generate an initial constraint factor matrix (28×28). The convolution operation integrates curvature and orientation parameters to generate spatially relevant weight values (with no limit on the range of values). For example, weight values are higher in regions with consistent extension directions.
[0106] Step 3.3.3: Normalize the initial constraint factors using the Sigmoid function to generate geological constraint factors.
[0107] The geological constraint factor is the normalized final weight matrix (size 28×28) with values in the range [0, 1], used to constrain the attention weights.
[0108] Input the initial constraint factor matrix (size 28×28).
[0109] Normalization is performed using the Sigmoid function, outputting a geological constraint factor matrix (size 28×28). Geological priors are reinforced in regions with high factor values (e.g., factor values are close to 1 when curvature is large), ensuring that the model generalizes in data-scarce scenarios.
[0110] Step 3.4: Integrate the spatial attention weights and geological constraint factors to generate an attention-weighted feature map.
[0111] Attention-weighted feature maps are feature maps with fused weights, which combine spatial attention weights (emphasizing key locations) and geological constraint factors (incorporating geological priors) to improve the identification of complex interlayers.
[0112] Weight fusion: Input spatial attention weight matrix (size 28×28) and geological constraint factor matrix (size 28×28).
[0113] Perform element-wise multiplication: generate a fused weight matrix (size 28×28). This operation ensures that spatial attention is constrained by geological parameters (e.g., weights are increased only in regions with consistent extension directions).
[0114] Feature weighting: Multiply the fused weight matrix positionally with the channel weighted feature map (size 28×28×1024) to generate an attention-weighted feature map (size 28×28×1024).
[0115] The output retains key features (such as crack edges) while suppressing irrelevant regions (such as background noise). For example, in low-contrast interlayer regions, the weighted feature values are significantly enhanced.
[0116] Step 4: Process the multi-scale fused feature map using a geologically prior-guided attention network to generate an attention-weighted feature map. The geologically prior-guided attention network integrates channel attention weights, spatial attention weights, and geological structural parameter constraint factors.
[0117] The Geological Prior-Guided Attention Network (MAF+GFGA) is a multimodal framework that integrates channel attention weights, spatial attention weights, and geological structural parameter constraints to generate a final attention-weighted feature map. This network leverages geological knowledge to guide attention, improving the model's accuracy in data-scarce scenarios.
[0118] The network structure should be: Input a multi-scale fused feature map (size 28×28×1024).
[0119] The complete process of executing steps 3.1-3.4 is as follows: first, generate channel attention weights, then generate spatial attention weights, then generate geological constraint factors, and finally fuse and output attention-weighted feature maps.
[0120] Geological a priori integration: During the fusion phase, the geological constraint factor (generated from step 3.3) dynamically adjusts the spatial attention weights (generated from step 3.2) to ensure that the attention distribution conforms to the geological characteristics of the interlayer (e.g., the weights are enhanced when the curvature is large).
[0121] Output attention-weighted feature map (size 28×28×1024), which enhances interlayer-related areas (such as crack boundaries) and weakens interference areas.
[0122] Step 5: Process the attention-weighted feature map based on the preset encoder-decoder structure to generate an initial segmentation prediction map.
[0123] The preset encoder-decoder structure is a neural network framework (such as the U-Net architecture), and the initial segmentation prediction map is the initial output of the model (size 224×224), which includes pixel-level prediction probabilities of the mezzanine (values in the range [0, 1].
[0124] Decoder structure processing: Input attention-weighted feature map (size 28×28×1024).
[0125] The feature map is upsampled stepwise using transposed convolutional layers: first upsampled to 56×56 resolution, then to 112×112, and finally to 224×224 resolution (the same as the input image).
[0126] Introducing residual connections into the decoder: This involves integrating high-level features from the encoder (such as...) After upsampling, it is added to the current feature to avoid information loss (formula: output feature = upsampled feature + decoded feature).
[0127] Prediction graph generation: Apply a 1×1 convolutional layer to output segmentation prediction: Generate an initial segmentation prediction map (size 224×224). Each pixel value represents the probability of the presence of the interlayer (0 for background, 1 for interlayer).
[0128] This output is used for subsequent loss calculation and optimization (step 6). For example, in the mezzanine edge region, the predicted value is close to 1 to ensure clear boundaries.
[0129] In a specific case, a multimodal attention framework (MAF) is constructed, including channel attention (SE Block) and spatial attention (CBAM), and a constraint factor γ is generated through geological structural parameters (such as interlayer curvature and extension direction).
[0130] MAF: Channel attention: Channel weight generation: for feature maps Perform global average pooling to obtain channel features. Channel weights are generated through a fully connected layer. .
[0131] Feature weighting: Multiplying α by F enhances the features of key channels. .
[0132] Spatial attention: Spatial feature extraction: for Max pooling and average pooling are performed to obtain spatial features. .
[0133] Spatial weight generation: Spatial weights are generated through convolutional layers. .
[0134] Feature weighting: combining β with Multiplication enhances key spatial regions. .
[0135] Geological structural parameter generation and constraint (GFGA): Geological feature extraction: Generate interlayer geometric parameters using a pre-trained geological feature extractor (GNN). , where D is the parameter dimension (such as curvature, extension direction).
[0136] Constraint factor calculation: Input G into the convolutional layer to generate constraint factors. By using γ to constrain the attention weight distribution, the ability to identify complex geological structures is improved.
[0137] Final attention mask: This mask represents the result of traditional multimodal attention. Multiply by the constraint factor γ to generate the final feature. .
[0138] By introducing parameters related to the curvature and extension direction of the interlayer, the distribution of attention weights is constrained, thereby improving the ability to recognize low-contrast regions. For example, in regions where the extension direction of the interlayer coincides with the main direction of the image, the attention weights are significantly enhanced.
[0139] Step 6: Calculate the difference between the initial segmentation prediction map and the true label using the adaptive loss function module to generate optimized gradient data.
[0140] The adaptive loss function module is a neural network component used to quantify the error between the model prediction (initial segmentation prediction map) and the true labels (manually labeled mezzanine regions), and to generate optimized gradient data (gradient vectors) for backpropagation to update model parameters. This module addresses class imbalance (e.g., small pixel proportion in the mezzanine region) and boundary blurring issues in weak mezzanine segmentation through dynamic weight allocation and edge enhancement.
[0141] Input data preparation: Initial segmentation prediction map: The prediction probability map (size 224×224) generated from step 5, where each pixel value represents the probability of the mezzanine being present (0 for background, 1 for mezzanine).
[0142] Real labels: Manually annotated binary mask (size 224×224), weak interlayer areas (such as cracks or faults) marked by geological experts.
[0143] These inputs ensure that error calculations are performed at the pixel level.
[0144] Step 6.1: Calculate the Dice loss and Focal loss values of the initial segmentation prediction map and the ground truth labels.
[0145] Dice loss measures the overlap between the prediction and the true label (suitable for class imbalance scenarios), while Focal loss enhances the penalty for hard samples (such as small object sandwiches) to prevent the model from ignoring the minority class.
[0146] Dice loss calculation: Based on pixel-level comparison, the formula is: .
[0147] in To predict probabilities, For real labels, For pixel index.
[0148] This emphasizes the overlap of the interlayer regions, alleviating the problem of background pixel dominance (e.g., in sparse interlayer regions, the loss value is higher to reinforce learning).
[0149] Focal loss calculation: To alleviate class imbalance and enhance optimization for small objectives (such as mezzanines), the formula is: , where α is the class weight, set to 0.25 (to balance positive and negative samples), and γ is the focus factor, usually set to 2 (to enhance the penalty for hard samples).
[0150] For example, when predicting probabilities When the value is close to 0 but the true label is 1 (indicating that the model is erroneously ignoring the mezzanine), the loss value increases significantly, improving the sensitivity to small targets.
[0151] This step outputs two loss values, which are used for subsequent dynamic weight allocation.
[0152] Step 6.2: Dynamically allocate Dice loss weights and Focal loss weights according to the interlayer density parameter to generate a weighted joint loss value.
[0153] The mezzanine density parameter represents the density of mezzanine pixels per unit area (a higher value indicates a denser mezzanine). Weights are dynamically assigned using the Sigmoid function to ensure the model adapts to different data regions (e.g., prioritizing Dice loss for high density and Focal loss for low density). The weighted joint loss is the total loss after merging the weights, optimizing training stability.
[0154] Step 6.2.1: Count the number of interlayer pixels per unit area and generate interlayer density parameters.
[0155] The interlayer density parameter is a scalar value that quantifies the interlayer distribution density. It is calculated by dividing the total number of interlayer pixels in the real label by the total area of the image.
[0156] Enter a real label image (size 224×224).
[0157] Count the number of pixels in the mezzanine (the number of pixels with a true label value of 1).
[0158] Calculate the density parameter ρ: ρ = (number of pixels in the interlayer) / (224 × 224). For example, in sparse interlayer regions, ρ may be less than 0.1; in dense regions, ρ may be greater than 0.3.
[0159] Step 6.2.2: Input the interlayer density parameter into the Sigmoid function to generate the dynamic weights of the Dice loss.
[0160] The dynamic weights of the Dice loss are weight values (ranging from 0 to 1) output by the Sigmoid function, used to adjust the contribution ratio of the Dice loss to the total loss. The Sigmoid function ensures a smooth transition of weights.
[0161] Input the interlayer density parameter ρ.
[0162] Step 6.2.3: Calculate the dynamic weight of Focal loss based on the dynamic weight of Dice loss.
[0163] Focal loss dynamic weights are complementary weights, calculated as 1 minus the Dice loss weights, ensuring that the total loss weights sum to 1.
[0164] For example, in low-density regions (ρ low), wFocal is close to 1, and Focal loss is enhanced to handle hard samples.
[0165] After outputting the dynamic weights, a weighted joint loss value is generated: Dynamic Weights and It is calculated using the interlayer density ρ (the number of interlayers per unit area), using the following formula: , where k is a coefficient that controls the gradient, usually set to 5 (to adjust the smoothness of weight changes). This is the interlayer density threshold, used for dynamic weight calculation; This improves the model's adaptability to regions with different densities.
[0166] Step 6.3: Extract the edge mask of the real label using the edge detection algorithm and calculate the edge perception loss value.
[0167] Edge-aware loss enhances the segmentation accuracy of the mezzanine boundary. A binary edge mask (containing only edge pixels) is generated by Canny edge detection, and cross-entropy loss is calculated.
[0168] Step 6.3.1: Perform Canny edge detection on the real label to generate a binary edge mask.
[0169] A binary edge mask is an image (224×224 pixels) that identifies the boundaries of the mezzanine, with edge pixels having a value of 1 and non-edge pixels having a value of 0. The Canny algorithm uses high and low thresholds to detect edges.
[0170] Input the actual label image.
[0171] Canny edge detection is applied: first, Gaussian filtering is used to remove noise, then the gradient magnitude and direction are calculated, and finally a binary mask is generated using dual thresholds (low threshold 0.1, high threshold 0.3).
[0172] Output a mask image, for example, setting the pixel value to 1 at the crack boundary.
[0173] Step 6.3.2: Extract the pixel prediction values of the edge regions in the initial segmentation prediction map to generate an edge prediction subset.
[0174] The edge prediction subset is the set of pixel prediction values corresponding to the edge mask positions in the initial segmentation prediction map.
[0175] Input the initial segmentation prediction map and the binary edge mask.
[0176] Traverse the mask image, and when the mask value is 1, extract the predicted value pi at the same position in the prediction image.
[0177] Generate a subset of data, for example, containing only the predicted probabilities of boundary pixels.
[0178] Step 6.3.3: Calculate the cross-entropy loss between the edge prediction subset and the real edge mask to generate the edge-aware loss value.
[0179] Edge extraction: Canny edge detection is performed on the ground truth label Y to enhance the segmentation accuracy of the sandwich edges and generate an edge mask. .
[0180] Edge loss calculation: Add edge weights to the cross-entropy loss. The formula is ,in Labels and predicted values for edge regions, This is the edge weighting coefficient, usually set to 0.7 (to balance edge and overall segmentation accuracy). This represents the true label and predicted value for the entire region.
[0181] Step 6.4: Combine the weighted joint loss value and the edge-aware loss value to generate optimized gradient data.
[0182] The optimized gradient data is the gradient vector of the loss function, used for backpropagation to update the model weights. Fusion is achieved through a weighted sum.
[0183] Input the weighted joint loss value and the edge-aware loss value.
[0184] Calculate the total loss: Total loss = weighted joint loss + edge sensing loss.
[0185] Optimized gradient data is generated by calculating gradients through automatic differentiation (such as the backward function in PyTorch or TensorFlow).
[0186] For example, gradient data includes the partial derivatives of each model parameter, which guides the direction of parameter updates.
[0187] Step 7: Iteratively update the model parameters based on the optimized gradient data to generate the final semantic segmentation model.
[0188] Iterative updates involve repeatedly applying gradient data to adjust model weights using an optimizer (such as AdamW); the final semantic segmentation model is a trained network that can accurately segment weak segments.
[0189] Optimizer configuration: Using the AdamW optimizer (combining adaptive learning rate and weight decay), the initial learning rate is set to... Weight decay coefficient .
[0190] Iteratively updated: Input the optimization gradient data (generated from step 6).
[0191] Perform backpropagation: Calculate the gradients of model parameters (such as convolution kernel weights).
[0192] Applying gradient descent: The formula for updating parameters is parameter = parameter - learning rate × gradient.
[0193] Learning rate scheduling is used: a cosine annealing strategy is employed, reducing the learning rate by 10% every 5 epochs to avoid local optima.
[0194] Training terminated: Set up an early stopping mechanism to monitor the loss on the validation set (e.g., stop if there is no improvement after 10 consecutive epochs).
[0195] After training for 100 epochs, output the final semantic segmentation model (save as a .pt or .h5 file).
[0196] This ensures that the model generalizes under data scarcity, for example, with improved mIoU on low-contrast images.
[0197] Step 8: Perform segmentation operation on the target weak interlayer image based on the final semantic segmentation model to generate an interlayer segmentation result image.
[0198] The target weak interlayer image is a new image to be segmented (such as a geological exploration or engineering inspection image); the segmentation operation is the model inference process; the interlayer segmentation result image is a binary output image (the same size as the input), which identifies the interlayer region.
[0199] Input preparation: Acquire the target image (such as satellite image of a mining area or tunnel site map) and perform the same preprocessing as in step 2 (normalize to [0, 1]).
[0200] Model Inference: Load the final semantic segmentation model.
[0201] Input image to model: processed through encoder-decoder structure, outputting a predicted probability map.
[0202] Result generation: Apply a threshold (usually 0.5) to the predicted probability map: pixels with a probability ≥ 0.5 are set to 1 (mezzanine), otherwise to 0 (background).
[0203] Generate a mezzanine segmentation result diagram (size 224×224), which can be directly visualized or used for risk assessment.
[0204] For example, it can be executed on embedded devices to support real-time applications such as engineering monitoring.
[0205] The above are embodiments of the method proposed in this application. Based on the same inventive concept, embodiments of this application also provide a weak mezzanine semantic segmentation device based on deep learning, the structure of which is as follows: Figure 2 As shown.
[0206] Figure 3 This is a schematic diagram of the internal structure of a weak mezzanine semantic segmentation device based on deep learning, provided as an embodiment of this application. Figure 3 As shown, the device includes: At least one processor 301; And a memory 302 that is communicatively connected to at least one processor; The memory 302 stores instructions executable by at least one processor, which are executed by at least one processor 301 to enable at least one processor 301 to: The process involves acquiring weak interlayer image data, including geological exploration images, engineering inspection images, and synthetic data; preprocessing the weak interlayer image data to generate normalized image data and enhanced image data; constructing a dynamic feature perception architecture to process the normalized image data and enhanced image data, generating a multi-scale fused feature map; the dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module; processing the multi-scale fused feature map through a geologically prior-guided attention network to generate an attention-weighted feature map; the geologically prior-guided attention network fuses channel attention weights, spatial attention weights, and geological structural parameter constraint factors; processing the attention-weighted feature map based on a preset encoder-decoder structure to generate an initial segmentation prediction map; calculating the difference between the initial segmentation prediction map and the true label through an adaptive loss function module to generate optimized gradient data; iteratively updating the model parameters based on the optimized gradient data to generate a final semantic segmentation model; and performing a segmentation operation on the target weak interlayer image based on the final semantic segmentation model to generate an interlayer segmentation result map.
[0207] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium based on deep learning-based weak mezzanine semantic segmentation, storing computer-executable instructions, wherein the computer-executable instructions are configured as follows: The process involves acquiring weak interlayer image data, including geological exploration images, engineering inspection images, and synthetic data; preprocessing the weak interlayer image data to generate normalized image data and enhanced image data; constructing a dynamic feature perception architecture to process the normalized image data and enhanced image data, generating a multi-scale fused feature map; the dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module; processing the multi-scale fused feature map through a geologically prior-guided attention network to generate an attention-weighted feature map; the geologically prior-guided attention network fuses channel attention weights, spatial attention weights, and geological structural parameter constraint factors; processing the attention-weighted feature map based on a preset encoder-decoder structure to generate an initial segmentation prediction map; calculating the difference between the initial segmentation prediction map and the true label through an adaptive loss function module to generate optimized gradient data; iteratively updating the model parameters based on the optimized gradient data to generate a final semantic segmentation model; and performing a segmentation operation on the target weak interlayer image based on the final semantic segmentation model to generate an interlayer segmentation result map.
[0208] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0209] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0210] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0211] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0212] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0213] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0214] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0215] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0216] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0217] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0218] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A weak sandwich semantic segmentation method based on deep learning, characterized in that, The method includes: Acquire weak interlayer image data, which includes geological exploration images, engineering inspection images, and composite data; The image data of the weak interlayer is preprocessed to generate normalized image data and enhanced image data; A dynamic feature perception architecture is constructed to process the normalized image data and the enhanced image data to generate a multi-scale fused feature map; the dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module. The multi-scale fused feature map is processed by a geological prior-guided attention network to generate an attention-weighted feature map; the geological prior-guided attention network fuses channel attention weights, spatial attention weights, and geological structural parameter constraint factors. The attention-weighted feature map is processed based on a preset encoder-decoder structure to generate an initial segmentation prediction map; The difference between the initial segmentation prediction map and the true label is calculated using the adaptive loss function module to generate optimized gradient data. Based on the optimized gradient data, the model parameters are iteratively updated to generate the final semantic segmentation model; Based on the final semantic segmentation model, a segmentation operation is performed on the target weak sandwich image to generate a sandwich segmentation result image. The weak interlayer image data is preprocessed to generate normalized image data and enhanced image data, specifically including: The encoder extracts multi-level feature maps from the input image, generating first-level feature maps, second-level feature maps, third-level feature maps, and fourth-level feature maps; An adaptive dilated convolution mechanism is used to process the fourth-level feature map to generate dynamic dilation rate parameters and multi-branch output features. By fusing the first-level feature map, the second-level feature map, the third-level feature map, and the multi-branch output features through the cross-scale feature interaction module, a multi-scale fused feature map is generated. In the decoder, residual connections are introduced to process the fourth-level feature map and the multi-scale fused feature map to generate a high-resolution feature map. The fourth-level feature map is processed using an adaptive dilated convolution mechanism to generate dynamic dilation rate parameters and multi-branch output features, specifically including: Perform global average pooling on the input feature map to generate local feature vectors; The local feature vectors are input into the learnable parameter matrix, and dynamic hole rate parameters are generated by Softmax normalization. The receptive field of the convolution kernel is adjusted according to the dynamic dilatation rate parameter, and a multi-branch dilated convolution operation is performed on the input feature map to generate multi-branch output features. The difference between the initial segmentation prediction map and the true label is calculated using an adaptive loss function module to generate optimized gradient data, specifically including: Calculate the Dice loss and Focal loss values between the initial segmentation prediction map and the ground truth labels; The weighted joint loss value is generated by dynamically allocating the Dice loss weight and Focal loss weight based on the interlayer density parameter. The edge mask of the real label is extracted using an edge detection algorithm, and the edge perception loss value is calculated. The weighted joint loss value and the edge-aware loss value are fused to generate optimized gradient data; The Dice loss weight and Focal loss weight are dynamically allocated based on the interlayer density parameter to generate a weighted joint loss value, specifically including: Count the number of interlayer pixels per unit area to generate interlayer density parameters; The interlayer density parameter is input into the Sigmoid function to generate dynamic weights for the Dice loss. Calculate the Focal loss dynamic weights based on the Dice loss dynamic weights.
2. The weak sandwich semantic segmentation method based on deep learning according to claim 1, characterized in that, A dynamic feature-aware architecture is constructed to process the normalized image data and enhanced image data, generating a multi-scale fusion feature map, specifically including: Extract the channel features of the multi-scale fused feature map and generate channel attention weights through a fully connected layer; Spatial features are extracted from the channel-weighted feature map, and spatial attention weights are generated through convolutional layers. A geometries of interlayer geometry are generated using a pre-trained geological feature extractor and then converted into geological constraint factors. The spatial attention weights and geological constraint factors are fused to generate an attention-weighted feature map.
3. The weak sandwich semantic segmentation method based on deep learning according to claim 1, characterized in that, A pre-trained geological feature extractor is used to generate interlayer geometric parameter maps, which are then converted into geological constraint factors, specifically including: Extract the interlayer curvature parameters and extension direction parameters from the input image to generate a geological structure parameter tensor; Perform a convolution operation on the geological structural parameter tensor to generate an initial constraint factor; The initial constraint factors are normalized using the Sigmoid function to generate geological constraint factors.
4. The weak sandwich semantic segmentation method based on deep learning according to claim 1, characterized in that, The edge mask of the real label is extracted using an edge detection algorithm, and the edge-aware loss value is calculated, specifically including: Perform Canny edge detection on the real label to generate a binary edge mask; Extract the pixel prediction values of the edge regions in the initial segmentation prediction map to generate an edge prediction subset; Calculate the cross-entropy loss between the edge prediction subset and the real edge mask to generate the edge-aware loss value.
5. A weak mezzanine semantic segmentation device based on deep learning, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Acquire weak interlayer image data, which includes geological exploration images, engineering inspection images, and composite data; The image data of the weak interlayer is preprocessed to generate normalized image data and enhanced image data; A dynamic feature perception architecture is constructed to process the normalized image data and the enhanced image data to generate a multi-scale fused feature map; the dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module. The multi-scale fused feature map is processed by a geological prior-guided attention network to generate an attention-weighted feature map; the geological prior-guided attention network fuses channel attention weights, spatial attention weights, and geological structural parameter constraint factors. The attention-weighted feature map is processed based on a preset encoder-decoder structure to generate an initial segmentation prediction map; The difference between the initial segmentation prediction map and the true label is calculated using the adaptive loss function module to generate optimized gradient data. Based on the optimized gradient data, the model parameters are iteratively updated to generate the final semantic segmentation model; Based on the final semantic segmentation model, a segmentation operation is performed on the target weak sandwich image to generate a sandwich segmentation result image. The weak interlayer image data is preprocessed to generate normalized image data and enhanced image data, specifically including: The encoder extracts multi-level feature maps from the input image, generating first-level feature maps, second-level feature maps, third-level feature maps, and fourth-level feature maps; An adaptive dilated convolution mechanism is used to process the fourth-level feature map to generate dynamic dilation rate parameters and multi-branch output features. By fusing the first-level feature map, the second-level feature map, the third-level feature map, and the multi-branch output features through the cross-scale feature interaction module, a multi-scale fused feature map is generated. In the decoder, residual connections are introduced to process the fourth-level feature map and the multi-scale fused feature map to generate a high-resolution feature map. The fourth-level feature map is processed using an adaptive dilated convolution mechanism to generate dynamic dilation rate parameters and multi-branch output features, specifically including: Perform global average pooling on the input feature map to generate local feature vectors; The local feature vectors are input into the learnable parameter matrix, and dynamic hole rate parameters are generated by Softmax normalization. The receptive field of the convolution kernel is adjusted according to the dynamic dilatation rate parameter, and a multi-branch dilated convolution operation is performed on the input feature map to generate multi-branch output features. The difference between the initial segmentation prediction map and the true label is calculated using an adaptive loss function module to generate optimized gradient data, specifically including: Calculate the Dice loss and Focal loss values between the initial segmentation prediction map and the ground truth labels; The weighted joint loss value is generated by dynamically allocating the Dice loss weight and Focal loss weight based on the interlayer density parameter. The edge mask of the real label is extracted using an edge detection algorithm, and the edge perception loss value is calculated. The weighted joint loss value and the edge-aware loss value are fused to generate optimized gradient data; The Dice loss weight and Focal loss weight are dynamically allocated based on the interlayer density parameter to generate a weighted joint loss value, specifically including: Count the number of interlayer pixels per unit area to generate interlayer density parameters; The interlayer density parameter is input into the Sigmoid function to generate dynamic weights for the Dice loss. Calculate the Focal loss dynamic weights based on the Dice loss dynamic weights.
6. A non-volatile computer storage medium for weak mezzanine semantic segmentation based on deep learning, storing computer-executable instructions, characterized in that, The computer-executable instructions are set as follows: Acquire weak interlayer image data, which includes geological exploration images, engineering inspection images, and composite data; The image data of the weak interlayer is preprocessed to generate normalized image data and enhanced image data; A dynamic feature perception architecture is constructed to process the normalized image data and the enhanced image data to generate a multi-scale fused feature map; the dynamic feature perception architecture includes an adaptive dilated convolution mechanism and a cross-scale feature interaction module. The multi-scale fused feature map is processed by a geological prior-guided attention network to generate an attention-weighted feature map; The geological prior-guided attention network integrates channel attention weights, spatial attention weights, and geological structural parameter constraint factors. The attention-weighted feature map is processed based on a preset encoder-decoder structure to generate an initial segmentation prediction map; The difference between the initial segmentation prediction map and the true label is calculated using the adaptive loss function module to generate optimized gradient data. Based on the optimized gradient data, the model parameters are iteratively updated to generate the final semantic segmentation model; Based on the final semantic segmentation model, a segmentation operation is performed on the target weak sandwich image to generate a sandwich segmentation result image. The weak interlayer image data is preprocessed to generate normalized image data and enhanced image data, specifically including: The encoder extracts multi-level feature maps from the input image, generating first-level feature maps, second-level feature maps, third-level feature maps, and fourth-level feature maps; An adaptive dilated convolution mechanism is used to process the fourth-level feature map to generate dynamic dilation rate parameters and multi-branch output features. By fusing the first-level feature map, the second-level feature map, the third-level feature map, and the multi-branch output features through the cross-scale feature interaction module, a multi-scale fused feature map is generated. In the decoder, residual connections are introduced to process the fourth-level feature map and the multi-scale fused feature map to generate a high-resolution feature map. The fourth-level feature map is processed using an adaptive dilated convolution mechanism to generate dynamic dilation rate parameters and multi-branch output features, specifically including: Perform global average pooling on the input feature map to generate local feature vectors; The local feature vectors are input into the learnable parameter matrix, and dynamic hole rate parameters are generated by Softmax normalization. The receptive field of the convolution kernel is adjusted according to the dynamic dilatation rate parameter, and a multi-branch dilated convolution operation is performed on the input feature map to generate multi-branch output features. The difference between the initial segmentation prediction map and the true label is calculated using an adaptive loss function module to generate optimized gradient data, specifically including: Calculate the Dice loss and Focal loss values between the initial segmentation prediction map and the ground truth labels; The weighted joint loss value is generated by dynamically allocating the Dice loss weight and Focal loss weight based on the interlayer density parameter. The edge mask of the real label is extracted using an edge detection algorithm, and the edge perception loss value is calculated. The weighted joint loss values are then fused together. The Dice loss weight and Focal loss weight are dynamically allocated based on the interlayer density parameter to generate a weighted joint loss value, specifically including: Count the number of interlayer pixels per unit area to generate interlayer density parameters; The interlayer density parameter is input into the Sigmoid function to generate dynamic weights for the Dice loss. Calculate the Focal loss dynamic weights based on the Dice loss dynamic weights.
Citation Information
Patent Citations
Image segmentation method based on multi-scale space adaptive hole convolution
CN115760687A
Camouflage object semantic segmentation method, device and equipment based on self-supervised dual construction model, and storage medium
CN120107584A