Double-branch multi-scale modeling semantic segmentation method based on tunnel scene
Through the dual-branch semantic segmentation model, combined with the residual network and spatial symmetric pyramid module, the problems of segmentation accuracy and computational complexity in tunnel image processing are solved, and efficient tunnel image segmentation effect is achieved.
Patent Information
- Application Number
- CN202510740596.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies find it difficult to strike a balance between segmentation accuracy, model parameter count, and inference speed in tunnel construction, resulting in reduced recognition accuracy in tunnel image processing, especially under complex backgrounds and uneven lighting conditions.
A dual-branch semantic segmentation model is adopted, combined with a residual network and a spatially symmetric pyramid module. Through a dynamic context-aware module and cross-layer feature fusion, feature information acquisition and fusion are enhanced, and computational complexity is reduced.
The accuracy and consistency of tunnel image segmentation are improved, the number of model parameters and computational overhead are reduced, and the method can adapt to complex backgrounds and uneven lighting conditions.
Smart Images

Figure CN120689612A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and specifically provides a double-branch multi-scale modeling semantic segmentation method based on tunnel scenes. Background Art
[0002] The continuous advancement of intelligent construction technology has placed higher demands on image processing for underground tunnel projects. During tunnel construction, complex geological conditions make it difficult to accurately predict the actual surrounding rock conditions and precisely determine the surrounding rock grade during construction, leading to increased construction difficulty and rising costs. By analyzing the geological characteristics of the tunnel face and combining them with surface geological survey data, the surrounding rock grade of the tunnel ahead can be effectively assessed, improving construction efficiency. Tunnel face geological information exhibits distinct multi-scale properties, and the rock masses at different faces exhibit significant differences in their apparent texture, grayscale distribution, and structural morphology, resulting in a high degree of inconsistency in the image geological features. These characteristics further complicate feature recognition for the model, resulting in reduced recognition accuracy, prone to target loss, inaccurate edge segmentation, and other phenomena. The effect is particularly pronounced in areas with blurred rock structure boundaries and weakened features.
[0003] Currently, mainstream approaches to achieving real-time semantic segmentation rely primarily on designing lightweight backbone networks and reducing model complexity by simplifying the decoder structure, thereby building an efficient segmentation framework. The core idea of these approaches is to leverage a concise network architecture to achieve a balance between speed and performance while maintaining a certain level of accuracy. However, overly streamlined network structures also present new challenges. In particular, the segmentation process struggles to effectively recover spatial detail information lost during downsampling, compromising the final segmentation accuracy.
[0004] In summary, the existing technology has the technical problem that the real-time scene image semantic segmentation network model is difficult to achieve a balance between segmentation accuracy, model parameter quantity and inference speed. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a dual-branch multi-scale modeling semantic segmentation method based on tunnel scenarios, which can effectively obtain feature information between different scales, improve the accuracy of segmentation, and reduce the number of model parameters and computational complexity.
[0006] In order to achieve the above object, the solution of the present invention is:
[0007] In step S1, images of the tunnel face under construction are collected to construct a pixel-level annotated dataset containing typical structural features. Sample diversity is enhanced through data enhancement, and standardized division and statistical analysis are completed.
[0008] In step S2, a context path of a dual-branch semantic segmentation model is established. A residual network is used to extract high-dimensional contextual feature information through fast downsampling. A spatially symmetric pyramid module is designed to enhance the perception of semantic features. Finally, the feature map obtained by upsampling the global average pooling output is fused with the high-dimensional contextual feature information after the spatial pyramid pooling module.
[0009] In step S3, the spatial path of the dual-branch semantic segmentation model is established, the high-dimensional feature map is downsampled, and the downsampled feature map is pooled. A dynamic context-aware enhancement module is added to the network to extract and fuse feature maps of different scales. The output feature map of the spatial branch is obtained and used as one of the inputs of the feature fusion module.
[0010] In step S4, the acquired context feature information is fused with the spatial feature information and finally input into the predicted image to achieve the semantic segmentation task.
[0011] Compared with the prior art, the present invention has the following beneficial effects:
[0012] The first improvement of this invention lies in the addition of a dynamic context-aware module to the spatial path of the dual-branch network, enabling the network to better handle complex backgrounds and uneven lighting conditions in tunnel face images. This module dynamically learns information about each pixel using multi-layer convolution operations, while simultaneously preserving local details and enhancing global semantic information. It also adjusts pixel features based on the current context, thereby improving the accuracy of segmentation results.
[0013] The second improvement of the present invention is the addition of a spatially symmetric pyramid module to the context path of the dual-branch network. This module uses dilated convolutions with different dilation rates to process input features, helping the model capture a wider range of information associations, improving the overall structural modeling capabilities, and thereby enhancing the accuracy and consistency of segmentation. To reduce the computational burden of the model, standard convolutions are replaced with symmetric convolutions, reducing computational overhead while maintaining expressiveness. At the same time, 1×1 group convolutions are introduced to group feature maps, further reducing the number of parameters.
[0014] The third improvement of this invention lies in the design of a cross-layer feature fusion module, which combines the detailed information of shallow features with the semantic information of deep features to build a comprehensive feature set that combines spatial resolution and semantic understanding capabilities. Features are extracted from each layer of the network. Upsampling is used to adjust the deep low-resolution features to the same spatial size as the shallow high-resolution features. Feature fusion is then achieved through channel splicing or point-by-point convolution. Finally, an attention mechanism is used to dynamically calibrate the weights of features at each layer. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 Schematic diagram of the structure of the dual-branch semantic segmentation algorithm used in the embodiment of the present invention.
[0016] Figure 2 Schematic diagram of a dynamic context awareness module used in an embodiment of the present invention.
[0017] Figure 3 Schematic diagram of a dynamic weight module used in an embodiment of the present invention.
[0018] Figure 4 Schematic diagram of a spatially symmetrical pyramid module used in an embodiment of the present invention.
[0019] Figure 5 Schematic diagram of various samples of data sets used in the embodiment of the present invention.
[0020] Figure 6 Schematic diagram of data amplification used in the embodiment of the present invention. DETAILED DESCRIPTION
[0021] To explain the technical solution of the present invention, the present invention is described in detail below with reference to specific embodiments and accompanying drawings.
[0022] like Figure 1 As shown in FIG, the present invention designs a dual-branch multi-scale modeling semantic segmentation method based on tunnel scenarios. The method includes the following steps:
[0023] In step S1, images of the tunnel face under construction are collected to construct a pixel-level annotated dataset containing typical structural features. Sample diversity is enhanced through data enhancement, and standardized division and statistical analysis are completed.
[0024] On-site construction personnel use technical means to implement quality control of tunnel excavation. They take photos and archive them every 10 meters along the excavation surface. Due to the inconsistent quality of the image resolution collected by different image acquisition hardware, normalization operation is performed to adjust the original image resolution to 448×448 pixels. Figure 5 shown.
[0025] The training process of the image semantic segmentation network model generally uses a rich data set to reduce the problem of overfitting during model training. In order to increase the data volume and diversify the types of data sets, mirror inversion, multi-angle rotation, noise introduction and brightness adjustment are used, such as Figure 6 shown.
[0026] In step S2, a context path of a dual-branch semantic segmentation model is established. A residual network is used to extract high-dimensional contextual feature information through fast downsampling. A spatially symmetric pyramid module is designed to enhance the perception of semantic features. Finally, the feature map obtained by global average pooling upsampling output is fused with the high-dimensional contextual feature information obtained by the spatially symmetric pyramid module.
[0027] When establishing the context path of the dual-branch semantic segmentation model, the ResNet18 residual network is used as the baseline network to extract high-dimensional contextual feature information through fast downsampling.
[0028] The network consists of 18 layers, including convolutional layers and fully connected layers.
[0029] The convolution operation layer is divided into 1 independent convolution module and 4 residual modules.
[0030] The residual module uses cross-layer connection technology to fuse the input features with the transformed output features. Each residual module consists of a 3×3 convolutional layer, a BN layer, and a ReLU.
[0031] The network outputs feature maps of different scales, and global average pooling operations are performed on the 1 / 16 feature map and the 1 / 32 feature map respectively.
[0032] A spatial symmetric pyramid module is established, and the output results after passing through the four residual modules are transmitted to the spatial symmetric pyramid module, which can effectively extract the edge and structural features of the region and improve the model's ability to recognize the target area.
[0033] In addition, using dilated convolutions with different expansion rates to process input features helps the model obtain information associations in a wider range.
[0034] The standard dilated convolution is replaced by symmetric dilated convolution, which reduces the computational overhead while maintaining the expressive power.
[0035] The input feature map is convolved with a 1×1 group to reduce the dimension and improve computational efficiency, generating two new feature representations.
[0036] In order to capture multi-scale features, symmetric convolution is used here through four convolution kernels with sizes of 3, 5, 7 and 9.
[0037] F l =Conv 1×1 (F)
[0038] Where: F is the input feature map, and F is obtained by 1×1 convolution l The feature map F l Input into the convolutional pyramid.
[0039]
[0040] F out =C 3×3 (C 3×3 (F i )⊕C 3×3 (F l ))
[0041] Where: SymConv 3×3 (F l ) is a 3×3 symmetric convolution, Cat represents the feature concatenation operation to obtain F i , where ⊕ represents element addition, and the final output feature F is obtained out .
[0042] In step S3, when establishing the spatial path of the dual-branch semantic segmentation model, the high-dimensional feature map is downsampled to reduce its size to 1 / 2 of the original size.
[0043] The pooling layer is applied to the downsampled feature map. At the same time, it is further downsampled to feature maps of 1 / 4 and 1 / 8 sizes.
[0044] A dynamic weight module is designed to perform weighted averaging of features at different scales, and dynamically change the size and weight of the convolution kernel according to the current image content to adapt to the feature expectations of different regions.
[0045] In order to capture the global dependency between pixels, the dot product of Query and Key is used to calculate the similarity between the two, and the result is normalized.
[0046]
[0047] Where: S ij Indicates the similarity between the i-th Query and the j-th Key;
[0048] Using the attention weight matrix A and the Value feature V, the features of all pixels are weighted and summed to obtain the context feature of each pixel.
[0049]
[0050] Where: C i Represents the contextual features of all pixels.
[0051] Each row corresponds to the global context representation of a pixel. In order to combine the context information with the original features, the context features are mapped back to the dimension of the input features through a linear projection and added element-by-element with the input features X to generate the final enhanced features.
[0052] A dynamic-aware context enhancement module is designed to perform preliminary context information extraction by applying 1×1 convolution to the input feature map.
[0053] The 3×3 convolution with hierarchical hole parameters enables the network to learn local features of different scales.
[0054] Adaptive average pooling is used in conjunction with the global attention mechanism to reduce spatial dimensions.
[0055] Then, a 1×1 convolution operation is implemented, and the normalization technique and activation function are combined to form a processing node. By introducing the interaction between local and global features, the final feature map is generated.
[0056] Define the local context feature F ld and F lw ,These features are processed by different convolution kernels on the input features.
[0057]
[0058]
[0059] Where: Cat(·) represents the channel splicing operation, Conv τ It is a set of 3×3 depth convolutions with hole rates r=1,2,3,4.
[0060] Next, by transforming the above-obtained feature F ld and F lw Fusion, generating local features F local .
[0061] F local =F ld +F lw
[0062] W=σ(F global )
[0063] Where: F local is a local feature, F global It is a global feature.
[0064] Here, the Sigmoid activation function is used to perform weighted operations on the obtained local features and global features so that their weight range is between [0, 1].
[0065] Specifically, the global feature F global Weighted superposition to local features F local Calculate the fusion weight W.
[0066]
[0067] Where: F final is the final output feature, Represents element-by-element multiplication operation, F l and F g are the adjusted local and global features respectively.
[0068] A cross-layer feature fusion module is established to combine the detailed information of shallow features with the semantic information of deep features, and to build a comprehensive feature with both spatial resolution and semantic understanding capabilities.
[0069] Upsampling is used to adjust the deep low-resolution features to the same spatial size as the shallow high-resolution features.
[0070] Then, channel splicing or point-by-point convolution is used to realize the feature fusion operation, and finally the weights of the features of each layer are dynamically calibrated by the attention mechanism.
[0071] F CLFF =Attention(Concat(F shallow ,Upsample(F deep )))
[0072] Where: F shallow and F deep They represent shallow and deep features respectively, Upsample(·) is the upsampling operation, Concat(·) is the channel concatenation operation, and Attention(·) is used to assign dynamic weights to the fused features.
[0073] The present invention introduces a weighted cross entropy loss function and adopts a method of setting weight coefficients for each category, so that the model can focus more on the target area with a small number of samples during training.
[0074]
[0075] Where: N is the total number of pixels in the image, C is the total number of categories, y i,c is the true label of pixel i. If the pixel belongs to category c, then y i,c =1, otherwise 0, p i,c is the probability predicted by the model that pixel i belongs to category c.
[0076]
[0077] Where: w c is the weight of category c.
[0078] The above embodiments are merely typical implementations of the present invention. Those skilled in the art may make appropriate adjustments or modifications to the technical solutions without departing from the core concept of the present invention. Any equivalent substitutions or simple modifications based on the basic concept of the present invention fall within the scope of protection of the present invention. The scope of protection of the present invention shall be subject to the claims.
Claims
1. A dual-branch multi-scale modeling semantic segmentation method based on tunnel scenes, characterized by: The steps include: Step S1: Collect tunnel face images from the tunnel construction process, construct a pixel-level annotation dataset containing typical structural features, improve sample diversity through data enhancement, and complete standardized division and statistical analysis. Step S2: Establish the context path of the dual-branch semantic segmentation model, use the residual network to extract high-dimensional contextual feature information through fast downsampling, design a spatially symmetric pyramid module to enhance the perception of semantic features, and finally fuse the feature map obtained by global average pooling upsampling output with the high-dimensional contextual feature information after the spatially symmetric pyramid module. Step S3: Establish the spatial path of the dual-branch semantic segmentation model, downsample the high-dimensional feature map, and perform pooling on the downsampled feature map. Add a dynamic-aware context enhancement module to the network, extract feature maps of different scales, and fuse them to obtain the output feature map of the spatial branch, which is used as one of the inputs of the feature fusion module. Step S4: Fuse the acquired contextual feature information with the spatial feature information, and finally input the predicted image to achieve the semantic segmentation task.
2. A dual-branch multi-scale modeling semantic segmentation method based on tunnel scenes according to claim 1, characterized in that: The specific experimental steps of step S1 are as follows: Workers captured images of the tunnel face during different construction phases to obtain on-site data on the construction environment. The images were normalized to a resolution of 448×448 pixels. Data augmentation was performed using methods such as mirror inversion, multi-angle rotation, noise introduction, and brightness adjustment. Manual annotation was performed using LabelMe annotation software, and the images were stored as .json files. These .json files were then converted into .png semantic label maps to complete the dataset. The dataset was randomly partitioned into training, validation, and test sets.
3. A dual-branch multi-scale modeling semantic segmentation method based on tunnel scenes according to claim 1, characterized in that: The specific experimental steps of step S2 are as follows: (1) When establishing the context path of the dual-branch semantic segmentation model, the ResNet18 residual network is used as the baseline network to extract high-dimensional contextual feature information through fast downsampling. The network consists of 18 layers, namely convolutional layers and fully connected layers. The convolution operation layer is divided into 1 independent convolution module and 4 residual modules. The residual module uses cross-layer connection technology to fuse the input features with the transformed output features. Each residual module consists of a 3×3 convolution layer, a batch normalization layer, and a ReLU layer. The network outputs feature maps of different scales, and global average pooling operations are performed on the 1 / 16 feature map and the 1 / 32 feature map respectively. (2) A spatial symmetric pyramid module is established. The output results after passing through the four residual modules are transferred to the spatial symmetric pyramid module, which can effectively extract the edge and structural features of the region and improve the model's ability to recognize the target region. In addition, the input features are processed using dilated convolutions with different expansion rates, which helps the model obtain information associations in a wider range. The standard dilated convolution is replaced by symmetric dilated convolution, which reduces the computational overhead while maintaining the expressive power. The input feature map uses a 1×1 group convolution to reduce the dimension and improve the computational efficiency, generating two new feature representations. In order to capture multi-scale features, four convolution kernels with sizes of 3, 5, 7, and 9 are used here. Symmetric convolution is used. F l =Conv 1×1 (F) Where: F is the input feature map, and F is obtained by 1×1 convolution l The feature map F l Input into the convolutional pyramid. Where: SymConv 3×3 (F l ) is a 3×3 symmetric convolution, Cat represents the feature concatenation operation to obtain F i ,here Indicates the addition of elements to obtain the final output feature F out .
4. A dual-branch multi-scale modeling semantic segmentation method based on tunnel scenes according to claim 1, characterized in that: The specific experimental steps of step S3 are as follows: (1) When building the spatial path of the dual-branch semantic segmentation model, the high-dimensional feature map is downsampled to reduce its size to 1 / 2 of its original size. Then, a pooling layer is applied to the downsampled feature map. At the same time, the feature map is further downsampled to 1 / 4 and 1 / 8 size. (2) A dynamic weight module is designed to perform weighted averaging of features at different scales. The size and weight of the convolution kernel are dynamically changed according to the current image content to adapt to the feature expectations of different regions. To capture the global dependencies between pixels, the dot product of the query and key is used to calculate the similarity between the two, and the result is normalized. Where: S ij Indicates the similarity between the i-th Query and the j-th Key; Using the attention weight matrix A and the Value feature V, the features of all pixels are weighted and summed to obtain the context feature of each pixel. Where: C i Represents the context features of all pixels, where each row corresponds to the global context representation of a pixel. In order to combine the context information with the original features, the context features are mapped back to the dimension of the input features through a linear projection and added element-by-element with the input features X to generate the final enhanced features. (3) Design a dynamic perception context enhancement module. By applying 1×1 convolution to the input feature map, preliminary context information extraction is performed. 3×3 convolution with hierarchical hole parameters is used to enable the network to learn local features of different scales. Adaptive average pooling is used in conjunction with the global attention mechanism to reduce the spatial dimension. Then, a 1×1 convolution operation is implemented. The processing nodes are composed of normalization technology and activation functions. The final feature map is generated by introducing the interaction between local and global features. The local context feature F is defined. ld and F lw ,These features are processed by different convolution kernels on the input features. Where: Cat(·) represents the channel splicing operation, Conv τ is a set of 3×3 depth convolutions with hole rates r=1, 2, 3, 4. Then, by combining the above-obtained feature F ld and F lw Fusion, generating local features F local . F local =F ld +F lw W=σ(F global ) Where: F local is a local feature, F global It is a global feature. Here, the Sigmoid activation function is used to perform weighted operations on the local features and global features obtained so that their weight range is between [0,1]. Specifically, the global feature F global Weighted superposition to local features F local Calculate the fusion weight W. Where: F final is the final output feature, Represents element-by-element multiplication operation, F l and F g are the adjusted local and global features respectively.
5. The method for semantic segmentation based on a dual-branch multi-scale modeling of a tunnel scene according to claim 1 is characterized in that: The specific experimental steps of step S4 are as follows: (1) Establish a cross-layer feature fusion module to combine the detailed information of shallow features with the semantic information of deep features, and build a comprehensive feature with both spatial resolution and semantic understanding capabilities. Use upsampling to adjust the deep low-resolution features to the same spatial size as the shallow high-resolution features; then use channel splicing or point-by-point convolution to achieve feature fusion operations, and finally rely on the attention mechanism to dynamically calibrate the weights of each layer of features. F CLFF =Attention(Concat(F shallow ,Upsample(F deep ))) Where: F shallow and F deep They represent shallow and deep features respectively, Upsample(·) is the upsampling operation, Concat(·) is the channel concatenation operation, and Attention(·) is used to assign dynamic weights to the fused features. (2) The present invention introduces a weighted cross entropy loss function and adopts a method of setting weight coefficients for each category, so that the model can focus more on the target area with a small number of samples during training. Where: N is the total number of pixels in the image, C is the total number of categories, y i,c is the true label of pixel i. If the pixel belongs to category c, then y i,c =1, otherwise 0, p i,c is the probability predicted by the model that pixel i belongs to category c. Where: w c is the weight of category c.
Citation Information
Cited By
Convolutional neural network-based surrounding rock grade identification method and system
CN121921307A
A method, system, and medium for measuring joint deformation based on multimodal information fusion.
CN122408649A