Image semantic segmentation method for unstructured terrain
By building a semantic segmentation network model including MobiliNetV2 network, SE module, ASPP module and CBAM module, the problem of unstructured terrain image segmentation in complex wild environments is solved, and the robot's perception ability and navigation stability of unstructured terrain are improved.
Patent Information
- Application Number
- CN202510365308.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-24
AI Technical Summary
In complex wild environments, prior art is difficult to accurately identify and segment unstructured terrain images, resulting in the impact of robot stability and security during movement and navigation.
A semantic segmentation method for unstructured terrain is adopted. By building a semantic segmentation network model, the model includes the MobiliNetV2 network, the SE module, the ASPP module and the CBAM module. Combined with attention mechanism and multi-scale feature extraction, the model's ability to extract different terrain features is improved.
This method effectively solves the problem of image semantic segmentation in real wild scenes with a high similarity between the terrain and non-terrain regions, improves the robot's perception of unstructured terrain, and ensures that it operates efficiently and stably under complex terrain.
Smart Images

Figure CN120198670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an image segmentation method. Background Art
[0002] With the rapid development of information technology, artificial intelligence technology and robotics technology, the application scenarios of mobile robots have expanded from static indoor environments to complex and changing outdoor environments. Different from structured indoor environments, outdoor scenes are full of uncertainties. Robots must face various complex road surfaces, such as soft sand, muddy wetlands, rugged mountain roads, etc. These unstructured terrains may seriously affect the movement stability of robots and even pose operation risks. In practical applications, the images captured by robot cameras not only contain terrain information (such as grasslands, sand), but may also be mixed with backgrounds such as trees, buildings, and the sky in the surrounding environment. This makes it difficult to accurately identify passable areas only relying on traditional classification or detection methods. Therefore, performing pixel-level semantic segmentation on the terrain and extracting accurate ground information has become one of the key technologies for robot autonomous navigation and environmental perception.
[0003] The core of pixel-level terrain recognition lies in classifying and labeling each pixel point in the image to distinguish different types of terrains and environmental objects. This technology not only helps robots perceive and understand the surrounding environment, but also provides key information during autonomous navigation, such as path planning, travel strategy adjustment, and speed control. In complex outdoor environments, accurate terrain segmentation can help robots more comprehensively perceive the surrounding world, ensure their efficient and stable operation in various complex terrains, and provide solid technical support for the wide application of intelligent robots in unstructured environments. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an image semantic segmentation method for unstructured terrains to solve the problem of semantic segmentation of unstructured terrain images in real outdoor scenes.
[0005] The image semantic segmentation method for unstructured terrains of the present invention includes the following steps:
[0006] 1) Obtain an unstructured terrain image dataset, and divide the data in the dataset into training data, validation data, and a test set;
[0007] 2) Build a semantic segmentation network model:
[0008] The semantic segmentation network model includes a MobiliNetV2 network for extracting features from the input image;
[0009] The semantic segmentation network model further includes a 1×1 convolutional layer connected to the shallow output of the MobiliNetV2 network and an SE module connected to the output of the 1×1 convolutional layer;
[0010] The semantic segmentation network model further includes an ASPP module connected to the deep output of the MobiliNetV2 network, a 1×1 convolutional layer connected to the output of the ASPP module, a CBAM module connected to the output of the 1×1 convolutional layer, and a first upsampling layer connected to the output of the CBAM module;
[0011] The semantic segmentation network model further includes a concat layer connected to the output of the first upsampling layer and the output of the SE module, a 3×3 convolutional block connected to the output of the concat layer, and a second upsampling layer connected to the output of the 3×3 convolutional block;
[0012] 3) Use the training data and validation data obtained in step 1) to train the semantic segmentation network model, and use the test set to test the trained semantic segmentation network model to obtain a qualified semantic segmentation network model;
[0013] 4) Use the tested and qualified semantic segmentation network model to perform semantic segmentation on the unstructured terrain image.
[0014] Further, the unstructured terrain image dataset obtained in step 1) is the RUGD dataset.
[0015] Further, the image semantic segmentation method for unstructured terrain further includes performing label remapping processing on the images in the RUGD dataset, and the label remapping rule is: label 0 maps to the background, label 1 maps to gravel roads and dirt, label 2 maps to grasslands, label 3 maps to the sky, and label 4 maps to concrete roads.
[0016] Further, the shallow output of the MobiliNetV2 network refers to the output of the first four layers of the MobiliNetV2 network, and the deep output of the MobiliNetV2 network refers to the output after the fourth layer.
[0017] The beneficial effects of the present invention:
[0018] The image semantic segmentation method for unstructured terrain of the present invention constructs a semantic segmentation network model that uses the lightweight MobiliNetV2 network as the feature extraction network and integrates the SE attention mechanism module and the CABM attention mechanism module, improving the model's feature extraction ability for different terrains, enabling the model to effectively handle the image semantic segmentation problem in real outdoor scenarios where there are many sundries and the similarity between terrain and non-terrain areas is relatively high; and the constructed semantic segmentation network model has fewer parameters and computational amounts, reducing resource occupancy. Brief Description of the Drawings
[0019] Figure 1 It is a schematic structural diagram of a semantic segmentation network model. Specific Embodiments
[0020] The present invention will be further described below in conjunction with the drawings and embodiments.
[0021] The method for image semantic segmentation for unstructured terrain in this embodiment includes the following steps:
[0022] 1) Obtain an unstructured terrain image dataset. Unstructured terrain refers to natural or artificial terrain with complex, irregular surface morphology and lack of obvious patterns. The distribution and morphology of the elements constituting the unstructured terrain have no fixed patterns, and these elements include dense grasslands, sparse vegetation, sandy land, muddy land, scattered gravel, and irregularly distributed blocky rocks, etc. Divide the data in the dataset into training data, validation data, and test sets; in this step, the ratio of dataset division is 8:1:1, that is, the dataset is divided into a training set (80%), a test set (10%), and a validation set (10%) to ensure the generalization ability of the model during training and effectively evaluate its performance.
[0023] The unstructured terrain image dataset obtained in this step is the RUGD dataset. Of course, in different embodiments, other unstructured terrain image datasets can also be selected. The images in the RUGD dataset cover a variety of complex terrain scenarios, and each image in it contains an original JPG image and a PNG mask image. The ground object information of different categories is marked in the image mask, and the categories and their corresponding RGB mask values are shown in Table 1:
[0024]
[0025]
[0026] For the original 25 categories in the RUGD dataset, perform label remapping processing on the images in the RUGD dataset. The label remapping rule is: label 0 maps to the background, label 1 maps to gravel roads and mud, label 2 maps to grasslands, label 3 maps to the sky, and label 4 maps to concrete roads. Optimize and adjust the category labels in the original data through label remapping to ensure semantic consistency of different categories, reduce the occupation of computing resources, and improve the learning effect of the model.
[0027] 2) Build a semantic segmentation network model:
[0028] The semantic segmentation network model includes a MobiliNetV2 network for extracting features from the input image. The MobiliNetV2 network is a lightweight feature extraction network, mainly composed of structures such as linear bottlenecks, inverted residuals, and depthwise separable convolutions. Its structure and main parameters are as follows:
[0029] MobiliNetV2 Structure and Main Parameters Among them, the depthwise separable convolution consists of two parts: Depthwise convolution and Pointwise convolution; Depthwise convolution is a convolution operation that is completely performed in the two-dimensional plane, and the relationship between the channels and the convolution kernels is one-to-one; Pointwise convolution is a normal convolution with a convolution kernel size of lx1, which is located after the Depthwise convolution and is used to fuse the features of multiple channels to enhance the network's expressive ability.
[0030] During the convolution operation process, if the number of input channels is Ci, the convolution kernel size is k×k, the number of output channels is C0, and the output feature size is H×W, the ratio of the number of parameters between the depthwise separable convolution and the standard convolution is shown in the following formula:
[0031]
[0032] The ratio of the amount of computation is shown in the formula:
[0033]
[0034] It can be seen from the two formulas that compared with the standard convolution, the depthwise separable convolution significantly reduces the computational complexity, effectively reduces the number of parameters, and improves the computational efficiency, thus meeting the requirements of the lightweight model for low computational cost and high operation speed.
[0035] The semantic segmentation network model also includes a 1×1 convolution layer connected to the shallow output (i.e., the output of the first four layers) of the MobiliNetV2 network and an SE (Squeeze-and-Excitation) module connected to the output of the 1×1 convolution layer. The SE module is a channel attention module, aiming to adaptively adjust the importance of feature channels to enhance the feature expression ability of the deep neural network. Its operation is mainly divided into four steps:
[0036] (1) Feature Map Transformation
[0037] Input X∈R H'×W'×C' Through an F tr transformation, it is mapped into U∈R H×W×C , F tr The transformation is a standard convolution operation, and the calculation formula is
[0038]
[0039] where v c represents the c-th convolutional kernel, and x s represents the s-th input under the coverage of the current convolutional kernel. C' represents the number of convolutional kernels, and u c is the c-th element of u. u represents the transformed feature map. is a 2D spatial kernel, representing the single channel of v c acting on the corresponding channel of X.
[0040] (2) Squeeze stage
[0041] This operation performs global average pooling on the feature map to generate a 1×1×C vector Z, where each channel is represented by a single value. The formula is as follows
[0042]
[0043] where H is the height of the feature map, W is the width of the feature map, and z c is the output value of the c-th channel after global average pooling;
[0044] (3) Excitation
[0045] In the Excitation stage, the obtained compressed vector is processed through two fully connected layers. The first fully connected layer compresses the C channels into C / r channels to reduce the computational load (r refers to the compression ratio), and then passes through a ReLU non-linear activation layer. The second fully connected layer restores the number of channels back to C channels, and then obtains the weight s through the Sigmoid activation. Finally, a vector s with a dimension of 1×1×C is obtained, and s is used to characterize the weights of the C feature maps in the feature map U.
[0046] s = F ex (z, W) = σ(g(z, W)) = σ(W2δ(W1z)) (5)
[0047] where δ refers to the ReLU function, and W1 and W2 are the weight matrices of the two fully connected layers respectively.
[0048] (4) Scale operation
[0049] The attention weights obtained previously are weighted to the features of each channel, and each feature map in the feature map U is multiplied by the corresponding weight to obtain the final output of the SE module
[0050]
[0051] The semantic segmentation network model further includes an ASPP module connected to the deep output (i.e., the output after the fourth layer) of the MobileNetV2 network, a 1×1 convolutional layer connected to the output of the ASPP module, a CBAM module connected to the output of the 1×1 convolutional layer, and a first upsampling layer connected to the output of the CBAM module.
[0052] After the input image passes through MobileNetV2, the deep feature map output by MobileNetV2 is input into the ASPP (Atrous Spatial Pyramid Pooling) module, that is, the atrous spatial pyramid pooling module, for multi-scale feature extraction. ASPP includes: a 1×1 standard convolutional layer for reducing the dimension of the input features and reducing the computational amount; three 3×3 atrous convolutional layers with different dilation rates (6, 12, 18) for capturing features of different scales; and a global average pooling branch for capturing global context information; the output of the global average pooling branch is reduced in dimension through a 1×1 convolution and bilinearly interpolated to match the size of the original feature map. The outputs of all branches of ASPP are concatenated in the channel dimension, and then the features are fused through a 1×1 convolution, and passed through batch normalization and the ReLU6 activation function.
[0053] CBAM (Convolutional Block Attention Module) is an attention mechanism that combines two dimensions of feature channels and feature space, which can improve the model accuracy in complex image scenarios. CBAM contains a channel attention module CAM and a spatial attention module SAM. The operation process of CBAM is shown in Equation (7). The input image is first processed by the channel attention module, multiplied by the input image, and then the multiplication result is processed by the spatial attention module and multiplied by the result of the previous step to obtain the final output image. In the formula, F is the input image, M C is the channel attention module, M C (F) is a one-dimensional array output by the channel attention module, M s is the spatial attention module, M s (F′) is a two-dimensional image output by the spatial attention module, is element-wise multiplication of arrays (the parts with insufficient dimensions are replicated), and F″ is the finally output image.
[0054]
[0055] (1) Channel attention module
[0056] The channel attention module shrinks the input image in the width and height directions by max-pooling and average-pooling to a size of C×1×1, then inputs the two results into an MLP (Multi-Layer Perceptron) respectively. The output results of the two MLPs are added together, passed through a Sigmoid layer, and the output result is obtained.
[0057] (2) Spatial attention mechanism module
[0058] The spatial attention module shrinks the input image in the channel direction by max-pooling and average-pooling to a size of 1×H×W. The two results are directly concatenated in the channel direction and then passed through a 1×1 convolutional layer to reduce the number of channels. After passing through a sigmoid layer, the output result is obtained.
[0059] The semantic segmentation network model further includes a concat layer connected to the output of the first upsampling layer and the output of the SE module, a 3×3 convolutional block connected to the output of the concat layer, and a second upsampling layer connected to the output of the 3×3 convolutional block. The first upsampling layer performs a 4-fold upsampling operation on the feature map extracted by the CBAM module through bilinear interpolation. The concat layer performs splicing and fusion on the feature map. The fused feature map is further extracted by a 3x3 convolution, and then the second upsampling layer performs a 4-fold upsampling operation to restore it to the original image size, and the prediction map is output.
[0060] 3) Use the training data and validation data obtained in step 1) to train the semantic segmentation network model, and use the test set to test the trained semantic segmentation network model to obtain a qualified semantic segmentation network model. In this embodiment, during the training process, the Stochastic Gradient Descent (SGD) optimization algorithm is adopted, the momentum is set to 0.9, the initial learning rate is set to 0.01, the weight decay coefficient (Weight Decay) is set to 1e-4, the learning rate decay strategy adopts poly (polynomial decay), and the exponential decay rate is set to 0.9. The loss function uses the cross-entropy loss function based on Softmax.
[0061] 4) Use the qualified semantic segmentation network model obtained through testing to perform semantic segmentation on the unstructured terrain image.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An image semantic segmentation method for unstructured terrain, characterized by: The following steps are involved: 1) Obtain an unstructured terrain image dataset and divide the data in the dataset into training data, verification data and test set; 2) Build a semantic segmentation network model: The semantic segmentation network model includes a MobiliNetV2 network for extracting features from an input image; The semantic segmentation network model also includes a 1×1 convolutional layer connected to the shallow output of the MobiliNetV2 network and a SE module connected to the output of the 1×1 convolutional layer; The semantic segmentation network model also includes an ASPP module connected to the deep output of the MobiliNetV2 network, a 1×1 convolutional layer connected to the output of the ASPP module, a CBAM module connected to the output of the 1×1 convolutional layer, and a first upsampling layer connected to the output of the CBAM module; The semantic segmentation network model also includes a concat layer connected to the output of the first upsampling layer and the output of the SE module, a 3×3 convolution block connected to the output of the concat layer, and a second upsampling layer connected to the output of the 3×3 convolution block; 3) Using the training data and verification data obtained in step 1) to train the semantic segmentation network model, and using the test set to test the trained semantic segmentation network model to obtain a qualified semantic segmentation network model; 4) Use the tested and qualified semantic segmentation network model to perform semantic segmentation on unstructured terrain images.
2. The image semantic segmentation method for unstructured terrain according to claim 1, characterized in that: The unstructured terrain image dataset obtained in step 1) is the RUGD dataset.
3. The image semantic segmentation method for unstructured terrain according to claim 2, characterized in that: It also includes label remapping of images in the RUGD dataset. The label remapping rules are as follows: label 0 maps the background, label 1 maps gravel roads and soil, label 2 maps grass, label 3 maps the sky, and label 4 maps concrete roads.
4. The image semantic segmentation method for unstructured terrain according to claim 1, characterized in that: The shallow output of the MobiliNetV2 network refers to the output of the first four layers of the MobiliNetV2 network, and the deep output of the MobiliNetV2 network refers to the output after the fourth layer.