Lightweight road segmentation method and system for unstructured road scene, terminal and storage medium
By using a uniform channel enhancement convolution module and a uniform channel fusion decoder in the road segmentation method, only feature enhancement and semantic segmentation are performed on some channels, which solves the problem that the prior art is difficult to apply in unstructured scenarios, and achieves a high-precision and low-cost road segmentation effect.
Patent Information
- Application Number
- CN202510297103.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing road segmentation method is difficult to apply in unstructured scenarios, mainly due to the high computational cost of relying on complex models, and it is difficult to deal with the problems of irregular terrain, blurred boundaries, and complex obstacles such as mud and pits on the road surface.
A lightweight road segmentation method based on a uniform channel enhancement convolution module and a uniform channel fusion decoder is adopted to enhance feature and semantic segmentation of some channels through uniform channel sampling rate, reducing the amount of model parameters and reducing the operation cost.
It realizes high-precision road segmentation in unstructured road scenarios, reduces computing costs, and is suitable for robot navigation systems with limited hardware resources.
Smart Images

Figure CN120147645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road segmentation, and particularly relates to a lightweight road segmentation method, system, terminal and computer-readable storage medium for unstructured road scenarios. Background Art
[0002] In the field of robot vision navigation, the environment perception method based on deep learning has become the mainstream research direction. By accurately dividing the passable area through road segmentation technology, high-precision and high-reliability road information can be provided for the navigation decision-making system. Existing road segmentation research mainly focuses on processing structured scenarios, where lane lines and road boundaries are clear, the environment is relatively simple and the geometric features are obvious, and the technology of road segmentation in structured scenarios has been relatively mature. However, in unstructured scenarios, road segmentation faces more severe challenges. The terrain structure of unstructured scenarios is irregular, the terrain boundaries are staggered and blurred, and there are complex road obstacles such as mud and potholes on the road surface. In addition, for unstructured road environments, the overall hardware resources of traveling robots are generally limited, so there is also a problem of limited model size.
[0003] Road segmentation is crucial in autonomous driving and robot navigation by assigning passable labels to each pixel in the scene image. Although existing methods have made significant progress in structured scenarios, they rely on complex models to improve segmentation accuracy, resulting in high computational costs and being difficult to apply to unstructured scenarios.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide a lightweight road segmentation method, system, terminal and computer-readable storage medium for unstructured road scenarios, aiming to solve the problem that existing road segmentation methods rely on complex models to improve segmentation accuracy, resulting in high computational costs and being difficult to apply to unstructured road scenarios.
[0006] To achieve the above-mentioned invention purpose, the present invention provides a lightweight road segmentation method for unstructured road scenarios. The lightweight road segmentation method for unstructured road scenarios includes: Obtain an unstructured road scenario image; Input the unstructured road scenario image into an encoder based on a uniform channel enhancement convolutional module. The encoder extracts features from the unstructured road scenario image and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain multi-scale encoded features; Input the multi-scale encoded features into a uniform channel fusion decoder. The uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features, and uses multiple channel sampling rates to uniformly sample some channels for semantic segmentation to obtain a road segmentation result.
[0007] Optionally, the encoder includes an embedding layer, a first uniform channel enhancement convolution module combination, a first merging layer, a second uniform channel enhancement convolution module combination, a second merging layer, a third uniform channel enhancement convolution module combination, a third merging layer, and a fourth uniform channel enhancement convolution module combination connected in sequence. The first uniform channel enhancement convolution module combination, the second uniform channel enhancement convolution module combination, the third uniform channel enhancement convolution module combination, and the fourth uniform channel enhancement convolution module combination are respectively composed of different numbers of uniform channel enhancement convolution modules.
[0008] Optionally, input the unstructured road scene image into an encoder based on a uniform channel enhancement convolution module. The encoder extracts features from the unstructured road scene image, and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain multi-scale encoded features. Specifically, it includes: Input the unstructured road scene image into the embedding layer. The embedding layer extracts features, performs spatial downsampling, and expands the number of channels on the unstructured road scene image to obtain a feature map. Input the feature map into the first uniform channel enhancement convolution module combination. The first uniform channel enhancement convolution module combination uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain a first enhanced feature map. Output the first enhanced feature map and input the first enhanced feature map into the first merging layer. The first merging layer extracts features, performs spatial downsampling, and expands the number of channels on the first enhanced feature map to obtain a processed first enhanced feature map. Input the processed first enhanced feature map into the second uniform channel enhancement convolution module combination. The second uniform channel enhancement convolution module combination uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain a second enhanced feature map. Output the second enhanced feature map and input the second enhanced feature map into the second merging layer. The second merging layer extracts features, performs spatial downsampling, and expands the number of channels on the second enhanced feature map to obtain a processed second enhanced feature map. Input the processed second enhanced feature map into the third uniform channel enhanced convolution module combination. The third uniform channel enhanced convolution module combination uses multiple channel sampling rates to uniformly sample some channels respectively for feature enhancement, and obtains a third enhanced feature map; Output the third enhanced feature map and input the third enhanced feature map into a third merging layer. The third merging layer performs feature extraction, spatial downsampling, and channel number expansion on the third enhanced feature map to obtain a processed third enhanced feature map; Input the processed third enhanced feature map into the fourth uniform channel enhanced convolution module combination. The fourth uniform channel enhanced convolution module combination uses multiple channel sampling rates to uniformly sample some channels respectively for feature enhancement, and obtains a fourth enhanced feature map; Among them, the multi-scale encoded features include the first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map at different scales.
[0009] Optionally, the working process of the uniform channel enhanced convolution module specifically includes: Obtain an original feature map; Use three different channel sampling rates to uniformly sample some channels from the channels of the original feature map respectively, and obtain a first feature map corresponding to a first sampling channel, a second feature map corresponding to a second sampling channel, and a third feature map corresponding to a third sampling channel; Perform channel splicing on the second feature map and the third feature map, and use a 3×3 convolution module for feature enhancement to obtain a preliminary enhanced feature map; Perform channel splicing on the preliminary enhanced feature map and the unenhanced first feature map, and use a 1×1 convolution module for channel feature information interaction and fusion to obtain an enhanced feature map.
[0010] Optionally, when inputting the multi-scale encoded features into a uniform channel fusion decoder, the uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features, and uses multiple channel sampling rates to uniformly sample some channels respectively for semantic segmentation to obtain a road segmentation result, which specifically includes: Input the first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map into the uniform channel fusion decoder; The uniform channel fusion decoder spatially upsamples the first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map to the same scale, and performs channel splicing to obtain a fused feature map; Using multiple channel sampling rates to uniformly sample some channels of the fused feature map for feature enhancement, an enhanced fused feature map is obtained; Using the enhanced fused feature map for per-pixel classification, an accessible label is assigned to each pixel of the unstructured road scene image, and a road segmentation result is obtained.
[0011] Optionally, the step of using multiple channel sampling rates to uniformly sample some channels of the fused feature map for feature enhancement, obtaining an enhanced fused feature map, specifically includes: Using two different channel sampling rates, uniformly sampling some channels from the channels of the fused feature map respectively, obtaining a fourth feature map corresponding to a fourth sampling channel and a fifth feature map corresponding to a fifth sampling channel; Performing channel splicing on the fourth feature map and the fifth feature map, and using a 3×3 convolution module for feature enhancement, obtaining an enhanced fused feature map.
[0012] Optionally, the step of using multiple channel sampling rates to uniformly sample some channels specifically is: When using multiple channel sampling rates to uniformly sample some channels, the entire channel range is covered by uniformly sampling some channels multiple times.
[0013] To achieve the above invention purpose, the present invention also provides a lightweight road segmentation system for an unstructured road scene. The lightweight road segmentation system for an unstructured road scene includes: An image acquisition module: used to acquire an unstructured road scene image; An encoding module: used to input the unstructured road scene image into an encoder based on a uniform channel enhancement convolution module. The encoder extracts features from the unstructured road scene image, and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement, obtaining multi-scale encoded features; A decoding module: used to input the multi-scale encoded features into a uniform channel fusion decoder. The uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features, and uses multiple channel sampling rates to uniformly sample some channels for semantic segmentation, obtaining a road segmentation result.
[0014] To achieve the above invention purpose, the present invention also provides a terminal. The terminal includes: a memory, a processor, and a lightweight road segmentation program for an unstructured road scene stored on the memory and executable on the processor. When the lightweight road segmentation program for an unstructured road scene is executed by the processor, the steps of the above-mentioned lightweight road segmentation method for an unstructured road scene are implemented.
[0015] To achieve the above-mentioned invention objective, the present invention also provides a computer-readable storage medium storing a lightweight road segmentation program for unstructured road scenes. When the lightweight road segmentation program for unstructured road scenes is executed by a processor, the steps of the above-mentioned lightweight road segmentation method for unstructured road scenes are implemented.
[0016] In the present invention, an unstructured road scene image is acquired; the unstructured road scene image is input into an encoder based on a uniform channel enhancement convolutional module. The encoder extracts features from the unstructured road scene image and uses multiple channel sampling rates to uniformly sample some channels respectively for feature enhancement, obtaining multi-scale encoded features; the multi-scale encoded features are input into a uniform channel fusion decoder. The uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features and uses multiple channel sampling rates to uniformly sample some channels respectively for semantic segmentation, obtaining a road segmentation result. Through the uniform channel enhancement convolutional module and the uniform channel fusion decoder, the present invention only uniformly samples some channels for feature enhancement and semantic segmentation, improves the segmentation performance while reducing the number of model parameters, effectively reduces the operation cost, realizes a good balance between segmentation accuracy and model lightweight, and is applicable to unstructured road scenes. Description of the Drawings
[0017] Figure 1 is a flowchart of a preferred embodiment of the lightweight road segmentation method for unstructured road scenes of the present invention; Figure 2 is a structural schematic diagram of a preferred embodiment of the lightweight road segmentation method for unstructured road scenes of the present invention; Figure 3 is a structural schematic diagram of the uniform channel enhancement convolutional module of the present invention; Figure 4 is a structural schematic diagram of the uniform channel fusion decoder of the present invention; Figure 5 is a structural diagram of a preferred embodiment of the lightweight road segmentation system for unstructured road scenes of the present invention; Figure 6 is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Embodiments
[0018] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] The application scope of semantic segmentation technology is constantly expanding, showing remarkable results in the practical applications of various fields. Especially driven by image processing technology, semantic segmentation provides important support for the development of autonomous driving technology. The core goals of autonomous driving technology are to improve travel safety, reduce traffic accidents, and optimize the travel experience. Compared with traditional manual driving, it has a wider range of environmental perception, a more scientific driving decision-making mechanism, and a faster emergency response ability, while effectively avoiding safety hazards such as fatigue driving and drunk driving. An autonomous driving system is mainly composed of four subsystems: environmental perception, positioning and path planning, behavior decision-making, and control execution. Among them, the road segmentation algorithm provides the system with high-precision environmental perception ability, laying an important foundation for path planning and behavior decision-making. Road segmentation aims to achieve semantic parsing at the pixel level of the image, that is, to accurately assign a passable label to each pixel in the image scene. This fine-grained image understanding method significantly improves the computer's ability to analyze visual information.
[0020] In the field of robot visual navigation, the environmental perception method based on deep learning has become the mainstream research direction. Accurately dividing the passable area through road segmentation technology can provide high-precision and high-reliability road information for the navigation decision-making system. Existing road segmentation research mainly focuses on processing structured scenarios, where the lane lines and road boundaries are clear, the environment is relatively simple and the geometric features are obvious, and the road segmentation technology is relatively mature in structured scenarios. However, in unstructured scenarios, road segmentation faces more severe challenges: the terrain structure of unstructured scenarios is irregular, the terrain boundaries are staggered and blurred, and there are complex road obstacles such as mud and potholes on the road surface; secondly, some terrains lack regular features and appear more randomly. In addition, for unstructured road environments, the general hardware resources of traveling robots are limited, so there is also a problem of limited model size.
[0021] Road segmentation is crucial in autonomous driving and robot navigation by assigning a passable label to each pixel in the scene image. Although existing methods have made significant progress in structured scenarios, their reliance on complex models to improve segmentation accuracy leads to high computational costs and is difficult to apply to unstructured scenarios.
[0022] To solve the above technical problems, the present invention provides a lightweight road segmentation method for unstructured road scenes, which acquires unstructured road scene images; inputs the unstructured road scene images into an encoder based on a uniform channel enhancement convolutional module, and the encoder extracts features from the unstructured road scene images and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain multi-scale encoded features; inputs the multi-scale encoded features into a uniform channel fusion decoder, and the uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features and uses multiple channel sampling rates to uniformly sample some channels for semantic segmentation to obtain a road segmentation result. Through the uniform channel enhancement convolutional module and the uniform channel fusion decoder, the present invention only uniformly samples some channels for feature enhancement and semantic segmentation, improves the segmentation performance while reducing the number of model parameters, effectively reduces the computing cost, realizes a good balance between segmentation accuracy and model lightweight, and is applicable to unstructured road scenes.
[0023] The following further illustrates the application content by describing the embodiments in conjunction with the accompanying drawings.
[0024] A preferred embodiment of the lightweight road segmentation method for unstructured road scenes of the present invention is as Figure 1 and Figure 2 shown, and specifically includes: S1. Acquire unstructured road scene images.
[0025] Specifically, the terrain structure of the unstructured scene is irregular, the terrain boundaries are staggered and blurred, and there are complex road obstacles such as mud and potholes on the road surface; secondly, some terrains lack regular features and appear more randomly. In addition, for the unstructured road environment, the overall hardware resources of the traveling robot are generally limited, so there is also a problem of limited model size. Existing road segmentation methods have made remarkable progress in structured scenes, but they rely on complex models to improve segmentation accuracy, resulting in high computing costs and being difficult to apply to unstructured scenes. Based on this, the present invention proposes a lightweight road segmentation method for unstructured road scenes, called channel uniform sampling fusion road segmentation, which realizes a good balance between segmentation accuracy and model lightweight. Specifically, during implementation, first, an unstructured road scene is photographed by an image acquisition device such as a camera to obtain an unstructured road scene image, and the unstructured road scene image is used as the input image of the model.
[0026] S2. Input the unstructured road scene images into an encoder based on a uniform channel enhancement convolutional module, and the encoder extracts features from the unstructured road scene images and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain multi-scale encoded features.
[0027] In an implementation of this embodiment, the encoder includes an embedding layer, a first uniform channel enhancement convolution module combination, a first merging layer, a second uniform channel enhancement convolution module combination, a second merging layer, a third uniform channel enhancement convolution module combination, a third merging layer, and a fourth uniform channel enhancement convolution module combination that are connected in sequence. The first uniform channel enhancement convolution module combination, the second uniform channel enhancement convolution module combination, the third uniform channel enhancement convolution module combination, and the fourth uniform channel enhancement convolution module combination are each composed of a different number of uniform channel enhancement convolution modules.
[0028] Specifically, the present invention proposes a uniform channel enhancement convolution module that only performs feature enhancement, i.e., feature update, on some channels to reduce redundant calculations for channels with similar features. This module is used to construct a lightweight and efficient encoder. As Figure 2 shown, first, the encoder is divided into four stages (it can be understood that each stage corresponds to each uniform channel enhancement convolution module combination). Each stage is composed of a different number of uniform channel enhancement convolution modules, that is, each uniform channel enhancement convolution module combination is composed of a different number of uniform channel enhancement convolution modules. Among them, the first uniform channel enhancement convolution module combination is composed of 1 uniform channel enhancement convolution module, the second uniform channel enhancement convolution module combination is composed of 2 uniform channel enhancement convolution modules, the third uniform channel enhancement convolution module combination is composed of 8 uniform channel enhancement convolution modules, and the fourth uniform channel enhancement convolution module combination is composed of 2 uniform channel enhancement convolution modules. There is an embedding layer before the first stage, and there is a merging layer before each subsequent stage. The embedding layer and the merging layer are both used for feature extraction, spatial downsampling, and channel number expansion, but the downsampling rates of the embedding layer and the merging layer are different. Before the first stage, the image size is reduced to one-fourth of the original, and the image size is reduced by half in each subsequent stage. It should be noted that Figure 2 in, H represents the image height, W represents the image width, C1, C2, C3, and C4 respectively represent the number of channels of the feature maps output by the embedding layer, the first merging layer, the second merging layer, and the third merging layer. Among them, the first merging layer, the second merging layer, and the third merging layer are all uniformly labeled as merging layers in Figure 2 . The uniform channel enhancement convolution module performs feature update on some channels through multiple scale channel sampling rates, and the computational complexity is greatly reduced compared to the standard convolution.
[0029] In an implementation of this embodiment, inputting the unstructured road scene image into the encoder based on the uniform channel enhancement convolution module, the encoder extracts features from the unstructured road scene image and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain multi-scale encoded features, specifically including: Input the unstructured road scene image into the embedding layer, and the embedding layer performs feature extraction, spatial downsampling, and channel number expansion on the unstructured road scene image to obtain a feature map; Input the feature map into the first uniform channel enhancement convolution module combination, and the first uniform channel enhancement convolution module combination uses multiple channel sampling rates to uniformly sample some channels respectively for feature enhancement to obtain a first enhanced feature map; Output the first enhanced feature map and input the first enhanced feature map into the first merging layer, and the first merging layer performs feature extraction, spatial downsampling, and channel number expansion on the first enhanced feature map to obtain a processed first enhanced feature map; Input the processed first enhanced feature map into the second uniform channel enhancement convolution module combination, and the second uniform channel enhancement convolution module combination uses multiple channel sampling rates to uniformly sample some channels respectively for feature enhancement to obtain a second enhanced feature map; Output the second enhanced feature map and input the second enhanced feature map into the second merging layer, and the second merging layer performs feature extraction, spatial downsampling, and channel number expansion on the second enhanced feature map to obtain a processed second enhanced feature map; Input the processed second enhanced feature map into the third uniform channel enhancement convolution module combination, and the third uniform channel enhancement convolution module combination uses multiple channel sampling rates to uniformly sample some channels respectively for feature enhancement to obtain a third enhanced feature map; Output the third enhanced feature map and input the third enhanced feature map into the third merging layer, and the third merging layer performs feature extraction, spatial downsampling, and channel number expansion on the third enhanced feature map to obtain a processed third enhanced feature map; Input the processed third enhanced feature map into the fourth uniform channel enhancement convolution module combination, and the fourth uniform channel enhancement convolution module combination uses multiple channel sampling rates to uniformly sample some channels respectively for feature enhancement to obtain a fourth enhanced feature map; Wherein, the multi-scale encoded features include the first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map at different scales (referring to different channel dimensions).
[0030] Specifically, such as Figure 2As shown in the figure, the input image, i.e., the unstructured road scene image, is input into the encoder based on the uniform channel enhancement convolution module, and is processed successively through the embedding layer, the first uniform channel enhancement convolution module combination, the first merging layer, the second uniform channel enhancement convolution module combination, the second merging layer, the third uniform channel enhancement convolution module combination, the third merging layer, and the fourth uniform channel enhancement convolution module combination. The first uniform channel enhancement convolution module combination, the second uniform channel enhancement convolution module combination, the third uniform channel enhancement convolution module combination, and the fourth uniform channel enhancement convolution module combination respectively output the first enhanced feature map with different channel dimensions ( Figure 2 denoted as F1 in Figure 2 ), the second enhanced feature map ( Figure 2 denoted as F2 in Figure 2 ), the third enhanced feature map (
[0031] denoted as F3 in ), and the fourth enhanced feature map ( denoted as F4 in ), and F1, F2, F3, and F4 are respectively input into the decoder.
[0032] Figure 3 Specifically, it has been studied in the prior art that there is a high similarity between different channels in the feature map, and calculating all channel features will undoubtedly increase the redundant calculation of the model; the features of the same layer have great similarity. It divides the features into three branches in the channel dimension and uses different-sized convolution kernels to extract multi-scale features. Essentially, it still processes the entire channel, but it is more lightweight than using different convolution kernels to process the entire channel; only using partial channel features can effectively improve the model performance and reduce the number of parameters. However, it only selects some channels with continuous indexes for calculation and does not consider the information of other channels, resulting in sub-optimization of the model. Based on this, in order to reduce the computational redundancy and at the same time perform more detailed and uniform processing in the channel dimension, the present invention proposes a uniform channel enhancement convolution module, as Figure 3As shown, the module samples the input feature map channels (referring to the original feature map, Figure 3 denoted as A1 in this example) using three different channel sampling rates (the channel sampling rates used in this embodiment are 2, 3, and 5), and updates the spatial features of the sampled channels (referring to the first feature map, the second feature map, and the third feature map, Figure 3 denoted as B1, B2, and B3 respectively in this example). Subsequently, to avoid loss of useful information, the enhanced sampled features (referring to the preliminary enhanced feature map, Figure 3 denoted as D1 in this example) are concatenated with the unenhanced features (referring to the first feature map B1) in the channel dimension, and the enhanced feature information is propagated from partial channels to the entire channels. A 1×1 pointwise convolution is applied to the above features to obtain the enhanced feature map ( Figure 3 denoted as E1 in this example). It should be noted that, Figure 3 C in this example represents the original number of channels.
[0033] S3. Input the multi-scale encoded features into a uniform channel fusion decoder. The uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features, and uses multiple channel sampling rates to uniformly sample partial channels for semantic segmentation to obtain the road segmentation result.
[0034] In an implementation manner of this embodiment, the step of inputting the multi-scale encoded features into a uniform channel fusion decoder, where the uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features, and uses multiple channel sampling rates to uniformly sample partial channels for semantic segmentation to obtain the road segmentation result, specifically includes: Input the first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map into the uniform channel fusion decoder; The uniform channel fusion decoder upsamples the first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map to the same scale (referring to the same image size) in the spatial dimension, and performs channel concatenation to obtain a fused feature map; Use multiple channel sampling rates to uniformly sample partial channels of the fused feature map for feature enhancement to obtain an enhanced fused feature map; Perform per-pixel classification using the enhanced fused feature map, and assign passable labels to each pixel of the unstructured road scene image to obtain the road segmentation result.
[0035] Specifically, the deep network has a large receptive field and strong ability to represent semantic information, but its feature map has a small resolution and weak ability to represent geometric information (lack of spatial geometric feature details). The shallow network has a small receptive field, strong ability to represent geometric information, and high resolution feature maps, but weak ability to represent semantic information. There are objects of different sizes in the image, and different objects have different features. Using shallow features, simple objects can be distinguished; using deep features, complex objects can be distinguished.
[0036] The multi-scale features obtained by the encoder (i.e., multi-scale coding features, including F1, F2, F3, and F4) contain rich semantic information and contextual information, which helps to enhance the global semantic understanding ability of the model. The standard feature fusion method simply fuses features of different scales directly, without considering that features of the same scale still have high similarity on different channels. Based on this, in order to reduce the computational cost caused by redundant information, the present invention proposes a uniform channel fusion decoder, whose specific structure is as follows: Figure 4 As shown in Figure 1, first, the feature scales of the four stages of encoder parsing are different, and the decoder needs to upsample them to the same scale and achieve splicing from the channel, that is, multi-scale feature matching and fusion ( Figure 4 Fuse is used in the representation), and the fusion feature map is obtained ( Figure 4 Represented by G1 in the figure). Since the feature maps output by each stage of the encoder also have redundant information on the channel, semantic segmentation of all channel features invisibly increases the number of parameters. In order to reduce the calculation of redundant information, a simple and effective strategy is to use only the features of some channels rather than all channels for semantic segmentation. In this process, how to accurately select some channels is particularly critical. The present invention uniformly and meticulously covers the entire channel through multiple channel sampling, and adopts a fusion of multiple different channel sampling rates to achieve the selection of some channels and discard the remaining channel information. The features of some channels obtained by multiple channel sampling are updated and the number of channels is reduced to obtain an enhanced fused feature map ( Figure 4 The invention proposes a uniform channel fusion decoder, which fuses multi-scale coding features and uniformly samples some channel features for semantic segmentation, further enhancing the model's perception of context and reducing the calculation of redundant features.
[0037] In an implementation of this embodiment, the step of uniformly sampling some channels of the fused feature map using multiple channel sampling rates to perform feature enhancement to obtain an enhanced fused feature map specifically includes: Using two different channel sampling rates, uniformly sampling part of the channels from the channels of the fused feature map, respectively, to obtain a fourth feature map corresponding to the fourth sampling channel and a fifth feature map corresponding to the fifth sampling channel; Perform channel concatenation on the fourth feature map and the fifth feature map, and use a 3×3 convolution module for feature enhancement to obtain an enhanced fused feature map.
[0038] Specifically, to make full use of the multi-scale features obtained by the encoder, the uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale features obtained by the encoder, and also uses multiple channel sampling rates to select some channels for semantic segmentation, further reducing the model's computation of redundant information. The specific process is as Figure 4 shown. Uniform channel sampling is performed on the fused feature map G1 using two different channel sampling rates (the channel sampling rates used in this embodiment are 3 and 5) to obtain the fourth feature map and the fifth feature map ( Figure 4 denoted as L1 and L2 respectively in the figure), and channel concatenation and feature enhancement are performed on the fourth feature map L1 and the fifth feature map L2 to obtain the enhanced fused feature map J1. It should be noted that Figure 2 、 Figure 3 and Figure 4 the numbers marked on each channel of the feature maps in the figure are channel index values. Since different channel sampling rates are used to evenly and meticulously cover the entire channel, the decoder can accurately capture context information from the sampled partial channels, thereby improving the segmentation performance.
[0039] In one implementation manner of this embodiment, the use of multiple channel sampling rates to respectively and uniformly sample some channels specifically means: When using multiple channel sampling rates to respectively and uniformly sample some channels, the entire channel range of the feature map (referring to the original number of channels C) is covered by uniformly sampling some channels multiple times.
[0040] Specifically, since different channel sampling rates are used to evenly and meticulously cover the entire channel, the model can accurately capture context information from the sampled partial channels, thereby improving the segmentation performance.
[0041] In one implementation manner of this embodiment, as Figure 2 shown, the model of the present invention for road segmentation includes an encoder and a decoder.
[0042] In summary, the present invention performs uniform sampling in the channel dimension, filters out partial channel information for feature update, and at the same time retains the information of other channels. Compared with traditional convolution, it can extract sufficient features with fewer floating-point calculations. On the decoder side, based on traditional multi-scale feature matching and fusion, only partial channels are uniformly sampled for road segmentation, which improves the segmentation performance while reducing the number of model parameters. The overall model is lightweight and efficient by stacking lightweight uniform channel enhancement convolution modules. In addition, a large number of experiments have demonstrated the superiority of the segmentation performance of the present invention in unstructured road environments. It should be noted that feature update is feature enhancement.
[0043] In addition, based on the above lightweight road segmentation method for unstructured road scenarios, the present invention also provides a lightweight road segmentation system for unstructured road scenarios. Among them, a preferred embodiment of the lightweight road segmentation system for unstructured road scenarios is as Figure 5 shown and specifically includes: Image acquisition module 01: used to acquire unstructured road scenario images; Encoding module 02: used to input the unstructured road scenario image into an encoder based on a uniform channel enhancement convolution module. The encoder extracts features from the unstructured road scenario image and uses multiple channel sampling rates to uniformly sample partial channels for feature enhancement to obtain multi-scale encoded features; Decoding module 03: used to input the multi-scale encoded features into a uniform channel fusion decoder. The uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale encoded features and uses multiple channel sampling rates to uniformly sample partial channels for semantic segmentation to obtain a road segmentation result.
[0044] In addition, based on the above lightweight road segmentation method and system for unstructured road scenarios, the present invention also correspondingly provides a terminal. Among them, a preferred embodiment of the terminal is as Figure 6 shown and specifically includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0045] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, and a Flash Card equipped on the terminal, etc. Further, the memory 20 may also include both the internal storage unit of the terminal and the external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as storing the program code of the terminal, etc. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a lightweight road segmentation program 40 for unstructured road scenarios is stored on the memory 20, and the lightweight road segmentation program 40 for unstructured road scenarios can be executed by the processor 10, thereby implementing the steps of the lightweight road segmentation method for unstructured road scenarios in the present application.
[0046] The processor 10 may be a Central Processing Unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing the lightweight road segmentation program 40 for unstructured road scenarios, etc.
[0047] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface.
[0048] In one embodiment, when the processor 10 executes the lightweight road segmentation program 40 for unstructured road scenarios in the memory 20, the steps of the lightweight road segmentation method for unstructured road scenarios as described above are implemented.
[0049] The present invention also correspondingly provides a computer-readable storage medium, wherein the computer-readable storage medium stores a lightweight road segmentation program for unstructured road scenarios, and when the lightweight road segmentation program for unstructured road scenarios is executed by the processor, the steps of the lightweight road segmentation method for unstructured road scenarios as described above are implemented.
[0050] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal including such an element.
[0051] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above-described embodiments of the method can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes of the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0052] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A lightweight road segmentation method for unstructured road scenes, characterized in that: The lightweight road segmentation method for unstructured road scenes includes: Acquire unstructured road scene images; The unstructured road scene image is input into an encoder based on a uniform channel enhancement convolution module, the encoder extracts features from the unstructured road scene image, and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement, thereby obtaining multi-scale coding features; The multi-scale coding features are input into a uniform channel fusion decoder, which performs multi-scale feature matching and fusion on the multi-scale coding features, and uses multiple channel sampling rates to uniformly sample part of the channels for semantic segmentation to obtain a road segmentation result.
2. The lightweight road segmentation method for unstructured road scenes according to claim 1, characterized in that: The encoder includes an embedding layer, a first uniform channel enhanced convolution module combination, a first merging layer, a second uniform channel enhanced convolution module combination, a second merging layer, a third uniform channel enhanced convolution module combination, a third merging layer and a fourth uniform channel enhanced convolution module combination, which are connected in sequence. The first uniform channel enhanced convolution module combination, the second uniform channel enhanced convolution module combination, the third uniform channel enhanced convolution module combination and the fourth uniform channel enhanced convolution module combination are respectively composed of different numbers of uniform channel enhanced convolution modules.
3. The lightweight road segmentation method for unstructured road scenes according to claim 2, characterized in that: The unstructured road scene image is input into an encoder based on a uniform channel enhancement convolution module, the encoder extracts features from the unstructured road scene image, and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain multi-scale coding features, specifically including: Inputting the unstructured road scene image into the embedding layer, the embedding layer performs feature extraction, spatial downsampling and channel number expansion on the unstructured road scene image to obtain a feature map; Inputting the feature map into the first uniform channel enhancement convolution module combination, the first uniform channel enhancement convolution module combination uses multiple channel sampling rates to uniformly sample part of the channels for feature enhancement, and obtains a first enhanced feature map; Outputting the first enhanced feature map and inputting the first enhanced feature map into a first merging layer, wherein the first merging layer performs feature extraction, spatial downsampling, and channel number expansion on the first enhanced feature map to obtain a processed first enhanced feature map; Inputting the processed first enhanced feature map into the second uniform channel enhanced convolution module combination, the second uniform channel enhanced convolution module combination uses multiple channel sampling rates to uniformly sample part of the channels for feature enhancement, and obtains a second enhanced feature map; Outputting the second enhanced feature map and inputting the second enhanced feature map into a second merging layer, wherein the second merging layer performs feature extraction, spatial downsampling, and channel number expansion on the second enhanced feature map to obtain a processed second enhanced feature map; Inputting the processed second enhanced feature map into the third uniform channel enhanced convolution module combination, wherein the third uniform channel enhanced convolution module combination uses multiple channel sampling rates to uniformly sample part of the channels for feature enhancement, thereby obtaining a third enhanced feature map; Outputting the third enhanced feature map and inputting the third enhanced feature map into a third merging layer, wherein the third merging layer performs feature extraction, spatial downsampling, and channel number expansion on the third enhanced feature map to obtain a processed third enhanced feature map; Inputting the processed third enhanced feature map into the fourth uniform channel enhanced convolution module combination, wherein the fourth uniform channel enhanced convolution module combination uses multiple channel sampling rates to uniformly sample part of the channels for feature enhancement, thereby obtaining a fourth enhanced feature map; The multi-scale coding features include the first enhanced feature map, the second enhanced feature map, the third enhanced feature map and the fourth enhanced feature map of different scales.
4. The lightweight road segmentation method for unstructured road scenes according to claim 3 is characterized in that: The working process of the uniform channel enhancement convolution module specifically includes: Get the original feature map; Using three different channel sampling rates, uniformly sampling some channels from the channels of the original feature map respectively, to obtain a first feature map corresponding to the first sampling channel, a second feature map corresponding to the second sampling channel, and a third feature map corresponding to the third sampling channel; Channel-joining the second feature map and the third feature map, and performing feature enhancement using a 3×3 convolution module to obtain a preliminary enhanced feature map; The preliminary enhanced feature map and the unenhanced first feature map are channel-joined, and a 1×1 convolution module is used to interact and fuse channel feature information to obtain an enhanced feature map.
5. The lightweight road segmentation method for unstructured road scenes according to claim 3, characterized in that: The multi-scale coding features are input into a uniform channel fusion decoder, the uniform channel fusion decoder performs multi-scale feature matching and fusion on the multi-scale coding features, and uses multiple channel sampling rates to uniformly sample part of the channels for semantic segmentation to obtain a road segmentation result, specifically including: Inputting the first enhanced feature map, the second enhanced feature map, the third enhanced feature map and the fourth enhanced feature map into a uniform channel fusion decoder; The uniform channel fusion decoder spatially upsamples the first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map to the same scale, and performs channel splicing to obtain a fused feature map; Uniformly sampling some channels of the fused feature map using multiple channel sampling rates to perform feature enhancement, thereby obtaining an enhanced fused feature map; The enhanced fusion feature map is used to perform pixel-by-pixel classification, and a passable label is assigned to each pixel of the unstructured road scene image to obtain a road segmentation result.
6. The lightweight road segmentation method for unstructured road scenes according to claim 5, characterized in that: The step of uniformly sampling some channels of the fused feature map using multiple channel sampling rates to perform feature enhancement to obtain an enhanced fused feature map specifically includes: Using two different channel sampling rates, uniformly sampling part of the channels from the channels of the fused feature map, respectively, to obtain a fourth feature map corresponding to the fourth sampling channel and a fifth feature map corresponding to the fifth sampling channel; The fourth feature map and the fifth feature map are channel-joined, and feature enhancement is performed using a 3×3 convolution module to obtain an enhanced fused feature map.
7. The lightweight road segmentation method for unstructured road scenes according to claim 1, characterized in that: The method of using multiple channel sampling rates to uniformly sample some channels is specifically as follows: When multiple channel sampling rates are used to uniformly sample some channels respectively, the entire channel range is covered by uniformly sampling some channels multiple times.
8. A lightweight road segmentation system for unstructured road scenes, characterized in that: The lightweight road segmentation system for unstructured road scenes includes: Image acquisition module: used to acquire unstructured road scene images; Coding module: used for inputting the unstructured road scene image into an encoder based on a uniform channel enhancement convolution module, wherein the encoder extracts features from the unstructured road scene image and uses multiple channel sampling rates to uniformly sample some channels for feature enhancement to obtain multi-scale coding features; Decoding module: used to input the multi-scale coding features into the uniform channel fusion decoder, which performs multi-scale feature matching and fusion on the multi-scale coding features, and uses multiple channel sampling rates to uniformly sample part of the channels for semantic segmentation to obtain road segmentation results.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a lightweight road segmentation program for unstructured road scenes stored in the memory and executable on the processor. When the lightweight road segmentation program for unstructured road scenes is executed by the processor, the steps of the lightweight road segmentation method for unstructured road scenes as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a lightweight road segmentation program for unstructured road scenes, and when the lightweight road segmentation program for unstructured road scenes is executed by a processor, the steps of the lightweight road segmentation method for unstructured road scenes as described in any one of claims 1-7 are implemented.