Semantic segmentation method for aided driving, electronic equipment and storage medium

By introducing scale convolution groups and spatial detail modules into the semantic segmentation model, and combining with the decoder to complement each other in features, the problems of high computational complexity and poor detail perception in the prior art are solved, and efficient and accurate semantic segmentation is achieved, which is suitable for assisted driving.

CN120564192APending Publication Date: 2025-08-29CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510665109.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing semantic segmentation methods have high computational complexity, poor detail perception, low robustness, low boundary segmentation accuracy, and poor robustness to complex backgrounds and lighting changes, resulting in limited real-time decision-making and degradation of segmentation performance in autonomous driving.

Method used

The semantic segmentation model using an encoder and decoder structure includes at least two scale convolution groups and spatial detail modules. Semantic features are extracted through scale convolution groups and spatial details modules are used to extract spatial features, combined with the decoder to complement each other, and finally generate a semantic segmented image.

Benefits of technology

It reduces the complexity of semantic segmentation, improves feature extraction efficiency, accurately captures tiny objects and complex edges in the image, improves segmentation accuracy and robustness, and is suitable for real-time assisted driving tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564192A_ABST
    Figure CN120564192A_ABST
Patent Text Reader

Abstract

The invention provides a semantic segmentation method for aided driving, electronic equipment and a storage medium, and the method comprises the steps: obtaining an environment image collected by an aided driving function in a vehicle, inputting the environment image into a pre-trained semantic segmentation model, determining the semantic features of the environment image according to each scale convolution group in the model, and obtaining the semantic features of the environment image. The method comprises the steps of obtaining a model, extracting the spatial features of an environment image from the output of each scale convolution group according to a spatial detail module in the model, and finally determining a semantic segmentation image corresponding to the environment image according to a decoder in the model, the semantic features and the spatial features, thereby achieving the semantic segmentation of the environment image collected by an auxiliary driving function. The multi-level features of the image can be effectively extracted, fine objects and complex edges in the image can be accurately captured, the attention of the model to details is improved, semantic segmentation is more accurate and comprehensive, and meanwhile the semantic segmentation efficiency is improved through feature multiplexing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent assisted driving technology, and specifically to a semantic segmentation method, electronic device, and storage medium for assisted driving. Background Art

[0002] Intelligent assisted driving technology is a key area of ​​development for modern automotive technology, encompassing a range of perception, decision-making, and control technologies. Semantic segmentation, which aims to classify pixels in images or videos, plays a crucial role in the environmental perception phase of intelligent assisted driving technology. It provides autonomous driving systems with a deep understanding of their environment, helping to improve safety, comfort, and efficiency.

[0003] Currently, there are numerous semantic segmentation methods for road scenes, but they suffer from the following major drawbacks: 1. High computational complexity, especially for models such as U-Net and DeepLabV3+, which require enormous computing resources and storage space during both training and inference. Furthermore, inference speed may be limited by hardware resources, impacting real-time decision-making in autonomous driving. 2. Inadequate processing of small objects and details. Due to the limited local perception capabilities of deep neural networks, it is difficult to accurately capture the features of small objects, especially in multi-scale environments. For example, in autonomous driving, small objects (such as animals and road obstacles) may be misclassified or missed. 3. Poor robustness to complex backgrounds and lighting changes. Models typically rely on large-scale training data to capture various variations in the environment. However, under different lighting, weather, or complex backgrounds, the robustness of the model may decrease, resulting in poor segmentation performance. 4. Blurred edges and poor precision. The segmentation accuracy of the model is often low at the edges of objects, and boundaries may become blurred. 5. Overfitting and insufficient generalization capabilities. Due to the complexity of deep neural networks, semantic segmentation models may overfit during training. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present application aims to provide a semantic segmentation method, electronic device and storage medium for assisted driving to solve problems such as high semantic segmentation complexity, poor detail perception, low robustness and low boundary segmentation accuracy.

[0005] This embodiment of the present application provides a semantic segmentation method for assisted driving, the method comprising: Obtaining an environmental image captured by an assisted driving function in a vehicle, and inputting the environmental image into a pre-trained semantic segmentation model, wherein the semantic segmentation model includes an encoder and a decoder, and the encoder includes at least two scale convolution groups and a spatial detail module; Determining semantic features of the environment image based on each of the scale convolution groups, and extracting spatial features of the environment image from outputs of each of the scale convolution groups based on the spatial detail module; A semantic segmentation image corresponding to the environment image is determined based on the decoder, the semantic features, and the spatial features.

[0006] Optionally, the scale convolution groups are sequentially connected, and determining the semantic features of the environment image based on the scale convolution groups includes: For each of the scale convolution groups, downsampling and feature extraction are performed on the environment image or the output of the previous scale convolution group based on the scale convolution group; The output of the last scale convolution group is determined as the semantic feature of the environment image.

[0007] Optionally, the scale convolution group includes a first spatial attention module, a downsampling unit, a convolution block, a first splicing unit, and a first channel attention module, and downsampling and feature extraction are performed on the environment image or the output of the previous scale convolution group based on the scale convolution group, including: Extracting spatial key information from the environment image or the output of the previous scale convolution group based on the first spatial attention module; Performing downsampling processing on the environment image or the output of the previous scale convolution group based on the downsampling unit to obtain a downsampling result, and performing convolution processing on the downsampling result based on the convolution block to obtain a convolution result; splicing the spatial key information, the convolution result, and the downsampling result based on the first splicing unit to obtain a splicing result; The weight of each channel in the splicing result is adjusted based on the first channel attention module.

[0008] Optionally, the convolution block includes a plurality of standard convolution layers or a plurality of depth residual modules connected in sequence, and performing convolution processing on the downsampling result based on the convolution block to obtain a convolution result includes: For each of the standard convolutional layers, performing convolution processing on the downsampling result or the output of the previous standard convolutional layer based on the standard convolutional layer; Alternatively, for each of the depth residual modules, convolution processing is performed on the downsampling result or the output of the previous depth residual module based on the depth residual module.

[0009] Optionally, the depth residual module includes a first point-by-point convolution unit, a channel grouping unit, a left branch depth convolution unit, a right branch depth convolution unit, a second point-by-point convolution unit, a second splicing unit, a connection unit, and a channel shuffling unit, and performs convolution processing on the downsampling result or the output of the previous depth residual module based on the depth residual module, including: Convolving the downsampling result or the output of the previous depth residual module based on the first point-by-point convolution unit, and grouping the output of the first point-by-point convolution unit along the channel dimension based on the channel grouping unit; Performing decomposition convolution on the output of the channel grouping unit based on the left depth convolution unit, and performing dilation convolution on the output of the channel grouping unit based on the right depth convolution unit, wherein the left depth convolution unit and the right depth convolution unit transmit information to each other during the convolution process; splicing the outputs of the left depth convolution unit and the right depth convolution unit based on the second splicing unit, and convolving the output of the second splicing unit based on the second point-by-point convolution unit; The output of the second point-by-point convolution unit is connected to the input of the first point-by-point convolution unit based on the connection unit, and the output of the connection unit is rearranged along the channel dimension based on the channel shuffling unit.

[0010] Optionally, the spatial detail module includes a first fusion unit, multiple convolutional layers, a second fusion unit, and a second spatial attention module. Extracting spatial features of the environment image from outputs of each of the scale convolution groups based on the spatial detail module includes: fusing the outputs of the scale convolution groups based on the first fusion unit, and convolving the outputs of the first fusion unit based on the convolution layers; The outputs of the convolutional layers are fused based on the second fusion unit, and spatial key information in the output of the second fusion unit is extracted based on the second spatial attention module to obtain the spatial features of the environmental image.

[0011] Optionally, the decoder includes a feature complementation module, a third point-by-point convolution unit, a second channel attention module, and an upsampling unit, and determines a semantic segmentation image corresponding to the environment image based on the decoder, the semantic features, and the spatial features, including: fusing the semantic features and the spatial features based on the feature complementation module, and convolving the output of the feature complementation module based on the third point-by-point convolution unit; Adjusting the weight of each channel in the output of the third point-by-point convolution unit based on the second channel attention module; The output of the second channel attention module is upsampled based on the upsampling unit to obtain a semantic segmentation image.

[0012] Optionally, the feature complementation module includes a fourth point-by-point convolution unit, a bilinear interpolation unit, a fifth point-by-point convolution unit, a third channel attention module, a third spatial attention module, a first element multiplication unit, a second element multiplication unit, and an element addition unit. The semantic feature and the spatial feature are fused based on the feature complementation module, including: Convolving the spatial features based on the fourth point-by-point convolution unit, and extracting spatial key information from the output of the fourth point-by-point convolution unit based on the third spatial attention module; Performing bilinear interpolation processing on the semantic features based on the bilinear interpolation unit; Convolving the output of the bilinear interpolation unit based on the fifth point-by-point convolution unit, and adjusting the weight of each channel in the output of the fifth point-by-point convolution unit based on the third channel attention module; Performing element-by-element multiplication on the output of the fourth point-by-point convolution unit and the third channel attention module based on the first element multiplication unit, and performing element-by-element multiplication on the output of the fifth point-by-point convolution unit and the third spatial attention module based on the second element multiplication unit; Based on the element-adding unit, outputs of the first element-multiplying unit and the second element-multiplying unit are added element by element to fuse the semantic features with the spatial features.

[0013] An embodiment of the present application further provides an electronic device, comprising: processor and memory; The processor is used to execute the steps of the semantic segmentation method for assisted driving provided in any embodiment of the present application by calling the program or instructions stored in the memory.

[0014] An embodiment of the present application also provides a computer-readable storage medium, which stores a program or instruction, and the program or instruction enables a computer to execute the steps of the semantic segmentation method for assisted driving provided in any embodiment of the present application.

[0015] In summary, the present application proposes a semantic segmentation method for assisted driving, which obtains the environmental image collected by the assisted driving function in the vehicle, and inputs the environmental image into a pre-trained semantic segmentation model, and then determines the semantic features of the environmental image according to the convolution groups of each scale in the model, and extracts the spatial features of the environmental image from the output of the convolution groups of each scale according to the spatial detail module in the model. Finally, according to the decoder and the semantic features and spatial features in the model, the semantic segmentation image corresponding to the environmental image is determined to realize the semantic segmentation of the environmental image collected by the assisted driving function. This method extracts semantic features through the convolution groups of each scale, which can effectively extract the multi-level features of the image and enhance the signal quality. Information flow is enhanced, and the feature extraction capability of the model is improved. Moreover, this method determines the spatial features through the output of the convolution groups of each scale, realizes feature reuse in the process of semantic information processing, and does not need to process the environmental image separately, further improves the efficiency of feature extraction, and solves the problem of high complexity of semantic segmentation in related technologies. In addition, through spatial features, small objects and complex edges in the image can be accurately captured, the segmentation accuracy is improved, and the problems of poor detail perception, low robustness and low boundary segmentation accuracy in related technologies are solved. This method also integrates spatial features and semantic features to obtain the final segmentation result, realizes feature complementarity, effectively improves the model's attention to details, and makes semantic segmentation more accurate and comprehensive. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 This is a flowchart of a semantic segmentation method for assisted driving provided in an embodiment of the present application; Figure 2 is a schematic diagram of a semantic segmentation model provided in an embodiment of the present application; Figure 3 Schematic diagram of a scale convolution group provided in an embodiment of the present application; Figure 4 is a schematic diagram of a depth residual module provided in an embodiment of the present application; Figure 5 is a schematic diagram of a spatial detail module provided in an embodiment of the present application; Figure 6 This is a schematic diagram of a model processing provided by an embodiment of the present application; Figure 7 is a schematic diagram of a decoder provided in an embodiment of the present application; Figure 8 is a schematic diagram of a feature complementary module provided in an embodiment of the present application; Figure 9 is a schematic diagram of a spatial attention module provided in an embodiment of the present application; Figure 10 It is a channel attention module provided by an embodiment of the present application; Figure 11 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.

[0019] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0020] As mentioned in the background technology, in response to the problems in the existing technology, this application proposes a semantic segmentation method for assisted driving. Figure 1 This is a flowchart of a semantic segmentation method for assisted driving provided by an embodiment of the present application. Figure 1 , the semantic segmentation method for assisted driving specifically includes: S110: Obtain an environmental image captured by an assisted driving function in the vehicle, and input the environmental image into a pre-trained semantic segmentation model.

[0021] Specifically, the vehicle's assisted driving function can capture an image of the vehicle's surroundings. This image can be captured during the environmental perception phase of the assisted driving function, depicting the vehicle's surroundings, nearby obstacles, and road conditions. Furthermore, the image can be fed into a pre-trained semantic segmentation model.

[0022] The semantic segmentation model includes an encoder and a decoder. The encoder includes at least two scale convolution groups and a spatial detail module. The semantic segmentation model can be trained using a pre-built sample set, which can include multiple sample images and their corresponding semantic segmentation labels.

[0023] In an embodiment of the present application, the encoder in the semantic segmentation model can be regarded as two branches, namely, a semantic branch and a spatial branch. The semantic branch is composed of at least two scale convolution groups, and the spatial branch is composed of a spatial detail module. The semantic branch is used to extract semantic features from the environmental image, and the spatial branch is used to reuse the multi-level features generated by the semantic branch in the semantic feature extraction process to generate spatial features.

[0024] It should be noted that since the semantic information in the shallow stage is highly similar to the spatial information, the spatial branch does not need to process the environmental image. The spatial detail module can directly use the features of the individual scale convolution groups in the semantic branch, and reduce the complexity of feature extraction by the model through feature reuse, thereby reducing the complexity of semantic segmentation and improving the efficiency of semantic segmentation, making semantic segmentation more suitable for real-time semantic tasks, especially in the scenario of vehicle assisted driving, which can ensure the real-time performance of subsequent assisted driving functions.

[0025] S120 , determining semantic features of the environment image based on the convolution groups of each scale, and extracting spatial features of the environment image from outputs of the convolution groups of each scale based on the spatial detail module.

[0026] In the semantic branch of the encoder, the convolutional groups within it are connected sequentially. Specifically, after entering the semantic segmentation model, the environment image can first enter the first convolutional group, be processed by the first convolutional group, and then continue to the next convolutional group for processing. This process is repeated until the last convolutional group completes processing.

[0027] In a specific embodiment, each scale convolution group is sequentially connected, and the semantic features of the environment image are determined based on each scale convolution group, including: For each scale convolution group, downsampling and feature extraction are performed on the environment image or the output of the previous scale convolution group based on the scale convolution group; the output of the last scale convolution group is determined as the semantic feature of the environment image.

[0028] Among them, the scale convolution group can be used for downsampling and feature extraction. For the first scale convolution group, its input is the environment image. For other scale convolution groups except the first scale convolution group, its input is the output of the previous scale convolution group connected to it.

[0029] Specifically, after the environment image enters the semantic segmentation model, the first scale convolution group can downsample and extract features of the environment image, and then the next scale convolution group connected to the first scale convolution group can downsample and extract features of the output of the first scale convolution group. This process is repeated, so that each intermediate scale convolution group downsamples and extracts features of the output of the previous scale convolution group, until the last scale convolution group downsamples and extracts features of the output of the previous scale convolution group, and the output of the last scale convolution group is used as the semantic feature of the environment image.

[0030] For example, Figure 2 is a schematic diagram of a semantic segmentation model provided in an embodiment of the present application, Figure 2 In the model shown, taking the spatial detail module reusing the outputs of the first and second scale convolution groups as an example, after the environment image is input into the semantic segmentation model, the environment image can first enter each scale convolution group and be processed in turn by each sequentially connected scale convolution group. The obtained semantic features and the spatial features obtained by the spatial detail module are then input into the decoder together, and the decoder obtains the semantic segmentation image. In the above implementation, each scale convolution group performs downsampling and convolution processing, which can reduce the resolution while improving the model inference speed and increasing the number of channels, thereby extracting richer semantic information.

[0031] In the embodiment of the present application, each scale convolution group can be regarded as a convolution stage. The purpose of setting multiple scale convolution groups is to sequentially construct multi-scale features from low to high in the environmental image through multiple scale convolution groups connected in sequence, thereby achieving multi-level feature extraction and ensuring model prediction accuracy. The number of scale convolution groups can be determined based on the resolution of the environmental image collected by the vehicle's assisted driving function, or based on the required accuracy of the image semantic segmentation required by the assisted driving function. For example, the number of scale convolution groups can be 3.

[0032] Specifically, for each scale convolution group, a spatial attention module and a channel attention module can be introduced into the scale convolution group to enrich the constructed features.

[0033] In some embodiments, the scale convolution group includes a first spatial attention module, a downsampling unit, a convolution block, a first splicing unit, and a first channel attention module, and downsampling and feature extraction are performed on the environment image or the output of the previous scale convolution group based on the scale convolution group, including the following steps: Step 11: extracting spatial key information from the environment image or the output of the previous scale convolution group based on the first spatial attention module; Step 12: Downsampling the environment image or the output of the previous scale convolution group based on the downsampling unit to obtain a downsampling result, and convolving the downsampling result based on the convolution block to obtain a convolution result; Step 13: splicing the spatial key information, the convolution result, and the downsampling result based on the first splicing unit to obtain a splicing result; Step 14: Adjust the weights of each channel in the splicing result based on the first channel attention module.

[0034] Among them, for any scale convolution group, it can be composed of a first spatial attention module, a downsampling unit, a convolution block, a first splicing unit and a first channel attention module.

[0035] For example, Figure 3 This is a schematic diagram of a scale convolution group provided in an embodiment of the present application, such as Figure 3 As shown in the figure, in the scaled convolution group, the downsampling unit is connected to the convolution block and the first splicing unit, the convolution block is connected to the first splicing unit, the first spatial attention module is connected to the first splicing unit, and the first splicing unit is connected to the first channel attention module. The input of the scaled convolution group will first pass through the first spatial attention module and the downsampling unit.

[0036] Specifically, the input of the first scale convolution group is the environment image, and the input of the other scale convolution groups except the first scale convolution group is the output of the previous scale convolution group. Therefore, in step 11, the first spatial attention module can extract spatial key information from the input of the scale convolution group, where the input of the scale convolution group is the environment image or the output of the previous scale convolution group.

[0037] Among them, the first spatial attention module can adjust the importance of features from the spatial dimension. For example, the first spatial attention module can identify key spatial positions in the input and suppress irrelevant background to reduce the response to the uninformative area and enhance the representation of objects within the key spatial positions.

[0038] Furthermore, in step 12, the downsampling unit may downsample the input of the scaled convolution group to obtain a downsampled result, where the input of the scaled convolution group is the environment image or the output of the previous scaled convolution group. Furthermore, the convolution block may convolve the downsampled result to obtain a convolution result. The convolution block may be composed of multiple convolutional layers.

[0039] Furthermore, the outputs of the downsampling unit, the convolution block, and the first spatial attention module may enter the first splicing unit. In step 13, the first splicing unit may splice the downsampling results, the convolution results, and the spatial key information to obtain a splicing result.

[0040] Furthermore, the output of the first concatenation unit can enter the first channel attention module. In step 14, the first channel attention module can adjust the channel weights of the concatenation result, that is, adjust the weights of each channel in the concatenation result. The first channel attention module can adjust the weights of each channel based on the contribution of each channel to enhance useful channels, suppress redundant channels, and improve feature discrimination. Through the above-mentioned first spatial attention module, downsampling unit, convolution block, first splicing unit and first channel attention module, it is possible to downsample the input of the scale convolution group, reduce the resolution, improve the inference speed and increase the number of channels, and dynamically adjust the importance of features from the spatial dimension and channel dimension, optimize the feature position and feature type respectively, and enhance the model's ability to capture key information.

[0041] It should be noted that the purpose of multiple scale convolution groups is to extract features from low-level to high-level features from the environment image in sequence. Therefore, the convolution blocks in each scale convolution group can adopt different structures so that each convolution block can accurately construct features of the corresponding scale.

[0042] In one example, the convolution block includes a plurality of standard convolution layers or a plurality of depth residual modules connected in sequence, and the convolution block is used to perform convolution processing on the downsampling result to obtain a convolution result, including: For each standard convolution layer, convolution processing is performed on the down-sampling result or the output of the previous standard convolution layer based on the standard convolution layer; or, for each depth residual module, convolution processing is performed on the down-sampling result or the output of the previous depth residual module based on the depth residual module.

[0043] A convolutional block can be composed of multiple standard convolutional layers connected sequentially, or multiple depth-based residual modules connected sequentially. A convolutional block composed of standard convolutional layers can extract lower-level features, while a convolutional block composed of depth-based residual modules can extract higher-level features.

[0044] For example, assuming that the number of scale convolution groups is 3, the three scale convolution groups are used to construct low-level features, intermediate features and high-level features of the environment image, respectively. Since the first scale convolution group is used to construct low-level features, in order to improve the efficiency of model feature analysis, the convolution block in the first scale convolution group can adopt a standard convolution layer, while the second scale convolution group and the third scale convolution group are used to construct intermediate and high-level features. Therefore, in order to ensure the richness of model feature extraction, the convolution blocks in the second scale convolution group and the third scale convolution group can adopt a deep residual module.

[0045] The number of standard convolutional layers or depth residual modules in a convolutional block can be determined based on the requirements of feature extraction and is not limited in this embodiment. For example, in the first scale convolution group, the number of standard convolutional layers within the convolution block can be 3, in the second scale convolution group, the number of depth residual modules within the convolution block can be 4, and in the third scale convolution group, the number of depth residual modules within the convolution block can be 14.

[0046] By setting a standard convolution layer in the first scale convolution group and a deep residual module in other scale convolution groups, we can ensure the richness of features while ensuring the efficiency of model feature extraction, thereby improving the semantic segmentation accuracy of the model.

[0047] Among them, the depth residual module can include multiple depth-level convolution and residual connection structures, and the depth residual module can be used to perform depth-level asymmetric convolution operations.

[0048] Optionally, the depth residual module includes a first point-by-point convolution unit, a channel grouping unit, a left branch depth convolution unit, a right branch depth convolution unit, a second point-by-point convolution unit, a second splicing unit, a connection unit, and a channel shuffling unit, and performs convolution processing on the downsampling result or the output of the previous depth residual module based on the depth residual module, including the following steps: Step 121: convolve the downsampling result or the output of the previous depth residual module based on the first point-by-point convolution unit, and group the output of the first point-by-point convolution unit along the channel dimension based on the channel grouping unit; Step 122: performing decomposition convolution on the output of the channel grouping unit based on the left depthwise convolution unit, and performing dilation convolution on the output of the channel grouping unit based on the right depthwise convolution unit, wherein the left depthwise convolution unit and the right depthwise convolution unit transmit information to each other during the convolution process; Step 123: splicing the outputs of the left depthwise convolution unit and the right depthwise convolution unit based on the second splicing unit, and convolving the output of the second splicing unit based on the second point-by-point convolution unit; Step 124: Connect the output of the second point-by-point convolution unit with the input of the first point-by-point convolution unit based on the connection unit, and rearrange the output of the connection unit along the channel dimension based on the channel shuffling unit.

[0049] Among them, the depth residual module is composed of a first point-by-point convolution unit, a channel grouping unit, a left-branch depth convolution unit, a right-branch depth convolution unit, a second point-by-point convolution unit, a second splicing unit, a connection unit and a channel shuffling unit.

[0050] like Figure 4 As shown, Figure 4This is a schematic diagram of a depth residual module provided by an embodiment of the present application, wherein the first point-by-point convolution unit is connected to the channel grouping unit, the channel grouping unit is connected to the left branch depth convolution unit and the right branch depth convolution unit, the left branch depth convolution unit and the right branch depth convolution unit are connected to the second splicing unit, the second splicing unit is connected to the second point-by-point convolution unit, the second point-by-point convolution unit and the input of the depth residual module are connected to the connection unit, and the connection unit is connected to the channel shuffling unit.

[0051] Among them, the first point-by-point convolution unit and the second point-by-point convolution unit can use 1×1 convolution kernel to linearly combine the input to avoid multiple superpositions resulting in excessive calculation and parameters. The left branch depth convolution unit and the right branch depth convolution unit can use two sets of convolution to extract features, refer to Figure 4 , the left depth convolution unit can adopt 3×1 and 1×3 decomposition convolution, and the right depth convolution unit can adopt 3×1 and 1×3 expansion convolution to minimize information redundancy.

[0052] Specifically, in step 121, the downsampling result or the output of the previous depth residual module can first pass through the first point-by-point convolution unit after entering the depth residual module. The first point-by-point convolution unit performs a 1×1 convolution on it, and then the output of the first point-by-point convolution unit flows into the channel grouping unit. The channel grouping unit groups the output of the first point-by-point convolution unit along the channel dimension, that is, splits the output channel of the first point-by-point convolution unit into multiple subgroups for subsequent separate processing, so that different branches can extract different features, avoid homogenization of all channels, and achieve the purposes of reducing computational complexity, improving efficiency, and enhancing feature diversity.

[0053] Furthermore, the output of the channel grouping unit can enter the left branch depth convolution unit and the right branch depth convolution unit. In step 122, the left branch depth convolution unit can use two sets of decomposition convolution kernels (without dilation rate) to perform decomposition convolution on the output of the channel grouping unit to extract local features, and the right branch depth convolution unit can use two sets of dilation convolution kernels to perform dilation convolution on the output of the channel grouping unit to increase the receptive field and extract rich context information.

[0054] It should be noted that during the processing of the left and right depth convolution units, information is exchanged between the two depth convolution units to achieve complementarity between different information. Figure 4As shown, after the left branch depth convolution unit uses 3×1 decomposition convolution for processing, the result can be input to the right branch depth convolution unit, so that the right branch depth convolution unit further performs 1×3 dilation convolution based on its own 3×1 dilation convolution result and the 3×1 decomposition convolution result. Moreover, after the right branch depth convolution unit uses 3×1 dilation convolution for processing, the result can be input to the left branch depth convolution unit, so that the left branch depth convolution unit further performs 1×3 decomposition convolution based on its own 3×1 decomposition convolution result and the 3×1 dilation convolution result.

[0055] Furthermore, the outputs of the left depth convolution unit and the right depth convolution unit can enter the second splicing unit. In step 123, the second splicing unit can splice the outputs of the left depth convolution unit and the right depth convolution unit, and then input the result into the second point-by-point convolution unit. The second point-by-point convolution unit can use a 1×1 convolution kernel to convolve the output of the second splicing unit.

[0056] Furthermore, the output of the second point-by-point convolution unit can enter the connection unit. In step 124, the connection unit can connect the output of the second point-by-point convolution unit with the input of the first point-by-point convolution unit, wherein the input of the first point-by-point convolution unit is the input of the depth residual module, that is, the downsampling result or the output of the previous depth residual module. After the connection unit is processed, the channel shuffling unit can rearrange the output of the connection unit along the channel dimension to rearrange the channel order so that the features of different groups can be fully mixed, enhance the information interaction between channels, and thus improve the model expression ability.

[0057] Through the above steps 121-124, the point-by-point convolution unit can reduce the amount of calculation and parameters, and can extract features through the left and right branches respectively, and achieve information complementarity. In addition, the connection unit can be used to connect the ends and construct a residual to prevent gradient shrinkage during training. The channel shuffling unit prevents information closure caused by separation between channels and realizes information interaction between different channels. The above-mentioned deep residual module can effectively extract multi-level deep features of the image, enhance information flow through the asymmetric residual structure, and improve the feature extraction capability of the model.

[0058] In an embodiment of the present application, during feature processing by each scale convolution group, the spatial detail module can extract spatial features from the output of each scale convolution group to extract spatial detail information through feature fusion at different stages or levels.

[0059] For example, taking the number of scale convolution groups as 3, the spatial detail module can fuse the output of the first scale convolution group (i.e., the first stage) with the output of the second scale convolution group (i.e., the second stage) to obtain spatial features.

[0060] The purpose of fusing the first and second stages as spatial features is to achieve higher image resolution, richer spatial details, and more comprehensive extracted content in the first stage. Therefore, the output of the second stage can be upsampled and fused with the first stage. Considering the moderate computational complexity of depth-level convolution, a 1×1 convolution kernel can be used to fully combine the features from both stages, reducing the number of channels to a certain value (e.g., 32). A pyramid structure is constructed using depth-level convolutions containing multiple kernels (e.g., 3×3, 5×5, and 7×7). The output of the upper convolution layer can be used as the input of the lower convolution layer to extract richer feature information. Finally, the outputs of all convolution kernels are summed to output the important spatial features.

[0061] In a specific embodiment, the spatial detail module includes a first fusion unit, multiple convolutional layers, a second fusion unit, and a second spatial attention module. The spatial detail module extracts spatial features of the environment image from the output of each scale convolution group, including the following steps: Step 21: Fuse the outputs of the convolution groups of each scale based on the first fusion unit, and convolve the output of the first fusion unit based on each convolution layer; Step 22: The outputs of each convolutional layer are fused based on the second fusion unit, and the spatial key information in the output of the second fusion unit is extracted based on the second spatial attention module to obtain the spatial features of the environment image.

[0062] The spatial detail module is composed of a first fusion unit, multiple convolutional layers, a second fusion unit, and a second spatial attention module. The first fusion unit can be composed of a splicing layer and a 1×1 convolutional layer.

[0063] Figure 5 is a schematic diagram of a space detail module provided in an embodiment of the present application, such as Figure 5 As shown, the first fusion unit (including splicing and 1×1 convolution) is connected to multiple convolutional layers (3×3, 5×5 and 7×7 convolutions), the upper convolution in each convolutional layer is connected to the lower convolution, and multiple convolutional layers are connected to the second fusion unit, which is connected to the second spatial attention module.

[0064] Specifically, in step 21, the first fusion unit may first select the outputs of at least two scale convolution groups, such as the first scale convolution group and the second scale convolution group, and then upsample the high-level features therein, and then splice them along the channel dimension.

[0065] For example, the output of the first scale convolution group is 67×256×512, and the output of the second scale convolution group is 131×128×256. The first fusion unit can first upsample the output of the second scale convolution group to enlarge it to 131×256×512, and then fuse it with the output of the first scale convolution group. The fused feature size is 198×256×512.

[0066] After the splicing is completed, the first fusion unit can also use the 1×1 convolution kernel to process the splicing result to complete the fusion of the output of convolution groups of different scales. Then the output of the first fusion unit will enter each convolution layer. For the top convolution layer, it performs convolution processing on the output of the first fusion unit. For other convolution layers except the top convolution layer, it adds the output of the first fusion unit and the convolution layer of the previous layer before performing convolution processing.

[0067] Furthermore, in step 22, the output of each convolutional layer can enter the second fusion unit, the second fusion unit can perform element-wise addition on the output of all convolutional layers, and then input it into the second spatial attention module, which can extract spatial key information from the output of the second fusion unit to obtain spatial features.

[0068] Through steps 21 and 22 above, multiple convolutional layers are used for deep convolution. The output of the upper convolution layer serves as the input of the lower convolution layer, which can extract richer feature information. Finally, the output of all convolutional layers is summed element-wise and processed by the spatial attention module to further obtain important pixel features, achieving accurate extraction of spatial features. The spatial detail module extracts spatial detail information at different scales through a pyramid structure, accurately capturing small objects and complex edges in the image, improving segmentation accuracy.

[0069] S130 : Determine a semantic segmentation image corresponding to the environment image based on the decoder, the semantic features, and the spatial features.

[0070] Specifically, after obtaining the semantic features and the spatial features, the semantic features and the spatial features can be further fused through a decoder to obtain a semantic segmentation result of the environment image, that is, a semantic segmentation image.

[0071] For example, Figure 6 This is a schematic diagram of a model processing provided by an embodiment of the present application, such as Figure 6As shown in the figure, the input environment image first enters the semantic branch on the left. The semantic branch consists of three convolutional stages, extracting features at 1 / 2, 1 / 4, and 1 / 8 the original size (of the environment image), respectively. The features obtained in one convolutional stage are ultimately used as semantic features. Furthermore, the spatial branch on the right uses the outputs of the first and second convolutional stages in the semantic branch. The output of the second convolutional stage is bilinearly interpolated (multiplied by 2) and then fused with the output of the first convolutional stage to obtain spatial features. Finally, the semantic and spatial features are fused to output a semantically segmented image. This dual-branch design improves the efficiency and accuracy of feature extraction, enabling simultaneous processing of both semantic and spatial information to optimize segmentation.

[0072] In a specific embodiment, the decoder includes a feature complementation module, a third point-by-point convolution unit, a second channel attention module, and an upsampling unit. Based on the decoder, the semantic features, and the spatial features, determining a semantic segmentation image corresponding to the environment image includes the following steps: Step 31: fusing semantic features and spatial features based on the feature complementation module, and convolving the output of the feature complementation module based on a third point-by-point convolution unit; Step 32: Adjust the weight of each channel in the output of the third point-by-point convolution unit based on the second channel attention module; Step 33: Upsample the output of the second channel attention module based on the upsampling unit to obtain a semantic segmentation image.

[0073] Among them, the decoder consists of a feature complementation module, a third point-by-point convolution unit, a second channel attention module and an upsampling unit. Figure 7 is a schematic diagram of a decoder provided in an embodiment of the present application, such as Figure 7 As shown, the feature complementation module is connected to the third point-by-point convolution unit, the third point-by-point convolution unit is connected to the second channel attention module, and the second channel attention module is connected to the upsampling unit.

[0074] Specifically, in step 31 , the semantic features and the spatial features may first enter a feature complementation module, and the feature complementation module fuses the semantic features with the spatial features.

[0075] In some embodiments, the feature complementation module includes a fourth point-by-point convolution unit, a bilinear interpolation unit, a fifth point-by-point convolution unit, a third channel attention module, a third spatial attention module, a first element-by-element multiplication unit, a second element-by-element multiplication unit, and an element-by-element addition unit. The semantic features and spatial features are fused based on the feature complementation module, including the following steps: Step 311: convolve the spatial features based on the fourth point-by-point convolution unit, and extract spatial key information from the output of the fourth point-by-point convolution unit based on the third spatial attention module; Step 312: performing bilinear interpolation processing on the semantic features based on the bilinear interpolation unit; Step 313: Convolve the output of the bilinear interpolation unit based on the fifth point-by-point convolution unit, and adjust the weight of each channel in the output of the fifth point-by-point convolution unit based on the third channel attention module; Step 314: perform element-wise multiplication on the output of the fourth point-by-point convolution unit and the output of the third channel attention module based on the first element-wise multiplication unit, and perform element-wise multiplication on the output of the fifth point-by-point convolution unit and the output of the third spatial attention module based on the second element-wise multiplication unit; Step 315: Based on the element-adding unit, the outputs of the first element-multiplying unit and the second element-multiplying unit are added element by element to fuse the semantic features with the spatial features.

[0076] Among them, the feature complementation module can be composed of a fourth point-by-point convolution unit, a bilinear interpolation unit, a fifth point-by-point convolution unit, a third channel attention module, a third spatial attention module, a first element multiplication unit, a second element multiplication unit and an element addition unit.

[0077] For example, Figure 8 is a schematic diagram of a feature complementary module provided in an embodiment of the present application, such as Figure 8 As shown, the spatial features can enter the fourth point-by-point convolution unit, the semantic features can enter the bilinear interpolation unit, the fourth point-by-point convolution unit is connected to the third spatial attention module and the first element multiplication unit, the first element multiplication unit is connected to the element addition unit, the third spatial attention module is connected to the second element multiplication unit, the bilinear interpolation unit is connected to the fifth point-by-point convolution unit, the fifth point-by-point convolution unit is connected to the second element multiplication unit and the third channel attention module, the third channel attention module is connected to the first element multiplication unit, and the second element multiplication unit is connected to the element addition unit.

[0078] Specifically, in step 311, the fourth point-by-point convolution unit can perform convolution on the low-level spatial features, and the output after convolution can enter the first element multiplication unit and the third spatial attention module. The third spatial attention module can extract spatial key information from the output of the fourth point-by-point convolution unit, and the output of the third spatial attention module can enter the second element multiplication unit.

[0079] Furthermore, in step 312, the bilinear interpolation unit may perform bilinear interpolation processing on the semantic features to upsample the semantic features to the same resolution as the spatial features, and the output of the bilinear interpolation unit may enter the fifth point-by-point convolution unit.

[0080] Further, in step 313, the fifth point-by-point convolution unit can use a 1×1 convolution kernel to convolve the output of the bilinear interpolation unit, and the output of the fifth point-by-point convolution unit can enter the third channel attention module and the second element multiplication unit. The third channel attention module can adjust the weights of each channel in the output of the fifth point-by-point convolution unit, and the output of the third channel attention module can enter the first element multiplication unit.

[0081] Furthermore, in step 314, the first element-wise multiplication unit may perform element-wise multiplication on the output of the fourth point-wise convolution unit and the output of the third channel attention module, and the second element-wise multiplication unit may perform element-wise multiplication on the output of the fifth point-wise convolution unit and the output of the third spatial attention module. The outputs of the two element-wise multiplication units enter the element-wise addition unit. In step 315, the element-wise addition unit may perform element-wise addition on the outputs of the first element-wise multiplication unit and the second element-wise multiplication unit to achieve the fusion of semantic features and spatial features.

[0082] Through steps 311-315 above, semantic features and spatial features can be processed separately by the spatial attention module and the channel attention module, and two element-wise multiplication units are combined to fuse the semantic features after the spatial features are processed by the spatial attention module, and the spatial features are fused after the semantic features are processed by the channel attention module. Finally, through element-wise unit fusion, information fusion at different scales, i.e., feature complementarity, can be achieved, improving the segmentation performance of the model. The above-mentioned feature complementarity module can optimize feature complementarity by fusing semantic features with spatial features, effectively improving the model's attention to detail, and making semantic segmentation more accurate and comprehensive.

[0083] After the feature complementation module fuses the spatial features with the semantic features, the output of the feature complementation module can enter the third point-by-point convolution unit, and the third point-by-point convolution unit can use a 1×1 convolution kernel to convolve the output of the feature complementation module.

[0084] Furthermore, the output of the third point-by-point convolution unit may enter the second channel attention module. In step 32, the second channel attention module may adjust the weights of each channel in the output of the third point-by-point convolution unit.

[0085] Furthermore, the output of the third channel attention module may enter the upsampling unit. In step 33, the upsampling unit may upsample the output of the second channel attention module (eg, using bilinear interpolation) to obtain a semantic segmentation image.

[0086] Through the above steps 31 to 33, the feature complementarity between semantics and space can be achieved through the feature complementarity module, and the number of channels can be restored through the third point-by-point convolution unit and the second channel attention module to further construct the channel information. Then, the semantic segmentation image is obtained by upsampling to ensure the accuracy of the semantic segmentation result.

[0087] In an embodiment of the present application, the first spatial attention module, the second spatial attention module, and the third spatial attention module can adopt the same structure. The spatial attention module can include a pooling splicing unit, a convolutional layer, and an activation layer, wherein the pooling splicing unit can process the input through maximum pooling (MaxPool) and average pooling (AvgPool), respectively, and then splice the results of the two pooling operations along the channel dimension, that is, it can be expressed as Cat(MaxPool;AvgPool).

[0088] For example, Figure 9 is a schematic diagram of a spatial attention module provided in an embodiment of the present application, such as Figure 9 As shown in the figure, the input of the spatial attention module can first enter the pooling splicing unit, namely Cat (MaxPool; AvgPool), and then enter the convolution layer after being processed by the pooling splicing unit, and realize convolution through the 7×7 convolution kernel. Finally, it enters the activation layer and performs nonlinear transformation through the activation function (such as Sigmoid) to realize the extraction of spatial key information.

[0089] In an embodiment of the present application, the first channel attention module, the second channel attention module, and the third channel attention module can adopt the same structure. The channel attention module may include a maximum pooling layer, an average pooling layer, two convolutional layers, a parameterized rectified linear unit, and a pooling activation layer. The pooling activation layer can be a nonlinear transformation (such as Sigmoid) after adding the results of maximum pooling and average pooling, i.e., Sigmoid (MaxPool + AvgPool).

[0090] For example, Figure 10 is a channel attention module provided by an embodiment of the present application, such as Figure 10 As shown in the figure, the input of the channel attention module can first undergo maximum pooling and average pooling, and then enter the first convolutional layer for 1×1 convolution, then be processed by the Parametric Rectified Linear Unit (PReLU), and enter the second convolutional layer for 1×1 convolution, and finally enter the pooling activation layer for nonlinear transformation to achieve dynamic adjustment of the weights of each channel.

[0091] The embodiment of the present application provides a semantic segmentation method for assisted driving. The method obtains an environmental image collected by the assisted driving function in the vehicle, and inputs the environmental image into a pre-trained semantic segmentation model, and then determines the semantic features of the environmental image according to the convolution groups of each scale in the model, and extracts the spatial features of the environmental image from the output of the convolution groups of each scale according to the spatial detail module in the model. Finally, the semantic segmentation image corresponding to the environmental image is determined according to the decoder and the semantic features and spatial features in the model, thereby realizing the semantic segmentation of the environmental image collected by the assisted driving function. The method extracts semantic features by the convolution groups of each scale, which can effectively extract the multi-level features of the image, enhance the information flow, and improve the semantic segmentation of the environmental image. The feature extraction capability of the model is improved. In addition, the method determines the spatial features through the output of the convolution groups of each scale, realizes the feature reuse in the process of semantic information processing, does not need to process the environmental image separately, reduces additional redundant calculations, improves the efficiency of semantic segmentation, and solves the problem of high complexity of semantic segmentation in related technologies. In addition, through spatial features, it can accurately capture small objects and complex edges in the image, improve segmentation accuracy, and solve the problems of poor detail perception, low robustness and low boundary segmentation accuracy in related technologies. The method also integrates spatial features and semantic features to obtain the final segmentation result, realizes feature complementarity, effectively improves the model's attention to details, and makes semantic segmentation more accurate and comprehensive.

[0092] The semantic segmentation method for assisted driving provided in the embodiments of the present application has the following technical effects: 1. Small number of model parameters: By optimizing the network structure and adopting efficient model design, the number of model parameters is significantly reduced. This improvement not only reduces the model's storage requirements but also reduces the consumption of computing resources. This enables the technology to run efficiently on devices with limited resources. It is particularly suitable for embedded systems and mobile devices, reducing hardware costs and improving deployment flexibility. 2. Good real-time performance: This method can optimize the model inference speed while ensuring segmentation accuracy. By streamlining the network layer and adopting efficient convolution operations and inference strategies, the processing speed of real-time semantic segmentation is greatly improved. This enables the technology to be applied in scenarios with high real-time requirements, ensuring that the system can quickly respond to dynamically changing environments and provide real-time decision support. 3. Good segmentation accuracy: This method improves segmentation accuracy by combining feature extraction and fusion techniques. Compared with related technologies, this method demonstrates higher accuracy in segmenting complex backgrounds, boundaries between different objects, and details in images. Through targeted optimization, the model can better handle complex scenes, improving the overall effect in application scenarios, especially in the field of autonomous driving (it can also be used in other fields), ensuring more accurate object recognition and region segmentation. 4. More accurate segmentation of small objects and edges: This method demonstrates significant advantages in segmenting small objects and edge areas. Through multi-scale feature extraction and refined edge processing strategies, the model can more accurately identify and segment small objects and complex edges. It is particularly suitable for application scenarios such as pedestrian and obstacle detection in autonomous driving systems, effectively improving the system's performance under these high-precision requirements.

[0093] In summary, the method provided in the embodiments of the present application can improve segmentation accuracy, increase inference speed, and realize refined segmentation processing. At the same time, due to the small number of parameters and fast inference speed, it can operate efficiently under limited device resources and has strong adaptability and flexibility; combined with its efficient edge processing and small object recognition capabilities, it brings higher potential for real-time semantic segmentation technology and has broad application prospects.

[0094] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 11 As shown, the electronic device 400 includes one or more processors 401 and a memory 402 .

[0095] The processor 401 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.

[0096] The memory 402 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may execute the program instructions to implement the semantic segmentation method for assisted driving and / or other desired functions of any embodiment of the present application described above. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage medium.

[0097] In one example, electronic device 400 may further include an input device 403 and an output device 404, which are interconnected via a bus system and / or other connection mechanisms (not shown). Input device 403 may include, for example, a keyboard, a mouse, etc. Output device 404 may output various information to the outside, including warning information, braking force, etc. Output device 404 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.

[0098] Of course, to simplify, Figure 11 Only some of the components related to the present application in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 400 may further include any other appropriate components according to specific application scenarios.

[0099] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the semantic segmentation method for assisted driving provided in any embodiment of the present application.

[0100] The computer program product may be written in any combination of one or more programming languages ​​to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0101] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, causes the processor to execute the steps of the semantic segmentation method for assisted driving provided in any embodiment of the present application.

[0102] The computer-readable storage medium may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0103] It should be noted that the terms used in this application are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates an exception, the words "one", "an", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.

[0104] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0105] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. The above is only the preferred implementation method of this application. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of this application, they can also make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of this application.

Claims

1. A semantic segmentation method for assisted driving, characterized in that: include: Obtaining an environmental image captured by an assisted driving function in a vehicle, and inputting the environmental image into a pre-trained semantic segmentation model, wherein the semantic segmentation model includes an encoder and a decoder, and the encoder includes at least two scale convolution groups and a spatial detail module; Determining semantic features of the environment image based on each of the scale convolution groups, and extracting spatial features of the environment image from outputs of each of the scale convolution groups based on the spatial detail module; A semantic segmentation image corresponding to the environment image is determined based on the decoder, the semantic features, and the spatial features.

2. The method according to claim 1, characterized in that The scale convolution groups are sequentially connected, and the semantic features of the environment image are determined based on the scale convolution groups, including: For each of the scale convolution groups, downsampling and feature extraction are performed on the environment image or the output of the previous scale convolution group based on the scale convolution group; The output of the last scale convolution group is determined as the semantic feature of the environment image.

3. The method according to claim 2, characterized in that The scale convolution group includes a first spatial attention module, a downsampling unit, a convolution block, a first splicing unit, and a first channel attention module, and performs downsampling and feature extraction on the environment image or the output of the previous scale convolution group based on the scale convolution group, including: Extracting spatial key information from the environment image or the output of the previous scale convolution group based on the first spatial attention module; Performing downsampling processing on the environment image or the output of the previous scale convolution group based on the downsampling unit to obtain a downsampling result, and performing convolution processing on the downsampling result based on the convolution block to obtain a convolution result; splicing the spatial key information, the convolution result, and the downsampling result based on the first splicing unit to obtain a splicing result; The weight of each channel in the splicing result is adjusted based on the first channel attention module.

4. The method according to claim 3, characterized in that The convolution block includes a plurality of standard convolution layers or a plurality of depth residual modules connected in sequence, and the convolution block performs convolution processing on the downsampling result to obtain a convolution result, including: For each of the standard convolutional layers, performing convolution processing on the downsampling result or the output of the previous standard convolutional layer based on the standard convolutional layer; Alternatively, for each of the depth residual modules, convolution processing is performed on the downsampling result or the output of the previous depth residual module based on the depth residual module.

5. The method according to claim 4, characterized in that The depth residual module includes a first point-by-point convolution unit, a channel grouping unit, a left branch depth convolution unit, a right branch depth convolution unit, a second point-by-point convolution unit, a second splicing unit, a connection unit, and a channel shuffling unit. The depth residual module performs convolution processing on the downsampling result or the output of the previous depth residual module based on the depth residual module, including: Convolving the downsampling result or the output of the previous depth residual module based on the first point-by-point convolution unit, and grouping the output of the first point-by-point convolution unit along the channel dimension based on the channel grouping unit; Performing decomposition convolution on the output of the channel grouping unit based on the left depth convolution unit, and performing dilation convolution on the output of the channel grouping unit based on the right depth convolution unit, wherein the left depth convolution unit and the right depth convolution unit transmit information to each other during the convolution process; splicing the outputs of the left depth convolution unit and the right depth convolution unit based on the second splicing unit, and convolving the output of the second splicing unit based on the second point-by-point convolution unit; The output of the second point-by-point convolution unit is connected to the input of the first point-by-point convolution unit based on the connection unit, and the output of the connection unit is rearranged along the channel dimension based on the channel shuffling unit.

6. The method according to claim 1, characterized in that The spatial detail module includes a first fusion unit, a plurality of convolutional layers, a second fusion unit, and a second spatial attention module. The spatial features of the environment image are extracted from the outputs of each of the scale convolution groups based on the spatial detail module, including: fusing the outputs of the scale convolution groups based on the first fusion unit, and convolving the outputs of the first fusion unit based on the convolution layers; The outputs of the convolutional layers are fused based on the second fusion unit, and spatial key information in the output of the second fusion unit is extracted based on the second spatial attention module to obtain the spatial features of the environmental image.

7. The method according to claim 1, characterized in that The decoder includes a feature complementation module, a third point-by-point convolution unit, a second channel attention module, and an upsampling unit. Based on the decoder, the semantic features, and the spatial features, determining a semantic segmentation image corresponding to the environment image includes: fusing the semantic features and the spatial features based on the feature complementation module, and convolving the output of the feature complementation module based on the third point-by-point convolution unit; Adjusting the weight of each channel in the output of the third point-by-point convolution unit based on the second channel attention module; The output of the second channel attention module is upsampled based on the upsampling unit to obtain a semantic segmentation image.

8. The method according to claim 7, characterized in that The feature complementation module includes a fourth point-by-point convolution unit, a bilinear interpolation unit, a fifth point-by-point convolution unit, a third channel attention module, a third spatial attention module, a first element multiplication unit, a second element multiplication unit, and an element addition unit. The semantic feature and the spatial feature are fused based on the feature complementation module, including: Convolving the spatial features based on the fourth point-by-point convolution unit, and extracting spatial key information from the output of the fourth point-by-point convolution unit based on the third spatial attention module; Performing bilinear interpolation processing on the semantic features based on the bilinear interpolation unit; Convolving the output of the bilinear interpolation unit based on the fifth point-by-point convolution unit, and adjusting the weight of each channel in the output of the fifth point-by-point convolution unit based on the third channel attention module; Performing element-by-element multiplication on the output of the fourth point-by-point convolution unit and the third channel attention module based on the first element multiplication unit, and performing element-by-element multiplication on the output of the fifth point-by-point convolution unit and the third spatial attention module based on the second element multiplication unit; Based on the element-adding unit, outputs of the first element-multiplying unit and the second element-multiplying unit are added element by element to fuse the semantic features with the spatial features.

9. An electronic device, characterized in that: The electronic device comprises: processor and memory; The processor is configured to execute the steps of the semantic segmentation method for assisted driving as described in any one of claims 1 to 8 by calling the program or instructions stored in the memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, which enables a computer to execute the steps of the semantic segmentation method for assisted driving as described in any one of claims 1 to 8.