Domain generalization semantic segmentation method, electronic equipment and storage medium
By extracting and stylizing the low-level features of three-channel color images and depth source domain images, combining class-level space soft sensitivity suppression and soft alignment loss function optimization, the problem of insufficient generalization ability of the semantic segmentation model in a variety of complex road scenarios is solved, and a more efficient semantic segmentation effect is achieved.
Patent Information
- Application Number
- CN202510250867.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-29
AI Technical Summary
The existing semantic segmentation models have poor generalization capabilities in a variety of complex road scenarios, especially in the high sensitivity suppression strategy and channel-level processing of depth images, such as information distortion, loss of key information and insufficient capture of sensitive areas.
By extracting the low-level features of the three-channel color images and depth source domain images, stylized processing and class-level spatial soft-sensitive suppression are performed, and the semantic segmentation model is optimized in combination with the soft alignment loss function to enhance the robustness and generalization performance of the depth features.
The generalization ability of semantic segmentation models in diverse and complex road scenarios is improved, ensuring high segmentation performance under low light conditions, accurately capturing category and spatial sensitivity characteristics, reducing noise impact, and enhancing the robustness of the model.
Smart Images

Figure CN120388376A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic segmentation, and in particular, to a domain generalization semantic segmentation method, an electronic device, and a storage medium. Background Art
[0002] Semantic segmentation technology aims to classify each pixel in an image. With the rapid development of deep neural networks, significant breakthroughs and progress have been made in this field. However, the current field still faces two key challenges. First, the performance improvement of these methods highly depends on high-quality labeled datasets, and creating such datasets is an extremely laborious and time-consuming process.
[0003] In the prior art, domain generalization technology (DG) is one of the effective methods to solve these two problems. In the DG task setting, the model can be well generalized to any unknown domain by using source domain data for training. In existing research, the assistance of depth information has been proven to be crucial for effectively learning domain-invariant representations. However, it is rarely used in the field of domain generalization. Among the methods in the field of domain generalization, the RGB-based high-sensitivity region suppression strategy enables the model to focus on weakly sensitive regions containing a large number of domain-invariant features, thereby improving the generalization performance. The RGB-based high-sensitivity region suppression strategy consists of two main parts: sensitivity detection and high-sensitivity suppression. However, the high-sensitivity suppression strategy for depth maps has not been studied, and in the sensitivity detection step, applying common enhancement strategies to depth images may cause distortion of depth information or loss of key information; the existing channel-based high-sensitivity suppression strategy can also be regarded as a coarse sensitivity suppression, which cannot capture the sensitivity specific to classes and spatial structures. At the same time, due to its binary operation, the existing sensitive hard suppression is difficult to capture the continuous changes and detailed information of sensitive regions, resulting in poor generalization of the model in image semantic segmentation and unable to maintain high segmentation performance in diverse road scenes.
[0004] Therefore, how to improve the cross-scene generalization ability of the semantic segmentation model in diverse and complex road scenes has become a core problem that urgently needs to be solved. Summary of the Invention
[0005] The present invention provides a domain generalization semantic segmentation method, an electronic device, and a storage medium, and its main purpose is to solve the problem of poor generalization ability of the semantic segmentation model in diverse and complex road scenes.
[0006] To achieve the above object, a domain generalization semantic segmentation method provided by the present invention includes: obtaining a three-channel color image of an image to be trained and a depth source domain image, and extracting a first low-level image feature of the three-channel color image and a second low-level image feature of the depth source domain image; performing a stylization process on the second low-level image feature according to the first low-level image feature and the second low-level image feature to obtain a stylized depth feature; performing a class-level spatial soft sensitivity suppression on the second low-level image feature according to the stylized depth feature to obtain an enhanced depth feature; optimizing a pre-constructed semantic segmentation model according to the enhanced depth feature and the stylized depth feature to obtain a target semantic segmentation model; and performing semantic segmentation on a target image by using the target semantic segmentation model to obtain a semantic segmentation result.
[0007] The present invention also provides a domain generalization semantic segmentation device, including: a low-level image feature extraction module, configured to obtain a three-channel color image of an image to be trained and a depth source domain image, and extract a first low-level image feature of the three-channel color image and a second low-level image feature of the depth source domain image; a feature stylization processing module, configured to perform a stylization process on the second low-level image feature according to the first low-level image feature and the second low-level image feature to obtain a stylized depth feature; a class-level spatial soft sensitivity suppression module, configured to perform a class-level spatial soft sensitivity suppression on the second low-level image feature according to the stylized depth feature to obtain an enhanced depth feature; a model optimization module, configured to optimize a pre-constructed semantic segmentation model according to the enhanced depth feature and the stylized depth feature to obtain a target semantic segmentation model; and an image semantic segmentation module, configured to perform semantic segmentation on a target image by using the target semantic segmentation model to obtain a semantic segmentation result.
[0008] The present invention also provides an electronic device, including: a memory communicatively connected to at least one processor; wherein, the processor is configured to execute a computer program stored on the memory; the memory stores a computer program executable by at least one processor, and when the computer program is executed by at least one processor, at least one processor is enabled to execute the above-mentioned domain generalization semantic segmentation method.
[0009] The present invention also provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the above-mentioned domain generalization semantic segmentation method is implemented.
[0010] In the embodiments of the present invention, by extracting the low-level image features of the three-channel color image and the depth source domain image, rich image information is provided; by performing stylization processing on the low-level image features of the depth source domain image, the stylized depth source domain image can retain part of the RGB information and retain the basic features, effectively improving the accuracy of subsequent model optimization; according to the stylized depth features, class-level spatial soft-sensitive suppression is performed on the second low-level image features, and enhanced depth features with improved generalization performance and greater robustness can be obtained, which is beneficial to identifying and enhancing the insensitive regions in the depth map containing domain-invariant features of each class. Then, using the enhanced depth features and the stylized depth features to optimize the pre-constructed semantic segmentation model, a target semantic segmentation model can be obtained, which can accurately capture the category sensitivity and spatial sensitivity features, while maximizing the insensitive features of each category, so as to learn more domain-invariant features to improve the model performance, and then achieve accurate semantic segmentation. Therefore, a domain generalization semantic segmentation method, an electronic device, and a storage medium proposed by the present invention can solve the problem that the generalization of image semantic segmentation is poor, which in turn leads to poor generalization ability of the semantic segmentation model in diverse and complex road scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic flowchart of the domain generalization semantic segmentation method provided by an embodiment of the present invention;
[0012] Figure 2 It is a schematic flowchart of performing stylization processing on the low-level image features of the depth source domain image provided by an embodiment of the present invention;
[0013] Figure 3 It is a schematic flowchart of performing class-level spatial soft-sensitive suppression on the second low-level image features provided by an embodiment of the present invention;
[0014] Figure 4 It is a schematic structural diagram of the domain generalization semantic segmentation method provided by an embodiment of the present invention;
[0015] Figure 5 It is a functional module diagram of the domain generalization semantic segmentation device provided by an embodiment of the present invention;
[0016] Figure 6 It is a schematic structural diagram of an electronic device for implementing the domain generalization semantic segmentation method provided by an embodiment of the present invention.
[0017] The realization, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] The embodiments of the present application provide a domain generalization semantic segmentation method. The execution subject of a domain generalization semantic segmentation method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, a domain generalization semantic segmentation method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0020] In the prior art, domain generalization semantic segmentation (DGSS) methods based on RGB-D in various complex road environments have not been fully explored. In some Unsupervised Domain Adaptation (UDA) methods, the effectiveness of depth maps that reflect the shape and spatial position of objects has been verified. Some works use depth estimation tasks to infer depth maps from color images, while other works innovatively input depth maps and RGB data into the model simultaneously. Although the ways they obtain and utilize depth maps are different, they all utilize the conclusion that depth maps are a kind of domain-invariant feature.
[0021] In related visualizations, depth images more clearly present the shape and spatial position of objects, with less texture interference, highlighting that depth maps contain rich domain-invariant information. Existing research work focusing on the RGB-D domain generalization semantic segmentation task is limited to two categories, "road" and "background", and cannot handle more realistic complex traffic conditions. In addition, existing research relies on the fusion of surface normal maps and requires additional and time-consuming processing using depth maps. Therefore, it is necessary and urgent to study the RGB-D domain generalization semantic segmentation task in multi-class complex scenarios.
[0022] Secondly, depth maps are not always perfect, especially in the real world. Different from the above RGB-D methods that roughly regard depth information as reliable domain-invariant information in cross-domain learning tasks, depth maps cannot fully represent domain-invariant information. For example, holes and errors caused by device accuracy and the environment still have some noise, and depth maps in the real world show fragmentation and discontinuity.
[0023] Therefore, how to reduce these noises and learn reliable domain-invariant information from depth maps has become the focus of attention.
[0024] The RGB-based high-sensitivity region suppression strategy enables the model to focus on weakly sensitive regions containing a large number of domain-invariant features, thereby improving generalization performance. This technique consists of two main parts: sensitivity detection and high-sensitivity suppression.
[0025] However, the high-sensitivity suppression strategy for depth maps has not been studied. Meanwhile, due to:
[0026] 1. Given that the geometric information in depth images mainly captures object distances and spatial relationships, while simple blurring and photometric transformations mainly aim to change colors and brightness. Therefore, using common data augmentation strategies in depth maps may lead to information distortion or loss of key information;
[0027] 2. Each channel in RGB images contains different color information, allowing for effective processing at the channel level. However, depth images have only one channel, and this channel mainly contains distance information, lacking the richness of color features. Instead, the spatial structure containing geometric shapes and the relative positions of objects plays a more crucial role. Moreover, aggregating information from surrounding pixels helps reduce the impact of noise. In addition, the feature distributions of each class are usually different, and the sensitivity to differences varies by class. Common sensitivity suppression is usually global, which may lead to the omission of sensitive features in certain classes;
[0028] 3. Existing sensitive hard suppression is difficult to capture the continuous changes and detailed information of sensitive regions due to its binary operation. High-sensitivity hard suppression uses a sensitivity threshold to distinguish sensitive regions from non-sensitive regions, and then binaryizes the sensitive regions to 0 or 1. This extreme operation ignores the gradual change of sensitivity across spatial regions, resulting in information loss, especially at the boundary between high-sensitivity and non-sensitive regions, where artifacts or distortions are easily introduced, and the detailed information of the transition region cannot be accurately captured. In addition, this method lacks flexibility and is difficult to adapt to diverse data distributions.
[0029] Therefore, based on the above analysis, this application proposes a high-sensitivity soft suppression method for depth maps for multi-class domain generalization semantic segmentation. First, the inter-modal stylization flow is used to increase the diversity of depth map data for sensitivity detection, enabling the stylization of depth features without an auxiliary dataset, significantly improving the processing efficiency and style controllability.
[0030] Secondly, to address the problem that existing methods cannot accurately capture class sensitivity and spatial sensitivity features, a class-level spatial soft sensitivity suppression strategy is proposed. This strategy maximally extracts the insensitive features of each class, thereby learning more domain-invariant features to improve the model performance. Finally, the model is optimized through a soft alignment loss function to ensure that the stylized depth map can retain part of the information of the RGB features while preserving the inherent properties of the original depth map, enabling the target semantic segmentation model to have strong robustness and excellent generalization ability even under low-light conditions, effectively improving the generalization ability of the semantic segmentation model in diverse and complex road scenes.
[0031] Referring to Figure 1 The flowchart of a domain generalization semantic segmentation method provided by an embodiment of the present invention is shown as follows, which specifically includes the following processes:
[0032] S1. Obtain the three-channel color image of the image to be trained and the depth source domain image, and extract the low-level image features of the three-channel color image and the depth source domain image.
[0033] In one embodiment, the image to be trained refers to images in different scenarios. For example, it can be real-world images or synthetic images, mainly including different urban road scene images. At the same time, the RGB source domain image of each image is collected as the three-channel color image, and the depth image is collected as the depth source domain image. Among them, the three-channel color image refers to an image that uses three color channels of red (Red), green (Green), and blue (Blue) to represent the color information of the image, and the depth source domain image refers to an image that uses the distance (depth) values of each point in the scene collected by the image collector as pixel values, directly reflecting the geometric shape of the visible surface of the scene. Through the three-channel color image and the depth source domain image, rich image information can be provided, thereby improving the accuracy of subsequent semantic segmentation.
[0034] In one embodiment, extracting the first low-level image feature of the three-channel color image and the second low-level image feature of the depth source domain image includes: respectively performing convolution processing on the three-channel color image and the depth source domain image to obtain convolution features; performing pooling processing on the convolution features to obtain the first low-level image feature of the three-channel color image and the corresponding second low-level image feature of the depth source domain image.
[0035] In one embodiment, the low-level features of an image refer to the basic information extracted from the image, usually including features such as color, texture, and edge that describe the basic attributes of the image, and can be extracted through a pre-constructed low-level feature extraction module. Among them, the low-level feature extraction module includes a convolution layer and a pooling layer for obtaining low-level image features.
[0036] S2. Perform stylization processing on the low-level features of the second image according to the low-level features of the first image and the low-level features of the second image to obtain stylized depth features.
[0037] In one embodiment, the stylization processing combines the semantic information of a three-channel color image and the spatial information of a depth source domain image, thereby stylizing the depth source domain image and generating diverse training data to help the model learn domain-invariant features.
[0038] Specifically, performing stylization processing on the low-level features of the second image according to the low-level features of the first image and the low-level features of the second image to obtain stylized depth features includes: calculating the variance of the inter-modal flow and the mean of the inter-modal flow according to the low-level features of the first image and the low-level features of the second image; calculating the stylized variance and the stylized mean according to the variance of the inter-modal flow and the mean of the inter-modal flow; and performing stylization processing on the low-level features of the second image according to the stylized variance and the stylized mean to obtain stylized depth features.
[0039] In one embodiment, a preset first formula is used to calculate the variance of the inter-modal flow and the mean of the inter-modal flow:
[0040]
[0041] Wherein, represents the mean of the inter-modal flow, μ(·) represents the mean, C represents a matrix of a preset size, and Z d represents the low-level features of the second image, represents the variance of the inter-modal flow, σ(·) represents the variance, and Z rgb represents the low-level features of the first image.
[0042] Specifically, the variance of the inter-modal flow refers to the variance of the elements in the feature map when the feature maps between different modalities in multi-modal data interact through a certain flow or mapping relationship. In this application, it is expressed as the difference in variance of the elements between the cropped image low-level features of a preset size matrix and the original image low-level features. Similarly, the mean of the inter-modal flow refers to the difference in mean of the elements between the cropped image low-level features and the original image low-level features.
[0043] Specifically, a random cropping operation is performed on the image low-level features using a matrix of size 64×64 to increase the diversity of the flow as much as possible. Then, the variance and the mean are calculated respectively between the cropped image low-level features and the pixel point feature values of the original second image low-level features, and then corresponding subtraction operations are performed to obtain the variance of the inter-modal flow and the mean of the inter-modal flow, which can increase the diversity of the flow as much as possible.
[0044] Furthermore, a preset second formula is used to calculate the stylized variance and the stylized mean:
[0045]
[0046] Among them, μ(Z d-style ) represents the stylized mean, μ(·) represents the mean, and Z d represents the low-level features of the second image, λ ∈ [0, 1], represents the mean of the inter-modal flow, and σ(Z d-style ) represents the stylized variance, and σ(·) represents the variance. represents the variance of the inter-modal flow.
[0047] In one embodiment, λ ∈ [0, 1] can be used to control the degree of depth map stylization. The low-level features of the second image are stylized through the stylized variance, the stylized mean, and the normalized low-level features of the second image to obtain stylized depth features.
[0048] Specifically, the low-level features of the second image are stylized using a preset stylization formula:
[0049] Z d-style = σ(Z d-style )Z d-nor + μ(Z d-style )
[0050] Among them, Z d-style represents the stylized depth features, μ(Z d-style ) represents the stylized mean, σ(Z d-style ) represents the stylized variance, and Z d-nor represents the low-level features of the second image after normalization processing.
[0051] In one embodiment, considering the noise in the depth map, the RGB image can provide or supplement the information that cannot be captured or lost in the depth source domain image, such as detailed textures, object edges, etc. Therefore, supplementing the depth map with RGB information during the image stylization process can complement the information that cannot be captured or lost in the depth source domain image.
[0052] Specifically, the process of stylizing the low-level features of the second image can be as Figure 2 shown, where represents element-wise multiplication, represents element-wise addition, represents element-wise subtraction.
[0053] In one embodiment, by stylizing the low-level features of the depth source domain image, the RGB information is used to stylize the depth source domain image for sensitivity detection, and the stylized depth source domain image retains part of the RGB information and retains its basic features, effectively improving the accuracy of subsequent model optimization.
[0054] S3. Perform class-level spatial soft sensitivity suppression on the low-level features of the second image according to the stylized depth features to obtain enhanced depth features.
[0055] In one embodiment, the stylized depth features not only obtain certain characteristics of the RGB features, but also retain the original attributes of the low-level features of the depth source domain image. Feature sensitivity is usually calculated from the channel dimension. However, for single-channel depth images, the information contained in the channel dimension is relatively limited compared to the spatial dimension.
[0056] More importantly, sensitivity calculation is often based on global features, which may cause sensitive features of certain specific categories to be ignored. In addition, due to the binarization operation of the high-sensitivity hard suppression strategy, it is difficult to capture the continuous changes and detailed information in the sensitive area, and it lacks flexibility in complex scenarios.
[0057] Therefore, in order to more accurately identify and enhance the weakly sensitive areas in the depth features that contain more domain-invariant features, by performing class-level spatial soft sensitivity suppression on the low-level features of the second image to obtain enhanced depth features, it can effectively weaken the interference of the sensitive part on the invariant features, thereby improving the multi-modal expression ability of the enhanced depth features.
[0058] In one embodiment, performing class-level spatial soft sensitivity suppression on the low-level features of the second image according to the stylized depth features to obtain enhanced depth features includes: calculating the depth feature difference corresponding to the low-level features of the second image according to the stylized depth features; calculating the class-level spatial sensitivity matrix according to the depth feature difference, and calculating the weight of the insensitive features according to the class-level spatial sensitivity matrix; performing feature enhancement on the low-level features of the second image according to the weight of the insensitive features to obtain enhanced depth features.
[0059] Specifically, calculate the absolute value of the difference between the low-level features of the depth source domain image and the stylized depth features to obtain the feature difference, which is used to reflect the difference between the original depth features and the stylized depth features, and can be calculated using the following preset feature difference formula:
[0060] Z d-diff = |Z d - Z d-style |
[0061] Where, Z d-diff represents the depth feature difference, Z d represents the low-level features of the second image, and Z d-style represents the stylized depth features.
[0062] In one embodiment, calculate the class-level spatial sensitivity matrix using the following sensitivity matrix calculation formula:
[0063]
[0064] Among them, represents the spatial sensitivity matrix of the k-th preset category in semantic segmentation, Mean represents the mean operation, and Softmax represents the Softmax activation function. represents the feature corresponding to the k-th preset category in the differential features.
[0065] Calculate the weight of the insensitive feature using the following weight formula for the insensitive feature:
[0066]
[0067] N g = 1 - S g
[0068] Among them, Sigmoid represents the Sigmoid activation function, Conv 1×1 represents a 1×1 convolution operation. represents the spatial sensitivity matrix of the k-th preset category in semantic segmentation, K represents the total number of preset categories in semantic segmentation, and N g represents the weight of the insensitive feature.
[0069] Among them, the preset category can be the label corresponding to a certain channel or region in the feature map and a specific category. For example, a region in the feature map has a high value, indicating that this region may contain an object belonging to the k-th class (such as a car). Each category represents a number, indicating that a car is the 0th class and a pedestrian is the 1st class during semantic segmentation.
[0070] In one embodiment, perform feature enhancement on the low-level features of the second image using a preset feature enhancement formula:
[0071] Z d-fine = Z d ⊙ N g + Z d
[0072] Among them, Z d-fine represents the enhanced depth feature, Z d represents the low-level features of the second image, N g represents the weight of the insensitive feature, and ⊙ represents element-wise multiplication.
[0073] Furthermore, the process schematic diagram of the class-level spatial soft-sensitive suppression strategy is as Figure 3 shown. In the figure, the class-level spatial sensitivity matrix uses the feature map corresponding to the k-th class label, and S g represents the depth-sensitive feature map, and with the aid of the preset label Y xExtract the features corresponding to the k-th type of label from the differential features respectively where k ∈ K represents the k-th feature label, B represents the batch dimension, C represents the channel dimension, H represents the image height, and W represents the image width. Secondly, normalize the spatial elements on each channel of each class of features through the Softmax operation to ensure that the sum of the pixel probabilities of this class in this channel is 1. This normalization process helps to capture the importance of each category in the spatial dimension and convert it into a probability weight. Then, obtain the class-level spatial sensitivity matrix by performing a mean operation on the channels After summing all category sensitivity matrices pixel by pixel, apply a 1×1 convolution and a Sigmoid activation function to normalize again to obtain the final global effective sensitivity matrix The more sensitive the feature, the higher its corresponding weight. Conversely, the less sensitive the feature, the lower the corresponding weight. Subsequently, by subtractingfrom the full tensor Subtract The weight of the insensitive feature can be obtained
[0074] Specifically, the insensitive features containing more domain-invariant information are selectively enhanced by element-wise multiplication, while the element-wise addition retains the overall architecture and intrinsic properties of the low-level image features of the depth source domain image, thereby directly enhancing the insensitive features to obtain enhanced depth features
[0075] In one embodiment, through class-level spatial soft sensitivity suppression, robust multi-modal features with improved generalization performance and greater robustness can be obtained, which is beneficial to identifying and emphasizing the insensitive regions in the depth map containing domain-invariant features of each class, and improving the generalization ability of subsequent semantic segmentation
[0076] S4. Optimize the pre-constructed semantic segmentation model according to the enhanced depth features and the stylized depth features to obtain the target semantic segmentation model
[0077] In one embodiment, the model optimization is to optimize the model for image semantic segmentation. Among them, the semantic segmentation model includes an image low-level feature extraction module, an encoding module, and a decoding module, as specifically shown in Figure 4 As shown Figure 4 where the solid line represents the process of model optimization, and the dashed line represents the process of the target semantic segmentation model performing image semantic segmentation after the model optimization is completed
[0078] Specifically, the pre-built semantic segmentation model is optimized according to the enhanced depth feature and the stylized depth feature to obtain the target semantic segmentation model, including: calculating the soft alignment loss according to the stylized depth feature; fusing the enhanced depth feature and the first image low-level feature to obtain a robust multi-modal feature; encoding the robust multi-modal feature by using the encoding module of the semantic segmentation model to obtain an encoded feature; decoding and predicting the encoded feature to obtain a prediction result, and calculating the cross-entropy semantic segmentation loss of the robust multi-modal feature according to the semantic prediction result; optimizing the parameters of the semantic segmentation model according to the loss function values of the soft alignment loss and the cross-entropy semantic segmentation loss until the loss function value is less than a preset loss threshold to obtain the target semantic segmentation model.
[0079] In one embodiment, the soft alignment loss is calculated by using the following preset soft alignment loss function:
[0080]
[0081] Wherein, represents the soft alignment loss, γ represents a preset scaling factor, N represents the total number of pixel points in the stylized depth feature, represents the feature value of the i-th pixel point in the first image low-level feature, represents the feature value of the i-th pixel point in the stylized depth feature.
[0082] As Figure 4 , the encoding module in the semantic segmentation model is composed of multiple convolutional layers with different convolutional scales and is used to extract semantic features of different scales. The decoding (Decoder) module can restore the encoded robust multi-modal feature to the original data or generate a new data form, that is, the encoded robust multi-modal feature can be restored to the prediction result of semantic segmentation. For example, the decoding module can include a fully connected layer for predicting the prediction result corresponding to the robust multi-modal feature.
[0083] Furthermore, fusing the first image low-level feature and the enhanced depth feature of the three-channel color image can obtain a more robust multi-modal feature that helps improve the generalization performance. Specifically, it is shown as the following formula:
[0084] Z rgb-d = Z d-fine ⊙ Z rgb + Z rgb
[0085] Wherein, Z rgb-d represents the robust multi-modal feature, Z d-fine represents the enhanced depth feature, ⊙ represents element-wise multiplication, and Z rgb represents the first image low-level feature.
[0086] Further, Figure 4 in which, IMSF represents stylization processing, CSSS represents class-level spatial soft sensitivity suppression strategy, and X rgb represents a three-channel color image, and X d represents a depth source domain image, and P x represents the prediction result of decoding and predicting the encoded features.
[0087] Specifically, the cross-entropy semantic segmentation loss is used to measure the difference between the model prediction and the true label, and the cross-entropy semantic segmentation loss function is obtained by calculating the classification error of each pixel through the following formula
[0088]
[0089] The obtained loss function value is used to guide the training of the semantic segmentation model until the loss function value is less than the preset loss threshold, and the target semantic segmentation model is obtained.
[0090] In one embodiment, model optimization includes adjusting the parameters in the semantic segmentation model, and the adjustment effect of the parameters is fed back through the loss function value until it is less than the preset loss threshold.
[0091] Specifically, the soft alignment loss can make the two modalities maintain a certain similarity while retaining their unique characteristics, and then the cross-entropy semantic segmentation loss function is used to obtain a more accurate target semantic segmentation model.
[0092] S5. Use the target semantic segmentation model to perform semantic segmentation on the target image to obtain a semantic segmentation result.
[0093] In one embodiment, the target image is an image that needs to be semantically segmented, where the target image includes a three-channel color image and a depth source domain image.
[0094] Specifically, using the target semantic segmentation model to perform semantic segmentation on the target image to obtain a semantic segmentation result includes: using the low-level feature extraction module in the target semantic segmentation model to extract the low-level image features of the target image; performing feature fusion on the low-level image features to obtain target image features; performing feature encoding and decoding processing on the target image features to obtain a semantic segmentation result.
[0095] In one embodiment, the low-level feature extraction module of the target semantic segmentation model is used to extract the low-level image features of the three-channel color image and the depth source domain image in the target image. Then, the low-level image features are multiplied element by element and then added element by element to the low-level image features of the three-channel color image to obtain the target image features. The encoding and decoding modules in the target semantic segmentation model are used to perform feature encoding and decoding processing on the target image features to obtain the prediction result of the target image semantic segmentation, that is, the semantic segmentation result.
[0096] Specifically, the target semantic segmentation model can accurately capture the category sensitivity and spatial sensitivity features, while maximizing the extraction of non-sensitive features for each category, so as to learn more domain-invariant features to improve the model performance, and then achieve accurate semantic segmentation.
[0097] As Figure 5 shown, it is a functional module diagram of a domain generalization semantic segmentation device provided by an embodiment of the present invention.
[0098] The domain generalization semantic segmentation device 500 in this embodiment can be installed in an electronic device. According to the implemented functions, the domain generalization semantic segmentation device 500 may include a low-level image feature extraction module 501, a feature stylization processing module 502, a class-level spatial soft sensitivity suppression module 503, a model optimization module 504, and an image semantic segmentation module 505. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0099] In this embodiment, the functions of each module / unit are as follows: The low-level image feature extraction module 501 is used to obtain the three-channel color image and the depth source domain image of the image to be trained, and extract the first low-level image features of the three-channel color image and the second low-level image features of the depth source domain image; the feature stylization processing module 502 is used to perform stylization processing on the second low-level image features according to the first low-level image features and the second low-level image features to obtain stylized depth features; the class-level spatial soft sensitivity suppression module 503 is used to perform class-level spatial soft sensitivity suppression on the second low-level image features according to the stylized depth features to obtain enhanced depth features; the model optimization module 504 is used to optimize the pre-constructed semantic segmentation model according to the enhanced depth features and the stylized depth features to obtain the target semantic segmentation model; the image semantic segmentation module 505 is used to perform semantic segmentation on the target image by using the target semantic segmentation model to obtain the semantic segmentation result.
[0100] Specifically, in one embodiment, each module in the domain generalization semantic segmentation device 500 uses the same technical means as a domain generalization semantic segmentation method in the accompanying drawings when in use, and can produce the same technical effects, which will not be elaborated here.
[0101] As Figure 6 shown, it is a schematic structural diagram of an electronic device for implementing a domain generalization semantic segmentation method provided by an embodiment of the present invention.
[0102] The electronic device 600 may include a processor 601, a memory 602, a communication bus 603, and a communication interface 604, and may further include a computer program stored in the memory 602 and executable on the processor 601, such as a domain generalization semantic segmentation program.
[0103] Among them, the processor 601 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips.
[0104] The memory 602 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 602 may be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device.
[0105] The communication bus 603 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 602 and at least one processor 601, etc.
[0106] The communication interface 704 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is usually used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface.
[0107] Figure 6 Only an electronic device with components is shown. It can be understood by those skilled in the art that Figure 6 the shown structure does not constitute a limitation on the electronic device 600, and it may include fewer or more components than those shown, or combine certain components, or have a different component arrangement.
[0108] It should be understood that the embodiments are only for illustrative purposes and are not limited by this structure in the scope of the patent application.
[0109] The present invention also provides a computer-readable storage medium. The readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a domain generalization semantic segmentation method according to any one of the above embodiments. It should be noted that the computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory). In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.
[0110] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus a software functional module.
[0111] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0112] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claims involved.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A domain generalization semantic segmentation method, characterized in that, The method includes: Obtain a three-channel color image of the image to be trained and a depth source domain image, and extract a first low-level image feature of the three-channel color image and a second low-level image feature of the depth source domain image; Perform stylization processing on the second low-level image feature according to the first low-level image feature and the second low-level image feature to obtain a stylized depth feature; Perform class-level spatial soft sensitivity suppression on the second low-level image feature according to the stylized depth feature to obtain an enhanced depth feature; Optimize a pre-constructed semantic segmentation model according to the enhanced depth feature and the stylized depth feature to obtain a target semantic segmentation model; Use the target semantic segmentation model to perform semantic segmentation on a target image to obtain a semantic segmentation result.
2. The domain generalization semantic segmentation method according to claim 1, characterized in that, The performing stylization processing on the second low-level image feature according to the first low-level image feature and the second low-level image feature to obtain a stylized depth feature includes: Calculate the variance of the inter-modal flow and the mean of the inter-modal flow according to the first low-level image feature and the second low-level image feature; Calculate a stylized variance and a stylized mean according to the variance of the inter-modal flow and the mean of the inter-modal flow; Perform stylization processing on the second low-level image feature according to the stylized variance and the stylized mean to obtain a stylized depth feature.
3. A domain generalization semantic segmentation method according to claim 2, characterized in that, The calculating the variance of the inter-modal flow and the mean of the inter-modal flow according to the first low-level image feature and the second low-level image feature includes: Calculate the variance of the inter-modal flow and the mean of the inter-modal flow by using a preset first formula, where the first formula is expressed as: Among them, represents the mean of the inter-modal flow, μ(·) represents the mean, C represents a matrix of a preset size, and Z d represents the low-level features of the second image, represents the variance of the inter-modal flow, σ(·) represents the variance, and Z rgb represents the low-level features of the first image.
4. The domain generalization semantic segmentation method according to claim 2, wherein, The calculating the stylized variance and the stylized mean according to the variance of the inter-modal flow and the mean of the inter-modal flow includes: Calculate the stylized variance and the stylized mean by using a preset second formula, where the second formula is expressed as: Among them, μ(Z d-style ) represents the stylized mean, μ(·) represents the mean, and Z d represents the low-level features of the second image. λ ∈ [0, 1], represents the mean of the inter-modal flow, σ(Z d-style ) represents the stylized variance, σ(·) represents the variance, represents the variance of the inter-modal flow.
5. The domain generalization semantic segmentation method according to claim 1, characterized in that, The performing class-level spatial soft sensitivity suppression on the second low-level image feature according to the stylized depth feature to obtain an enhanced depth feature includes: Calculate a depth feature difference corresponding to the second low-level image feature according to the stylized depth feature; Calculate a class-level spatial sensitivity matrix according to the depth feature difference, and calculate the weight of the insensitive feature according to the class-level spatial sensitivity matrix; Perform feature enhancement on the second low-level image feature according to the weight of the insensitive feature to obtain an enhanced depth feature.
6. The domain generalization semantic segmentation method according to claim 5, characterized in that The calculating the class-level spatial sensitivity matrix according to the depth feature difference and calculating the weight of the insensitive feature according to the class-level spatial sensitivity matrix includes: Calculate the class-level spatial sensitivity matrix by using a preset sensitivity matrix calculation formula, where the sensitivity matrix calculation formula is expressed as: Among them, represents the spatial sensitivity matrix of the k-th preset category in semantic segmentation, Mean represents the mean operation, and Softmax represents the Softmax activation function. represents the feature corresponding to the k-th preset category in the differential feature. Calculate the weight of the insensitive feature by using a preset weight formula of the insensitive feature, where the weight formula of the insensitive feature is expressed as: N g = 1 - S g Among them, Sigmoid represents the Sigmoid activation function, and Conv 1×1 represents a 1×1 convolution operation, represents the spatial sensitivity matrix of the k-th preset category in semantic segmentation, K represents the total number of preset categories in semantic segmentation, and N g represents the weight of insensitive features.
7. The domain generalization semantic segmentation method according to claim 1, wherein The optimizing a pre-constructed semantic segmentation model according to the enhanced depth feature and the stylized depth feature to obtain a target semantic segmentation model includes: Calculate a soft alignment loss according to the stylized depth feature and the first low-level image feature; Fuse the enhanced depth features and the first low-level image features to obtain robust multi-modal features; Encode the robust multi-modal features using the encoding module of the semantic segmentation model to obtain encoded features; Decode and predict the encoded features to obtain a prediction result, and calculate the cross-entropy semantic segmentation loss of the robust multi-modal features according to the semantic prediction result; Optimize the parameters of the semantic segmentation model according to the loss function values of the soft alignment loss and the cross-entropy semantic segmentation loss until the loss function value is less than a preset loss threshold to obtain a target semantic segmentation model.
8. The domain generalization semantic segmentation method according to claim 1, wherein The semantic segmentation of the target image using the target semantic segmentation model to obtain a semantic segmentation result includes: Extract the low-level image features of the target image using the low-level feature extraction module in the target semantic segmentation model; Fuse the low-level image features to obtain target image features; Perform feature encoding and decoding processing on the target image features to obtain a semantic segmentation result.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The processor is configured to execute a computer program stored on the memory; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the domain generalization semantic segmentation method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the domain generalization semantic segmentation method according to any one of claims 1 to 8 is implemented.