Domain adaptive semantic segmentation method and system based on lightweight visual converter, terminal and storage medium
By applying a lightweight visual converter in unsupervised domain adaptation semantic segmentation, combining CNN, BN/LN, additive attention mechanism and rectangular self-calibration module, the problem of inaccurate segmentation results is solved, and efficient migration and practicality of the semantic segmentation model between different domains is achieved.
Patent Information
- Application Number
- CN202510297078.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, lightweight vision converters are not applied to unsupervised domain adaptation semantic segmentation, resulting in inaccurate segmentation results.
The domain adaptation semantic segmentation method based on lightweight vision converters is adopted, and the semantic segmentation of images is performed through an unsupervised domain adaptation model, including a downsampling module based on CNN, a fusion module based on BN and LN, a visual converter model variant based on additive attention mechanism, and a rectangular self-calibration module based on RCM.
It effectively solves the problem of the migration of semantic segmentation model between different domains, saves a lot of manpower, reduces the cost of data annotation, and improves the practicality of segmentation.
Smart Images

Figure CN120147644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic segmentation, and in particular to a domain adaptation semantic segmentation method, system, terminal and computer-readable storage medium based on a lightweight vision transformer. Background Art
[0002] In recent years, computer vision technology has made breakthrough progress, and semantic segmentation has important applications in many fields, such as medical image analysis, autonomous driving, remote sensing image interpretation, video surveillance, etc. However, semantic segmentation models usually require a large amount of pixel-level annotated data for training, which is a time-consuming and laborious task. In addition, semantic segmentation models often suffer from the problem of domain shift, that is, in different data sets or scenarios, the generalization ability of the model will decline, resulting in a decrease in segmentation performance. To solve the problem of domain shift, a common method is to use the annotated data of the target domain for fine-tuning or retraining. However, the disadvantage of this method is that it requires a large amount of manual annotation, which is both time-consuming and resource-consuming.
[0003] Therefore, unsupervised domain adaptation (UDA) for semantic segmentation is an alternative to avoid data annotation problems: by jointly leveraging labeled images from different source datasets (the label spaces of the two datasets must be compatible), a well-performing model is learned from the unlabeled target dataset. The main idea of unsupervised domain adaptation for semantic segmentation is to reduce or eliminate domain shift by aligning the data distributions of the source domain and the target domain, thereby improving the segmentation performance of the model on the target domain. Unsupervised domain adaptation semantic segmentation methods can be divided into two categories: Adversarial-training methods and Self-training methods. The purpose of Adversarial-training methods is to align the distributions of the source and target domains at the input, feature, output, or patch level in the GAN framework. Using multiple scales or class information for the discriminator can optimize the alignment. In Self-training methods, the network is trained using pseudo-labels of the target domain. Most UDA methods pre-compute pseudo-labels offline, train the model, and then repeat this process. Alternatively, pseudo-labels can be computed online during training. To avoid training instability, pseudo-label prototypes or consistency regularization methods based on data augmentation or domain confusion are adopted. Since datasets are usually imbalanced and follow a long-tail distribution, this makes the model tend to learn common classes. Strategies to address this problem include resampling, loss re-weighting, and transfer learning. In UDA, re-weighting and class-balanced sampling are adopted for image classification. Since it has been proven that resampling is very effective for UDA training of Transformers, DAFormer implements a new UDA network structure that includes a Transformer encoder and a multi-level context-aware feature fusion decoder, but there are still few studies on applying Transformers to UDA currently. Although a large number of UDA methods have proposed new adaptation strategies, most of them are based on classical network architectures (DeeplabV2, FCN8s). Traditional UDA methods mainly use the DeepLabV2 or FCN8s network architectures with ResNet or VGG Backbone to evaluate their contributions.
[0004] Therefore, there is still room for improvement and development in the prior art. Summary of the Invention
[0005] The main purpose of the present invention is to provide a domain adaptation semantic segmentation method, system, terminal, and computer-readable storage medium based on a lightweight vision transformer, aiming to solve the problem in the prior art that the lightweight vision transformer is not applied to unsupervised domain adaptation for semantic segmentation of images, resulting in inaccurate segmentation results.
[0006] To achieve the above object, the present invention provides a domain adaptation semantic segmentation method based on a lightweight vision transformer. The domain adaptation semantic segmentation method based on a lightweight vision transformer is implemented by an unsupervised domain adaptation model. The unsupervised domain adaptation model includes a downsampling module based on CNN, a fusion module based on the combination of BN and LN, a variant of the vision transformer model based on the additive attention mechanism, and a rectangular self-calibration module based on RCM. The domain adaptation semantic segmentation method based on a lightweight vision transformer includes: Obtain the original image, and use the downsampling module based on CNN to perform downsampling processing on the original image to obtain a first feature map; Use the fusion module based on the combination of BN and LN to perform normalization operation and adaptive parameter fusion processing on the first feature map to obtain a second feature map; Use the variant of the vision transformer model based on the additive attention mechanism to perform feature weighting processing on the second feature map by using the additive attention mechanism to obtain a third feature map; Use the rectangular self-calibration module based on RCM to perform spatial feature reconstruction and pyramid context extraction processing on the third feature map, output the semantic segmentation result of the original image, and divide each pixel in the original image into different semantic categories.
[0007] Optionally, in the domain adaptation semantic segmentation method based on a lightweight vision transformer, the training of the unsupervised domain adaptation model uses rare class sampling and learning rate warm-up.
[0008] Optionally, in the domain adaptation semantic segmentation method based on a lightweight vision transformer, the rare class sampling is used to balance the data distribution or adjust the loss weight, and the learning rate warm-up is used to gradually increase the learning rate; During the rare class sampling process, the frequency of each category in the source domain is denoted as : ; where is the source domain dataset, and are the height and width of the image respectively, is the indicator function. If the sample is at the spatial position and belongs to the category , the function value is 1, otherwise it is 0; The sampling probability of the category is defined as: ; where is a temperature parameter used to adjust the smoothness of the probability distribution. denotes summation over all classes and is the total number of classes in the source domain dataset. is the class frequency; In the warm-up stage, before the iteration number reaches , the learning rate for the -th iteration is: ; where is the base learning rate, representing the learning rate benchmark value used for normal training of the model after the warm-up stage. represents the total number of steps in the learning rate warm-up stage.
[0009] Optionally, in the domain adaptation semantic segmentation method based on the lightweight vision transformer, the processing process of the CNN-based downsampling module specifically includes: using an improved fused mobile inverted convolution module and a max pooling layer: ; ; ; ; where represents a 3×3 depth convolution, represents a Gaussian error linear unit, represents a squeeze-and-excitation module, represents a 1×1 convolution, represents the input, represents the output; In the fused mobile inverted convolution module, perform a 3×3 depth convolution on the input, expand the number of channels by 4 times, and perform normalization and ReLU activation function; Perform SE processing on the output after convolution, learn the importance of each channel through global average pooling and two fully connected layers, and weight each channel; Perform a 1×1 convolution on the weighted output, restore the number of channels to the original size, and perform normalization and residual connection. Add a max pooling layer and a normalization layer after the fused mobile inverted convolution module.
[0010] Optionally, in the domain adaptation semantic segmentation method based on the lightweight vision transformer, the fusion module based on BN and LN includes a batch normalization branch and a layer normalization branch. The processing process of the fusion module based on BN and LN specifically includes: The batch normalization branch calculates the mean and variance along the (M, h, w) axes, extracts batch-level statistical information, and is used to reflect the overall distribution characteristics of the current batch of data. By capturing the distribution law of the data within the batch in the spatial dimension (h, w) and the batch dimension (M); The layer normalization branch calculates the mean and variance along the (P, h, w) axes, extracts channel-level statistical information, and is used to focus on the data distribution characteristics of each channel (P) itself; Key statistical information is obtained from the batch dimension and the channel dimension through the batch normalization branch and the layer normalization branch respectively; A learnable parameter α is introduced to dynamically adjust the contributions of BN and LN, where 0 ≤ α ≤ 1. The learnable parameter α is learned through backpropagation and is automatically optimized according to the dataset characteristics or task requirements; A unified scaling and translation operation is performed on the fused feature z: z = γ⋅z + β, where γ and β are learnable parameters.
[0011] Optionally, in the domain adaptation semantic segmentation method based on the lightweight vision transformer, the processing process of the vision transformer model variant based on the additive attention mechanism specifically includes: Regarding the input matrix as a set of multiple vectors, a global query vector is obtained by performing weighted averaging on the input vectors, which is used to compress the information of the global context. The calculation formula of the global query vector is: ; ; where, represents the final output result of the global query vector, which is the weighted summation comprehensive value of multiple input elements , represents the total number of input elements, represents the th input element , and is used to measure the contribution degree to the final output , represents the input vector , represents a learnable parameter vector, which is optimized through model training and is used to measure the importance of the input , represents the weight vector and the input vector Vector inner product; For the key matrix, it is also regarded as the geometry of multiple vectors. Perform the Hadamard product with the global query vector, and perform weighted averaging on the interaction vectors to obtain a global key vector. The calculation formula for the global key vector is: ; ; Where, represents the final output result of the global key vector, represents the th input element weight coefficient, used to measure 's contribution to the final output , represents the vector inner product of the weight vector and the input vector ; Perform element-wise multiplication between the global key vector and each value vector to calculate the key-value interaction vector. Apply a linear transformation layer to learn the hidden representation, and combine them into a matrix and add it to the original query matrix to form the final output.
[0012] Optionally, in the domain adaptation semantic segmentation method based on the lightweight vision transformer, where, In addition, to achieve the above object, the present invention also provides a domain adaptation semantic segmentation system based on a lightweight vision transformer, where the domain adaptation semantic segmentation system based on the lightweight vision transformer includes: A CNN-based downsampling module for performing downsampling processing on the acquired original image to obtain a first feature map; A fusion module based on BN and LN for performing normalization operations and adaptive parameter fusion processing on the first feature map to obtain a second feature map; A variant of the vision transformer model based on the additive attention mechanism for performing feature weighting processing on the second feature map using the additive attention mechanism to obtain a third feature map; A rectangular self-calibration module based on RCM for performing spatial feature reconstruction and pyramid context extraction processing on the third feature map, outputting the semantic segmentation result of the original image, and dividing each pixel in the original image into different semantic categories.
[0013] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and a domain adaptation semantic segmentation program based on a lightweight vision transformer stored on the memory and executable on the processor. When the domain adaptation semantic segmentation program based on the lightweight vision transformer is executed by the processor, the steps of the above-mentioned domain adaptation semantic segmentation method based on the lightweight vision transformer are implemented.
[0014] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a domain adaptation semantic segmentation program based on a lightweight vision transformer. When the domain adaptation semantic segmentation program based on the lightweight vision transformer is executed by a processor, the steps of the above-mentioned domain adaptation semantic segmentation method based on the lightweight vision transformer are implemented.
[0015] In the present invention, an original image is obtained, and the original image is downsampled by a CNN-based downsampling module to obtain a first feature map; the first feature map is normalized and adaptively parameter-fused by a fusion module based on BN and LN to obtain a second feature map; a variant of the vision transformer model based on the additive attention mechanism uses the additive attention mechanism to perform feature weighting on the second feature map to obtain a third feature map; a rectangular self-calibration module based on RCM performs spatial feature reconstruction and pyramid context extraction on the third feature map, and outputs the semantic segmentation result of the original image, dividing each pixel in the original image into different semantic categories. The present invention can effectively solve the problem of migrating the semantic segmentation model between different domains, save a large amount of manpower, reduce the cost of data annotation, and improve the practicality of segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of a preferred embodiment of the domain adaptation semantic segmentation method based on the lightweight vision transformer of the present invention; Figure 2 is a schematic diagram of the network framework structure in a preferred embodiment of the domain adaptation semantic segmentation method based on the lightweight vision transformer of the present invention; Figure 3 is a schematic diagram of the processing process of the CNN-based downsampling module in a preferred embodiment of the domain adaptation semantic segmentation method based on the lightweight vision transformer of the present invention; Figure 4 is a schematic diagram of the processing process of the rectangular self-calibration module based on RCM in a preferred embodiment of the domain adaptation semantic segmentation method based on the lightweight vision transformer of the present invention; Figure 5 is a generated effect diagram after image semantic segmentation in a preferred embodiment of the domain adaptation semantic segmentation method based on the lightweight vision transformer of the present invention; Figure 6 It is a structural diagram of a preferred embodiment of the domain adaptation semantic segmentation system based on a lightweight vision transformer of the present invention; Figure 7 It is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed implementation manners
[0017] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] The present invention adopts a downsampling module based on CNN, a fusion module based on the fusion of BN and LN, a variant of the vision transformer model based on the additive attention mechanism, and a rectangular self-calibration module based on RCM. The novel downsampling module that borrows spatial feature contraction from CNN aims to reduce dimensions while adding cross-channel interaction; the fusion module based on BN and LN technologies utilizes channel and batch dependencies and adaptively combines the advantages of BN and LN based on a specific dataset or task to adaptively perform normalized output; the efficient additive attention mechanism module calculates the attention matrix as a vector combination, greatly reducing the complexity of the original matrix multiplication calculation; the rectangular self-calibration module based on RCM explicitly models the rectangular key regions, respectively capturing the global context information of the image in two directions and improving the feature representation ability.
[0019] The domain adaptation semantic segmentation method based on a lightweight vision transformer according to a preferred embodiment of the present invention, as Figure 1 shown, the domain adaptation semantic segmentation method based on a lightweight vision transformer is implemented through an unsupervised domain adaptation model. The unsupervised domain adaptation model includes a downsampling module based on CNN, a fusion module based on the fusion of BN and LN, a variant of the vision transformer model based on the additive attention mechanism, and a rectangular self-calibration module based on RCM. The domain adaptation semantic segmentation method based on a lightweight vision transformer includes the following steps: Step S10: Obtain an original image, and perform downsampling processing on the original image by using the downsampling module based on CNN to obtain a first feature map.
[0020] Specifically, as Figure 2As shown, to train the UDA (Unsupervised Domain Adaptation model), the framework adopts two simple yet effective training strategies, namely rare class sampling and learning rate warm-up. Rare class sampling can enhance the model's recognition ability for minority classes by balancing data distribution or adjusting loss weights; learning rate warm-up can ensure the stability in the initial stage of training and promote model convergence by gradually increasing the learning rate. To mitigate the confirmation bias of self-training for common classes, the framework performs rare class sampling on the source domain to improve the quality of pseudo-labels.
[0021] During the rare class sampling process, the frequency of each class in the source domain is denoted as : ; where is the source domain dataset, and are the height and width of the image respectively, is the indicator function, and if the sample is at the spatial position and belongs to the class , the function value is 1, otherwise it is 0.
[0022] The sampling probability of class is defined as: ; where is the temperature parameter used to adjust the smoothness of the probability distribution, represents the summation over all classes , is the total number of classes in the source domain dataset, is the frequency of class .
[0023] Therefore, classes with smaller frequencies will have higher sampling probabilities, and rare classes will be sampled more frequently than common classes, which is beneficial for the resampled classes to be closer to balance. The second strategy is the learning rate warm-up for UDA. Linearly increasing the learning rate at the beginning of Transformer training can improve the generalization ability of the network by avoiding large variances in the adaptive learning rate that distort the gradient distribution at the beginning of training. During the warm-up phase, before the number of iterations reaches , the learning rate for the th iteration is: ; where represents the learning rate corresponding to the current step in the training process, that is, the learning rate when the model reaches the The learning rate actually used during the step is the base learning rate, which represents the learning rate benchmark value used for the normal training of the model after the warm-up phase. refers to the current training step, that is, the number of iterations that have been executed during the training process. represents the total number of steps in the learning rate warm-up phase, that is, the total number of iterations experienced from the start of training to the end of warm-up.
[0024] In the present invention, a new downsampling module that borrows spatial feature contraction from CNN is adopted. This module aims to reduce the dimension while adding cross-channel interaction. In CNN, the role of cross-channel interaction is that it can enhance the network's ability to express features, enabling the network to better capture and utilize information in different channels. In this way, the network can learn more rich and complex feature representations, thereby improving the performance and generalization ability of the model. In addition, cross-channel interaction also helps to reduce the number of parameters and computational amount of the model, because it allows the network to achieve more effective information exchange and feature fusion with fewer parameters.
[0025] The processing process of the CNN-based downsampling module specifically includes: Specifically, an improved fusion mobile inverted convolution module and a max pooling layer need to be used: ; ; ; ; Among them, represents a 3×3 depth convolution, represents a Gaussian error linear unit, represents a squeeze-and-excitation module, represents a 1×1 convolution, represents the input, represents the output.
[0026] As Figure 3 shown, in the fusion mobile inverted convolution module, a 3×3 depth convolution is performed on the input, the number of channels is expanded by 4 times, and normalization and ReLU activation function are performed; the output after convolution is SE processed by global average pooling and two fully connected layers to learn the importance of each channel and weight each channel; finally, a 1×1 convolution is performed on the weighted output to restore the number of channels to the original size, and normalization and residual connection are performed. In this method, downsampling follows the above module, and a max pooling layer and a normalization layer are added after the fusion mobile inverted convolution module to expect to improve the generalization ability and robustness of the overall model. Step S20: Use a fusion module based on BN and LN to perform normalization operation and adaptive parameter fusion processing on the first feature map to obtain a second feature map.
[0027] Specifically, the core idea of BN is to normalize small batches of data at each layer of the network. The specific operation is to independently calculate the mean and variance of the batch (M) and spatial (h, w) dimensions for each channel (P): The normalized data is reconstructed by scaling and translation parameters. The advantage of BN is that it uses batch statistics to accelerate convergence and performs well in large batch tasks such as image classification. However, it is sensitive to batch size. When the batch is too small, the statistics are unstable, and the performance is significantly reduced in dynamic input scenarios such as target detection and video processing.
[0028] LN independently calculates the mean and variance of the channel (P) and spatial (h, w) dimensions for each sample, and the normalized data is also reconstructed through learnable parameters. The advantage of LN is that it is insensitive to batch size and is suitable for small batches or dynamic input scenarios. However, it ignores the statistical information between batches, has insufficient generalization ability in spatial data processing such as images, and the amount of calculation increases significantly with the increase in the number of channels.
[0029] This method proposes an adaptive joint normalization method to achieve the complementary advantages of BN and LN through the following steps: The fusion module based on BN and LN includes a batch normalization branch and a layer normalization branch. The processing process of the fusion module based on BN and LN specifically includes: The batch normalization branch calculates the mean and variance along the (M, h, w) axis and extracts batch-level statistical information. This information is used to reflect the overall distribution characteristics of the current batch data. By capturing the distribution patterns of data within the batch in the spatial dimensions (h, w) and the batch dimension (M), it provides a batch-level statistical basis for the normalization operation and helps the model learn the common characteristics of batch data.
[0030] The layer normalization branch calculates the mean and variance along the (P, h, w) axis, extracts channel-level statistical information, and focuses on the data distribution characteristics of each channel (P). It can accurately describe the differences in feature expression between different channels, help the model to explore the uniqueness and correlation between channels, and provide support for feature optimization in the channel dimension. Through the two branches of dual-axis independent normalization, key statistical information can be obtained from the batch dimension and the channel dimension respectively, laying a solid foundation for subsequent feature fusion and model performance improvement.
[0031] Adaptive parameter fusion: A learnable parameter α is introduced to dynamically adjust the contribution of BN and LN, 0 ≤ α ≤ 1. The learnable parameter α is learned through back propagation and automatically optimized according to the characteristics of the dataset or task requirements.
[0032] Uniform parameter reconstruction: Perform uniform scaling and translation operations on the fused feature z: z = γ⋅z + β, where γ and β are learnable parameters.
[0033] By dynamically fusing BN and LN, the model can adaptively process input data in different scenarios, improving performance in complex scenarios such as small batches and dynamic inputs. And by fully utilizing the complementarity of batch statistics and channel statistics, the robustness and generalization ability of feature representation are enhanced.
[0034] Step S30: Use a variant of the vision transformer model based on the additive attention mechanism to perform feature weighting on the second feature map using the additive attention mechanism to obtain a third feature map.
[0035] Specifically, due to the high complexity and large size of the vision transformer model (Vision Transformer, hereinafter referred to as ViT) based on the self-attention mechanism, it is difficult to be actually deployed. Therefore, it is necessary to try to use a more lightweight ViT. This project intends to optimize the model size and complexity using a variant of ViT based on additive attention. Specifically, the input matrix needs to be regarded as a set of multiple vectors, and by performing weighted averaging on the input vectors, a global query vector is obtained for compressing the information of the global context. The calculation formula for the global query vector is: ; ; where represents the final output result of the global query vector, which is the weighted summation comprehensive value of multiple input elements , represents the total number of input elements, that is, the number of participating in the calculation, represents the th input element , is the weight coefficient of the th input element, used to measure the contribution degree of to the final output , represents the dimension of the input vector , is a learnable parameter vector, optimized through model training, used to measure the importance of the input , represents the vector inner product of the weight vector
[0036] For the key matrix, it is also regarded as the geometry of multiple vectors. Perform the Hadamard product with the global query vector, and perform weighted averaging on the interaction vectors to obtain a global key vector. The calculation formula for the global key vector is: ; ; Among them, represents the final output result of the global key vector, represents the th input element 's weight coefficient, which is used to measure 's contribution to the final output , represents the vector inner product of the weight vector and the input vector .
[0037] Similarly, perform element-wise multiplication between the global key vector and each value vector to calculate the key-value interaction vector. Immediately apply a linear transformation layer to learn the hidden representation, and combine them into a matrix and add it to the original query matrix to form the final output. In this way, by modeling the association between inputs through element-wise multiplication, the time complexity can be reduced to , which is significantly improved compared to the original complexity .
[0038] Step S40: Use the rectangle self-calibration module based on RCM to perform spatial feature reconstruction and pyramid context extraction on the third feature map, output the semantic segmentation result of the original image, and divide each pixel in the original image into different semantic categories.
[0039] Specifically, as Figure 4As shown, the processing process of the RCM-based rectangular self-calibration module specifically includes: The RCM (Rectangular Context Module) first captures the global context information of the image in the horizontal and vertical directions through horizontal pooling and vertical pooling respectively, generating two independent axis vectors. Subsequently, a broadcast addition operation is performed on the two axis vectors to explicitly model the rectangular key area. This step allows the RCM to capture the rectangular region of interest in the image. To make the rectangular region of interest closer to the object, this method designs a shape self-calibration function. The shape self-calibration function uses horizontal strip convolution and vertical strip convolution to independently calibrate the attention area. The horizontal strip convolution adjusts the elements of each row to make the horizontal shape closer to the object. The vertical strip convolution is similar to the horizontal operation to further calibrate the shape. The weights of the two convolutions are learnable. By training to learn appropriate weights, the rectangular region is adjusted to match the object. The attention features after shape self-calibration will be fused with the original input features. The local details of the input features are extracted through 3×3 depthwise separable convolution, and the calibrated attention features are weighted by multiplication to obtain the refined output features, thereby enhancing the attention to the foreground object. To further improve the feature representation ability, the RCM adds batch normalization and a multi-layer perceptron (MLP) after rectangular self-attention to refine the features. Finally, the feature reuse is further enhanced through residual connection. The generated effect diagram is as shown in Figure 5 shown.
[0040] After obtaining the original image, the present invention processes it using a CNN-based downsampling module. This module performs downsampling on the original image through convolutional layer operations. While reducing the spatial resolution of the image, it enhances the network's ability to express features through cross-channel interactions. After being processed by this module, the image is converted into a feature map with more advanced features. Then, a fusion module based on BN and LN processes the obtained feature map. This module adopts an adaptive joint normalization method. It performs normalization calculations independently on two axes, that is, on the batch axis and the channel axis respectively, and then conducts adaptive parameter fusion to make the features in the feature map more stable and representative. After that, a Vision Transformer variant module based on the additive attention mechanism processes the normalized feature map. This module weights the features at different positions in the feature map through an optimized attention mechanism, enabling the model to pay more attention to important feature regions. It uses the additive attention mechanism when calculating attention, further reducing the computational complexity of the method compared to traditional attention calculation methods. In this way, the model can better capture long-range dependencies in the feature map, perform global modeling on the feature map, and successfully optimize the feasibility of algorithm deployment. Finally, a rectangular self-calibration module based on RCM processes the feature map after attention processing. This module successfully performs spatial feature reconstruction and pyramid context extraction. It uses rectangular regions to self-calibrate the feature map, processes the feature map through rectangular windows of different scales, captures spatial feature information at different scales, and enhances the model's ability to understand objects of different scales in the image. After being processed by this module, the semantic segmentation result of the final output image is obtained, and each pixel in the image is assigned to different semantic categories.
[0041] From an academic perspective, unsupervised domain adaptation semantic segmentation is a challenging and promising research direction that involves knowledge in multiple fields such as machine learning, computer vision, image processing, and optimization theory. The research on unsupervised domain adaptation semantic segmentation can promote the development of these fields, facilitate the emergence of new theories and methods, and improve academic levels and influence. The research on unsupervised domain adaptation semantic segmentation can also provide reference and inspiration for other related tasks, such as unsupervised domain adaptation object detection and unsupervised domain adaptation image classification.
[0042] In terms of application, unsupervised domain adaptation semantic segmentation can effectively solve the problem of migrating semantic segmentation models between different domains, saving a large amount of manpower, reducing the cost of data annotation, and improving the practicality of segmentation. Unsupervised domain adaptation semantic segmentation can be applied to many practical scenarios, such as medical image segmentation, autonomous driving, remote sensing image segmentation, etc., bringing convenience and value to people's lives and work. Unsupervised domain adaptation semantic segmentation can also promote the development of related industries, such as medical, transportation, agriculture, etc., and contribute to the progress and development of society.
[0043] Furthermore, as Figure 6 shown, based on the above-mentioned domain adaptation semantic segmentation method based on a lightweight vision transformer, the present invention also correspondingly provides a domain adaptation semantic segmentation system based on a lightweight vision transformer, wherein the domain adaptation semantic segmentation system based on a lightweight vision transformer includes: A CNN-based downsampling module 51 for performing downsampling processing on the acquired original image to obtain a first feature map; A fusion module 52 based on the fusion of BN and LN for performing normalization operation and adaptive parameter fusion processing on the first feature map to obtain a second feature map; A variant of the vision transformer model 53 based on the additive attention mechanism for performing feature weighting processing on the second feature map by using the additive attention mechanism to obtain a third feature map; A rectangular self-calibration module 54 based on RCM for performing spatial feature reconstruction and pyramid context extraction processing on the third feature map, outputting the semantic segmentation result of the original image, and dividing each pixel in the original image into different semantic categories.
[0044] Furthermore, as Figure 6 shown, based on the above-mentioned domain adaptation semantic segmentation method and system based on a lightweight vision transformer, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively.
[0045] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in some other embodiments, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit of the terminal and external storage devices. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a domain adaptation semantic segmentation program 40 based on a lightweight vision transformer is stored on the memory 20, and this domain adaptation semantic segmentation program 40 based on a lightweight vision transformer can be executed by the processor 10, thereby implementing the domain adaptation semantic segmentation method based on a lightweight vision transformer in the present application.
[0046] The processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips in some embodiments, and is used to run the program codes stored in the memory 20 or process data, such as executing the domain adaptation semantic segmentation method based on a lightweight vision transformer, etc.
[0047] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface. The processor 10, the memory 20, and the display 30 of the terminal communicate with each other through a system bus.
[0048] In one embodiment, when the processor 10 executes the domain adaptation semantic segmentation program 40 based on a lightweight vision transformer in the memory 20, the steps of the above-mentioned domain adaptation semantic segmentation method based on a lightweight vision transformer are implemented.
[0049] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a domain adaptation semantic segmentation program based on a lightweight vision transformer, and when the domain adaptation semantic segmentation program based on a lightweight vision transformer is executed by a processor, the steps of the domain adaptation semantic segmentation method based on a lightweight vision transformer as described above are implemented.
[0050] In summary, the present invention provides a domain adaptation semantic segmentation method, system, terminal, and computer-readable storage medium based on a lightweight vision transformer. The method includes: obtaining an original image, performing downsampling processing on the original image using a CNN-based downsampling module to obtain a first feature map; performing a normalization operation and adaptive parameter fusion processing on the first feature map using a fusion module based on BN and LN to obtain a second feature map; using a variant of the vision transformer model based on the additive attention mechanism to perform feature weighting processing on the second feature map using the additive attention mechanism to obtain a third feature map; using a rectangular self-calibration module based on RCM to perform spatial feature reconstruction and pyramid context extraction processing on the third feature map, and outputting the semantic segmentation result of the original image, dividing each pixel in the original image into different semantic categories. The present invention can effectively solve the problem of migration of the semantic segmentation model between different domains, save a large amount of manpower, reduce the cost of data annotation, and improve the practicality of segmentation.
[0051] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such a process, method, article, or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article, or terminal including that element.
[0052] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0053] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A domain adaptation semantic segmentation method based on lightweight visual converter, characterized in that: The domain adaptation semantic segmentation method based on the lightweight visual converter is implemented by an unsupervised domain adaptation model, wherein the unsupervised domain adaptation model includes a downsampling module based on CNN, a fusion module based on BN and LN, a visual converter model variant based on an additive attention mechanism, and a rectangular self-calibration module based on RCM. The domain adaptation semantic segmentation method based on the lightweight visual converter includes: Acquire an original image, and perform downsampling processing on the original image using a downsampling module based on CNN to obtain a first feature map; Using a fusion module based on BN and LN to perform normalization and adaptive parameter fusion processing on the first feature map to obtain a second feature map; Using a variant of the visual transformer model based on an additive attention mechanism, a feature weighting process is performed on the second feature map using an additive attention mechanism to obtain a third feature map; The third feature map is subjected to spatial feature reconstruction and pyramid context extraction processing by using a rectangular self-calibration module based on RCM, and a semantic segmentation result of the original image is output, so as to divide each pixel in the original image into different semantic categories.
2. The domain adaptation semantic segmentation method based on lightweight visual converter according to claim 1 is characterized in that: The unsupervised domain adaptation model is trained with rare class sampling and learning rate warmup.
3. The domain adaptation semantic segmentation method based on lightweight visual converter according to claim 2 is characterized in that: The rare class sampling is used to balance data distribution or adjust loss weight, and the learning rate warm-up is used to gradually increase the learning rate; In the rare class sampling process, each category in the source domain The frequency is recorded as : ; in, is the source domain dataset, and are the height and width of the image, is the indicator function, if the sample In spatial position Belongs to category , then the function value is 1, otherwise it is 0; category The sampling probability Defined as: ; in, is the temperature parameter, which is used to adjust the smoothness of the probability distribution. For all categories To sum, is the total number of categories in the source domain dataset, For Category frequency; In the warm-up phase, when the number of iterations reaches Before, The learning rate of the iteration for: ; in, Is the basic learning rate, which indicates the learning rate benchmark value used for normal training of the model after the warm-up phase. Represents the total number of steps in the learning rate warmup phase.
4. The domain adaptation semantic segmentation method based on lightweight visual converter according to claim 1, characterized in that: The processing process of the CNN-based downsampling module specifically includes: Use an improved fused shifted inverted convolution module and a max pooling layer: ; ; ; ; in, represents a 3×3 depthwise convolution, represents the Gaussian error linear unit, represents the extrusion and excitation module, represents 1×1 convolution, Indicates input, Indicates output; In the fused mobile inverted convolution module, a 3×3 depth convolution is performed on the input to expand the number of channels by 4 times, and normalization and ReLU activation function are performed; The output after convolution is SE Processing, learning the importance of each channel through global average pooling and two fully connected layers, and weighting each channel; A 1×1 convolution is performed on the weighted output to restore the number of channels to the original size, and normalization and residual connection are performed. A maximum pooling layer and a normalization layer are added after the fusion of the mobile inverted convolution module.
5. The domain adaptation semantic segmentation method based on lightweight visual converter according to claim 1, characterized in that: The fusion module based on BN and LN includes a batch normalization branch and a layer normalization branch. The processing process of the fusion module based on BN and LN specifically includes: The batch normalization branch calculates the mean and variance along the (M, h, w) axis and extracts batch-level statistical information to reflect the overall distribution characteristics of the current batch data by capturing the distribution law of the data within the batch in the spatial dimension (h, w) and the batch dimension (M); The layer normalization branch calculates the mean and variance along the (P, h, w) axis and extracts channel-level statistical information to focus on the data distribution characteristics of each channel (P) itself; The batch normalization branch and the layer normalization branch are used to obtain key statistical information from the batch dimension and the channel dimension respectively; Introduce a learnable parameter α to dynamically adjust the contribution of BN and LN, 0 ≤ α ≤ 1. The learnable parameter α is learned through back propagation and automatically optimized according to the characteristics of the data set or task requirements; Perform uniform scaling and translation operations on the fused feature z: z = γ⋅z+β, where γ and β are learnable parameters.
6. The domain adaptation semantic segmentation method based on lightweight visual converter according to claim 1, characterized in that: The processing of the variant of the visual transformer model based on the additive attention mechanism specifically includes: The input matrix is regarded as a collection of multiple vectors. By taking a weighted average of the input vectors, a global query vector is obtained to compress the global context information. The calculation formula of the global query vector is: ; ; in, Represents the final output result of the global query vector, which is a sum of multiple input elements The weighted sum of the comprehensive values, Represents the total number of input elements, Indicates Input elements The weight coefficient is used to measure For the final output The degree of contribution Represents the input vector The dimension of Represents a learnable parameter vector, optimized through model training, used to measure the input The importance of Represents the weight vector With the input vector The vector inner product of For the key matrix, it is also regarded as the geometry of multiple vectors. The Hadamard product is performed with the global query vector, and the interaction vector is weighted averaged to obtain a global key vector. The calculation formula of the global key vector is: ; ; in, Represents the final output result of the global key vector, Indicates Input elements The weight coefficient is used to measure For the final output The degree of contribution Represents the weight vector With the input vector The vector inner product of An element-wise product is performed between the global key vector and each value vector, the key-value interaction vector is calculated, a linear transformation layer is applied to learn the hidden representation, and the matrix is combined into a matrix and added to the original query matrix to form the final output.
7. The domain adaptation semantic segmentation method based on lightweight visual converter according to claim 1, characterized in that: The processing process of the rectangular self-calibration module based on RCM specifically includes: Through horizontal pooling and vertical pooling, the global context information of the image in the horizontal and vertical directions is captured respectively, and two independent axis vectors are generated. The two axis vectors are broadcasted and added to explicitly model the rectangular key area. Design a shape self-calibration function, which uses horizontal strip convolution and vertical strip convolution to independently calibrate the attention area; The attention features after shape self-calibration are fused with the original input features, and the local details of the input features are extracted through 3×3 depth-wise separable convolution. The calibrated attention features are weighted by product to obtain the refined output features. Batch normalization and multi-layer perceptron MLP are added after rectangular self-attention to refine the features.
8. A domain-adaptive semantic segmentation system based on lightweight visual converter, characterized in that: The domain adaptation semantic segmentation system based on lightweight visual converter includes: A downsampling module based on CNN is used to downsample the acquired original image to obtain a first feature map; A fusion module based on BN and LN, used for performing normalization operation and adaptive parameter fusion processing on the first feature map to obtain a second feature map; A variant of the visual transformer model based on an additive attention mechanism, configured to perform feature weighting processing on the second feature map using the additive attention mechanism to obtain a third feature map; The RCM-based rectangular self-calibration module is used to perform spatial feature reconstruction and pyramid context extraction processing on the third feature map, output the semantic segmentation result of the original image, and divide each pixel in the original image into different semantic categories.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a domain-adaptive semantic segmentation program based on a lightweight visual converter stored in the memory and executable on the processor. When the domain-adaptive semantic segmentation program based on a lightweight visual converter is executed by the processor, the steps of the domain-adaptive semantic segmentation method based on a lightweight visual converter as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a domain-adaptive semantic segmentation program based on a lightweight visual converter. When the domain-adaptive semantic segmentation program based on a lightweight visual converter is executed by a processor, the steps of the domain-adaptive semantic segmentation method based on a lightweight visual converter as described in any one of claims 1 to 7 are implemented.