Self-adaption difficult scene semantic segmentation method based on test time domain
Through a method based on test time domain adaptation, using predicted entropy to screen low-entropy samples for style transfer and using instance-level standardization layer, the domain offset and computing resource consumption problems of semantic segmentation models in complex scenarios are solved, and efficient semantic segmentation effect is achieved.
Patent Information
- Application Number
- CN202510337677.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The existing semantic segmentation model is difficult to adapt efficiently in unmanned driving environments due to domain offset and privacy security issues in complex real-life scenarios, and the existing methods consume high computing resources and cannot adapt to the dynamically changing target domain.
Using a method based on test time domain adaptation, low-entropy samples are screened through predictive entropy for style transfer, combined with an instance-level normalization layer to replace the batch normalization layer, dynamically adjust the statistics to adapt to the change of the target domain, and intermediate domains are generated to reduce domain offset.
Efficient semantic segmentation is achieved in the dynamically changing target domain, reducing computing resource consumption, and improving the model's adaptability and segmentation accuracy.
Smart Images

Figure CN120259658A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of time-domain adaptive, and specifically relates to a semantic segmentation method for difficult scenarios based on test-time domain adaptation. Background Art
[0002] The most significant application scenario of semantic segmentation is autonomous driving, but there are often many obstacles in complex real-world scenarios. A semantic segmentation model trained on data under clear weather conditions on the European continent may suffer significant performance degradation due to different weather, motion blur during driving, and different street scene locations. This phenomenon of domain shift caused by the difference in data distribution between the source domain and the target domain is very common in autonomous driving scenarios. In addition, the source domain data may contain sensitive information, and accessing this data may raise privacy and security issues, especially in fields such as healthcare and finance. Moreover, the source domain data itself is usually large in scale, and with limited storage and computing resources, accessing this data will increase the burden and affect efficiency. Therefore, in real-world scenarios, considering privacy and efficiency, accessing the data and labels of the source domain does not conform to the actual application situation.
[0003] Previous research can be addressed using methods of domain generalization (DG) and domain adaptation (DA), both of which are important methods in transfer learning. However, both of these transfer learning methods have their own drawbacks. The performance of DG highly depends on the diversity and quality of the source domain. If the source domain data is insufficient to cover the possible distributions of the target domain, the performance of the model will be greatly reduced. DA usually requires methods of adversarial training and domain alignment, which need additional computing resources, and the training process may be more complex and time-consuming than traditional machine learning methods. Considering the high requirements for the model inference rate in autonomous driving scenarios and the reality of not accessing the source domain, the methods of domain generalization and domain adaptation are not suitable.
[0004] Existing test-time adaptation methods usually can only adapt to a single invariant target domain. Before the model starts to adapt to another target domain, the model parameters are reset to the source model state. This method of parameter resetting is time-consuming and lacks practical application value. In the real world, the target domain is usually constantly changing. For example, in an autonomous driving system, the driving environment may change continuously (such as sunny, cloudy, rainy, snowy weather, etc.). Therefore, how to enable the model to have adaptive capabilities in a dynamically changing environment has become the focus of research. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention proposes a semantic segmentation method for difficult scenarios based on test-time domain adaptation. The method includes: obtaining an image to be semantically segmented, and preprocessing the image; inputting the preprocessed image into a semantic segmentation model for difficult scenarios to obtain a segmentation result;
[0006] Training the semantic segmentation model for difficult scenarios includes: obtaining image samples and preprocessing the images in the samples; screening the preprocessed images using the intermediate domain generation method of prediction entropy to obtain the source domain and the intermediate domain; inputting the images in the source domain and the intermediate domain into an improved neural network for feature extraction, and inputting the extracted features into the segmentation model to obtain the semantic segmentation results for difficult scenarios; constructing the loss function of the model based on the semantic segmentation results for difficult scenarios, adjusting the model parameters, and completing the training of the model when the loss function converges.
[0007] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any of the above-mentioned difficult scenario semantic segmentation methods based on test-time domain adaptation.
[0008] To achieve the above object, the present invention also provides a difficult scenario semantic segmentation device based on test-time domain adaptation, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the difficult scenario semantic segmentation device based on test-time domain adaptation executes any of the above-mentioned difficult scenario semantic segmentation methods based on test-time domain adaptation.
[0009] Advantages of the present invention:
[0010] The difficult scenario semantic segmentation method based on test-time domain adaptation of the present invention reduces the adaptation difficulty of continuous domain shift and mixed domain shift by performing style transfer on the high-entropy samples of the test data stream using the selected low-entropy samples. Considering that the statistics of the batch normalization layer are based on the training data and cannot adapt to the distribution changes of the test data, an instance-level normalization layer is introduced to dynamically adjust the weights between the global mean and global variance trained in the source domain and the mean and variance of the samples input into the model, so as to better adapt. Description of the drawings
[0011] Figure 1 It is a flowchart of the test-time domain adaptation method for generating the intermediate domain based on prediction entropy of the present invention;
[0012] Figure 2 It is a flowchart of the test-time adaptation method based on the adaptive instance-level normalization layer of the present invention;
[0013] Figure 3 It is a comparison diagram of the test-time domain adaptation segmentation effect of the present invention. Detailed implementation manners
[0014] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] A semantic segmentation method for difficult scenarios based on test-time domain adaptation, as Figure 1 shown, the method includes: obtaining an image to be semantically segmented, and preprocessing the image; inputting the preprocessed image into a semantic segmentation model for difficult scenarios to obtain a segmentation result; training the semantic segmentation model for difficult scenarios includes: obtaining image samples, and preprocessing the images in the samples; using an intermediate domain generation method based on prediction entropy to screen the preprocessed images to obtain a source domain and an intermediate domain; inputting the images in the source domain and the intermediate domain into an improved neural network for feature extraction, and inputting the extracted features into the segmentation model to obtain a semantic segmentation result for difficult scenarios; constructing a loss function of the model according to the semantic segmentation result for difficult scenarios, adjusting the model parameters, and completing the training of the model when the loss function converges.
[0016] In this embodiment, an intermediate domain generation method based on prediction entropy is adopted, and the proposed adaptive instance-level normalization layer is used to replace the batch normalization layer of the original model. Specifically, it includes: first, during the adaptation process of the test data stream, sample screening is performed based on prediction entropy, and low-entropy samples are gradually screened out during the test-time domain adaptation process. The minimum low-entropy sample at this moment is used as a temporary source domain and participates in the subsequent test-time domain adaptation. According to the selected low-entropy samples, as the temporary source domain, style transfer is performed on the high-entropy samples, so that the statistical features of the high-entropy samples shift towards the source domain method, thereby generating an intermediate domain.
[0017] In this embodiment, screening the image using an intermediate domain generation method based on prediction entropy includes: screening low-entropy samples through a minimum entropy sample buffer in the test data stream, using the selected low-entropy samples as the samples in the test data stream whose feature distribution is closest to the source domain before the current time step, and performing arbitrary style transfer on the image samples after the time step.
[0018] For a neural network, when the input is image samples from different domains, the batch normalization layer will have limitations. When the test data and the training data have different distributions, these fixed statistics learned from the source domain will lead to performance degradation. Therefore, it is necessary to redesign the normalization layer to gradually adapt the statistics (mean and variance) input to the normalization layer for images from different domains, so as to achieve a significant performance improvement when only using a small amount of unlabeled test data. As Figure 2As shown, the present invention introduces an instance-level normalization layer to replace the batch normalization layer and disables the exponential moving average. The sliding mean and sliding variance required for updating the normalization are only related to the global mean and global variance of the model trained in the source domain, as well as the mean and variance of the samples currently input into the model. This ensures that when the target domain changes, the sliding mean and sliding variance of the previous samples in the normalization layer do not affect the normalization calculation for the new target domain.
[0019] In this embodiment, processing the input image using an improved neural network includes: performing arbitrary style transfer on the input image at the current time step with the lowest entropy image before this time step to obtain an intermediate domain image after style transfer, and then inputting it into the segmentation model after replacing the batch normalization layer with an adaptive instance-level normalization layer. An appropriate balance parameter α is selected through the predicted entropy of this image, thereby adjusting the weights between the sliding mean and sliding variance trained in the source domain and the mean and variance of the samples input into the model. Finally, the weighted sliding mean and sliding variance that need to be normalized for this sample are obtained by addition, and appropriate normalization is performed.
[0020] The segmentation model processes the feature maps by: extracting multi-level feature maps from the input image through a pre-trained convolutional neural network feature extractor. These feature maps contain semantic information at different scales, which helps to capture details and global context in the scene. In the test phase, the model is fine-tuned through a test-time domain adaptation method to adapt to the distribution of the target domain.
[0021] The goal of the present invention is to continuously compare the predicted entropy of samples in an online manner to select the optimal low-entropy samples for a continuously changing target domain during inference. In the intermediate domain generation method based on predicted entropy, there is an image buffer that stores only a pair of the minimum entropy sample and the value of the minimum entropy. The test data stream x T {x T 1, x T 2, x T 3, …, x T n} of the target domain is provided in sequence. When adaptation starts, the sample x T 1 is input into the model f θ After obtaining the prediction result p1, the average predicted entropy E1 is calculated, and then the value of the average predicted entropy of the sample x T 1 is saved into the image buffer. Subsequently, when it comes to the sample x T 2, the sample pair x T 2 in the image buffer is style-transferred and then input to obtain the prediction result p2. The average predicted entropy E2 calculated from p2 is compared with the value E1 of the minimum entropy saved in the image buffer. If E2 is less than the value E1 saved in the buffer, the style-transferred x TThe value of the predicted entropy E2 of 2 and p2 is saved to the image buffer, replacing the previous x T 1 and p1. Conversely, if E2 is greater than E1, the data in the image buffer is not changed, and x T 1 for subsequent x T 3 is input into the model f after style transfer θ for inference. When t = n, the sample x T n is input into the network. At this time, the sample with the minimum entropy stored in the buffer can be considered as the local optimal sample for the intermediate domain generation method based on predicted entropy at the moment t = n. This sample is closer to the source domain and more similar to the data distribution of the source domain for the first n target domain samples
[0022] After the style - transferred sample is input into the segmentation model, it enters the instance - level normalization layer that replaces the batch normalization layer. This standard layer stores the global mean and global variance trained on the source domain. When the sample enters the instance - level normalization layer, the mean and variance of the sample input to the model are calculated. According to the difference in the data distributions of the two, the parameter α is obtained to dynamically adjust the weights between the global mean and global variance trained on the source domain and the mean and variance of the sample input to the model, so as to add them to obtain the weighted sliding mean and sliding variance for standardizing the sample, thereby adapting to the new data distribution and reducing the problems of covariate shift and data distribution difference
[0023] In the test data stream, the predicted entropy of the image sample input into the segmentation model is used as the loss. Since the predicted entropy of high - entropy samples is not accurate and may lead to noisy gradients, when the predicted entropy of the image sample is greater than half of the maximum entropy (the maximum entropy value when the distribution is uniform, equal to the logarithm of the number of classes), the model does not perform backpropagation and directly outputs the prediction result. The output result is as Figure 3 shown
[0024] In an embodiment of the present invention, the present invention further includes a computer - readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements any one of the above - mentioned semantic segmentation methods for difficult scenarios based on test - time domain adaptation
[0025] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above - mentioned method embodiments can be completed by hardware related to a computer program. The aforementioned computer program can be stored in a computer - readable storage medium. When the program is executed, it executes the steps including the above - mentioned method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disk, or optical disk and other various media that can store program codes
[0026] A semantic segmentation device for difficult scenarios based on test-time domain adaptation, comprising a processor and a memory; the memory is used for storing a computer program; the processor is connected to the memory and is used for executing the computer program stored in the memory, so that the semantic segmentation device for difficult scenarios based on test-time domain adaptation executes any of the above-mentioned semantic segmentation methods for difficult scenarios based on test-time domain adaptation.
[0027] Specifically, the memory includes: various media such as ROM, RAM, magnetic disks, USB flash drives, memory cards or optical discs that can store program codes.
[0028] Preferably, the processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0029] The above-mentioned embodiments further elaborate on the purpose, technical solutions and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A semantic segmentation method for difficult scenarios based on test-time domain adaptation, characterized in that Including: Obtain an image to be semantically segmented, and preprocess the image; Input the preprocessed image into a difficult scene semantic segmentation model to obtain a segmentation result; Training the difficult scene semantic segmentation model includes: obtaining image samples and preprocessing the images in the samples; Adopt the intermediate domain generation method of prediction entropy to screen the preprocessed images to obtain a source domain and an intermediate domain; input the images in the source domain and the intermediate domain into an improved neural network for feature extraction, input the extracted features into a segmentation model to obtain a difficult scene semantic segmentation result; set model convergence conditions, compare the difficult scene semantic segmentation result with the convergence conditions, if the convergence conditions are not met, adjust the model parameters, if the convergence conditions are met, complete the training of the model.
2. The semantic segmentation method for difficult scenarios based on test time domain adaption according to claim 1, wherein Preprocessing the image includes: performing filtering and enhancement processing on the image, and cropping the enhanced image to obtain the preprocessed image.
3. A semantic segmentation method for difficult scenarios based on test time domain adaptation according to claim 1, characterized in that Adopting the intermediate domain generation method of prediction entropy to screen the image includes: screening low-entropy samples through a minimum entropy sample buffer in the test data stream, taking the screened low-entropy samples as the samples whose feature distribution before the current time step in the test data stream is closest to the source domain, and performing arbitrary style transfer on the image samples after the time step.
4. A semantic segmentation method for difficult scenarios based on test time domain adaptation according to claim 1, characterized in that The improved neural network includes: replacing the batch normalization layer in the neural network with an adaptive instance-level normalization layer.
5. A semantic segmentation method for difficult scenarios based on test time domain adaptation according to claim 4, characterized in that, Processing the input image by using the improved neural network includes: performing arbitrary style transfer on the input image of the current time step and the lowest entropy image before this time step to obtain an intermediate domain image after style transfer; inputting the intermediate domain image into a segmentation model after replacing the batch normalization layer with an adaptive instance-level normalization layer, selecting an appropriate balance parameter α through the prediction entropy of the image, so as to adjust the weights between the sliding mean and sliding variance trained in the source domain and the mean and variance of the samples input into the model, and finally adding them to obtain the weighted sliding mean and sliding variance that need to be normalized for the sample, and normalizing the weighted sliding mean and sliding variance.
6. A semantic segmentation method for difficult scenarios based on test time domain adaptation according to claim 4, characterized in that The segmentation model processes the feature map by: extracting multi-level feature maps from the input image by using a feature extractor; in the test stage, fine-tuning the model through a test-time domain adaptation method to adapt to the distribution of the target domain.
7. A semantic segmentation method for difficult scenarios based on test time domain adaptation according to claim 1, characterized in that The convergence condition is: the prediction entropy after inputting the image sample into the segmentation model. When the prediction entropy of the image sample is greater than half of the maximum entropy, the model does not perform backpropagation and directly outputs the prediction result.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor to implement the difficult scene semantic segmentation method based on test-time domain adaptation according to any one of claims 1 to 7.
9. A semantic segmentation device for difficult scenarios based on test-time domain adaptation, characterized in that, Including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that a difficult scene semantic segmentation device based on test-time domain adaptation executes the difficult scene semantic segmentation method based on test-time domain adaptation according to any one of claims 1 to 7.