Controllable diffusion-assisted pipeline to improve unsupervised domain adaptation for semantic segmentation

The controllable diffusion-assisted pipeline addresses the challenge of semantic segmentation under adverse weather by generating high-fidelity synthetic images using a UDA segmentor and UDAControlNet, enhancing model performance and adaptability without labeled target data.

WO2025162656A1PCT designated stage Publication Date: 2025-08-07HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/087158
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2024-12-18
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Conventional semantic segmentation methods degrade significantly under adverse weather conditions due to limited training data and lack of cross-domain adaptation frameworks, posing safety risks for autonomous systems and being time-consuming and error-prone.

Method used

A controllable diffusion-assisted pipeline integrating a UDA segmentor and a Text-to-image Diffusion Model (UDAControlNet) to generate high-fidelity synthetic images by leveraging target priors, multi-scale training, and a residual condition fusion module, enabling closed-loop learning without labeled target data.

Benefits of technology

Enhances robustness and adaptability of semantic segmentation systems by generating pseudo-target images that align with target domain characteristics, bridging the domain gap and improving performance in adverse weather scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024087158_07082025_PF_FP_ABST
    Figure EP2024087158_07082025_PF_FP_ABST
Patent Text Reader

Abstract

A controllable diffusion-assisted pipeline to improve Unsupervised Domain Adaptation (UDA) for semantic segmentation comprising being configured to train an Unsupervised Domain Adaptation (UDA) model for semantic segmentation, the pipeline comprising a UDA segmentor (ℳ) which is pretrained on a labelled source domain image set (X S ) associated with a semantic label map (Y s ) and an unlabelled target domain image set (X t ), and a Text-to-image Diffusion Model, a trainable diffusion network (UDAControlNet), for image generation, where the UDA segmentor (ℳ) is configured to predict a target prior (I) which is a set of labels for the unlabeled target domain image set (X t ) and the UDAControlNet is configured to utilize the target prior (I) as the main condition for training.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]CONTROLLABLE DIFFUSION-ASSISTED PIPELINE TO IMPROVE UNSUPERVISED DOMAIN ADAPTATION FOR SEMANTIC SEGMENTATIONTECHNICAL FIELDThe present disclosure relates generally to the field of computer vision and machine learning; and more specifically, to acontrollable diffusion-assisted pipeline to improve Unsupervised Domain Adaptation (UDA) for semantic segmentation under adverse weather conditions. BACKGROUNDWith the rapid advancement in autonomous driving and robotic navigation systems, semantic segmentation has become apivotal component for scene understanding and decision-making. The conventional semantic segmentation methods havedemonstrated promising results under favorable weather conditions. However, the performance of the conventional methodssignificantly degrades under adverse weather and illumination conditions, such as rain, snow, fog, and nighttime scenarios.Such performance degradation poses substantial safety risks for autonomous systems (e.g., robotic navigation systems)operating in challenging environmental conditions. The collection and annotation of data under adverse weather conditionspresents significant challenges. Furthermore, the data acquisition in adverse weather conditions exposes human annotators tosafety risks, and also, the reduced visibility in such weather conditions (i.e., snowy, foggy, night-time, etc.) makes theannotation process more time-consuming and error-prone. Moreover, the high costs associated with collecting and annotating large-scale datasets under various adverse weather conditions have hindered the development of robust semantic segmentation models for autonomous systems. Unsupervised Domain Adaptation (UDA) has emerged as a promising approach to address the challenges of semanticsegmentation under adverse weather conditions. The UDA operates by transferring knowledge from a labelled source domainimage set having images of clear weather conditions to an unlabeled target domain image set having images of adverse weatherconditions. Traditional UDA approaches often utilize Generative Adversarial Networks (GANs) to synthesize cross-domain data for training. However, these approaches face limitations in handling multiple weather conditions simultaneously and oftenproduce low-fidelity synthetic images due to training from scratch with limited data. Thus, there exists a technical problem ofgenerating high-fidelity synthetic images under multiple adverse weather conditions due to limited training data and a lack of cross-domain adaptation frameworks for semantic segmentation tasks. Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with conventional UDA methods for semantic segmentation under adverse weather conditions. SUMMARYThe present disclosure provides a controllable diffusion-assisted pipeline to improve Unsupervised Domain Adaptation (UDA)for semantic segmentation under adverse weather conditions. The present disclosure provides a solution to the existing problemof generating high-fidelity synthetic images under multiple adverse weather conditions due to limited training data and a lackof cross-domain adaptation frameworks for semantic segmentation tasks. An aim of the present disclosure is to provide asolution that at least partially overcomes the problems encountered in the prior art, which often struggled with the diverseweather conditions and complex scenes encountered in urban driving scenarios. The object of the present disclosure is achievedby integrating target priors, multi-scale training, and a residual condition fusion (RCF) module with a diffusion network(UDAControlNet) which is configured to generate controllable high-fidelity and cross-domain synthetic images datasets withenhanced quality. The object of the present disclosure is achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims.In one aspect, the present disclosure provides a controllable diffusion-assisted pipeline to improve Unsupervised DomainAdaptation (UDA) for semantic segmentation comprising being configured to train an Unsupervised Domain Adaptation(UDA) model for semantic segmentation, the pipeline comprising a UDA segmentor (ℳ) which is pretrained on a labelledsource domain image set (^^) associated with a semantic label map (^^) and an unlabelled target domain image set (^^), anda Text-to-image Diffusion Model, UDAControlNet, for image generation, where the UDA segmentor (ℳ) is configured topredict a target prior (^^) which is a set of labels for the unlabeled target domain image set (^^) and the UDAControlNet isconfigured to utilize the target prior (^^) as the main condition for training.The controllable diffusion-assisted pipeline addresses the key challenge of the domain gap between source and target data andthe absence of labelled target domain data in UDA. The pipeline overcomes this challenge by leveraging the target prior thatis the predicted semantic labels for the unlabeled target domain images generated by the pre-trained UDA segmentor. By usingthe target prior as the primary training condition for the UDAControlNet, the UDAControlNet can capture the overall class- wise distribution of the target domain, enabling the generation of pseudo-target images conditioned on source domain labels. These generated images can then be used to further train the UDA segmentation model, improving its performance on the targetdomain without requiring any labelled target data. This closed-loop learning approach enhances the overall robustness andadaptability of the semantic segmentation system, making the UDAControlNet well-suited for real-world deployment scenarios with limited labelled data availability.In an implementation form, the UDAControlNet is further configured to utilize edge detection for preparing a set of sketches^^ = ^(^^), and train the UDAControlNet conditioned also on the set of sketches as a secondary condition.The utilization of edge detection sketches as the secondary training condition of the UDAControlNet allows the model to better capture the structural information of the target domain images. By incorporating both the semantic target prior and the structural edge sketches, the UDAControlNet can generate more faithful target domain images that align with the actual visual characteristics of the target environment, improving the overall performance of the unsupervised domain adaptation for semantic segmentation.In a further implementation form, the UDAControlNet is further configured to utilize edge detection based on a pre-trainededge detection module (ℋ). The use of the pre-trained edge detection module in the UDAControlNet provides an efficient and reliable way to capture thestructural information of the target domain images. By leveraging the pre-trained edge detection module, the UDAControlNetcan generate high-quality edge sketches without the requirement to train the edge detection component from scratch,streamlining the overall pipeline and improving its effectiveness in unsupervised domain adaptation for semantic segmentation.In a further implementation form, the UDAControlNet is further configured to prioritize input conditions on the semanticmodality. By emphasizing the semantic aspects, the UDAControlNet can generate target domain images that are better aligned with the desired semantic segmentation objectives, improving the overall performance of the unsupervised domain adaptation pipeline. In a further implementation form, the UDAControlNet is further configured to prioritize input conditions on the semanticmodality utilize a residual condition fusion, RCF, module by, in each training iteration, input ^^ (^^ ∈ ^^) to a semantic encoder(^^^^) to produce segmental conditions (^^^^^ ), input ℎ^ (where ℎ^ ∈ ^^) to a structure encoder (^^^^) to produce structuralconditions (^^^^^ ), applying an attention module, (^) to the structural conditions, (^^^^^ ) to remove non-salient noisy artifactsand enhance salient structures in the structural conditions (^^^^^ ) and then fusing the conditions by elementwise adding thestructural conditions (^^^^^ ) to the semantic conditions (^ ^^^^ ) to provide early condition fusion (^ ^^ ), processing the earlycondition fusion (^ ^ ^), in a convolution layer (^) and applying a skip connection from the segmental conditions (^^^^ ^), to theearly condition fusion (^^^ ), and enlarging the target domain by providing a low-resolution version and / or a high-resolution ofthe input images (^^), and label with the same corresponding label (^^).The utilization of the RCF module leads to effectively prioritize the input conditions on the semantic modality by separatelyencoding the semantic and structural conditions from the target prior and target images, respectively, and applying attention toenhance the salient structural features and fusing the semantic and structural conditions through elementwise addition and utilizing the skip connection to retain the semantic information. This multi-stage condition fusion process ensures that the semantic information takes precedence, while still incorporating relevant structural cues. Additionally, the use of multi-scale training of the UDAControlNet, by including both low-resolution and high-resolution versions of the input images, further enhances its ability to handle small objects and distant features, which is required for robust performance in adverse driving conditions. In a further implementation form, the UDAControlNet is further configured to enhancing the default Blip prompts by mappingthe target prior (^^) into extra label-guided prompts based on a class-wise mapping.The mapping of the target prior (^^) into extra label-guided prompts allows for more targeted and informative prompts that gobeyond generic descriptions, potentially leading to more accurate and nuanced image generation. In a further implementation form, the UDAControlNet is further configured to enhancing the prompts based on a target sub- domain indicating a weather condition by name. By incorporating weather-specific conditions (such as night or foggy) directly into the prompts, the UDAControlNet improves the semantic relevance and contextual understanding of the image generation process. This targeted approach leads to mitigate the generic nature of default prompts, allowing for more precise and condition-aware image generation in challenging environmental scenarios. In a further implementation form, the UDAControlNet is further configured to applying DDIM sampling to acquire pseudo target images conditioned mainly on source labels to increase the data diversity. By utilizing DDIM sampling with source labels to generate pseudo target images, the UDAControlNet effectively increases data diversity and manifests the ability to overcome domain shift challenges. This approach allows for the synthetic expansion of training data, potentially improving the model's generalization and performance across different domains by creating additional varied training samples.In a further implementation form, the UDAControlNet is further configured to augmenting the unlabeled target domain imageset (^^) conditioned mainly on the target prior (^^) using both the final and the initial checkpoint of the UDAControlNet.By leveraging both the initial and final checkpoints of the UDAControlNet to augment the unlabeled target domain imagesbased on the target prior information, the approach can effectively generate more diverse and representative synthetic data. This approach leads to bridge the domain gap by creating additional training samples that capture the nuanced characteristics of the target domain, potentially improving the model's performance and generalization capabilities. By combining the baseline loss with the segmentation loss computed from dynamically generated pseudo labels, the UDA Segmentor (ℳ) can effectively leverage both the original data distribution and the newly created pseudo-labelled information.This approach enables more comprehensive learning, potentially improving the model's ability to adapt and generalize acrossdifferent domains by simultaneously minimizing initial classification errors and refining segmentation performance. The proposed approach allows for a more reliable and selective supervision strategy by dynamically validating source label pixels through predicted label matching. This approach enables the UDA Segmentor to use only high-confidence source labels for segmentation loss calculation, potentially reducing noise and improving the overall learning process by ensuring that only well-aligned labels contribute to the model's training. By implementing the double consistency check to validate predicted labels against source labels, the UDA Segmentor can enhance the reliability and accuracy of its label matching process. This approach provides an additional layer of verification that helps to reduce potential misclassifications and ensures a more robust method of determining label correspondence, ultimately improving the model's learning precision. By using the confidence threshold (λ) to validate label matching, the UDAControlNet can dynamically filter and accept predictions that are sufficiently close to source labels. This approach allows for a flexible and adaptive method of label validation, enabling the model to handle subtle variations while maintaining a high level of prediction accuracy and reliability. It is to be appreciated that all the aforementioned implementation forms can be combined. It has to be noted that all devices, elements, circuitry, units and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims. Additional aspects, advantages, features and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow. BRIEF DESCRIPTION OF THE DRAWINGS The summary above, as well as the following detailed description of illustrative embodiments, is better understood when readin conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions ofthe disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods andinstrumentalities disclosed herein. Moreover, those skilled in the art will understand that the drawings are not to scale. Whereverpossible, like elements have been indicated by identical numbers. Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein: FIG.1 is a block diagram of a controllable diffusion-assisted pipeline to improve unsupervised domain adaptation (UDA) for semantic segmentation, in accordance with an embodiment of the present disclosure;FIG. 2 illustrates a training process for training of a controllable diffusion model, in accordance with an embodiment of the present disclosure; FIG.3 illustrates a training process based on label-guided prompts and multi-scale images for training of a controllable diffusion model, in accordance with an embodiment of the present disclosure; FIG.4 illustrates a training process based on introducing structure guidance during training of a controllable diffusion model, in accordance with an embodiment of the present disclosure; FIG. 5 illustrates pseudo target data generation after training of a controllable diffusion model, in accordance with an embodiment of the present disclosure; FIG.6 illustrates a process for improving the domain adaptation in semantic segmentation under adverse weather conditions, in accordance with an embodiment of the present disclosure; and FIG.7 illustrates an exemplary implementation scenario of controllable diffusion-assisted unsupervised domain adaptation for cross-weather semantic segmentation, in accordance with an embodiment of the present disclosure. In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow,the non-underlined number is used to identify a general item at which the arrow is pointing.DETAILED DESCRIPTION OF EMBODIMENTS The following detailed description illustrates embodiments of the present disclosure and ways in which they can beimplemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art wouldrecognize that other embodiments for carrying out or practicing the present disclosure are also possible. FIG.1 is a block diagram of a controllable diffusion-assisted pipeline to improve unsupervised domain adaptation (UDA) forsemantic segmentation, in accordance with an embodiment of the present disclosure. With reference to FIG. 1, there is showna block diagram 100 of a controllable diffusion-assisted pipeline 102 to improve Unsupervised Domain Adaptation (UDA) for semantic segmentation. The controllable diffusion-assisted pipeline 102 includes a UDA segmentor 104, a text-to-imagediffusion model 106, and a trainable diffusion network (UDAControlNet) 108. The UDA segmentor 104 is pretrained on alabelled source domain image set 110 which is associated with a semantic label map 112, and an unlabeled target domain imageset 114. Further, the UDA segmentor 104 is configured to predict a target prior 116. There is further shown a Residual ConditionFusion (RCF) module 118, a semantic encoder 120, a structure encoder 122, an attention module 124, and a convolution layer126 which are used in training of the UDAControlNet 108.The controllable diffusion-assisted pipeline 102 may be referred to as a method designed to improve Unsupervised DomainAdaptation (UDA) for semantic segmentation. The controllable diffusion-assisted pipeline 102 employs the UDAControlNet108 which is used in conjunction with the text-to-image diffusion model 106 (e.g., stable diffusion model) for generating high-fidelity, weather-controllable pseudo target domain data under adverse weather conditions. This synthetic data can then be used to enhance the training of the UDA segmentation model, effectively bridging the gap between the source and target domains and improving performance on the target domain.The UDA segmentor 104 (may be represented as ℳ) refers to a key component of the controllable diffusion-assisted pipeline102, designed to perform semantic segmentation in an Unsupervised Domain Adaptation (UDA) framework. The UDA segmentor 104 is initially pretrained on the labelled source domain image set 110 and is subsequently refined to adapt to theunlabeled target domain image set 114 by leveraging pseudo-target data generated within the controllable diffusion-assistedpipeline 102. The UDA segmentor 104 is engineered to overcome cross-domain challenges by leveraging techniques like self- training, pseudo-labelling, and sophisticated loss functions that enable learning across different environmental conditions while maintaining high semantic segmentation accuracy.The text-to-image diffusion model 106 refers to an advanced generative model within the controllable diffusion-assistedpipeline 102 that synthesizes high-fidelity pseudo-target images by translating semantic and structural conditions into realistic visual representations. The text-to-image diffusion model 106 operates by progressive denoising a latent representation, guidedby semantic prompts, structural guidance, and enhanced conditions like multi-resolution inputs and domain-specific prompts.For example, the text-to-image diffusion model 106 utilizes techniques, such as Denoising Diffusion Implicit Models (DDIM)sampling to generate detailed and weather-controllable target domain images under adverse weather conditions, like fog, rain,snow or night.The UDAControlNet 108 is an advanced deep-learning model designed to enhance diffusion-based image generation for UDAin semantic segmentation under adverse weather conditions. The architecture of the UDAControlNet 108 enables the generationof objects at diverse scales, differentiates overlapped instances, and augments semantic label information across weather sub-domains like night, foggy, and rainy conditions. The UDAControlNet 108 may be referred to as a control mechanism used inconjunction with the text-to-image diffusion model 106 to provide more precise guidance and control during the imagegeneration process. The UDAControlNet 108 is designed to add additional conditioning to the text-to-image diffusion model 106, allowing for more nuanced control over image generation. While the text-to-image diffusion model 106 (e.g., Stable Diffusion) generate images from text prompts, the UDAControlNet 108 adds an extra layer of control by incorporating additional input signals or constraints.The labelled source domain image set 110 (may also be represented as ^^) refers to a collection of annotated images capturedunder clear and standard weather conditions, specifically designed to provide comprehensive semantic ground truth labels for training the UDA segmentor 104. Examples of the labelled source domain image set 110 may include, but are not limited to, an image set comprising high-resolution urban driving scenes with precise pixel-wise annotations for objects such as vehicles, pedestrians, road markings, traffic signs, buildings, and vegetation, captured under optimal lighting and weather conditions.The aforementioned image set may be used as a foundational dataset for generating pseudo-target domain images across variousadverse weather conditions.The semantic label map 112 (may also be represented as ^^) refers to a detailed pixel-wise annotation that assigns specific classlabels to every pixel within an image, providing a comprehensive semantic understanding of the scene's visual composition.Examples of the semantic label map 112 may include, but are not limited to, color-coded representations where different pixelregions are classified into predefined semantic categories, such as road, sidewalk, building, car, pedestrian, traffic sign, sky,vegetation, with each color uniquely representing a distinct object class or scene element.The unlabeled target domain image set 114 (may also be represented as ^^) refers to a collection of raw images captured underchallenging and adverse weather conditions, lacking semantic annotations and representing the challenging target domain forthe UDA segmentor 104. Examples of the unlabeled target domain image set 114 may include, but are not limited to, drivingscene images captured during night time, foggy, rainy, or snowy conditions, and the like. Such images have degraded visibility,complex scene structures, and varying illumination levels that significantly differ from the source domain's clear weather images.The target prior 116 (may also be represented as ^^) refers to a set of preliminary semantic predictions generated by a pretrainedUDA segmentor (i.e., the UDA segmentor 104), providing noisy but informative label information for the unlabeled targetdomain image set 114 captured under adverse weather conditions. Examples of the target prior 116 may include, but are notlimited to, predicted semantic segmentation labels obtained from an already existing Mutual Information Consistency (MIC)model applied to the unlabeled target domain image set 114, which capture the overall class-wise distribution and semanticstructure of scenes under challenging weather scenarios.The RCF module 118 refers to a specialized architectural component designed to integrate multiple input modalities, such assemantic and structural conditions, during the training of the UDA diffusion model (i.e., the UDAControlNet 108). The RCFmodule 118 processes and combines the information from semantic predictions and structural guidance to prioritize semanticalignment while enhancing salient structural features. The controllable diffusion-assisted pipeline 102 to improve Unsupervised Domain Adaptation (UDA) for semantic segmentation comprising being configured to train an Unsupervised Domain Adaptation (UDA) model for semanticsegmentation, the controllable diffusion-assisted pipeline 102 comprising the UDA segmentor 104 (ℳ) which is pretrained onthe labelled source domain image set 110 (^^) associated with the semantic label map 112 (^^) and the unlabelled target domain image set 114 (^^). The labelled source domain image set 110 (^^) may be referred to as a collection of images and their corresponding annotated semantic segmentation labels, which belong to the source domain. The source domain typically represents the "clear" or "normal" conditions, where data is more readily available and can be accurately annotated. The semantic label map (or semantic segmentation labels) 112 (^^) provide pixel-level annotations, where each pixel in the imageis assigned a class label (e.g., road, building, vehicle, pedestrian, etc.). The unlabeled target domain image set 114 (^^) may bereferred to as a collection of images that belong to the target domain (i.e., adverse weather conditions), but do not have anyassociated semantic labels or ground truth annotations. The pretraining of the UDA segmentor 104 (ℳ) using the labelled source domain image set 110 (^^) associated with the semantic label map 112 (^^) and the unlabelled target domain image set114 (^^) generates the target prior 116 (^^^). The target prior 116 (^^^) may be referred to as predicted semantic labels for theunlabeled target domain image set 114 (^^). The target prior information is then used as a key input to the UDAControlNet108, which tries to learn to generate synthetic target domain images conditioned on this prior knowledge. The controllable diffusion-assisted pipeline 102 further comprises the text-to-image diffusion model 106, the UDAControlNet108, for image generation. The UDA segmentor 104 (ℳ) is configured to predict the target prior 116 ^^^^^ which is the set oflabels for the unlabeled target domain image set 114 (^^), and the UDAControlNet 108 is configured to utilize the target prior116 ^^^^^ as the main condition for training. The UDA segmentor 104 (ℳ) first processes the unlabelled target domain imageset 114 (^^) and generates predicted labels, creating the target prior 116 ^^^^^. The target prior 116 ^^^^^ may provide valuablesemantic information about the target domain. During training of the UDAControlNet 108, the unlabelled target domain imageset 114 (^^) and the target prior 116 ^^^^^ are utilized, and the use of any source-specific features is avoided to ensure thatgenerated images follow the target domain distribution only. The UDAControlNet 108 is configured to leverage the text-to-image diffusion model 106 (e.g., stable diffusion model) to generate high-fidelity, weather-controllable pseudo target domain images. The generated synthetic images are conditioned on source domain labels (i.e., the labelled source domain image set110 (^^)), taking advantage of the shared label space between the source and target domains (i.e., the unlabeled target domainimage set 114 (^^)).FIG. 2 illustrates a training process for training of a controllable diffusion model, in accordance with an embodiment of thepresent disclosure. FIG. 2 is described in conjunction with elements from FIG. 1. With reference to FIG. 2, there is shown atraining process 200 for training of the controllable diffusion model that is the UDAControlNet 108. The training process 200includes key steps, namely, an offline preparation step 202, utilizing an encoder 204 (^), the semantic encoder 120 (^^^^) andthe UDAControlNet 108 (may be mathematically denoted as ^), and a stable diffusion model 206 (may be mathematicallydenoted as θ).The training process 200 starts with the offline preparation step 202 in which the UDA segmentor 104 (ℳ) is pretrained usingthe labelled source domain image set 110 (^^) associated with the semantic label maps 112 (^^) and the unlabeled target domainimage set 114. During this initial phase, the UDA segmentor 104 (ℳ) generates the target prior 116 (^^) which may be referredto as predicted labels for the unlabeled target domain image set 114 (^^). The labelled source domain image set 110 (^^) maybe a set of raw Red-Green-Blue (RGB) images used in pre-training of the UDA segmentor 104 (ℳ). By virtue of pre-trainingthe UDA segmentor 104 (ℳ), the UDA segmentor 104 (ℳ) may also be referred to as a pre-trained UDA segmentor.Thereafter a subset (i.e., ^^ ∈ ^^ ) of the target prior 116 (^^) is used as an input to the semantic encoder 120 (^^^^) to generatesegmental conditions (^^^^^ ) which are also used in training of the UDAControlNet 108. Furthermore, there is shown theencoder 204 (^) which may be configured to encode the unlabeled target domain image set 114 (^^) and the labelled sourcedomain image set 110 (^^) into latent representation ^^(^). The encoder 204 (^) acts as a feature extraction module that convertsthe input image sets (i.e., the unlabeled target domain image set 114 (^^) and the labelled source domain image set 110 (^^))into a format suitable for use by the stable diffusion model 206 (θ). The stable diffusion model 206 (θ) is an example of thetext-to-image diffusion model 106. During training of the UDA model, the encoded latent representation ^^(^) undergoesprogressive noising through Gaussian noise addition Єτ~^(μτ, στ2) and is further used as an input to the UDAControlNet 108(^) and the stable diffusion model 206 (θ). The stable diffusion model 206 (θ) is configured to generate high-fidelity imagesfor the unlabeled target domain taking the encoded latent representation ^^(^) with the added Gaussian noise Єτ~^(μτ, στ2)over multiple time steps ^.FIG. 3 illustrates a training process based on label-guided prompts and multi-scale images for training of a controllable diffusionmodel, in accordance with an embodiment of the present disclosure. FIG. 3 is described in conjunction with elements fromFIGs. 1 and 2. With reference to FIG. 3, there is shown a training process 300 for training of the controllable diffusion modelthat is the UDAControlNet 108. The training process 300 is based on utilizing the label-guided prompts and multi-resolution, cropped and resized versions of raw RGB images. The training process 300 includes key steps, namely, the offline preparationstep 202, utilizing the encoder 204, the semantic encoder 120 (^^^^) and the UDAControlNet 108 (^), and the stable diffusionmodel 206 (θ). The training process 300 is the similar to the training process 200 (shown and described in FIG. 2) except withthe use of label-guided prompts and multi-resolution, cropped and resized versions of raw RGB images for an enhanced training of the UDAControlNet 108 (^). Conventionally, Blip, which is a prompt generator used by a conventional model, namely ControlNet. The Blip prompt generator can provide captions for many natural images to improve the training quality of the diffusion model. However, the Blip prompt generator can produce inaccurate captions for driving images in adverse conditions. This happens either by mentioning objects that do not exist or sometimes provides only a few vague words to describe an image. Consequently, the use of noisy information may degrade the training quality of the text-to-image diffusion models and make the generated output less aligned with the input conditions. Therefore, Blip is considered as a frozen tool and fine-tuning of the Blip is impractical. The only feasible solution is to make the Blip less influential during prompt generation. In order to achieve this solution, thetarget prior 116 (^^) is mapped class-wisely into extra prompt to provide more semantic guidance in the description.In accordance with an embodiment, the UDAControlNet 108 is further configured to enhancing the default Blip prompts bymapping the target prior 116 into extra label-guided prompts based on a class-wise mapping. The UDAControlNet 108 (^) isconfigured to enhance the quality and alignment of default prompts generated by Blip, particularly in scenarios involvingdiverse and adverse conditions. To achieve this, the target prior 116 (^^) is mapped into additional guided prompt through aclass-wise mapping strategy. The class-wise mapping introduces detailed semantic context into the prompts, enabling the stablediffusion model 206 to consider specific features of the input conditions. To further enrich the prompt, target sub-domain details (e.g., "night" or "foggy") are injected into the default Blip prompts and an enhanced prompt (mathematically may be denoted as ^^^ ) is generated, which can be represented as, ^ ^^ = ^^^ − ^^^^^^ + ^^^^ ^^^^^^ + ^^^^^ − ^^^^^^^^.The enhanced prompt’s reliance on Blip output is reduced and its correlation with the input semantic conditions is built. Theenhanced prompt (^^^ ) is processed using Clip, which supports the alignment of the textual and visual features more effectively.Furthermore, a dropout technique is applied to the enhanced prompt (^ ^ ^) with a low probability during training to furtherencourage the stable diffusion model 206 to learn from the input conditions. The use of the enhanced prompt (^^^ ) makes thetraining of the UDAControlNet 108 (^) more efficient for data under adverse conditions. In accordance with an embodiment, the UDAControlNet 108 is further configured to enhancing the prompts based on a targetsub-domain indicating a weather condition by name. The UDAControlNet 108 (^) incorporates a mechanism applied to theenhanced prompts (^^^) by incorporating target sub-domain information that explicitly identifies weather conditions by name (e.g., "foggy," "night," or "rainy"). By explicitly including the target sub-domain (weather condition) in the prompts, the UDAControlNet 108 (^) enhances the descriptive power of the captions, making them more informative and aligned with the scene. For example, if the target sub-domain is "foggy," the enhanced prompt (^^^ ) might read: "a foggy image showing carson the road with poor visibility." This additional layer of context ensures that the prompts are weather-specific and semanticallyenriched. The enhanced prompt (^^^) is passed through Clip, which aligns the textual sub-domain information with the visual features of the input image. By embedding weather-specific details, the prompts provide clearer guidance during trainingcausing the stable diffusion model 206 to learn for generating the outputs that are more consistent with the target sub-domain.Additionally, the multi-scale training is introduced to ensure the UDAControlNet 108 (^) learns to handle both local and globalfeatures. Due to the existence of small objects that are far away from the camera and their degraded visibility in adverse drivingconditions, the conventional ControlNet is observed to struggle in handling those small objects. Therefore, in order to resolve the issue of handling small objects with degraded visibility conditions, the proposed model that is the UDAControlNet 108 (^)is trained on multiple scales of raw input images. In the multi-scale training of the UDAControlNet 108 (^), multiple-resolution,cropped and resized versions of raw RGB images are used. In an exemplary scenario, a low-resolution version of a raw inputimage can be created through resizing, while a high-resolution version of the raw input image can be created through randomcropping. Both versions may be included in the mini-batch for each training iteration, enabling the UDAControlNet 108 (^)to learn from different image scales simultaneously.FIG. 4 illustrates a training process based on introducing structure guidance during training of a controllable diffusion model,in accordance with an embodiment of the present disclosure. FIG.4 is described in conjunction with elements from FIGs.1, 2and 3. With reference to FIG. 4, there is shown a training process 400 for training of the controllable diffusion model that isthe UDAControlNet 108. There is further shown a pre-trained edge detection module 402 and a skip connection 404. Thetraining process 400 is similar to the training process 300 except the introduction of structure guidance via the RCF module 118 (of FIG.1) during training of the UDAControlNet 108.To further enhance the training of the UDAControlNet 108 (^), the stable diffusion model 206 (^) leverages both semanticconditions and structural guidance through sketches (^^), the latter being generated by the pre-trained edge detection (HED) module 402. To effectively combine these inputs, the UDAControlNet 108 (^) employs the RCF module 118, which prioritizes semantic guidance while integrating structural information. The semantic input (^^^) is encoded into segmental conditions (^^^^^) using the semantic encoder 120 (^^^^), while the structural sketches (^^) are encoded into the structural features (^^^^^) using the structure encoder 122 (^^^^). The structural features (^^^^^ ) are refined through the attention module 124 (^) which isconfigured to filter out non-salient noisy artifacts and enhance the required salient structures. The refined structural features(^^^^^) are then combined with the segmental conditions (^^^^^) through element-wise addition, followed by processing with theconvolution layer 126 (^). To maintain the dominance of semantic information, the skip connection 404 is applied, adding theoriginal segmental conditions (^^^^^) to the fused representation, resulting in the fused feature condition (^^^). The overall process of generating the fused feature condition can be represented by Equation (1) ^^ ^^^ ^^ ^^^^ = ^^ ⊕ ^(^ ^^ ⊙ (^ ⊕ ^(^^^^^ )) ⊕ ^^ ) (1)Moreover, regarding the training of the controllable diffusion model that is the UDAControlNet 108 (^), given an encodedfeature (^^(^)) from the target sub-domain, the diffusion algorithm progressively creates noisy encoded features (^^^(^)), where^ represents uniformly sampled timestep from {0, ... , ^}. The UDAControlNet 108 (^) is configured to predict the addedGaussian noise using the enhanced prompt (^^^), the fused conditions (^^^), and the timestep (τ), with the objective of optimizingL2-norm, a loss function. The L2-norm, also known as Euclidean norm function, used for noise prediction. The stable diffusionmodel 206 (^) is guided by a diffusion loss function ^^^^^. The diffusion loss function ^^^^^can be computed using Equation (2) Unlike many popular datasets, the complex scene structure and object appearances of autonomous driving data, as well as their harsh weather conditions make it technically challenging to generate high quality data from the noisy target prior label alone(i.e., the target prior 116 (^^)). Therefore, the set of sketches (^^) are used in addition to the target prior 116 (^^) for anenhanced training of the UDAControlNet 108 (^). In accordance with an embodiment, the UDAControlNet 108 is further configured to prioritize input conditions on the semanticmodality utilize the residual condition fusion (RCF) module 118 by, in each training iteration, input ^^ (^^ ∈ ^^) to the semanticencoder 120 (^^^^) to produce segmental conditions (^^^^^ ), input ℎ^ (where ℎ^ ∈ ^^) to the structure encoder 122 (^^^^) toproduce structural conditions (^^^^^), applying the attention module 124 (^) to the structural conditions (^^^^^) to remove non- salient noisy artifacts and enhance salient structures in the structural conditions (^^^^^), fusing the conditions by elementwise adding the structural conditions (^^^^^) to the semantic conditions (^^^^^) to provide early-condition fusion (^^^ ), processing theearly condition fusion (^^^ ) in the convolution layer 126 (^) and applying the skip connection 404 from the segmentalconditions (^^^^^) to the early condition fusion (^^^) and enlarging the target domain by providing a low-resolution versionand / or a high-resolution of the unlabeled target domain image set 114 (^^) and label with the same corresponding label.For each training iteration, the RCF module 118 processes semantic labels ^^ (^^ ∈ ^^) using the semantic encoder 120 (^^^^)to produce semantic conditions (^^^^^ ) as shown in Equation (4). Similarly, the structural input (ℎ^) is encoded into structuralconditions (^^^^^ ) using the structure encoder 122 (^^^^ ), can be represented as ^^^^^ = ^^^^(ℎ^). These structural conditions arerefined by applying the attention module 124 (^) to remove non-salient noisy artifacts and enhance important salient structurefeatures, represented as ^^^^^(refined) = ^(^^^^^). The refined structural conditions and semantic conditions are elementwise added to form fused conditions (^^^ ), expressed as ^^^= ^^^^^ + ^^^^^(refined). The fused conditions ^^^are processed throughthe convolution layer 126 (^), represented as ^ ^^ = ^(^^^). The skip connection 404 from the semantic conditions (^^^^^) to the fused conditions (^^^ ) ensures semantic dominance. The target domain is enlarged by employing both low-resolution andhigh-resolution versions of the unlabeled target domain image set 114 (^^) and the input semantic labels (^^) to capture varyingfeature scales, enabling more comprehensive training of the UDAControlNet 108 (^). The combination of the structural andsemantic inputs is required for addressing the challenges of complex scene geometries and adverse visual conditions in domainadaptation. However, maintaining semantic dominance ensures that the UDAControlNet 108 (^) remains focused on label-aligned segmentation tasks. The RCF module 118 allows the stable diffusion model 206 (^) to integrate structural informationwithout overshadowing semantic inputs, improving the generation of pseudo-target domain data while preserving its semanticaccuracy. The inclusion of multi-resolution inputs ensures the UDAControlNet 108 (^) captures both local and global features,enabling an enhanced domain adaptation performance.Thus, the controllable diffusion-assisted pipeline 102 addresses the challenge of semantic segmentation in adverse weather conditions, where collecting and annotating data is difficult and expensive. To address this, the UDAControlNet 108, a novelframework, configured to use large text-to-image diffusion models (e.g. the stable diffusion model 206) to generate synthetictraining data. This approach aims to bridge the gap between clear weather (i.e., the labelled source domain image set 110 (^^)and adverse weather (i.e., the unlabeled target domain image set 114 (^^) semantic segmentation by creating controllable,realistic synthetic images that can improve model performance across different environmental conditions. The key innovationis using diffusion models to generate high-quality, condition-specific training data, effectively expanding the model's (i.e., the stable diffusion model 206) ability to perform semantic segmentation under challenging weather scenarios. The key challenge in UDA for semantic segmentation under adverse weather conditions is the lack of labelled target domain data. This makes it difficult to directly transfer knowledge from the labelled source domain image set 110 (^^) to the unlabeled target domain image set 114 (^^), as the distribution shift between the domains can significantly degrade performance. To address this challenge, the proposed framework (i.e., the UDAControlNet 108) leverages the capabilities of large-scale text-to- image diffusion models (i.e., the stable diffusion model 206), to generate high-quality, controllable pseudo target domain data. The core ideas include utilizing target domain priors (i.e., the target prior 116 (^^)), enhanced prompts, the RCF module 118, during training of the UDAControlNet 108, which has been described in detail, for example, in FIGs.2, 3, and 4, respectively. The multi-scale training of the UDAControlNet 108 has also been described in detail, for example, in FIG.3. After training of the UDAControlNet 108, the framework is able to generate high-fidelity, weather-controllable pseudo target domain data, described in detail, for example, in FIG.5. This synthetic data can then be used to enhance the training of the UDA segmentation model (i.e., the UDA segmentor 104), effectively bridging the gap between the source and target domains and improvingperformance on the target domain. The key advantage of this novel framework is to overcome the fundamental challenge oflacking target domain labels, enabling powerful diffusion-based data generation for UDA in adverse weather scenarios. FIG. 5 illustrates pseudo target data generation after training of a controllable diffusion model, in accordance with an embodiment of the present disclosure. FIG.5 is described in conjunction with elements from FIGs.1, 2, 3, and 4. With referenceto FIG. 5, there is shown a pseudo target data generation process 500 utilizing the UDAControlNet 108 (^). There is furthershown a decoder 502 (may be represented as ^) and the UDAControlNet 108 (^).After training of the UDAControlNet 108, the trained UDAControlNet enables the stable diffusion model 206 for pseudo targetdata generation. The pseudo target data generation process 500 begins with utilizing the semantic label maps 112 (^^), whichconsist of semantic information and structural guidance (^^) (e.g., edge maps). To generate pseudo target data resembling thetarget domain, the UDAControlNet 108 (^) is fueled by rich hidden knowledge of the source domain (i.e., the labelled sourcedomain image set 110 (^^)), which further enables the UDAControlNet 108 (^) to facilitate the stable diffusion model 206 toconditioned outputs with high fidelity. By virtue of using the target prior 116 (^^) during training of the UDAControlNet 108, the UDAControlNet 108 is trained to provide the outputs resembling with the target domain distribution. However, since a common label space is shared between the source and target domains, synthesizing pseudo target images from more accuratesource labels thus, becomes feasible, especially with additional guidance of the structure information that is less sensitive acrossdomains. Therefore, conditioned on source labels, DDIM sampling is applied to the UDAControlNet 108 (more specifically,trained UDAControlNet) to acquire pseudo target images conditioned on source labels, according to Equation (3) Different from training, ^^~^(0, ^) and the time step ^ gradually reduces from ^ to 1 throughout the diffusion denoisingprocess. The decoder 502 (^) adopted from the source domain to project the resulting latent representations into pixel space.By randomly specifying the target sub-domain (^) in our enhanced prompt, the UDAControlNet 108 is configured to generatecontrolled high-fidelity outputs under various adverse conditions, such as snowy, rainy, night, foggy. Thus, a dataset (^^^^^ ∣^^ , ^^) is obtained which enriches data diversity and is further used to refine the UDA segmentor 104 (ℳ). To emphasize themain input condition, the dataset (^^^^^ ∣ ^^) is used instead of (^^^^^).The decoder 502 (^) refers to a component used in generative models like stable diffusion models and its various variants,responsible for converting latent representations back into interpretable outputs such as images or text. The decoder 502 (^) isused to project the resulting latent representations from the DDIM sampling step into the desired pixel space, which are furtherused in generation of high-fidelity outputs resembling the unlabeled target domain image set 114 (^^).In accordance with an embodiment, the UDAControlNet 108 is further configured to applying DDIM sampling to acquirepseudo target images conditioned mainly on source labels to increase the data diversity. The DDIM sampling may be referredto as a process used to synthesize high-quality images. The UDAControlNet 108 utilizes DDIM sampling to create pseudotarget images through a systematic process. Firstly, a random latent representation ^^~^(0, I) is initialized. The stablediffusion model 206 then, iteratively refines this representation by reducing the time step τ from ^ to 1, using a denoisingprocess that gradually transforms the noisy latent input into a high-quality image. The process is guided by source labels, whichprovide semantic and structural information, and enhanced Blip prompts, which specify the desired conditions (e.g., weathertypes) for further customization. The decoder 502 (^) from the stable diffusion model 206 (^) translates the refined latentrepresentation into pixel space, producing the final pseudo target images. These images, along with their corresponding sourcelabels, form the pseudo target dataset (^^^^^ ∣ ^^, ^^), which is then used to train and refine the UDA segmentor 104 (ℳ),resulting in an enhanced performance and adaptability of the UDA segmentor 104 (ℳ).In accordance with an embodiment, the UDAControlNet 108 is further configured to augmenting the target domain datasetconditioned mainly on the target prior 116 (^^) using both the final and the initial checkpoint of the UDAControlNet 108. Initialcheck point of the UDAControlNet 108, and the stable diffusion model 206 are configured to generate pseudo target domainFIG. 6 illustrates a process for improving the domain adaptation in semantic segmentation under adverse weather conditions,in accordance with an embodiment of the present disclosure. FIG. 6 is described in conjunction with elements from FIGs. 1,2, 3, 4, and 5. With reference to FIG. 6, there is shown a process 600 for improving the domain adaptation in semanticsegmentation under adverse weather conditions. The process 600 is related to performance enhancement of the UDAControlNet108 via refinement training using the generated pseudo target data.After training of the UDAControlNet 108 (as shown and described in FIGs. 2, 3 and 4) and pseudo target data generation usingthe UDAControlNet 108 and the stable diffusion model 206 (as shown and described in FIG. 5), the remaining step is toimprove domain adaptation, for which the process 600 is used. The training of the UDA segmentor 104 (ℳ) using the generatedpseudo target data can further raise the upper limit of domain adaptive semantic segmentation in adverse weather conditions. can be mapped back to the same semantic label through the UDA segmentor 104 (ℳ), this indicates that the mapping is meaningful and the semantic prediction is likely correct. The supervised cross-entropy segmentation loss applied on the generated data (^^^^^) can be written according to Equation (5)ℒ^^^^^^^ = −^^^^^^, ^^ [^^^log(ℳ(^^^^^))](^,^,^)(5) Therefore, the combined loss of the UDAControlNet 108 can be represented as In accordance with an embodiment, the UDA segmentor 104 (ℳ) is further configured to determining that the predicted label(^^^^^) matches a source label (^^) by performing a double consistency check. The double consistency check is a method usedto verify the reliability of the predicted label (^^^^^). The double consistency check involves few steps: (i) The source groundtruth labels (^^) are used to generate pseudo-target images through a domain mapping process. (ii) The UDA segmentor 104(ℳ) predicts labels for the pseudo-target data. These predicted labels (^^^^^) are then checked to see if they match the originalsource label (^^). If the pseudo-target image generated from (^^) can be mapped back to (^^) through the UDA segmentor 104(ℳ), it validates the consistency of the mapping and predictions. This forms a "label-to-label" consistency check, ensuring that the semantic prediction is accurate and reliable. When the predicted label (^^^^^) satisfies the double consistency check, the source label (^^) is deemed suitable for supervision. The segmentation loss (ℒ^^^^^^^) is then calculated to update the UDAControlNet 108. If the check fails, the label is ignored to avoid introducing noise into the training process. In accordance with an embodiment, the UDAControlNet 108 is further configured to determining that the predicted label matches a source label (^^) by determining if the predicted label differs from the label (^^) by less than a confidence threshold(λ). The UDAControlNet 108 checks if the difference between the predicted label (^^^^^) and the source GT label (^^) fallsbelow the confidence threshold (λ), the discrepancy is assumed to be due to undertraining or minor errors, and the ground truth label (^^) is retained for supervision. This helps to refine the model’s predictions while avoiding the inclusion of noisy data. On the other hand, if the confidence difference exceeds (λ), the prediction is deemed unreliable, and the corresponding data is excluded from the training process. This mechanism ensures that only high-quality and trustworthy data contribute to the training, improving the model’s segmentation accuracy in both the source and target domains.FIG. 7 illustrates an exemplary implementation scenario of controllable diffusion-assisted unsupervised domain adaptation forcross-weather semantic segmentation, in accordance with an embodiment of the present disclosure. FIG. 7 is described inconjunction with elements from FIGs. 1, 2, 3, 4, 5, and 6. With reference to FIG. 7, there is shown an implementation scenario700 of controllable diffusion-assisted unsupervised domain adaptation for cross-weather semantic segmentation. The implementation scenario 700, may be for example, a robot navigation system. The robot navigation system is composed of various modules, such as data processing, perception, task analysis, motion planning, and execution. The perception module is crucial, as it requires to accurately identify and segment the surrounding environment to enable effective navigation,particularly in adverse weather conditions. However, due to the shortage of labelled data for training, the segmentationcomponent of the perception module may lack the required robustness to handle challenging environmental conditions, suchas rain, fog, or low visibility. This can lead to inaccurate semantic information being fed into the downstream modules,potentially compromising the overall navigation performance. This is where the ControlUDA framework can be leveraged toenhance the robustness of the segmentation component. The ControlUDA pipeline can serve as an automatic data generation and labelling module, training a UDA segmentation model (i.e., UDAControlNet 108) that can accurately perform segmentation even in adverse weather conditions. While the UDA segmentation model may be computationally heavy and not suitable for real-time execution on the embedded system, a knowledge distillation process can be employed to transfer the acquired knowledge to the lightweight, embedded segmentation module without adding any additional complexity to the overall system. By integrating the ControlUDA framework, the robot navigation system can benefit from improved semantic understanding of the environment, even in challenging conditions, without the requirement to make significant changes to the existing system architecture. This can lead to more reliable and robust navigation capabilities, enhancing the overall performance and effectiveness of the robot navigation system. Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments and / or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the present disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.

Claims

CLAIMS 1. A controllable diffusion-assisted pipeline (102) to improve Unsupervised Domain Adaptation, UDA, for semanticsegmentation comprising being configured to train an Unsupervised Domain Adaptation, UDA, model for semanticsegmentation, the controllable diffusion-assisted pipeline (102) comprisinga UDA segmentor, ℳ (104) which is pretrained on a labelled source domain image set, ^^ (110) associatedwith a semantic label map, ^^ (112) and an unlabeled target domain image set, ^^ (114), anda Text-to-image Diffusion Model (106), and a trainable diffusion network, UDAControlNet (108), forimage generation, wherein the UDA segmentor, ℳ (104) is configured to predict a target prior, ^^ (116) which is a set of labels for theunlabeled target domain image set, ^^ (114) and the UDAControlNet (108) is configured to utilize the target prior,^^ (116) as the main condition for training.

2. The controllable diffusion-assisted pipeline (102) according to claim 1, wherein the UDAControlNet (108) is furtherconfigured to utilize edge detection for preparing a set of sketches ^^ = ^(^^), andtrain the UDAControlNet (108) conditioned also on the set of sketches as a secondary condition.

3. The controllable diffusion-assisted pipeline (102) according to claim 2, wherein the UDAControlNet (108) is furtherconfigured to utilize edge detection based on a pre-trained edge detection module, ℋ (402).

4. The controllable diffusion-assisted pipeline (102) according to any preceding claim, wherein the UDAControlNet(108) is further configured to prioritize input conditions on the semantic modality.

5. The controllable diffusion-assisted pipeline (102) according to claim 4, wherein the UDAControlNet (108) is furtherconfigured to prioritize input conditions on the semantic modality utilize a residual condition fusion, RCF, module (118)by, in each training iteration, input ^^ (^^ ∈ ^^) to a semantic encoder, ^^^^ (120) to produce segmental conditions, ^^^^^, input ℎ^ (where ℎ^ ∈ ^^) to a structure encoder, ^^^^ (122) to produce structural conditions, ^^^^^, applying an attention module, ^ (124) to the structural conditions, ^^^^^to remove non-salient noisy artifacts and enhance salient structures in the structural conditions, ^^^^^and then fusing the conditions by elementwise adding the structural conditions, ^^^^^to the semantic conditions, ^ ^^^ ^to provide early condition fusion, ^^^, processing the early condition fusion, ^^^ , in a convolution layer, ^ (126) andapplying a skip connection (404) from the segmental conditions, ^ ^^^^ , to the early condition fusion, ^^^, and enlarging the target domain by providing a low-resolution version and / or a high-resolution of the input images, ^^, and label with the same corresponding label, ^^.

6. The controllable diffusion-assisted pipeline (102) according to claim 5, wherein the UDAControlNet (108) is furtherconfigured to enhancing the default Blip prompts by mapping the target prior, ^^(116) into extra label-guided prompts based on a class-wise mapping.

7. The controllable diffusion-assisted pipeline (102) according to claim 5 or 6, wherein the UDAControlNet (108) isfurther configured to enhancing the prompts based on a target sub-domain indicating a weather condition by name.

8. The controllable diffusion-assisted pipeline (102) according to any preceding claim, wherein the UDAControlNet(108) is further configured to applying DDIM sampling to acquire pseudo target images conditioned mainly on source labels to increase the data diversity.

9. The controllable diffusion-assisted pipeline (102) according to claim 8, wherein the UDAControlNet (108) is furtherconfigured to augmenting the unlabeled target domain image set, ^^ (114) conditioned mainly on the target prior, ^^ (116)using both the final and the initial checkpoint of the UDAControlNet (108).

10. The controllable diffusion-assisted pipeline (102) according to any preceding claim, wherein the UDA Segmentor,ℳ (104) is further configured tocalculate a baseline loss, ℒ^^^^,generating new pseudo labels, ^^^,calculating a segmentation loss,based on the new pseudo labels, ^^^and train based on the sum of the baseline loss, ℒ^^^^, and the segmentation loss, ℒ^^^^ ^^^.

11. The controllable diffusion-assisted pipeline (102) according to claim 10, wherein the UDA Segmentor, ℳ (104) isfurther configured to determining whether a source label pixel can be used for supervision by outputting a predicted label, ^^^^^ anddetermining if the predicted label, ^^^^^matches a source label, ^^, and if so, relying on the source label, ^^ and calculating the segmentation loss, ℒ^^^^^^^.

12. The controllable diffusion-assisted pipeline (102) according to claim 11, wherein the UDA Segmentor, ℳ (104) isfurther configured to determining that the predicted label, ^^^^^ matches the source label, ^^ by performing a doubleconsistency check.

13. The controllable diffusion-assisted pipeline (102) according to claim 11 or 12, wherein the UDAControlNet (108) isfurther configured to determining that the predicted label, ^^^^^ matches the source label, ^^ by determining if thepredicted label, ^^^^^ differs from the source label, ^^ by less than a confidence threshold, λ.