Text-based bayesian zero-shot domain adaptation training method for image segmentation

CN118865407BActive Publication Date: 2026-09-11HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411151077.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-09-11
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

[0009]本发明目的是为了解决现有的零样本域适应方法主要集中于优化经验风险最小化目标,通常依赖于基于有限提示的离散增强训练,难以充分捕捉目标域的复杂性,从而削弱了迁移模型的有效性的问题;本发明提供了一种用于图像分割的基于文本的贝叶斯零样本域适应训练方法

Benefits of technology

[0052] The framework of this invention mainly consists of a residual distribution model, a semantic segmentation head, and a loss function for the first training stage. Losses in the second training phase The trainable residual distribution model is represented as a normal distribution, with its mean and variance learned through backpropagation. The entire residual distribution model is trained end-to-end using a text-based loss function, which more accurately aligns the learned distribution with the actual residual distributions between the target and source domains. This efficiently adapts the model to the target domain images, improving its performance in the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118865407B_ABST
    Figure CN118865407B_ABST
Patent Text Reader

Abstract

A text-based Bayesian zero-shot domain adaptation training method for image segmentation belongs to the zero-shot domain adaptation field in computer vision. The existing zero-shot domain adaptation methods mainly focus on optimizing the empirical risk minimization target, usually rely on discrete augmented training based on limited prompts, and are difficult to fully capture the complexity of the target domain, thereby weakening the effectiveness of the migration model. The present application regards the parameter learning process in zero-shot domain adaptation as a variational inference problem from the Bayesian perspective. Specifically, the residual between the source domain and the target domain is probabilistically modeled, and uncertainty related to the domain gap is introduced, thereby reducing the dependence of the model on specific weights and improving the performance of the model in the target domain. The present application is mainly used for semantic segmentation of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of zero-sample domain adaptation in computer vision. Background Technology

[0002] The development of autonomous driving technology is rapid, but the complex and ever-changing road conditions in its application environments pose significant challenges to its practical application. In real-world driving scenarios, vehicles may encounter various complex road conditions and traffic situations, including but not limited to nighttime driving, rainy or snowy weather, rural roads, and urban rush hour traffic. These situations not only increase the difficulty of autonomous driving systems but also place higher demands on their perception, decision-making, and control capabilities.

[0003] However, while autonomous driving systems perform well in common road conditions, they may underperform in some rare scenarios. Data for these rare scenarios is often difficult to collect because they occur infrequently and involve complex and variable conditions. Examples include driving data under extreme weather conditions and traffic conditions in specific geographical locations. Collecting and labeling this data often requires significant human and material resources, resulting in high costs and low efficiency. Therefore, improving the performance of autonomous driving systems in these rare scenarios, even with sufficient data, has become a pressing issue.

[0004] To address this issue, Zero-Shot Domain Adaptation (ZSDA) has gradually attracted researchers' attention. The core idea of ​​ZSDA is to construct a model that performs well even in unseen data domains, enabling it to generalize well to data in the target domain even in the absence of specific scenario data. The advantage of this method is that it does not require target domain data, reducing the cost of data acquisition and annotation, thereby improving the performance of autonomous driving systems in complex scenarios.

[0005] Existing ZSDA methods primarily focus on optimizing the Empirical Risk Minimization (ERM) objective. The ERM objective aims to improve the overall performance of the model by minimizing the error on the training data. However, this approach typically relies on discrete prompts, guiding the model to make predictions in unseen data domains by providing specific, discrete hints. While this approach performs well in some cases, its effectiveness often falls short when facing complex target domains. Specifically:

[0006] First, discrete cue words struggle to capture the complexity of the target domain. Rare situations in autonomous driving scenarios often exhibit high complexity and diversity, which simple discrete cue words cannot fully describe. For example, in nighttime driving, factors such as lighting conditions, traffic conditions, and road conditions all significantly impact driving safety, and the interactions between these factors are complex and variable, which a single cue word cannot comprehensively cover. This leads to predictive biases in models when facing these complex situations, reducing their performance.

[0007] Secondly, ERM-based methods are often influenced by the training data during optimization, making them prone to overfitting. This means the model performs well on the training data but poorly on unseen data domains. This further limits its application in autonomous driving scenarios.

[0008] In summary, existing zero-shot domain adaptation methods (ZSDA methods) mainly focus on optimizing the goal of minimizing empirical risk. They usually rely on discrete augmentation training based on limited cues, which makes it difficult to fully capture the complexity of the target domain, thus weakening the effectiveness of the transfer model. These problems urgently need to be solved. Summary of the Invention

[0009] The purpose of this invention is to address the problem that existing zero-shot domain adaptation methods mainly focus on optimizing the goal of minimizing empirical risk, and usually rely on discrete augmentation training based on limited clues, which makes it difficult to fully capture the complexity of the target domain, thereby weakening the effectiveness of the transfer model; this invention provides a text-based Bayesian zero-shot domain adaptation training method for image segmentation.

[0010] A text-based Bayesian zero-shot domain adaptation training method for image segmentation, comprising the following steps:

[0011] From a Bayesian perspective, the learning process of parameters for domain adaptation models in image segmentation tasks is as follows: A residual distribution model is obtained by probabilistically modeling the residuals between the source and target domain images using a learnable distribution for domain adaptation; the residual distribution model and the semantic segmentation head constitute the image segmentation model; the residual distribution model includes t residual distributions with different text descriptions.

[0012] The first training phase trains the residual distribution model:

[0013] Set the optimization objective of the residual distribution model and construct the loss function for the first training phase. Afterwards, based on Supervise the training of the residual distribution model using N labeled data points.

[0014] The second training phase involves training both the residual distribution model and the semantic segmentation head simultaneously.

[0015] Constructing the second training phase loss based on The image segmentation model is trained by using N labeled data to train the residual distribution model and the semantic segmentation head.

[0016] Preferably, there are N labeled data. In the middle, x i For the i-th input data, I i For the i-th source domain image, y i Let r be the ground truth value of the source domain image after the i-th semantic annotation. * It is the theoretically optimal residual distribution. Let p be a set consisting of t target domain image text descriptions. s Provide a text description of the source domain image;

[0017] The optimization objective r for each residual distribution in the residual distribution model * for:

[0018]

[0019] P(y i |x i ,r j ) for in r j and labeled data x i Under the constraints y i The conditional probability, For x i ,y i Let r be the mathematical expectation of the random variable. j Let L be the set of L features sampled from the j-th residual distribution, where each feature is the residual between the target domain image feature and the corresponding source domain image feature, j = 1, 2, ..., t.

[0020] Preferably,

[0021] in, For distance loss, The loss is related to distribution optimization.

[0022] Preferably,

[0023] Where λ is the weight. The sum of cross-entropy loss and dice loss. Let d be the KL divergence. φ (r j ) is r j Follows a standard normal distribution, d γ (r j ) is rj It follows a learnable Gaussian distribution, j = 1, 2, ..., t.

[0024] Preferably,

[0025]

[0026] To compare the losses, Let d be the KL divergence. φ (r j ) is r j Follows a standard normal distribution, d γ (r j ) is r j It follows a learnable Gaussian distribution, j = 1, 2, ..., t.

[0027] Preferably,

[0028]

[0029] for and Cosine loss between, v s These are deep features of the source domain image. To synthesize a deep feature set of the target domain image using residual distribution, To utilize v s The target domain image deep features are synthesized with text description features, where || ||1 is the Manhattan distance and || ||2 is the Euclidean distance.

[0030] Preferably,

[0031]

[0032]

[0033]

[0034] r j ~d γ (r j );

[0035] in, To encode textual descriptive features into feature vectors, Let be the set of textual description features of t target domains. These are textual description features of the source domain. To perform deep feature extraction on F, where F represents the shallow features of the source domain image, and r is the depth of the feature extraction process. 1 to r t The set that constitutes the composition.

[0036] Preferably,

[0037]

[0038] Where τ is the temperature coefficient. for The i-th deep feature of the target domain image synthesized using the residual distribution. for The text description features of the i-th target domain. for The textual description features of the k-th target domain.

[0039] Preferably, based on Supervise the given N labeled data The specific process of training the residual distribution model includes:

[0040] A1. Using image shallow feature extractor For each input data x i Shallow feature extraction is performed to obtain the shallow features F of the source domain image;

[0041] A2. Obtaining r through the residual distribution model j , The mean is μ j The variance is Σ j Gaussian distribution;

[0042] A3, r 1 to r t The results are added to F and then fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain a set of deep features of the target domain image synthesized using residual distribution.

[0043] F is fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain the deep features v of the source domain image. s ;

[0044] A4. Through a text encoder For p s Encode the text description features of the source domain to obtain them. via text encoder right Encode the data to obtain a set of textual description features for t target domains.

[0045] A5, Through In Supervision and The cosine distance between them In Supervision and The degree of feature similarity between them, and the supervision d φ (r j ) and d γ (r j The similarity between the parameters is used to update the parameters of the residual distribution model, thereby enabling the training of the residual distribution model.

[0046] Preferably, based on Using the given N labeled data The specific process of training the residual distribution model and the semantic segmentation head is as follows:

[0047] B1. Using image shallow feature extractor For each input data x i Shallow feature extraction is performed to obtain the shallow features F of the source domain image;

[0048] B2. Obtaining r through the residual distribution model j , The mean is μ j The variance is Σ j Gaussian distribution;

[0049] B3, r 1 to r t The results are added to F and then fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain a set of deep features of the target domain image synthesized using residual distribution.

[0050] B4. Deep feature set of target domain image synthesized using residual distribution through semantic segmentation head pairs. Make predictions to determine the values ​​of each labeled data x. i The corresponding segmentation results, through Supervise each labeled data x i The corresponding segmentation result and the semantically annotated ground truth y of the source domain image i The difference between the parameters is used to update the parameters of the residual distribution model and the semantic segmentation head, thereby enabling the training of the residual distribution model and the semantic segmentation head.

[0051] The beneficial effects of this invention are:

[0052] The framework of this invention mainly consists of a residual distribution model, a semantic segmentation head, and a loss function for the first training stage. Losses in the second training phase The trainable residual distribution model is represented as a normal distribution, with its mean and variance learned through backpropagation. The entire residual distribution model is trained end-to-end using a text-based loss function, which more accurately aligns the learned distribution with the actual residual distributions between the target and source domains. This efficiently adapts the model to the target domain images, improving its performance in the target domain.

[0053] This invention proposes a text-based Bayesian zero-shot domain adaptation training method (ProGBA) for image segmentation. This method treats the parameter learning process in zero-shot domain adaptation as a variational inference problem from a Bayesian perspective. By probabilistically modeling the residuals between the source and target domains, it introduces uncertainty related to the domain gap, thereby reducing the model's dependence on specific weights and improving the model's performance in the target domain. Specifically, this invention obtains a residual distribution model by probabilistically modeling the residuals between the source and target domains. This residual distribution model undergoes a loss during the first training phase. Losses in the second training phase Supervised training enables effective coverage of the feature space of the residuals between the source and target domain features. For example, sampling from the residual distribution corresponding to the text description "driving in the rain" yields residual features corresponding to different amounts of precipitation, such as "heavy rain," "moderate rain," and "light rain." In contrast, existing domain adaptation methods trained based on the goal of minimizing empirical risk can only obtain residual features corresponding to one amount of precipitation for the text description of "driving in the rain."

[0054] During the second stage of training, this invention utilizes the regularization capability of Bayesian methods to optimize the domain-adaptive representation space. Specifically, the residual features, added to the shallow features F of the source domain image, are obtained by applying the residual distribution r... j Random sampling is used to obtain data. This noisy random sampling introduces uncertainty related to the domain gap during the second-stage training process, thereby reducing the model's dependence on specific weights. This helps to mitigate the risk of overfitting in image segmentation models and improves the model's performance in the target domain.

[0055] The core of this invention lies in two main innovations:

[0056] First, domain transfer from the source domain to the target domain is modeled as a probability distribution, rather than a fixed discrete change. This modeling approach reduces the risk of overfitting by capturing the randomness and uncertainty of domain transfer, allowing the model to adapt more broadly to changes between different domains, thereby improving the model's generalization ability. In practical applications, this means that the model can handle unseen target domains more effectively and exhibit stronger robustness.

[0057] Secondly, this invention proposes a novel loss function based on Evidence Lower Bound (ELBO). This loss function promotes a learnable residual distribution that closely approximates the actual domain gaps, thus more accurately reflecting the relationships between domains during optimization. In this way, the model can effectively capture domain gaps and adapt to the target domain during training, achieving more efficient zero-shot domain adaptation. Attached Figure Description

[0058] Figure 1 This is a schematic diagram illustrating the principle of the first training phase;

[0059] Figure 2 This is a schematic diagram illustrating the principle of the second training phase.

[0060] In the attached diagram, For shallow image feature extractors, For shallow image feature extractors, For text encoders. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0063] This invention aims to enhance the perception and parsing capabilities of models in autonomous driving scenarios for unseen target domain images, thereby improving the overall safety of the system. It proposes a text-based Bayesian zero-shot domain adaptation training method (ProGBA) for image segmentation. This method treats the parameter learning process in zero-shot domain adaptation as a variational inference problem from a Bayesian perspective. Specifically, by probabilistically modeling the residuals between the source and target domains, it introduces uncertainties related to the domain gaps, thus reducing the model's dependence on specific weights and improving the model's performance in the target domain. Specifically:

[0064] See Figure 1 and Figure 2 This embodiment describes a text-described guided Bayesian zero-shot image segmentation model training method, which includes the following steps:

[0065] From a Bayesian perspective, the learning process of parameters for domain adaptation models in image segmentation tasks is as follows: A residual distribution model is obtained by probabilistically modeling the residuals between the source and target domain images using a learnable distribution for domain adaptation; the residual distribution model and the semantic segmentation head constitute the image segmentation model; the residual distribution model includes t residual distributions with different text descriptions.

[0066] The first training phase trains the residual distribution model:

[0067] Set the optimization objective of the residual distribution model and construct the loss function for the first training phase. Afterwards, based on Supervise the given N labeled data Training the residual distribution model;

[0068] x i For the i-th input data, I i For the i-th source domain image, y i Let r be the ground truth value of the source domain image after the i-th semantic annotation. * It is the theoretically optimal residual distribution. Let p be a set consisting of t target domain image text descriptions. s Provide a text description of the source domain image;

[0069] To enhance the learning distribution so that it more closely approximates the actual distribution of the residuals between the target and source domains, an optimization objective r is defined for each residual distribution in the residual distribution model. * for:

[0070]

[0071] P(y i |x i ,r j ) for in r j and labeled data x i Under the constraints y i The conditional probability, For x i ,y i Let r be the mathematical expectation of the random variable. j Let L be the set of L features sampled from the j-th residual distribution, where each feature is the residual between the target domain image feature and the corresponding source domain image feature, j = 1, 2, ..., t;

[0072] The second training phase involves training both the residual distribution model and the semantic segmentation head simultaneously.

[0073] Constructing the second training phase loss based on Using the given N labeled data The residual distribution model and semantic segmentation head are trained to complete the training of the image segmentation model.

[0074] In practical applications, the trained image segmentation model extracts shallow and deep features from the image to be semantically segmented in sequence, and then sends it to the trained semantic segmentation head for semantic segmentation to obtain the semantically segmented image.

[0075] This implementation introduces an innovative zero-shot domain adaptation training method. It constructs the residual distribution of source and target domain features using textual descriptions, setting the domain adaptation learning framework as a variational inference problem. The goal of this method is to address the limitations of existing methods when dealing with complex target domains, especially when target domain training data is unavailable.

[0076] For text branches in the source domain, this invention predefines relevant text prompts p. s For example: driving during the day; for the target domain text branch, create corresponding target domain text prompts. For example, nighttime driving, where T represents the number of target domains. This invention approximates the distance between the target domain and source domain image features by subtracting the target domain text cue features from the source domain text features at the feature level, and then uses this approximation with the deep features v of the source domain image. s The deep features of the synthesized target domain image are obtained by addition.

[0077] To ensure that the enhanced image features do not deviate significantly from their original values ​​and to maintain the integrity of the image's semantic content, this invention further introduces distance loss. Constructing the loss in the first training phase

[0078]

[0079] in, For distance loss, The loss is related to distribution optimization.

[0080]

[0081] for and Cosine loss between, v s These are deep features of the source domain image. To synthesize a deep feature set of the target domain image using residual distribution, To utilize v s The target domain image deep features are synthesized with text description features, where || ||1 is the Manhattan distance and || ||2 is the Euclidean distance.

[0082] This invention uses a learnable distribution to probabilistically model the residuals between the source and target domains, represented by the distribution d. γ Expected sampling from d γ r j Able to reflect source domain hints p s and target prompts The semantic distance between them. Its effect on features. For image shallow feature extractor In performing shallow feature extraction, this invention assumes that the intermediate image features in the target domain consist of two parts: image features F from the source domain, and residual enhancement r between the image features of the source and target domains. j .

[0083] Based on this assumption, the zero-sample-domain adaptive training method of the present invention learns the latent distribution d. γ , and have Where μ j and Σ j These are two learnable parameters. Finally, the features of the target domain can be represented by r. r is r 1 to r t The set constituted

[0084]

[0085] This invention uses To approximate the target domain image features by "synthesizing" the residuals of the source and target domain text description features. To utilize v s Deep features of the target domain image synthesized with textual descriptive features:

[0086]

[0087]

[0088]

[0089] r j ~d γ (r j );

[0090] in, To encode textual descriptive features into feature vectors, Let be the set of textual description features of t target domains. These are textual description features of the source domain. To perform deep feature extraction on F, where F represents the shallow features of the source domain image, and r is the depth of the feature extraction process. 1 to rt The set that constitutes the composition.

[0091]

[0092] To compare the losses, Let d be the KL divergence. φ (r j ) is r j Follows a standard normal distribution, d γ (r j ) is r j It follows a learnable Gaussian distribution, j = 1, 2, ..., t.

[0093] In this preferred implementation, the loss for constructing the first training phase is given. The specific process ensures that the residual distribution of the learned data closely approximates the distribution of the contrast between the features of the source and target domains.

[0094] Furthermore, construct the loss function for the second training phase.

[0095]

[0096] Where λ is the weight. The sum of cross-entropy loss and dice loss. Let d be the KL divergence. φ (r j ) is r j Follows a standard normal distribution, d γ (r j ) is r j It follows a learnable Gaussian distribution, j = 1, 2, ..., t.

[0097] As a preferred option, λ is set to 0.01 by default.

[0098] In this preferred implementation, the loss for constructing the second training phase is given. The specific expression is constructed to ensure that the residual distribution does not lose the knowledge learned in the first stage while supervising the semantic segmentation head to output the correct semantic segmentation result.

[0099] See Figure 1 ,

[0100]

[0101] Where τ is a temperature coefficient, its function is to adjust the smoothness of the loss function distribution. for The i-th deep feature of the target domain image synthesized using the residual distribution. for The text description features of the i-th target domain. for The textual description features of the k-th target domain.

[0102] In this preferred embodiment, the following is given: The specific expression makes the deep features of the target domain image synthesized using the residual distribution closer to the deep features of the target domain image synthesized using the deep features of the source domain image and the text description features.

[0103] See Figure 1 In the first training phase, during the training of the residual distribution model, based on Supervise the given N labeled data The specific process of training the residual distribution model includes:

[0104] A1. Using image shallow feature extractor For each input data x i Shallow feature extraction is performed to obtain the shallow features F of the source domain image;

[0105] A2. Regarding the features of the target domain image, this invention first samples and enhances the features r from the learnable target domain distribution. j Specifically, r is obtained through the residual distribution model. j , The mean is μ j The variance is Σ j Gaussian distribution; μ j and Σ j These are learnable parameters;

[0106] A3, r 1 to r t The results are added to F and then fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain a set of deep features of the target domain image synthesized using residual distribution.

[0107] F is fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain the deep features v of the source domain image. s ;

[0108] A4. Through a text encoder For p s Encode the text description features of the source domain to obtain them. via text encoder right Encode the data to obtain a set of textual description features for t target domains.

[0109] A5, Through In Supervision and The cosine distance between them In Supervision and The degree of feature similarity between them, and the supervision d φ (r j ) and d γ (r j The similarity between the parameters is used to update the parameters of the residual distribution model, thereby enabling the training of the residual distribution model.

[0110] As a preferred option, the AdamW optimization algorithm is used to train the residual distribution model, which performs end-to-end training on the learnable distribution. Figure 1 and Figure 2 Chinese text encoder Image shallow feature extractor and image deep feature extractor The parameters are frozen and not used in training.

[0111] Image shallow feature extractor and image deep feature extractor All are image encoders, using either ResNet50 or Vision Transformer for image encoding, and text encoders. This can be achieved using Vision Transformer.

[0112] In this preferred embodiment, the process of training the target domain distribution model in the first training stage is given. In this process, the training of the segmentation head is decoupled from the training of the residual distribution model, reducing the training difficulty of the residual distribution model, and thus obtaining a residual distribution model that is close to the residual distribution of the real source domain image features and target domain image features.

[0113] See Figure 2 During the second training phase, while simultaneously training the residual distribution model and the semantic segmentation head, based on Using the given N labeled data The specific process of training the residual distribution model and the semantic segmentation head is as follows:

[0114] B1. Using image shallow feature extractor For each input data x i Shallow feature extraction is performed to obtain the shallow features F of the source domain image;

[0115] B2. Obtaining r through the residual distribution model j , The mean is μ jThe variance is Σ j Gaussian distribution; μ j and Σ j These are learnable parameters;

[0116] B3, r 1 to r t The results are added to F and then fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain a set of deep features of the target domain image synthesized using residual distribution.

[0117] B4. Deep feature set of target domain image synthesized using residual distribution through semantic segmentation head pairs. Make predictions to determine the values ​​of each labeled data x. i The corresponding segmentation results, through Supervise each labeled data x i The corresponding segmentation result and the semantically annotated ground truth y of the source domain image i The difference between the parameters is used to update the parameters of the residual distribution model and the semantic segmentation head, thereby enabling the training of the residual distribution model and the semantic segmentation head.

[0118] In this preferred embodiment, the image shallow feature extractor and image deep feature extractor All are image encoders, using either ResNet50 or Vision Transformer for image encoding, and text encoders. VisionTransformer can be used.

[0119] In this preferred embodiment, a second training stage is provided to simultaneously train the target domain distribution model and the semantic segmentation head to accelerate the convergence of the semantic segmentation head and obtain a semantic segmentation head with better segmentation performance.

[0120] Principle Analysis: The framework of this invention mainly consists of a residual distribution model, a semantic segmentation head, and a loss function for the first training phase. Losses in the second training phase The trainable residual distribution model is represented as a normal distribution, with its mean and variance learned through backpropagation. The entire residual distribution model is trained end-to-end using a text-based loss function, which more accurately aligns the learned distribution with the actual residual distributions between the target and source domains. This efficiently adapts the model to the target domain images, improving its performance in the target domain.

[0121] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.

Claims

1. A text-based Bayesian zero-shot domain adaptation training method for image segmentation, characterized in that, From a Bayesian perspective, the learning process of parameters for domain adaptation models in image segmentation tasks involves: probabilistically modeling the residuals between the source and target domain images using a learnable distribution to obtain a residual distribution model for domain adaptation; the residual distribution model and the semantic segmentation head constitute the image segmentation model; the residual distribution model includes... The first training phase trains the residual distribution model: Set the optimization objective of the residual distribution model and construct the loss function for the first training phase. Afterwards, based on Supervise the training of the residual distribution model using N labeled data points. N labeled data middle, For the i-th source domain image, for The ground truth of the source domain image after semantic annotation. It is the theoretically optimal residual distribution. For the reason A set consisting of text descriptions of target domain images. Source domain image text description; The optimization objective of each residual distribution in the residual distribution model for: ; In order to be in and labeled data Under the premise of constraints The conditional probability, For Let be the mathematical expectation of the random variable. Let L be the set of L features sampled from the j-th residual distribution, where each feature is the residual between the target domain image feature and the corresponding source domain image feature, j=1,2,... ; based on Supervise the given N labeled data The specific process of training the residual distribution model includes: A1. Using image shallow feature extractor For each Shallow feature extraction is performed to obtain the shallow features F of the source domain image; A2. Obtained through residual distribution model , , The mean is variance is Gaussian distribution; A3, will The results are added to F and then fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain a set of deep features of the target domain image synthesized using residual distribution. ; F is fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain the deep features of the source domain image. ; A4. Through a text encoder right Encode the text description features of the source domain to obtain them. via text encoder right Encoding is performed to obtain a set consisting of textual description features of t target domains. ; A5, Through In Supervision and The cosine distance between them In Supervision and The degree of feature similarity between them, and supervision and The similarity between the parameters is used to update the parameters of the residual distribution model, thereby enabling the training of the residual distribution model; The second training phase involves training both the residual distribution model and the semantic segmentation head simultaneously. Constructing the second training phase loss ,based on The image segmentation model is trained by using N labeled data to train the residual distribution model and the semantic segmentation head.

2. The text-based Bayesian zero-shot domain adaptation training method for image segmentation according to claim 1, characterized in that, ; in, For distance loss, The loss is related to distribution optimization.

3. The text-based Bayesian zero-shot domain adaptation training method for image segmentation according to claim 1, characterized in that, ; in, As weight, The sum of cross-entropy loss and dice loss. Let KL divergence be the KL divergence. for Standard normal distribution for Learnable Gaussian distribution, j=1,2,... .

4. The text-based Bayesian zero-shot domain adaptation training method for image segmentation according to claim 2, characterized in that, ; To compare the losses, Let KL divergence be the KL divergence. for Standard normal distribution for Learnable Gaussian distribution, j=1,2,... .

5. The text-based Bayesian zero-shot domain adaptation training method for image segmentation according to claim 2, characterized in that, ; for and Cosine loss between These are deep features of the source domain image. To synthesize a deep feature set of the target domain image using residual distribution, To utilize Deep features of the target domain image synthesized with textual descriptive features. For Manhattan distance, .

6. The text-based Bayesian zero-shot domain adaptation training method for image segmentation according to claim 5, characterized in that, ; ; , ; in, To encode textual descriptive features into feature vectors, Let be the set of textual description features of t target domains. These are textual description features of the source domain. To Perform deep feature extraction. These are shallow features of the source domain image. for The set that constitutes the composition.

7. The text-based Bayesian zero-shot domain adaptation training method for image segmentation according to claim 4, characterized in that, ; ; in, For temperature coefficient, for The i-th deep feature of the target domain image synthesized using the residual distribution. for The text description features of the i-th target domain. for The textual description features of the k-th target domain.

8. The text-based Bayesian zero-shot domain adaptation training method for image segmentation according to claim 1, characterized in that, based on Using the given N labeled data The specific process of training the residual distribution model and the semantic segmentation head is as follows: B1. Using image shallow feature extractor For each input data Shallow feature extraction is performed to obtain the shallow features F of the source domain image; B2. Obtaining through the residual distribution model , , The mean is variance is Gaussian distribution; B3, will The results are added to F and then fed into the deep feature extractor of the image. Deep feature extraction is performed to obtain a set of deep features of the target domain image synthesized using residual distribution. ; B4. Deep feature set of the target domain image synthesized using residual distribution through semantic segmentation head pairs. Make predictions to forecast the labeled data. The corresponding segmentation results, through Supervise each labeled data The corresponding segmentation results and the ground truth of the semantically annotated source domain image The difference between the parameters is used to update the parameters of the residual distribution model and the semantic segmentation head, thereby enabling the training of the residual distribution model and the semantic segmentation head.