Domain generalization semantic segmentation method based on semantic consistency and style diversity

By introducing semantic consistency and style diversity means in domain generalized semantic segmentation, using CLIP encoder and semantic query enhancer, combined with text-driven style transformation and coordinated weighting loss, the semantic feature confusion and noise problems in the existing methods are solved, and better model promotion capabilities and segmentation consistency are achieved.

CN120014272APending Publication Date: 2025-05-16XIAMEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510093896.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

While improving the model's generalization semantic segmentation method, the existing domain generalization semantic segmentation method is easy to confuse the semantic features within the domain of different categories, resulting in classification misjudgment, and may introduce domain-independent noise, which will damage the domain-invariant representation and segmentation consistency.

Method used

A domain generalized semantic segmentation method based on semantic consistency and style diversity is proposed. Visual and text features are extracted through CLIP vision encoder and text encoder, and cross-modal semantic association is established using semantic query enhancer. The text-driven style transformation module guides the transformation of the low-frequency amplitude spectrum of image features, and strengthens the separation and aggregation of features between domains through collaborative weighted style comparison loss and style aggregation loss.

Benefits of technology

The best performance is achieved significantly better than existing methods, showing good generalization capabilities on various cross-domain datasets, while maintaining the model's training overhead and fast inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014272A_ABST
    Figure CN120014272A_ABST
Patent Text Reader

Abstract

The invention discloses a domain generalization semantic segmentation method based on semantic consistency and style diversity. The method comprises the following steps: S1, performing visual and text feature extraction based on a CLIP visual encoder and a text encoder; s2, based on a semantic query intensifier, establishing cross-modal semantic association and aggregating related semantic features by using semantic consistency between image-text modalities to enhance initial object query; s3, guiding the transformation of the low-frequency amplitude spectrum of the image features by using the text embedding difference based on a text-driven style transformation module; s4, strengthening the separation of the features between the fields and the aggregation of the features in the fields by cooperatively weighting the style comparison loss and the style aggregation loss; s5, performing mask prediction, category prediction and query refinement layer by layer by using semantic query based on a mask decoder; according to the method, the optimal performance obviously superior to that of an existing method is achieved on each cross-domain data set, meanwhile, the training overhead of the model is kept low, the reasoning speed is high, and the method has obvious practical value and application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of domain generalization semantic segmentation, and in particular relates to a domain generalization semantic segmentation method based on semantic consistency and style diversity. Background Art

[0002] Semantic segmentation is a core task in computer vision that involves assigning a semantic class label to each pixel in an image. Traditionally, semantic segmentation is trained and evaluated under the assumption that datasets are independent and identically distributed. However, this assumption often does not hold in the real world due to variations in lighting, weather conditions, and geographical differences. Domain generalization semantic segmentation focuses on learning only from source domain data to enhance the model's ability to generalize to unknown target domains, thereby better addressing the challenges posed by real-world applications.

[0003] Existing domain generalization semantic segmentation methods are generally divided into two categories: feature normalization and domain randomization. Methods based on feature normalization include normalization and whitening, which aim to restrict the distribution of different features to the same space or remove domain-specific style features, thereby promoting the learning of domain-invariant semantic content representation. However, this often confuses the intra-domain semantic features of different categories, resulting in classification misjudgments. In addition, due to the non-orthogonality of content and style, removing style also leads to the loss of semantic content. Methods based on domain randomization attempt to transform the source domain style into multiple other domains to increase style diversity. However, these methods rely on artificially created auxiliary domains (e.g., ImageNet). This may introduce domain-independent noise, which may damage the domain-invariant representation and lead to segmentation ambiguity. Therefore, it is of great significance to propose a new domain generalization semantic segmentation method based on semantic consistency prediction and style diversity generalization. Summary of the invention

[0004] To solve the above problems, the present invention proposes a domain generalization semantic segmentation method based on semantic consistency and style diversity.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A domain generalization semantic segmentation method based on semantic consistency and style diversity, comprising the following steps:

[0007] S1, extract visual and text features based on CLIP visual encoder and text encoder;

[0008] S2, based on the semantic query enhancer, the semantic consistency between the image and text modalities is utilized to establish cross-modal semantic associations and aggregate relevant semantic features to enhance the initial object query;

[0009] S3, the text-driven style transfer module uses text embedding differences to guide the transformation of the low-frequency amplitude spectrum of image features;

[0010] S4. Strengthen the separation of inter-domain features and the aggregation of intra-domain features through co-weighted style contrast loss and style aggregation loss;

[0011] S5. Mask decoder is used to perform mask prediction, category prediction and query refinement layer by layer using semantic query.

[0012] Preferably, the specific process of step S1 is:

[0013] S11. In the visual feature extraction stage, the feature extractor uses the CNN-based CLIP visual encoder as the backbone network to extract multi-scale feature maps. Where i∈{2,3,4,5}, D is the number of channels; H is the height of the image feature; W is the width of the image feature;

[0014] S12. In the text feature extraction stage, CLIP text encoder is used to encode the given category vocabulary to obtain text embedding C is the number of categories and D is the number of channels.

[0015] Preferably, the specific process of step S2 is:

[0016] S21, the last layer of image features F5 is processed by attention pooling to obtain dense visual features Among them, H' and W' represent and Used to aggregate global context information while maintaining feature scale; dense visual features F v With text embedded E t Semantic similarity graph between The calculation formula is:

[0017]

[0018] Wherein, T represents the transpose operation;

[0019] S22, by performing a maximum operation along the category dimension of the semantic similarity graph S, determine the most relevant category index for each pixel Embed E from text using category index G t Retrieve the corresponding category embeddings from , and build a semantic aggregation graph tailored for pixel-level classification The calculation formula is:

[0020]

[0021] S a =E t [G],

[0022] Among them, E t [G] indicates along E t The first dimension selects the index in G; the semantic similarity graph S and the semantic aggregation graph S a They are mapped to the same channel dimension as the object query Q through a multi-layer perceptron respectively;

[0023] S23. A set of learnable object queries Sequentially with the semantic similarity graph S and the semantic aggregation graph S a Generate semantic queries with both semantic perception and cross-domain consistent prediction and discrimination capabilities through cross-attention mechanism interaction Where N is the number of queries and D is the number of channels.

[0024] Preferably, the specific process of step S3 is:

[0025] S31, introduce a triple domain prompt set P = {p1, p2, p3}, where p1 is a general domain prompt; p2 is a conditional domain prompt; p3 is a specific domain prompt;

[0026] S32. For each image in each batch, generate a domain-specific cue p3 for the current image by combining the category of each pixel from the true mask with the randomly selected conditional domain cue p2.

[0027] S33, the general domain hint p1 is constructed directly from the category of each pixel in the true mask. These two sets of hints are fed into the text encoder separately, and the resulting embeddings are averaged along the hint dimension to obtain the corresponding domain-specific embeddings and universal domain embedding Domain-Different Embedding Obtained by element subtraction, the calculation formula is:

[0028] E d =E s -E g =ε(p3)-ε(p1),

[0029] where ε(·) represents the CLIP text encoder; a domain style adapter is introduced to embed the domain difference into E d Mapping to style difference features Used to convert pixel-level text embedding into image feature space;

[0030] S34, apply fast Fourier transform to transform image features Decomposition into amplitude spectrum and phase spectrum The calculation formula is:

[0031]

[0032]

[0033] Among them, u and v are the horizontal and vertical coordinates of the frequency coordinate in the frequency domain; Re(·) and Im(·) represent the Fourier spectrum, respectively. The real and imaginary parts of

[0034] S35. Construct low-frequency mask M l ∈{0, 1}, which is used to concentrate the low-frequency components at the center of the spectrum and is defined as:

[0035]

[0036] Among them, x and y are the horizontal and vertical coordinates of the pixel coordinates in the spatial domain, respectively; H is the height of the image feature; W is the width of the image feature; α is the proportion of the low-frequency component; for the style difference feature F d The numerical constraints of the hyperbolic tangent function are performed, the activated style difference features are weighted and added to the low-frequency components of the amplitude spectrum, and then combined with the amplitude spectrum of the high-frequency components to obtain a composite amplitude spectrum. The calculation formula is:

[0037]

[0038] Among them, β is the style control strength; is the low-frequency amplitude spectrum; is the composite amplitude spectrum;

[0039] S36, the composite amplitude spectrum and the original phase spectrum Merge and transform using inverse fast Fourier transform to obtain the style visual features after style transformation

[0040]

[0041] in, is the inverse fast Fourier transform;

[0042] S37. Repeat steps S34-S36 for the last three layers of image features.

[0043] Preferably, the specific process of step S4 is:

[0044] S41, a set of learnable domain libraries initialized by domain-conditional cue embeddings Used to store the style vector for each domain, where K is the number of conditional domain cues; style contrast loss after global average pooling of global average pooling style visual features The calculation between the domain library D is:

[0045]

[0046] in, is the style contrast loss; L is the number of feature layers; l is the feature layer index; D + is the positive sample domain library; τ is the temperature parameter; j is the negative sample domain library index; is the negative sample domain library; N is the number of negative domain samples;

[0047] S42, projecting the domain-specific embedding into the image feature space to obtain the domain-specific style feature Fs, normalizing the domain-specific style feature Fs and the style visual feature V, and using the L2 loss function to constrain the relationship between the two. The calculation formula is:

[0048]

[0049] in, is the style aggregation loss; h is the image height index; w is the image width index; w l is the weight coefficient that changes with the number of layers; is the visual feature of style; F s,(h,w) is a domain-specific style feature;

[0050] S43. According to the change of one loss, the weight of another loss is adaptively adjusted. The calculation formula of the collaborative weighted strategy is:

[0051]

[0052] Among them, w1 is the weight coefficient of style contrast loss or style aggregation loss; w init is the initial weight coefficient; λ is the hyperparameter used for scaling changes; is the loss difference value.

[0053] Preferably, the specific process of step S5 is:

[0054] S51. In each layer of the mask decoder, the semantic query First perform self-attention, then perform cross-attention operation with multi-scale image features;

[0055] S52. Semantic Query Generate category predictions through linear layers, and perform dot products with the highest-level feature maps to generate mask predictions;

[0056] S53, repeat steps S1-S2 multiple times to implement semantic query of continuous refinement.

[0057] After adopting the above technical scheme, the present invention has the following beneficial effects: a new framework for semantic consistency prediction and style diversity generalization proposed by the present invention aims to explore the semantic consistency and style diversity of domain generalization semantic segmentation; the semantic query enhancer uses the semantic consistency between image and text modalities to establish cross-modal semantic associations and aggregate relevant semantic features, and by enhancing the semantic discrimination ability of object queries in the mask decoder, it helps to robustly predict the semantic consistency between different domains; the introduced text-driven style transformation module is used to mine the style diversity of text modalities, and by using the difference between the text embedding vectors of specific domain prompts and general domain prompts as domain difference embedding and mapping them to style difference features, the module controllably guides the transformation of the low-frequency amplitude spectrum of image features, thereby realizing cross-domain style transformation and enhancing inter-domain style diversity; in order to prevent the collapse of similar domain feature space, the style collaborative optimization mechanism strengthens the separation of inter-domain features and the aggregation of intra-domain features through collaborative weighted style contrast loss and style aggregation loss. Therefore, the method achieves the best performance significantly better than the existing methods on various cross-domain datasets, while keeping the model's training overhead low and inference speed fast, and has significant practical value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flow chart of the present invention;

[0059] Figure 2 It is a schematic diagram of the network structure of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] like Figure 1 to Figure 2 As shown, a domain generalization semantic segmentation method based on semantic consistency and style diversity includes the following steps:

[0062] S1, extract visual and text features based on CLIP visual encoder and text encoder;

[0063] The specific process of step S1 is:

[0064] S11. In the visual feature extraction stage, the feature extractor uses the CNN-based CLIP visual encoder as the backbone network to extract multi-scale feature maps. Where i∈{2,3,4,5}, D is the number of channels; H is the height of the image feature; W is the width of the image feature;

[0065] S12. In the text feature extraction stage, CLIP text encoder is used to encode the given category vocabulary to obtain text embedding C is the number of categories, D is the number of channels;

[0066] S2, based on the semantic query enhancer, the semantic consistency between the image and text modalities is utilized to establish cross-modal semantic associations and aggregate relevant semantic features to enhance the initial object query;

[0067] The specific process of step S2 is:

[0068] S21, the last layer of image features F5 is processed by attention pooling to obtain dense visual features Among them, H' and W' represent and Used to aggregate global context information while maintaining feature scale; dense visual features F v With text embedded E t Semantic similarity graph between The calculation formula is:

[0069]

[0070] Wherein, T represents the transpose operation;

[0071] S22, by performing a maximum operation along the category dimension of the semantic similarity graph S, determine the most relevant category index for each pixel Embed E from text using category index G t Retrieve the corresponding category embeddings from , and build a semantic aggregation graph tailored for pixel-level classification The calculation formula is:

[0072]

[0073] S a =E t [G],

[0074] Among them, E t [G] indicates along E t The first dimension selects the index in G; the semantic similarity graph S and the semantic aggregation graph S a They are mapped to the same channel dimension as the object query Q through a multi-layer perceptron respectively;

[0075] S23. A set of learnable object queries Sequentially with the semantic similarity graph S and the semantic aggregation graph S a Generate semantic queries with both semantic perception and cross-domain consistent prediction and discrimination capabilities through cross-attention mechanism interaction Where N is the number of queries and D is the number of channels;

[0076] S3, the text-driven style transfer module uses text embedding differences to guide the transformation of the low-frequency amplitude spectrum of image features;

[0077] The specific process of step S3 is:

[0078] S31, introduce a triple domain prompt set P = {p1, p2, p3}, where p1 is a general domain prompt; p2 is a conditional domain prompt; p3 is a specific domain prompt;

[0079] S32. For each image in each batch, generate a domain-specific cue p3 for the current image by combining the category of each pixel from the true mask with the randomly selected conditional domain cue p2.

[0080] S33, the general domain hint p1 is constructed directly from the category of each pixel in the true mask. These two sets of hints are fed into the text encoder separately, and the resulting embeddings are averaged along the hint dimension to obtain the corresponding domain-specific embeddings and universal domain embedding Domain-Different Embedding Obtained by element subtraction, the calculation formula is:

[0081] E d =E s -E g =ε(p3)-ε(p1),

[0082] where ε(·) represents the CLIP text encoder; a domain style adapter is introduced to embed the domain difference into E d Mapping to style difference features For converting pixel-level text embeddings into image feature space:

[0083] S34, apply fast Fourier transform to transform image features Decomposition into amplitude spectrum and phase spectrum The calculation formula is:

[0084]

[0085]

[0086] Among them, u and v are the horizontal and vertical coordinates of the frequency coordinate in the frequency domain; Re(·) and Im(·) represent the Fourier spectrum, respectively. The real and imaginary parts of

[0087] S35. Construct low-frequency mask M l∈{0, 1}, which is used to concentrate the low-frequency components at the center of the spectrum and is defined as:

[0088]

[0089] Among them, x and y are the horizontal and vertical coordinates of the pixel coordinates in the spatial domain, respectively; H is the height of the image feature; W is the width of the image feature; α is the proportion of the low-frequency component; for the style difference feature F d The numerical constraints of the hyperbolic tangent function are performed, the activated style difference features are weighted and added to the low-frequency components of the amplitude spectrum, and then combined with the amplitude spectrum of the high-frequency components to obtain a composite amplitude spectrum. The calculation formula is:

[0090]

[0091] Among them, β is the style control strength; is the low-frequency amplitude spectrum; is the composite amplitude spectrum;

[0092] S36, the composite amplitude spectrum and the original phase spectrum Merge and transform using inverse fast Fourier transform to obtain the style visual features after style transformation

[0093]

[0094] in, is the inverse fast Fourier transform;

[0095] S37, repeat steps S34-S36 for the last three layers of image features;

[0096] S4. Strengthen the separation of inter-domain features and the aggregation of intra-domain features through co-weighted style contrast loss and style aggregation loss;

[0097] The specific process of step S4 is:

[0098] S41, a set of learnable domain libraries initialized by domain-conditional cue embeddings Used to store the style vector for each domain, where K is the number of conditional domain cues; style contrast loss after global average pooling of global average pooling style visual features The calculation between the domain library D is:

[0099]

[0100] in, is the style contrast loss; L is the number of feature layers; l is the feature layer index; D +is the positive sample domain library; τ is the temperature parameter; j is the negative sample domain library index; is the negative sample domain library; N is the number of negative domain samples;

[0101] S42, projecting the domain-specific embedding into the image feature space to obtain the domain-specific style feature Fs, normalizing the domain-specific style feature Fs and the style visual feature V, and using the L2 loss function to constrain the relationship between the two. The calculation formula is:

[0102]

[0103] in, is the style aggregation loss; h is the image height index; w is the image width index; w l is the weight coefficient that changes with the number of layers; is the visual feature of style; F s,(h,w) is a domain-specific style feature;

[0104] S43. According to the change of one loss, the weight of another loss is adaptively adjusted. The calculation formula of the collaborative weighted strategy is:

[0105]

[0106] Among them, w is the weight coefficient of style contrast loss or style aggregation loss; w init is the initial weight coefficient; λ is the hyperparameter used for scaling changes; is the loss difference value;

[0107] S5, mask prediction, category prediction and query refinement layer by layer using semantic query based on mask decoder;

[0108] The specific process of step S5 is:

[0109] S51. In each layer of the mask decoder, the semantic query First perform self-attention, then perform cross-attention operation with multi-scale image features;

[0110] S52. Semantic Query Generate category predictions through linear layers, and perform dot products with the highest-level feature maps to generate mask predictions;

[0111] S53, repeat steps S1-S2 multiple times to implement semantic query of continuous refinement.

[0112] Performance Test:

[0113] In the domain generalization semantic segmentation task, the model is trained on the source domain dataset and evaluated on other datasets as the target domain using mIoU and average mIoU metrics. Experiments are conducted on two synthetic datasets and four real-world datasets to evaluate the generalization ability of the method. For the synthetic datasets, GTAV(G) is a game synthetic dataset containing 24,966 images with a resolution of 1914×1052. SYNTHIA(S) is a large-scale synthetic dataset containing 9,400 images with a resolution of 1280×760. For the real-world datasets, Cityscapes(C), BDD-100K(B), and Mapillary(M) contain 2,975, 7,000, and 18,000 training images and 500, 1,000, and 2,000 validation images, respectively.

[0114] Table 1: Performance comparison of different domain generalization semantic segmentation methods in a single domain setting

[0115]

[0116] Table 2: Performance comparison of different domain generalization semantic segmentation methods in multi-domain settings

[0117]

[0118] It can be seen from Table 1 that, whether trained on real datasets or synthetic datasets, after introducing the semantic query enhancer, the text-driven style transformation module and the style collaborative optimization mechanism, the method proposed in the present invention consistently achieves significantly better performance than the best existing methods.

[0119] It can be seen from Table 2 that in the multi-domain scenario setting, the method proposed in the present invention still shows good generalization ability.

[0120] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A domain generalization semantic segmentation method based on semantic consistency and style diversity, characterized by: The following steps are involved: S1, extract visual and text features based on CLIP visual encoder and text encoder; S2, based on the semantic query enhancer, the semantic consistency between the image and text modalities is utilized to establish cross-modal semantic associations and aggregate relevant semantic features to enhance the initial object query; S3, the text-driven style transfer module uses text embedding differences to guide the transformation of the low-frequency amplitude spectrum of image features; S4. Strengthen the separation of inter-domain features and the aggregation of intra-domain features through co-weighted style contrast loss and style aggregation loss; S5. Mask decoder is used to perform mask prediction, category prediction and query refinement layer by layer using semantic query.

2. The domain generalization semantic segmentation method based on semantic consistency and style diversity as claimed in claim 1, characterized in that: The specific process of step S1 is: S11. In the visual feature extraction stage, the feature extractor uses the CNN-based CLIP visual encoder as the backbone network to extract multi-scale feature maps. Where i∈{2,3,4,5}, D is the number of channels; H is the height of the image feature; W is the width of the image feature; S12. In the text feature extraction stage, CLIP text encoder is used to encode the given category vocabulary to obtain text embedding C is the number of categories and D is the number of channels.

3. The domain generalization semantic segmentation method based on semantic consistency and style diversity as claimed in claim 1, characterized in that: The specific process of step S2 is: S21, the last layer of image features F5 is processed by attention pooling to obtain dense visual features Among them, H' and W' represent and Used to aggregate global context information while maintaining feature scale; dense visual features F v With text embedded E t Semantic similarity graph between The calculation formula is: Wherein, T represents the transpose operation; S22, by performing a maximum operation along the category dimension of the semantic similarity graph S, determine the most relevant category index for each pixel Embed E from text using category index G t Retrieve the corresponding category embeddings from , and build a semantic aggregation graph tailored for pixel-level classification The calculation formula is: S a =E t [G], Among them, E t [G] indicates along E t The first dimension selects the index in G; the semantic similarity graph S and the semantic aggregation graph S a They are mapped to the same channel dimension as the object query Q through a multi-layer perceptron respectively; S23. A set of learnable object queries Sequentially with the semantic similarity graph S and the semantic aggregation graph S a Generate semantic queries with both semantic perception and cross-domain consistent prediction and discrimination capabilities through cross-attention mechanism interaction Where N is the number of queries and D is the number of channels.

4. The domain generalization semantic segmentation method based on semantic consistency and style diversity as claimed in claim 1, characterized in that: The specific process of step S3 is: S31, introduce a triple domain prompt set P = {p1, p2, p3}, where p1 is a general domain prompt; p2 is a conditional domain prompt; p3 is a specific domain prompt; S32. For each image in each batch, generate a domain-specific cue p3 for the current image by combining the category of each pixel from the true mask with the randomly selected conditional domain cue p2. S33, the general domain hint p1 is constructed directly from the category of each pixel in the true mask. These two sets of hints are fed into the text encoder separately, and the resulting embeddings are averaged along the hint dimension to obtain the corresponding domain-specific embeddings and universal domain embedding Domain-Different Embedding Obtained by element subtraction, the calculation formula is: E d =E s -E g =ε(p3)-ε(p1), Among them, ε(·) represents the CLIP text encoder; at the same time, a domain style adapter is introduced to map the domain difference embedding E to the style difference feature Used to convert pixel-level text embedding into image feature space; S34, apply fast Fourier transform to transform image features Decomposition into amplitude spectrum and phase spectrum The calculation formula is: Among them, u and v are the horizontal and vertical coordinates of the frequency coordinate in the frequency domain; Re(·) and Im(·) represent the Fourier spectrum, respectively. The real and imaginary parts of S35. Construct low-frequency mask M l ∈{0,1}, which is used to concentrate the low-frequency components at the center of the spectrum and is defined as: Among them, x and y are the horizontal and vertical coordinates of the pixel coordinates in the spatial domain, respectively; H is the height of the image feature; W is the width of the image feature; α is the proportion of the low-frequency component; for the style difference feature F d The numerical constraints of the hyperbolic tangent function are performed, the activated style difference features are weighted and added to the low-frequency components of the amplitude spectrum, and then combined with the amplitude spectrum of the high-frequency components to obtain a composite amplitude spectrum. The calculation formula is: Among them, β is the style control strength; is the low-frequency amplitude spectrum; is the composite amplitude spectrum; S36, the composite amplitude spectrum and the original phase spectrum Merge and transform using inverse fast Fourier transform to obtain the style visual features after style transformation in, is the inverse fast Fourier transform; S37. Repeat steps S34-S36 for the last three layers of image features.

5. The domain generalization semantic segmentation method based on semantic consistency and style diversity as claimed in claim 1, characterized in that: The specific process of step S4 is: S41, a set of learnable domain libraries initialized by domain-conditional cue embeddings Used to store the style vector for each domain, where K is the number of conditional domain cues; style contrast loss after global average pooling of global average pooling style visual features The calculation between the domain library D is: in, is the style contrast loss; L is the number of feature layers; l is the feature layer index; D + is the positive sample domain library; τ is the temperature parameter; j is the negative sample domain library index; is the negative sample domain library; N is the number of negative domain samples; S42, projecting the domain-specific embedding into the image feature space to obtain the domain-specific style feature Fs, normalizing the domain-specific style feature Fs and the style visual feature V, and using the L2 loss function to constrain the relationship between the two. The calculation formula is: in, is the style aggregation loss; h is the image height index; w is the image width index; w l is the weight coefficient that changes with the number of layers; is the visual feature of style; F s,(h,w) is a domain-specific style feature; S43. According to the change of one loss, the weight of another loss is adaptively adjusted. The calculation formula of the collaborative weighted strategy is: Among them, w1 is the weight coefficient of style contrast loss or style aggregation loss; w init is the initial weight coefficient; λ is the hyperparameter used for scaling changes; is the loss difference value.

6. The domain generalization semantic segmentation method based on semantic consistency and style diversity as claimed in claim 1, characterized in that: The specific process of step S5 is: S51. In each layer of the mask decoder, the semantic query First perform self-attention, then perform cross-attention operation with multi-scale image features; S52. Semantic Query Generate category predictions through linear layers, and perform dot products with the highest-level feature maps to generate mask predictions; S53, repeat steps S1-S2 multiple times to implement semantic query of continuous refinement.

Citation Information

Cited By

  • Medical image field generalization segmentation method and system

    CN120355929A