Open Hybrid Domain Adaptation Image Segmentation Method Based on Object-Level Difference Memory

By introducing an object-level differential memory (OLDM) method in the image segmentation model, the problem of poor adaptability of image segmentation models in different scenarios is solved, and better multi-scene segmentation performance and open domain image segmentation effect are achieved.

CN116612286BActive Publication Date: 2025-05-30TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310747258.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-25
Publication Date
2025-05-30
Estimated Expiration
2043-06-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively adapt to the image segmentation model in different scenarios, resulting in poor segmentation performance of the model in multiple scenarios.

Method used

An open hybrid domain adaptation image segmentation method based on object-level differential memory (OLDM) is adopted. By constructing an OLDM module, the differences between the target domain image features and the source domain features are stored, and weighted fusion is performed in the segmentation task to achieve domain adaptation.

Benefits of technology

The segmentation performance of the image segmentation model in multiple scenarios is improved, especially in the open unknown domain image segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612286B_ABST
    Figure CN116612286B_ABST
Patent Text Reader

Abstract

The present invention relates to an open hybrid domain adaptation image segmentation method based on object-level differential memory, comprising the following steps: constructing a semantic segmentation network, roughly training the model using source domain images, calculating the centroids of features of different classes in the source domain images as class keys for different classes; feeding the target domain images and semantic segmentation results into the OLDM module; using OLDM for semantic segmentation optimization; using OLDM for pseudo-label update, calculating the loss function and performing backpropagation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and computer vision, relates to image segmentation technology, and specifically relates to an open hybrid domain adaptation image segmentation method based on object-level difference memory. Background Art

[0002] Image segmentation refers to the process of subdividing a digital image into multiple image sub-regions. It means splitting the input picture into multiple image sub-regions and has applications in areas such as security, healthcare, transportation, and the Internet.

[0003] A semantic segmentation model trained on a dataset in one scenario cannot well adapt to the data in another scenario. Therefore, it is necessary to transfer the scenario to achieve good segmentation of data in multiple scenarios by the model. Thus, a domain adaptation for two-dimensional image semantic segmentation is proposed, which transfers and processes information in different image domains, and finally obtains a large amount of data for training the semantic segmentation model to enhance the model's segmentation performance. Summary of the Invention

[0004] The present invention provides an open hybrid domain adaptation image segmentation method, called an open hybrid domain adaptation image segmentation method based on object-level difference memory (OLDM), which uses compensation content of different classes to narrow the distance between the source domain and target domain images. During the training process of the segmentation network, OLDM is constructed to store the keys generated from the target domain image features and the differences from the source domain features. During the test of the segmentation task, key-value pairs are selected from each class in OLDM and weighted and fused into the original target domain features to complete the domain adaptation from the target domain to the source domain. The present invention is implemented through the following technical solutions:

[0005] An open hybrid domain adaptation image segmentation method based on object-level difference memory, comprising the following steps:

[0006] Step 1: Construct a semantic segmentation network, and use the source domain image I s to roughly train the model, and calculate the centroid A l of the features of different classes in the source domain image as the class keys of different classes;

[0007] Step 2: Input the image of the target domain t and the semantic segmentation result R

[0008] (1) into the OLDM module. The method for updating OLDM is as follows: Input the target domain image

[0009] (2) into the semantic segmentation network; Obtain the features t of the target domain image and the semantic segmentation result R of the entire image;

[0010] (3) Utilize the target domain features and the centroid A of the source domain features l to initialize and update the instance key N in the OLDM l,m and the difference feature D l,m The update method is as follows: Calculate the similarity w l,m between the instance features of different categories and the saved instance key N l,m , and update it in a weighted manner according to the similarity;

[0011] Step 3: Optimize semantic segmentation using OLDM, the method is as follows:

[0012] (1) Input the target domain image into the semantic segmentation network;

[0013] (2) Obtain the features of the target domain image and the semantic segmentation result R of the entire image t ;

[0014] (3) For each pixel feature F t (x, y), according to its preliminary semantic segmentation result R t (x, y), select the categories corresponding to the largest K scores; Calculate the similarity w k,m between the instance features of different categories and the saved instance key N k,m , for each instance key N k,m , obtain its corresponding instance content D k,m , k is the category index in the categories corresponding to the largest K scores selected; Calculate the final compensation content according to the similarity and the instance content, and add it to the pixel feature F t (x, y) to generate the final compensated feature and the segmentation score map

[0015] Step 4: Update the pseudo-labels using OLDM, calculate the loss function and backpropagate, the method is as follows:

[0016] (1) Set the threshold γ according to the input segmentation score map 2 . For the points where the score of the category with the largest score in the score map is less than γ 2 , do not update, otherwise update, and thus calculate the pseudo-label Y t (x, y) for update;

[0017] (2) For the pixel points that need to update the pseudo-label Y t (x, y), calculate the loss function and backpropagate.

[0018] Furthermore, the method of step one is as follows:

[0019] (1) Construct a semantic segmentation network and input the source domain image into this network;

[0020] (2) Obtain the feature F s (x, y) of the source domain image and the semantic segmentation result R s of the entire image. For each pixel, classify it into the category l with the highest score in the segmentation result, and calculate the centroid A l of the source domain image features for each category as the category key for different categories. The centroid A l is updated as follows:

[0021] For each category, select all pixels I s (x, y) with the true result being this category. The corresponding source domain image feature after encoding is F s (x, y). Update the corresponding centroid A l in the following way;

[0022] A l ←A l +λF s (x, y)

[0023] where λ is the weight coefficient for adding the feature F s (x, y) to the centroid A l .

[0024] Furthermore, in step two, the instance key N l,m and the difference feature D l,m are updated as follows:

[0025] Step 1: Determine the category corresponding to the highest score of the pixel point (x, y) in the target domain image through the preliminary semantic segmentation result R t , and set a threshold γ. For the points where the score of the category with the largest score among the pixel points (x, y) in the target domain image is less than γ, do not update the instance key, otherwise update it;

[0026] Step 2: Calculate the similarity w l,m between the instance features of different categories and the saved instance key N l,m ,

[0027]

[0028] where C is the result of multiplying the square of the norm of the instance key N l,m by the square of the norm of the instance feature F t (x, y);

[0029] Step 3: Update the instance key in a weighted manner:

[0030] N l,m ←N l,m +w l,m F t (x, y)

[0031] Step 4: Update the differential features in a weighted manner:

[0032] D l,m ←D l,m +w l,m (A l -F t (x ,y ))。

[0033] Furthermore, for sub-step (3) of Step 4, the method is as follows:

[0034] Step 1: Determine the categories corresponding to the top K values of the scores of the corresponding pixel points (x, y) through the preliminary semantic segmentation result R t ;

[0035] Step 2: Calculate the similarity w k,m between the instance features of different categories and the saved instance key N k,m ,

[0036]

[0037] where C is the result of multiplying the square of the norm of the instance key N k,m by the square of the norm of the instance feature F t (x, y), that is:

[0038] C = |N k,m | 2 × |Ft ( (x, y)| 2

[0039] Step 3: Calculate the compensated feature using the following formula

[0040]

[0041] (4) Input the obtained compensated feature into the decoder to generate the final segmentation score map

[0042] Furthermore, in Step 4, the calculation method of the loss function is as follows:

[0043] Step 1: Calculate the loss function of the source domain image Use the cross - entropy function CELoss as the loss function to calculate the score map R obtained from the source - domain images s and the ground - truth label Y s to calculate the difference loss:

[0044]

[0045] Step 2: Calculate the loss function of the target - domain images Use the cross - entropy function CELoss as the loss function to calculate the score map R obtained from the target - domain images t and the compensated score map and the generated pseudo - label Y t to calculate the difference loss;

[0046]

[0047] Step 3: Sum the two loss functions to obtain the final loss function:

[0048]

[0049] The beneficial effects of the technical solution provided by the present invention are as follows:

[0050] 1. During the image segmentation process of the present invention, since the target - domain images used contain multiple domains (such as different weather conditions, different times, etc.), OLDM stores them during the segmentation process, providing richer information for the segmentation task.

[0051] 2. During the image segmentation process of the present invention, OLDM stores the feature differences of multiple target domains. When performing fusion, the features of multiple target domains are fused, and the images in the open unknown domain also have good segmentation effects. Brief Description of the Drawings

[0052] Figure 1 It is a compensation flow chart for the open - mixed - domain adaptation image segmentation method based on object - level difference memory;

[0053] Figure 2 It is an explanatory diagram of the compensation method for the open - mixed - domain adaptation image segmentation method based on object - level difference memory;

[0054] Figure 3 It is a quantitative comparison between the method of the present invention and other existing optimal image segmentation methods. Detailed Embodiments

[0055] The technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings. All other embodiments obtained by those of ordinary skill in the art based on the technical solutions in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] Open mixed-domain adaptation images have differential information between different domains. How to preserve and utilize this differential information is the key content of domain adaptation image segmentation. The source-domain image is a training image with labels, while the target-domain image is a training image without labels. At the same time, the open-domain image is invisible during training but exists in testing. OLDM is a method based on differential memory. First, it encodes the source-domain image and the target-domain image to obtain the feature information of the images, subtracts them to obtain the difference in feature information, and uses the difference in feature information and the corresponding target-domain image features as the differential feature D l,m and the instance key N l,m , and saves them in the form of key-value pairs, where l represents different categories and m represents different entries in the save. During subsequent training and testing, compensation is performed by adding the difference in feature information to the target-domain image features to reduce the information difference between different domains, so as to achieve the effect that the target-domain and open-domain images approach the source-domain image, and thus the model has good performance on both the target domain and the open domain.

[0057] The specific implementation method includes the following steps:

[0058] Step 1: Calculate the centroids of different categories of the source-domain image

[0059] Use the source-domain image I s to roughly train the model, and at the same time calculate the centroid A of the features of different categories of the source-domain image l as the category key of different categories. The specific steps are as follows:

[0060] (1) Generate a network for semantic segmentation results and input the source-domain image into this network. The network uses ResNet101 as the backbone, and a skip connection method is added between the encoder and the decoder.

[0061] (2) Obtain the features of the source-domain image and the semantic segmentation result R of the entire image s , for each pixel, assign it to the category l with the highest score in the segmentation result, and calculate the centroid A of the source-domain image features category by category l , and the centroids of different categories are updated by taking the average. The update method is as follows:

[0062] For each category, select all pixels I s (x, y) with the true result being this category, and the corresponding feature after encoding is F s (x, y), and update the corresponding centroid A in the following way l .

[0063] Ax ←A l +λF s (x, y)

[0064] where λ is the weight coefficient of feature F s (x, y) is added to the centroid A l .

[0065] Step 2: Save and update the OLDM content

[0066] For the image of the target domain and the semantic segmentation result R t The specific method of feeding them into the OLDM module to update the OLDM content is as follows:

[0067] (1) Input the target domain image into the network with ResNet101 as the backbone.

[0068] (2) Obtain the features of the target domain image and the semantic segmentation result R of the whole image t .

[0069] (3) Use the target domain features and the centroid A of the source domain features l to initialize and update the instance key N l,m and the difference feature D l,m in the OLDM.

[0070] The update method of the instance key N l,m is as follows:

[0071] Step 1: Determine the category corresponding to the highest score of the pixel point (x, y) of the target domain image through the preliminary semantic segmentation result R t and set a threshold γ. For the points where the score of the category with the largest score among the pixel points (x, y) of the target domain image is less than γ 1 , do not update the instance key, otherwise update it.

[0072] R t (x, y, l) = max{R t (x, y, 1), …, R t (x, y, L)} > γ 1

[0073] Step 2: Calculate the similarity w l,m between the instance features of different categories and the saved instance key N l,m ,

[0074]

[0075] Among them, C is the instance key N l,m The norm of and the instance feature F t (x, y) after taking the square of the norm and multiplying, that is:

[0076] C = |N l,m | 2 × |F t (x, y)| 2

[0077] Step 3: Update the instance key in a weighted manner

[0078] N l,m ← N l,m + w l,m F t (x, y)

[0079] The difference feature D l,m The update method is as follows:

[0080] Step 1: Update the difference feature in a weighted manner

[0081] D l,m ← D l,m + w l,m (A l - F t (x, y))

[0082] Step three: Use OLDM for semantic segmentation optimization

[0083] (1) Input the target domain image into the network with ResNetl01 as the backbone.

[0084] (2) Obtain the features of the target domain image and the semantic segmentation result R of the whole image t .

[0085] (3) For each pixel feature F t (x, y), according to its preliminary semantic segmentation result R t (x, y), select the K categories corresponding to the largest K scores. Similarly, calculate the similarity w k,m between different category instance features and the saved instance key N k,m , for each instance key N k,m , obtain its corresponding instance content D k,m , k is the category index in the K categories corresponding to the largest K scores selected. Calculate the final compensation content according to the similarity and the instance content, and add it to the pixel feature F t (x, y) to generate the final compensated feature Compensated feature The calculation method is as follows:

[0086] Step 1: First, through the preliminary semantic segmentation result R t Determine the categories corresponding to the top K values of the scores of the corresponding pixel points (x, y).

[0087] R t (x, y, k) ∈ max K {R t (x, y, 1), …, R t (x, y, L)}

[0088] Step 2: Calculate the similarity w k,m between the instance features of different categories and the saved instance key N k,m ,

[0089]

[0090] where C is the result of multiplying the norm of the instance key N k,m by the square of the norm of the instance feature F t (x, y), that is:

[0091] C = |N k,m | 2 × |F t (x, y)| 2

[0092] Step 3: Calculate the compensated feature using the following formula

[0093]

[0094] (4) Input the obtained compensated feature into the decoder to generate the final segmentation score map

[0095] Step Four: Use OLDM to update the pseudo-label

[0096] (1) Set the threshold γ according to the input segmentation score map 2 . For the points where the score of the category with the highest score in the score map is less than γ 2 , do not update, otherwise update, and thus calculate the pseudo-label Y t (x, y) is updated. The calculation method of the pseudo-label Y t (x, y) is as follows:

[0097] Step 1: Select the category l with the highest score according to the input segmentation score map .

[0098] Step 2: Calculate the pseudo-labels using the following formula:

[0099]

[0100] Y t (x, y, l) is the pseudo-label of the pixel point of category l at the coordinate point (x, y).

[0101] (2) For those pixel points that need to be updated, calculate the loss function and backpropagate. The calculation method of the loss function is as follows:

[0102] Step 1: Calculate the loss function of the source domain image Use the cross-entropy function CELoss as the loss function to calculate the score map R obtained from the source domain image s and the difference loss with the true label Y s .

[0103]

[0104] Step 2: Calculate the loss function of the target domain image Use the cross-entropy function CELoss as the loss function to calculate the score map R obtained from the target domain image t and the difference loss with the compensated score map and the generated pseudo-label Y t .

[0105]

[0106] Step 3: Sum the two loss functions to obtain the final loss function.

[0107]

[0108] The feasibility of the method of the present invention is verified below with specific examples. See the following description for details:

[0109] Using the method of the present invention as the source domain images on the GTA5 dataset and the SYNTHIA dataset, and verifying on the C-Driving, Cityscapes, KITTI, and WildDash datasets. The GTA5 dataset contains 24,966 street view images, and the SYNTHIA dataset contains 200,000 high-definition images from video streams and 20,000 high-definition images from independent snapshots. The laboratory uses the C-Driving, Cityscapes, KITTI, and WildDash datasets as the target domain images and tests them on the test sets of these datasets, where the C-driving dataset provides open-domain pictures.

[0110] The experiment uses the mean intersection over union (MIoU) to quantitatively evaluate the restoration results.

[0111] According to Figure 3 The results of the restoration of images by the proposed method and the existing optimal semantic segmentation method under different conditions show that the method of the present invention has improvements in the metrics, which proves the effectiveness of the method of the present invention. At the same time, the OLDM of the present invention has an obvious improvement effect on the overall semantic segmentation.

Claims

1. An open hybrid domain adaptation image segmentation method based on object-level difference memory, comprising the following steps: Step 1: Construct a semantic segmentation network and use the source domain image I s Roughly train the model and calculate the centroid A of the features of different classes in the source domain image l As the class keys for different classes; Step 2: For the image in the target domain and the semantic segmentation result R t Feed them into the object-level difference memory OLDM module. The method for updating OLDM is as follows: (1) Input the target domain image into the semantic segmentation network; (2) Obtain the features of the target domain image and the semantic segmentation result R of the entire image t ; (3) Construct an OLDM for storing the keys generated by the target domain image features and the differences from the source domain features, and utilize the target domain features and the centroid A of the source domain features l , initialize and update the instance key N in the OLDM l,m and the difference feature D l,m . The update method is as follows: calculate the similarity w l,m between the instance features of different categories and the saved instance key N l,m , and perform the update in a weighted manner according to the similarity; Step 3: Optimize semantic segmentation using OLDM, the method is as follows: (1) Input the target domain image into the semantic segmentation network; (2) Obtain the features of the target domain image and the semantic segmentation result R of the entire image t ; (3) For each instance feature F t (x, y), according to its preliminary semantic segmentation result R t (x, y), select the categories corresponding to the top K scores; calculate the similarity w between different category instance features and the saved instance key N k,m ; for each instance key N k,m , obtain its corresponding instance content D k,m ; k is the category index in the categories corresponding to the top K scores selected; calculate the final compensation content based on the similarity and the instance content, and add it to the instance feature F k,m (x, y) to generate the final compensated instance feature t and the segmentation score map ​ Step 4: Update the pseudo labels using OLDM, calculate the loss function and backpropagate, the method is as follows: (1) According to the input segmentation score map Set the threshold γ 2 , for the points where the score of the category with the largest score in the score map is less than γ 2 , do not update, otherwise update, and thus calculate the pseudo-label Y t (x, y) is updated; (2) For the pseudo-label Y that needs to be updated t For the pixel points (x, y), calculate the loss function And perform backpropagation.

2. The open hybrid domain adaptation image segmentation method according to claim 1, wherein, the method of Step 1 is as follows: (1)Construct a semantic segmentation network and input the source domain image into this network; (2) Obtain the feature F of the source domain image s (x, y), as well as the semantic segmentation result R of the entire image s , for each pixel, classify it into the category l with the highest score in the segmentation result, and calculate the centroid A of the source domain image features for each category l As the category key for different categories, the centroid A l The update method is as follows: For each category, select all pixels I whose true result is this category s (x, y), and the corresponding source domain image feature after encoding is F s (x, y), and update the corresponding centroid A in the following way l ; A l ←A l +λF s (x,y) where λ is the feature F s (x, y) is added to the centroid A l as the weight coefficient.

3. The open hybrid domain adaptation image segmentation method according to claim 1, wherein, In step two, instance key N l,m and the difference feature D l,m are updated as follows: Step 1: Through the preliminary semantic segmentation result R t Determine the category corresponding to the highest score of the pixel point (x, y) of the target domain image, and set a threshold γ. For the points where the score of the category with the largest score among the pixel points (x, y) of the target domain image is less than γ, do not update the instance key, otherwise update it; Step 2: Calculate the similarity w of the instance features of different categories and the saved instance key N l,m l,m ,​ where C is the instance key N l,m is the norm of the instance feature F t the result of multiplying the squared norm of (x, y) Step 3: Update the instance key in a weighted manner, N l,m ←N l,m +w l,m F t (x,y) Step 4: Update the difference features in a weighted manner: D l,m ←D l,m +w l,m (A l -F t (x,y))。 4. The open hybrid domain adaptation image segmentation method according to claim 1, wherein, sub-step (3) of Step 4, the method is as follows: Step 1: Through the preliminary semantic segmentation result R t Determine the categories corresponding to the top K values of the scores of the corresponding pixel points (x, y); Step 2: Calculate the similarity w of the instance features of different categories and the saved instance key N k,m of k,m , where C is the instance key N k,m The norm of and the instance feature F t (x, y) after taking the square of the norm and multiplying the results, i.e.: C = |N k,m | 2 ×|F t (x,y)| 2 Step 3: Calculate the compensated instance features using the following formula (4) Input the obtained compensated instance features into the decoder to generate the final segmentation score map 5. The open hybrid domain adaptation image segmentation method according to claim 1, wherein, In step 4, the loss function is calculated as follows: Step 1: Calculate the loss function of the source domain image Use the cross-entropy function CELoss as the loss function to calculate the score map R obtained from the source domain image s and the difference loss with the true label Y s : Step 2: Calculate the loss function of the target domain image Use the cross-entropy function CELoss as the loss function to calculate the score map R obtained from the target domain image t and the score map obtained after compensation and the generated pseudo-label Y t for the differential loss; Step 3: Sum the two loss functions to obtain the final loss function:

Citation Information

Patent Citations

  • Unsupervised cross-domain self-adaptive medical image segmentation method based on deep adversarial learning

    AU2020103905A4

  • Image segmentation method and device and computer equipment

    CN113706558A