Dynamic visual segmentation method based on sub-domain self-adaption

Through the sub-domain adaptation method, the domain discriminator and cross-entropy loss are used to optimize the model structure, which solves the feature extraction difficulties and class confusion problems of the segmentation model in different scenarios and realizes efficient semantic segmentation of the model in new scenarios.

CN120726323AActive Publication Date: 2025-09-30BEIJING SPACEFLIGHT TUOPUGAO SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510797318.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-30
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Existing segmentation models face problems such as difficult feature extraction and inaccurate segmentation when faced with large differences in training scenarios, especially the class confusion phenomenon caused by global domain adaptation strategies. In addition, large models have high computing power resource requirements in real-time dynamic segmentation tasks.

Method used

A subdomain adaptation method is adopted to align subdomain features of the same category between the source domain and the target domain. By reconstructing the classifier into a domain discriminator and combining cross-entropy and symmetric cross-entropy loss training models, the model structure is optimized and the influence of noise is reduced, thus achieving subdomain feature alignment under the same category.

Benefits of technology

It enhances the generalization ability of the model in new scenarios, optimizes the real-time and robustness of the model, alleviates the classification bias in the global domain adaptation process, and improves the real-time and accuracy of semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726323A_ABST
    Figure CN120726323A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic vision segmentation method based on subdomain self-adaption, and belongs to the field of machine vision. The method comprises the following steps: firstly, acquiring a source domain data set image sample, a label and a target domain image sample, and constructing a verification set; and training the model by using source domain data in combination with cross entropy loss, and verifying with average pixel accuracy. Reconstructing a source domain classifier into a domain discriminator, setting the number of iterations, extracting and classifying target domain sample features through a trained encoder and classifier of a source domain, distributing domain labels to similar samples of the source domain and the target domain, training the encoder and the domain discriminator by using binary cross entropy loss, and inverting the labels for retraining. And finally, selecting a sample of which a target domain classification result is consistent with a source domain, integrating the sample with source domain data into a new data set, and further training and verifying the model by using symmetric cross entropy loss. Through sub-domain alignment, the domain adaptation effect is improved, class confusion is relieved, and the model generalization ability and the reasoning efficiency are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine vision, and in particular relates to a dynamic vision segmentation method based on subdomain adaptation. Background Art

[0002] Semantic segmentation is a highly influential technology in the field of computer vision. Compared to image classification and object detection algorithms, it can accurately separate various objects in an image at the pixel level. However, this technology also faces generalization challenges when applied to environments that differ significantly from the training scenarios, resulting in difficulties in feature extraction and inaccurate segmentation.

[0003] Segment Anything Model (SAM), a large-scale visual model technology, has excellent generalization capabilities in zero-shot tasks. However, this technology requires high computing power, especially for real-time dynamic segmentation tasks. Domain adaptation, a branch of transfer learning, has proven effective in addressing the problem of decreased model accuracy caused by large differences in image scenes and can compensate for the poor real-time performance of large models. However, global domain adaptation strategies can cause the model to experience class confusion when faced with object categories with similar features, mistakenly segmenting two objects with similar features that are close in distance as a single object. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present invention provides a dynamic visual segmentation method based on subdomain adaptation. Based on the model pre-trained in the source domain, domain adaptation training is performed in subdomains of the same category between the source domain and the target domain to achieve subdomain feature alignment under the same category, thereby alleviating the classification bias problem caused by cross-category feature confusion in the global domain adaptation process.

[0005] In order to achieve the above objectives, the present invention performs subdomain adaptation training on the encoder of the feature extraction part of the model and the classifier of the specific object category judgment part in the image.

[0006] A dynamic visual segmentation method based on subdomain adaptation includes the following steps:

[0007] Step S1, data collection:

[0008] From the source domain dataset D s Get all image samples X s and its corresponding label Y s , from the target domain dataset D t Get the target domain image sample X t ;

[0009] Step S2: Construct the source domain verification set V sand the target domain validation set V t :

[0010] Take 1 / 10 samples of each category in the source domain dataset D s and their annotation files as the validation set V s . The target domain dataset D t Collect 5 pictures for each category and perform manual annotation. The annotation files and image samples are used as the target domain validation set V t ;

[0011] Step S3, model pre-training:

[0012] Use the source domain dataset D s and cross-entropy loss to pre-train the encoder E and classifier C of the model. In each round of iterative training, the source domain validation set V s is used to perform the mean pixel accuracy MPA s test. When the mean pixel accuracy MPA s in the source domain validation set V s does not improve for 5 consecutive rounds, stop the pre-training. Step S4, construct the domain discriminator D:

[0013] Copy the classifier C trained in the source domain and reconstruct the classifier C. Set it to a binary classification structure. The reconstructed classifier is used as the domain discriminator D;

[0014] Step S5, start the iterative process of domain adaptation, and set the initial value of the iteration number e to 1;

[0015] ​​​​​​​​​​​​​​​​​​​​​​​​​Set it to 0, and re-input the two parts into the model to train the encoder E and domain discriminator D;

[0019] Step S10: Use encoder E and classifier C to retrain target domain dataset D. t The image samples in step S7 are subjected to feature extraction and classification, and the classification results and the image samples consistent with the source domain dataset D are taken. s Form a new dataset D ST , and feed it into the model using symmetric cross entropy loss Continue training the encoder E and classifier C.

[0020] Step S11: Use encoder E and classifier C to test the target domain set V. t Perform verification tests and record the average pixel accuracy MPA t If MPA is iterated for 5 consecutive rounds t If the value of does not increase, the iteration is terminated early; step S12, obtaining the model after domain adaptation is completed.

[0021] Furthermore, the cross entropy loss in step S3 The calculation method is:

[0022]

[0023] Where n is the source domain dataset D s The total number of samples, is the source domain dataset D s The true label of the i-th sample in, is the i-th sample, Represents the encoder pair sample The extracted eigenvalues, Represents the classification result of the classifier on the feature values ​​extracted by the encoder.

[0024] Furthermore, the average pixel accuracy MPA in step S3 s The calculation method is:

[0025]

[0026] Where K represents the total number of categories, TP c Indicates the number of pixels correctly predicted to be of category c, FN c Indicates the number of pixels that are actually of category c but are predicted to be of other categories.

[0027] Furthermore, the binary cross entropy loss in step S8 The calculation method is:

[0028]

[0029] Where n is the source domain dataset D s The total number of samples, is the source domain dataset D s The true label of the i-th sample in, is the predicted probability distribution of the model for the i-th sample, is the i-th sample.

[0030] Furthermore, the symmetric cross entropy loss in step S10 The calculation method is:

[0031]

[0032] in is the cross entropy loss, is the inverse cross entropy loss;

[0033] The inverse cross entropy loss is calculated as:

[0034]

[0035] Therefore, the symmetric cross entropy loss in step S10 The calculation method can be transformed into:

[0036]

[0037] The beneficial effects of the present invention are as follows: a dynamic visual segmentation method based on subdomain adaptation is provided, which enhances the generalization of the model in new scenarios through domain adaptation with subdomain alignment. The model structure is optimized by reconstructing the classifier as a domain discriminator, the model inference speed is optimized, and the real-time performance of semantic segmentation is enhanced. At the same time, in order to address the noise impact problem in the domain adaptation training process, a symmetric cross-entropy loss is used to reduce the impact of noise in the training process, thereby enhancing the robustness of the model and achieving alignment of subdomain features under the same category, thereby alleviating the classification bias problem caused by cross-class feature confusion in the global domain adaptation process.

[0038] The present invention will be further explained in detail below with reference to the accompanying drawings and specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flow chart of the dynamic visual segmentation method based on subdomain adaptation of the present invention;

[0040] Figure 2 Schematic diagram of the model training process in the dynamic visual segmentation method based on subdomain adaptation of the present invention;

[0041] Figure 3 It is a schematic diagram of a method for reconstructing a classifier into a domain discriminator in the dynamic visual segmentation method based on subdomain adaptation of the present invention. DETAILED DESCRIPTION

[0042] A dynamic visual segmentation method based on subdomain adaptation, such as Figure 1-2 As shown in the figure, the subdomain adaptation technology is used to improve the generalization and real-time performance of the model in semantic segmentation tasks, effectively solving the problems of model generalization, real-time inference speed and category confusion in computer vision semantic segmentation. Figure 2 In the decoder part, it is a conventional inverse encoding operation and will not be described in detail. The method includes the following steps:

[0043] Step S1, data collection:

[0044] From the source domain dataset D s Get all image samples X s and its corresponding label Y s , from the target domain dataset D t Get the target domain image sample X t ;

[0045] The label Y s is the image sample X s Annotation file that labels each pixel with a category.

[0046] Step S2: Construct the source domain verification set V s and the target domain validation set V t :

[0047] Take the source domain dataset D s 1 / 10 samples of each category and their annotation files are used as the validation set. t In the above example, 5 images are collected for each category and manually annotated. The annotated files and image samples are used as the target domain verification set V. t ;

[0048] Step S3, model pre-training:

[0049] Using the source domain dataset D s and cross entropy loss Pre-train the encoder E and classifier C of the model, and use the source domain verification set V for each round of iterative training s Perform the average pixel accuracy test. When the model is tested on the source domain validation set V for 5 consecutive rounds, s The average pixel accuracy MPA s When there is no improvement, stop pre-training.

[0050] The cross entropy loss is calculated as:

[0051]

[0052] where n is the total number of samples in the source domain dataset D s , is the true label of the i-th sample in the source domain dataset D s , is the i-th sample represents the feature value extracted by the encoder for the sample , represents the classification result of the classifier for the feature value extracted by the encoder )]]is the cross-entropy loss

[0053] The calculation method of the average pixel accuracy is as follows:

[0054]

[0055] ]>]where K represents the total number of categories, TP c represents the number of pixel points correctly predicted as category c, FN c represents the number of pixel points that are actually category c but are predicted as other categories, MPA s represents the average pixel accuracy

[0056] Step S4, construct the domain discriminator D:

[0057] Duplicate the classifier C trained in the source domain and reconstruct the classifier C, set it as a binary classification structure, and use this classifier as the domain discriminator D. The reconstruction method is as Figure 3 shown, adding a fully connected layer on the basis of the original classifier C

[0058] Step S5, start the iterative process of domain adaptation, and set the initial value of the iteration number e to 1;

[0059] Step S6, set the number of executions to i times; if e < i, execute Step S7; if e >= i, go to Step S12;

[0060] Step S7, use the encoder E and classifier C trained in the source domain to extract features and classify the image samples in the target domain dataset D t ;

[0061] Step S8, assign domain labels to the image samples of the same category in the source domain and target domain datasets. The source domain part of the source domain label Y sd is set to 0, and the target domain part of the target domain label Y td is set to 1, and these two parts are re-input into the model and the binary cross-entropy loss is used to train the encoder E and the domain discriminator D. The calculation method of the binary cross-entropy loss is as follows:<s

[0062]

[0063] Where n is the source domain dataset D s The total number of samples, is the source domain dataset D s The true label of the i-th sample in, is the predicted probability distribution of the model for the i-th sample, is the i-th sample, is the binary cross entropy loss.

[0064] Step S9: Reverse the domain label in step S8, and the source domain label Y of the source domain part sd Set to 1, the source domain label Y of the target domain part td Set it to 0, and re-input the two parts into the model to train the encoder E and domain discriminator D;

[0065] Step S10: Use encoder E and classifier C to retrain target domain dataset D. t The image samples in step S7 are subjected to feature extraction and classification, and the classification results and the image samples consistent with the source domain dataset D are taken. s Form a new dataset D ST , and feed it into the model using symmetric cross entropy loss Continue training the source domain encoder E and classifier C. The symmetric cross entropy loss is calculated as follows:

[0066]

[0067] in is the cross entropy loss, is the inverse cross entropy loss, and:

[0068]

[0069] Therefore:

[0070]

[0071] The reason for using symmetric cross entropy here is that when the image data of the target domain is put into the source domain for joint training, there will inevitably be noise data in the target domain. Therefore, it is necessary to use the robustness of the symmetric cross entropy loss to reduce the impact of the noise data.

[0072] Step S11: Use encoder E and classifier C to test the target domain set V. t Perform verification tests and record the average pixel accuracy as MPA t ; If MPA is iterated for 5 consecutive rounds t If the value of does not increase, the iteration is terminated early. t The calculation method is the same as that of MPA in step S3 s The calculation method is consistent;

[0073] Step S12: Obtaining a model after domain adaptation is completed.

[0074] Based on the model pre-trained in the source domain, this method adopts a discriminative adversarial domain adaptation method to narrow the difference in feature distribution extracted by the model in the source domain and the target domain. Then, pseudo-labeling is used on the classification results made by the classifier to achieve subdomain alignment between the same categories to alleviate the class confusion phenomenon in the global domain adaptation process. This not only enhances the generalization ability of the model in the target domain, but also retains the real-time performance of fast segmentation.

[0075] Finally, it should be noted that the above is only used to illustrate the technical solution of the present invention and is not limiting. Although the present invention is described in detail with reference to the preferred arrangement scheme, ordinary technicians in this field should understand that the technical solution of the present invention (such as the use of various formulas, the sequence of steps, etc.) can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A dynamic visual segmentation method based on subdomain adaptation, characterized in that: It includes the following steps: Step S1, data collection: From the source domain dataset D s Get all image samples X s and its corresponding label Y s , from the target domain dataset D t Get the target domain image sample X t ; Step S2: Construct the source domain verification set V s and the target domain validation set V t : Take the source domain dataset D s 1 / 10 samples of each category and their annotation files are used as the validation set V s , target domain dataset D t In the above example, 5 images are collected for each category and manually annotated. The annotated files and image samples are used as the target domain verification set V. t ; Step S3, model pre-training: Using the source domain dataset D s and cross entropy loss Pre-train the encoder E and classifier C of the model, and use the source domain verification set V for each round of iterative training s Perform average pixel accuracy MPA s Test, when the model is tested on the source domain validation set V for 5 consecutive rounds s The average pixel accuracy MPA s When there is no improvement, stop pre-training; Step S4, constructing domain discriminator D: Duplicate the classifier C trained in the source domain and reconstruct the classifier C, set it as a binary classification structure, and the reconstructed classifier is used as the domain discriminator D; Step S5, start the iterative process of domain adaptation, and set the initial value of the iteration number e to 1; Step S6, set the execution number to i times; if e < i, execute Step S7; if e >= i, transfer to Step S12; Step S7: Use the encoder E and classifier C trained in the source domain to train the target domain dataset D. t Feature extraction and classification of image samples in the image; Step S8: Get the source domain dataset D s and the target domain dataset D t Assign domain labels to image samples of the same category in the source domain label Y sd Set to 0, target domain label Y td Set to 1 and re-enter the two parts into the model using binary cross entropy loss Train the encoder E and domain discriminator D; Step S9: Reverse the domain label in step S8, and the source domain label Y sd Set to 1, target domain label Y td Set it to 0, and re-input the two parts into the model to train the encoder E and domain discriminator D; Step S10: Use encoder E and classifier C to retrain target domain dataset D. t The image samples in step S7 are subjected to feature extraction and classification, and the classification results and the image samples consistent with the source domain dataset D are taken. s Form a new dataset D ST , and feed it into the model using symmetric cross entropy loss Continue training encoder E and classifier C; Step S11: Use encoder E and classifier C to test the target domain set V. t Perform verification tests and record the average pixel accuracy MPA t If MPA is iterated for 5 consecutive rounds t If the value of does not increase, the iteration is terminated early; Step S12, obtain the model after domain adaptation is completed.

2. The dynamic visual segmentation method based on subdomain adaptation according to claim 1, characterized in that: The calculation method of cross-entropy loss in Step S3 is: Where n is the source domain dataset D s The total number of samples, is the source domain dataset D s The true label of the i-th sample in, is the i-th sample, Represents the encoder pair sample The extracted eigenvalues, Represents the classification result of the classifier on the feature value extracted by the encoder, is the cross entropy loss.

3. The dynamic visual segmentation method based on subdomain adaptation according to claim 1, characterized in that: The calculation method of average pixel accuracy in Step S3 is: Where K represents the total number of categories, TP c Indicates the number of pixels correctly predicted to be of category c, FN c MPA represents the number of pixels that are actually in category c but are predicted to be in other categories. s Indicates the average pixel accuracy.

4. The dynamic visual segmentation method based on subdomain adaptation according to claim 1, characterized in that: The calculation method of binary cross-entropy loss in Step S8 is: Where n is the source domain dataset D s The total number of samples, is the source domain dataset D s The true label of the i-th sample in, is the predicted probability distribution of the model for the i-th sample, is the i-th sample, is the binary cross entropy loss.

5. The dynamic visual segmentation method based on subdomain adaptation according to claim 1, characterized in that: The calculation method of symmetric cross-entropy loss in Step S10 is: in is the cross entropy loss, is the inverse cross entropy loss, and: Therefore:

Citation Information

Patent Citations

  • Robust field adaptive image learning method based on self-training noise label correction

    CN114283287A

  • Representative feature alignment-based domain adaptation target detection method

    CN114529753A

  • Unsupervised subdomain adaptation method based on variational auto-encoder

    CN117972392A

  • Latent code for unsupervised domain adaptation

    EP3767536A1

  • Method, device, and storage medium for targeted adversarial discriminative domain adaptation

    US20240185555A1