Cross-Domain Few-Shot Image Semantic Segmentation Method Based on Memory Mechanism
By introducing memory mechanisms into the image semantic segmentation model, imitating human experience learning and memory processes, the model's adaptation problem in the sparse labeled data and cross-domain scenarios is solved, and better generalization and transfer capabilities are achieved.
Patent Information
- Application Number
- CN202210707799.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-20
AI Technical Summary
When faced with sparse annotation data and cross-domain scenarios, existing image semantic segmentation models are difficult to achieve rapid migration and efficient adaptation, resulting in reduced segmentation effect and poor generalization.
A cross-domain small sample segmentation method based on memory mechanism is adopted to mimic human experience learning and memory mechanism. By constructing metatasks and memory modules, the model accumulates experience and stores knowledge during the training process, so as to perform selective loading and adaptation in new scenarios.
It effectively improves the adaptability of the segmentation model to new scenarios, reduces the dependence on a large amount of labeled data, and improves the generalization and migration capabilities of the model.
Smart Images

Figure CN115359250B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and small-sample image semantic segmentation. The core technology involves meta metric learning and memory mechanism. The research is carried out for the image segmentation scenario with scarce labeled data. The invention can effectively improve the generalization adaptability of the segmentation model to the environment and alleviate the dependence on a large number of image annotations. Background Art
[0002] Image semantic segmentation aims to assign accurate label information to each pixel point in an image. To adapt to the complex application scenarios, the accuracy of deep segmentation models is constantly increasing. Correspondingly, the number of parameters of the models is gradually becoming large. The large number of parameters puts higher requirements on the data required for training the models. Not only the amount of training data needs to be guaranteed, but also to improve the generalization, the collected data needs to cover various scenarios. However, it is very difficult to collect such a large amount of labeled segmentation data in a specific scenario. And when the segmentation model faces a new scenario, the effect will decrease significantly, that is, it cannot achieve fast migration. Facing the requirements of various application scenarios, how to train an efficient segmentation model with limited labeled data has become an urgent problem in the industrial and academic fields.
[0003] In the existing transfer learning of image semantic segmentation, after the model is trained, it needs to use new scenario data for transfer training. During the transfer process, the model will gradually fit the new knowledge. However, due to the increasing number of parameters of the model at the present stage, if there is not enough transfer label data, the obtained model is very likely to overfit, and the segmentation ability for the new scenario will decrease significantly. Moreover, the existing segmentation methods cannot adapt to the changing environment. However, for humans, we can recognize a new category or get familiar with a new scenario through very few pictures. This is mainly because humans have the ability to collect, store, and transfer knowledge. During the growth process, we continuously accumulate experience. When facing a new scenario or task, we review the experience collected in the past for comparative learning. In this way, even if the learning resources provided by the new scenario are limited, humans can quickly complete the adaptation. At the same time, memory is an important mechanism for human development, which enables the accumulation of experience and the transmission of information. For machine learning models, during the transfer learning process, they mainly fit the data distribution and there is no explicit memory mechanism, which is very difficult for fitting the new environment with scarce label data.
[0004] After the above analysis, the scarcity of data labels in the new scenario and the variability of the new scenario data distribution are the two main issues that this invention focuses on. These are the key factors restricting the use of large-scale segmentation models, making it difficult for the segmentation model to learn, resulting in a decline in the segmentation effect, and the ever-present scene deviation also poses a significant challenge to the generalization of the model. To address these two issues, this method mimics humans to design and develop a cross-domain image semantic segmentation model with a memory mechanism. It solves the cross-domain segmentation problem in the few-shot scenario by mimicking the human experience learning and memory mechanism. Specifically, regarding the problem of scarce labeled data, this invention uses an empirical learning strategy to replace the original model learning process. By simulating the human learning process in reality, the model is forced to continuously collect knowledge between image segmentation tasks and convert it into model parameters. Compared with the traditional training learning process, this is beneficial for the model to obtain more sensitive parameters and has higher transferability. Regarding the problem of the variable data distribution in the new scenario, this invention imitates the human perception mechanism and designs a module based on the memory mechanism to collect generalization knowledge during the training process. Through the memory module, the model can load the knowledge accumulated during the training process into the new scenario. By fusing the memory knowledge with the data features of the new scenario, the impact brought by the variable environment can be effectively diluted, and the segmentation model can focus more on the target segmentation task. Summary of the Invention
[0005] In view of this, this invention proposes a cross-domain few-shot segmentation method based on the memory mechanism. The purpose of this invention is to overcome the problem of low adaptability of image segmentation to the few-shot cross-domain scenario and solve the image segmentation problem in the case of scarce data labels and cross-domain scenarios. By mimicking the human learning process, this invention modifies the learning process of the segmentation model into empirical learning. The model continuously accumulates experience during the learning process and converts the experience into highly adaptable parameters. At the same time, based on the human perception and cognition foundation, a memory module is designed to alleviate the impact of cross-domain data distribution on model learning. The entire invention framework realizes the process of "accumulating experience and transferring memory".
[0006] During the training process of this invention, a large number of image segmentation tasks that are similar but different from each other are constructed for the segmentation model to learn. Designing these tasks is not only used to optimize the model, but also to extract transferable memory knowledge. The model continuously collects knowledge between tasks during the learning process and stores this knowledge in the memory module. As the training progresses, the tasks experienced by the model gradually increase, and the knowledge it stores becomes gradually richer and more meaningful. After the training is completed, the model parameters and the learned memory module are fixed. During the migration process, the model selectively loads the knowledge in the memory module. This knowledge is responsible for reducing the data distribution difference between images in the new scenario and, together with the image features of the new scenario, helps the model adapt to the new scenario, and finally completes the smooth migration and efficient segmentation of the model.
[0007] According to the above main idea, the specific implementation of the method of the present invention includes the following steps, including two stages: training and migration. The training stage includes steps 1-6
[0008] The present invention is an image segmentation method that does not rely on a large amount of labeled data and has high robustness for cross-scene parsing. For this purpose, a large number of image segmentation tasks need to be constructed first. We call these meta-tasks, which are used to simulate the scenes that constantly appear in the human learning process. Subsequently, the segmentation model fits the meta-tasks in batches, and finally obtains a set of highly adaptable parameters, which are called meta-knowledge. In this process, paired vector groups are used to collect the distribution of image features during the training process, and the memory of the model is formed through a read-write mechanism. Finally, when migrating and segmenting new scenes, the obtained memory is loaded to alleviate the impact brought by data cross-domain and improve the generalization of the method.
[0009] Training stage:
[0010] Step 1: Imitate and construct a large number of meta-tasks
[0011] To imitate the human learning process, the present invention first performs sampling with replacement on the image segmentation dataset. Each sampling obtains a segmentation meta-task composed of different image samples. The meta-task is used to imitate the possible future image distribution for the model to learn more generalized parameters. The specific sampling can refer to step 1 in the implementation manner. Each task includes a support image set and a query image set at the same time. The support image set is used to extract the knowledge learned in each task, and the query image set is used to adjust the parameters of the segmentation model to convert the knowledge into highly adaptable parameters of the segmentation model. Each task is used for training the model parameters.
[0012] Step 2: Initialize the memory module
[0013] The memory module is mainly used to imitate the human cognitive mechanism, store and collect the channel mean and variance of the image features in each meta-task, and finally migrate them to the learning of new scenes. At the beginning of training, we randomly initialize a group of vectors as the meta-mean and meta-variance. The two groups of vectors are used together to collect different domain distributions in the features. The specific initialization process can refer to step 2 in the implementation manner.
[0014] Step 3: Store the mean and variance of the support images in each meta-task into the memory module
[0015] The meta-mean and meta-variance in the memory module need to collect the feature mean and variance for subsequent migration. During the learning process of the model, it is necessary to continuously fit the training data. For each image sample, the segmentation model extracts the mean and variance of its corresponding features and disassembles them into the meta-mean and meta-variance of the memory module. The storage process can refer to step 3 in the implementation manner.
[0016] Step 4: Match the query image set using the support image set features
[0017] The training process of the support image set in the meta-task is similar to the human learning process, collecting experience by continuously solving segmentation meta-tasks. For each meta-task, the segmentation model extracts category features on its support image set. This process is used to imitate the human few-shot learning process, and the extracted features are used to segment the query image set. The segmentation model learns a large number of meta-tasks, collecting shared knowledge between tasks. This knowledge is reflected in the features of each image sample. Through learning a large number of tasks, the image features tend to be more shared features, which are shared by all images. For details, refer to Step 4 in the implementation manner.
[0018] Step 5: Transform the learned meta-knowledge into highly adaptable parameters of the segmentation model on the query image set
[0019] In the support image set training stage of Steps 3 - 4, more robust expressive features are obtained and the memory module is updated. Subsequently, we need to transform the knowledge obtained by the model into highly adaptable parameters of the model. This is mainly achieved through the query image set in each meta-task. Using the enhanced features of the support image set obtained in Step 5, similarity segmentation is performed on the query image set, and the final result is used to calculate the loss and back-callback the segmentation model parameters. The network constraints of the present invention refer to Step 5 in the implementation manner.
[0020] Transfer stage:
[0021] Step 6: Segment the new scene using the trained segmentation model
[0022] By continuously repeating Steps 3 - 5, the parameters of the segmentation model gradually tend to be highly adaptable as it continuously solves meta-tasks, and the meta-means and meta-variance groups stored in the memory module also gradually collect different data distributions. After training, the model parameters and the memory module are fixed. In the new scene transfer, first, the mean and variance of each image data in the new scene are obtained, and then their similarity with the meta-means and meta-variances stored in the memory module is calculated. The most similar meta-means and meta-variances are selected for loading. The image features are reconstructed using the loaded new means and variances to obtain fused and enhanced features, and the fused features are used for segmentation of the new scene. Refer to Step 6 in the implementation manner.
[0023] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects: The present invention proposes a cross-domain few-shot segmentation method based on a memory mechanism, which improves the segmentation effect of traditional segmentation networks in scenarios with insufficient data by mimicking human experience learning. At the same time, compared with the prior art, the memory mechanism proposed by this method can reduce the differences in data distributions in different scenarios and effectively extend the model obtained from single-domain training to multi-domain segmentation, which has fundamental supporting significance for segmentation tasks in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is the overall flow chart of the method involved in the present invention;
[0025] Figure 2 It is the overall architecture diagram of the algorithm involved in the present invention;
[0026] Figure 3 Configuration table of each layer structure of the backbone network;
[0027] Figure 4 Configuration table of the memory module structure;
[0028] Figure 5 Comparison of the transfer segmentation effects of the present invention and other different models on Pascal; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further details the present invention with reference to specific examples and detailed drawings. However, the described embodiments are only intended to facilitate the understanding of the present invention and do not impose any limitations on it. Figure 1 It is the method flow chart of the present invention. As Figure 1 shown, the method includes the following steps. The training stage includes steps 1 to 5, and the testing stage includes step 6.
[0030] Step 1: Mimic and construct a large number of meta-tasks
[0031] Each meta-task consists of a support segmentation image set and a query segmentation image set. The images in the two image sets have the same image category and a certain number of images. Its essence is an image segmentation task. The implementation process needs to use the features of each image in the support segmentation image set to match the query segmentation image set, so as to imitate the process of humans solving few-shot image segmentation. The present invention improves the adaptability of model parameters by continuously solving meta-tasks. First, the present invention trains a segmentation model on a segmentation data set. We choose the currently most used COCO image segmentation data set for training. Each meta-task consists of a support image set and a query image set. The support image set is used to extract the image features to be learned, and the query image set is used to train the highly adaptable parameters. To construct the support image set of each meta-task, K image samples are randomly extracted from the COCO data set as the reference for the segmentation model to learn which is used to help the segmentation model extract the image category features, where represents the sample image, Y = {Y s 1 ,..., Y s K} is the corresponding image pixel label information. Each meta-task also includes a query image set with the same category as the support images which is used to learn the parameters of the segmentation model and is also obtained by sampling. Among them and are the image to be segmented and the corresponding pixel label. According to this principle, a large number of meta-tasks are constructed by sampling with replacement.
[0032] In this embodiment, taking one meta-task as an example, the images in the support image set are a small sample set. Five images containing dogs are extracted as the support image set. The segmentation model learns the features of dogs from these five images containing dogs. The query image set contains an unlimited number of images containing dogs. This meta-task is to use the support image set to train the segmentation model, and then input the query image set into the trained segmentation model to enable the segmentation model to obtain the ability to recognize dogs; a large number of meta-tasks are constructed, and the categories included in each meta-task are different. These meta-tasks simulate the future image distribution, and the adaptability of the segmentation model parameters is improved through these meta-tasks.
[0033] Step 2: Initialize the memory module
[0034] In each meta-task, the support images interact with a common memory module, which is used to store the underlying information of the images during the training process, including texture, style, etc. Storing this underlying information can effectively improve the model's adaptability to new scenarios. The memory module is a set of trainable vectors used to collect information on image features during the training process. The memory module includes a meta-mean vector group and a meta-variance vector group, which are mainly used to collect the channel means and variances of image features to reflect domain information. We initialize the memory module randomly as follows:
[0035]
[0036] where M and E are the meta-mean vector group and the meta-variance vector group respectively, responsible for collecting image feature information during the training process. C is the number of feature channels of the image features, m j and e j are the j-th meta-mean and meta-variance in the memory module, N represents the number of meta-means in the memory module, and the number of meta-means and meta-variances are equal.
[0037] Step 3: Store the means and variances of the support images in each meta-task into the memory module
[0038] When processing each meta-task, for each image in its support segmentation image set, the segmentation model first extracts the corresponding feature f b C×H×W , where C, H, and W represent the size of the channel, the height, and the width of the corresponding feature respectively. First, calculate the mean and variance of each image feature. The mean μ b and variance v b of the b-th image are calculated as follows:
[0039]
[0040]
[0041] where, represents the image feature on the channel,
[0042] To implement the function of the memory module to collect information, calculate the similarity between the mean of each support image feature and each meta-mean, and calculate the similarity between the variance of each support image feature and each meta-variance. Among them,
[0043] the similarity b between the mean μ j of the b-th support image feature and the j-th meta-mean m is calculated as follows:
[0044]
[0045] The variance v of the b-th support image feature b and the j-th meta-variance e j The similarity between is calculated as follows:
[0046]
[0047] where B represents the number of images in each support segmentation image set,
[0048] Finally, the paired meta-means and meta-variances are updated through aggregation. Among them, the updated j-th meta-mean m j ′ and the j-th meta-variance e j ′ are expressed by the following formulas:
[0049]
[0050]
[0051] λ is the aggregation weight, which is set to 0.9 in all experiments. The style information obtained from the images is assigned to the most similar m j , e j . Thus, the distribution of the image features of the support image set in each meta-task is successfully stored in the memory module.
[0052] Step 4: Matching the query image set using the support image set features
[0053] The segmentation model needs to calculate the segmentation results of the query image set in each meta-task for subsequent loss calculation. We use the segmentation model to obtain the image feature f of each image in the query image set q C×H×W , for each f in step 3 b , the present invention first uses the label Y corresponding to the image b to screen f b to obtain the prototype expression of each support image:
[0054]
[0055] where T is the foreground pixel points screened using Y b , and |T| is the total number of foreground pixel points. The extracted prototype p b needs to segment the query image set. The specific segmentation method is as follows: First, the prototype p b is concatenated with the query image set feature f q , which provides prior information for the query image set feature; subsequently, the standard 1×1 convolution R is used to discover the relationship of the concatenated feature, map it to a single-channel relationship heat map, and finally perform upsampling to obtain the final matching result with each prototype
[0056]
[0057] Step 5: Calculate the loss using the matching results and train the model
[0058] This step transforms the experience learned by the model on each meta-task into highly adaptable parameters of the model. This is mainly achieved through the backpropagation of the model. For our overall segmentation model, the meta-mean and meta-variance in the memory module represent various style patterns. To obtain a richer style distribution, we hope that the meta-mean and meta-variance in the memory module are as independent as possible. Therefore, we developed an orthogonal loss to constrain our memory module as follows:
[0059]
[0060] where represents the cosine similarity between m i and m j , and represents the cosine similarity between e i and e j . L orth cooperates with the final segmentation loss to jointly complete the training of the model parameters. For the segmentation loss, the present invention uses the classical binary cross-entropy loss function:
[0061]
[0062] where represents the matching relationship between the q-th query image and the b-th support image at (h, w), that is the binary matching relationship at the pixel point (h, w), is the image label. At the same time, use L orth added to L to train the model and complete the learning of the highly adaptable parameters of the model.
[0063] Step 6: Use the trained segmentation model to perform segmentation in a new scenario
[0064] Steps 1 - 5 complete the training of the model. In the image segmentation of a new scenario, we need to fine-tune the trained segmentation model. Since we are facing a small-sample scenario with few image annotations, the present invention first uses the given images in the new scenario to fine-tune the model according to Step 5, and then fixes the model parameters. After completing the fine-tuning process of the image semantic segmentation model in the new scenario, the segmentation of the new scenario can be performed.
[0065] As can be seen from Table 3, the method proposed in the present invention has a better segmentation effect than the latest method on the cross-domain object segmentation dataset.
[0066] Table 1
[0067]
[0068] Table 2
[0069]
[0070] Table 3
[0071]
Claims
1. Cross-domain few-shot image semantic segmentation method based on a memory mechanism, characterized in that: It includes two stages: training and transfer. During the training process, a large number of image segmentation tasks that are similar but different from each other, namely meta-tasks, are constructed for the segmentation model to learn; During the learning process, the segmentation model continuously collects knowledge between tasks and stores this knowledge in the memory module; after the training is completed, the parameters of the segmentation model and the learned memory module are fixed; During the transfer process, the segmentation model selectively loads the knowledge in the memory module; finally, the transferred model is used for object segmentation; The training stage includes: Step 1: Imitate and construct a large number of meta-tasks: Sampling with replacement on the image segmentation dataset. Each time a sample is obtained, it is a meta-task, which is used to imitate the possible future image distribution for the segmentation model to learn more general parameters. The image segmentation dataset is composed of labeled image samples; each meta-task includes a support image set and a query image set at the same time. The image categories in the two image sets are the same, and the support image set consists of a small number of images; Step 2: Randomly initialize the memory module: The memory module consists of two sets of vectors, representing the meta-mean and meta-variance respectively. The two sets of vectors are used in combination to collect the feature distributions of each image in the support image set extracted by the segmentation model during the training process. The memory module is used to imitate the cognitive mechanism of humans and improve the adaptability of the segmentation model to new scenarios; each support image in each meta-task interacts with the memory module; Step 3: Store the mean and variance of the support images in each meta-task into the memory module; Step 4: Use the features of the support image set to match the query image set, and the matching results are used for subsequent loss calculation: For the image features of each support image obtained in Step 3, first filter the image features using the corresponding labels of the images to obtain the prototype expression of each support image. Among them, the prototype expression of the b-th support image is as follows: where T is the use of label Y b the foreground pixel points screened out, |T| is the total number of foreground pixel points, representing the pixel point features at position t in image b; The extracted prototypes are used to segment the query image set to obtain the matching results. The specific method for the prototype of the b-th support image to segment the q-th query image is as follows: First, the prototype p of the b-th support image b is concatenated with the feature f of the q-th query image q ; Subsequently, the standard 1×1 convolution R is used to discover the relationship of the concatenated features, map them to a single-channel relationship heat map, and finally perform upsampling to obtain the final matching result with the b-th prototype Step 5: Calculate the loss using the matching results and train the segmentation model: Overall loss function = L orth + L Among them, Among them, N represents the number of meta-means in the memory module, and the number of meta-means is equal to the number of meta-variances. represents m i and m j cosine similarity, and represents e i and e j cosine similarity; L is the binary cross-entropy loss function: where represents the matching relationship between the q-th query image and the b-th support image at (h, w), is the image label; H and W respectively represent the height and width of the corresponding feature; In this step, the experience learned by the segmentation model on each meta-task is converted into highly adaptable parameters of the segmentation model through the backpropagation of the model; The transfer stage includes: Step 6: Use the trained segmentation model to perform segmentation in the new scenario.
2. The cross-domain few-shot image semantic segmentation method based on the memory mechanism according to claim 1, characterized in that: The specific process of the said Step 3 is as follows: When processing each meta-task, for each image in its support image set, the segmentation model first extracts the corresponding features. C, H, and W respectively represent the size of the feature channels, the height and width of the corresponding features; then, calculate the mean and variance of each support image feature; Next, calculate the similarity between the mean of each support image feature and each meta-mean, and calculate the similarity between the variance of each support image feature and each meta-variance; finally, update the paired meta-mean and meta-variance through aggregation. Thus, the distribution of the image features of the support image set in each meta-task is successfully stored in the memory module; The specific content of the said Step 6 is as follows: Steps 1-5 complete the training of the segmentation model. In the new scene image segmentation, the trained segmentation model is fine-tuned. Specifically, the given images in the new scene are used, and the segmentation model is fine-tuned again according to step 5. Subsequently, the model parameters are fixed. After completing the fine-tuning process of the image semantic segmentation model in the new scene, the segmentation of the new scene can be carried out.
3. The cross-domain few-shot image semantic segmentation method based on a memory mechanism according to claim 2, wherein: Further, the mean μ of the b-th support image feature b and the variance v b are calculated as follows: Among them, represents the image features on the channel.
4. The cross-domain few-shot image semantic segmentation method based on the memory mechanism according to claim 3, wherein: Further, the mean value μ of the b-th support image feature b and the mean value m of the j-th element j The similarity is calculated as follows: Among them, sim represents the cosine similarity, and B represents the number of images in each support segmentation image set.
5. The cross-domain few-shot image semantic segmentation method based on the memory mechanism according to claim 4, wherein: Furthermore, The variance v of the b-th support image feature b and the j-th meta-variance e j The similarity between is calculated as follows:
6. The cross-domain few-shot image semantic segmentation method based on the memory mechanism according to claim 5, characterized in that: Further, the formula for the updated j-th meta-mean m j ' and the j-th meta-variance e j ' is expressed as follows: λ is the aggregation weight.