Small sample segmentation method and system based on semantic transmission and context distribution modeling
This few-sample segmentation method, which utilizes semantic transfer and contextual distribution modeling, solves the problem of existing technologies struggling to balance object structure and fine-grained boundaries, achieving high-precision segmentation even in complex backgrounds and with significant differences between samples.
Patent Information
- Application Number
- CN202510327479.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing few-sample segmentation methods struggle to simultaneously consider both the overall structure and fine-grained boundaries of the target object, resulting in insufficient generalization performance. Furthermore, they neglect the semantic information of the pre-trained model, leading to inaccurate segmentation results in complex backgrounds or when there are significant differences between samples.
We employ a semantic transfer and context distribution modeling approach. The concept-aware generation module extracts semantic information of the target category, and the context distribution mining module models the internal similarity of objects in spatial and channel dimensions. We combine multi-dimensional mining of query features to generate enhanced query features.
It improves the accuracy and stability of small sample segmentation, and can accurately segment target objects in complex backgrounds and with large differences between samples, while maintaining semantic consistency and spatial integrity.
Smart Images

Figure CN120182605B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and particularly relates to a small sample segmentation method and system based on semantic transmission and context distribution modeling. BACKGROUND
[0002] The statements in this section merely provide background technology related to the present application and do not necessarily constitute prior art.
[0003] A common semantic segmentation task requires a large number of samples for training, and the model after training can recognize all learned classes of objects. The core challenge of the small sample segmentation task is to generalize the model trained based on the basic class to new classes without additional training. To solve this problem, the Episodic Training Strategy is widely used. This framework simulates the scene of small sample testing during the training process, so that the model can segment the target objects in the query sample by referring to the support sample (Support Samples) and its prompt information. This method organizes the overall training process into a number of small tasks, each of which simulates a task containing a small amount (usually only 1-5) of labeled support set and unlabeled query set. Through repeated solving of these tasks, the model can quickly learn and generalize to adapt to new tasks under limited data.
[0004] Existing small sample segmentation methods can be roughly divided into two categories: prototype-based methods and similarity-based methods. The prototype-based method generates a prototype from the support image and its mask through Masked Average Pooling (MAP), thereby compressing the representation of the target class, and then uses a simple classifier or a more complex decoder to transfer this information to achieve segmentation of the query image. In this method, the prototype captures the core concept-level semantic features of the target class, which helps to accurately locate the target object in the query image. However, due to the loss of spatial context information, local details may not be adequately represented, leading to challenges in accurate segmentation. In contrast, similarity-based methods focus on pixel-level and spatial information, and emphasize context details. This method usually calculates a similarity map between multi-scale support foreground pixels and all query pixels, and explores the potential relationship of objects in the image through additional multi-layer perceptron (MLP) or Transformer encoder. Although the similarity-based method can effectively maintain spatial integrity and reduce local errors, it tends to ignore core semantic concepts, which may lead to macro-level errors, especially when there is a significant visual difference between the support sample and the query sample. Some researchers have introduced a learnable label to encode semantic information to enhance the generalization ability of the model. However, such methods often ignore the potential of classification labels used in the pre-training stage of the Transformer model.
[0005] In summary, the existing small sample segmentation method at least has the following shortcomings and deficiencies: (1) the current model usually only uses global semantic prototype or local similarity information, it is difficult to take into account the overall structure and fine-grained boundary of the target object; (2) due to the existing method, when facing complex background or large difference between samples, the generalization performance is insufficient, and it is difficult to stably reproduce the target semantics under small sample conditions; (3) in the current method based on the Transformer model, additional training parameters are usually used to capture the semantic information of the sample, ignoring the role of the original parameters of the model. SUMMARY
[0006] In order to solve the problems of the prior art, the present application provides a small sample segmentation method and system based on semantic transmission and context distribution modeling, which combines semantic transmission and context distribution modeling modules, captures class semantic representation and spatial and channel feature overall similarity distribution information, and enhances the information transmission between support samples and query samples.
[0007] In order to achieve the above purpose, the present application adopts the following technical scheme:
[0008] In a first aspect, the present application provides a small sample segmentation method based on semantic transmission and context distribution modeling.
[0009] A small sample segmentation method based on semantic transmission and context distribution modeling includes the following processes:
[0010] Generate a support image with only foreground from a support image and a corresponding foreground mask, process the support image with only foreground, and extract the class core semantic representation of the target region through a learnable class label;
[0011] Connect the class core semantic representation with the image block label of the support image and the query image to obtain support features and query features;
[0012] Use the foreground mask to segment the support features into foreground features and background features corresponding to the support image, model the spatial and channel distribution between the foreground features, the background features and the query features, and generate enhanced query features by mining the features similar to the overall object in the query features in multiple dimensions;
[0013] Average pool the foreground features and the background features to form foreground prototypes and background prototypes, respectively, and predict the segmentation result of the query image based on the cosine similarity between the foreground prototypes, the background prototypes and the enhanced query features.
[0014] In a second aspect, the present application provides a small sample segmentation system based on semantic transmission and context distribution modeling.
[0015] A small sample segmentation system based on semantic transmission and context distribution modeling, comprising:
[0016] A concept perception generation module is configured to generate a foreground-only support image according to a support image and a corresponding foreground mask, process the foreground-only support image, and extract a class core semantic representation of a target region through a learnable class label.
[0017] A concept feature integration module is configured to connect the class core semantic representation with image block labels of the support image and a query image to obtain support features and query features.
[0018] A context distribution mining module is configured to divide the support features into foreground features and background features corresponding to the support image using the foreground mask, model spatial and channel distributions between the foreground features, the background features and the query features, mine features similar to the whole object in the query features through multi-dimensional mining, and generate enhanced query features.
[0019] A segmentation result generation module is configured to perform average pooling on the foreground features and the background features to form foreground prototypes and background prototypes, and predict a segmentation result of the query image based on cosine similarity between the foreground prototypes, the background prototypes and the enhanced query features.
[0020] In a third aspect, the present application provides a computer device, comprising a processor and a computer readable storage medium.
[0021] The processor is adapted to execute a computer program.
[0022] The computer readable storage medium has a computer program stored therein, and the computer program is executed by the processor to implement the small sample segmentation method based on semantic transmission and context distribution modeling according to the first aspect of the present application.
[0023] In a fourth aspect, the present application provides a computer readable storage medium having a computer program stored therein, and the computer program is adapted to be loaded and executed by a processor to implement the small sample segmentation method based on semantic transmission and context distribution modeling according to the first aspect of the present application.
[0024] In a fifth aspect, the present application provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the small sample segmentation method based on semantic transmission and context distribution modeling according to the first aspect of the present application.
[0025] Compared with the prior art, the present application has the following beneficial effects:
[0026] 1. The application innovatively proposes a brand-new small sample segmentation (FSS) framework-CCFormer, which integrates semantic transmission and context distribution modeling modules. It captures the overall similarity distribution information of class semantic representation and spatial and channel features, and enhances the information transmission between support samples and query samples.
[0027] 2. The application innovatively proposes a concept perception generation (CPG) module, which effectively utilizes the key category information perception ability of the pre-trained model to generate semantic information representation for the target category. In addition, the application designs a concept feature integration (CFI) module, which integrates these category representations at an early stage of the processing flow, providing more opportunities for information interaction, so that the model can accurately locate the foreground features of the query sample.
[0028] 3. The application innovatively proposes a context distribution mining (CDM) module, which provides support for fine object internal relationship transmission between support samples and query samples. By using Brown distance covariance (BDC) to transmit information about the spatial and channel object internal similarity relationship between the two, the model can accurately segment the details of the foreground object.
[0029] The advantages of the additional aspects of the application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0030] The drawings accompanying the specification of this application form a part thereof, serve to provide further understanding of the application, and together with the description of the exemplary embodiments of the application and their description serve to explain the application, and do not constitute an improper limitation of the application.
[0031] Figure 1 A schematic diagram of the principle framework of the small sample segmentation method based on semantic transmission and context distribution modeling provided for embodiment 1 of the application;
[0032] Figure 2 A schematic diagram of the context distribution mining (CDM) module provided for embodiment 1 of the application;
[0033] Figure 3 A schematic diagram of the PASCAL-5i dataset visualization instance provided for embodiment 1 of the application;
[0034] Figure 4 A schematic diagram of the small sample segmentation system based on semantic transmission and context distribution modeling provided for embodiment 2 of the application;
[0035] Figure 5 A schematic diagram of the computer device provided for embodiment 3 of the application. DETAILED DESCRIPTION
[0036] The application will be further described below with reference to the drawings and examples.
[0037] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0038] The embodiments in the application and the features in the embodiments can be combined with each other without conflict.
[0039] Example 1:
[0040] With the development of deep neural networks and the availability of large-scale labeled datasets, semantic segmentation has made significant progress and achieved high-precision results. However, there are still challenges in obtaining a large number of pixel-level annotations, which usually require professional annotators to complete. To address these challenges, the few-shot segmentation (FSS) technology proposes to realize the segmentation of new class targets only through a small number of labeled samples. However, existing few-shot segmentation methods often only focus on extracting global semantic information or capturing local affinity relationships, resulting in problems such as inaccurate overall positioning or blurred edge details in the segmentation results.
[0041] Therefore, the present implementation proposes a small sample segmentation method (CCFormer) based on semantic transmission and context distribution modeling, which fully gives play to and combines the advantages of semantic transmission and context distribution modeling. By using the characteristics of the Transformer structure, the CCFormer proposes a concept perception generation (CPG) module. The module first performs a mask multiplication operation on the support samples to separate out only the target class region. Then, a network loaded with the pre-training backbone network weight is used to extract a classification label from the target region of the support sample, which encapsulates the semantic information of the target class. To ensure that this concept representation is effectively transmitted to the query sample segmentation process, the CCFormer further introduces a concept feature integration (CFI) module, which replaces the original model's classification label with the classification label generated by the CPG. This classification label contains key semantic information of the target class and will interact with the input image features in the backbone network processing process. By introducing the target class concept in the early stage of the backbone network and interacting with the image features, our method maintains semantic consistency in the extraction process of support set and query set features, so as to more accurately locate the foreground features in the query sample. In order to improve the transmission of context-level information in the query sample segmentation process, the CCFormer introduces a context distribution modeling (CDM) module in the decoder. Based on the assumption that objects of the same class have consistent internal similarity distribution, the module uses the Brownian distance covariance (BDC) matrix to model the overall similarity of the segmentation features. Specifically, the CDM focuses on the overall similarity between the support sample foreground and all query objects in the spatial and channel dimensions, so that the model can segment more complete objects.
[0042] More specifically, the specific method of small sample segmentation based on semantic transmission and context distribution modeling according to the present implementation comprises Figure 1 As shown in the figure, it includes a support image X s and its corresponding foreground mask M s for generating a support image with only foreground; the CPG module processes the support image with only foreground, and extracts the class core semantic representation t′ s of the target region through a learnable class label. Next, the CFI module connects t′ s with the image block label of the support image and the query image, and realizes interaction in the encoder layer. This interaction emphasizes the features similar to the target object in the support feature s and the query feature q; then, Ms Split the support feature s into foreground feature s fg and background feature s bg The CDM module further models the spatial and channel distribution between s fg , s bg and q respectively, by mining the features in the query that are similar to the object as a whole and generating enhanced query feature q'. Finally, s fg and s bg are averaged-pooled to form foreground prototype s and background prototype s The classifier then predicts the segmentation result y q of the query sample based on the cosine similarity between these prototypes and q'. During training, query prototype s q and s are also generated from q' and y to segment the support sample to produce result y s , which is only used for joint training to improve the segmentation of the query image and is not calculated during the inference test phase.
[0043] The concept perception generation module of the present implementation specifically includes:
[0044] Similar to many recent small sample segmentation, the present application uses a pre-trained Transformer model as the backbone network. These pre-trained models usually use a learnable class token to capture class information and mainly rely on this token for image classification. Although it has the potential to serve as a semantic carrier, it has not been fully utilized in existing FSS methods. The main goal of the present application is to capture key information in the foreground of the support sample, which is an important prerequisite for effectively transferring semantic information to the query image in the FSS task.
[0045] To this end, the present application proposes a CPG module that extracts target class information from the support set by using the class token in the pre-trained backbone network to form a target semantic representation. This representation is then fed into the subsequent CFI module for more targeted feature extraction and enhancement.
[0046] Specifically, CPG extracts the semantic representation of the target class from the support set, thereby achieving effective information transfer. However, since the support sample has foreground and background regions, the class token may capture unnecessary background noise while extracting foreground semantic information. To solve this problem, the present application performs Hadamard product on the support image and its corresponding mask M s to obtain a pure foreground region, so that the network obtains purer foreground features and ensures that it only retains key class information. The input obtained by this operation is denoted as xms , which is defined as follows:
[0047]
[0048] where denotes Hadamard product, and next, x ms Through the image block embedding module F pe Perform feature mapping, which is composed of a convolutional layer (whose convolution kernel size and stride are the same as the block size) and subsequent normalization and dimension reconstruction rearrangement operations, and the specific process is as follows:
[0049] t mx = F pe (x ms ) (2);
[0050] wherein, denotes the embedded image token feature, h and w are the height and width of each patch respectively, and c is the number of channels. Before the forward propagation of the model, the present application splices t mx with the learnable class label t s and sends it into the backbone network F bkb composed of several Transformer encoding layers, so that the class label and the foreground region token are fully interacted, so as to gradually refine the semantic information of the target class. This process can be represented as:
[0051]
[0052] Here, Concat represents the splicing operation, t′ s is the class label output by the final CPG module, which inherits the key semantic information from the support set and is the key information carrier for semantic transmission in the subsequent stage.
[0053] The conceptual feature integration module of the present implementation mode specifically includes:
[0054] After the CPG module generates high-quality semantic information representation t′ s from the support set, the next step is to use it for information interaction between the support sample and the query sample. Existing research shows that multi-head self-attention (MHSA) has good effect in feature interaction. Based on this, the present application integrates MHSA in this module, and uses the semantic information representation t′ s as a semantic hint to interact with t′ s in the process of extracting support and query features by the model, so as to enhance information flow and maintain the semantic consistency of the features.
[0055] In the original Transformer architecture, learnable class labels are typically used to extract key class information from images. However, in this method, we obtain the semantic information representation t′ of the target class from support samples. s Therefore, the goal of CFI is to enable image block tokens to be directly obtained from t′. s Semantic information is obtained from the data, and the Transformer coding layer MHSA is used to achieve information exchange. For support samples and query samples, this process can be represented as follows:
[0056]
[0057] in, and represents the patch embedding features of the support and query samples, respectively, where s and q represent the output support and query features, corresponding to... and
[0058] In this process, the learnable class labels in the standard Transformer are t′ s The replacement brings two advantages: First, by introducing the target class representation t′ early in feature extraction. s MHSA can be used to effectively interact with image patch tokens, enhancing the model's focus on the target region; secondly, due to t′ s Information originates from supporting samples and participates in the feature extraction of both supporting and query samples, enabling stronger information coupling and consistency between supporting and query features, thus laying the foundation for modeling the overall relationships within objects in subsequent stages.
[0059] The context distribution modeling module in this implementation specifically includes:
[0060] After CFI is completed, we obtain a result after t′ s Enhanced features are essentially an interaction method between support set and query set features in prototype-based few-sample segmentation. However, directly using the features output by CFI for prediction ignores the spatial information of foreground features in the support samples, leading to poor prediction results. To better integrate the spatial information of foreground features, existing methods often perform explicit fusion at the decoder stage, such as multi-level feature enhancement, cross-attention mechanisms between foreground features of support samples and query features, or using additional Transformer layers for information interaction. However, these methods typically rely on additional MLPs or attention mechanisms to model the relationship between pixel similarities, making it difficult to directly focus on the similarity distribution within the complete target.
[0061] Based on the assumption that the internal pixels or semantic similarity in the same category object follow consistent distribution in spatial and channel dimensions, CDM explores the co-occurrence characteristics of foreground objects in channel and spatial dimensions through Brownian covariance matrix. BDC, as a nonlinear similarity matrix, is a low-cost and effective alternative compared to complex decoders that rely on multilayer perceptron (MLP) or additional encoder layers. Complementary to the concept-level semantic emphasized in CPG and CFI, CDM supports the alignment of the similarity distribution of foreground between the query image and the support image, enhancing semantic guidance while retaining stronger spatial consistency, thus better integrating semantic cues into the final feature representation.
[0062] Brownian covariance matrix BDC is proposed by Székely et al. based on characteristic function theory, which can be understood as the Euclidean distance between the product of marginal characteristic functions and joint characteristic functions, and is defined as follows:
[0063]
[0064] wherein, and are random vectors, φ XY (t,s) is a joint characteristic function, φ X (t) and φ Y (s) are marginal characteristic functions, ||·|| represents the Euclidean norm, Γ is the gamma function. In the case of discrete observations, let wherein represents the pairwise distance of the components of the variable of the observation sample X, is defined similarly to . At this time, ρ(X,Y) can be written in the following form:
[0065]
[0066] wherein, tr(·) represents the trace of the matrix, A=(a kl ) and B=(b kl ) are the centered distance matrices of and , and each is obtained by subtracting the row mean, column mean and overall mean, and B is the same. Finally, ρ(X,Y) can also be expressed in the form of inner product as follows:
[0067]
[0068] wherein, a and b are the upper triangular part vectorization results of the A and B matrices, respectively.
[0069] The spatial and channel dimension overall co-occurrence relationship is enhanced, specifically including:
[0070] BDC metric characterizes the distribution of correlation between variables, and its output matrix reflects the local association between inputs. This property makes BDC particularly suitable for evaluating the co-occurrence of pixels in foreground features, such as capturing the overall similar objects between support and query samples in small sample segmentation. Therefore, CDM uses this property to model the dependence of support foreground and query features in channel and spatial dimensions to obtain more comprehensive similar object co-occurrence information. To adapt to the demand for local fine segmentation in small sample segmentation, CDM retains the distribution of object internal similarity throughout the process, maintains the dimension of the output BDC matrix to directly use formula (8), and thus measures the similarity between distributions without vectorization.
[0071] As shown in Figure 2 , input features s fg , s bg and q are first processed by convolutional layers respectively, and then the spatial dimension is flattened to The process is as follows:
[0072] s′ fg = Flatten(Conv(s fg )) (9);
[0073] s′ bg = Flatten(Conv(s bg )) (10);
[0074] q′ = Fatten(Conv(q)) (11);
[0075] where Conv represents convolution operation, and Flatten reshapes the dimension from c x h x w to c x hw. Subsequently, CDM calculates the BDC matrix for s′ fg , s′ bg and q′ respectively:
[0076]
[0077] where BDC is used to calculate the BDC matrix, and the obtained characterize the similarity distribution of the respective input features in the channel dimension. To quantify the similarity of the distribution, CDM performs matrix multiplication on B q and R :
[0078]
[0079] where represents matrix multiplication. Next, to determine whether each channel of the query feature matches the channel of the target class, CDM performs matrix multiplication on R fgand R bg The last dimension is summed to obtain the matching score, and the maximum value of the two is taken:
[0080] A=Max(Concat(Σ -1 R fg ,Σ -1 R bg )) (14);
[0081] Where Σ -1 (·) represents the sum of the last dimension of the feature, and Max is the maximum value selection operation. The output As a foreground channel similarity mask, it is used to highlight the channels more similar to the support foreground. Subsequently, A is applied to q to enhance the channels corresponding to the target class, and matrix multiplication is performed with q to enhance q:
[0082]
[0083] Where T represents matrix transposition, and the spatial dimension also performs similar operations to obtain Finally, the two are added to the original query feature, and layer normalization operation is performed:
[0084]
[0085] Where LN represents layer normalization, and Reshape restores the feature to the original spatial dimension. By enhancing the query feature in the channel and spatial dimensions, CDM can better capture the overall co-occurrence relationship within the object, thereby improving the fine description ability of the target region.
[0086] After passing through the CPG, CFI and CDM modules, the query feature is gradually strengthened in macro positioning, target feature integrity and local details, forming the enhanced feature representation At the same time, the support feature highlights the target class region and is further divided into foreground and background parts. The present application performs average pooling on the non-zero regions in s fg and s bg , to obtain the foreground and background prototypes of the support set, denoted as and Subsequently, the classifier F cls generates the final query image segmentation result y according to the similarity of these prototypes and q , which can be represented as:
[0087]
[0088] In the training stage, in order to further enhance the expression ability of the support feature s and indirectly improve the accuracy of the query segmentation result, we will yq Using the foreground mask of the query set, similar segmentation operations are performed on the support samples to obtain the support segmentation result y. s For all segmentation results, CCFormer uses the cross-entropy loss function to optimize the model. The overall objective function is defined as follows:
[0089] L = L ce (y q M q )+L ce (y s M s (18);
[0090] Where L ce Representing cross-entropy loss, when the model is extended to multi-support sample scenarios, the CDM module will process the data obtained from multiple inferences. and Perform mean processing, and then execute the same follow-up operations as in a single inference.
[0091] like Figure 3 As shown, from left to right, the images are: support set image and foreground object, query set image and corresponding segmented object, QPENet prediction result, HSNet prediction result, FPTrans prediction result, and the prediction result of the method of this invention. It can be seen that prototype-based methods (such as QPENet and FPTrans) can roughly locate the "bottle" target in the first row; however, HSNet, based on similarity, only predicts the upper half of the bottle, revealing its insufficient understanding of the target's semantics. In contrast, the method CCFormer of this invention combines the advantages of both types of methods, not only accurately locating the entire target but also obtaining more refined segmentation results. When the foreground of the support image differs significantly from the target in the query image (as shown in the second row), QPENet and FPTrans struggle to find the target location, and HSNet can only identify some similar objects; while CCFormer effectively addresses this challenge by using CPG and CFI modules for semantic interaction during the feature extraction stage and further mining contextual information using the CDM module, enabling accurate target location and segmentation. In the visualization results in rows 3 to 6, all methods can roughly predict the foreground region, but CCFormer consistently demonstrates higher accuracy and stability in terms of segmentation fineness.
[0092] In summary, the CCFormer method proposed in the present application combines the advantages of the two paradigms, which can accurately locate the target at the macro level and retain the details and accuracy of segmentation at the micro level. Through the introduction of the concept perception generation (CPG) and concept feature integration (CFI) modules in the feature extraction stage, the CCFormer can extract and deliver key semantic information from a small amount of support samples. The context distribution mining (CDM) module further improves the accuracy of the segmentation result by deeply modeling the object internal correlation of the foreground target in the channel and spatial dimensions. The experimental results on two public datasets show that the CCFormer can achieve more accurate and fine segmentation results in scenes with large target appearance differences or complex backgrounds, significantly outperforming the existing optimal technology.
[0093] Embodiment 2
[0094] As Figure 4 shown, the present implementation provides a small sample segmentation system based on semantic transmission and context distribution modeling, which includes:
[0095] The concept perception generation module is configured to generate a support image with only foreground according to the support image and the corresponding foreground mask, process the support image with only foreground, and extract a class core semantic representation of the target region through a learnable class label.
[0096] The concept feature integration module is configured to connect the class core semantic representation with the image block label of the support image and the query image to obtain support features and query features.
[0097] The context distribution mining module is configured to use the foreground mask to divide the support features into foreground features and background features corresponding to the support image, model the spatial and channel distributions between the foreground features, the background features, and the query features, mine the features similar to the whole object in the query features through multi-dimensional mining, and generate enhanced query features.
[0098] The segmentation result generation module is configured to perform average pooling on the foreground features and the background features to form foreground prototypes and background prototypes, respectively, and predict the segmentation result of the query image based on the cosine similarity between the foreground prototypes, the background prototypes, and the enhanced query features.
[0099] The specific working method is described in Embodiment 1 and will not be repeated here.
[0100] It can be understood that the above-mentioned various units can be combined into one or several other units respectively or entirely, or some of the units can be further split into a plurality of units with smaller functions to constitute, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions, and in actual application, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the system can also include other units, and in actual application, these functions can also be assisted by other units, and can be implemented by multiple units in cooperation.
[0101] According to another embodiment of the present application, the system described in the embodiment can be constructed and the method of the embodiment 1 of the present application can be implemented by running a computer program (including program codes) capable of performing each step involved in the corresponding method described in the embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), a Read Only Memory (ROM), etc., the computer program can be recorded on a computer readable recording medium, loaded into the above-mentioned computing device through the computer readable recording medium, and run therein.
[0102] Embodiment 3:
[0103] As shown in Figure 5 The present implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer readable storage medium 1003. The processor 1001, the communication interface 1002, and the computer readable storage medium 1003 can be connected through a bus or other means.
[0104] The communication interface 1002 is configured to receive and send data, the computer readable storage medium 1003 can be stored in the memory of the electronic device, the computer readable storage medium 1003 is configured to store a computer program, the computer program includes program instructions, and the processor 1001 is configured to execute the program instructions stored in the computer readable storage medium 1003.
[0105] The processor 1001 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, which is suitable for implementing one or more instructions, and is particularly suitable for loading and executing one or more instructions to implement a corresponding method flow or a corresponding function.
[0106] The processor 1001 is configured to perform the following processes:
[0107] generating a foreground-only support image according to the support image and the corresponding foreground mask, processing the foreground-only support image, extracting a class core semantic representation of the target region through a learnable class label;
[0108] connecting the class core semantic representation with image block labels of the support image and the query image to obtain support features and query features;
[0109] segmenting the support features into foreground features and background features corresponding to the support image using the foreground mask, modeling spatial and channel distributions among the foreground features, the background features, and the query features, mining features similar to the object as a whole in the query features through multi-dimensional mining, and generating enhanced query features;
[0110] averaging and pooling the foreground features and the background features to form foreground prototypes and background prototypes, respectively, and predicting a segmentation result of the query image based on cosine similarities between the foreground prototypes, the background prototypes, and the enhanced query features.
[0111] The specific working method is described in Embodiment 1, which will not be repeated here.
[0112] Embodiment 4
[0113] The present implementation provides a computer readable storage medium (Memory), which is a memory device in an electronic device, used for storing programs and data. It can be understood that the computer readable storage medium here can include a built-in storage medium in the electronic device, and of course can also include an expansion storage medium supported by the electronic device. The computer readable storage medium provides a storage space, which stores a processing system of the electronic device.
[0114] In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium here can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory; optionally, it can also be at least one computer readable storage medium located away from the aforementioned processor.
[0115] In one embodiment, the computer readable storage medium stores one or more instructions; the processor loads and executes the one or more instructions stored in the computer readable storage medium to implement the following processes:
[0116] generating a foreground-only support image from the support image and the corresponding foreground mask, processing the foreground-only support image, and extracting a class core semantic representation of the target region through a learnable class label;
[0117] connecting the class core semantic representation with image block labels of the support image and the query image to obtain support features and query features;
[0118] segmenting the support features into foreground features and background features corresponding to the support image using the foreground mask, modeling spatial and channel distributions among the foreground features, the background features, and the query features, and generating enhanced query features by mining features similar to the whole object in the query features in multiple dimensions;
[0119] performing average pooling on the foreground features and the background features to form foreground prototypes and background prototypes, respectively, and predicting a segmentation result of the query image based on cosine similarities between the foreground prototypes, the background prototypes, and the enhanced query features.
[0120] The specific working method is described in Embodiment 1, which will not be repeated here.
[0121] Embodiment 5
[0122] The present implementation provides a computer program product or a computer program including computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the following processes:
[0123] generating a foreground-only support image from the support image and the corresponding foreground mask, processing the foreground-only support image, and extracting a class core semantic representation of the target region through a learnable class label;
[0124] connecting the class core semantic representation with image block labels of the support image and the query image to obtain support features and query features;
[0125] segmenting the support features into foreground features and background features corresponding to the support image using the foreground mask, modeling spatial and channel distributions among the foreground features, the background features, and the query features, and generating enhanced query features by mining features similar to the whole object in the query features in multiple dimensions;
[0126] performing average pooling on the foreground features and the background features to form foreground prototypes and background prototypes, respectively, and predicting a segmentation result of the query image based on cosine similarities between the foreground prototypes, the background prototypes, and the enhanced query features.
[0127] The specific working method is described in the embodiment 1, which will not be repeated here.
[0128] Those skilled in the art can appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0129] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, all or part can be realized in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions can be stored in a computer-readable storage medium or transmitted by a computer-readable storage medium. Computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server, data center, etc. containing one or more available media sets. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk (SSD)) and the like.
[0130] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A small sample segmentation method based on semantic transfer and context distribution modeling, characterized in that, The method comprises the following processes: generating a foreground-only support image according to a support image and a corresponding foreground mask, processing the foreground-only support image, and extracting a class core semantic representation of a target region through a learnable class label; connecting the class core semantic representation with image block labels of the support image and the query image to obtain support features and query features; The support features are segmented into foreground features and background features corresponding to the support image using the foreground mask, and the spatial and channel distributions among the foreground features, the background features and the query features are respectively modeled, and features similar to the whole object in the query features are mined in multiple dimensions to generate enhanced query features performing average pooling on the foreground features and the background features to form foreground prototypes and background prototypes, respectively, and predicting a segmentation result of the query image based on cosine similarity between the foreground prototypes, the background prototypes and enhanced query features; performing convolution layer processing on the foreground features, the background features and the query features corresponding to the support image, respectively, and performing spatial dimension flattening processing on results after the convolution layer processing; calculating Brown covariance matrices for the foreground features, the background features and the query features after the spatial dimension flattening processing, respectively; performing matrix multiplication on the Brownian covariance matrix of the foreground feature and the Brownian covariance matrix of the query feature to obtain performing matrix multiplication on the Brownian covariance matrix of the foreground feature and the Brownian covariance matrix of the query feature to obtain ; to and summing over the last dimension to obtain a match score and taking the maximum of the two; applying a maximum value to the query features to enhance channels corresponding to a target class, performing matrix multiplication on the query features and the query features to enhance the query features to obtain a channel dimension enhancement result; performing spatial dimension enhancement on the foreground features, the background features and the query features after the spatial dimension flattening processing to obtain a spatial dimension enhancement result; adding the channel dimension enhancement result, the spatial dimension enhancement result and the query features and performing a normalization operation to obtain enhanced query features.
2. The small sample segmentation method based on semantic transmission and context distribution modeling according to claim 1, wherein, performing Hadamard product on the support image and the corresponding mask to obtain a foreground-only support image; performing feature mapping on the foreground-only support image through an image block embedding module; splicing the feature mapping result with a learnable class label and inputting the same into a backbone network composed of a plurality of Transformer encoding layers to make the class label fully interact with the foreground region to obtain a class core semantic representation.
3. The small sample segmentation method based on semantic transmission and context distribution modeling according to claim 1, wherein, connecting the class core semantic representation with image block labels of the support image and the query image to obtain support features and query features, comprising: ; ; wherein, and are patch embedding features for support and query samples, respectively, and represent output support and query features, corresponding to and , denotes an image patch embedding module, denotes a class core semantic representation, is a processed class core semantic representation.
4. The small sample segmentation method based on semantic transmission and context distribution modeling according to claim 1, wherein, based on the foreground prototype and the background prototype and the cosine similarity between the enhanced query features , the segmentation result of the query image is predicted , comprising: ; wherein represent a classifier.
5. The small sample segmentation method based on semantic transmission and context distribution modeling according to any one of claims 1-4, wherein, during training, generating foreground prototypes and background prototypes corresponding to the query image from the enhanced query features and a segmentation result of the query sample, segmenting the support image, and predicting a segmentation result of the support image; adopting a cross-entropy loss function to optimize the model, and an overall objective function is: ; wherein, denotes a cross-entropy loss, denotes a segmentation result of the query image, denotes a segmentation result of the support image, denotes a foreground mask corresponding to the support image, denotes a foreground mask corresponding to the support image.
6. A small sample segmentation system based on semantic transfer and context distribution modeling, characterized in that, comprising: a concept perception generation module configured to generate a foreground-only support image according to a support image and a corresponding foreground mask, process the foreground-only support image, and extract a class core semantic representation of a target region through a learnable class label; a concept feature integration module configured to connect the class core semantic representation with image block labels of the support image and the query image to obtain support features and query features; The context distribution mining module is configured to: segment the support features into foreground features and background features corresponding to the support image using the foreground mask, model the spatial and channel distributions between the foreground features, the background features, and the query features, and mine features similar to the object as a whole in the query features through multi-dimensional mining and generate enhanced query features The segmentation result generation module is configured to: average-pool the foreground feature and the background feature to form a foreground prototype and a background prototype respectively, and predict a segmentation result of the query image based on cosine similarity between the foreground prototype and the background prototype and an enhanced query feature; The foreground feature, the background feature and the query feature corresponding to the support image are respectively processed by a convolution layer, and the results after the convolution layer processing are subjected to spatial dimension flattening processing; Brown covariance matrices are calculated for the foreground feature, the background feature and the query feature after the spatial dimension flattening processing respectively; performing matrix multiplication on the Brownian covariance matrix of the foreground feature and the Brownian covariance matrix of the query feature to obtain performing matrix multiplication on the Brownian covariance matrix of the background feature and the Brownian covariance matrix of the query feature to obtain ; to and the last dimension is summed to obtain a match score, and the maximum of both is taken; A maximum value is applied to the query feature to enhance the channel corresponding to the target class, and matrix multiplication is performed on the query feature to enhance the query feature to obtain a channel dimension enhancement result; The foreground feature, the background feature and the query feature after the spatial dimension flattening processing are subjected to spatial dimension enhancement to obtain a spatial dimension enhancement result; The channel dimension enhancement result, the spatial dimension enhancement result and the query feature are added and subjected to normalization operation to obtain an enhanced query feature.
7. A computer device, comprising: It comprises: a processor and a computer readable storage medium; a processor adapted to execute a computer program; a computer readable storage medium having a computer program stored therein, wherein the computer program is executed by the processor to implement the small sample segmentation method based on semantic transmission and context distribution modeling according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the small sample segmentation method based on semantic transmission and context distribution modeling according to any one of claims 1 to 5.
9. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program is executed by the processor to implement the small sample segmentation method based on semantic transmission and context distribution modeling according to any one of claims 1 to 5.
Citation Information
Patent Citations
Small sample semantic segmentation method and system based on background information mining
CN117409413A
Metalearning fault diagnosis and identification method based on Brown covariance
CN119577536A