Semantic segmentation method for small sample intestinal polyp image

Through the semantic segmentation method of small-sample intestinal polyp image with mixed feature decoupling and multimodal feature alignment, the problems of low localization accuracy of polyps lesions and excessive dependence on training samples in the prior art are solved, and efficient polyps region segmentation and model generalization are achieved.

CN120339601APending Publication Date: 2025-07-18LANZHOU UNIV SECOND HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510241452.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing semantic segmentation model of intestinal polyps image is not highly localized for newly emerging polyps lesions, is not very generalized, and is too dependent on training samples, resulting in high segmentation costs.

Method used

A semantic segmentation method of small-sample intestinal polyp image with mixed feature decoupling and multimodal feature alignment is used to extract features through a backbone network with shared weights, and multimodal interaction and alignment is performed by combining text-coded features and visual features of the diagnostic report. A prototype set is generated and similarity calculation is performed to locate the polyp region.

Benefits of technology

Reliance on training data is reduced, the positioning accuracy and segmentation performance of the model on newly emerging polyps regions is improved, and the robustness and reliability of the model is enhanced, especially the segmentation ability of rare lesion images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339601A_ABST
    Figure CN120339601A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample intestinal polyp image semantic segmentation method, which comprises the following steps: feature extraction: mapping intestinal polyp images input by a support branch and a query branch to a depth feature space by using a group of backbone networks sharing weight; carrying out multi-modal feature alignment; performing general semantic alignment among the branches; prototype generation: capturing a prototype set of classes on the separated target foreground supporting features, query decoupling features and branch supporting multi-mode alignment features, calculating the similarity between query mixed features and each prototype on the prototype set position by position, and quickly positioning and segmenting a polyp region in a query picture according to the maximum similarity. According to the method, through decoupling of the mixed features, mutual interference among different semantic features in the mixed features is reduced, and the perceptual ability of the segmentation model to polyps and surrounding tissue structures is enhanced; meanwhile, the robustness, the reliability and the generalization of the semantic segmentation model of the intestinal polyp image can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image semantic segmentation, and particularly to a small-sample intestinal polyp image semantic segmentation method. Background Art

[0002] In the medical field, experts mainly rely on rich clinical experience to locate and segment the intestinal polyp area. However, due to the variability of polyps and the high similarity between the lesion area and the surrounding tissues, the workload of experts is significantly increased. In addition, due to the differences between individual patients and the subjectivity of experts, segmentation false alarms are inevitably caused. In recent years, deep learning has received high attention in clinical diagnosis tasks. Especially with the emergence of large models such as vision and language, the application of deep learning in clinical practice has shown a highly increasing trend, and deep learning technology is gradually changing the way doctors work.

[0003] With the rapid development of deep learning, the performance of the intestinal polyp image semantic segmentation model has made breakthrough progress. However, the performance of the existing polyp image semantic segmentation model overly relies on a large number of pixel-by-pixel annotated training samples, and the positioning accuracy of newly emerging polyp lesion areas is not high and the generalization ability is not strong. At the same time, the acquisition and pixel-by-pixel annotation cost of intestinal polyp images are extremely high, especially the acquisition and pixel-by-pixel annotation of some rare polyp lesion images are more difficult.

[0004] A small-sample intestinal polyp image semantic segmentation method proposed by the present invention reduces the mutual interference between different semantic features inside the mixed features through mixed feature decoupling, and enhances the perception ability of the segmentation model for polyps and surrounding tissue structures. In addition, the present invention can further improve the robustness, reliability and generalization ability of the intestinal polyp image semantic segmentation model. Summary of the Invention

[0005] To overcome the problems that the performance of the existing technology in the intestinal polyp image semantic segmentation method overly relies on a large number of annotated training samples and the generalization ability of the model for newly emerging polyp lesion areas is not high, the present invention proposes a small-sample intestinal polyp image semantic segmentation method.

[0006] The technical solution of the present invention is implemented as follows. A small-sample intestinal polyp image semantic segmentation method includes the steps:

[0007] S1, Feature extraction. Use a group of backbone networks with shared weights to map the intestinal polyp pictures input by the support branch and the query branch into the deep feature space.

[0008] S2, Multimodal feature alignment. Establish multimodal feature interaction and alignment between the text encoding features of the diagnostic report, the support target foreground features, and the query mixed features.

[0009] S3. General semantic alignment between branches. By establishing cross-branch semantic interaction and alignment between the target foreground features and the query mixed features, the general semantics between branches are mined to improve the decoupling effect of the mixed features.

[0010] S4. Prototype generation. Capture the prototype set of the class on the separated target foreground features of the support, the decoupled features of the query, and the multi-modal alignment features of the support branch, and calculate the similarity between the query mixed features and each prototype in the prototype set at each position. According to the maximum similarity, quickly locate and segment the polyp area in the query image.

[0011] Furthermore, the group of backbone networks with shared weights in step S1 adopts a pre-trained CLIP vision large model to extract the visual encoding features of the support image and the query image, and separate the encoded support mixed features using the real support image mask to extract the target foreground features of the support.

[0012] Even further, the text encoding features of the diagnostic report in step S2 are obtained by mapping the diagnostic report to the text feature space using a pre-trained CLIP language large model, and multi-modal feature interaction and alignment are established between the target foreground features of the support and the query mixed features using the text encoding features of the diagnostic report respectively. The multi-modal feature interaction and alignment of different branches include the following steps:

[0013] S21: Establish multi-modal interaction and alignment between the target foreground visual features of the support and the text encoding features of the diagnostic report. The calculation of the multi-modal feature interaction and alignment on the support branch is shown in formula (1).

[0014]

[0015] where, F s ct The multi-modal alignment features between vision and text in the support branch; F t represents the text encoding features of the diagnostic report; F s f represents the target foreground visual features of the support; d is the feature dimension, and T represents the transpose operation.

[0016] S22: Establish multi-modal interaction and alignment between the query mixed features and the text encoding features of the diagnostic report to promote the decoupling of the query mixed features and reduce the mutual interference of different semantics within the mixed features. The calculation of the multi-modal feature interaction and alignment on the query branch is shown in formula (2).

[0017]

[0018] where, Represents the multimodal alignment features between vision and text in the query branch.

[0019] Further, the alignment of the common semantics between branches in step S3 is to use the cross-attention mechanism to establish the visual feature alignment between the support target foreground features and the query hybrid features. The calculation of the common semantics alignment between branches is shown in formula (3).

[0020]

[0021] Further, the prototype generation in step S4 includes the steps:

[0022] S41: Generate the support target foreground prototype by using global average pooling on the support target foreground features. The calculation of the support target foreground prototype is shown in formula (4).

[0023]

[0024] Among them, P s f Represents the support target foreground prototype, h and w represent the length and width of the feature map, τ[·] represents the truth function, and M s (i) represents the mask of the target foreground, and c represents the target foreground class label value.

[0025] S42: Generate the target foreground prototype by using global average pooling on the multimodal feature interaction and alignment features of the support branch. The calculation of the multimodal interaction and alignment prototype on the support branch is shown in formula (5).

[0026]

[0027] S43: Generate the target foreground prototype by using global average pooling on the multimodal feature interaction and alignment features of the query branch. The calculation of the multimodal interaction and alignment prototype on the query branch is shown in formula (6).

[0028]

[0029] S44: Generate the target foreground prototype by using global average pooling on the common semantics alignment features between the two branches. The calculation of the common semantics alignment prototype is shown in formula (7).

[0030]

[0031] S45: Concatenate all foreground prototypes to obtain the final foreground prototype set And calculate the query hybrid feature F position by position q with each prototype p kThe similarity between ∈P, and each pixel is classified according to the maximum similarity value. The calculation of the maximum similarity is shown in formula (8).

[0032]

[0033] S46: Calculate the cross-entropy loss L between the predicted mask and the true query mask of the query image q , and use the cross-entropy loss to promote the optimization of the prototype set. The calculation of the cross-entropy loss is shown in formula (9).

[0034]

[0035] where H and W are the length and width of the query image, and M q is the true query mask, is the predicted query mask, and M q is only used in the training phase and is not visible in the test phase.

[0036] The beneficial effects of the present invention are that, compared with the prior art, the present invention not only reduces the dependence of the model on training data, but also improves the localization and segmentation performance of the model for newly emerged or unknown polyp regions, and also helps to improve the interpretability of the deep learning model; in addition, the small-sample intestinal polyp image semantic segmentation method disclosed in the present invention provides a new idea for the semantic segmentation of medical images of rare diseases and the semantic segmentation of medical images where the lesion region is similar to the surrounding tissue structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a schematic flowchart of a small-sample intestinal polyp image semantic segmentation method of the present invention;

[0038] Figure 2 is a schematic flowchart of multi-modal feature interaction and alignment in a small-sample intestinal polyp image semantic segmentation method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.

[0040] The present invention builds a meta - learning training network that supports a support branch and a query branch, and establishes the interaction and alignment of multi - modal features with the text encoding features of the diagnostic report and the visual features of a small number of intestinal polyp images. While promoting the decoupling of the target region and the background region of the polyp image to be segmented, it enhances the alignment ability between the features of the polyp lesion regions across branches.

[0041] Please refer to Figure 1 , a method for semantic segmentation of small - sample intestinal polyp images according to the present invention, includes the steps:

[0042] S1, Feature extraction: Use a backbone network with a set of shared weights to map the intestinal polyp images input to the support branch and the query branch into the deep feature space.

[0043] S2, Multi - modal feature alignment: Establish the interaction and alignment of multi - modal features between the text encoding features of the diagnostic report, the support target foreground features, and the query mixed features.

[0044] S3, General semantic alignment between branches: By establishing the cross - branch semantic interaction and alignment between the support target foreground features and the query mixed features, mine the general semantics between branches and improve the decoupling effect of the mixed features.

[0045] S4, Prototype generation: Capture the prototype set of classes on the separated support target foreground features, query decoupled features, and multi - modal alignment features of the support branch, and calculate the similarity between the query mixed features and each prototype in the prototype set position by position. Then, quickly locate and segment the polyp region in the query image according to the maximum similarity.

[0046] In step S1, the backbone network with a set of shared weights adopts a pre - trained CLIP visual large model to extract the visual encoding features of the support image and the query image, and separates the encoded support mixed features using the real support image mask to extract the support target foreground features.

[0047] In step S2, the text encoding features of the diagnostic report are obtained by using a pre - trained CLIP language large model to map the diagnostic report into the text feature space, and the interaction and alignment of multi - modal features between the support target foreground features and the query mixed features are established respectively using the text encoding features of the diagnostic report. The interaction and alignment of multi - modal features of different branches include the steps:

[0048] S21: Establish the interaction and alignment of multi - modal features between the support target foreground visual features and the text encoding features of the diagnostic report. The calculation of the multi - modal feature interaction and alignment on the support branch is shown in formula (1).

[0049]

[0050] S22: Establish multimodal interaction and alignment between the query hybrid features and the text encoding features of the diagnostic report, promote the decoupling of the query hybrid features, and reduce the mutual interference of different semantics within the hybrid features. The calculation of the multimodal feature interaction and alignment on the query branch is shown in Formula (2).

[0051]

[0052] Among them, F s f represents the visual features of the supporting target foreground, F t represents the text encoding features of the diagnostic report, d is the feature dimension, and T represents the transpose operation.

[0053] In step S3, the alignment of the general semantics between branches uses the cross-attention mechanism to establish the visual feature alignment between the supporting target foreground features and the query hybrid features. The calculation of the general semantic alignment between branches is shown in Formula (3).

[0054]

[0055] In step S4, the prototype generation includes the following steps:

[0056] S41: Generate the supporting target foreground prototype using global average pooling on the supporting target foreground features. The calculation of the supporting target foreground prototype is shown in Formula (4).

[0057]

[0058] S42: Generate the target foreground prototype using global average pooling on the multimodal feature interaction and alignment features of the support branch. The calculation of the multimodal interaction and alignment prototype on the support branch is shown in Formula (5).

[0059]

[0060] S43: Generate the target foreground prototype using global average pooling on the multimodal feature interaction and alignment features of the query branch. The calculation of the multimodal interaction and alignment prototype on the query branch is shown in Formula (6).

[0061]

[0062] S44: Generate the target foreground prototype using global average pooling on the general semantic alignment features between the two branches. The calculation of the general semantic alignment prototype is shown in Formula (7).

[0063]

[0064] Among them, h and w represent the length and width of the feature map, τ[·] represents the truth function, and M s (i) represents the mask of the target foreground, c represents the label value of the target foreground class;

[0065] S45: Concatenate all foreground prototypes to obtain the final foreground prototype set And calculate the query mixed feature F position by position q With each prototype p k ∈P, calculate the similarity between them, and classify each pixel according to the maximum similarity value. The calculation of the maximum similarity is shown in formula (8),

[0066]

[0067] S46: Calculate the cross-entropy loss L between the predicted mask and the true query mask of the query image q , and use the cross-entropy loss to promote the optimization of the prototype set. The calculation of the cross-entropy loss is shown in formula (9),

[0068]

[0069] Among them, H and W are the length and width of the query image, and M q Is the true query mask, Is the predicted query mask, and M q Is only used in the training phase and is not visible in the test phase.

[0070] The small-sample intestinal polyp image semantic segmentation method based on hybrid feature decoupling and multi-modal feature alignment technology proposed by the present invention reduces the mutual interference between different semantic features inside the hybrid feature through hybrid feature decoupling, and enhances the perception ability of the segmentation model for polyps and surrounding tissue structures. In addition, the present invention can further improve the robustness, reliability and generalization of the intestinal polyp image semantic segmentation model.

[0071] The above is the preferred implementation mode of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches can be made, and these improvements and retouches are also regarded as the protection scope of the present invention.

Claims

1. A semantic segmentation method for small-sample intestinal polyp images, characterized in that, Including the steps: S1, Feature extraction: Using a backbone network with a set of shared weights to map the intestinal polyp images input by the support branch and the query branch into the deep feature space; S2, Multimodal feature alignment: Establishing multimodal feature interaction and alignment between the text encoding features of the diagnostic report, the support target foreground features, and the query mixed features; S3, Cross-branch general semantic alignment: By establishing cross-branch semantic interaction and alignment between the support target foreground features and the query mixed features, mining the general semantics between branches, and enhancing the decoupling effect of the mixed features; S4, Prototype generation: Capturing the prototype set of the class on the separated support target foreground features, query decoupled features, and multimodal alignment features of the support branch, calculating the similarity between the query mixed features and each prototype in the prototype set position by position, and quickly locating and segmenting the polyp area in the query image according to the maximum similarity.

2. The small-sample intestinal polyp image semantic segmentation method according to claim 1, characterized in that The backbone network with a set of shared weights described in step S1 uses a pre-trained CLIP vision large model to extract the visual encoding features of the support image and the query image, and separates the encoded support mixed features using the real support image mask to extract the support target foreground features.

3. The small-sample intestinal polyp image semantic segmentation method according to claim 2, characterized in that The text encoding features of the diagnostic report described in step S2 are to map the diagnostic report into the text feature space using a pre-trained CLIP language large model, and establish multimodal feature interaction and alignment between the support target foreground features and the query mixed features respectively using the text encoding features of the diagnostic report. The multimodal feature interaction and alignment of different branches include the steps: S21: Establishing multimodal interaction and alignment between the support target foreground visual features and the text encoding features of the diagnostic report. The calculation of the multimodal feature interaction and alignment on the support branch is shown in formula (1). Among them, F s ct To support the multimodal alignment feature between vision and text in the branch, F t represents the text encoding feature of the diagnostic report, F f s Indicates support for the target foreground visual feature, d is the feature dimension, and T represents the transpose operation; S22: Establishing multimodal interaction and alignment between the query mixed features and the text encoding features of the diagnostic report to promote the decoupling of the query mixed features and reduce the mutual interference of different semantics inside the mixed features. The calculation of the multimodal feature interaction and alignment on the query branch is shown in formula (2).

4. The small-sample intestinal polyp image semantic segmentation method according to claim 1, wherein, The cross-branch general semantic alignment described in step S3 is to use the cross-attention mechanism to establish visual feature alignment between the support target foreground features and the query mixed features. The calculation of the cross-branch general semantic alignment is shown in formula (3).

5. The small-sample intestinal polyp image semantic segmentation method according to claim 1, characterized in that, The prototype generation described in step S4 includes the steps: S41: Generating the support target foreground prototype using global average pooling on the support target foreground features. The calculation of the support target foreground prototype is shown in formula (4). Among them, P s f represents supporting the target foreground prototype, h and w represent the length and width of the feature map, τ[·] represents the truth function, M s (i) represents the mask of the target foreground, and c represents the target foreground class label value; S42: Generating the target foreground prototype using global average pooling on the multimodal feature interaction and alignment features of the support branch. The calculation of the multimodal interaction and alignment prototype on the support branch is shown in formula (5). S43: Generating the target foreground prototype using global average pooling on the multimodal feature interaction and alignment features of the query branch. The calculation of the multimodal interaction and alignment prototype on the query branch is shown in formula (6). S44: Generate a target foreground prototype by using global average pooling on the general semantic alignment features between the two branches. The calculation of the general semantic alignment prototype is shown in Equation (7). where h and w represent the length and width of the feature map, τ[·] represents the truth function, M s (i) represents the mask of the target foreground, and c represents the label value of the target foreground class; S45: Stitch all foreground prototypes to obtain the final foreground prototype set And calculate the query hybrid feature F position by position q With each prototype p k ∈ P, and classify each pixel according to the maximum similarity value. The calculation of the maximum similarity is shown in formula (8). S46: Calculate the cross-entropy loss L between the predicted mask and the true query mask of the query image q , and use the cross-entropy loss to promote the optimization of the prototype set. The cross-entropy loss is calculated as shown in formula (9). where H and W are the length and width of the query image, and M q is the true query mask, is the predicted query mask, and M q is only used in the training phase and is not visible in the testing phase.