Multi-modal medical image segmentation method and system based on decoupling contrast learning, terminal and storage medium
Through the method based on decoupling and contrast learning, the feature alignment and decoupling of multimodal medical images are achieved, solving the problem that features cannot be effectively fused in multimodal imaging, and improving the accuracy and robustness of image segmentation.
Patent Information
- Application Number
- CN202510536228.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, multimodal imaging cannot effectively align the features of each modal, and cannot fully retain the unique information of each modal, affecting the image segmentation performance.
Using a method based on decoupling and contrast learning, the shared and unique features of white light images and narrowband images are extracted respectively through multi-scale distribution alignment and feature decoupling, and stitching and fusion are performed in a unified subspace to output the lesion mask image.
It improves the accuracy and robustness of medical image segmentation, makes full use of the complementary information of multimodal images, and achieves better image segmentation performance.
Smart Images

Figure CN120495312A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image segmentation technology, and in particular to a multimodal medical image segmentation method, system, terminal and computer-readable storage medium based on decoupled contrast learning. Background Art
[0002] Endoscopic imaging has high clarity, high resolution, and true color representation. Among them, white light imaging (WLI) and narrow band imaging (NBI) are two widely used endoscopic imaging modes. WLI can provide rich morphological information, capturing macroscopic features such as tissue texture, color, and boundary structure; while NBI enhances the visibility of superficial vascular structures by utilizing the absorption characteristics of hemoglobin in specific blue and green light bands, thereby more effectively detecting microvascular changes in the early stage of tumor transformation. Although these two imaging modes have their own advantages, it is usually difficult to fully capture the multidimensional characteristics of tumors by relying on a single image source. Although WLI can present tissue morphology in detail, its sensitivity to early vascular changes is limited; although NBI performs well in enhancing vascular patterns, it lacks structural context information and is more susceptible to noise and low contrast.
[0003] Therefore, single-modality methods often fail to meet the high requirements for segmentation accuracy and robustness in complex clinical scenarios. Multimodal imaging has become a cutting-edge direction for solving the above-mentioned problems by fusing the complementary information of different imaging modalities. Combining the macroscopic structural information provided by WLI with the fine microvascular features of NBI, multimodal methods are expected to achieve more comprehensive and accurate modeling of tumor characteristics. However, this fusion process also brings many technical challenges, such as significant differences in image distribution, resolution and acquisition conditions between modalities. Therefore, to achieve efficient fusion, it is necessary to effectively align the features of each modality and fully retain the unique information of each modality. This process is of considerable technical complexity.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide a multimodal medical image segmentation method, system, terminal and computer-readable storage medium based on decoupled contrastive learning, aiming to solve the problem in the prior art that multimodal imaging cannot effectively align the features of each modality and cannot fully retain the unique information of each modality, thereby affecting the image segmentation performance.
[0006] To achieve the above object, the present invention provides a multimodal medical image segmentation method based on decoupled contrastive learning, which comprises the following steps:
[0007] receiving a multimodal image pair, the multimodal image pair comprising a paired white light image and a narrowband image, processing the white light image through a plurality of coding blocks to obtain a final coded feature of the WLI modality, and processing the narrowband image through a plurality of coding blocks to obtain a final coded feature of the NBI modality;
[0008] The final encoded features of the WLI modality are processed by a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the WLI modality and a shared feature portion of the WLI modality, respectively. The final encoded features of the NBI modality are processed by a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the NBI modality and a shared feature portion of the NBI modality, respectively.
[0009] The shared feature parts of the WLI modality and the shared feature parts of the NBI modality are spliced along the feature dimension and projected into a unified subspace through linear mapping to obtain spliced shared features. The spliced shared features, the feature parts unique to the WLI modality, and the feature parts unique to the NBI modality are fused to obtain fused features. The fused features are predicted by a segmentation decoder to output a lesion mask image.
[0010] In addition, to achieve the above-mentioned object, the present invention further provides a multimodal medical image segmentation system based on decoupled contrastive learning, wherein the multimodal medical image segmentation system based on decoupled contrastive learning comprises:
[0011] a distribution alignment module configured to receive a multimodal image pair, the multimodal image pair comprising a paired white-light image and a narrow-band image, process the white-light image through a plurality of coding blocks to obtain a final coding feature of the WLI modality, and process the narrow-band image through a plurality of coding blocks to obtain a final coding feature of the NBI modality;
[0012] a feature decoupling module, configured to process the final encoded features of the WLI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the WLI modality and a shared feature portion of the WLI modality, and to process the final encoded features of the NBI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the NBI modality and a shared feature portion of the NBI modality, respectively;
[0013] A feature fusion module is used to splice the shared feature parts of the WLI modality and the shared feature parts of the NBI modality along the feature dimension, and project them into a unified subspace through linear mapping to obtain spliced shared features. The spliced shared features, the feature parts unique to the WLI modality, and the feature parts unique to the NBI modality are fused to obtain fused features. The fused features are predicted by a segmentation decoder to output a lesion mask image.
[0014] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a multimodal medical image segmentation program based on decoupled contrast learning stored in the memory and runnable on the processor, wherein the multimodal medical image segmentation program based on decoupled contrast learning, when executed by the processor, implements the steps of the multimodal medical image segmentation method based on decoupled contrast learning as described above.
[0015] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal medical image segmentation program based on decoupled contrast learning, and when the multimodal medical image segmentation program based on decoupled contrast learning is executed by a processor, the steps of the multimodal medical image segmentation method based on decoupled contrast learning as described above are implemented.
[0016] In the present invention, a multimodal image pair is received, wherein the multimodal image pair includes a paired white light image and a narrowband image, and the white light image is processed by multiple coding blocks to obtain the final coding features of the WLI modality, and the narrowband image is processed by multiple coding blocks to obtain the final coding features of the NBI modality; the final coding features of the WLI modality are processed by a unique feature encoder and a common feature encoder to obtain the unique feature part of the WLI modality and the shared feature part of the WLI modality respectively, and the final coding features of the NBI modality are processed by a unique feature encoder and a common feature encoder. After processing by the common feature encoder, the NBI modality-specific feature part and the NBI modality-shared feature part are obtained respectively; the shared feature part of the WLI modality and the shared feature part of the NBI modality are spliced along the feature dimension and projected into a unified subspace by linear mapping to obtain the spliced shared features, the spliced shared features, the WLI modality-specific feature part and the NBI modality-specific feature part are fused to obtain fused features, the fused features are predicted by the segmentation decoder, and the lesion mask image is output. The present invention performs distribution alignment between shallow features of multimodality at the feature level, improves the compatibility of cross-modal features, and further strengthens the distinction between modality-common features and modality-specific features based on contrastive learning of feature decoupling to model higher-order multimodal relationships, achieves better image segmentation performance, and improves the accuracy of medical image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flowchart of a preferred embodiment of the multimodal medical image segmentation method based on decoupled contrast learning of the present invention;
[0018] Figure 2 Schematic diagram of a preferred embodiment of the multimodal medical image segmentation method based on decoupled contrastive learning of the present invention, in which the overall framework follows the three-step strategy of "distribution alignment - feature decoupling - feature fusion";
[0019] Figure 3 1 is a structural diagram of a preferred embodiment of a multimodal medical image segmentation system based on decoupled contrastive learning according to the present invention;
[0020] Figure 4 FIG. 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0022] Medical image segmentation is one of the core technologies in computer-aided diagnosis and has made significant progress with the development of deep learning. Convolutional neural networks excel at extracting local spatial features, and typical architectures such as Unet have been widely used. To further improve segmentation performance, a large number of variants have been proposed, such as Attention-Unet and Unet++. In recent years, the Transformer architecture has gradually attracted attention for its ability to model long-range dependencies and capture global context. Representative works include Unetr and MedT. To combine the advantages of CNN and Transformer, hybrid models such as TransUne and SwinUnet have emerged. These methods combine convolutional encoders with Transformer modules to achieve richer feature representation and stronger generalization capabilities. These advances mark the gradual evolution of segmentation models from CNN-based architectures to Transformer-based designs, promoting the deep integration of the two.
[0023] Meanwhile, in the field of laryngopharyngeal tumor segmentation, most research still focuses on single-modality imaging. While these methods have demonstrated promising results to a certain extent, they lack explicit modeling of the differences and complementarities between modal features. Because WLI and NBI provide distinct and complementary diagnostic information, single-modality approaches often face challenges with accuracy and robustness in complex clinical scenarios. Early multimodal medical image segmentation methods primarily relied on simple fusion strategies, such as concatenating images from multiple modalities as input. However, these methods often fail to fully exploit the complementary information between modalities. Some research focuses on integrating complementary information between modalities to generate more informative images. SwinFusion fuses CNN and Swin Transformer architectures to simultaneously model long-range dependencies within and between modalities. MACTFusion proposes a lightweight cross-modal Transformer framework for adaptive fusion of local and global features. These methods share the common characteristic of compressing multimodal input into a single enhanced image. However, this compression often results in the loss of modality-specific cues and weakens complementary information, thus compromising subsequent segmentation accuracy.
[0024] In the field of endoscopic image segmentation, most existing research is limited to single-modality image analysis. The method proposed in this paper is applied to multi-modal endoscopic medical image segmentation for the first time, achieving efficient extraction and utilization of complementary information in multi-modal image pairs.
[0025] In paired endoscopic image pairs, the acquisition process of white light images and narrowband images has differences in spectral characteristics, lighting conditions, and acquisition settings. These differences lead to distribution mismatch and modality-specific deviations, which hinder effective joint learning. Simply fusing them directly can easily cause information confusion and mutual interference, thereby affecting segmentation performance. To address this problem, the present invention first aligns the distribution of shallow features of multimodal images at the feature level, improving the compatibility of cross-modal features.
[0026] After distribution alignment, semantic inconsistencies may still exist due to inherent information differences between cross-modal features. These differences typically manifest as redundant and complementary semantic features across modalities. To better utilize these semantic cues, this paper proposes a contrastive learning framework based on feature decoupling to further strengthen the distinction between modality-shared and modality-specific features, thereby modeling higher-order multimodal relationships and achieving superior image segmentation performance.
[0027] To address these issues, the present invention employs a multi-scale distribution alignment strategy in the early encoding phase to effectively align features from different modalities, thereby fully integrating the complementary strengths of WLI and NBI in terms of structural information and vascular characteristics. Furthermore, a multimodal contrastive learning method with feature decoupling capabilities is introduced, effectively distinguishing between shared and unique features between modalities. This allows for more refined semantic fusion, ensuring that the fused features possess both excellent robustness and good clinical interpretability.
[0028] The technical solution of the present invention is: receiving a multimodal image pair, wherein the multimodal image pair includes a paired white light image X w and narrowband imaging X n , the white light image X w After multiple encoding blocks, the final encoding feature F of the WLI modality is obtained w , the narrowband image X n After multiple coding blocks are processed, the final coding feature F of the NBI modality is obtained n ; The final encoding feature F of the WLI modality w After being processed by the unique feature encoder and the common feature encoder, the unique feature parts Z of the WLI modality are obtained respectively. w,p and the shared features of the WLI modality Z w,s , the final encoding feature F of the NBI modality n After being processed by the unique feature encoder and the common feature encoder, the NBI modality-specific feature parts Z and Z are obtained respectively. n,p and the shared feature part Z of the NBI modality n,s ; The shared feature part Z of the WLI modalityw,s and the shared feature part Z of the NBI modality n,s Splicing is performed along the feature dimension and projected into a unified subspace through linear mapping to obtain the spliced shared feature Z sh , the spliced shared feature Z sh , the characteristic part Z of the WLI mode w,p and the NBI modality-specific feature part Z n,p Perform fusion processing to obtain the fusion feature Z fused , the fusion feature Z fused After prediction processing by the segmentation decoder, the lesion mask image Y is output.
[0029] The multimodal medical image segmentation method based on decoupled contrast learning described in the preferred embodiment of the present invention is as follows: Figure 1 and Figure 2 As shown, the multimodal medical image segmentation method based on decoupled contrastive learning includes the following steps:
[0030] Step S10: Receive a multimodal image pair, where the multimodal image pair includes a paired white light image and a narrowband image, process the white light image through multiple coding blocks to obtain a final coding feature of the WLI modality, and process the narrowband image through multiple coding blocks to obtain a final coding feature of the NBI modality.
[0031] Specifically, the present invention uses the TransUnet encoder to extract multimodal features. The encoder consists of 12 Transformer encoding blocks, namely shallow encoding blocks and deep encoding blocks. The shallow encoding blocks are used to extract shallow encoding features, while the deep encoding blocks are used to extract deep encoding features. In deep learning networks, shallow encoding features mainly capture the statistical distribution properties of the data, while deep encoding features encode abstract high-level semantic information. Therefore, the present invention chooses to perform multi-scale distribution alignment of features in the shallow stage of the encoder to reduce significant statistical distribution differences.
[0032] Given an input multimodal image pair {X w ,X n}, that is, paired white light image and narrow band image, X w represents white light image, X n Indicates a narrowband image. If and They represent the feature map sequences extracted by the white light encoder and the narrowband encoder in the L shallow stages, respectively. The form of each feature map is Where B represents the batch size, N represents the number of image patches, and D represents the embedding dimension of each image patch. represents the field of real numbers;
[0033]
[0034] Among them, f w represents the features after WLI modal concatenation, f n Represents the features after NBI modality splicing, [;] represents splicing along the embedding dimension; feature f w , f n Covering coarse to fine semantic information, they can capture both global context and local details across different modalities.
[0035] In order to promote robust and stable distribution alignment across modalities, the present invention constructs a global feature representation that covers both statistical features and visual saliency. Based on this goal, two complementary components are defined in the global feature representation. The first part is used to capture the overall statistical distribution of the input features, and the second part is used to selectively emphasize spatially salient areas.
[0036] In order to model statistical features, the present invention applies a global average pooling operation on the multi-scale feature map in the spatial dimension:
[0037]
[0038] in, They are respectively used as global features of WLI and NBI features to capture statistical distribution information, and n represents a certain image block.
[0039] At the same time, in order to model spatial and semantic correlations in a global context, this paper constructs a complementary representation by introducing a weighted aggregation mechanism to highlight key information regions. Specifically, f w and f n This input is a function that calculates scores per image patch, generating raw importance scores:
[0040]
[0041] in, and Represents the independent image block level linear projection function applied to WLI and NBI features, respectively. and represent the WLI modality score and NBI modality score, respectively.
[0042] Will Normalized by the Softmax function, the corresponding attention weight is obtained:
[0043]
[0044] in, represents the importance score of the nth image block in the WLI modality, b represents the number of samples in a batch, represents the importance score corresponding to the rth image block in the WLI modality, Express The normalized value, represents the importance score of the nth image block in the NBI modality, represents the importance score corresponding to the r-th image block in the NBI modality, Express The normalized value r is used to represent the traversal because it is necessary to traverse all image blocks.
[0045] These weights guide the aggregation of spatial features by assigning higher importance to more informative image patches, thereby generating an attention-aware global representation:
[0046]
[0047] in, represents the features of the nth image block in the WLI modality, represents the weighted features of the WLI modality, represents the features of the nth image block in the NBI modality, represents the weighted features of the NBI modality, It is used to highlight semantically important regions, thus helping to build more discriminative global feature representations.
[0048] Finally, the present invention integrates statistical features with attention-weighted features to construct a complete global feature representation:
[0049]
[0050] in, represents the unified characteristics of the WLI modality, represents the unified features of NBI modalities, It combines global distribution structure and spatial saliency, is more semantically robust, and provides a solid foundation for subsequent cross-modal distribution alignment.
[0051] After obtaining the global feature representation, the present invention uses Maximum Mean Discrepancy (MMD) to align the feature distribution to alleviate the differences between modalities. MMD can effectively measure the distribution differences in high-dimensional space, thereby ensuring the consistency of features in terms of mean and high-order statistical information:
[0052]
[0053] Among them, L DA represents the loss, σ represents the bandwidth of the kernel function, represents the unified features of the WLI modality of the i-th batch, represents the unified features of the j-th batch of WLI modalities, represents the unified features of the p-th batch of NBI modalities, represents the unified features of the qth batch of NBI modalities.
[0054] By minimizing L DA This can achieve alignment between WLI and NBI features, thereby enhancing the cross-modal feature fusion effect and prompting the model to learn more robust feature representations.
[0055] After cross-modal distribution alignment, the final encoding feature F of the WLI modality is output w and the final encoded features F of the NBI modality n .
[0056] Step S20: After the final encoding features of the WLI modality are processed by a unique feature encoder and a common feature encoder, the feature part unique to the WLI modality and the shared feature part of the WLI modality are obtained respectively; after the final encoding features of the NBI modality are processed by a unique feature encoder and a common feature encoder, the feature part unique to the NBI modality and the shared feature part of the NBI modality are obtained respectively.
[0057] Specifically, in order to better utilize these semantic clues, the present invention further decouples the aligned features into a shared subspace and a modality-specific subspace, thereby explicitly modeling complementary information and modality-specific representations, thereby improving the effectiveness of feature fusion. w and F n Represent the final encoded features output by the encoders of WLI and NBI modalities respectively. In order to capture the shared features and unique features of each modality at the same time, the present invention introduces different encoders represents the common feature encoder of the WLI modality, The unique feature encoder representing the WLI modality, represents the common feature encoder of NBI modality, The NBI modality-specific feature encoder is used to extract the shared features and modality-specific features of the modality:
[0058]
[0059] Among them, Z w,s Represents the shared feature part of the WLI modality, Z w,p Represents the characteristic part unique to the WLI modality, Z n,s represents the shared feature part of the NBI modality, Zn,p Indicates the characteristic part unique to the NBI modality.
[0060] In order to promote semantic consistency between modalities, the present invention imposes a semantic constraint L Align , encourages similarity between shared features of WLI and NBI. From the perspective of cosine similarity, this is equivalent to minimizing the angle between shared feature vectors, which ideally should be close to 0 o :
[0061]
[0062] in, represents the shared feature part of the b-th sample WLI modality, represents the shared feature part of the b-th sample NBI modality.
[0063] In order to maximize the discrimination between the modality-specific features of WLI and NBI, a discrimination constraint L is imposed. Diff , minimizing the similarity between them and encouraging their directions in the embedding space to approach 180 o In this case, the cosine similarity tends to -1, indicating maximum semantic opposition. The corresponding constraints are defined as:
[0064]
[0065] in, represents the characteristic part of the WLI modality of the bth sample, Represents the characteristic part of the b-th sample NBI modality.
[0066] Furthermore, to decouple the shared features from the specific features in each modality, the present invention imposes an orthogonality constraint L on the shared components and modality-specific components of WLI and NBI. Orth , which geometrically corresponds to a 90° difference between the respective eigenvectors. o , ensuring that shared features and specific features capture different semantic subspaces. The corresponding constraint expression is:
[0067]
[0068] In order to further enhance the discriminability of decoupled feature representation, after the initial decoupling, the present invention further introduces a contrastive learning strategy based on multimodal decoupled features. w,s ,Z w,p ,Z n,s ,Z n,p}, for each sample b, the shared features of different modalities and should capture similar semantic information, and thus be considered as positive pairs and close their distance. Modality-specific features with both modalities and Distinguish them and treat them as negative sample pairs, thereby widening the distance between them. This design promotes the compactness of the shared feature space, while enhancing the separation between shared representations and specific representations, thereby achieving more thorough feature decoupling. To this end, the present invention proposes a disentangled-aware contrastive learning method (DACL), which can not only encourage the alignment of shared features across modalities, but also effectively widen the distance between multiple negative samples, including modality-specific features and irrelevant shared features:
[0069]
[0070] Among them, L DACL represents the contrastive learning constraint, represents modality-specific features, represents unrelated shared features, Represents cosine similarity, a can be represented b can represent or τ represents the temperature coefficient, which is used to control the concentration of the distribution. represents the shared feature part of the mth traversal NBI modality, represents the characteristic part of the mth traversal WLI modality, Represents the characteristic part of the mth traversal NBI modality.
[0071] By minimizing L DACL , which can bring the shared features between different modalities from the same sample closer together, while at the same time distance them from modality-specific features and irrelevant shared features, thereby promoting stronger feature decoupling. By combining all constraints with the contrastive learning strategy, the decoupling regularization term L FD The definition is as follows:
[0072] L FD =αL Align +βL Diff +γL Orth +δL DACL ; (17)
[0073] Among them, α, β, γ and δ all represent weights. In order to ensure numerical stability and avoid imbalance in the proportion of each loss term during training, the present invention introduces a weight adjustment mechanism for each loss function. Align , LDiff and L Orth The values of are limited to the interval [0, 1]. To ensure their equivalent contribution to the overall loss, the present invention uniformly sets their weights to In contrast, L DACL It often presents a larger positive value and is more difficult to converge during the optimization process. In order to suppress its dominant effect on the total loss function and maintain the stability and balance of the training process, the present invention assigns it a relatively small weight, namely δ = 0.01.
[0074] After decoupling the multimodal features, the final output decoupled feature set {Z w,s ,Z w,p ,Z n,s ,Z n,p}.
[0075] Step S30: splice the shared feature part of the WLI modality and the shared feature part of the NBI modality along the feature dimension, and project them into a unified subspace through linear mapping to obtain the spliced shared features; fuse the spliced shared features, the feature part unique to the WLI modality, and the feature part unique to the NBI modality to obtain fused features; and predict the fused features through a segmentation decoder to output a lesion mask image.
[0076] Specifically, after obtaining the decoupled feature set {Z w,s ,Z w,p ,Z n,s ,Z n,p}After that, the present invention first aligns the shared features of different modalities to enhance the representation of shared features. w,s and the shared feature part Z of the NBI modality n,s Splicing is performed along the feature dimension and projected into a unified subspace through linear mapping to obtain the spliced shared feature Z sh :
[0077] Z sh =f sh ([Z w,s ; Z n,s ]); (18)
[0078] in, Represents a linear mapping layer. Subsequently, the concatenated shared features The WLI modality-specific feature part Z w,p and the NBI modality-specific feature part Z n,p Perform fusion processing to obtain the fusion feature Z fused :
[0079] Z fused =f′ sh (Z sh )+f′ w (Z w,p )+f′ n (Z n,p ); (19)
[0080] in, f′ sh 、f′ w and f′ n Represents an independent nonlinear transformation acting on the feature dimension of each image block, with both input and output spaces being These transformations can further model the shared features and modality-specific features to obtain fused features. It integrates the complementary semantic information of the two modalities, thereby improving the overall representation ability. fused After prediction processing by the segmentation decoder D, the lesion mask image Y is output:
[0081] Y=D(Z fused );(20)
[0082] Furthermore, the multimodal medical image segmentation method based on decoupled contrastive learning of the present invention is applied to a multimodal two-dimensional medical image segmentation framework (i.e. Figure 2 The overall framework follows the three-step strategy of "distribution alignment-feature decoupling-feature fusion" to achieve robust segmentation of endoscopic lesion areas using white light imaging and narrow-band imaging). The multimodal two-dimensional medical image segmentation framework is trained using a total loss function (i.e., integrating multimodal distribution alignment, feature decoupling, and segmentation supervision into a unified overall training objective). The total loss function is defined as follows:
[0083] L total =L DA +L FD +L CE +L Dice ;(twenty one)
[0084] Among them, L DA It is used to align the distribution information between WLI and NBI, reduce the modality difference, and ensure consistent statistical properties across modalities in shallow features. FD Used to impose decoupling constraints, including alignment of shared features, differentiation of modality-specific features, and intra-modality orthogonality between shared and specific features. CE represents the cross entropy loss, L Dice represents the Dice loss, L CE and L DiceUsed to provide pixel-level supervision and guide the model to achieve accurate lesion area segmentation.
[0085] This paper proposes a multimodal 2D medical image segmentation framework for recognition tasks. This framework incorporates a multiscale distribution alignment strategy to effectively align feature representations from different modalities during the initial encoding phase, fully exploiting and integrating the complementary strengths of WLI and NBI images in terms of structural information and vascular features. Furthermore, the framework incorporates a multimodal contrastive learning method with feature decoupling capabilities, which explicitly distinguishes between shared and proprietary semantic features between modalities. This decoupling mechanism enables finer-grained semantic fusion, thereby improving the robustness and clinical interpretability of the fused features.
[0086] Compared with the prior art, the present invention can bring the following technical effects:
[0087] (1) Existing multimodal medical image fusion methods usually use simple feature splicing or attention mechanisms for fusion, which makes it difficult to take into account both modality consistency and individual differences. The three-stage multimodal fusion method of distribution alignment-feature decoupling-feature fusion proposed in this paper first distributes low-level features, then explicitly decouples high-level features from shared and specific features, and finally fuses the decoupled representations. This ensures the consistency and complementarity between modalities from a mechanism perspective, effectively improving the semantic expression ability of the fused representation.
[0088] (2) Most existing multimodal medical image fusion methods only perform feature alignment at specific scales or specific levels, making it difficult to fully capture the statistical differences between modalities at multiple scales. The multi-scale distribution alignment strategy proposed in this paper introduces statistical consistency constraints at multiple levels of the encoder, effectively alleviating the distribution mismatch problem between low-level perception and high-level semantics of different modalities, thereby improving the semantic consistency and discriminative ability of the fused features.
[0089] (3) Most current multimodal contrastive learning methods focus only on the aggregation of shared information between modalities, ignoring the structural differences between shared features and specific features, resulting in incomplete decoupling. This paper proposes a decoupling-aware multimodal contrastive learning method, designs a special loss function, and simultaneously enhances the aggregation of cross-modal shared features and the distinction between shared and specific features, effectively promoting the discriminative decoupling of features and improving the robustness and interpretability of the fused representation.
[0090] The feasibility of the solution of the present invention has been verified by theory and experiments. In experiments on three private multimodal endoscopic image datasets, the present invention has achieved the best segmentation effect at present. The segmentation performance in terms of Dice, IoU, Sensitivity, G-mean and other indicators exceeds the existing deep learning model based on this task.
[0091] Furthermore, the feature encoder used in the present invention can be replaced by a pre-trained basic model with stronger segmentation capability (such as an arbitrary segmentation model), replacing full parameter fine-tuning with efficient fine-tuning of partial parameters.
[0092] Further, if Figure 3 As shown, based on the above-mentioned multimodal medical image segmentation method based on decoupled contrastive learning, the present invention also provides a multimodal medical image segmentation system based on decoupled contrastive learning, wherein the multimodal medical image segmentation system based on decoupled contrastive learning includes:
[0093] a distribution alignment module 51 configured to receive a multimodal image pair, the multimodal image pair comprising a paired white light image and a narrowband image, process the white light image through multiple coding blocks to obtain a final coding feature of the WLI modality, and process the narrowband image through multiple coding blocks to obtain a final coding feature of the NBI modality;
[0094] a feature decoupling module 52 for processing the final encoded features of the WLI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the WLI modality and a shared feature portion of the WLI modality, and for processing the final encoded features of the NBI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the NBI modality and a shared feature portion of the NBI modality, respectively;
[0095] The feature fusion module 53 is used to splice the shared feature part of the WLI modality and the shared feature part of the NBI modality along the feature dimension, and project them into a unified subspace through linear mapping to obtain the spliced shared features, fuse the spliced shared features, the feature part unique to the WLI modality, and the feature part unique to the NBI modality to obtain fused features, and predict the fused features through a segmentation decoder to output a lesion mask image.
[0096] Further, if Figure 4 As shown, based on the above-mentioned multimodal medical image segmentation method and system based on decoupled contrast learning, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.
[0097] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a multimodal medical image segmentation program 40 based on decoupled contrast learning is stored on the memory 20, and the multimodal medical image segmentation program 40 based on decoupled contrast learning can be executed by the processor 10, thereby realizing the multimodal medical image segmentation method based on decoupled contrast learning in the present application.
[0098] In some embodiments, the processor 10 can be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the multimodal medical image segmentation method based on decoupled contrast learning.
[0099] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The processor 10, memory 20, and display 30 of the terminal communicate with each other via a system bus.
[0100] In one embodiment, when the processor 10 executes the multimodal medical image segmentation program 40 based on decoupled contrastive learning in the memory 20 , the steps of the multimodal medical image segmentation method based on decoupled contrastive learning are implemented.
[0101] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal medical image segmentation program based on decoupled contrast learning, and when the multimodal medical image segmentation program based on decoupled contrast learning is executed by a processor, the steps of the multimodal medical image segmentation method based on decoupled contrast learning as described above are implemented.
[0102] In summary, the present invention provides a multimodal medical image segmentation method, system, terminal, and computer-readable storage medium based on decoupled contrastive learning. The method includes: receiving a multimodal image pair, the multimodal image pair including a paired white-light image and a narrow-band image; processing the white-light image through multiple coding blocks to obtain a final coding feature of the WLI modality; processing the narrow-band image through multiple coding blocks to obtain a final coding feature of the NBI modality; processing the final coding feature of the WLI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the WLI modality and a shared feature portion of the WLI modality, respectively. After the final encoded features of the NBI modality are processed by a unique feature encoder and a common feature encoder, the unique feature part of the NBI modality and the shared feature part of the NBI modality are obtained respectively; the shared feature part of the WLI modality and the shared feature part of the NBI modality are spliced along the feature dimension and projected into a unified subspace by linear mapping to obtain the spliced shared features, the spliced shared features, the unique feature part of the WLI modality and the unique feature part of the NBI modality are fused to obtain fused features, the fused features are predicted by a segmentation decoder, and a lesion mask image is output. The present invention performs distribution alignment between shallow features of multimodality at the feature level, improves the compatibility of cross-modal features, and further strengthens the distinction between modality common features and modality-specific features based on contrastive learning of feature decoupling to model higher-order multimodal relationships, achieves better image segmentation performance, and improves the accuracy of medical image segmentation.
[0103] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.
[0104] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0105] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A multimodal medical image segmentation method based on decoupled contrastive learning, characterized in that: The multimodal medical image segmentation method based on decoupled contrast learning includes: receiving a multimodal image pair, the multimodal image pair comprising a paired white light image and a narrowband image, processing the white light image through a plurality of coding blocks to obtain a final coded feature of the WLI modality, and processing the narrowband image through a plurality of coding blocks to obtain a final coded feature of the NBI modality; The final encoded features of the WLI modality are processed by a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the WLI modality and a shared feature portion of the WLI modality, respectively. The final encoded features of the NBI modality are processed by a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the NBI modality and a shared feature portion of the NBI modality, respectively. The shared feature parts of the WLI modality and the shared feature parts of the NBI modality are spliced along the feature dimension and projected into a unified subspace through linear mapping to obtain spliced shared features. The spliced shared features, the feature parts unique to the WLI modality, and the feature parts unique to the NBI modality are fused to obtain fused features. The fused features are predicted by a segmentation decoder to output a lesion mask image.
2. The multimodal medical image segmentation method based on decoupled contrastive learning according to claim 1, characterized in that: The receiving of a multimodal image pair, wherein the multimodal image pair includes a paired white light image and a narrowband image, processing the white light image through a plurality of coding blocks to obtain a final coding feature of the WLI modality, and processing the narrowband image through a plurality of coding blocks to obtain a final coding feature of the NBI modality, specifically includes: Given an input multimodal image pair {X w ,X n }, X w represents white light image, X n Indicates a narrowband image. If and They represent the feature map sequences extracted by the white light encoder and the narrowband encoder in the L shallow stages, respectively. The form of each feature map is Where B represents the batch size, N represents the number of image patches, and D represents the embedding dimension of each image patch. represents the field of real numbers; Among them, f w represents the features after WLI modal concatenation, f n Represents the features after NBI modality splicing, [;] represents splicing along the embedding dimension; Construct a global feature representation that covers both statistical features and visual saliency. Define two complementary components in the global feature representation: the first component is used to capture the overall statistical distribution of input features, and the second component is used to selectively emphasize spatially salient regions. Apply global average pooling operation on the multi-scale feature map in the spatial dimension: in, As global features of WLI and NBI features, they are used to capture statistical distribution information, and n represents a certain image block; Constructing complementary representations, by introducing a weighted aggregation mechanism to highlight key information areas, w and f n This input is a function that calculates scores per image patch, generating raw importance scores: in, and Represents the independent image block level linear projection function applied to WLI and NBI features, respectively. and represent the WLI modality score and NBI modality score, respectively; Will Normalized by the Softmax function, the corresponding attention weight is obtained: in, represents the importance score of the nth image block in the WLI modality, b represents the number of samples in a batch, represents the importance score corresponding to the rth image block in the WLI modality, Express The normalized value, represents the importance score of the nth image block in the NBI modality, represents the importance score corresponding to the r-th image block in the NBI modality, Express Normalized value; Generating attention-aware global representations: in, represents the features of the nth image block in the WLI modality, represents the weighted features of the WLI modality, represents the features of the nth image block in the NBI modality, represents the weighted features of the NBI modality, Used to highlight semantically important regions, which helps to build more discriminative global feature representations; Integrate statistical features with attention-weighted features to construct a complete global feature representation: in, represents the unified characteristics of the WLI modality, Represents the unified features of NBI modalities; After obtaining the global feature representation, the maximum mean difference is used to align the feature distribution to alleviate the differences between modalities: Among them, L DA represents the loss, σ represents the bandwidth of the kernel function, represents the unified features of the WLI modality of the i-th batch, represents the unified features of the j-th batch of WLI modalities, represents the unified features of the p-th batch of NBI modalities, represents the unified features of the qth batch of NBI modalities; By minimizing L DA To achieve alignment between WLI and NBI features, output the final encoding feature F of the WLI modality w and the final encoded features F of the NBI modality n .
3. The multimodal medical image segmentation method based on decoupled contrastive learning according to claim 2, characterized in that: The coding block includes a shallow coding block and a deep coding block; The shallow coding block is used to extract shallow coding features, and the deep coding block is used to extract deep coding features; The shallow coding features are used to capture the statistical distribution properties of the data, and the deep coding features are used to encode abstract high-level semantic information.
4. The multimodal medical image segmentation method based on decoupled contrastive learning according to claim 2, characterized in that: The method of processing the final encoded features of the WLI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the WLI modality and a shared feature portion of the WLI modality, and processing the final encoded features of the NBI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the NBI modality and a shared feature portion of the NBI modality, specifically includes: Introducing different encoders represents the common feature encoder of the WLI modality, The unique feature encoder representing the WLI modality, represents the common feature encoder of NBI modality, The NBI modality-specific feature encoder is used to extract the shared features and modality-specific features of the modality: Among them, Z w,s Represents the shared feature part of the WLI modality, Z w,p Represents the characteristic part unique to the WLI modality, Z n,s represents the shared feature part of the NBI modality, Z n,p It represents the characteristic part unique to the NBI modality; Based on the semantic consistency between modalities, a semantic constraint L is imposed. Align : in, represents the shared feature part of the b-th sample WLI modality, Represents the shared feature part of the b-th sample NBI modality; Based on maximizing the discrimination between the modality-specific features of WLI and NBI, a discrimination constraint L is imposed. Diff : in, represents the characteristic part of the WLI modality of the bth sample, Represents the characteristic part of the b-th sample NBI modality; Based on decoupling the shared features and specific features in each modality, an orthogonality constraint L is imposed on the shared components and modality-specific components of WLI and NBI. Orth : A multimodal contrastive learning method based on decoupled perception is used to align shared features across modalities and increase the distance between multiple negative samples, including modality-specific features and irrelevant shared features: Among them, L DACL represents the contrastive learning constraint, represents modality-specific features, represents unrelated shared features, Represents cosine similarity, τ represents the temperature coefficient, which is used to control the concentration of the distribution. represents the shared feature part of the mth traversal NBI modality, represents the characteristic part of the mth traversal WLI modality, represents the characteristic part of the mth traversal NBI modality; Combine all constraints with contrastive learning strategy to decouple the regularization term L FD The definition is as follows: L FD =αL Align +βL Diff +γL Orth +δL DACL ;(17) Among them, α, β, γ and δ all represent weights. δ = 0.01; Output decoupling feature set {Z w,s ,Z w,p ,Z n,s ,Z n,p }.
5. The multimodal medical image segmentation method based on decoupled contrastive learning according to claim 4, characterized in that: The semantic constraint is used to promote semantic consistency between modalities, the discriminability constraint is used to maximize the discriminability between modality-specific features of WLI and NBI, the orthogonality constraint is used to decouple shared features and specific features in each modality, and the contrastive learning constraint is used to bring together shared features between different modalities from the same sample and to increase the distance between shared features and modality-specific features as well as irrelevant shared features to promote stronger feature decoupling.
6. The multimodal medical image segmentation method based on decoupled contrastive learning according to claim 4, characterized in that: The shared feature part of the WLI modality and the shared feature part of the NBI modality are spliced along the feature dimension and projected into a unified subspace through linear mapping to obtain spliced shared features, the spliced shared features, the feature part unique to the WLI modality, and the feature part unique to the NBI modality are fused to obtain fused features, the fused features are predicted by a segmentation decoder, and a lesion mask image is output, specifically comprising: Get the decoupled feature set {Z w,s ,Z w,p ,Z n,s ,Z n,p }After that, the shared features of different modalities are aligned to enhance the representation of shared features; The shared feature part Z of the WLI modality w,s and the shared feature part Z of the NBI modality n,s Splicing is performed along the feature dimension and projected into a unified subspace through linear mapping to obtain the spliced shared feature Z sh : WITH sh =f sh ([WITH w,s ;WITH n,s ]);(18) in, Represents a linear mapping layer; The shared feature Z after splicing sh , the characteristic part Z of the WLI mode w,p and the NBI modality-specific feature part Z n,p Perform fusion processing to obtain the fusion feature Z fused : Z fused =f′ sh (Z sh )+f′ w (Z w,p )+f′ n (Z n,p ); (19) in, f′ sh 、f′ w and f′ n Represents an independent nonlinear transformation acting on the feature dimension of each image block, with both input and output spaces being The fusion feature Z fused After prediction processing by the segmentation decoder D, the lesion mask image is output: Y=D(Z fused );(20) Where Y represents the lesion mask image.
7. The multimodal medical image segmentation method based on decoupled contrastive learning according to claim 6, characterized in that: The multimodal medical image segmentation method based on decoupled contrastive learning is applied to a multimodal two-dimensional medical image segmentation framework. The multimodal two-dimensional medical image segmentation framework is trained using a total loss function, which is defined as follows: L total =L DA +L FD +L CE +L Dice ;(21) Among them, L DA Used to align the distribution information between WLI and NBI, L FD Used to impose decoupling constraints, L CE represents the cross entropy loss, L Dice represents the Dice loss, L CE and L Dice Used to provide pixel-level supervision and guide the model to achieve accurate lesion area segmentation.
8. A multimodal medical image segmentation system based on decoupled contrastive learning, characterized in that: The multimodal medical image segmentation system based on decoupled contrastive learning includes: a distribution alignment module configured to receive a multimodal image pair, the multimodal image pair comprising a paired white-light image and a narrow-band image, process the white-light image through a plurality of coding blocks to obtain a final coding feature of the WLI modality, and process the narrow-band image through a plurality of coding blocks to obtain a final coding feature of the NBI modality; a feature decoupling module, configured to process the final encoded features of the WLI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the WLI modality and a shared feature portion of the WLI modality, and to process the final encoded features of the NBI modality through a unique feature encoder and a shared feature encoder to obtain a unique feature portion of the NBI modality and a shared feature portion of the NBI modality, respectively; A feature fusion module is used to splice the shared feature parts of the WLI modality and the shared feature parts of the NBI modality along the feature dimension, and project them into a unified subspace through linear mapping to obtain spliced shared features. The spliced shared features, the feature parts unique to the WLI modality, and the feature parts unique to the NBI modality are fused to obtain fused features. The fused features are predicted by a segmentation decoder to output a lesion mask image.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a multimodal medical image segmentation program based on decoupled contrast learning stored in the memory and executable on the processor. When the multimodal medical image segmentation program based on decoupled contrast learning is executed by the processor, the steps of the multimodal medical image segmentation method based on decoupled contrast learning as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a multimodal medical image segmentation program based on decoupled contrast learning. When the multimodal medical image segmentation program based on decoupled contrast learning is executed by a processor, the steps of the multimodal medical image segmentation method based on decoupled contrast learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Image data compression method and device and storage medium
CN121357335A