A semi-supervised feature fusion segmentation system based on OCT-OCTA multi-modal
The OCT-OCTA multimodal semi-supervised feature fusion segmentation system solves the problems of insufficient data and high computational resources in retinal vessel segmentation, achieving efficient and accurate vessel segmentation results, especially with significant advantages when the amount of data is small.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2025-10-13
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies for retinal vessel segmentation suffer from problems such as insufficient data information, high computational resource requirements for models, long training time, and inconsistent and inaccurate segmentation results, especially when data acquisition is difficult under supervised learning, resulting in poor performance.
A semi-supervised feature fusion and segmentation system based on OCT-OCTA multimodal approaches is adopted. The geometric shape and contextual features of retinal vessels are extracted by a pre-trained large model with an embedded adapter. The system combines intramodal information enhancement and intermodal wavelet extraction fusion modules, utilizes cross-attention mechanism for feature fusion, and improves model robustness through contrastive consistency learning.
It improves the accuracy and connectivity of retinal vessel segmentation, reduces the probability of incorrect segmentation, and enhances segmentation precision and visual experience, especially with a small amount of labeled data.
Smart Images

Figure CN121305066B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information technology, and in particular relates to a semi-supervised feature fusion and segmentation system based on OCT-OCTA multimodal. Background Technology
[0002] Fundus photography provides a non-invasive way to examine the human eye. The acquired fundus images can reveal various normal fundus structures or eye damage caused by diseases, such as retinal vascular lesions and diabetic retinopathy (DR). Fundus images provide important reference information for ophthalmologists and are also a crucial method for the early diagnosis of cardiovascular diseases like DR. In recent years, deep learning technology, characterized by high-performance computing, has made significant progress in image processing and has become a hot research topic in medical image analysis. Retinal vascular structures are complex and variable. Traditional manual segmentation methods are not only labor-intensive and time-consuming but also easily influenced by the subjective factors of doctors, leading to inconsistent and inaccurate segmentation results. Deep learning technology can automatically learn features in images, achieving high-precision segmentation of retinal vessels through extensive training data. Furthermore, deep learning models typically have faster inference speeds, significantly improving segmentation efficiency. This section categorizes various methods for retinal vessel segmentation based on their network architecture: convolutional neural networks, fully convolutional networks, and U-Net.
[0003] Convolutional Neural Networks (CNNs), as a deep learning technique, initially showed broad application potential in digital image classification, encompassing plant species identification, plant disease diagnosis, and retinal vessel detection. The core of a CNN architecture includes a feature learning module and a classifier component. The former is responsible for automatically extracting visual features from digital images, while the latter uses these features to determine whether objects in the image belong to a specific category. Maji et al. proposed an ensemble CNN method to process fundus color images for accurate retinal vessel segmentation. This study employs an ensemble strategy to reduce the risk of overfitting and uses the DRIVE database as a benchmark dataset to validate the performance and effectiveness of the proposed method. Maninis et al. proposed a unified retinal image analysis framework that provides retinal vessel and optic disc segmentation. This framework, based on CNN, is called Deep Retinal Image Understanding (DRIU). DRIU uses a basic network architecture on which two sets of specialized layers are trained to solve retinal vessel and optic disc segmentation. Its effectiveness was validated on the DRIVE and STARE databases. Fu et al. designed a method specifically for retinal vessel segmentation to address the boundary detection problem. This method utilizes a convolutional neural network (CNN) architecture to learn discriminative features that can more effectively characterize important latent patterns related to blood vessels and background in fundus images. To generate high-quality blood vessel segmentation results, the researchers further employed a fully connected conditional random field (CRF) for segmentation processing, thereby outputting a high-precision blood vessel probability map and ultimately achieving binarized segmentation. Ngo and Han et al. proposed a novel maximum entropy technique to improve the generalization ability of the training process for predicting blood vessels from retinal fundus images. The authors developed a multi-level CNN model for automatic blood vessel segmentation. This network has two input branches with different input image resolutions. In the second branch, a maximization technique is used to reduce the resolution of the input image, improving the generalization ability of the training process. The advantage of CNNs lies in their wide applicability and good generalization performance. However, the main limitation of these models is their high dependence on large-scale image datasets during the training phase, which not only leads to large memory consumption and long training time, but also comes with stringent requirements for computing resources such as CPUs and GPUs.
[0004] To overcome the inherent limitations of Convolutional Neural Network (CNN) architectures, Long et al. pioneered the Fully Convolutional Network (FCN) in 2014. The core of the FCN architecture lies in abandoning traditional fully connected layers and instead constructing entirely from convolutional layers. This innovative design allows FCN to generate segmented images from input images of arbitrary sizes, greatly improving the model's adaptability and flexibility under different imaging conditions. Furthermore, FCN uses manually labeled ground truth values as supervision signals to achieve pixel-level accurate predictions of targets. This strategy not only deepens image understanding capabilities but also successfully extends image-level classification tasks to more refined pixel-level classification, thus driving significant progress in image segmentation technology. Dasgupta and Singh designed a fully convolutional neural network architecture for blood vessel segmentation tasks. The authors reconstructed the blood vessel segmentation task as a multi-label inference problem, which is optimized through learning using a joint loss function. The core contribution of this research lies in revealing the advantages of the joint optimization strategy, namely, combining a fully convolutional neural network framework with a structured prediction method to achieve accurate segmentation of retinal vessels in fundus images. Lyu et al. proposed a novel retinal vessel segmentation method that aims to capture the strongest spatial correlations between vessels by introducing separable spatial and channel flows and a dense neighbor vessel prediction strategy. Geometric transformation techniques and overlapping patching strategies are employed during the training and prediction phases to augment and expand the dataset. Khan et al. proposed a fully convolutional network (FCN) architecture based on an encoder-decoder structure specifically for vessel segmentation, named VessSeg. VessSeg consists of 13 convolutional layers configured similarly to the VGG16 network, coupled with a corresponding decoder network. In the decoder network, the final decoder output is passed to two softmax classifiers for pixel-by-pixel classification. The output of the softmax classifier represents the probability value of each pixel belonging to the vessel or non-vessel class.
[0005] In the field of biomedical image segmentation, U-Net is one of the most prominent deep learning architectures recently. The U-Net architecture consists of two core components: an encoder and a decoder. The encoder is responsible for extracting spatial features from the image, while the decoder uses these encoded features to construct segmentation feature maps. The U-Net architecture integrates traditional convolutional layers and upsampling techniques, aiming to learn and integrate local details and global structural features of the image. Yang et al. proposed an improved U-Net network as the foundation network for a multi-task segmentation framework, aiming to improve the accuracy of blood vessel segmentation in fundus images. This method can simultaneously extract features from both thick and thin blood vessels, thereby strengthening the correlation between these two types of blood vessels under different segmentation networks. An effective loss function was developed to adapt to two different blood vessel segmentation tasks, solving the problem of an imbalance between the proportion of thick and thin blood vessels. Dong et al. proposed a cascaded residual attention U-Net network called CRAUNet, designing a multi-scale fusion channel attention (MFCA) module for implementing coarse-to-fine retinal vessel segmentation analysis. Furthermore, a DropBlock regularization method was proposed to reduce overfitting. Zhang et al. proposed a deep network architecture called Bridge-net, designed to effectively utilize contextual information from retinal vessels. This bridging network combines recurrent neural networks (RNNs) with convolutional neural networks (CNNs) to provide contextual information and generate a probabilistic map of retinal vessels. Furthermore, the authors introduced a block-based loss weight mapping method to address the imbalance problem present in the image. Summary of the Invention
[0006] To address the issue of limited information in single-modality retinal and fundus vascular data, this invention provides a semi-supervised feature fusion segmentation system based on OCT-OCTA multimodality. This system fully integrates edge and contextual information from both OCT and OCT-OCTA modalities, effectively reducing the probability of predicting erroneous vascular pixels and improving the accuracy and connectivity of vascular segmentation.
[0007] A semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal features includes a preprocessing module, an encoder module, a cross-modal feature alignment and fusion module, and a contrast consistency learning module. The cross-modal feature alignment and fusion module includes an intra-modal OCT information enhancement unit, an intra-modal OCTA information enhancement unit, and an inter-modal wavelet extraction and fusion unit.
[0008] The preprocessing module is used to perform data enhancement on the OCT image and the OCT A image respectively, so as to obtain the OCT enhanced image and the OCT A enhanced image.
[0009] The encoder module is used to extract the geometric shape features of the OCT-enhanced image and the OCTA-enhanced image respectively, and obtain the corresponding OCT feature map and OCTA feature map;
[0010] The intramodal OCT information enhancement unit is used to perform channel and spatial dimension weighted enhancement on the OCT feature map, and then the resulting enhanced OCT feature map is concatenated with the original OCT feature map to obtain the first concatenated feature map. ;
[0011] The intramodal OCTA information enhancement unit is used to perform channel and spatial dimension weighted enhancement on the OCTA feature map, and then the resulting enhanced OCTA feature map is concatenated with the original OCTA feature map to obtain a second concatenated feature map. ;
[0012] The intermodal wavelet extraction and fusion unit is used to extract the spatial and frequency domain information of the OCT feature map and the OCTA feature map using Haar wavelet transform, and to fuse the spatial and frequency domain information corresponding to the two feature maps through a cross-attention mechanism. Then, the fused feature map is concatenated with the OCT enhanced feature map and the OCTA enhanced feature map to obtain the third concatenated feature map. ;
[0013] The contrastive consistency learning module is used for... , , Regularization is then applied to obtain the final segmentation result.
[0014] Furthermore, the encoder module extracts the geometric features of any enhanced image to obtain any feature map using the following method:
[0015] The current input enhanced image is downsampled and flattened into a two-dimensional tensor by convolution calculation with a specific stride.
[0016] Generate a corresponding position encoding vector for the three-dimensional position information of each pixel in the current input enhanced image, and then add it to the corresponding two-dimensional tensor to obtain a two-dimensional encoding tensor with embedded position information;
[0017] The two-dimensional encoded tensor is input into a Transformer unit composed of multiple cascaded Transformer blocks for feature extraction, resulting in a feature map corresponding to the current input enhanced image.
[0018] Furthermore, the method for feature extraction of the input tensor received by any Transformer block is as follows:
[0019] The input tensor is subjected to layer normalization to obtain a normalized tensor;
[0020] A multi-head self-attention mechanism is used to extract semantic space information of different dimensions of the normalized tensor. Then, the semantic space information of different dimensions is added to the normalized tensor to obtain the original self-attention tensor.
[0021] The original self-attention tensor is subjected to layer normalization to obtain a normalized self-attention tensor.
[0022] Linear processing is performed on the normalized self-attention tensor to obtain the linear self-attention tensor;
[0023] Add the linear self-attention tensor to the original self-attention tensor to obtain the fused self-attention tensor;
[0024] The dimensionality of the fused self-attention tensor is reduced to obtain the dimensionality-reduced self-attention tensor.
[0025] An adapter is used to enhance the features of the dimension-reduced self-attention tensor, and the enhancement result is the final output.
[0026] Furthermore, the adapter's method for feature enhancement of the dimensionality-reduced self-attention tensor is as follows:
[0027] First, the dimension-reduced self-attention tensor is reshaped into a three-dimensional image;
[0028] The 3D image undergoes three consecutive adaptation operations, and the output of the last adaptation operation is reshaped into a 2D image. The 2D image is the final enhanced output. The adaptation operation specifically involves:
[0029] Use 3 A 3-dimensional convolution kernel performs a convolution operation on the current input 3D image;
[0030] The application layer normalization operation adjusts the input of each layer of the 3D image after the convolution operation;
[0031] The nonlinear expressive power of the 3D image after layer normalization is increased by using activation functions;
[0032] The final output of the adaptation operation is obtained by superimposing the 3D image before the convolution operation and the 3D image after the nonlinear representation is enhanced by adding the residuals.
[0033] Furthermore, only the network parameters in the adapter are used as adjustable parameters to be trained, while the network parameters of the other modules remain frozen during training. The total loss function Loss used when training the adapter is as follows:
[0034] Loss=L sup1 +L sup2 +L mut1 +L mut2
[0035] Among them, Lsup1 L is a loss function constructed from the predicted segmentation results and the segmentation label values obtained from labeled OCT images. sup2 L is a loss function constructed from the predicted segmentation values and segmentation label values obtained from labeled OCTA images. mut1 The predicted value of the segmentation result obtained from the unlabeled OCT image is compared with the value obtained from the segmentation result obtained from the unlabeled OCT image. The loss function L is constructed from the predicted values of the obtained segmentation results. mut2 The predicted values of the segmentation results obtained from the unlabeled OCTA image are compared with those obtained from the... The loss function is constructed from the predicted values of the obtained segmentation results.
[0036] Furthermore, L sup1 L sup2 L mut1 L mut2 The calculation method is as follows:
[0037]
[0038]
[0039]
[0040]
[0041] Where CE represents cross-entropy loss, Dice represents Dice loss, Y1 represents the predicted segmentation result obtained from the labeled OCT image, Y represents the segmentation result label value determined jointly by the labeled OCT image and the OCTA image, Y2 represents the predicted segmentation result obtained from the labeled OCTA image, and Y3 represents the predicted segmentation result obtained from the labeled OCTA image. The obtained segmentation prediction values are Y1' and Y2', respectively, which represent the segmentation prediction values obtained from the unlabeled OCT image and the unlabeled OCT image.
[0042] Furthermore, the method for obtaining the fused feature map by the intermodal wavelet extraction fusion unit is as follows:
[0043] The Haar wavelet transform is used to transform the OCT feature map and the OCTA feature map into the following components: OCT low frequency component, OCT horizontal high frequency component, OCT vertical high frequency component, OCT diagonal high frequency component, OCTA low frequency component, OCTA horizontal high frequency component, OCTA vertical high frequency component, and OCTA diagonal high frequency component, respectively.
[0044] The low-frequency components of OCT are used as global features representing the OCT feature map; the horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components of OCT are concatenated, and then 1... 1. Convolutional low-dimensional mapping yields OCT high-frequency features that characterize the local detail features of the OCT feature map;
[0045] The low-frequency components of OCTA are used as the global features of the OCTA feature map; the horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components of OCTA are concatenated, and then 1... 1. Convolutional low-dimensional mapping yields high-frequency OCTA features that characterize the local detail features of the OCTA feature map;
[0046] A multi-head attention mechanism is used to extract and fuse semantic space information from different dimensions of OCT high-frequency features and OCTA high-frequency features, respectively, to obtain the first fused image of OCT and the first fused image of OCTA.
[0047] A multi-head attention mechanism is used to extract and fuse semantic space information from different dimensions of the OCT feature map and the OCT feature map, respectively, to obtain the OCT second fused image and the OCT second fused image.
[0048] A cross-fusion attention mechanism is used to combine OCT low-frequency features, OCTA low-frequency features, the first fused OCT image, the first fused OCTA image, the second fused OCT image, and the second fused OCTA image. Then, the combined image is processed through a 3D model. 3. Convolution, batch normalization, and activation functions are used to effectively fuse features to obtain the final fused feature map.
[0049] Furthermore, both the intramodal OCT information enhancement unit and the intramodal OCT information enhancement unit employ convolutional attention modules to perform channel and spatial dimension weighted enhancement on the OCT feature maps and OCT feature maps.
[0050] Beneficial effects:
[0051] 1. This invention provides a semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal data, aiming to provide the model with prior features of different modalities of retinal fundus vascular data by using retinal fundus vascular data of different modalities. This invention integrates the shape features of retinal vessels and global context into the segmentation network by embedding a pre-trained large model with an adapter, thereby enhancing its adaptability to vascular segmentation tasks. The segmentation network includes an intramodal information enhancement module and an intermodal wavelet extraction and fusion module. The former is used to independently fuse weighted information from each modality, while the latter extracts high and low frequency components through wavelet transform, captures details and global contours, and uses cross-attention to mix intermodal features. In addition, a contrast consistency learning module establishes regularization between different perturbation outputs, further improving the robustness and segmentation performance of the model.
[0052] 2. This invention provides a semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal approaches. It fully integrates edge and contextual information from both OCT and OCT-OCTA modalities. While accurately capturing tiny pixels at the intersections of small blood vessels, it not only preserves overall global information, resulting in better continuity of blood vessel segmentation, but also effectively reduces the probability of incorrect segmentation while preserving pixels at the bends of blood vessels to a large extent. This significantly improves the accuracy and visual experience of segmentation, providing a favorable guarantee for the subsequent realization of multimodal semi-supervised blood vessel segmentation in fundus images.
[0053] 3. This invention provides a semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal model. It is the first solution to apply semi-supervised learning and multimodal fusion to the field of retinal vessel segmentation. Experiments on the OCT-500 dataset with different annotation ratios show that this invention has significant advantages over other segmentation system models and achieves higher semantic segmentation accuracy with a smaller amount of data. Attached Figure Description
[0054] Figure 1 A schematic diagram of the structure of the semi-supervised feature fusion and segmentation system based on OCT-OCTA multimodal provided by the present invention;
[0055] Figure 2 This is a schematic diagram of the structure of the fine-tuning EfficientSAM encoder module provided by the present invention;
[0056] Figure 3 This is a schematic diagram of the wavelet extraction and fusion module provided by the present invention;
[0057] Figure 4 This is a schematic diagram of the Haar wavelet transform provided by the present invention;
[0058] Figure 5This is a schematic diagram comparing the segmentation results of different segmentation networks trained using 20% OCT labeled data, as provided by the present invention.
[0059] Figure 6 This is a schematic diagram comparing the segmentation results of different segmentation networks trained using 20% OCTA labeled data, as provided by the present invention. Detailed Implementation
[0060] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0061] This invention aims to achieve high-quality segmentation of retinal vessel images, producing high-precision segmentation results. However, the geometry of retinal vessels is complex and tortuous, requiring a lengthy process for specialized physicians to obtain the gold standard, resulting in high annotation costs and significant difficulties in acquiring large amounts of paired training data. Furthermore, the high privacy requirements of medical images limit data acquisition and dissemination, leading to poor image segmentation performance of supervised learning with limited labeled data. With the rapid development of non-invasive fundus imaging technologies (OCT, OCTTA), combining the unique salient features of different modalities to improve network segmentation performance is a major direction for the development of retinal vessel segmentation.
[0062] Existing work utilizes deep learning methods based on semi-supervised training to segment retinal vessels, but current technology lacks an effective network that combines the semi-supervised paradigm with multimodal data fusion enhancement. To fully leverage retinal fundus vessel data from different modalities and provide more supplementary information for semi-supervised methods, this invention proposes a Cross-Modality Feature Fusion Network (CMNet). This network rationally utilizes prior information from different modalities, significantly enhancing the model's ability to recognize target features and providing a more comprehensive reference for ophthalmic diagnosis. Specifically, CMNet consists of three main modules: an EfficientSAM encoder with an embedded adapter, a cross-modal feature alignment module, and a contrast consistency learning module. The pre-trained large model with the embedded adapter incorporates the unique geometric features and spatial information of retinal vessels while maintaining excellent generalization ability, enhancing the large model's adaptability to vessel segmentation tasks. The cross-modal feature alignment and fusion module performs accurate extraction and deep integration of inter-modal and intra-modal data. The intramodal information enhancement module performs independent information fusion and weighting for each modality, while the intermodal wavelet extraction and fusion module uses wavelet transform to convert spatial features into high- and low-frequency components, better capturing detailed information and global contours. Cross-attention is used to mix intermodal features. Finally, the contrastive consistency learning module fully utilizes different types of consistency to improve the model's robustness and segmentation performance.
[0063] Specifically, a semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal approaches, such as... Figure 1 As shown, it includes a preprocessing module, an encoder module, a cross-modal feature alignment and fusion module, and a contrast consistency learning module. The cross-modal feature alignment and fusion module includes an intra-modal OCT information enhancement unit, an intra-modal OCT information enhancement unit, and an inter-modal wavelet extraction and fusion unit.
[0064] The preprocessing module is used to perform data enhancement on the OCT image and the OCT A image respectively, so as to obtain the OCT enhanced image and the OCT A enhanced image.
[0065] The encoder module is used to extract the geometric shape features of the OCT-enhanced image and the OCT-enhanced image respectively, thereby obtaining the corresponding OCT feature map and OCT-enhanced feature map; wherein, the method by which the encoder module extracts the geometric shape features of any enhanced image to obtain any feature map is as follows:
[0066] The current input augmented image is downsampled and flattened into a two-dimensional tensor through convolution calculation with a specific stride; a corresponding position encoding vector is generated for the three-dimensional position information of each pixel in the current input augmented image, and then added to the two-dimensional tensor to obtain a two-dimensional encoded tensor with embedded position information; the two-dimensional encoded tensor is input into a Transformer unit composed of multiple cascaded Transformer blocks for feature extraction to obtain the feature map corresponding to the current input augmented image.
[0067] Furthermore, the method for feature extraction of the input tensor received by any Transformer block is as follows:
[0068] The input tensor is subjected to layer normalization to obtain a normalized tensor. A multi-head self-attention mechanism is used to extract semantic space information of different dimensions of the normalized tensor, and then the semantic space information of different dimensions is added to the normalized tensor to obtain the original self-attention tensor. The original self-attention tensor is subjected to layer normalization to obtain a normalized self-attention tensor. The normalized self-attention tensor is linearized to obtain a linear self-attention tensor. The linear self-attention tensor is added to the original self-attention tensor to obtain a fused self-attention tensor. The fused self-attention tensor is dimensionality reduced to obtain a dimensionality-reduced self-attention tensor. An adapter is used to enhance the features of the dimensionality-reduced self-attention tensor, and the enhanced result is the final output.
[0069] It should be noted that the adapter's method for feature enhancement of the dimensionality-reduced self-attention tensor is as follows:
[0070] First, the dimensionality-reduced self-attention tensor is reconstructed into a 3D image. Then, three consecutive adaptation operations are performed on the 3D image, and the output of the last adaptation operation is reconstructed into a 2D image. The 2D image represents the final enhanced output. Specifically, the adaptation operation involves:
[0071] Use 3 A 3-dimensional convolution kernel is used to perform a convolution operation on the current input 3D image; layer normalization is applied to adjust the input of each layer of the 3D image after the convolution operation; the nonlinear expressive power of the 3D image after the layer normalization operation is increased by an activation function; the 3D image before the convolution operation and the 3D image after the nonlinear expressive power are superimposed by adding the residuals to obtain the final output result of the adaptation operation.
[0072] The intramodal OCT information enhancement unit is used to perform channel and spatial dimension weighted enhancement on the OCT feature map, and then the resulting enhanced OCT feature map is concatenated with the original OCT feature map to obtain the first concatenated feature map. ;
[0073] The intramodal OCTA information enhancement unit is used to perform channel and spatial dimension weighted enhancement on the OCTA feature map, and then the resulting enhanced OCTA feature map is concatenated with the original OCTA feature map to obtain a second concatenated feature map. ;
[0074] The intermodal wavelet extraction and fusion unit is used to extract the spatial and frequency domain information of the OCT feature map and the OCTA feature map using Haar wavelet transform, and to fuse the spatial and frequency domain information corresponding to the two feature maps through a cross-attention mechanism. Then, the fused feature map is concatenated with the OCT enhanced feature map and the OCTA enhanced feature map to obtain the third concatenated feature map. Specifically, the method for obtaining the fused feature map by the intermodal wavelet extraction fusion unit is as follows:
[0075] The Haar wavelet transform is used to transform the OCT feature map and the OCTA feature map into the following components: OCT low frequency component, OCT horizontal high frequency component, OCT vertical high frequency component, OCT diagonal high frequency component, OCTA low frequency component, OCTA horizontal high frequency component, OCTA vertical high frequency component, and OCTA diagonal high frequency component, respectively.
[0076] The low-frequency components of OCT are used as global features representing the OCT feature map; the horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components of OCT are concatenated, and then 1... 1. Convolutional low-dimensional mapping yields OCT high-frequency features that characterize the local detail features of the OCT feature map;
[0077] The low-frequency components of OCTA are used as the global features of the OCTA feature map; the horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components of OCTA are concatenated, and then 1... 1. Convolutional low-dimensional mapping yields high-frequency OCTA features that characterize the local details of the OCTA feature map;
[0078] A multi-head attention mechanism is used to extract and fuse semantic space information from different dimensions of OCT high-frequency features and OCTA high-frequency features, respectively, to obtain the first fused image of OCT and the first fused image of OCTA.
[0079] A multi-head attention mechanism is used to extract and fuse semantic space information from different dimensions of the OCT feature map and the OTA feature map, respectively, to obtain the OCT second fused image and the OTA second fused image.
[0080] A cross-fusion attention mechanism is used to combine OCT low-frequency features, OCTA low-frequency features, the first fused OCT image, the first fused OCTA image, the second fused OCT image, and the second fused OCTA image. Then, the combined image is processed through a 3D model. 3. Convolution, batch normalization, and activation functions are used to effectively fuse features to obtain the final fused feature map.
[0081] The contrastive consistency learning module is used for... , , Regularization is then applied to obtain the final segmentation result.
[0082] Thus, it can be seen that the pre-trained large model encoder of the embedded adapter used in this invention performs initial feature extraction on data from different modalities; the cross-modal feature alignment and fusion module effectively weights and fuses data between and within modalities; the contrast consistency learning module improves the generalization ability and robust performance of the network through a joint loss function; and for the first time, it achieves efficient and accurate segmentation of retinal vascular images through multimodal fusion technology and a semi-supervised training paradigm.
[0083] Based on the above, the data processing flow of the segmentation network of the present invention can be summarized into three main steps.
[0084] Step 1: Use the encoder module, which is fine-tuned with an embedded adapter, to extract features from the feature maps of the two modalities.
[0085] It should be noted that this invention employs an encoder module based on fine-tuned EfficientSAM to extract features from the feature maps of two modalities. Specifically, the structure of the fine-tuned EfficientSAM encoder is as follows: Figure 2 As shown:
[0086] In fundamental visual tasks, convolutional neural networks or Transformer structures are typically used for image feature extraction. However, with the development of large-scale model technology in recent years, more and more downstream tasks are using large models or components of large models as auxiliary modules for feature extraction. The Segment Anything Model (SAM) image segmentation large-scale model recently proposed by the Meta team achieves zero-shot segmentation capabilities through active learning training and the large-scale dataset SA-1B, and can also serve as a feature extractor, providing a basic paradigm for downstream tasks. Utilizing the powerful segmentation generalization performance of large models to process medical image tasks has also become a research hotspot. However, its huge computational cost also limits its application in a wider range of scenarios to some extent.
[0087] Preferably, this embodiment uses a lightweight version of the SAM model, namely the encoder of EfficientSAM, which acts as the feature extractor.
[0088] In this embodiment, the lightweight encoder from EfficientSAM is embedded as a feature extractor into the network model. While EfficientSAM learns adaptability to different image distributions from the SAM encoder and performs well in natural image segmentation tasks, the SA-1B dataset contains a relatively small proportion of medical images, and the retinal vessel segmentation task involves numerous segmentation targets with weak boundaries, low contrast, small size, and irregular shapes. Therefore, directly applying EfficientSAM has limited effectiveness. To improve its generalization performance in retinal vessel images, a learnable adapter layer is inserted into the image encoder of EfficientSAM. This allows the encoder to better adapt to the specific features and distribution of retinal vessel images with only a few parameter adjustments, thereby significantly improving its performance in segmentation tasks.
[0089] The specific structure is as follows: First, the input feature map is downsampled and flattened into a two-dimensional tensor through convolution with a specific stride, facilitating subsequent self-attention calculation. However, three-dimensional feature maps typically have a spatial structure, where features at each location have relative relationships with features at other locations. The flattening operation converts the three-dimensional feature map with positional information into two-dimensional information, destroying the original spatial structure and making it difficult for the model to capture the relative positional relationships between features. Therefore, positional information needs to be embedded in the vector, assigning a fixed encoding vector to each location in the sequence, which is directly added to the two-dimensional tensor. In other words, this invention directly references the model weights in the original EfficientSAM during training for the flattening operation and positional encoding part, freezing these model weights during training and not participating in gradient calculation and backpropagation. After absolute positional encoding, the data is input into a Transformer block that is repeated multiple times for feature extraction. To enable the model to quickly adapt to different tasks and avoid catastrophic forgetting that may result from full model fine-tuning, an adapter with only a few parameters is inserted into the Transformer block, allowing the model to be fine-tuned for specific tasks while maintaining the pre-trained weights, enabling it to learn the unique feature information of retinal vessels during training. Specifically, the parameters of the Transformer block in the original EfficientSAM are frozen, and a trainable module is added after the regular multi-head self-attention structure. First, the features are dimensionality-reduced to further reduce computation, and then an adapter is used to further refine the features. The dimensionality-reduced 2D tensor is first reconstructed into a 3D image, using 3D... A 3x3 convolutional kernel performs a convolution operation on the feature map to further extract features. After convolution, layer normalization is applied to adjust the input of each layer, which helps stabilize and accelerate the training process, improving model performance and generalization ability. Then, an activation function is used to increase the model's non-linear expressive power. The feature map before convolution is then superimposed on the output by adding residuals, alleviating the gradient vanishing problem in deep networks. This combination of convolution normalization and activation function is repeated three times, and the feature map is then reshaped into two-dimensional data for subsequent module computations.
[0090] Step 2: Weighted enhancement and information fusion of intra-modal features and inter-modal information are performed through the cross-modal feature alignment and fusion module.
[0091] The fine-tuned EfficientSAM encoder module proposed in step 1 performs independent feature extraction on OCT and OCT image data, without involving information interaction between modalities. In contrast, the cross-modal feature alignment and fusion module in step 2, while enhancing information within each modality, also extracts high- and low-frequency detail information from different modalities through wavelet transform and fuses information of different fine granularities from different modalities through a cross-attention mechanism. This not only enhances the model's ability to capture intramodal features but also further improves the model's understanding of complex data structures through information interaction between modalities.
[0092] First, the feature maps of different modalities are input into different intramodal information enhancement modules to enhance the feature representation of a single modality and extract key information at different locations in space.
[0093] Preferably, in this embodiment, a Convolutional Block Attention Module (CBAM) is selected as the feature extraction part of the intramodal information enhancement module.
[0094] While inputting feature maps of different modalities into the intramodal information enhancement module, this section proposes a wavelet extraction and fusion module to fully utilize the high-resolution structural information of OCT images and the dynamic blood flow information of blood vessels provided by OCTA images, and to better utilize multimodal data for retinal vessel identification and analysis. This module aims to decompose the image into high-frequency and low-frequency components at different scales using wavelet transform, extract edge texture information and overall contour information, and then perform intermodal information cross-fusion on the original feature map and the high- and low-frequency components through cross-self-attention. The structure is as follows: Figure 3 As shown.
[0095] Specifically, for input First, proceed with step 1. 1. Convolution operations are used to increase the non-linearity of features, thereby producing new features with invariant dimension. Then, the Haar wavelet transform is used to transform the features into a low-frequency component, a horizontal high-frequency component, a vertical high-frequency component, and a diagonal high-frequency component. The specific operation is as follows: Figure 4 The A, H, V, and D components are shown in the figure.
[0096] For each channel The first-order Haar wavelet transform is calculated as follows:
[0097] (2)
[0098] in This represents the low-frequency approximation coefficients of the channel, and Represents the high-frequency detail coefficients of the channel, where i is the row index and j is the column index. i ranges from 1 to... j ranges from 1 to Where H and W are the width and height of the feature map. Then, Haar wavelet transform is applied to each column of the approximation coefficients and detail coefficients.
[0099] (3)
[0100] in, Represents the approximation coefficient for a single channel. This represents the horizontal detail factor for a single channel. This represents the vertical detail factor for a single channel. This represents the diagonal detail coefficient for a single channel, where i ranges from 1 to... j ranges from 1 to Finally, the three high-frequency components are spliced together, and then 1... 1. Convolution is used to perform low-dimensional mapping, resulting in the final high-frequency and low-frequency features.
[0101] (4)
[0102] in, and Representing high-frequency and low-frequency features respectively, with dimensions of , This indicates batch normalization calculation. Indicates 1 1. Convolution Operation. The wavelet high- and low-frequency feature decomposition module decomposes spatial domain features into high-frequency features with local details and low-frequency features with global characteristics, providing a more comprehensive feature set for the segmentation network to consider frequency domain information. The high- and low-frequency features are then concatenated and input into the subsequent cross-fusion attention module.
[0103] Frequency domain features and spatial domain features capture different aspects and attributes of an image. The next step is to align the spatial and frequency domain information between modalities to ensure consistency and complementarity. This invention designs a cross-fusion attention module that uses each other's query matrices to calculate attention scores during multi-head attention operations, enabling full interaction of information between different modalities. Finally, semantic alignment and feature selection of frequency and spatial domain features are achieved through spatial-frequency domain feature fusion.
[0104] Specifically, for input The three-dimensional feature map is converted into two-dimensional data through a flattening operation. Finally, it is multiplied with the transformation matrix to obtain the query matrix, key matrix and value matrix. The query matrices of the two modalities are exchanged before attention calculation is performed.
[0105] (5)
[0106] in Representative 1 1. Convolution operation, the attention calculation formula is as follows (4.6), where... It is the length of the flattened sequence, and its size is... .
[0107] (6)
[0108] After fusing information between different modalities using cross-fusion attention, it is also necessary to combine the frequency domain and spatial domain information. Therefore, the feature maps from different domains are concatenated and then processed using 3D modeling. 3. Effective feature fusion is achieved through a combination of convolution, batch normalization, and activation functions.
[0109] After the feature maps from the fine-tuned EfficientSAM encoder module undergo intra-modal and inter-modal information extraction and fusion enhancement, feature data from two individual modalities and one fused modal are obtained. The feature maps from the individual modalities are then concatenated with the feature map without intra-modal information enhancement, allowing for interaction between shallow overall contour information and deep detail information, thus enhancing the network's expressive power. Furthermore, the three outputs from the cross-modal feature alignment and fusion module are concatenated to further strengthen the fusion of data from different modalities. Finally, the three output feature maps are passed through their respective independent decoders. Three output results are then generated.
[0110] Step 3: The consistency learning module uses the loss function to establish regularization under different consistency levels, thereby improving the model's robustness and segmentation performance.
[0111] Because various modalities possess both correlation and difference information, data from different modalities, after being output through different network modules, also exhibit both differences and correlations. Predicting their consistency is the driving force behind semi-supervised cross-modal knowledge interaction learning. Specifically, for labeled outputs... and We employ conventional supervised training, using a combination of Dice loss and cross-entropy loss as the loss function, and calculating the loss function from the decoded feature maps of individual OCT modalities.
[0112] (7)
[0113] Where CE represents cross-entropy loss, Dice represents Dice loss, Y1 represents the predicted segmentation result obtained from the labeled OCT image, and Y represents the segmentation result label value determined jointly by the labeled OCT image and the OCT image.
[0114] Calculate the loss function from the decoded feature map of the OCTA modality:
[0115] (8)
[0116] Y2 represents the predicted value of the segmentation result obtained from the labeled OCTA image;
[0117] For unlabeled data, the multimodal fusion output will be used. and single-mode output and Mutual supervision is conducted between them, and the role of regularization methods in training is brought into full play by model consistency and data consistency.
[0118] (9)
[0119] (10)
[0120] Y3 indicates that, according to The obtained segmentation prediction values are Y1' and Y2', respectively, which represent the segmentation prediction values obtained from the unlabeled OCT image and the unlabeled OCT image.
[0121] Adding the above loss functions together yields the final total loss function Loss=L sup1 +L sup2 +L mut1 +L mut2 .
[0122] The above are all the steps of the semi-supervised feature fusion segmentation network based on OCT-OCTA multimodal architecture.
[0123] In other words, this invention first processes the dual-modal input data through image preprocessing and data augmentation, then independently inputs each modality into the EfficientSAM encoder module for fine-tuning. This module enables fine-tuning of a large-model encoder with only a small number of parameters, learning the unique geometric features of retinal vessels, and combining the powerful generalization ability of the large model to perform preliminary extraction of vascular pixels from different modalities (Step 1). Subsequently, the processed feature maps are input into a cross-modal feature alignment and fusion module (Step 2). This module includes an intra-modal information enhancement module and an inter-modal wavelet extraction and fusion module. The intra-modal information enhancement module independently weights and enhances the feature information of each modality in terms of channel and spatial dimensions. The inter-modal wavelet extraction and fusion module uses Haar wavelet transform to extract frequency domain information from feature maps of different modalities and calculates the fusion of inter-modal information through a cross-attention mechanism, further enhancing the ability to identify and capture vascular edges and global vascular pixels. Finally, a contrastive consistency learning module establishes regularization between different perturbation outputs, improving the robustness of the network model to image inputs from different modalities (Step 3).
[0124] Existing deep learning-based retinal vessel segmentation techniques struggle to achieve ideal segmentation results in the absence of ground truth annotations, and the few semi-supervised networks available for retinal vessel segmentation can only handle a single input modality. To address the limited information available in single-modality retinal fundus vessel data, this invention proposes a cross-modal feature fusion network based on dual-modal outputs of OCT and OCTTA. This network aims to fuse and complement information from different modalities, thereby improving the accuracy and connectivity of vessel segmentation.
[0125] Therefore, the semi-supervised multimodal blood vessel image segmentation method provided by this invention can effectively utilize inputs from both OCT and OCTA modalities, and perform efficient and accurate segmentation of blood vessels through a designed network module. This is the first solution for segmentation using a semi-supervised training method based on OCT and OCTA dual-modal data, providing a new system for disease diagnosis, pathological analysis, and mesoscopic tissue localization. Comparative experiments were conducted using 10% labeled data from the OCTA-500 dataset with other single-modal semi-supervised methods. Detailed comparisons of the segmentation results are shown in [the provided text]. Figure 5 and Figure 6 Among them.
[0126] in, Figure 5 The results show the results obtained by training different segmentation systems using 10% OCT-annotated data. Figure 5The area highlighted in the middle box clearly demonstrates the segmentation capabilities of the proposed CMNet for connected regions of thick blood vessels and its ability to capture light-colored blood vessels. Traditional segmentation systems cannot effectively segment entire thick blood vessels, nor can they accurately identify the light-colored regions in the center of the vessel. However, the CMNet of this invention can utilize information from both OCT and OCTTA modalities, and through the modules proposed in this invention, accurately extract and fully fuse edge and contextual information, significantly improving the performance of blood vessel segmentation. Figure 5 The area marked by the red box shows that the CMNet of this invention can effectively control false positives and reduce the probability of predicting incorrect blood vessel pixels while extracting a large number of foreground blood vessel pixels.
[0127] Figure 6 This shows OCTA images using different segmentation systems with a 20% annotation ratio. From Figure 6 As shown in the area marked in the middle frame, traditional segmentation systems exhibit incomplete segmentation and insufficient connectivity at the initial intersections of small blood vessels. MC-Net+ failed to correctly segment the small blood vessels at the intersections, while SS-Net could identify most of the vessel edges, but the segmentation results contained holes and insufficient connectivity. In contrast, the CMNet of this invention can accurately capture the tiny pixels of small blood vessels and their intersections while preserving overall global information, resulting in better continuity in blood vessel segmentation. Figure 6 Another box marks a location showing a separate, fine blood vessel with a highly curved shape. Most traditional segmentation systems exhibit varying degrees of false positives when segmenting this vessel and have weak ability to identify and judge the curves. In contrast, the segmentation system CMNet of this invention retains more pixels at the curves of the blood vessel while controlling erroneous segmentation, significantly improving segmentation accuracy and user experience.
[0128] In summary, the fine-tuning EfficientSAM encoder module proposed in this invention enables fine-tuning of a large-scale encoder model with only a small number of parameters, learning the unique geometric features of retinal vessels, and combining the powerful generalization ability of the large model to perform preliminary extraction of vascular pixels from different modalities. The intramodal information enhancement module independently weights and enhances the feature information of each modality in terms of channel and spatial dimensions. The intermodal wavelet extraction and fusion module uses Haar wavelet transform to extract frequency domain information from feature maps of different modalities and calculates the fusion of intermodal information through a cross-attention mechanism, further enhancing the ability to identify and capture vascular edges and global vascular pixels. The contrast consistency learning module establishes regularization between different perturbation outputs, improving the robustness of the network model to image inputs from different modalities. Experiments on the OCTA-500 dataset also fully demonstrate the advantages of this model over other single-modal models.
[0129] Based on the above description of the embodiments, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related software. The program can be programmed as Matlab code, C++, or Python code, and when executed, it can include the processes of the embodiments of the above methods. Those skilled in the art can understand and implement this without any creative effort.
[0130] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal approaches, characterized in that, It includes a preprocessing module, an encoder module, a cross-modal feature alignment and fusion module, and a contrast consistency learning module. The cross-modal feature alignment and fusion module includes an intra-modal OCT information enhancement unit, an intra-modal OCTA information enhancement unit, and an inter-modal wavelet extraction and fusion unit. The preprocessing module is used to perform data enhancement on the OCT image and the OCT A image respectively, so as to obtain the OCT enhanced image and the OCT A enhanced image. The encoder module is used to extract the geometric shape features of the OCT-enhanced image and the OCTA-enhanced image respectively, and obtain the corresponding OCT feature map and OCTA feature map; The intramodal OCT information enhancement unit is used to perform channel and spatial dimension weighted enhancement on the OCT feature map, and then the resulting enhanced OCT feature map is concatenated with the original OCT feature map to obtain the first concatenated feature map. ; The intramodal OCTA information enhancement unit is used to perform channel and spatial dimension weighted enhancement on the OCTA feature map, and then the resulting enhanced OCTA feature map is concatenated with the original OCTA feature map to obtain a second concatenated feature map. ; The intermodal wavelet extraction and fusion unit is used to extract the spatial and frequency domain information of the OCT feature map and the OCTA feature map using Haar wavelet transform, and to fuse the spatial and frequency domain information corresponding to the two feature maps through a cross-attention mechanism. Then, the fused feature map is concatenated with the OCT enhanced feature map and the OCTA enhanced feature map to obtain the third concatenated feature map. ; The contrastive consistency learning module is used for... , , Regularization is performed to obtain the final segmentation result; The method for obtaining the fusion feature map by intermodal wavelet extraction fusion unit is as follows: The Haar wavelet transform is used to transform the OCT feature map and the OCTA feature map into the following components: OCT low frequency component, OCT horizontal high frequency component, OCT vertical high frequency component, OCT diagonal high frequency component, OCTA low frequency component, OCTA horizontal high frequency component, OCTA vertical high frequency component, and OCTA diagonal high frequency component, respectively. The low-frequency components of OCT are used as global features representing the OCT feature map; the horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components of OCT are concatenated, and then 1...
1. Convolutional low-dimensional mapping yields OCT high-frequency features that characterize the local detail features of the OCT feature map; The low-frequency components of OCTA are used as the global features of the OCTA feature map; the horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components of OCTA are concatenated, and then 1...
1. Convolutional low-dimensional mapping yields high-frequency OCTA features that characterize the local details of the OCTA feature map; A multi-head attention mechanism is used to extract and fuse semantic space information from different dimensions of OCT high-frequency features and OCTA high-frequency features, respectively, to obtain the first fused image of OCT and the first fused image of OCTA. A multi-head attention mechanism is used to extract and fuse semantic space information from different dimensions of the OCT feature map and the OTA feature map, respectively, to obtain the OCT second fused image and the OTA second fused image. A cross-fusion attention mechanism is used to combine OCT low-frequency features, OCTA low-frequency features, the first fused OCT image, the first fused OCTA image, the second fused OCT image, and the second fused OCTA image. Then, the combined image is processed through a 3D model.
3. Convolution, batch normalization, and activation functions are used to effectively fuse features to obtain the final fused feature map; Both the intramodal OCT information enhancement unit and the intramodal OCT / OCTA information enhancement unit employ convolutional attention modules to perform channel and spatial dimension weighted enhancement on the OCT and OCT feature maps.
2. The semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal as described in claim 1, characterized in that, The encoder module extracts the geometric shape features of any enhanced image to obtain any feature map using the following method: The current input enhanced image is downsampled and flattened into a two-dimensional tensor by convolution calculation with a specific stride. Generate a corresponding position encoding vector for the three-dimensional position information of each pixel in the current input enhanced image, and then add it to the corresponding two-dimensional tensor to obtain a two-dimensional encoding tensor with embedded position information; The two-dimensional encoded tensor is input into a Transformer unit composed of multiple cascaded Transformer blocks for feature extraction, resulting in a feature map corresponding to the current input enhanced image.
3. The semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal as described in claim 2, characterized in that, The method for any Transformer block to extract features from the input tensors it receives is as follows: The input tensor is subjected to layer normalization to obtain a normalized tensor; A multi-head self-attention mechanism is used to extract semantic space information of different dimensions of the normalized tensor. Then, the semantic space information of different dimensions is added to the normalized tensor to obtain the original self-attention tensor. The original self-attention tensor is subjected to layer normalization to obtain a normalized self-attention tensor. Linear processing is performed on the normalized self-attention tensor to obtain the linear self-attention tensor; Add the linear self-attention tensor to the original self-attention tensor to obtain the fused self-attention tensor; The dimensionality of the fused self-attention tensor is reduced to obtain the dimensionality-reduced self-attention tensor. An adapter is used to enhance the features of the dimension-reduced self-attention tensor, and the enhancement result is the final output.
4. The semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal as described in claim 3, characterized in that, The adapter method for feature enhancement of the dimension-reduced self-attention tensor is as follows: First, the dimension-reduced self-attention tensor is reshaped into a three-dimensional image; The 3D image undergoes three consecutive adaptation operations, and the output of the last adaptation operation is reshaped into a 2D image. The 2D image is the final enhanced output. The adaptation operation specifically involves: Use 3 A 3-dimensional convolution kernel performs a convolution operation on the current input 3D image; The application layer normalization operation adjusts the input of each layer of the 3D image after the convolution operation; The nonlinear expressive power of the 3D image after layer normalization is increased by using activation functions; The final output of the adaptation operation is obtained by superimposing the 3D image before the convolution operation and the 3D image after the nonlinear representation is enhanced by adding the residuals.
5. The semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal as described in claim 3, characterized in that, Only the network parameters in the adapter are used as adjustable parameters to be trained, while the network parameters of the other modules are frozen during the training process. The total loss function Loss used when training the adapter is as follows: Loss=L sup1 +L sup2 +L mut1 +L mut2 Among them, L sup1 L is a loss function constructed from the predicted segmentation results and the segmentation label values obtained from labeled OCT images. sup2 L is a loss function constructed from the predicted segmentation values and segmentation label values obtained from labeled OCTA images. mut1 The predicted value of the segmentation result obtained from the unlabeled OCT image is compared with the value obtained from the segmentation result obtained from the unlabeled OCT image. The loss function L is constructed from the predicted values of the obtained segmentation results. mut2 The predicted values of the segmentation results obtained from the unlabeled OCTA image are compared with those obtained from the... The loss function is constructed from the predicted values of the obtained segmentation results.
6. The semi-supervised feature fusion segmentation system based on OCT-OCTA multimodal as described in claim 5, characterized in that, L sup1 L sup2 L mut1 L mut2 The calculation method is as follows: Where CE represents cross-entropy loss, Dice represents Dice loss, Y1 represents the predicted segmentation result obtained from the labeled OCT image, Y represents the segmentation result label value determined jointly by the labeled OCT image and the OCTA image, Y2 represents the predicted segmentation result obtained from the labeled OCTA image, and Y3 represents the predicted segmentation result obtained from the labeled OCTA image. The obtained segmentation prediction values are Y1' and Y2', respectively, which represent the segmentation prediction values obtained from the unlabeled OCT image and the unlabeled OCT image.
Citation Information
Patent Citations
Multi-modal retina fundus image classification method
CN116824217A
Semi-supervised polyp segmentation method based on hierarchical granularity contrast learning
CN120107961A