Melanoma lesion area segmentation method based on CLIP multi-mode fusion network

By constructing the MA-CLIP model of the CLIP multimodal fusion network, the problems of insufficient data annotation and feature fusion in melanoma lesion region segmentation were solved, achieving accurate lesion segmentation and boundary identification, improving segmentation accuracy and robustness, and providing a reliable reference for lesion range for clinical diagnosis.

CN120807924APending Publication Date: 2025-10-17XIJING UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510916312.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies for melanoma lesion area segmentation have problems such as insufficient data annotation, limitations of single-modal information, low efficiency of cross-modal feature fusion, and insufficient multi-scale feature extraction, resulting in insufficient segmentation accuracy and boundary fuzziness.

Method used

A MA-CLIP model based on the CLIP multimodal fusion network is constructed. A large-scale multimodal dataset is generated by manually annotating text-image pairs. A cross-modal attention mechanism and a multi-scale hollow spatial pyramid pooling module are used in conjunction with the BAM module to accurately segment lesion regions, realizing bidirectional interaction between image and text features and multi-scale feature extraction.

Benefits of technology

It significantly improved the Dice coefficient, boundary F1 score, and crossover ratio for melanoma segmentation, providing accurate reference for lesion extent and promoting the practical application of AI-assisted diagnostic technology in dermatology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807924A_ABST
    Figure CN120807924A_ABST
Patent Text Reader

Abstract

The invention discloses a melanoma lesion area segmentation method based on a CLIP multi-modal fusion network. The method comprises the following steps: 1, constructing an MA-CLIP model; 2, a BLIP language model is finely adjusted through a manually-labeled text-image pair, a large-scale multi-modal data set is constructed, a training set, a test set and a verification set are divided, and preprocessing is carried out; 3, training the MA-CLIP model; 4, evaluating the performance of the MA-CLIP model and optimizing parameters; and 5, inputting a to-be-segmented melanoma clinical image into the trained MA-CLIP model, and outputting a segmentation result. According to the method, the problem of insufficient traditional medical data annotation can be solved, accurate guidance of clinical semantics on image segmentation is realized, and the recognition precision and boundary segmentation capability of a focus in a complex form are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of medical image analysis and computer vision, and particularly relates to a melanoma lesion region segmentation method based on a CLIP multi-modal fusion network. BACKGROUND

[0002] Melanoma, as a highly malignant skin tumor, its accurate segmentation of the lesion region is a key link for clinical diagnosis and treatment. At present, the image segmentation technology based on deep learning has been widely applied in the medical field, but the traditional method is mostly dependent on a single image mode, and it is difficult to combine the semantic description of the lesion (such as the clinical features of "irregular edge", "blue and white stripes", etc.), which limits the segmentation accuracy, especially in the processing of complex morphological lesions and fuzzy boundary areas.

[0003] With the development of multi-modal fusion technology, some studies attempt to combine text semantics and image features for segmentation, but there are still the following limitations:

[0004] (1) The data set is small in size and high in annotation cost, which is difficult to cover the diverse morphology of melanoma;

[0005] (2) Single modal information limitation: only relying on image pixel features, unable to directly use clinical text descriptions such as "melanoma often accompanied by uneven pigmentation" and "edge jagged", resulting in insufficient ability to distinguish similar textures (such as benign nevus and malignant lesions)

[0006] (3) Low efficiency of cross-modal feature fusion, and the guiding effect of semantic information on the lesion area is not fully utilized; the existing multi-modal model does not design a specific attention mechanism, and the text semantic features cannot accurately locate the lesion area in the image, resulting in the problem of "semantic drift" (such as the text description "irregular shape" cannot correspond to the specific edge area in the image);

[0007] (4) Insufficient multi-scale feature extraction, unable to balance global semantics and local details; traditional segmentation network is difficult to effectively fuse semantic information of different levels (such as high-level semantic "tumor category" and low-level feature "edge contour") in the up-sampling process, resulting in blurred segmentation boundary and missed detection of small lesions. SUMMARY

[0008] The purpose of the present application is to provide a melanoma lesion region segmentation method based on a CLIP multi-modal fusion network, which can solve the problem of insufficient traditional medical data annotation, realize the accurate guidance of clinical semantics to image segmentation, and improve the recognition accuracy and boundary segmentation ability of complex morphological lesions.

[0009] To achieve the above purpose, the present application provides the following technical scheme:

[0010] The melanoma lesion region segmentation method based on the CLIP multi-modal fusion network comprises the following steps:

[0011] Step 1, constructing a MA-CLIP model;

[0012] The MA-CLIP model comprises an image and text feature extraction module, a cross-modal attention (CMA) module, a multi-scale atrous spatial pyramid pooling (MS-ASPP) module, and a BAM module;

[0013] Step 2, fine-tuning the BLIP language model through manually labeled text-image pairs, constructing a large-scale multi-modal dataset, dividing the training set, test set and validation set, and performing preprocessing;

[0014] Step 3, training the MA-CLIP model;

[0015] Step 3.1, extracting semantic features and visual features from clinical text and images respectively through the Transformer text encoder and Vision Transformer backbone network of the CLIP model, and finally outputting hierarchical image features and text semantic guidance vectors;

[0016] Step 3.2, the image features and text semantic vectors enter the CMA module for bidirectional interaction to generate fusion features with visual information and semantic representation;

[0017] Step 3.3, the fusion features output by the CMA module enter the MS-ASPP module to extract multi-scale context through multiple parallel branches;

[0018] Step 3.4, the multi-scale fusion features output by the MS-ASPP module enter the BAM module for optimization to overcome the fuzzy defects of traditional models in lesion boundary delineation;

[0019] Step 3.5, calculating the total loss

[0020] Step 3.6, the features output by the BAM module are restored to the original image size through three layers of progressive upsampling, a 1x1 convolution is used to generate a probability map which is threshold segmented to output the lesion segmentation results with accurate boundaries and complete regions;

[0021] Step 3.7, updating the parameters of the MA-CLIP model;

[0022] Step 4, evaluating the performance of the MA-CLIP model using the validation set and optimizing the parameters;

[0023] Step 5, inputting the melanoma clinical image to be segmented into the trained MA-CLIP model to output the segmentation result.

[0024] Further, the specific content of step 2 includes:

[0025] Step 2.1, select melanoma clinical images and their pixel-level annotations from the ISIC2018 dataset, and write text descriptions containing location, color, and shape to form an initial text-image pair dataset;

[0026] Step 2.2, fine-tune the BLIP pre-trained model using the initial text-image pair dataset to optimize the image-text feature alignment capability, and automatically generate text descriptions based on the fine-tuned BLIP pre-trained model for the annotated images to construct a large-scale multi-modal dataset;

[0027] Step 2.3, divide into training set, validation set, and test set, then perform data preprocessing, for image data, first unify the size to 256x256 pixels to ensure consistent input dimensions, then normalize the RGB channel values to the [0,1] interval; for text descriptions, use the CLIP text encoder to convert natural language descriptions into 768-dimensional semantic vectors to adapt the feature dimensions to the features extracted by the image branch, and finally output standardized image feature vectors and text semantic vectors.

[0028] Further, the specific steps of step 3.1 include:

[0029] For the text branch, the clinical text description is processed by the CLIP Transformer text encoder to convert it into an initial sequence vector; then the semantic features are extracted through multiple Transformer blocks to capture the semantic information in the text, outputting a 768-dimensional text embedding vector; finally, a 256-dimensional semantic guidance vector is generated through global pooling operation;

[0030] For the image branch, the input image is processed by the CLIP Vision Transformer (ViT-B / 16) backbone network, first split the image into 16x16 patches and encode them through Patch Embedding to generate low-level texture features; then extract middle-level lesion shape features through multi-head self-attention layers, and finally obtain high-level semantic features through global pooling; and each layer of features is reduced to 256 channels through 1x1 convolution to align the feature dimensions with the text features, and the multi-scale information is preserved through skip connection to avoid loss of spatial details in high-level semantics.

[0031] Further, the specific content of step 3.2 includes:

[0032] In one aspect, the image features are converted into a query matrix Q_img, a key matrix K_text, and a value matrix V_text through learnable matrices W_q, W_k, and W_v, where the dimension of the image features is 256xHxW, H represents the height of the feature map, and W represents the width of the feature map; an image-to-text attention map Attn_img is generated by calculating attention scores and weighted summation, allowing the image features to focus on the key semantics of the text description;

[0033] On the other hand, the 256-dimensional text semantic vector is converted into a query matrix Q_text, a key matrix K_img, and a value matrix V_img through W_q', W_k', and W_v', generating a text-to-image attention map Attn_text to accurately locate the lesion area in the image;

[0034] Finally, the original image features and the two types of attention maps are added to obtain the fused features F_fused, achieving accurate alignment of text semantics and visual features.

[0035] Further, the specific content of step 3.3 includes constructing a multi-branch parallel structure, 1x1 convolution is used to extract local texture details, 3x3 dilated convolution is set with expansion rates of 6, 12, and 18 respectively, corresponding receptive fields are 13x13, 25x25, and 37x37, and global average pooling is used to obtain global semantics;

[0036] Skip connections are added between the convolution branches with expansion rates of 6 and 12, and 12 and 18, small-scale edge features and large-scale semantic features are fused through 1x1 convolution, and the recognition ability of complex boundaries and small lesions is improved;

[0037] The outputs of each branch are weighted through a channel attention block to suppress background noise channels and strengthen lesion-related feature channels.

[0038] Further, the specific content of step 3.4 includes:

[0039] First, two layers of 3x3 dilated convolution are used to extract pixel-level boundary clues, capturing details including "jagged" and "fuzzy transition", where the expansion rates of the dilated convolution are 1 and 2 respectively;

[0040] Then, global average pooling and 1x1 convolution are applied to the boundary feature map to generate a boundary attention map A_edge through a sigmoid function, and the features are weighted according to the boundary attention map to strengthen the feature response of the "fuzzy transition zone" boundary details;

[0041] Finally, the output feature F_out = fused feature x A_edge + boundary feature F_edge is obtained, which strengthens the boundary region through attention weight and preserves semantic information.

[0042] Further, the total loss of step 3.5 is:

[0043]

[0044] where λ1=0.3, λ2=0.3, λ3=0.4, represents the Dice loss, where P represents the segmentation result predicted by the model, G represents the true segmentation label, and ε is a very small positive number (for example, 1e -7 , used to prevent the denominator from being zero; represents the cross-entropy loss, where y i represents the true label of the i-th sample, and p i represents the probability prediction value of the model that the i-th sample belongs to the positive class; represents the boundary loss, where B represents the pixel set of the boundary region, y i represents the true label of the i-th pixel, and p i represents the probability prediction value of the model that the i-th pixel belongs to the target region.

[0045] Further, the specific content of step 3.6 includes:

[0046] First, a three-layer progressive upsampling strategy is adopted: the first layer fuses the feature map with high-level semantic features and global pooling results through 4 times upsampling, realizes the preliminary positioning of the lesion overall region, the second layer fuses the middle-level morphological features and the dilated convolution feature with the dilated rate of 12 through 8 times upsampling, refines the lesion contour, and the third layer fuses the low-level texture features and the dilated convolution feature with the dilated rate of 6 through 16 times upsampling, completes the optimization of the detail part; in the upsampling process, the features of each layer are connected through jump connection to realize cross-scale information fusion, and avoid the loss of details caused by upsampling;

[0047] After the feature map is restored to the original input size, a pixel-level classification probability map is generated through 1*1 convolution, and then the probability map is converted into a binary lesion mask through threshold segmentation, and a pixel-level classification confidence map is output at the same time, which clearly distinguishes the melanoma lesion from the normal skin area, and finally outputs a segmentation result containing accurate boundaries and complete regions.

[0048] Compared with the prior art, the present application has the following beneficial effects:

[0049] (1) The application realizes accurate lesion segmentation under semantic guidance by constructing a text-image multi-modal data set and a cross-modal attention mechanism. The method refines the BLIP language model through manually annotated text-image pairs to generate a large-scale multi-modal data set, solving the problem of insufficient traditional medical data annotation. The method extracts hierarchical features by combining the visual and text encoders of CLIP, realizes bidirectional interaction of image and text features through the CMA module, and accurately locates the image lesion area of clinical semantics such as "irregular edge" and "uneven pigment". With the help of the multi-scale hollow convolution of the MA-SPP module to capture the context information of different receptive fields, and the BAM boundary refinement module to strengthen the pixel-level boundary details, the final segmentation result containing complete area and clear boundary is generated through progressive upsampling. The application can significantly improve the Dice coefficient and boundary F1 score of melanoma segmentation, providing accurate lesion range reference for clinical diagnosis, while ensuring inference efficiency through lightweight design, promoting the practical application of AI-assisted diagnosis technology in dermatology. Compared with the classic model, the core indicators of the application are comprehensive: F1 score is improved by 19%, sensitivity is improved by 17%, and intersection over union is improved by 23%, realizing overall superiority in "precision-efficiency-explainability" and providing a new paradigm for technology upgrade.

[0050] (2) Multi-modal fusion accuracy: Break through the limitation of single mode, innovate to fuse clinical text semantics and image features, use CLIP to build correlation, and accurately map the lesion area to the pathological description. Compared with image input only, the miou and F1 of BLIP after fine-tuning are significantly improved, helping to break through visual interference and achieve more accurate recognition and segmentation.

[0051] (3) Model architecture adaptability: For clinical pain points, design multi-level modules: CMA bidirectional interaction aligns semantic and image features, solving the problem of semantic understanding; MASPP+BAM uses hollow convolution to cover multiple scales and attention to strengthen boundaries. Compared with U-net, the intersection over union is higher, and the segmentation is more consistent with the real lesion.

[0052] (4) Clinical landing practicality: Real scene value is high, pseudo-labels are generated by BLIP to solve the problem of annotation scarcity and improve the recognition rate of rare subtypes; lightweight architecture ensures inference speed to meet real-time needs, semantic-image visualization alignment enhances doctor trust, and promotes clinical translation. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 is the model structure diagram of the application;

[0054] Figure 2 is a CMA module schematic diagram;

[0055] Figure 3 is a MS-ASPP module schematic diagram;

[0056] Figure 4 is a schematic diagram of the BAM module;

[0057] Figure 5 is a numerical result graph of the performance comparison experiment of different models;

[0058] Figure 6 is a bar graph of the performance comparison experiment results of different models;

[0059] Figure 7 is a numerical result graph of the ablation experiment of each module of the model;

[0060] Figure 8 is a bar graph of the ablation experiment results of the model separation module. DETAILED DESCRIPTION

[0061] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0062] The specific steps of the melanoma lesion region segmentation method based on the CLIP multi-modal fusion network according to the embodiment are as follows:

[0063] Step 1, constructing a MA-CLIP model

[0064] As shown in Figure 1 , the MA-CLIP model includes image and text feature extraction, a cross-modal attention (CMA) module, a multi-scale atrous spatial pyramid pooling (MS-ASPP) module, and a BAM module.

[0065] Step 2, fine-tuning the BLIP language model through manually labeled text-image pairs, constructing a large-scale multi-modal data set, dividing the training set, test set and validation set, and preprocessing

[0066] Firstly, 200 clinical images of melanoma and their pixel-level annotations were selected from the ISIC2018 dataset, and a text description containing location, color, and shape (such as "left upper side brown and black irregular lesion") was written to form an initial 200 text-image pair dataset. Subsequently, the BLIP pre-trained model was fine-tuned using this dataset to optimize the image-text feature alignment capability, and then based on the fine-tuned model, text descriptions were automatically generated for 2500 labeled images to construct a large-scale multi-modal dataset. Finally, the data was divided into training set, validation set, and test set according to the ratio of 7:2:1, and then data preprocessing was performed. For image data, the size was first adjusted to 256x256 pixels to ensure consistent input dimensions, and then the RGB channel values were normalized to the [0,1] interval to eliminate pixel differences caused by different device acquisitions. At the same time, to enhance the generalization of the model, random flipping, rotation, brightness, and contrast adjustment were also applied. For text descriptions, the CLIP text encoder was used to convert natural language descriptions such as "irregular dark brown lesion with abnormal surface texture" into 768-dimensional semantic vectors, making the feature dimensions compatible with the image branch extracted features, and finally output standardized image feature vectors and text semantic vectors, laying the foundation for subsequent cross-modal feature interaction.

[0067] Step 3, training the MA-CLIP model

[0068] Step 3.1, extract semantic features and visual features from clinical text and images respectively through the Transformer text encoder of the CLIP model and the Vision Transformer backbone network, and finally output hierarchical image features and text semantic guide vectors

[0069] During the image and text dual-branch feature extraction (encoding layer) stage, the model simultaneously performs deep feature mining on text and image data, and finally outputs hierarchical image features and text semantic guide vectors, providing basic features for cross-modal fusion.

[0070] For the text branch, the clinical text description is processed by the Transformer text encoder of the CLIP to convert it into an initial sequence vector; then the semantic features are extracted through multiple layers of Transformer blocks to capture semantic information such as "uneven pigmentation" and "jagged edges" in the text, outputting a 768-dimensional text embedding vector; finally, a 256-dimensional semantic guide vector is generated through global pooling operation to accurately condense the features of key lesion descriptions.

[0071] For the image branch, the input image is processed by CLIP's Vision Transformer (ViT-B / 16) backbone network. The image is first divided into 16×16 patches and encoded using the Patch Embedding method to generate low-level texture features; then the mid-level lesion morphological features (such as irregular contours) are extracted through a multi-head self-attention layer, and finally high-level semantic features (such as "melanoma" category information) are obtained through global pooling. The features of each layer are reduced to 256 channels through 1×1 convolution and aligned with the text feature dimension. Multi-scale information is retained through skip connections to avoid the loss of spatial details in high-level semantics.

[0072] Step 3.2: Image features and text semantic vectors enter the CMA module for two-way interaction to generate fusion features that combine visual information and semantic representation.

[0073] In the cross-modal attention fusion (interaction layer) stage, based on the hierarchical image features extracted by the dual branches and the text semantic guidance vector, a bidirectional attention mechanism (Cross-Model Attention, CMA) is constructed to achieve deep interaction, such as Figure 2 shown.

[0074] First, image features (256×H×W, where H represents the height of the feature map and W represents the width) are converted into a query matrix Q_img, a key matrix K_text, and a value matrix V_text via the learnable matrices W_q, W_k, and W_v. By calculating attention scores and taking a weighted sum, an image-to-text attention map Attn_img is generated, allowing the image features to focus on the key semantics of the text description. Second, a text semantic vector (256 dimensions) is converted into a query matrix Q_text, a key matrix K_img, and a value matrix V_img via W_q', W_k', and W_v'. This generates a text-to-image attention map Attn_text, allowing the text semantics to accurately locate lesion areas in the image. Finally, the original image features are combined with these two attention maps to generate the fused features F_fused, achieving precise alignment of text semantics with visual features. For example, when the text mentions "blue and white stripes," the feature activation of the corresponding image region is significantly enhanced, effectively addressing the semantic gap between single-modal features and providing high-quality fused features for subsequent accurate segmentation.

[0075] Step 3.3: The fusion features output by the CMA module enter the MS-ASPP module, and multi-scale context is extracted through multiple parallel branches.

[0076] The fusion features output by the CMA module enter the MS-ASPP module. The schematic diagram of the MS-ASPP module is as follows Figure 3As shown, different dilated convolution rates (such as setting the dilated rate to 6, 12, and 18) are used to construct a multi-branch parallel structure to capture multi-scale context information such as small lesion details and large area semantics; at the same time, a skip connection is added between the convolution branches with different dilated rates to efficiently fuse the edge features extracted by the small dilated rate and the semantic features extracted by the large dilated rate, so that the model can accurately identify subtle lesions and grasp the overall distribution of the lesions.

[0077] Specifically, the entire structure contains four parallel branches: among them, the 1x1 convolution is used to extract local texture details; the 3x3 dilated convolution is set to dilated rates of 6, 12, and 18, respectively, and the corresponding receptive fields are 13x13, 25x25, and 37x37, respectively, and the global average pooling obtains the global semantics. A skip connection is added between the convolution branches with dilated rates of 6 and 12, and 12 and 18, and the small-scale edge features and large-scale semantic features are fused through 1x1 convolution (such as splicing the fine edges extracted by the dilated rate of 6 and the morphological features extracted by the dilated rate of 12), which improves the recognition ability of complex boundaries and small lesions. The outputs of each branch are weighted by the channel attention block to suppress background noise channels and enhance lesion-related feature channels.

[0078] Step 3.4. The multi-scale fusion features output by the MS-ASPP module enter the BAM module for optimization to overcome the fuzzy defects of traditional models in lesion boundary delineation

[0079] The multi-scale fusion features output by the MS-ASPP enter the BAM module, and a schematic diagram of the BAM module is as shown in Figure 4 As shown, first, two layers of 3x3 dilated convolution (dilated rate 1, 2) are used to extract pixel-level boundary clues and capture details such as "jagged" and "fuzzy transition"; then, global average pooling and 1x1 convolution are applied to the boundary feature map to generate a boundary attention map A_edge (a higher value represents a higher possibility of being a lesion boundary) through a sigmoid function, and the features are weighted according to the boundary attention map to enhance the feature response of the "fuzzy transition zone" and other boundary details; finally, the output feature F_out = fusion feature x A_edge + boundary feature F_edge. Through attention weighting, the boundary region is enhanced while the semantic information is preserved, solving the "boundary fuzziness" problem of traditional models.

[0080] Step 3.5. Calculate the total loss

[0081] After obtaining the multi-scale and boundary-optimized features, the model enters the loss calculation and parameter optimization section, and the specific logic is as follows: considering the region, boundary, and classification accuracy, the total loss function is defined as

[0082]

[0083] wherein λ1=0.3, λ2=0.3, λ3=0.4, denotes the Dice loss, wherein P denotes the segmentation result predicted by the model, G denotes the true segmentation label, and ε is a very small positive number (for example, 1e -7 to prevent the denominator from being zero; denotes the cross-entropy loss, wherein y i denotes the true label of the i-th sample, and p i denotes the probability prediction value of the model that the i-th sample belongs to the positive class; denotes the boundary loss, wherein B denotes the pixel set of the boundary region, and y i denotes the true label of the i-th pixel, and p i denotes the probability prediction value of the model that the i-th pixel belongs to the target region.

[0084] Step 3.6, the features output by the BAM module are restored to the original image size through three layers of progressive upsampling, a probability map is generated through 1x1 convolution, and a segmentation result of the lesion with accurate boundaries and complete regions is output through threshold segmentation

[0085] In the progressive upsampling and segmentation output (decoding layer) stage, the model gradually restores and classifies the features optimized in multiple scales and boundaries, and finally generates a clinically usable segmentation result.

[0086] Firstly, a three-layer progressive upsampling strategy is adopted: the first layer fuses the feature map, high-level semantic features, and global pooling results through 4 times upsampling, realizing preliminary positioning of the overall region of the lesion. The second layer fuses the middle-level morphological features and the dilated convolution feature with an expansion rate of 12 through 8 times upsampling, refining the outline of the lesion. The third layer fuses the low-level texture features and the dilated convolution feature with an expansion rate of 6 through 16 times upsampling, completing the optimization of the details. During the upsampling process, the features of each layer are connected through a skip connection to realize cross-scale information fusion and avoid the loss of details caused by upsampling.

[0087] After the feature map is restored to the original input size, a pixel-level classification probability map is generated through 1x1 convolution, and then the probability map is converted into a binary lesion mask through threshold segmentation (the default threshold is 0.5), while a pixel-level classification confidence map is output, clearly distinguishing between melanoma lesions and normal skin regions, and finally outputting a segmentation result containing accurate boundaries and complete regions, providing a visual lesion range reference for clinical diagnosis.

[0088] Step 3.7, update the parameters of the MA-CLIP model

[0089] Step 4, evaluate the performance of the MA-CLIP model using the validation set and optimize the parameters

[0090] Step 5, inputting the melanoma clinical image to be segmented into the trained MA-CLIP model, and outputting a segmentation result

[0091] The effects of the present application can be further illustrated by the following experiments.

[0092] To verify the segmentation effect of the present application on melanoma clinical images, the AdamW optimizer (learning rate 1e-4, weight decay 0.01) was used for 100 iterations. On this basis, in order to further prove the superiority of the method of the present application, comparative experiments and ablation experiments were carried out, and the results are shown in Figures 5-8 , to evaluate the contribution of each module to the overall performance.

[0093] As can be seen from Figure 5 , Figure 6 , the present application performs outstandingly in the lesion segmentation task, with the core indicators F1 value reaching 0.837, sensitivity (SEN) being 0.870, and Jaccard similarity (IoU) being 0.728, all of which are significantly better than other methods. These results show that the present application not only effectively improves the accuracy of lesion recognition, but also significantly improves the overlap between the segmented region and the true lesion region. Although the present application is slightly lower than some methods (such as Att R2U-Net) in specificity (SPE) and accuracy (ACC), its overall performance is still outstanding. Through the multi-branch parallel structure and the design of different dilated convolution rates, the present application can capture multi-scale context information and strengthen lesion-related features through channel attention mechanism, thereby showing significant advantages in complex boundary and small lesion recognition.

[0094] As can be seen from Figure 7 , Figure 8 , the present application has achieved significant performance improvement in medical image segmentation tasks by combining BLIP fine-tuning and multi-modal fusion strategy. Specifically, miou (mean intersection over union) reached 0.7208, and F1 value reached 0.8379, both of which are better than other input methods, fully demonstrating that the integration of text semantic information and optimization strategy can break through the limitations of single modality, accurately guide segmentation, and adapt to clinical needs, providing a practical path for the landing of multi-modal technology for medical image segmentation. In summary, the present application successfully realizes the effective combination of text semantic information and image information through BLIP fine-tuning and multi-modal fusion strategy, significantly improving the accuracy and robustness of medical image segmentation.

Claims

1. A melanoma lesion region segmentation method based on CLIP multimodal fusion network, characterized in that: The following steps are involved: Step 1: Construct the MA-CLIP model; The MA-CLIP model includes image and text feature extraction, cross-modal attention (CMA) module, multi-scale atrous spatial pyramid pooling (MS-ASPP) module and BAM module; Step 2: Fine-tune the BLIP language model using manually annotated text-image pairs, build a large-scale multimodal dataset, divide it into training, test, and validation sets, and perform preprocessing. Step 3: Train the MA-CLIP model; Step 3.1: Extract semantic features from clinical text and visual features from images through the CLIP model's Transformer text encoder and Vision Transformer backbone networks, and ultimately output hierarchical image features and text semantic guidance vectors. Step 3.2: The image features and text semantic vectors enter the CMA module for bidirectional interaction to generate fusion features that combine visual information and semantic representation. Step 3.3: The fused features output by the CMA module enter the MS-ASPP module to extract multi-scale context through multiple parallel branches; Step 3.4: The multi-scale fusion features output by the MS-ASPP module are optimized in the BAM module to overcome the fuzzy defects of the traditional model in depicting lesion boundaries. Step 3.5: Calculate the total loss In step 3.6, the features output by the BAM module are restored to the original image size through three layers of progressive upsampling. A probability map is generated through 1×1 convolution and threshold segmentation is performed to output the lesion segmentation results with precise boundaries and complete areas. Step 3.7, update the parameters of the MA-CLIP model; Step 4: Use the validation set to evaluate the performance of the MA-CLIP model and optimize the parameters; Step 5: Input the melanoma clinical image to be segmented into the trained MA-CLIP model and output the segmentation result.

2. The melanoma lesion region segmentation method based on CLIP multimodal fusion network according to claim 1, characterized in that: The specific contents of step 2 include: Step 2.1: Select melanoma clinical images and their pixel-level annotations from the ISIC2018 dataset, and write text descriptions containing location, color, and morphology to form an initial text-image pair dataset; Step 2.2: Fine-tune the BLIP pre-trained model using the initial text-image dataset to optimize image-text feature alignment. Automatically generate text descriptions for the annotated images based on the fine-tuned BLIP pre-trained model to construct a large-scale multimodal dataset. Step 2.3: Divide the data into training, validation, and test sets, and then perform data preprocessing. For image data, first resize it to 256×256 pixels to ensure consistent input dimensions, and then normalize the RGB channel values ​​to the [0,1] range. For text descriptions, use the CLIP text encoder to convert the natural language description into a 768-dimensional semantic vector, so that its feature dimension is adapted to the features extracted by the image branch, and finally output the standardized image feature vector and text semantic vector.

3. The melanoma lesion region segmentation method based on CLIP multimodal fusion network according to claim 2, characterized in that: The specific steps of step 3.1 include: For the text branch, the clinical text description is converted into an initial sequence vector through the CLIP Transformer text encoder. Then, semantic features are extracted through multi-layer Transformer blocks to capture the semantic information in the text and output a 768-dimensional text embedding vector. Finally, a global pooling operation is performed to generate a 256-dimensional semantic guidance vector. For the image branch, the input image is processed by CLIP's Vision Transformer (ViT-B / 16) backbone network. The image is first divided into 16×16 patches and encoded using the Patch Embedding method to generate low-level texture features. The mid-level lesion morphological features are then extracted through a multi-head self-attention layer. Finally, high-level semantic features are obtained through global pooling. The features of each layer are reduced to 256 channels through 1×1 convolution and aligned with the text feature dimension. Multi-scale information is retained through skip connections to avoid the loss of spatial details in high-level semantics.

4. The melanoma lesion region segmentation method based on CLIP multimodal fusion network according to claim 3, characterized in that: The specific contents of step 3.2 include: On the one hand, image features are converted into query matrix Q_img, key matrix K_text, and value matrix V_text through learnable matrices W_q, W_k, and W_v. The dimension of image features is 256×H×W, where H represents the height of the feature map and W represents the width of the feature map. By calculating the attention score and taking the weighted sum, the image-to-text attention map Attn_img is generated, allowing the image features to focus on the key semantics of the text description. On the other hand, the 256-dimensional text semantic vector is converted into the query matrix Q_text, key matrix K_img, and value matrix V_img through W_q', W_k', and W_v', generating the text-to-image attention map Attn_text, so that the text semantics can accurately locate the lesion area in the image; Finally, the original image features are added to the two types of attention maps to obtain the fused features F_fused, which achieves precise alignment of text semantics and visual features.

5. The melanoma lesion region segmentation method based on CLIP multimodal fusion network according to claim 4, characterized in that: The specific contents of step 3.3 include: constructing a multi-branch parallel structure, using 1×1 convolution to extract local texture details, setting the expansion rates of 3×3 dilated convolution to 6, 12, and 18, respectively, corresponding to receptive fields of 13×13, 25×25, and 37×37, respectively, and performing global average pooling to obtain global semantics; Adding skip connections between convolution branches with dilation rates of 6 and 12, and 12 and 18, and fusing small-scale edge features with large-scale semantic features through 1×1 convolution to improve the recognition of complex boundaries and small lesions; The output of each branch is weighted by the channel attention block to suppress the background noise channel and enhance the lesion-related feature channel.

6. The melanoma lesion region segmentation method based on CLIP multimodal fusion network according to claim 5, characterized in that: The specific contents of step 3.4 include: First, two layers of 3×3 dilated convolutions are used to extract pixel-level boundary clues, capturing details such as "jaggies" and "blurred transitions", where the dilation rates of the dilated convolutions are 1 and 2 respectively; Then, global average pooling and 1×1 convolution are applied to the boundary feature map, and the boundary attention map A_edge is generated through the sigmoid function. The features are weighted according to the boundary attention map to enhance the feature response of the boundary details of the "fuzzy transition zone"; The final output feature F_out = fusion feature × A_edge + boundary feature F_edge, which strengthens the boundary area through attention weight while retaining semantic information.

7. The melanoma lesion region segmentation method based on CLIP multimodal fusion network according to claim 6, characterized in that: The total loss of step 3.5 for: Among them, λ1=0.3, λ2=0.3, λ3=0.4, represents the Dice loss, Where P represents the segmentation result predicted by the model, G represents the true segmentation label, and ε is a very small positive number (e.g. 1e -7 , used to prevent the denominator from being zero; represents the cross entropy loss, where y i represents the true label of the i-th sample, p i Represents the model's predicted probability that the i-th sample belongs to the positive class; represents the boundary loss, Where B represents the pixel set in the boundary area, and y i The true label of the i-th pixel, p i Represents the model's predicted probability that the i-th pixel belongs to the target area.

8. The melanoma lesion region segmentation method based on CLIP multimodal fusion network according to claim 7, characterized in that: The specific contents of step 3.6 include: First, a three-layer progressive upsampling strategy is adopted: the first layer uses 4x upsampling to fuse the feature map with high-level semantic features and global pooling results to achieve preliminary localization of the overall lesion area; the second layer uses 8x upsampling to fuse mid-level morphological features with dilated convolution features with a dilation rate of 12 to refine the lesion outline; the third layer uses 16x upsampling to fuse low-level texture features with dilated convolution features with a dilation rate of 6 to optimize the details; during the upsampling process, features at each layer are fused across scales through skip connections to avoid detail loss caused by upsampling. After the feature map is restored to the original input size, a pixel-level classification probability map is generated through 1×1 convolution, and then the probability map is converted into a binary lesion mask through threshold segmentation. At the same time, a pixel-level classification confidence map is output to clearly distinguish melanoma lesions from normal skin areas. The final output is a segmentation result containing precise boundaries and complete areas.

Citation Information

Cited By

  • Photoacoustic multi-mode segmentation method and system based on boundary information and medium

    CN121010606A

  • Image segmentation method for identifying mineral boundary in table ore zone

    CN121415079A

  • Method and device for detecting red and swollen airway mucosa, medium and program product

    CN121482001A

  • Medical ultrasound image segmentation method based on geometric guidance and concept perception fusion

    CN122023813A

  • A method for intelligent identification and quantification of microscopic residual oil

    CN122414499A